
Big data has been enjoying a lot of hype, with promises it will help deliver from to...
Loading summary
Tim Harford
This is the short edition of More or Less, first broadcast on the BBC World Service.
BBC Announcer
Thank you for downloading from the BBC. For details of our complete range of podcasts and our terms of use, go to bbcworldservice.com podcasts
Narrator/Host
hello and welcome to More or Less on the BBC World Service. I'm Tim Harford.
David Lazer
All of it.
Narrator/Host
Add all that information together and it's called Big Data. It is immensely valuable to a lot of people for good and possibly for ill. Now back to health and the real gold rush in this, as in so many other areas, is all about big data more and more. One area where big data is about to make quite a big difference. A lot of people are buzzing with excitement about the promise of what they're calling big Data. Big data itself is a vague term. It sometimes refers to the vast data sets produced by scientific instruments such as radio telescopes or the Large Hadron Collider. But another meaning of big data, the one which interests us for the next few minutes, is the digital information we are constantly producing as a by product of searching online, tweeting, posting to Facebook, paying by credit card, or wandering around with a mobile phone, constantly revealing our location. It looks like computers processing huge data sets are going to give us all the answers that social scientists, marketers and spies could possibly want. But here on More or Less, we want to make the case for caution.
Tim Harford
There's this enormous mythology that somehow the larger the data, the closer it is to truth. And I think it's at that level of mythology that we need to be most careful and most critical.
Narrator/Host
This is Kate Crawford, an academic and a researcher at Microsoft. We'll hear more from her later. But first, a couple of cautionary tales. Five years ago, a team of researchers from Google announced an impressive discovery. They'd found a way to track the spread of influenza across the United States by analyzing what we search for on the Internet.
David Lazer
Google Flu Trends was a lot faster at detecting the spread of the flu than traditional surveillance systems that required monitoring at hospitals.
Narrator/Host
This is David Lazer, professor of political science and computer and information science at Northeastern University. Google Flu Trends could give you flu case figures within 24 hours. The official figures from the US authorities took up to a fortnight. Google Flu Trends was fast, cheap and effective. It was also theory free. Instead of developing some model of what people with flu might search for, the Google team just looked at historical correlations between flu and their top 50 million search terms. Then they let the algorithms do the work.
David Lazer
And this created a great deal of attention. There were headlines And I think it has been held up as one of the exemplars of the potential of big data.
Narrator/Host
But there was a problem.
David Lazer
It started going off kilter and systematically so, which was a bit odd.
Narrator/Host
In the season 2011-2012, Google Flu Trends overestimated the flu by 50%. By the following year, it was predicting two cases of flu for every one that actually materialised.
David Lazer
If you say that there are more than twice as many cases as there really are, that's a big miss.
Narrator/Host
So what went wrong? Perhaps it was TV coverage about a flu epidemic that scared healthy people into searching online. Or perhaps Google Search itself got too clever for Google Flu Trends, automatically suggesting search terms and changing what people ended up looking for. No one's sure what happened, which is part of the problem. Without a theory for why people were searching for flu terms, Google could only spot patterns, and patterns weren't enough. No doubt Google Flu Trends will bounce back, but unless we learn the lessons of this episode, we will find ourselves repeating it. I've been looking into this as part of my day job at the Financial Times. And what worries me is that for all the genuine promise of these new data sets, we risk forgetting some very old statistical lessons. Google Flu Trends has already shown that finding patterns isn't enough. Knowing what causes those patterns matters too. And every time I hear people boasting about the size of their data sets, it reminds me of an old statistical story.
Pathe News Narrator
The battle is on. The Republican National Convention has nominated Governor Alfred Mossman Landon, the Kansas Coolidge, as its candidate for president.
Narrator/Host
As Pathe News reminds us, in 1936, the President of the United States, Franklin Delano Roosevelt, a Democrat, was seeking re election.
Pathe News Narrator
The Republicans think Alf Landon is the man who will win in November. He is our next president if he he can beat Roosevelt.
Narrator/Host
A very popular and respected magazine, the Literary Digest set itself the task of forecasting the result. This was a vast postal opinion poll. They sent ballots to a quarter of the total electorate, 10 million people. A quarter of those contacted. 2.4 million people sent responses. Eventually, the Literary Digest announced its prediction. Alfred Landon would win with a solid margin of 55% to 41%. But it was Roosevelt who crushed Landon by 61% to 37%. The Literary Digest was very wrong. Even worse, a fellow called George Gallup, who the opinion poll pioneer, conducted a much smaller survey and was far closer to the eventual result. So what did Gallup understand that the Literary Digest didn't? The Literary Digest went for size but neglected sampling bias. They got their vast mailing list from the phone book and the list of car registrations. But Americans who owned phones and had cars in 1936 weren't representative of the of the voting population. George Gallup, on the other hand, carefully selected a representative sample. The Literary Digest thought the bigger the sample, the better the result. But bigger isn't always better, as the authorities in the American city of Boston recently found out.
Tim Harford
One of the things you learn as a resident of Boston is that there's a lot of bad weather.
Narrator/Host
This is Kate Crawford again. You might remember that she's at Microsoft.
Tim Harford
There's actually a big problem with potholes in the road. They end up patching around 20,000 potholes a year, and they're always trying to think of more efficient ways to figure out where the potholes are. So I first heard about this new app that was being released by the city of Boston called Street Bump, that you could download to your smartphone. And what it would do is it would track your accelerometer data, which is the way that your phone is moving in space, along with your gps, which gives the coordinates of where you are, so that it could actually passively detect every time you would hit a pothole as you were driving, driving around the streets of Boston. And this was actually a very clever idea.
Narrator/Host
Well, great. Everyone who has the app is sending back data. But Kate asked herself, who's missing?
Tim Harford
People in lower income groups and older citizens. That's people over the age of 60 are less likely to have smartphones. Therefore, in the areas where we have those populations living, we're actually getting less data about their roads. And that might mean then that a city could say, well, we're not getting data about any kinds of potholes there. We don't need to send out the road repair crews.
Narrator/Host
In this particular case, the city of Boston was wise to the bias and took steps to correct for it. But for Kate Crawford, the story represents something bigger.
Tim Harford
What's so interesting about this story is that I think it's a kind of parable for big data, is that when we start looking to smartphones and apps, we always have to think about who is being left out, who is not in that data set, who is not being represented.
Narrator/Host
Statisticians are scrambling to develop new methods to seize the opportunity of big data. Now, such new methods. Methods are essential, but they'll work by building on the old statistical lessons, not by ignoring them. It's dangerous to assume the results of big data analysis are 100% accurate. It's important to understand why we have the results. We get rather than merely picking out patterns. And we need to ask ourselves, who's missing from this data? Big data has arrived, but big answers have not. Well, that's all we have time for today. Please keep your questions and your comments coming. We're at more or lessbc.co.uk and as always, there's further information and a downloadable edition of the show available@bbcworldservice.com more or less. We'll be back next week. Until then, goodbye.
BBC Announcer
There are dozens of different podcasts now available from the BBC, including news, documentaries, science, business, arts, sport. The details of them all go to bbcworldservice.com podcasts.
In this episode of More or Less on the BBC World Service, host Tim Harford explores the promises and pitfalls of "Big Data," specifically examining how large data sets are used—and sometimes misused—in fields ranging from public health to urban planning. Through real-world examples, the episode highlights the importance of statistical wisdom, the risks of sampling bias, and the pressing need to understand who is represented—and who is overlooked—when making data-driven decisions.
"There's this enormous mythology that somehow the larger the data, the closer it is to truth. And I think it's at that level of mythology that we need to be most careful and most critical." ([01:29])
"Google Flu Trends was a lot faster at detecting the spread of the flu than traditional surveillance systems..." ([02:07])
"...it has been held up as one of the exemplars of the potential of big data." ([02:55])
"If you say that there are more than twice as many cases as there really are, that's a big miss." ([03:26])
"Finding patterns isn't enough. Knowing what causes those patterns matters too." ([04:13])
"The Literary Digest thought the bigger the sample, the better the result. But bigger isn't always better..." ([06:23])
"Who's missing?" ([07:29])
Lower-income and older residents—less likely to own smartphones—were left out, causing underreporting of potholes in their neighborhoods ([07:37]).
"...when we start looking to smartphones and apps, we always have to think about who is being left out, who is not in that data set, who is not being represented." ([08:05])
"Big data has arrived, but big answers have not." ([08:22])
Tim Harford (on Big Data mythology, [01:29]):
"There's this enormous mythology that somehow the larger the data, the closer it is to truth. And I think it's at that level of mythology that we need to be most careful and most critical."
David Lazer (Google Flu Trends’ failure, [03:26]):
"If you say that there are more than twice as many cases as there really are, that's a big miss."
Historical Reflection (on Literary Digest, [06:23]):
"The Literary Digest thought the bigger the sample, the better the result. But bigger isn't always better..."
Kate Crawford (on representation in Big Data, [08:05]):
"When we start looking to smartphones and apps, we always have to think about who is being left out, who is not in that data set, who is not being represented."
Host Summary (on Big Data’s current state, [08:22]):
"Big data has arrived, but big answers have not."
This episode of More or Less deftly illustrates both the power and perils of Big Data, urging listeners to value critical thinking over technological hype. Whether examining failed disease tracking, historical polling errors, or city planning via smartphone apps, host Tim Harford and his guests repeatedly emphasize the need to understand what data means—and, crucially, to ask who is missing from it.