
With an avalanche of 2. 5 quintillion bytes of data generated daily, could this be used...
Loading summary
BBC Announcer
Thank you for downloading from the BBC. For details of our complete range of podcasts and our terms of use, go to bbcworldservice.com podcasts
Ruth Alexander
hello, and welcome to More or Less on the BBC World Service. I'm Ruth Alexander. Our everyday lives generate 2.5 quintillion bytes of data every 24 hours, according to IBM. And 90% of the data that exists was created in just the last two years, it says. So could we be using this daily avalanche of statistics to make our lives better? Big data, as it's called, could even save lives, according to Kenneth Cukier, the data editor of the Economist magazine and co author of the book Big Data A Revolution that Will Transform How We Work, live and Think. So what exactly I began by asking him, is big data?
Kenneth Cukier
Well, there's no one single definition of big data, but a way to understand it is by its characteristics. And it basically means that we have vastly more data than we ever have before. And we have new techniques and tools by which we can learn from a large body of data, things that we never could when we only had smaller amounts to extract new forms of economic value. And in a way, it turns data into a new resource, a new economic input.
Ruth Alexander
So where is this data?
Kenneth Cukier
The data is everywhere. We're collecting more data about things that we always collected data on before. So every time we go to a website, everything we do, all of our interactions, even sometimes where our mouse scrolls over but we don't click, is being recorded. Now, there's another facet to it. We are taking things that we never considered as being informational and rendering it into a data form. So, for example, the way that you're sitting and other people are sitting, you don't think of that as data, but if I put sensors under your chair and measured it, I could probably create an index for which is almost like a fingerprint. The way that you're sitting is unique to you. And the service here could be an anti theft device in cars. Put it into a car seat and the car will know if the unauthorized driver is driving.
Interviewer
So you're predicting a revolution. Who's leading this revolution? Who's using this data and how are they transforming our lives?
Kenneth Cukier
Well, the first company to probably think about is Google. They are really the paragon of a big data mindset. The way that their algorithm works is that when you visit where you click, the algorithm records and learns from. So if many, many visitors are searching for the same term and they're clicking on the eighth link on a page, the algorithm knows to lift it up. The system is constantly learning because it's watching what people are doing and adjusting based on that.
Ruth Alexander
But this is just good for marketeers, isn't it?
Interviewer
Advertisers?
Kenneth Cukier
Absolutely not. It's going to transform all aspects of society and our lives. But just think of healthcare. Go into a clinic and the doctor has to use his wisdom and his judgment and experience to many times he is completely wrong. Instead, we should take an approach whereby we are looking into a database of correlations of all of our electronic medical records of what treatments work with what sort of people, where are adverse side effects for drugs, and depending on the machine to do things that we can't do and can do it better, and depending on our judgment and our creativity to do what we do very well and blend the two.
Interviewer
Has this happened yet?
Kenneth Cukier
Well, no. And one of the biggest reasons why is that the healthcare data is not either in electronic form, hadn't been because you never had the techniques to extract information from it before, or because there's actually rules protecting privacy that actively discourage the sharing of medical information. This is heinous. Future generations will be bewildered to wonder how medical service in the 21st century could have made decisions about human health without looking at the data of all of us and learning from correlations. Correlations just simply mean that we find associations or links between two things. So an example would be we might find out that everyone who takes one sort of drug and another sort of drug in combination becomes very, very drowsy. It might be a side effect that we didn't know about before. It might be too infrequent when it happens for an individual doctor to recognize. But through big data correlations, the machine would groom all of the data and identify it and then we'd know.
Interviewer
I suppose the key thing here though, is that we're talking about correlations, about links. People are often tempted to see correlations as causations though, aren't they? Do you see a danger that if we mine big data, people might be looking to link two things which are actually spurious? I mean, I can use a silly example and say that you could look at a massive database and see that people of a certain height are more likely to have the cancer star sign or the Scorpio star sign?
Kenneth Cukier
Well, that's right. What you're referring to is spurious correlations. And with big data, it's still a problem, but it was a problem prior to big data. You're essentially saying that we're looking for needles in haystacks and Instead of finding a needle, we're finding a strand of hay. Still a problem. But the point is that that doesn't apply when we're learning from data, learning from large data sets. So an example is machine translation. The way it works is through statistics. It's basically presuming that a word in one language is a suitable alternative for a word in another language. And the only way to make that system work very well is to have more data. The more data we have in it, the better we can score the probability that these words are the correct ones. Hence, in those instances, spurious correlation is not a big problem.
Interviewer
There's a lot of enthusiasm about big data. There's a lot of hyp.
Ruth Alexander
Is there a dark side?
Kenneth Cukier
Yes, there is. And it's one you might call propensity. The idea is that algorithms will be making decisions based on what we are likely to do, but penalizing us prior to us committing the crime or the infraction. It's a little bit like in the movie with Tom Cruise called Minority Report, in which people are being locked up for crimes they were predicted to do, called pre crime, not for what they've done. So what would this look like in practice? You could imagine that the state might find that someone has a 98% likelihood to shoplift in the next 12 months. And so instead of locking the person up, which I don't think is that plausible, I think we're a wiser society. You'll have a social worker knock on the door and say, I want to give this young gentleman an after school job or I want to give him some counseling. Now, it may look like a positive intervention by the state. However, it also could be seen as a negative one. He'll be stigmatized. And of course, the person also, using basic statistics, could make the claim that I will be part of the 2% that will exercise moral choice and not shoplift. I need to have the sanctity of my free will.
Interviewer
Do you think there's the potential that big data could bring about policies which are essentially racist, for example, because in that big data a correlation has been found.
Kenneth Cukier
The answer, of course, is yes, but I would emphasize that that already happens today in a world of small data. So an example is if you wanted to give loans to people and you used postal codes, it's not allowed in the United States to be the primary signal, because that encodes race, because certain races live in certain areas and others in other places. Likewise, this already is a matter of law in the European Union not to have happen the insurance Industry has looked at the statistics and found out that men get into more car accidents than women. European law has said that you must charge the same premium, even though they recognize that the statistics, if you will, are sexist. So I think that even in a big data world, we have the potential for this happening, but we also have to rely on our social values to prevent it from being enshrined in public policy.
Interviewer
Big data, much of it comes from the developed world. Is there a danger that some other countries, developing countries, fall behind in this?
Kenneth Cukier
Well, these gaps already exist. There's a sanitation gap.
Interviewer
Maybe they'll get bigger.
Kenneth Cukier
Yeah, maybe they'll get bigger. If you look at the way that computing has evolved, it hasn't. It's actually shrunk. You see very high end mobile phones in developing countries now. They're increasing their penetration of technology at a faster rate than we could ever predict because the cost is just becoming so accessible. So I don't think that big data is going to exacerbate these same divides.
Interviewer
Looking at the whole world does show how big data is to some extent limited, though, because some of the biggest problems of disease are in continents like Africa, for example, where there will be limited data to draw on.
Kenneth Cukier
Yes, but I'm not bothered by that because first the cost of collecting data is much lower and the tools are so much more accessible now that people in those very countries can start collecting the data and learning from it.
Ruth Alexander
Kenneth Cukier, thank you for listening to More or Less. If there's a number you'd like us to investigate or explain, please email us at more or lessbc.co.uk. and you can listen to more editions of the programme via our BBC.co.uk more or less.
Unknown/Comic Relief
You like potatoes and I like potato. You like tomatoes and I like tomato. Potato, potato. Data, let's call the whole thing off.
BBC Announcer
There are dozens of different podcasts now available from the BBC, including news, documentaries, science, business, arts and sports. For details of them all, go to bbcworldservice.com podcasts.
Episode Title: Can Big Data Save Lives?
Host: Ruth Alexander (BBC Radio 4)
Guest: Kenneth Cukier (Data Editor, The Economist; co-author of Big Data)
Date: March 25, 2013
This episode of More or Less explores the concept of "big data," its transformative potential, and its possible impact on society—including whether harnessing vast datasets can literally save lives. Host Ruth Alexander interviews Kenneth Cukier about the revolutionary possibilities, practical limitations, ethical concerns, and societal risks associated with big data.
The conversation is thoughtful, inquisitive, and analytical, with Ruth Alexander asking probing questions and Kenneth Cukier providing enthusiastic but balanced insights—flagging both opportunities and limitations.
For listeners, this episode offers a clear-eyed overview of big data: celebrating its promise but also rigorously examining ethical and practical challenges, especially around healthcare, discrimination, and social policy.