Posts

Showing posts with the label Big data

China and Big Data

In an earlier blog , we saw how the Chinese government was instrumental in kick-starting China’s foray into AI. But that started in 2014, so how has China made so much progress in such a short time? Given that most of the R&D around AI was done in the West – the US, UK, and Canada - how did China emerge in this #2 position so quickly?   It helped greatly that most of the AI algorithms are public knowledge. They are not trade secrets or protected by patents. Combine that with the fact that these AI’s get better the more data they have access to. Not surprisingly, China with its huge population, generates enormous data.   Further, unlike the West, China doesn’t care about privacy. No, this isn’t just because the government says so. Rather, most Chinese don’t mind sharing their data with Alibaba, TikTok, Baidu or WeChat either. They find the conveniences and features they get in return to be worth it.   Plus, China has learnt AI by doing, i.e., by its entrep...

Aid and Big Data

ROI. It stands for “Return On Investment”. To put it crudely, it’s the “How much do you get in return?” question. A must-ask question in business, but considered a mark of selfishness in other fields. Like aid, philanthropy or charity. That is why this comment by a reader of the Dish was so thought-provoking: “Aid itself is a brilliant idea, and one with far-reaching and lasting effects, but there are basically no metrics for ROI, and nobody is acting to direct it intelligently. This happens because of the phenomenon Easterly notes, wherein people truly donate to feel good, not to actually effect change (which requires much more work).” And when doing good becomes an end in itself (without caring about the results), people sometimes stop asking whether they are even focusing on the most important problem: “The fact of the matter is that diarrhea is still the number-one infectious disease killer in the developing world, with HIV/AIDS so far off in the distance as to be virt...

Models v/s Patterns

Image
In response to my blog on the three generations of the Internet , my dad had commented: “Are we all stuffing ourselves with data and information, with very little time and inclination left for sharpening our innate, marvelous tool that evolution has led us to - the ability for digging meaning out of abundant data?...Will we have humankind reduced its ability to mind's ability for keen insights, failing to appropriately sharpening our grand mind potential?” Such questions have been asked and debated for years (ironically) on the Internet! Is correlation good enough? For example, Big Data would allow algorithms to tell you where the planet would be without ever discovering Kepler’s laws of planetary motion; but is that the same as knowledge? On the other hand, can laws only be found for the inanimate universe? And is Big Data the (only) way to go when it comes to predicting humans? Is George Box’s statement (“All models are wrong, but some are useful.” ) so true for humans th...

When Literature and Big Data Combine

“Literature is the opposite of data,” wrote the novelist Stephen Marche. Such a statement made sense even a few decades back, but today? Let’s take a look. Today, Dana Mackenzie’s article says, “the scientific method is tiptoeing into the English department”. Huge amounts of literature have been digitized, and once digitized, surely somebody will start hurling algorithms to find…well, something. In 2011, Google’s N-gram server allowed you to search Google Books for frequency of words or word combinations in the books in its database. There are, of course, obvious limitations to the significance of such raw counts (other than perhaps trending when words caught on or died). Enter topic modeling: “A topic-modeling algorithm infers, for each word in a document, what topic that word refers to.” Does the word “black” mean color? Race? Something bad? The algorithm “produces “bags” of words that belong together”, and leaves it to the human reader to decide the meaning from the c...

When N = All

Statistics is all about analyzing data for patterns. In the past, the size and quality of the sample set was critical. Sometimes, even a problem. Enter Big Data. Or as Kenneth Neil Cukier and Viktor Mayer-Schoenberger wrote in their article, The Rise of Big Data wrote: “But if we collect all the data -- “n = all,” to use the terminology of statistics -- the problem disappears.” Sure, “n = all” is an exaggeration. But it is true that the size of data samples has gone through the roof over the last decade or so. And no, Big Data doesn’t just refer to the size of the data: “Big data is also characterized by the ability to render into data many aspects of the world that have never been quantified before; call it “datafication.”” A few examples would help understand “datafication” better: take location and friendship. They got datafied due to GPS and Facebook respectively! And if datafication is here, can algorithms be far behind? Google Translate is based on statisti...

Friends List, Big Data and Loans

In the West, they have credit ratings for individuals. There are systems that keep track of your repayment record (electricity bills, credit card bills, EMI’s paid, outstanding loans etc). Your track record is then referred to when you apply for that loan or an increase on your card limit. India too has started building similar systems (like CIBIL). But what about poorer countries with no such systems? Or immigrants with no credit record? The Economist reports  that lenders are beginning to look at social networks to refine the credit ratings of potential borrowers: -          Like your LinkedIn contacts could act as a cross-reference about your job (do you have many contacts from that job you claim to have? How good do your contacts think you are at your job (this acts as a hint of how long it might take you to land a new job should you get laid off); -          Or your Facebook data ca...