In this blog post, we’ll explore what big data is and how analyzing vast amounts of data produces insights that differ from traditional information processing methods.
What Is Big Data?
Have you ever heard of terms like “Big Data,” “Smart Data,” “Data Mining,” or “Machine Learning”? Although their approaches differ, they all share a common thread: they are academic and technological fields that utilize massive amounts of data to gain new insights and value that were previously impossible to obtain. While the terms themselves may feel somewhat unfamiliar, we are already living in the age of big data. Search algorithms that present search results in an order most relevant to the user, as well as AI-based translation services, spell-checkers, and recommendation systems, are prime examples of big data applications.
Big data generally refers to “data so vast in terms of volume, frequency of generation, and format that it is difficult to collect, store, retrieve, and analyze using traditional data processing methods.” Big data is typically characterized by the so-called “3Vs”: Volume, Variety, and Velocity. Its key characteristics are that the volume of data is enormous; it includes not only structured data such as text but also various unstructured data types like videos, photos, audio, and location information; and it is generated in real time and rapidly analyzed and utilized.
How Does Big Data Analysis Generate New Insights?
So, how can big data analysis yield new insights? In the past, technical limitations meant that only small amounts of data could be collected and analyzed. Therefore, to test a single hypothesis, researchers primarily relied on random sampling techniques that prioritized precision. However, with significant advancements in data storage technology and computing power, these limitations have been largely overcome.
Today’s big data analysis collects as much data as possible and analyzes overall patterns while allowing for some margin of error in individual data points. Much like an Impressionist painting—where individual brushstrokes may appear somewhat irregular when viewed up close, but when viewed from a step back, a magnificent picture is revealed—this approach sacrifices some micro-level accuracy in exchange for macro-level insights.
The core of the macro-level insights gained through big data lies in predictions based on correlation. Correlation refers to the statistical relationship between two sets of data; a high correlation means that when one data value changes, the other is highly likely to change as well. In other words, unlike the causality we are familiar with, correlation does not imply necessity; it merely indicates a certain degree of probability. However, the fact that we can predict the probability of one phenomenon occurring based on another is highly significant.
A prime example is Google’s former flu prediction service (Google Flu Trends). Noting that search terms can reflect users’ health conditions or interests, Google analyzed search history and actual flu outbreak data to identify search terms that showed a high correlation with the spread of the flu. Based on this, Google sought to predict which regions were most likely to see the flu spread. Although there is only a correlation—not a causal relationship—between search terms and flu outbreaks, this initiative demonstrated the new potential of big data analysis to understand social phenomena in near real time. However, it was later confirmed that there were limitations to the accuracy of these predictions, such as overestimating the actual scale of outbreaks, and the field is now evolving toward a combined approach that utilizes both big data analysis and traditional epidemiological investigation methods.
What new changes will big data bring?
In his book ‘The Third Wave’, futurist Alvin Toffler likened the Agricultural Revolution, the Industrial Revolution, and the Information Revolution to three massive waves, describing them as revolutionary turning points that fundamentally transformed society and culture. The Information Revolution ushered in the Information Age through the selective collection of information and analysis based on causality. However, today, an approach that utilizes as much data as possible is becoming more important than selective data collection, and analysis that discovers new patterns and predictions based on correlation is gaining greater significance than analysis focused solely on identifying causality.
The volume of data is expected to continue growing in the future. Accordingly, the importance of big data technology—which effectively analyzes vast amounts of data to derive meaningful insights—will only increase. Big data provides new insights across various fields and will continue to establish itself as a core technology that brings about ongoing changes to society, industry, and our daily lives.