Site icon DataFLOQ

Small Data vs. Big Data : Back to the basics

Intel Data Center GPU codenamed Crescent Island architectural slide showcasing Xe3P AI optimized GPU IP, up to 480GB LPDDR5x memory capacity, and a 350W air-cooled PCIe form factor.

Intel’s Crescent Island GPU targets AI inference economics by prioritizing massive LPDDR5x memory capacity over costly HBM architectures within a practical 350W air-cooled design.

Small data is data in a volume and format that makes it accessible, informative and actionable.

The Small Data Group offers the following explanation:

Small data connects people with timely, meaningful insights (derived from Big Data and/or local sources), organized and packaged often visually to be accessible, understandable, and actionable for everyday tasks.

This definition applies to the data we have, as well as the end-user apps and analyst workbenches for turning Big Data sets into actionable small data. The key action words here are connect, organize, and package, and the value is rooted in making insights available to all (accessible), easy to apply (understandable), and focused on the task at hand (actionable).

 

The term small data contrasts with Big Data, which usually refers to a combination of structured and unstructured data that may be measured in petabytes or exabytes. Big Data is often said to be characterized by 3Vs: the volume of data, the variety of types of data and the velocity at which it is processed, all of which combine to make Big Data very difficult to manage. Small data, in contrast, consists of usable chunks.

The idea of Big Data is compelling: Want to uncover hidden patterns about customer behavior, predict the next election, or see where to focus ad spend? Theres an app for that. And to listen to the pundits, we should all be telling our kids to become data scientists, since every company will need to hire an army of them to survive the next wave of digital disruption.

Yet all the steam coming out of the Big Data hype machine seems to be obscuring our view of the big picture: in many cases Big Data is overkill. And most cases Big Data is useful only if we (those of us who aren’t data scientists) can do something with it in our everyday jobs, which is where small data enters the picture.

At its core, the idea of small data is that businesses can get actionable results without acquiring the kinds of systems commonly used in Big Data analytics.


A company might invest in a whole lot of server storage, and use sophisticated analytics machines and data mining applications to scour a network for lots of different bits of data, including dates and times of user actions, demographic information and much more. All of this might get funneled into a central data warehouse, where complex algorithms sort and process the data to display it in detailed reports. While these kinds of processes have benefited businesses in a lot of ways, many enterprises are finding that these measures require a lot of effort, and that in some cases; similar results can be achieved using much less robust data mining strategies.

Small data is one of the ways that businesses are now drawing back from a kind of obsession with the latest and newest technologies that support more sophisticated business processes. Those promoting small data contend that its important for businesses to use their resources efficiently and avoid overspending on certain types of technologies.

Why Small Data?

The Future of Small Data

Rufus Pollock, of the Open Knowledge Foundation, says the hype around Big Data is misplaced – small, linked data is where the real value lies.

The discussions around Big Data miss a much bigger and more important picture: the real opportunity is not Big Data, but small data. Not centralized “big iron”, but decentralized data wrangling. Not “one ring to rule them all” but “small pieces loosely joined”.

The real revolution is the mass democratization of the means of access, storage and processing of data, it not about large organizations running parallel software on tens of thousands of servers, but about more people than ever being able to collaborate effectively around a distributed ecosystem of information, an ecosystem of small data.

For many problems and questions, small data in itself is enough. The data on my household energy use, the times of local buses, government spending these are all small data. Everything processed in Excel is small data. And when we want to scale up the way to do that is through componentized small data: by creating and integrating small data “packages” not building Big Data monoliths, by partitioning problems in a way that works across people and organizations, not through creating massive centralized silos.

This next decade belongs to distributed models not centralized ones, to collaboration not control, and to small data not Big Data.

Exit mobile version