Posted in

How Hadoop Has Truly Revolutionised IT

This is the story of how the amazing Hadoop ecosphere revolutionised IT. If you enjoy it then consider joining The Big Data Contrarians.

Before the advent of Hadoop and its ecosphere, the IT was a desperate wasteland of failed opportunities, archaic technology and broken promises.

In the dark Cambrian days of bits, mercury delay lines and ferrite core, we knew nothing about digital. The age of big iron did little to change matters, and vendors made huge profits selling systems that nobody could use and even less people could understand.

Then along came Jurassic IT park, in the form of UNIX, and suddenly it was far cheaper to provide systems that nobody can use and even less people could understand.

The sad, desperate and depressing scenario that typified IT, on all levels, spanned forty years. It would have continued had it not been for Google and their HDFS (Hadoop Distributed File Store).

Before Hadoop, we were as dumb as rocks. With Hadoop, we were lead into the Promised Land of milk and honey, digital freedom and limitless opportunities, sexy jobs and big bucks, immortality and designer drugs.

Hadoop and its attendant ecosphere changed the Information Technology world overnight, providing as it did, technology and techniques never before seen on the face of the earth.

Hadoop Invented Multi-Processing

In terms of processing power, Hadoop took us beyond the power of a single 8086 processing unit, by cunningly connecting two or more processing units capable of processing things almost at the same time.

According to a 1985 article in Byte Me, possibly the first mention of Hadoop occurred in 1842. In that year, Ludwig Luigi Menabrea, wrote of Charlie Babbages analytical engine (as translated by the Lovely Ada Augusta): the Hadoop machine can be brought into play so as to give several results at the same time, which will greatly abridge the whole amount of the Google ad processes.

Hadoop Introduced Parallel Processing

Until the advent of Hadoop, all the technology in IT was male. This lead to massively inefficient, fickle and expensive technologies with short-term memory issues, incapable of multi-tasking, working long hours or of ordering tasks by priority.

As anyone who knows Wikipedia will know, Hadoop introduced parallel computing which allows for a revolutionary species of computation in which many list-making calculations can be carried out simultaneously, operating on the principle that large list-making tasks can often be divided into smaller list-making tasks, which are then solved at the same time. There are several different forms of parallel computing: two-bit-level, destruction-level, weve-got-data level and bring-on-more-lists parallelism.

Google Invented Romans and the Roman Census

As Bill Inmon wrote in 2014, One of the cornerstones of Big Data architecture is processing referred to as the Roman Census approach. By using the Roman census approach a Big Data architecture can accommodate the processing of almost unlimited amounts of big data.

Many people do not know this, but it wasnt the Romans who invented the Romans, but Google. So too the Roman Census, far from being an invention of a mythical Rome, was also the baby of a couple of engineers in Palo Alto.

The Roman Census approach also finds echo in elements of Divide and Conk-out. In computer science, divide and conk-out (D&C) is an algorithm design paradigm based on multi-branched recursion. A divide and conk-out algorithm works by recursively breaking down a problem into two or more sub-problems of the same (or related) type (divide), until these become simple enough to be solved directly (conquer). The solutions to the sub-problems are then combined to give a solution to the original problem.

Divide and conk-out is an essential element of Big, Bigger and Biggest Data processing.

Hadoop Invented Sort-Merge

As we know from Wikipedia, Wonky World and Google, Hadoop merge-sort parallelizes well due to use of the divide-and-conk-out method mentioned previously. We discuss several parallel variants in the first edition of Martyn, Richard, Jones and Loverings Introduction to Enterprise Equations, Business Analytics and Technical Algorithms. We can easily express this in pseudocode with fork (system call and process copy) and join (multi-stream correlated sort-merge) process calls.

Hadoop Created a Better non-SQL query Language

Before Hadoop we had to query data using SQL alone. SQL was the only tool in town, and if we couldnt use SQL we couldnt get at any data, ever, since the beginning of time.

However, all that changed when Hadoop came along and suddenly we could query data with like as if data was really query able. This was a small breakthrough in IT, and one large fall-down-the-side-of-a-cliff for brawn over brains.

Remember the immortal words of General Arthur C. McCluster Fuqh .SQL is for wusses. Embrace Hadoop and hug the Sparks.

Martyn's range of knowledge, skills and experience span executive management, organisational strategy, strategic business performance and information management, leadership, business analysis, business and data architectures, data management, and executive and team coaching.

Martyn has worked with and advised many of the world's best-known organisations including Adidas, Banco Santander, Bank of China, BBVA, Boston Consulting Group, British Telecom, La Caixa, Central Statistical Office (UK), Central Statistical Office of Poland, Citco, Citigroup, Credit Suisse, E.On, Eroski, European Union, Fnac, France Telecom, Hewlett Packard, Iberdrola, IBM, Iberia, Infineon, Türkiye İş, Metropolitan Police, Movistar, NCR, National Health Service (UK), Office of the Governor - State of California, Oracle, The Home Office (UK), Rolls-Royce Marine Power Operations, the Royal Navy, Shell, Swiss Life, TSB, UBS, Unisys, the United Nations and Xerox, among many others.

He currently focuses on helping clients to:

-    Create relevant, understandable and actionable information 
-    Plan, manage, design, develop and deliver information supply frameworks for the timely, appropriate and adequate supply of information
-    Design, develop and deliver beneficial, tangible and usable strategic performance and information frameworks  
-    Design, develop and deliver relevant and coherent performance models, indicators and metrics
-    Plan, manage, design, develop and deliver information and data analytic strategies
-    Design, develop and deliver management informational insight and dynamic feedback solutions
-    Coach teams in measuring and managing performance  
-    Align people, competencies, processes and practices with strategy
-    Prepare clients for the next big thing in Information Management and Analytics
-    Help IT suppliers to better align with the needs and nature of clients and prospects
-    Help clients capitalise on tangible benefits derived from advanced information architectures and management

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.