Amazon has an unrivalled bank of data on online consumer purchasing behaviour that it can mine from its 152 million customer accounts. Since many years, Amazon uses that data to built a recommender systems that suggest products to people who visit Amazon.com. Already in 2003 they used item-item similarity methods from collaborative filtering, which was at that time state-of-the-art. Since then Amazon has evolved and improved its recommender engine and today they master this to perfection. They use customer click-stream data and historical purchase data of all those 152 million customers and each user is shown customized results on customized web pages.
Amazon uses big data also to offer a superb service to its customers. This could be the effect of the purchase of Zapos in 2009, but it clearly helps that it ensures that customer representatives have all the information they need the moment a customer needs support. They can do this because they use all the data they have collected from their customers to build and constantly improve the relationship with its customers. This is something many e-tailers can learn from.
But Amazon is expanding its usage of Big Data since it notices that the competition is nearing closer. As such, Amazon added a remote computing services, via Amazon Web Services (AWS), to their already massive product and service offering. AWS was launched in 2002, but only recently they added Big Data services and they now offer tools to support data collection, data storage, data computation along with data collaboration and data sharing. All are available in the cloud. The Amazon Elastic MapReduce provides a managed, easy to use analytics platform built around the powerful Hadoop framework that is used by large companies, including Dropbox, Netflix and Yelp.
Are you looking for Big Data Jobs or Candidates? Please go to our WORK section
However, there is more. Amazon also uses Big Data to monitor, track and secure its 1.5 billion items in its retail store that are laying around it 200 fulfilment centres around the world. Amazon stores the product catalogue data in S3. This is a simple web service interface that can be used to store any amount of data, at any time, from anywhere on the web. It can write, read and delete objects up to 5 TB of data each. The catalogue stored in S3 receives more than 50 million updates a week and every 30 minutes all data received is crunched and reported back to the different warehouses and the website.
At AWS, Amazon also hosts public big data sets at no cost. All available big data sets can be used and seamlessly integrated in AWS cloud-based solutions. Everyone can now use this public data, such as the data from mapping the Human Genome Project.
Just recently, MIT reported on a new project developed by Amazon: they are now packaging information on what it knows about consumers and sell this to marketers who can use it to advertise products tailored to what people really want. In contrast to Google and Facebook, who might have more overall data about consumers, Amazon has a clear understanding of what people actually buy and therefore what they are looking for and what they need. This is much more valuable information and this could definitely grow Amazons advertising revenue in the coming years.
In the past few years, Amazon has definitely moved away from a pure e-commerce player to a giant online player who offers much more than just products. It focuses massively on big data and is changing from an online retailer into a big data company.