Showing posts with label Hadoop. Show all posts
Showing posts with label Hadoop. Show all posts

Thursday, 17 September 2015

Data Lake Showdown: Object Store or HDFS?

The explosion of data is causing people to rethink their long-term storage strategies. Most agree that distributed systems, one way or another, will be involved. But when it comes down to picking the distributed system–be it a file-based system like HDFS or an object-based file store such as Amazon S3–the agreement ends and the debate begins.
The Hadoop Distributed File System (HDFS) has emerged as a top contender for building a data lake. The scalability, reliability, and cost-effectiveness of Hadoop make it a good place to land data before you know exactly what value it holds. Combine that with the ecosystem growing around Hadoop and the rich tapestry of analytic tools that are available, and it’s not hard to see why many organizations are looking at Hadoop as a long-term answer for their big data storage and processing needs.
At the other end of the spectrum are today’s modern object storage systems, which can also scale out on commodity hardware and deliver storage costs measured in the cents-per-gigabyte range. Many large Web-scale companies, including Amazon, Google, and Facebook, use object stores to give them certain advantages when it comes to efficiently storing petabytes of unstructured data measuring in the trillions of objects.
But where do you use HDFS and where do you use object stores? In what situations will one approach be better than the other? We’ll try to break this down for you a little and show the benefits touted by both.
Why You Should Use Object-Based Storage
According to the folks at Storiant, a provider of object-based storage software, object stores are gaining ground among large companies in highly regulated industries that need greater assurances that no data will be lost.
“They’re looking at Hadoop to analyze the data, but they’re not looking at it as a way to store it long term,” says John Hogan, Storiant’s vice president of engineering and product management. “Hadoop is designed to pour through a large data set that you’ve spread out across a lot of compute. But it doesn’t have the reliability, compliance, and power attributes that make it appropriate to store it in the data lake for the long term.”
Object-based storage systems such as Storiant’s offer superior long-term data storage reliability compared to Hadoop for several reasons, Hogan says. For starters, they use a type of algorithm called erasure encoding that spreads the data out across any number of commodity disks. Object stores like Storiant’s also build spare drives into their architectures to handle unexpected drive failures, and rely on the erasure encoding to automatically rebuild the data volumes upon failure.
If you use Hadoop’s default setting, everything is stored three times, which delivers five 9s of reliability, which used to be the gold standard for enterprise computing. Hortonworks architect Arun Murthy, who helped develop Hadoop while at Yahoo, pointed out at the recent Hadoop Summit that if you only storing everything twice in HDFS, that it takes one 9 off the reliability, giving you four 9s. That certainly sounds good. Source

Thursday, 9 July 2015

Opportunities in Data Management With Hadoop

Every day, every minute, millions of pictures videos and other forms of data are being dumped on to the internet via websites like Facebook, you tube etc. Ever wondered where this data is being stored to be used effectively year after year? The growing number of data sources like social media are challenging the big data technologies. Being the latest sensation, media giants like Google, Facebook and Yahoo have decided to choose Hadoop for their data management predicaments.
Any enterprise wishing to leverage its data and analytics is advised to install Hadoop framework; open source software that allows processing of large data over clusters of computers.
History of Hadoop
Hadoop was created back in 2005 by computer scientists Doug Reed Cutting and Mike Cafarella. Hadoop was named by Doug after his son's stuffed toy elephant and is now being managed by Apache Software Foundation. In 2006 Dough joined Yahoo! which dedicated a team to develop Hadoop. By 2008, Hadoop was being used by other companies beside Yahoo! like Facebook, New York Times and Last.fm.
The Hadoop architecture is made up of the Hadoop Common, Hadoop distributed file system (HDFS) and a MapReduce engine. MapReduce and HDFS are designed to handle any node failures. The architecture distributes data into chunks across many servers for the programmers to easily analyze and visualize easily.
Demand for Hadoop
The market for Hadoop is projected to rise from a $1.5 billion in 2012 to an estimated $16.1 billion by 2020 as per report by Allied Market Research. The profits are predicted to be made by the Commercial Hadoop companies like Amazon Web Services, Cloudera, Hortonworks etc.
The reason for the success for this platform is its low cost implementation which helps companies to adopt this technology more conveniently. It is also adept at automatically handling node failures and data replications and does all the hard work.
It is clear that data management industry has expanded from software and web into retail, hospitals, government etc. This creates a huge demand for scalable and cost effective platforms of data storage like Hadoop. Hence it comes as a no surprise that a skill in Hadoop is most desired as of now. The future for data storage is endless, as it is highly unlikely that the companies will stop storing their data or find an alternative to do so anytime soon.
Training in Hadoop basics is sure to go long way and will pay off in the long run as companies are willing to offer competitive salaries for candidates with desired skill-sets. Banking on this demand will definitely prove beneficial.