Tuesday, 30 June 2015

Big Data-Hadoop and Its Impact on Business Intelligence Systems

Recently my work necessitated me look into the new features added in informatica 9.1, but I never thought the journey will take me to explore further on this and write a blog Let's see how I traversed through different new aspects that are getting very much related to data management and Business Intelligence. First we will look what is Bigdata and its position now.

People would always think how the organizations like Yahoo, Google, Facebook store large amounts of data of the users. We should take a note that Facebook stores more photos than Google's Picassa. Any guesses??

What is Hadoop
The answer is BigData Hadoop and it is a way to store large amounts of data in petabytes and zettabytes. This storage system is called as Hadoop Distributed File System. Hadoop was developed by Doug Cutting based on ideas suggested by Google's papers. Mostly we get large amounts of machine generated data. For example, the Large Hadron Collider to study the origins of universe produces 15 petabytes of data every year for each experiment carried out.

MapReduce
The next thing which comes to our mind is how quick we can access these large amounts of data. Hadoop uses MapReduce, which first appeared in research papers of Google. It follows 'Divide and Conquer'. The data is organized as key value pairs. It processes the entire data that is spread across countless number of systems in parallel chunks from a single node. Then it will sort and process the collected data.

With a standard PC server, Hadoop will connect to all the servers and distributes the data files across these nodes. It used all these nodes as one large file system to store and process the data, making it a 100% unadulterated distributed file system. Extra nodes can be added if data reaches the maximum installed capacity, making the setup highly scalable. It is very cheap as it is open source and doesn't require special processors like used in traditional servers. Hadoop is also one of the NoSQL implementations.

Hadoop in Real time
The Tennessee Valley Authority(TVA) uses smart-grid field devices to collect data on its power-transmission lines and facilities across the country. These sensors send in data at a rate of 30 times per second - at that rate, the TVA estimates it will have half a petabyte of data archived within a few years. TVA uses Hadoop to store and analyze data. In India Power Grid Corporation of India intends to install these smart devices in their grids for collecting data to reduce transmission losses. It is better they also emulate TVA. Recently Facebook moved to 30 Petabyte Hadoop, which sounds incredible and hard to digest the fact we are using such a myriad volume of data.

Data Warehouse and Business Intelligence [http://hexaware.com/business-intelligence-analytics.htm] Products supporting Hadoop and MapReduce
1) Greenplum
2) Informatica
3) Teradata
5) Pentaho
6) Talend
If Hadoop and other NoSQL implementations are widely used, the limitations of traditional SQL systems can be resolved like storing unstructured data. With the volume of data increasing exponentially, commercialization of Hadoop will happen in a large scale and data integrator tools will play a key role in mining data for business.

Readers share your experiences if any of you have worked with Hadoop on other ETL and BI Tools, tools that are available in the market.

Tuesday, 16 December 2014

Overview of Hadoop Applications

Hadoop is nothing but a source of software framework that is generally used in the processing immense and bulk data simultaneously across many servers. In the recent years, it has turned out to be one of most viable option for enterprises, which has the never-ending requirement to save and manage all the data. Web based businesses such as Facebook, Amazon, eBay, and Yahoo have used high-end Hadoop applications to manage their large data sets. It is believed that Hadoop Training is still relevant to both small organizations as well as big time businesses.
Hadoop is able to process a huge chunk of data in a lesser time which enabled the companies to analyze that this was not possible before within that stipulated time. Another important advantage of the Hadoop applications is the cost effectiveness, which cannot be availed in any other technologies. One can avoid the high cost involved in the software licenses and the fees that has to be upgraded periodically when using anything apart from Hadoop. It is highly recommended for businesses, which have to work with huge amount of data, to go for Hadoop applications as it helps in fixing any issues.

Actually, Hadoop applications are made up of two parts; one is the HDFS, which means the Hadoop Distributed File System while the other is the Hadoop map reduce that helps in the processing of data and scheduling of job depending upon the priority, which is a technique that initially originated in Google search engine. Along with these two primary components, there are nine other parts, which are decided as per the distribution one uses along with other complementary tools. There are three most common functions of Hadoop applications. The first function is the storage and analysis of all the data, which does not require the loading of the relational database management system. Secondly, it is used in the conversion of huge repository of semi-structured and unstructured data, for example a log file in the form of a structured data. Such complicated data are hard to understand in SQL tools like analyzing the graph and data mining.

Hadoop applications are mostly used in the web-related businesses wherein one has to work with big log files and data from the social network sites. When it comes to media or the advertising world, enterprises use Hadoop, which enables the best performance of ad offer analysis and help understand online reviews. Before using any Hadoop tool, it is advisable to read through the Hadoop map tutorials available online.

Friday, 28 November 2014

PMO using Big Data techniques on mygov.in to translate popular mood into government action

NEW DELHI: The Prime Minister's Office is using Big Data techniques to process ideas thrown up by citizens on its crowd sourcing platform mygov. in, place them in context of the popular mood as reflected in trends on social media, and generate actionable reports for ministries and departments to consider and implement.
The Modi government has roped in global consulting firm PwC to assist in the data mining exercise, and now wants to elevate Mygov.in platform from a one-way flow of citizens' ideas to a dialogue where the government keeps them abreast of some of the actions that emerge from their brainstorming.

"There is a large professional data analytics team working behind the scenes to process and filter key points emerging from debates on mygov.in, gauge popular mood about particular issues from social media sites like Twitter and Facebook," said a senior official aware of the development, adding that these are collated into special reports about possible action points that are shared with the PMO and line ministries. Ministries are being asked to revert with an action taken report on these ideas and policy suggestions currently being generated on 19 different policy challenges such as expenditure reforms, job creation, energy conservation, skill development and government initiatives such as Clean India, Digital India and Clean Ganga.

With the PM inviting Indian communities in America and Australia to join the online platform, which he has termed a 'mass movement towards Surajya', the traffic handling capacity of mygov.in is being scaled up consistently, the official said.

PwC executive director Neel Ratan said that the firm is 'helping the government' process the citizen inputs coming through on Mygov. in in.

"There is a science and art behind it. We have people constantly looking at all ideas coming up, filtering them and after a lot of analysis, correlating it to sentiments coming through on the rest of social media," he said, stressing this is throwing up interesting trends and action points, being relayed to ministries. "It's turning out to be fairly action-oriented. I think it is distinctly possible that 30-50 million people would be actively contributing to Mygov.in over the next year and a half, given its current pace of growth," Ratan said.

PwC's global leader in government and public services Jan Sturesson told ET the participative governance model being adopted through mygov.in could become a model for the developed world.

"The biggest issue for governments today is how to be relevant. If all citizens are treated with dignity and invited to collaborate, it can be easier for administrations to have a direct finger on the pulse of the nation rather than lose it in transmission through multiple layers of bureaucracy," he said, not ruling out the possibility of using the mygov.in for quick referendums on contemporary policy dilemmas in a couple of years.
"The problem in the West has been that the US, Australia and UK follow a public management philosophy that treats citizens as consumers. That's ridiculous, because a consumer pays the bill and complains, while a citizen engages differently and takes responsibility," said Sturesson.
Within the 19 broad citizen engagement themes on mygov.in, there are multiple discussion groups focused on specific subsectors and themes. When it was launched in July, the site enabled brainstorming among its registered users around seven policy challenges. Users are allowed to sign up for four discussion groups in areas of interest apart from a group dealing with issues in their immediate vicinity.

Article Source - http://articles.economictimes.indiatimes.com/2014-11-26/news/56490626_1_mygov-digital-india-modi-government

Saturday, 22 November 2014

Learning Hadoop with Linux World India



The best Big Data Hadoop Training in Jaipur is offered by LinuxWorld India. When it comes to technological training, we are the best in the business. The organization was established in 2005. Our training comprises of all the expert professionals and teachers who impart tremendous knowledge. When it comes to Cisco certifications, we are the first preference. We bring together a blend of network training and solutions with the help of our authorized training partner. We believe in providing the knowledge to its betterment. With the best study in Hadoop, you can build your super computer. We provide the best classroom and training of Version 2 Hadoop. We are the first ones to bring this facility to India.

The training fees required for this course is Rs. 25500 and the course module will be delivered to you by us. After completing this course with us you will be able a master in Hadoop and will produce a framework of MapReduce. You will also be an expert in writing complex programs related to MapReduce. Hadoop is a tool that was built on Java, and it focuses on improving the performance of hardware. By studying this, you can also create a cluster of data that will help you to program different models.

Thursday, 13 November 2014

Big Data Training in Jaipur



Intelligence is a leading an advanced Hadoop Training in Jaipur. They are especially known for the Data Warehousing and Hadoop in Jaipur. It includes the core competencies such as ABAP, DW, DI, HANA, SAP BO, AIX, Solaris, Red Hat Certification, Linux, PLSQL, Oracle SQL, Business Objects, Data stage, Informatics, Congos, Software Testing, ISTQB certification, Amazon Web Services, Mobile application of the developer using the IOS and Android. They also provide you with the best online training for all of the above technologies. Their main objective is to provide production based knowledge for every training module. They also have experienced trainers and well-equipped lab for every technology. 
The Hadoop administrators are one of the finest and famous among the world that are more in demand and are highly compensated in the technical role. Many of them are nowadays going for the CCAH qualification. They get to learn all the Hadoop topics such as first, to determine the infrastructure for your cluster and to correct the hardware. Second, learn the internals of the HDFS, Map Reduce and YARN. Third, learn the deployment to integrate with the data center and proper cluster configuration. Fourth, learn how to load data into the cluster from the RDBMS with the help of Scoop and from the dynamically generated files with the help of the Flume. 
Fifth, to configure the fair scheduler in order to provide service level agreements for the multiple users of the cluster. Sixth, solving the Hadoop issues, tuning, diagnosing and troubleshooting problems. Seventh, best practices for maintaining and preparing the production in the Apache Hadoop. This training helps you with good jobs, you get the complete technology specified, and they conduct campus drives. This course is best for the IT managers and the system administrators who have the basic knowledge about the Linux experience. Despite the Apache Hadoop, knowledge is not required.

To More about the Big Data Hadoop Training in Jaipur ,Please Click This