Showing posts with label cluster. Show all posts
Showing posts with label cluster. Show all posts

Thursday, July 11, 2013

Apache Hadoop: Solution for Bigdata

Nowadays “Bigdata” is the most hitting word all over the business world, peoples are not just talking about bigdata but finding business out of it. What exactly bigdata is? Simplest definition of bigdata is nothing but a data comes with high velocity with different varieties and huge volumes. The purpose of publishing this paper to not just to talk about bigdata but how to integrate bigdata in our current solution, how to find more business insights around the bigdata and hidden bigdata dimensions around your business. 
Apache Hadoop is the open source framework provided by Apache foundation to deal with bigdata, the power of Apache Hadoop is to provide cost efficient and effective solution to businesses for focusing more on exactly what matters: extracting business values from bigdata. In this paper we will be addressing more about the technical details about Hadoop Ecosystem architecture and integration with real time application to process and analysis and to find out the various hidden dimensions of bigdata, which helps our business to grow up.

Apache Hadoop as a Team:
Consider a regular scenario; you have a project team, one project manager and ten resources under him. 
If a client comes to your project manager and asked him to sort out the ten files, each file of 100 pages record.  What will be best approach your project manager will follow? 
Exactly! what you are thinking is right, Project manager will distribute the ten files among ten resources and keep the only record track with him. This approach will reduce to work load about 1/10th, ultimately increases speed and efficiency. 





Hadoop Team Structure:
This is what hadoop is, data storage and processing team. Hadoop has data storage and processing components. Hadoop follows master-slave architecture 

Physical structure of Hadoop cluster is same as above project team we have a Manager called namenode and team members called datanodes and Data storage is the responsibility of  datanodes(slaves), controlled by name node at master level and data processing is the responsibility of task tracker(slave) and controller over task tracker is job tracker at master level.



You can see in the diagram and do map with the project team that you have already and see how interesting it is. Try to map everything with the real world things you can find many possible ways and solutions out of it.  

Thursday, March 7, 2013

BigData : The data growth is 100 times bigger than population.


There are approximately 490,000 babies born and Over 150,000 People Die every day worldwide.

1. Near about 175 Million People Log Into Facebook Every Day, more on 250 million photos uploaded per day, 2.7 billion likes and comments per day.but more than this,
2. Twitter has confirmed that there are over 250,000,000 tweets posted every single day on the network.
3. Over 800 million unique users visit YouTube each month, Over 4 billion hours of video are watched each month, 72 hours of video are uploaded to every minute.
4. In 2011, YouTube had more than 1 trillion views or around 140 views for every person on Earth

This is present, what about the future, data is increasing 100 times more than human growth speed.

More recently, multiple analysts have estimated that data will grow 800% over the next five years. Computer World states that unstructured information might account for more than 70%–80% of all data in organizations

Volume,Variety and Velocity are the three measure pillars of BigData to achieve this data becomes Unstructured Data.
Unstructured Data : The Data that either does not have a data model or does not have relational tables. Unstructured data is typically text-heavy, but may contain data such as pictures, videos, songs, and of course the tabular data.

However, unstructured content is largely created by humans: inconsistent, emotional, careless, opinionated, lazy, driven, over-worked, always unique, humans. Appreciating this difference in the origins of the data that we seek to analyze is the first step to producing actionable insight and business advantage.

Then only Hadoop is one of the best the solution for BigData.

Hadoop:
The Apache Hadoop is a open source framework that allows for the distributed processing of large data sets(BigData) across clusters of computers using simple programming models.

Hadoop changes the economics and the dynamics of large scale computing. Its impact can be boiled down to four salient characteristics.

Eighty percent of the world’s data is unstructured, and most businesses don’t even attempt to use this data to their advantage. Imagine if you could afford to keep all the data generated by your business? Imagine if you had a way to analyze that data?

For more details about hadoop : http://hadoop.apache.org/

Followers