Big data | Blogs de Big Data | Learn Big Data

Contents

Consider the following fact:

Facebook currently has more than one billion active users every month

Let's take a few seconds to think about what information Facebook usually stores about its users. Some of this is:

  • Basic demographic data (as an example, date of birth, sex, Current location, previous location, University)
  • Pact of activities and user updates (their photos, comments, I like it, applications you have used, games you have played, messages, chats, etc.)
  • Your social network (your friends, their circles, how are you related, etc.)
  • User interests (Read books, movies seen, places, etc.)

big-data

Using this and a lot of other information (as an example, what a user clicked on, what did you read and how long did you spend on it), Facebook does the following in real time:

  • Recommend people you know and mutual connections with them.
  • Use your current and past activities to understand what interests you
  • Reach out to you with updates, activities and announcements that might interest you more.

At the same time of these, there are close time activities (updated in batches and not in real time) as the number of people who talk about a page, people reached in a week.

Now, imagine the type of data infrastructure required to run Facebook, the size of your data center, the processing power required to meet user requirements. The magnitude can be exciting or terrifying, depending on how you look at it.

IBM's upcoming infographic highlights the magnitude of the requirements / data processing for some similar institutions:

big-data_ibm

This type of size and scale was not heard by any analyst until a few years ago and the data infrastructure in which some of these institutions had invested was not prepared to handle this scale.. This is often called a Big Data problem..

Then, What is big data?

Big data is data that is too big, complex and dynamic for any conventional data tool to capture, store, manage and analyze. Traditional tools were designed with scale in mind. As an example, when an organization would like to invest in a Business Intelligence solution, implementation partner would come, study business requirements and then create a solution to meet these requirements.

If the requirement for this organization increases over time or if you want to run a more granular analysis, had to reinvest in data infrastructure. The cost of resources involved in scaling up the resources that are regularly used to increase exponentially. At the same time, there would be a limitation on the size to which it could scale (as an example, machine size, CPU, RAM, etc.). These traditional systems could not support the scale required by some of the Internet companies..

How is big data different from traditional data?

Luckily or unfortunately, there is no size limit / parametric to choose whether the data is “big data” or not. Big data is typically characterized on the basis of what is popularly known as 3 Vs:

  • Volume – Nowadays, there are institutions that produce terabytes of data in a day. With increasing data, you will need to leave some data unanalyzed, if you want to use traditional tools. As the size of the data grows even more, will leave more and more data unanalyzed. This means leaving value on the table. It has all the information about what the client is doing and saying, But can't understand! – a sure sign that you are dealing with larger data than your system supports.
  • Variety – Although the volume is only the beginning, variety is what makes traditional tools very difficult. Traditional tools work best with structured data. Require data to have a particular structure and format to make sense. Despite this, the flood of data from emails, customer feedback, social media forums, the customer journey in the web portal and call centers are not structured by nature or, in the best case, they are semi-structured.
  • Speed – The rate at which the data is generated is as critical as the other two factors. The speed with which a company can analyze data would eventually become a competitive advantage for them.. It is its speed of analysis that enables Google to predict the location of flu patients in near real time. Therefore, if you can't analyze data at a speed faster than your input stream, you may need a big data solution.

Individually, each of these V's can still be fixed with the help of traditional solutions. As an example, if most of your data is structured, you can still get from 80% al 90% of commercial value through traditional tools. Despite this, if you face a challenge with the three Vs, you will know that it is about “big data”.

tres V

When do you need a big data solution?

Although the 3 V will tell you if you are dealing with "big data" or not, whether or not you need a big data solution depending on your needs. Below are scenarios in which big data solutions are inherently more suitable:

  • When it comes to semi-structured or unstructured big data from multiple sources
  • You need to analyze all your data and cannot work with sampling it.
  • The procedure is iterative in nature (as an example, searches in Google search engine, search for graphics on Facebook)

How does the big data answer work?

Although the limitations of traditional solutions are clear, How do big data solutions solve them? Big Data solutions operate on a fundamentally different architecture that is based on the following characteristics (illustrative below):

  • Data distribution and parallel processing: Big data solutions work on distributed storage and parallel processing. Briefly, files are divided into several small blocks and stored in different drives (calls backstage). After, processing occurs in parallel in these blocks and the results are merged again. The first part of the operation is typically called Distributed file system (DFS) while the second part is called Small map.
  • Tolerance for failure: By the nature of its design, big data response has built-in redundancy. As an example, Hadoop creates 3 copies of each data block in at least 2 racks. Therefore, even if a complete rack fails or is not enabled, the answer still works. Why is it incorporated? This feature enables big data solutions to scale to even inexpensive entry-level hardware instead of expensive SAN disks..
  • Scalability and flexibility: This is the genesis of the complete paradigm of big data solutions. Puede agregar o quitar racks fácilmente del cluster sin preocuparse por el tamaño para el que se diseñó esta solución.
  • Cost effectiveness: Due to the use of basic hardware, the cost of creating this infrastructure is much less than buying expensive servers with fault resistant disks (as an example, SAN)

HDFS illustrative_v2

In summary, What if this was all in the cloud?

Although developing a big data architecture is profitable, finding the right resources is difficult, which increases the cost of implementation.

Imagine a situation where a cloud service provider also takes care of all your IT concerns / infrastructure. You focus on conducting analysis and delivering results to the company instead of organizing the racks and worrying about the scope of their use.

All you have to do is pay according to your usage. Nowadays, end-to-end solutions are available in the market, where you can not only store your data in the cloud, but also consult and analyze them in the cloud. You can query terabytes of data in seconds and leave all the worrying about this infrastructure to someone else!!

Although I have provided an overview of big data solutions, this by no means covers the whole spectrum. The purpose is to start the journey and be prepared for the revolution that is underway..

If you like what you have just read and want to continue learning about analytics, may subscribe to our emails or like ours page the Facebook

Subscribe to our Newsletter

We will not send you SPAM mail. We hate it as much as you.

Datapeaker