Full transcript
0:03[Music]
0:08have you ever played with building
0:09blocks data Engineers are like expert
0:11Builders but instead of blocks they work
0:14with information they put together
0:16structures that make data neat and
0:18organized like creating a super cool
0:20puzzle where everything fits
0:23perfectly a data engineer is someone who
0:26designs and builds system to gather and
0:28store and manage lots of
0:31information on that note hello everyone
0:33welcome to this session to the data
0:35engineer full course join us on the
0:38journey and discover how data Engineers
0:40shape the future on decision making and
0:43Innovation we'll Begin by exploring the
0:46data engineering and the steps to become
0:48data engineer next we'll provide the
0:50overview of Big Data followed by the
0:53guidance on becoming the data engineer
0:55and insights into the associated salary
0:59asure is essential for data engineer
1:01because it offers a range of specialized
1:03tools and service designed to manage
1:06analyze and scale data effectively so as
1:09we will cover Azure topics such as Azure
1:12data Factory Azure database services and
1:14Azure SQL database and aure data link
1:18advancing further will explore Advanced
1:20Data modeling using PBI and aure
1:22including assure data bricks as we
1:25progress to Advanced topics we'll
1:27explore Haro Essentials covering HD
1:30Yan map reduce spark Pig hi and EDP for
1:34its
1:35ecosystem following this will progress
1:38to understanding Kafka streams to
1:40conclude the course we'll discuss the
1:42essential her interview question and
1:44answers to advance your career in data
1:46engineering interviews before we begin
1:49please consider subscribing to our
1:50YouTube channel and hit the Bell icon to
1:53stay updated on the latest tech content
1:55from edura also visit the Eda website
1:58for the data engineer master master
2:00program the link to which is given in
2:02the description box
2:04[Music]
2:08below today I'll share you an engaging
2:11story about two friends Alex and Bob hey
2:15have I told you about the incredible
2:16data engineering journey of a
2:18multinational e-commerce company no I
2:21don't think so what happened Alex well
2:24let me share this story with you Bob
2:26this company had lots of valuable data
2:28but struggled with scattered systems and
2:31outdated databases oh that's a tough
2:34situation what did they do about it Alex
2:37they embarked on a data engineering
2:38initiative B it was quite a journey I
2:41can imagine so what changes did they
2:43make Alex they integrated their data
2:46into a centralized system and automated
2:48data cleaning and transformation it made
2:51a huge difference Bob that sounds so
2:54promising did it impact their operations
2:56Alex definitely Bob they gained faster
2:59access to ACC data and generated
3:01realtime analytical reports impressive
3:04did they address data governance as well
3:06Alex yes they implemented data
3:08governance practices for data quality
3:10privacy and compilance that's
3:13commendable it's a great example of how
3:15data engineering can transform a
3:17business what an inspiring story Alex
3:20yeah I thought you would find it
3:22fascinating Bob data engineering truly
3:24has the power to drive success and
3:25Innovation together yeah thanks for
3:28sharing the story Alex it reinforces the
3:30importance of data Engineering in
3:32today's data driven World hello everyone
3:35this is Saia from edua and in this video
3:37we will be diving into the fascinating
3:39world of data engineering so data
3:42engineering is all about harnessing the
3:44power of data to drive meaningful
3:46insights in today's rapidly evolving
3:48digital landscape organizations are
3:51grappling with massive amounts of data
3:53and that's where data and sharing comes
3:55into play now let's take a quick look at
3:57the agenda for this video we will start
4:00by the introduction of what data
4:01insuring is and why it is important for
4:03so many businesses outside then we will
4:06delve into the key components of data
4:07engineering such as data inje data
4:10integration data transformation and data
4:12storage next we will discuss some of the
4:15responsibilities of data engineers and
4:17the comparison between data engineer
4:19data analyst and data scientist after
4:22that we will delve into the installation
4:24process of popular tools and
4:26Technologies which is used in data
4:28engineering finally we will wrap up with
4:30a glimpse of what data pipelines are so
4:33let's understand what is data
4:35engineering First Data engineering
4:37involves the design development and
4:39maintenance of system and infrastructure
4:41to handle large volumes of data
4:43effectively it focuses on the extraction
4:46transformation loading and storage of
4:48data which ensure its quality and
4:49scalability of analysis now let's
4:52understand what are the key components
4:54of data engineering data engineering
4:56helps to collect data from the disparate
4:58sources and integrate into a unified
5:01format which allows organization to have
5:03a comprehensive view of their data it
5:06also involves processing and
5:08transforming data into a suitable format
5:11which ensures data quality consistency
5:13and applications it includes the
5:16development of data pipelines and
5:18processing system that enable the
5:19efficient processing and Analysis of
5:21data this includes batch processing for
5:24large scale data transformation and
5:26realtime processing for immediate
5:28insights and action
5:30data engineering plays a crucial role in
5:32enabling data driven World by providing
5:35clean integrated and accessible data
5:37data insuring empowers organizations to
5:40make an informed decision based on
5:42accurate insights and Analysis overall
5:45it is used to handle the complexity of
5:46data processing storage and integration
5:49which ensures that the organization can
5:51leverage the full potential of their
5:53data assets for strategic operations and
5:56Innovations now after the thorough
5:59understanding of what data engineering
6:00is and its key features also we will Del
6:03into the importance of data
6:05engineering so here are some of the key
6:08reasons why data engineering should be
6:10important or why data engineering is
6:12important so our first reason is it
6:15focuses on building scalable and
6:18efficient data processing systems by
6:20optimizing data pipelines leveraging
6:23distributed computing Technologies and
6:25employing performance tuning techniques
6:27data Engineers ensured did organization
6:30can handle large volumes of data and
6:33Achieve faster processing times it
6:35provides the foundation for advanced
6:38analytics and machine learning
6:40initiatives by structuring and preparing
6:42data in a switchable format data
6:44Engineers unable data scientists and
6:46analyst to extract valuable insights
6:49build predictive models and develop
6:51machine learning algorithms data
6:54insuring enables organization to process
6:56and analyze data as it arrives this is
6:59crucial in scenarios such as fraud
7:02detection recommendation systems iot
7:05applications and monitoring systems that
7:07require immediate insights and actions
7:10as we already know the importance of
7:12data engineering we will cover the real
7:14word applications based on this so our
7:17first application is e-commerce sites so
7:20data insuring is used to collect and
7:21process large volumes of customer data
7:24transaction data and product data this
7:27enables e-commerce companies to predict
7:29recommendations optimize pricing
7:31strategies and improve Inventory
7:33management our second example is social
7:36media sites like data insuring plays a
7:39crucial role in collecting processing
7:41and analyzing social media data which
7:43allows companies to monitor brand
7:45sentiment track user interactions and
7:48derive insights for targeted marketing
7:51campaigns so our third example is in
7:53finance and banking sector data insuring
7:56is used to handle financial data
7:58including trans action records customer
8:00information and Market data it enables
8:03fraud detection risk assessment
8:05algorithmic trading and personalized
8:08Financial Services also okay so now the
8:11fourth application is in the healthcare
8:13sector data insuring is employed to
8:15manage and analyze patient records
8:17Medical Imaging data and clinical child
8:20data it enables Healthcare Providers and
8:22researchers to gain insights improve
8:24patient outcomes and develop predictive
8:27models so this examples highlight the
8:29diverse application and the importance
8:31of data Engineering in various
8:33Industries demonstrating how it enables
8:35organization to Leverage The Power of
8:37data for operation and Innovation now
8:40the data insuring process involves
8:42several key steps including data
8:44injection transformation storage
8:47processing and integration let's discuss
8:50each step in more detail so our first
8:52step is data inje data inje refers to
8:56the process of collecting and importing
8:58data from various sources into a data
9:00system or data pipeline this can involve
9:03extracting data from databases files
9:06apis streaming platforms or the other
9:08resources the goal is to GA relevant
9:11data and make it available for the
9:12further processing and Analysis Second
9:15Step would be the data transformation
9:18once the data is ingested it often needs
9:20to be transformed into a suitable format
9:22for analysis or storage this data
9:25transformation involves cleaning
9:27validating and restructuring the data to
9:29ensure consistency and usability this
9:32step may include tasks such as data
9:34filtering aggregation normalization data
9:37type conversion or the application of
9:39business rules so our third step would
9:41be data storage so after the data is
9:44transformed it needs to be stored in a
9:46structured manner right so this
9:48typically involves using databases or
9:50data storage system that provide
9:52efficient storage and retrial
9:54capabilities popular choices for data
9:56storage include relational databases
9:58data data warehouses data laks or
10:01distributed file systems the selection
10:03depends on the specific requirements of
10:05the project such as data volume SS
10:07patterns and the analytical records so
10:10our next step would be data processing
10:13data processing involves performing
10:15computations and analysis on the stored
10:17data this STP can include tasks such as
10:20data aggregation data Improvement data
10:22summarization statistical calculation or
10:25machine learning algorithms data
10:27processing can be done through various
10:28to tools such as the SQL queries data
10:31processing engines or the custom scripts
10:34Next Step would be the data integration
10:36so in this step data integration
10:39involves combining data from multiple
10:41sources to create a unified and
10:43comprehensive view this step is crucial
10:45when dealing with heterogeneous data
10:47sources or when different style produced
10:49data this needs to be Consolidated data
10:52integration can be achieved through data
10:54consolidation or by using extract
10:56transform load process to combine and
10:59merge data from various sources next our
11:02last step would be the data governance
11:04so data governance is nothing but a
11:06framework in a set of process which
11:08ensures the effective management and
11:10utilization of data assets within an
11:14organization it involves establishing
11:16policies procedures and guidelines for
11:18data management data quality and data
11:21usage all the steps in data Engineering
11:24Process are typically iterative and may
11:26require continuous monitoring optimizing
11:29ation and maintenance to ensure data
11:31quality reliability and performance data
11:35Engineers play a crucial role in
11:37designing and implementing efficient and
11:39scalable data pipelines to support
11:41datadriven applications and
11:43analytics Now we move ahead with the key
11:46responsibilities of data Engineers which
11:48typically includes designing and
11:51developing scalable and efficient data
11:53pipelines which extract transform and
11:56load data from various sources to
11:59targeted systems next would be the
12:01managing and optimizing data storage
12:03infrastructure for performance
12:05scalability and reliability selecting
12:07and configuring appropriate data storage
12:10technology such as relational databases
12:12data warehouses or no SQL databases to
12:14meet data storage and retrial
12:16requirements data Engineers are also
12:19responsible for applying data
12:20transformation techniques to clean
12:22filter Aggregate and structure data for
12:25analysis or consumption by Downstream
12:27systems they're also responsible for
12:30implementing data processing task using
12:32programming languages or data processing
12:34Frameworks to manipulate and transform
12:37data efficiently now after analyzing the
12:40key responsibilities of data insuring we
12:42will proceed with our next topic and
12:44understand the key differences between
12:46these three roles that is data engineer
12:48data analyst and data scientist these
12:51three are the distinct role within the
12:53field of data science each with its own
12:55set of responsibilities and skill
12:57requirements so our our first profile is
12:59of data engineer so data engineer are
13:02responsible for Designing building and
13:04maintaining the infrastructure and
13:06system that enable data storage
13:08processing and retrial as we all know
13:11they focus on creating and managing the
13:12data pipelines and architecture
13:15necessary for efficient data collection
13:17transformation and storage data
13:19Engineers also work closely with
13:21software engineers and database
13:23administrator to ensure data is
13:25accessible reliable and scalable they
13:28typically work work with tools like
13:29Hardo spark SQL ETL Frameworks and
13:33cloud-based platforms for data
13:34processing and storage then our next
13:37profile would be the data analyst so
13:40data analysts are focused on analyzing
13:41and interpreting data to derive
13:43meaningful insights they work with
13:46structured and unstructured data to
13:47identify patterns strengths and
13:49correlations data analysts are
13:51proficient in statistical analysis data
13:54visualization tools and data quaring
13:56techniques as well they often use tools
13:58tools like SQL Excel table or powerbi to
14:02analyze data and create reports and
14:04dashboards after dat our next profile
14:07would be the data scientist data
14:09scientist possess a blend of skills from
14:11mathematics statistics programming and
14:13domain knowledge they leverage their
14:15expertise to develop and Implement
14:17complex algorithms and models to solve
14:20integrate data patterns or extract
14:22insights from large data sets data
14:25scientists employ techniques like
14:26machine learning predictive modeling and
14:28statistical analysis to build predictive
14:30models uncover patterns and make
14:32predictions also they also collaborate
14:34with stakeholders to Define business
14:36problems and design experiments to
14:38together their data so to summarize this
14:41data Engineers focus on the
14:43infrastructure and data pipelines data
14:46analysts work on analyzing and Reporting
14:48data while data scientists concentrate
14:51on Advanced modeling and extracting
14:54insights so to make all this happen data
14:57Engineers rely on power powerful tools
14:59and Technologies platforms like Apache
15:02spark Apache Kafka SQL and no SQL
15:05databases and cloud services such as AWS
15:08and gcp provides the building blocks for
15:10efficient data engineering this tools
15:12help process vast amounts of data
15:15facilate realtime data streaming and
15:17ensure secure and scalable data storage
15:20now let's understand what are the data
15:23pipelines are so here is the glimpse of
15:25what data pipelines are and why we use
15:27data Pipelines
15:29so data pipelines are the series of
15:31steps that extract transform and load
15:34data from source to destination itself
15:36okay so they enable the efficient and
15:38automated flow of data through different
15:41stages ensuring data quality consistency
15:43and avability
15:45let's understand each step one by one so
15:48the first step is extraction the
15:50extraction phase of ETL involves
15:53retrieving data from various sources
15:55such as databases files apis or
15:57streaming platform platforms the goal is
16:00to extract the relevant data needed for
16:02further processing and Analysis this
16:05process typically includes establishing
16:07connection to the data sources
16:09performing data queries or using data
16:11extraction tools to pull the required
16:14data into the data pipeline Next Step
16:16would be the transformation the
16:18transformation phase of ETL focuses on
16:20cleaning validating and reshaping the
16:23exctic data to ensure its quality and
16:26consistency this step involves applying
16:28various operations and rules to the data
16:30such as data filtering data type
16:32conversions data aggregation data
16:35enrichment or data
16:36normalization the transformation process
16:39aims to make the data suitable for
16:41analysis storage or integration into the
16:43destination system Next Step would be
16:45the loading the loading pH of ETL
16:48involves storing the transform data into
16:50the targeted system such as databases
16:52data warehouses or data Lakes then the
16:55data is loaded in a structured format
16:57that aligns with the schema or format of
17:00the destination system this phase may
17:03include tasks such as data mapping
17:05schema matching data passing or indexing
17:08to optimize data storage and retrieval
17:11after that our next step would be the or
17:13maybe we call this our last step would
17:15be the monitoring and handling
17:17monitoring and handling refer to the
17:19ongoing monitoring management and the
17:21maintenance of data pipelines on okay so
17:24now let's discuss about batch processing
17:27and realtime streaming pipeline lines so
17:29in a batch processing pipeline there is
17:31a delay between time data is collected
17:33and when it is processed this delay can
17:36range from minutes to hours or even days
17:38depending on the schedule intervals on
17:41the other hand realtime streaming
17:43pipeline aim to process data as it
17:45arrives which results in a minimal
17:48latency data is processed analyzed in
17:50near real time or within a very low
17:52delay the second point is batch
17:55processing pipelines are designed to
17:56handle large volumes of data efficiently
17:59they can process and analyze massive
18:01amounts of historical data in a batch
18:03mode on the other hand realtime
18:05streaming pipelines focus on processing
18:07data as it arrives making them more
18:10suitable for handling data streams with
18:12continuous High Velocity data updates
18:15now the third point is batch processing
18:17pipelines typically utilize batch
18:19processing Frameworks like Apache spark
18:22or Hadoop map redu this Frameworks
18:25process data in chunks or batches which
18:27allows for parall processing and
18:29optimize resource utilizations also
18:32while realtime streaming pipelines often
18:34use streaming Frameworks like Apache
18:36Kafka Apache Flink or Apache stom this
18:39Frameworks enable continuous processing
18:41of data streams supporting low latency
18:43operations and realtime
18:45analytics now the fourth point is batch
18:48processing pipelines are commonly used
18:50for tasks that involve historical
18:52analysis generating periodic reports or
18:55data preparation for machine learning
18:56models they are Well Suited for
18:58scenarios where processing time is not
19:00critical but analyzing large volumes of
19:02data is essential on the other hand
19:05realtime streaming pipelines are ideal
19:07for application that require realtime
19:09monitoring immediate response or instant
19:12insights based on live data use cases
19:14include like fraud detection realtime
19:16recommendation systems network
19:18monitoring or iot sensor data analysis
19:21so our last point is batch processing
19:24pipelines often require significant
19:25Computing resources during the
19:27processing phase as they process large
19:29volumes of data in a batch mode whereas
19:32realtime streaming pipelines also
19:34requires Computing resources but are
19:36more focused on low latency processing
19:38and continuous data streams requiring
19:41efficient resources allocation and
19:43management so it's important to note
19:45that there can be the overlap between
19:47batch processing and realtime streaming
19:49pipelines and hybrid architectures
19:51combining both approaches are common so
19:54that wraps up our Deep dive into the
19:56world of data engineering we have
19:57covered the essential aspects of data
19:59pipelines governance and security giving
20:02you a comprehensive understanding of how
20:04data is ingested transformed stored and
20:07[Music]
20:12processed now how to become a data
20:15engineer the road map to become a data
20:18engineer can go like from being
20:21proficient in programming languages to
20:23learning Automation and scripting then
20:26understanding your databases
20:28and mastering data processing techniques
20:31to studying cloud computing and
20:34internalizing
20:35infrastructure now we'll see how to get
20:38started with becoming a data engineer to
20:41get started with the learning process
20:43you can look into our Eda YouTube
20:45channel to start with even with no prior
20:48knowledge of data engineering one can
20:50just go through the videos and get an
20:53understanding of the whole
20:55subject we even have edureka blogs to to
20:58help you with detailed information on
21:00the topics and can give you a clearer
21:03picture apart from these we even have
21:06premium courses which can help you
21:10understand the topic at ease with the
21:13personal trainer and 24 hours access to
21:17Lifetime content you can learn here at
21:19your own pace with a live trainer who's
21:22extremely efficient and knowledgeable
21:24and experienc in the particular field
21:27these courses can even get you
21:28certificates which you can add in your
21:30CV for better opportunity in your job
21:33market so that's it for today I hope
21:36this video helped you and will help you
21:38decide how to become a data
21:42[Music]
21:46engineer now I feel Sor of it's the best
21:49time to tell the story about how data
21:51evolved and how big data came fine RMA
21:54so we'll move forward so sort of what
21:57can you notice here RMA I see how
22:00technology has evolved earlier we had
22:02landline phones but now we have
22:04smartphones we have Android we have IOS
22:06that are making our lives smarter as
22:08well as our phone smarter apart from
22:10that we were also using bulky desktops
22:12for processing MBS of data now if you
22:14can remember we were using floppies and
22:16you know how much data it can store
22:18right then came hard dis for storing TBS
22:20of data and now we can store data on
22:22cloud as well and similarly nowadays
22:25even self-driving cars have come up and
22:28I know you must be thinking why are we
22:29telling that now if you notice due to
22:32this enhancement of Technology we're
22:34generating a lot of data so let's take
22:36the example of your phones have you ever
22:39noticed how much data is generated due
22:41to your fancy smartphones your every
22:44action even one video that you send
22:46through WhatsApp or any other messenger
22:48app that generates data now this is just
22:50an example you have no idea how much
22:53data you're generating because of every
22:55action you do now the deal is this data
22:57is not in a format that our relational
22:59database can handle and apart from that
23:02even the volume of data has also
23:04increased exponentially now I was
23:06talking about self-driving cars so
23:08basically these cars have sensors that
23:10records every minute details like the
23:12size of the obstacle the distance from
23:14the obstacle and many more and then it
23:17decides how to react now you can imagine
23:19how much data is generated for each
23:21kilometer that you drive on that car I
23:24completely agree with you RMA so let's
23:26move forward and focus on ious other
23:28factors behind the evolution of data I
23:31think you guys must have heard about iot
23:33if you can recall in the previous slide
23:35we were discussing about self-driving
23:36cars it is nothing but an example of iot
23:39let me tell you what exactly it is iot
23:42connects your physical device with
23:44internet and makes the device smarter so
23:46nowadays if you have noticed we have
23:48Smart ACS TVs Etc so we'll take the
23:51example of smart air conditioners so
23:53this device actually monitors your body
23:55temperature and the outside temperature
23:57and accordingly decides what should be
23:59the temperature of the room now in order
24:01to do this it has to First accumulate
24:04data from where it can accumulate data
24:06from internet through sensors that are
24:08monitoring your body temperature and the
24:10surroundings so basically from various
24:13sources that you might not even know
24:14about it is actually fetching that data
24:17and accordingly it decides what should
24:19be the temperature of your room now we
24:20can actually see that because of iot we
24:22are generating huge amount of data now
24:25there's one startat also that is there
24:26in front of your screen so if you notice
24:28by 2020 we'll have 50 billion iot
24:31devices so I don't think so I need to
24:34explain much that how iot is generating
24:36huge amount of data so we'll move
24:38forward and focus on one more factor
24:39that is social media now when we talk
24:42about social media I think RMA can
24:44explain this better right RMA yeah s but
24:47I'm pretty sure that even you use it so
24:50let me tell you that social media is
24:52actually one of the most important
24:54factor in the evolution of big data so
24:57nowadays everyone is using Facebook
24:59Instagram YouTube and a lot of other
25:02social media websites so these social
25:04media sites have so much data for
25:07example it will have your personal
25:09details like your name age and apart
25:11from that even each picture that you
25:14like or react to also generates data and
25:16even the Facebook pages that you go
25:18around liking that is also generating
25:20data and nowadays you can see that most
25:23people are sharing videos on Facebook so
25:25that is also generating huge amount of
25:28data and the most challenging part here
25:30is that the data is not present in a
25:32structured Manner and at the same time
25:35it is huge in size isn't that right sort
25:38of can't agree more the point you made
25:40about the form of data is actually one
25:42of the biggest factor for the evolution
25:44of big data so due to all these reasons
25:46that we have discussed have not only
25:48increased the amount of data but it has
25:50also shown us that data is actually
25:52getting generated in various formats for
25:54example data is generated with videos
25:56that is actually unstructured same goes
25:58for images as well so there are numerous
26:00or you can say millions of ways in which
26:02data is getting generated nowadays
26:05absolutely and these are just few
26:07examples that we have given you there
26:09are many other driving factors for the
26:11evolution of data so these are few more
26:14examples because of which data is
26:17evolving and converting to Big Data
26:19we'll discuss about the retail part I'm
26:21pretty sure that all of you must have
26:22visited websites like Amazon flip cart
26:24Etc and rishma I know you visited a lot
26:27of times
26:28yeah I do and suppose RMA wants to buy
26:30shoes so she won't just directly go buy
26:32shoes she'll search for a lot of shoes
26:35so somewhere her search history will be
26:36stored and I know for sure that this
26:39won't be the first time that she's
26:40buying something so there will be her
26:42purchase history as well along with her
26:44personal details and there are numerous
26:46ways in which she might not even know
26:47that she's generating data and obviously
26:49Amazon was not present earlier so at
26:52that time there is no way that such huge
26:53amount of data was generated similarly
26:56the data has evolved due to other
26:57reasons as well like Banking and finance
26:59media and entertainment etc etc so now
27:02the deal is what exactly is Big Data how
27:05do we consider data as big data so let's
27:07move forward and understand what exactly
27:10it
27:11is okay now let us look at the proper
27:14definition of Big Data even though we've
27:16put forward our own definitions already
27:19so sort of why don't you take us through
27:21it yes RMA sure so big data is a term
27:24for collection of data sets so large and
27:26complex that it becomes difficult to
27:28process using onhand database system
27:31tools or traditional data processing
27:34applications okay so what I understand
27:37from this is that our traditional
27:39systems are a problem because they're
27:41too oldfashioned to process this data or
27:44something no RMA the real problem is
27:47there is too much data to process when
27:49the traditional systems were invented in
27:51the beginning we never anticipated that
27:53we would have to deal with such enormous
27:55amount of the data it's like like a
27:57disease infected on you you don't change
27:59your body orientation when you get
28:01infected with a disease right RMA you
28:03cure it with medicines couldn't agree
28:06more sort of now the question is how do
28:09we consider some data as Big Data how do
28:11we classify some data as Big Data how do
28:14we know which kind of data is going to
28:16be hard for us to process well Sor of we
28:19have the five vs to tell us
28:21that so let's take a closer look at what
28:24are those so starting with the first V
28:27it's the volume of data it's
28:29tremendously large so if you look at the
28:31stats here you can see the volume of
28:33data is rising exponentially so now
28:35we're dealing with just 4.4 zettabytes
28:38of data and by 2020 just in three years
28:41is expected that the data will rise up
28:44to 44 zettabytes which is like equal to
28:4644 trillion gigabytes so that's really
28:50really
28:51huge it is because all these humongous
28:54all this humongous data is coming from
28:56multiple sources and that is the second
28:59V which is nothing but variety we deal
29:01with so many different kinds of files at
29:03all once there are MP3 files videos Json
29:06CSV tsv and many more now these are all
29:09structured unstructured and
29:10semi-structured all together now let me
29:12explain you this with the diagram that
29:14is there on your screen so over here we
29:16have audio we have video files we have
29:18PNG files we have Json log files emails
29:22various formats of data now this data is
29:24classified into three forms one is
29:26structured format now in structure
29:27format you have a proper schema for your
29:29data so you know what all columns will
29:31be there and basically you know the
29:33schema about your data so it is
29:35structured it is in a structured format
29:37or you can say in a tabular format now
29:39when we talk about semi-structured files
29:41these are nothing but Json XML and CSV
29:43files where schema is not defined
29:45properly now when I go to unstructured
29:47format we have log files here audio
29:50files videos and images so these are all
29:52considered as unstructured files and Sor
29:56it is also because of of the speed of
29:58accumulation of all this variety of data
30:00alog together which brings us to our
30:02third V which is velocity so if you look
30:05here earlier we were using Mainframe
30:07systems huge computers but less data
30:10because there were less people working
30:11with computers at that time but as
30:14computers evolved and we came to the
30:15client server model the time came for
30:18the web applications and the internet
30:20boomed and as it grew among the masses
30:22the web applications got increased over
30:24the internet and everyone started using
30:26all this applications and not only from
30:29their computers and also from mobile
30:31devices so more users more appliances
30:34more apps and H hands a lot of data and
30:37when you talk about people generating
30:39data or Internet RMA the one kind of
30:42application that strikes first in my
30:43mind is social media so you tell me how
30:46much data you generate alone with your
30:47Instagram posts and stories uh it will
30:50be quite a boast if I only talk about
30:52myself here so let's talk including
30:55every social media user so if you see
30:57the stats in front of your screen you
31:00can see that for every 60 seconds there
31:03are 100,000 tweets actually more than
31:06100,000 tweets generated in Twitter
31:08every minute similarly there are 695,000
31:12status updates on Facebook when you talk
31:15about messaging there are 11 million
31:17messages generated every minute and
31:20similarly there are
31:25698,000 email and that equals to almost
31:301,820 terabytes of data and obviously
31:34the number of mobile users are also
31:36increasing every minute and there are
31:382117 plus new mobile users every 60
31:42seconds gez that's a lot of data I don't
31:45even want to go ahead and calculate the
31:46total it would actually scare me yeah
31:49that's a lot now the bigger problem is
31:52how to extract the useful data from here
31:55and that's when we come to our next week
31:57that is value so over here what happens
32:00first you need to mine the useful
32:02content from your data basically you
32:04need to make sure that you have only
32:05useful fields in your data set after
32:07that you perform certain analytics or
32:09you say you you perform certain analysis
32:11on that data that you have cleaned and
32:13you need to make sure that whatever
32:15analysis you have done it is of some
32:17value that is it will help you in your
32:19business to grow it can basically find
32:21out certain insights which were not
32:23possible earlier so you need to make
32:25sure that whatever big data that has
32:26been generated or whatever data that has
32:28been generated it makes sense it will
32:31actually help your business to grow and
32:32it has some value to it now getting the
32:34value of this data is one big challenge
32:37let me tell you why and that brings us
32:39to our next V which is veracity now this
32:42big data has a lot of
32:43inconsistencies obviously when you're
32:45dumping such huge amount of data some
32:48data packets are bound to lose in the
32:50process now what we need to do we need
32:52to fill up these missing data and then
32:54start mining again and then process it
32:56and then come up with a good Insight if
32:59possible so if you can notice there's a
33:01diagram in front of your screen so over
33:03here we have this field which is not
33:04defined similarly this field and if you
33:06can notice here when we talk about this
33:08minimum value you see the other minimum
33:10values and when you talk about this it
33:12is it is way more than the other fields
33:15present in this particular column
33:16similarly goes for this particular
33:18element as well okay so obviously
33:21processing data like this is one
33:23problematic thing and now I get it why
33:26Big Data is a problem statement well we
33:29have only five vs now but maybe later on
33:31we'll have more so there are good
33:33chances that big data might be even more
33:36big okay so there are a lot of problems
33:39in dealing with big data but there are
33:41always different ways to look at some
33:43things so let us get some positivity in
33:45the environment now and let us
33:48understand how can we use Big Data as an
33:50opportunity yes RMA and I would say the
33:53situation is similar to the proverb when
33:55life throws you lemons make
33:57lemonade yeah so let us go through the
34:00fields where we can use Big Data as a
34:02boon and there are certain unknown
34:05problems solved only because we started
34:07dealing with big data and the Boon that
34:09you're talking about RMA is big data
34:11analytics first thing with big data we
34:13figured out how to store our data cost
34:16effectively we were spending too much
34:18money on storage before until Big Data
34:20came into the picture we never thought
34:22of using commodity Hardware to store and
34:25manage a data which is both reliable and
34:27feasible as compared to the costly
34:29servers now let me give you a few
34:31examples in order to show you how
34:33important big data analytics is nowadays
34:35so when you go to a website like Amazon
34:37or YouTube or Pandora Netflix any other
34:40website so they'll actually provide you
34:42certain fields in which they'll
34:43recommend some products or some videos
34:45or some movies or some songs for you
34:47right so how do you think they do that
34:49so basically whatever data that you are
34:51generating on these kind of websites
34:53they make sure that they analyze it
34:54properly and let me tell you guys that
34:57data is not small it is actually big
34:59data now they analyze that big data and
35:02they make sure that whatever you like or
35:04whatever your preferences are
35:05accordingly they'll generate
35:06recommendations for you and when I go to
35:09YouTube I don't know if you guys have
35:10noticed it but I'm pretty sure you must
35:12have done that so when I go to YouTube
35:14YouTube knows what song or what video
35:16that I want to watch next similarly
35:18Netflix knows what kind of movies are
35:20like and when I go to Amazon it actually
35:23shows me what all products that I would
35:25prefer to buy right so so how do you
35:27think it happens it happens only because
35:28of big data analytics okay so there is
35:31one more example that just popped into
35:33my mind I'll share with you guys so
35:36there was this time when the Hurricane
35:38Sandy was about to hit on New Jersey in
35:41United States so what happened then the
35:44Walmart used big data analytics to
35:47profit from it now I'll tell you how
35:49they did it so what Walmart did is that
35:52they studied the purchase patterns of
35:55different customers when a Hur hurri is
35:57about to strike or any kind of natural
35:58Calamity is about to strike on a
36:00particular area and when they made an
36:02analysis of it so they found out that
36:05people tend to buy emergency stuff like
36:08flashlight life jackets and a little bit
36:11of other stuff and interestingly people
36:13also buy a lot of strawberry poptart
36:17strawberry poptarts are you serious yeah
36:20now I didn't do that analysis so I
36:22Walmart did that and apparently it is
36:25true so what they did is so they stuffed
36:28all their stores with a lot of
36:29strawberry Pop-Tarts and emergency stuff
36:32and obviously it was sold out and they
36:34earned a lot of money during that time
36:37but my question here RMA is people want
36:39to die eating strawberry poptarts like
36:42what was the idea behind strawberry
36:43poptarts I'm pretty unsure about it but
36:45yeah since you have given us a very
36:47interesting example and Walmart did that
36:49analysis we didn't do it so yeah so it
36:51is a very good example in order to
36:53understand how big data analytics can
36:55help your business to grow and find
36:57better insights from the data that you
36:59have yeah and also if you want to know
37:01why strawberry poptarts maybe later on
37:04we can start making an analysis by
37:05gathering some more data also yeah that
37:08can be possible okay so now let's move
37:11ahead and take a look at a case study by
37:14IBM how they have used big data
37:16analytics to profit their company so if
37:19you have noticed that earlier the data
37:21that was collected from The Meters that
37:23you have in your home that measures the
37:24electricity consumed it is actually
37:26sending data after one month but
37:29nowadays what IBM did they came up with
37:31this thing called smart meter and that
37:33smart meter used to collect data after
37:35every 15 minutes so whatever energy that
37:38you have consumed after every 15 minutes
37:40it will send that data and because of it
37:42big data was generated so we have some
37:44stats here which says that we have 96
37:47million reads per day for every million
37:50meters which is pretty huge this data
37:52the amount of data that is generated is
37:54pretty huge now IBM actually realize the
37:57data that they're generating it is very
37:59important for them to gain something
38:01from that data so for that what they
38:03need to for that what they need to do
38:04they need to make sure that they analyze
38:06this data so they realize that big data
38:08analytics can solve a lot of problems
38:11and they can get better business Insight
38:13through that so let us move forward and
38:14see what type of analysis they did on
38:16that data so before analyzing that data
38:19they came to know that energy
38:20utilization and billing was only
38:23increasing now after analyzing Big Data
38:25they came to know that during Peak load
38:27the users require more energy and during
38:30off peak times the users require less
38:32energy so what advantage they must have
38:34got from this analysis one thing that I
38:36can think of right now is they can tell
38:39the industries to use their Machinery
38:41only during the off peak times so that
38:43the load will be pretty much balanced
38:45and you can even say that time of use
38:47pricing encourages cost savy retail like
38:50industrial heavy machines to be used off
38:52peak time so yeah they can save money as
38:55well because off peak times pricing will
38:58be less than the peak time prices right
39:01so this is just one analysis now let us
39:03move forward and see the IBM Suite that
39:05they developed so over here what happens
39:09you first dump all your data that you
39:11get in this data warehouse after that it
39:13is very important to make sure that your
39:15user data is secure then what happens
39:18you need to clean that data as I've told
39:20you earlier as well there might be many
39:21feeds that you don't require so you need
39:23to make sure that you have only useful
39:24material or useful data in your data set
39:27and then you perform certain analysis
39:30and in order to use this Suite that IBM
39:32offered you efficiently you have to take
39:34care of a few things the first thing is
39:36that you have to be able to manage the
39:38smart meter data now there is a lot of
39:41data coming from all this million Smart
39:43Meters so you have to be able to manage
39:45that large volume of data and also be
39:48able to retain it because maybe later on
39:50you might need it for some kind of
39:52regulatory requirements or something and
39:55next thing you should keep in mind is is
39:56to monitor the distribution grid so that
39:59you can improve and optimize the overall
40:01grid reliability so that you can
40:03identify the abnormal conditions which
40:05are causing any kind of problem and then
40:07you also have to take care of optimizing
40:10the unit commitment so by optimizing the
40:12unit commitment the companies can
40:14satisfy their customers even more they
40:16can reduce the power outages that is the
40:19they can reduce the power outages so
40:21that their customers don't get angry
40:23more identify problems and then reduce
40:25it obviously
40:27and then you have also to optimize the
40:29energy trading so it means that you can
40:31advise your customers when they should
40:33use their appliances in order to
40:35maintain that balance in the power load
40:38and then you also have to forecast and
40:40schedule loads so companies must be able
40:43to predict when they can profitably sell
40:45the Excess power and when they need to
40:48hedge the supply and continuing from
40:51this now let's talk about how Encore
40:53have made use of the IBM solution so
40:56Encore is an electric delivery company
40:59and it is the largest electrical
41:01distribution and transmission company in
41:03Texas and it is one of the six largest
41:06in the United States they have more than
41:08three million customers and their
41:10service area covers almost
41:13117,000 square miles and they began the
41:17advanced speeder program in 2008 and
41:20they have deployed almost 3.25 million
41:23meters serving customers of North and
41:27Central Texas so when they were
41:29implementing it they kept three things
41:31in mind the first thing was that it
41:34should be instrumented so this solution
41:36utilizes the smart electricity meters so
41:39that they can accurately measure the
41:41electricity usage of a household in
41:44every 15 minutes because like we
41:46discussed that the smart meters were
41:47sending out data every 15 minutes and it
41:50provided data inputs that is essential
41:52for consumption insights next thing is
41:55that it should be interconnected so now
41:58the customers have access to the
41:59detailed information about the
42:01electricity they are consuming and it
42:03creates a very Enterprise wide view of
42:06all the meter assets and it helped them
42:08to improve the service delivery the next
42:11thing is to make your customers
42:13intelligent now since it is getting
42:15monitored already about how each of the
42:17household or each customer is consuming
42:19the power so now they're able to advise
42:22the customers about maybe to tell them
42:24to wash their clothes at night because
42:27they're using a lot of appliances during
42:28the daytime so maybe they could divide
42:30it up so that they can use some
42:32appliances at off peak hours so that
42:34they can even save more money and this
42:37is beneficial for both of them for both
42:39the customers and the company as well
42:42and they have gained a lot of benefits
42:44by using the IBM solution so what are
42:47the benefits they got is that it enabled
42:50oncore to identify and fix outages
42:53before the customers get inconvenience
42:55that means they were able to ident
42:56identify the problem before it even
42:57occurred and it also improved the
42:59emergency response on events of severe
43:02weather events and views of outages and
43:05it also provides the customers the data
43:07needed to become a active participant in
43:09the power consumption management and it
43:12enabled every individual household to
43:14reduce their electrical consumption by
43:16almost 5 to 10% and this is how oncore
43:20used the IBM solution and made huge
43:23benefits out of it just by using big
43:25data analytics that IBM performed but
43:27let me just interrupt right now so since
43:30RMA told us in the beginning as well
43:32that there are no free lunches in life
43:34right so this is an opportunity but
43:36there are many problems to encase this
43:38opportunity right so let us focus on
43:40those problems one by
43:42one so the first problem is storing
43:44colossal amount of data so let's discuss
43:47few starts that are there in front of
43:49your screen so data generated in the
43:51past 2 years is more than the previous
43:53history in total so guys what are we
43:55doing toop generating so much amount of
43:58data and it said that by 2020 total
44:01Digital Data will grow to 44 Zab bytes
44:04approximately and there's one more stat
44:06that amazes me is about 1.7 MB of new
44:11information will be created every second
44:13for every person by 2020 so storing this
44:16huge data in traditional system is not
44:19possible the reason is obvious the
44:21storage will be limited for one system
44:24for example you have a with a storage
44:26limit of 10 terab but your company is
44:29growing really fast and data is
44:31exponentially increasing now what you'll
44:33do now at one point you'll exhaust all
44:35the storage so investing in huge servers
44:38is definitely not a cost effective
44:41solution so RMA what do you think what
44:43can be the solution to this problem uh
44:45according to me a distributed file
44:48system will be a better way to store
44:50this huge data because with this we'll
44:52be uh saving a lot of money let me tell
44:54you how because due to this distributed
44:57system you can actually store your data
45:00in commodity Hardware instead of
45:02spending money on high-end servers don't
45:04you agree sort of completely now we know
45:07storing is a problem but let me tell you
45:09guys it is just one part of the problem
45:11let's see few more okay so since we saw
45:15that the data is not only huge but it is
45:18present in various formats as well like
45:21unstructured semi-structured and
45:23structured so you not only need to store
45:26this huge data but you also need to make
45:28sure that a system is present to store
45:31this varieties of data generated from
45:33various sources and now let's focus on
45:35the next problem now let's focus on the
45:38diagram so over here you can notice that
45:40the hard disk capacity is increasing but
45:43the disc transfer performance or speed
45:45is not increasing at that rate let me
45:47explain you this with an example if you
45:50have only 100 MVPs input output Channel
45:54and you are processing say 1 ter of data
45:56now how much time will it take maybe
45:59calculate it'll be somewhere around 2.91
46:03hours right so it'll be somewhere around
46:062.91 hours and I have taken an example
46:09of 1 terabytes what if you're processing
46:11some Zab bytes of data so you can
46:13imagine how much time will it take now
46:15what if you have four input output
46:17channels for the same amount of data
46:20then it'll take
46:22approximately 72 hours or converted to
46:25minutes so it be around 43 minutes
46:27approximately right and now imagine
46:30instead of 1 TB you have zabt of data
46:32for me more than storage accessing and
46:34processing speed for huge data is a
46:37bigger problem okay so RMA has a very
46:39good example to discuss yeah so since
46:41you were talking about accessing the
46:43data and you told us already about how
46:46Amazon at different websites and YouTube
46:48they make those recommendations so if
46:51there was no solution for it if it would
46:53take so much time to access the data the
46:55recommendation system won't work at all
46:58and they make a lot of money just for
47:00recommendation system because a lot of
47:02people go there and click over there and
47:04buy that product right so let's consider
47:06that that it is taking like hours or
47:08maybe years of time in order to process
47:10my that big amount of data so let's say
47:13that at one time I purchased an iPhone
47:165s from Amazon and after two years I'm
47:20again browsing onto Amazon and since it
47:23took so much time to access the data and
47:25I already switched over to a new iPhone
47:29and they are recommending me the old
47:31iPhone case for 5S so obviously that
47:33won't work I won't go there and click it
47:35because I've already changed my phone
47:37right so that will be a huge problem for
47:40Amazon the recommendation system won't
47:42work anymore and I know that RMA changes
47:45her phone every year so if she has
47:47bought a phone and people are
47:50recommending if she has bought a phone
47:51now and someone's recommending the case
47:54for that phone after 2 years years
47:56doesn't make sense to me at all yeah
47:58only it will work if I have both the two
48:01phones at the same time but yeah I don't
48:03want to waste money on purchasing new
48:05iPhone case for my old phone so
48:08basically it won't be fair if we don't
48:10discuss the solution to these problems
48:12RMA we can't leave our viewers with just
48:15the problems right it won't be fair what
48:17is the solution had doop Hadoop is a
48:20solution so let's introduce Hadoop now
48:24okay so now what is had
48:26so Hadoop is a framework that allows you
48:29to first store big data in a distributed
48:31environment so that you can process it
48:34parallel there are basically two parts
48:37one is hdfs that is Hadoop distributed
48:39file system for storage it allows you to
48:42store data of various formats across a
48:45cluster and the second part is map
48:47reduce now it is nothing but a
48:49processing unit of Hadoop it allows
48:51parallel processing of data that is
48:54stored across the hdfs now let us dig
48:57deep in hdfs and understand it better
49:01yeah so hdfs creates an abstraction of
49:04resources um let me simplify it for you
49:07so similar to virtualization you can see
49:09hdfs logically as a single unit for
49:11storing big data but actually you're
49:14storing your data across multiple
49:16systems or you can say in a distributed
49:18fashion so here you have a Master Slave
49:21architecture in which the name node is a
49:23master node and the data node are slaves
49:26and the name note contains the metadata
49:28about the data that is stored in the
49:30data nodes like which data block is
49:33stored in which data node where are the
49:36replications of the data block kept and
49:38etc etc so the actual data is stored in
49:42the data nodes and I also want to add
49:44that we actually replicate the data
49:46blocks that is present in the data nodes
49:49and by default the replication factor is
49:51three so it means that there are three
49:53copies of each file so s of can you tell
49:56us why do we need that replication sure
49:59RMA since we are using commodity
50:01Hardwares right and we know failure rate
50:03of these Hardwares are pretty high so if
50:06one of the data notes fail I won't have
50:08that data block and that's the reason we
50:10need to replicate the data block now
50:12this replication Factor depends on your
50:14requirements right now let us understand
50:17how actually Hadoop provided the
50:18solution to the big data problems that
50:20we have discussed so RMA can you
50:23remember what was the first problem yeah
50:25it was storing the big data so how hdfs
50:29solved it let's discuss it so hdfs
50:32provides a distributed way to store Big
50:34Data we've already told you that so your
50:37data is stored in blocks in data nodes
50:39and you then specify the size of each
50:42block so basically if you have a 512 MB
50:45of data and you have configured hdf as
50:48such that it will create 128 megabytes
50:51of data block so hdfs will so hdfs will
50:54divide the data in four blocks because
50:5652 divid by 128 is four and it will
51:00store it across different data nodes and
51:02it will also replicate the data blogs on
51:05the different data nodes so now we are
51:07using commodity hardware and storing is
51:09not a challenge so what are your
51:11thoughts on it sort of I will also add
51:13one thing RMA it also solves the scaling
51:16problem it focuses on horizontal scaling
51:18instead of vertical now you can always
51:20add some extra data nodes to your hdfs
51:22cluster ads and when required instead of
51:25scaling the resources of your data nodes
51:27so you're not actually increasing the
51:29resources of your data nodes you're just
51:30adding few more data nodes when you
51:32require let me summarize it for you so
51:35basically for storing one TB of data I
51:37don't need a one TB system I can instead
51:40do it on multiple 128 GB systems or even
51:43less now RMA what was the second
51:46challenge with big data so the next
51:49problem was storing variety of data and
51:52that problem was also addressed by hdfs
51:55so with htfs you can store all kinds of
51:58data whether it's structured
51:59semi-structured or unstructured it is
52:02because in htfs there is no pre- dumping
52:04scheme of validation so you can just
52:06dump all the kinds of data that you have
52:08in one place and it also follows a write
52:11once and read many model and due to this
52:14you can just write the data once and you
52:16can read it multiple times for finding
52:18out insights and if you can recall the
52:21third challenge was accessing the data
52:23faster and this is one of the major
52:26challenge with big data and in order to
52:29solve it we're moving processing to data
52:31and not data to processing so what it
52:34means sort of just go ahead and explain
52:37it yes RMA I will so over here let me
52:40explain you what do you mean by actually
52:42moving process to data so consider this
52:45as our master and these are our slaves
52:48so the data is stored in these slaves so
52:50what happens one way of processing this
52:52data is what I can do is I can send this
52:54data to my Master node and I can process
52:57it over here but what will happen if all
53:00of my slaves will send the data to my
53:02master node it'll cause Network
53:04congestion plus input output Channel
53:06congestion and at the same time my
53:08master node will take a lot of time in
53:10order to process this huge amount of
53:12data so what I can do I can send this
53:14process to data that means I can send
53:17the logic to all these slaves which
53:19actually contain the data and perform
53:22processing in the slaves itself so after
53:24that what will happen the small chunks
53:26of the result that will come out will be
53:28sent to our name node so in that way
53:30there won't be any network congestion or
53:32input output congestion and it will take
53:34comparatively very less time so this is
53:37what actually means sending process to
53:41[Music]
53:44data so who's a big data engineer now
53:47every data driven business needs to have
53:49a framework in place for the data
53:51science and data analytics Pipeline and
53:54a data engineer is the one who's
53:56responsible for building and maintaining
53:58this framework now these Engineers must
54:00ensure that there is an uninterrupted
54:02flow of data between servers and
54:04applications so in simple words a data
54:07engineer builds tests maintains data
54:10structures and architectures for data
54:12ingestion processing and deployment of
54:15large-scale data intensive applications
54:18now data Engineers work in tandem with
54:20data Architects data analysts and data
54:22scientists so they must all share comp
54:25these insights to other stakeholders in
54:28the company through data visualization
54:30and storytelling but what does a big
54:32data engineer do exactly now the most
54:35crucial part of a big data engineer is
54:37to design develop construct install test
54:40and maintain the complete data
54:42management and processing systems they
54:45are basically the ones who handle the
54:47complete endtoend infrastructure for
54:49data management and processing they
54:51build a pipeline for data collection and
54:54storage and funnel the data to data
54:56analysts and scientists so basically
54:59what they do is they create the
55:00framework to make data consumable for
55:03data scientists and analysts so they can
55:06use the data to derive insights from it
55:09know that the data Engineers are the
55:11Builders of data systems and not those
55:14who mine for insights so the data
55:16engineer Works more behind the scenes
55:19and must be comfortable with other
55:20members of the team producing Business
55:22Solutions from this data now all there
55:25responsibilities revolve around this
55:27they need to take care of a lot of
55:28things while performing these activities
55:31hence one of the most sought after
55:33skills in data engineering is the
55:35ability to design and build data
55:37warehouses this is where all the raw
55:39data is collected stored and retrieve
55:41from without data warehouses all the
55:44tasks that a data scientist does will
55:46become obsolete it is either going to
55:48get too expensive or very very large to
55:51scale now data Engineers should always
55:53keep in mind that the system which he or
55:56she builds needs to be scalable robust
55:59and fault tolerant so that the system
56:02can be scaled up without increasing the
56:04number of data sources and can handle a
56:06huge amount of heterogeneous data
56:09without any failure now imagine a
56:11situation wherein the source of data is
56:13doubled or tripled but the system cannot
56:15scale up will it not cost a lot more
56:18time and resources to build the same
56:20system again which is suitable for this
56:22kind of intake exactly this is why the
56:26Big Data Engineers have a role here next
56:29he or she is the one that handles the
56:32extract transform and load process which
56:35is basically the blueprint for how the
56:38collected raw data is processed and
56:40transformed into Data ready for analysis
56:43now you're going to acquire a lot of
56:45data from different sources how do you
56:48bring them together to one platform ETL
56:51is your answer apart from all this a
56:54data engine engineer should always aim
56:57at deriving insights by acquiring data
57:00from new sources some of the
57:02responsibilities of a data engineer also
57:04include improving data foundational
57:07procedures integrating new data
57:09management Technologies and the software
57:12into existing systems and building data
57:14collection pipelines and finally one of
57:17the major roles of a data engineer is to
57:20include performance tuning and make the
57:22whole system way more efficient which is
57:24pretty self-explanatory if you ask me
57:27now most of us have some idea about who
57:30a big data engineer is but there's still
57:33some confusion about their
57:36responsibilities now this ambiguity
57:38further increases when we gain more
57:41information about the role now let me
57:43help you debunk all your queries about
57:46it so let's talk about some big data
57:48engineer
57:49responsibilities first up we have data
57:52ingestion now this is associated with
57:54the task of getting data out of the
57:56source systems and ingesting it into a
57:58data Lake now a data engineer would need
58:01to know how to efficiently extract the
58:03data from a source including multiple
58:05approaches for both batch and real-time
58:07extraction as well as needing to know
58:10about the incremental data loading
58:12fitting within small Source windows and
58:14parallelization of data loading as well
58:17now another small subtask of data
58:19ingestion is data
58:21synchronization but because it's such a
58:23big issue in the Big Data world we are
58:26going to talk about it now since Hadoop
58:28and other big data platforms don't
58:30support incremental loading of data a
58:33data engineer would need to know how to
58:34deal with detecting changes in the data
58:37source merge and sync change data from
58:40sources into the Big Data environment
58:42next we have data transformation this is
58:45basically the T in the extract transform
58:48and load that we had discussed earlier
58:50it is basically focused on integration
58:52and transformation of data for a
58:54specific use case now a major skill set
58:56here is the knowledge of SQL as it turns
58:59out not much has changed in terms of the
59:01type of data Transformations that people
59:03are doing now compared to purely
59:05relational environments now imagine all
59:08this data that you've acquired from
59:10various sources what would you have to
59:12do to make them all palatable in the
59:14same platform you need to transform that
59:17data and this is what a data engineer
59:19does here and finally we have
59:22performance optimization which is one of
59:24the tougher areas because anyone can
59:26build a slow performing system the
59:28challenge is to build data pipelines
59:30that are both scalable and efficient so
59:33the ability and understanding of how to
59:35optimize the performance of an
59:37individual data Pipeline and the overall
59:39systems are a higher level of data
59:41engineering skill now for example Big
59:44Data platforms continue to be
59:45challenging with regard to query
59:47performance and have added complexity to
59:49a data engineer's job in order to
59:51optimize performance of queries and
59:53creation of reports BS the data engineer
59:56needs to know how to denormalize
59:58partition and index data models he also
1:00:01needs to understand tools and Concepts
1:00:04regarding in-memory models and olap
1:00:06cubes now let's quickly move ahead and
1:00:09look at the required skills to fulfill
1:00:11these
1:00:12responsibilities now we'll be going
1:00:14through these skills in a clockwise
1:00:16order so starting with big data
1:00:18Frameworks now with the rise of big data
1:00:20in the early 21st century a new
1:00:22framework was born and that is Hadoop
1:00:24all thanks to Doug cutting for
1:00:26introducing this framework it not only
1:00:28stores big data in a distributed manner
1:00:31but also processes the data parallell
1:00:33there are several tools in the Hadoop
1:00:35ecosystem which cater differently for
1:00:37different purposes and Professionals for
1:00:39a big data engineer mastering Big Data
1:00:42tools is a must some of the tools which
1:00:45you will need to Master first of all you
1:00:47have hdfs which is the storage part of
1:00:49Hadoop being the foundation of Hadoop
1:00:52knowledge of hdfs is a must to start
1:00:54working working with Hadoop framework
1:00:56next we have yarn which performs
1:00:58resource management by allocating
1:01:00resources to different applications and
1:01:02scheduling jobs Now map ruce is a
1:01:05parallel processing Paradigm which
1:01:08allows data to be processed parallely on
1:01:10top of the hdfs next we have pig and
1:01:13Hive now Hive is a data warehousing tool
1:01:16on top of hdfs which caters to
1:01:18professional from an SQL background to
1:01:21perform analytics on top of hdfs whereas
1:01:23apachi pig is a high level platform
1:01:26which is used for data transformation on
1:01:28top of Hado now Hive is generally used
1:01:31by data analyst for creating reports
1:01:33whereas pig is used by researchers for
1:01:35programming both are pretty easy to
1:01:37learn if you already familiar with SQL
1:01:40next we have Flume and scoop Flume is a
1:01:43tool which is used to import
1:01:45unstructured data to hdfs and scoop is
1:01:48used to Import and Export structured
1:01:50data from our dbms now next we have
1:01:53zookeeper which acts as a coordinator
1:01:56among the distributed Services running
1:01:58in a Hadoop environment it basically
1:02:00helps to configure management and
1:02:02synchronized services and finally we
1:02:05have Uzi which is basically a scheduler
1:02:07which binds multiple logical jobs
1:02:10together and helps in accomplishing a
1:02:12complete task next up we have realtime
1:02:15processing Frameworks now real-time
1:02:17processing with quick actions is the
1:02:20need of R either it is a credit card
1:02:22fraud detection system or a Rec
1:02:24recommendation system now imagine if you
1:02:26wanted a red dress today and Amazon
1:02:29decides to suggest it to you a month
1:02:31later now wouldn't that be completely
1:02:33useless for you in this case you need
1:02:36realtime processing it is very important
1:02:38for a data engineer to have knowledge of
1:02:41realtime processing Frameworks now
1:02:43Apachi spark is one of the distributed
1:02:46real-time processing Frameworks which is
1:02:48used in the industry rigorously it can
1:02:50be easily integrated with Hadoop
1:02:52leveraging hdfs as well next we have
1:02:55dbms now a database management system
1:02:58stores organizes and manages a large
1:03:01amount of information within a single
1:03:03software application now data Engineers
1:03:06need to understand the database
1:03:07management system to manage data
1:03:09efficiently and allow users to perform
1:03:11multiple tasks with ease this will help
1:03:14data engineers in improve data sharing
1:03:16data security data access and better
1:03:19data integration with minimized data
1:03:22inconsistencies these are the fun
1:03:24mentals that data Engineers should know
1:03:27prior to building a scalable robust and
1:03:29fall tolerance system next we have SQL
1:03:32based Technologies now there are various
1:03:34relational databases that are used in
1:03:36the industry such as Oracle DB Microsoft
1:03:39SQL Server Etc now data Engineers must
1:03:43have at least the knowledge of one such
1:03:45database now knowing SQL is also a must
1:03:49this structured query language as SQL is
1:03:52also known as used to structure
1:03:54manipulate and manage data stored in
1:03:57relational databases as data Engineers
1:03:59work closely with
1:04:01rdbms's they need to have a strong
1:04:03command on SQL now next we have no SQL
1:04:07Technologies as the requirements of
1:04:09organizations have grown Beyond
1:04:11structured data no SQL databases have
1:04:15been introduced into this environment it
1:04:17can store large volumes of structured
1:04:19semi-structured or unstructured data
1:04:21with quick iteration and agile structure
1:04:24as per application requirements some of
1:04:26the most prominently used databases are
1:04:29hbas Cassandra and mongodb now hbase is
1:04:33a column oriented nosql database on top
1:04:36of hdfs which is great for scalable and
1:04:39distributed Big Data stores it is also
1:04:41great for applications with optimized
1:04:43read and range based scan and it
1:04:46provides consistency and partitioning
1:04:48out of capap now Cassandra is a highly
1:04:51scalable database with incremental
1:04:53scalability and the best part about
1:04:55Cassandra is the minimal Administration
1:04:58and no single point of failure it's good
1:05:01for applications with fast and random
1:05:03read and writs it provides available and
1:05:06partitioning out of capap and finally we
1:05:09have mongodb which is basically a
1:05:12document oriented nosql database which
1:05:15is a schema free database it gives full
1:05:17index support for high performance and
1:05:20replication for fall tolerance it has a
1:05:23Master Slave sort of architecture and
1:05:25provides CP out of capap it is
1:05:28rigorously used by web applications and
1:05:31semi-structured data handling next we're
1:05:34going to discuss programming and
1:05:35scripting languages so various
1:05:37programming languages can serve for the
1:05:39same purpose so knowledge of one
1:05:41programming language is enough I'm
1:05:43saying this because the flavor of
1:05:44language may change but the logic
1:05:46Remains the Same if you're a beginner
1:05:48you can go ahead with python as it is an
1:05:51easy language to learn due to its syntax
1:05:53and good good Community Support whereas
1:05:55R has a steep learning curve which is
1:05:58developed by statisticians and it is
1:06:00mostly used by analysts and data
1:06:02scientists the next skill we're going to
1:06:05discuss is an important one it is ETL or
1:06:08data warehousing now data warehousing is
1:06:11very important when it comes to managing
1:06:13a huge amount of data coming in from
1:06:15heterogeneous sources where you need to
1:06:17apply extract transform and load now
1:06:20data warehousing is used for analytics
1:06:23and Reporting and is a very very crucial
1:06:25part of every business intelligence
1:06:27solution because this is the part which
1:06:30is going to take you most time now it is
1:06:32very important for a big data engineer
1:06:34to Master One data warehousing or ETL
1:06:36tool after mastering one it becomes
1:06:39pretty easy to learn new tools and as
1:06:41the fundamentals remain the same now
1:06:43Informatica click View and talent are
1:06:46very well-known tools used in the
1:06:48industry Informatica and talent Open
1:06:50studio are data integration tools with
1:06:53ETL architecture
1:06:54the major benefit of talent is its
1:06:57support from the Big Data Frameworks if
1:06:59you're new to data warehousing and ETL
1:07:01tools I would definitely recommend you
1:07:03start with talent because after learning
1:07:06this any data warehousing tools will
1:07:08become a piece of cake in finally we
1:07:10have our operating systems now intimate
1:07:13knowledge of Unix Linux and Solaris is
1:07:17very helpful as many mathematical tools
1:07:19are going to be based off of these
1:07:21systems due to their unique demands for
1:07:24root access to hardware and operating
1:07:26system functionality above and beyond
1:07:28that of Microsoft's Windows or Mac OS
1:07:31now some level of understanding of how
1:07:33to act upon this data is also very
1:07:35valuable for data Engineers for this
1:07:38reason some knowledge of statistical
1:07:39analysis and the basics of data modeling
1:07:42are also hugely valuable knowledge of
1:07:45machine learning in Cloud also will
1:07:46serve as a big plus while machine
1:07:48learning is technically something
1:07:50relegated to a data scientist knowledge
1:07:52in this area is helpful to construct
1:07:54Solutions usable by your cohorts now
1:07:57this knowledge has the added benefit of
1:07:59making you extremely marketable in this
1:08:02space as being able to put on both hats
1:08:05in which case makes you a really
1:08:06formidable
1:08:07[Music]
1:08:13tool here you can see the job
1:08:16distribution per salary range in India
1:08:19for a data engineer as we can see people
1:08:22who get paid more than5 inom are about
1:08:2533% 730 perom are 26% 8 lak 70,000 about
1:08:3220% Which are very high salary brackets
1:08:36apart from that the average salary for
1:08:38data engineer is almost 8 lakhs inim and
1:08:41for a senior data engineer is almost 16
1:08:44lakhs in anim if we look at the same
1:08:47numbers in the US there are 32% of
1:08:50professionals who make more than $90,000
1:08:53a a year and 27% professionals who make
1:08:57$105,000 a year the average salary in
1:09:00the US for a data engineer is way more
1:09:03than
1:09:04$90,000 and for a senior data engineer
1:09:07it is
1:09:09$124,000 per anom now as we have
1:09:12discussed the salary of a big data
1:09:14engineer let's look at a few factors in
1:09:17the form of skills and technology that
1:09:19they know on which their salary depends
1:09:23here we've carefully curated a table
1:09:25which lists out the skills and the
1:09:28average salary which can be encashed
1:09:30through them you can see services such
1:09:33as AWS data analysis data mining
1:09:36warehousing machine learning and even
1:09:39programming languages like Java and R
1:09:43apart from that you can see bi tools and
1:09:46statistical tools like Tablo database
1:09:49architecture ETL and structured query
1:09:52languages now another another influence
1:09:54on the salary is experience because
1:09:57experience is also a very important
1:10:00factor in deciding the Big Data engineer
1:10:03salary the distribution of salary is
1:10:06like so an entry-level data engineer
1:10:09makes about
1:10:11$85,000 a year people who have 5 to8
1:10:14years of work experience bag nearly
1:10:18$113,000 a year and people who are
1:10:20experienced I'm talking like 10 years of
1:10:22Industry experience experience get over
1:10:26$118,000 a year now this salary must be
1:10:30coming from somewhere presenting to you
1:10:33the companies that hire in this job role
1:10:36as you can see there are some very big
1:10:38names like Amazon Google Bosch Microsoft
1:10:42and IBM who hire Big Data Engineers now
1:10:46companies that hire Big Data
1:10:48professionals are companies that are
1:10:50invested in the future and the worldwide
1:10:53big data market revenues for software
1:10:55and services are projected to increase
1:10:58from $42 billion in 2018 to $13 billion
1:11:03in
1:11:042027 attaining a compound annual growth
1:11:07rate of about
1:11:0910.5% that sort of a growth needs some
1:11:12kind of work after going through
1:11:15multiple job descriptions we found that
1:11:17the Big Data engineer salary has many
1:11:21[Music]
1:11:22variables
1:11:27so let's go ahead and also understand
1:11:29the average salary of a data engineer so
1:11:31in the US the average salary of a data
1:11:34engineer is
1:11:36$133,000 but remember guys this is
1:11:39basically an average salary which
1:11:41basically means that there are people
1:11:42who are above this salary or can even be
1:11:44a person who is below this salary okay
1:11:47but later in this session I'm also going
1:11:49to talk about some of the job
1:11:51description which can tell you the kind
1:11:52of salary that you can earn once you
1:11:55start applying for a data engineer
1:11:56profile in India the salary is around
1:11:596.5 lakhs per anom or a 7 lakh perom and
1:12:03the same goes for an Indian jobs as well
1:12:05that the salary can be higher than this
1:12:07average package or it can be lower than
1:12:10the package as well now that you have
1:12:12seen the basic or an average scale of a
1:12:15aor data engineer let's go ahead and see
1:12:18the job description of an aor data
1:12:20engineer all right guys so thinking
1:12:22again about what a data engineer does
1:12:25and aor data engineer is responsible for
1:12:28Designing implementing and maintaining
1:12:30data management and data processing
1:12:32system on the Microsoft Azor Cloud
1:12:35platform now they work with large and
1:12:37complex data sets and are responsible
1:12:39for ensuring that data is stored
1:12:41processed and secure efficiently and
1:12:43effectively now when you basically
1:12:45Define a job role like that you have
1:12:47many jobs description which are floating
1:12:49in the market right now how can you
1:12:52identify which job description you have
1:12:55to apply to now let's go ahead and
1:12:57understand that so talking about the job
1:13:00description which basically exists in
1:13:02the market guys you will see a job
1:13:04description which is an entry-level job
1:13:06description and then on the next slide I
1:13:09will show a job description which is a
1:13:11mid or a senior level job description
1:13:13all right so heading back to the entry
1:13:15level so if you see the first job
1:13:17description which is an aor data
1:13:19engineer the salary is anywhere around 6
1:13:22lakh perom to4 14 lakhs perom and these
1:13:26are the skills that you require now as I
1:13:29mentioned before these are the expected
1:13:31skills required for an aor data engineer
1:13:35now what are the skills which are
1:13:37expected now you are expected to know no
1:13:39SQL or a cosmos database skill now they
1:13:42are expected to know data Lake data
1:13:45factory data warehouse and strong
1:13:47experience in building pipelines in Azor
1:13:49data or in aor data Lake then you should
1:13:52be able to analyze and understand
1:13:54complex data you should be able to
1:13:56understand business requirement and
1:13:58actively provides input from data
1:14:00perspective now at the same time if you
1:14:02look at these skill sets these skill
1:14:04sets are all the scales at which you
1:14:06will basically know after you study for
1:14:09or clear the aor data engineer
1:14:11certification so once you're done with
1:14:13the certification once you are done with
1:14:14the skill sets which is just there in
1:14:16the certification you can easily go
1:14:18ahead and apply for a job that lies in
1:14:20the salary range which is for an zor
1:14:22data engineer
1:14:23now talking about a mid senior level
1:14:25data engineer profile the list is quite
1:14:28long as you can see now you can see that
1:14:31over here apart from all these skill
1:14:32sets a lot of other things are also
1:14:34mentioned here as well for example you
1:14:36should have some four or five plus years
1:14:38of experience in implementing or
1:14:41designing solution using Azor Big Data
1:14:43Technologies then you should have an
1:14:45experience with an Hands-On in Azor data
1:14:47factory aor devops aor data Lake storage
1:14:50Etc now you should have a knowledge of
1:14:53big data pipeline then design and build
1:14:55modern data pipelines and maintain the
1:14:57data warehouse schematics layouts
1:14:59architecture and relational or
1:15:01non-relational database for data access
1:15:03and advanced analytics now you should
1:15:06also know Java jQuery SQL or Scala or
1:15:10any preferred programming language right
1:15:12so over here what we recommend to our
1:15:14Learners is you should go ahead and
1:15:15Learn Python because although they have
1:15:17mentioned only these programming
1:15:19language but companies are very much
1:15:21flexible if the target profile has any
1:15:23programming experience but a strong one
1:15:25in any of the programming languages
1:15:27right next thing that they expect you to
1:15:29know is Advanced skill using one or most
1:15:32common language for example like python
1:15:34batch Etc so this python will basically
1:15:38serve as a dual purpose that is all a
1:15:40scripting language and a programming
1:15:42language as well the next thing that
1:15:44they expect you to know is the ETL
1:15:46process using big data Technologies such
1:15:49as spark Kafka Hadoop and others now the
1:15:52next thing that they they expect you to
1:15:53know is the ETL process using big data
1:15:56technology such as spark Kafka doops and
1:16:00others and then you need to understand
1:16:01the data visualization experience using
1:16:04python python here is a plus guys so you
1:16:07need to understand this and then you
1:16:08actually need to learn tableu or a
1:16:10powerbi any one of the tools or
1:16:13technology is a plus point so you either
1:16:15have to have a skill on tblo or a
1:16:19powerbi so this again is an important
1:16:21skill to have and this again coincide
1:16:24with what a data analyst does right
1:16:27because he is also responsible for data
1:16:29visualization to some extent then you
1:16:31have a solid understanding and
1:16:33experience implementing cloud data
1:16:34platform in Microsoft aor devops so if
1:16:38you're a guy who wants to start off you
1:16:39can start off with the freshia profile
1:16:42and after having experience in the
1:16:43freshia profile and learn all the skill
1:16:45sets like big data zor these are the
1:16:48skill sets that if you gain you can
1:16:50actually apply to the senior or midlevel
1:16:52in particular particular okay so Guys
1:16:54these are the few job description which
1:16:56are related to the data engineer profile
1:16:59now that we are clear with why who and
1:17:01what are the career opportunities and
1:17:03salary of an aor data engineer we shall
1:17:05see the path towards becoming an aor
1:17:08data engineer so first of all you will
1:17:10have to talk about the different aor
1:17:12storage which are out there now you'll
1:17:14have to learn about these like the blob
1:17:16storage table storage file storage and
1:17:19the que storage now you will have to
1:17:21learn about the relational database
1:17:22option options as well which are there
1:17:24in the market such as the SQL database
1:17:27SQL DB Warehouse analysis Services now
1:17:30you will also have to learn about no SQL
1:17:33you will have to learn about big data
1:17:35services in Azor like data Lake
1:17:37analytics data Lake storage now you also
1:17:40have to learn about the data factories
1:17:42aor function stream analytics iot hubs
1:17:45even hubs Etc now apart from that you
1:17:48will also need to learn redist cache and
1:17:50aor search so these are the servic that
1:17:53are basically required for you to
1:17:55understand in order to clear the
1:17:56certification and go ahead and become an
1:17:59aor data engineer now these are also
1:18:01Conde with the job description that we
1:18:03had a look earlier right now all the
1:18:05services which are mentioned there is
1:18:07basically a part of what you have
1:18:08learned in order to correct the
1:18:10certification now apart from this we
1:18:12also recommend going through the open
1:18:14source Services of a doop such as spark
1:18:18hi and the Hadoop itself right now let
1:18:21us go step by step to reach our goal in
1:18:23becoming an aor data engineer so in
1:18:26order to become an aor data engineer you
1:18:28will also need to have a strong
1:18:29foundation in data engineering and cloud
1:18:32computing here are some steps you can
1:18:34take to develop the skills and knowledge
1:18:36needed for a career as an aor data
1:18:38engineer so the first one is learn the
1:18:41basics of data engineering now in order
1:18:44to become an Azor data engineer you
1:18:46should first develop a strong foundation
1:18:48in data engineering Concepts such as
1:18:50data modeling data pipelines data
1:18:52processing and data storage now you can
1:18:55learn these Concepts through online
1:18:57courses or books or by working on
1:18:59practical projects now the second is
1:19:02learning or programming language now as
1:19:05an Reser data engineer you will need to
1:19:07proficient in at least one programming
1:19:09languages as I've already mentioned
1:19:11earlier now python is a very popular
1:19:13choice for data engineering but you can
1:19:15also consider learning languages such as
1:19:17Java C or Scala now the next step is you
1:19:21need to learn an SQL now as a data
1:19:23engineer you will be working with large
1:19:25amount of data you've already known that
1:19:27since the name itself as an aor data
1:19:30engineer so you will need to be a
1:19:31proficient in SQL to extract and
1:19:34transform data now the next thing you
1:19:36need to learn is Azor data storage
1:19:38options now as you know Azor offers a
1:19:41range of data storage options such as
1:19:43Azor blob storage Azor data L Storage
1:19:46Azor Cosmos DB Azor SQL database and
1:19:49Azor signups analytics formally SQL
1:19:52database warehouse now you should
1:19:54familiarize yourself with the features
1:19:56and capabilities of these storage
1:19:58options now the next thing you need to
1:20:00learn is Azor data processing
1:20:02Technologies now Azor provides several
1:20:04Technologies for processing data such as
1:20:06Azor stream analytics Azor data bricks
1:20:09Azor data Factory and Azor HD insights
1:20:12you should learn how to use these
1:20:14Technologies to build data pipelines for
1:20:16ingestion transforming and processing
1:20:19data so the next thing you need to learn
1:20:21is aor data management and and security
1:20:24so you as an Azor data engineer should
1:20:26learn about the Azor tools for managing
1:20:28and securing data such as Azor data
1:20:31catalog aor data link security and aor
1:20:34private link so what is the next step
1:20:37the next step is get certified as you
1:20:40know getting certified is very crucial
1:20:42and it's very important in today's
1:20:45generation because all the organization
1:20:48as I mentioned earlier if I just have to
1:20:49go back in my slide I will show you the
1:20:51job description where they have actually
1:20:53mentioned that you need to actually pass
1:20:56the certification all right this is the
1:20:59certification you required that is a dp2
1:21:01200 and a DP 2011 but as for now you
1:21:04don't need dp2 200 or DP 2001 you only
1:21:08have to give one exam which I'll be
1:21:10further talking about it that is the
1:21:12dp23 all right guys so moving ahead
1:21:15again so you need to earn an aor data
1:21:18engineer associate certification by
1:21:20passing the dp23 exam
1:21:23right now what is the next thing you
1:21:25need to do is you need to gain practical
1:21:27experience now the best way to learn aor
1:21:29data engineering is by working on
1:21:31practical projects now we all know this
1:21:34right now you can find data engineering
1:21:36projects on online platform such as
1:21:38keigle or you can work on projects in
1:21:40your own organization or you can enroll
1:21:43with edura and they provide you a tons
1:21:45and tons of Life projects which are
1:21:47developed by the instructor who have
1:21:50already worked as a data engineer in
1:21:52their their organization now the next
1:21:54thing you need to learn is continue
1:21:56learning now as a data engineer you will
1:21:58need to keep your skills and knowledge
1:22:00up to date as Technologies and best
1:22:02practices evolve right now making sure
1:22:05to stay current by learning about new
1:22:06aor data engineering features and
1:22:08participating in professionals
1:22:10development activities all right guys so
1:22:13now let us understand the things that
1:22:14you need to know about this
1:22:15certification that I've just talked
1:22:17about that is your dp23 so guys if
1:22:20you're applying for an aor data engineer
1:22:22associate
1:22:23there used to be two exam that you have
1:22:25to give which I've just mentioned now on
1:22:27the screen in the previous slide right
1:22:29one is implementing an Azor data
1:22:31solution and the next exam is design an
1:22:34Azor data solution now after clearing
1:22:36both these exam it will give you the aor
1:22:38data engineering associate certification
1:22:41and these two exam basically have the
1:22:42code like I've mentioned dp2 200 and DP
1:22:462011 right but now this exam have been
1:22:49retired on 23rd February 2021 now you
1:22:53guys just have to give one exam to clear
1:22:55the data engineer certification and this
1:22:58exam is
1:22:59dp23 now what they have done is they
1:23:02have Club both these exam and they have
1:23:04now included its syllabus in just one
1:23:06examine they're asking questions from it
1:23:09right so earlier what you have to do is
1:23:11you have to pay for two exams then
1:23:13prepare for two exams and give them and
1:23:15then only you can get the certification
1:23:18but now just by passing one exam you can
1:23:20clear out the data engineer
1:23:22certification ation right now as part of
1:23:24the new exams now the skill sets have
1:23:26been updated as you can see on the
1:23:28screen so basically these are the
1:23:30distribution of topics and they'll be
1:23:32covered in this exam now first of all
1:23:34you will be asked most of the questions
1:23:36on design and Implement data storage now
1:23:39this is going to have 40 to 45% of
1:23:42weightage okay now I'm going to tell you
1:23:45what are the topics that comes under the
1:23:47design and Implement data storage so
1:23:49make sure you write it down okay guys so
1:23:51the first is design a data storage
1:23:53structure then the second is design a
1:23:55partitioning strategy then the next
1:23:57question is design the serving layer
1:24:00then you have the Implement physical
1:24:02data storage structures then Implement
1:24:04logical data structure and lastly
1:24:06implement the serving layer now after
1:24:09this you will have design and develop
1:24:12data processing now this is of 25 to 30%
1:24:15weightage now there are four points
1:24:17coming under this design and develop
1:24:19data processing now the first is you
1:24:20need to learn about the inest and
1:24:21transform data second is design and
1:24:24develop a batch processing solution the
1:24:27third is design and develop a stream
1:24:29processing solution and lastly we have
1:24:31the manage batches and pipelines all
1:24:34right so coming to the third is design
1:24:37and Implement data security which is of
1:24:4010 to 15% so there are two main topics
1:24:43here that is design security for data
1:24:45policies and standards and the second is
1:24:47Implement data security Now talking
1:24:50about the last is you have the Monitor
1:24:52and optimize data storage and data
1:24:54processing which is again of 10 to 15%
1:24:57now even here also we have two main
1:25:00topics that you need to be covered that
1:25:01is the monitor data storage and data
1:25:03processing and the second is optimize
1:25:05and troubleshoot data storage and data
1:25:08processing now if you are with me till
1:25:10at this point you shall have understood
1:25:12what are all the things that you have to
1:25:13learn in order to clear the exams and
1:25:16become an aor data engineer right since
1:25:19we have already discussed what are the
1:25:20path and what are the skills to learn
1:25:22learn to prepare ourself so like I said
1:25:24earlier you now just have to give this
1:25:26dp23 exam to get the Microsoft
1:25:29certification associate data engineering
1:25:31certification okay so just one exam and
1:25:34now you will get the certification for
1:25:36it all right so now let's move on guys
1:25:39and now let's talk about how you guys
1:25:41can get started in clearing this exam
1:25:43and developing those skills and go ahead
1:25:46and apply for the job and become a
1:25:47successful aor data engineer so we have
1:25:51mentioned a lot of things that you have
1:25:52to learn but how exactly you should go
1:25:55forward and start learning these things
1:25:57let's go ahead and clear that out so
1:26:00first of all guys what we can do for you
1:26:02is you can basically refer to a lot of
1:26:04blocks that we have written on or we
1:26:06frequently updat videos on YouTube as
1:26:09well such as this video which has info
1:26:11about the engineering certification or
1:26:13how to become an Azor data engineer now
1:26:16we frequently put more videos for such
1:26:18topics now you can go through them and
1:26:20basically get a jump start into how you
1:26:22you can prepare now my recommendation to
1:26:25you will be to plan out your working
1:26:27plan ask after or before working shift
1:26:30of yours now you should at least spend 3
1:26:32to four hours every day for the next two
1:26:34or 3 months in order to clear this
1:26:37certification exam otherwise if you
1:26:39don't invest this much amount of time
1:26:41guys it is going to be very difficult to
1:26:44crack the exam because there are a lot
1:26:46of things to learn especially if you're
1:26:48not from a data engineer domain now it
1:26:50is going to be a little difficult
1:26:52because you will have to read
1:26:53documentation you'll have to do Hands-On
1:26:56right now at some point you will get
1:26:58stuck and you'll have some issues you
1:27:00have to figure out what went wrong or
1:27:02what's wrong with the hands on I know
1:27:04this because I've also been there guys
1:27:06but getting stuck is the most beautiful
1:27:09part of learning anything because that
1:27:11is where you start the actual research
1:27:13and about how things work right so spend
1:27:17at least 2 to three hours every day
1:27:18either after your workshift or before
1:27:21your workshift to basic learn these
1:27:22Technologies right and now for those
1:27:25people who feel like they do not have
1:27:26the time or they do not want to invest
1:27:28time in researching and they want
1:27:30someone to help them out in getting this
1:27:32exam cleared and become a successful AER
1:27:35data engineer so guys we at a directa
1:27:38also offer a course of Microsoft Azor
1:27:40certification training of the Azor data
1:27:42engineer associate certification course
1:27:45and we also have a master program as
1:27:47well all right which will basically help
1:27:49you in clearing the aor data engineer
1:27:51associat certification right now if you
1:27:54need a helping hand and you need someone
1:27:56or you need to be taught by someone
1:27:57who's already cleared this examination
1:27:59and is already working as a data
1:28:01engineer in the industry then this is
1:28:04the right course for you
1:28:06[Music]
1:28:10guys so why as dat Factory because again
1:28:13we know that we have been generating
1:28:16data at an exponential rate especially
1:28:18since last 5 years and in 2015 we were
1:28:22generating data at the rate of almost
1:28:233.4 exabytes per month and now we are
1:28:27generating at the rate of almost more
1:28:29than 44 exabytes of data per month but
1:28:33that's almost 15 times increase in the
1:28:35amount of of data in just a span of 5
1:28:39years and it is going to be increased by
1:28:41almost 30% more because of the current
1:28:44lockdown fa by all the entire globe and
1:28:46that's why the entire consumption of
1:28:48data is again has been increased
1:28:51tremendously right that is that exactly
1:28:53is what we have the that's why we need
1:28:56to have the most optimized solution for
1:28:58data driven Solutions out there
1:28:59especially on the cloud computing
1:29:01platforms right and mod and modern data
1:29:04handling requires us to move from on
1:29:06promise to database to Cloud database
1:29:09services and that to quickly and then
1:29:11here we have to make sure again and
1:29:13that's why the data needs processing and
1:29:16goes through a series of steps making
1:29:18the process tedious because again when
1:29:20we are trying to load the data if we are
1:29:22trying to migrate other data from our
1:29:24database servers from our on premise
1:29:26storage Services then that has to go
1:29:28through a series of steps and that makes
1:29:30the entire process much more complicated
1:29:32it makes it much more slower as well and
1:29:35plus since it requires a good amount of
1:29:37investment both in times of both in
1:29:39terms of time and money it becomes we
1:29:41can see not feasible at all and data
1:29:44Factory simply help us in automating
1:29:46this entire process and does serve the
1:29:48cost that exactly is why we have data
1:29:51Factory
1:29:53now what exactly is data Factory what
1:29:55exactly is data Factory here so data
1:29:57Factory is basically a cloud-based
1:29:59integration service through which we can
1:30:03Define the entire we can design the data
1:30:05the entire workflow of data in cloud and
1:30:08making sure that we we can use for
1:30:11orchestration and for the automation
1:30:13purposes so for example suppose if we
1:30:15have now if you want we can create a
1:30:17complete pipeline we can Define what is
1:30:19a source and how the pipeline should be
1:30:21structured how the data is coming in
1:30:24where to where it should be stored and
1:30:26that to on a regular manner so we can
1:30:28Define the entire Pipeline and we can
1:30:30automate the entire data movement that
1:30:32means if we make any changes to any
1:30:34particular file or any objects in the
1:30:36oromis that will be automatically
1:30:38replicated or we can say moved into the
1:30:41cloud services as well by the help of
1:30:43data
1:30:44Factory right and using data Factory
1:30:47here we can create and schedule the
1:30:48entire data workflows called as
1:30:50pipelines that can inest data from
1:30:52disparate s data source for example if
1:30:54we have multiple data source defined
1:30:56here we can simply injest data from
1:30:58multiple sources and then we can process
1:31:00by simple a single pipeline it is
1:31:03basically used for processing large
1:31:05amount of large volume of data and these
1:31:09are done by using by integrating it with
1:31:11services that we have SEO SG inside Hado
1:31:14spark SEO data L analytics and is your
1:31:17machine
1:31:18learning so basically if we are looking
1:31:20to make sure that we have we have the
1:31:23optimized data available and optimized
1:31:25stream of data available for analytics
1:31:28and for machine learning then we have to
1:31:30do that by using data Factory to
1:31:31maintain that
1:31:33consistency and if you talk about the
1:31:35entire workflow here first of all the
1:31:37pipelines are all datadriven workflows
1:31:40in SEO data Factory typically performs
1:31:42the following four steps it simply help
1:31:44us in connecting and collecting the
1:31:46entire data set so first of all if we
1:31:48have multiple sources if you have wide
1:31:51variety of sources here we have to make
1:31:53sure we are able to set the source for
1:31:55each and every each and every
1:31:58connector we have to make sure that we
1:32:00are able to connect you to multiple
1:32:03services and then we can through
1:32:05multiple apis or if you have multiple
1:32:07sources of data says for example if we
1:32:09have data from our own CRM coming in if
1:32:11we have data from our own Erp tool from
1:32:13multiple social media analytic PL
1:32:15dashboard here we have to connect and
1:32:16collect all the different data sets then
1:32:19we have to transform and enrich because
1:32:21data always consists of multiple
1:32:23anomalities right they may be multiple
1:32:26missing values they will be some
1:32:28incorrect data formats available they
1:32:30the the data may be incomplete or it may
1:32:33be maybe multiple mistakes in that data
1:32:35corre so you have to make sure that we
1:32:37take care of entire data transformation
1:32:39that means if you want to convert the
1:32:40data format from one to the other like
1:32:42we have the ETL so again we can take
1:32:44care of the transformation and the
1:32:47enrichement of data that means data
1:32:49pre-processing and making sure it is
1:32:51much clear for and it is usable by the
1:32:54antical platforms for which we want to
1:32:55use it then we have to publish it to
1:32:58make it usable and then we have to
1:33:01continuously monitor it so that the
1:33:03entire process can be monitored in case
1:33:05there have been some breakdowns in any
1:33:07of the Clusters from which we are
1:33:09collecting data then that can be
1:33:10monitored and if something can be done
1:33:12it can be it can be processed as well we
1:33:15can do that now let's understand each
1:33:18and every concepts for data Factory here
1:33:20so if you talk about the enti Concepts
1:33:22when first of all we do have pipeline we
1:33:24do have something as pipeline so
1:33:26pipeline is a logical grouping of
1:33:28activities for example Suppose there are
1:33:3110 different sequence like we discussed
1:33:33first of all we have to connect then we
1:33:35can transform then we have to process it
1:33:37then we have to make it available for
1:33:39the analytical platforms so there if
1:33:41there are multiple sequences that needs
1:33:43to be followed again that then that is
1:33:45something that we Define as a part of
1:33:47pipeline alog together and then we have
1:33:49data sets so obviously without data data
1:33:51sets the entire data pipeline is of is
1:33:54of no use so data set simply represent
1:33:56the data structure within the available
1:33:59data store so there have multiple data
1:34:01stores available what exactly those are
1:34:03we are going to discuss step by step and
1:34:06then we have activities for activity
1:34:08simply represents a processing step in a
1:34:10pipeline for example there are 10
1:34:12different steps here for example we have
1:34:14to first of all connect to the source
1:34:16then we have to work on transforming
1:34:18data Exel then we have to work on
1:34:20pre-processing then we have to work on
1:34:22setting up the connection so each now
1:34:24this entire process itself is called as
1:34:27pipeline where we it consists of
1:34:28multiple sequence steps and each and
1:34:31every individual step here each and
1:34:32every individual step these are termed
1:34:34as activities so a pipeline is what we
1:34:36can say pipeline is simply a collection
1:34:38of different activities so we can so
1:34:41that activity is focused on completing
1:34:43one different task and pipeline simply
1:34:46defines the structure or we can say
1:34:48sequence of those STS and it simply
1:34:50makes sure that that sequence is
1:34:51followed whenever the entire pipeline is
1:34:53being
1:34:56implemented so let's erase this
1:34:59up and next we have Link services so it
1:35:03simply the information needed to connect
1:35:04to the external sources like for example
1:35:06we have apis if you have third party
1:35:08vendors and we have third party sources
1:35:10then again we do need an active APS for
1:35:12that and that's why these are all termed
1:35:14as link sources from which we can Source
1:35:16the entire assets here these are all
1:35:18additional links
1:35:20available and as we know data Factory is
1:35:23basically used for on premise itself if
1:35:25we are looking to to connect this if we
1:35:27are looking to connect this to multiple
1:35:29on premise data sets here then that
1:35:30exactly is what we use it for let's
1:35:33understand what exactly is data L
1:35:35service so data L as we know data l so
1:35:38data L as you know is simply an
1:35:40Enterprise wide hyperscale repository so
1:35:43here okay it is simply an Enterprise
1:35:44wide hyperscale repository for big data
1:35:46analytics workload and now it simply
1:35:49holds now it has a capability of
1:35:52petabyte so it can hold data for any
1:35:55size it will allow us to do again using
1:35:57that we can we can do multiple
1:36:00operational and exploratory analytics as
1:36:02well for example if we have multiple
1:36:04sources of data so for example if we
1:36:06have data sources like from on premise
1:36:08we have sensors data like for example if
1:36:10we when we talking about aviation
1:36:12industry then we have multiple we have
1:36:14tons of data sets coming it from
1:36:15different sensors especially for and
1:36:17same way for M for any manaing sectors
1:36:20as well same way if we have any data
1:36:22connected for any websites for example
1:36:24we are talking about any so any uh any
1:36:27stream of data coming in for analytics
1:36:29for any special websites for example
1:36:31suppose we have any big e-commerce
1:36:32solution for Amazon flip cart Airbnb so
1:36:35they have tons of data available on the
1:36:37platforms if we had data coming in from
1:36:39different devices from be it can also be
1:36:42a part of the I complete iot networks
1:36:45they can be data in the format of videos
1:36:47multiple social social media streams
1:36:49coming in for example we have streams of
1:36:51Facebook on Twitter redit on Reddit on
1:36:55multiple platforms so if you have
1:36:56multiple social media streams coming in
1:36:58in the formats of post or if you or
1:37:00let's say if you want to understand the
1:37:02real time I can say if you use case is
1:37:05to work on studying the user sentiments
1:37:09right then that in that case we have to
1:37:10make sure we are we are pitching in the
1:37:13multiple social media streams coming in
1:37:15if we have data in the format of images
1:37:17on application then these all Concepts
1:37:19have to be these all sources needs to be
1:37:21connected as a part of data Factory
1:37:24right and that's why here we can in now
1:37:26once we have these data sources
1:37:28available then we can connect it to ad
1:37:30analytics we can connect this to HG
1:37:32insight to R spark and machine learning
1:37:35purposes now if you have been aware if
1:37:38we have been aware of the fact such as
1:37:41now we can say data warehousing we can
1:37:43say data Lake works like a data
1:37:45warehouse for example if you're working
1:37:47on the realtime analytics right and if
1:37:50you have data sources from multiple
1:37:51platforms so instead of connecting each
1:37:53and every Source manually or we can say
1:37:55one by one to spark we can store the
1:37:57entire data and a send a centralized
1:37:59location and then we can we only need to
1:38:01connect a center location with a single
1:38:04connection to spark that's it so it
1:38:06simply improvises the entire entire
1:38:09performance as well just like we have
1:38:10red shift available in awss same way
1:38:14here we have data Le now let's also
1:38:17understand multiple data L Concepts
1:38:19let's understand multiple data Le
1:38:20Concepts here so data Lake as we know
1:38:23again has multiple components inside it
1:38:25like we have analytics now in data Lake
1:38:28components we have anal we have
1:38:29components for analytics such as a
1:38:31Insight we have AO data L available we
1:38:35can use data store for the complete
1:38:37storage purposes or for analytics we can
1:38:40indicate this with a inside and then we
1:38:42have hard inside and then we have AO
1:38:45data
1:38:46available and then when we are trying
1:38:49again in terms of Analytics we have to
1:38:51make sure we are we do remember these
1:38:53three main key points analytics we can
1:38:55do on data of any size there is no
1:38:58limitation users all users are
1:39:01productive on day one and then we have
1:39:04to make sure that ready it is all scale
1:39:07up exactly as per our Enterprise
1:39:09requirement we have to make sure of that
1:39:11part in terms of type of data stored
1:39:14here it supports all data types we have
1:39:16structured semi-structured and
1:39:18unstructured data types supported we
1:39:21have CSC files we have XML files right
1:39:23we have the emails we have Jon's files
1:39:25right so again these all are a part of
1:39:28sem instructure data when we don't have
1:39:30a direct format available on which we
1:39:32can start and perform the entire sorting
1:39:35and that is example for semi structure
1:39:36data set and then we have unstructured
1:39:39data set like we have images videos
1:39:42audio clips these all are part of semi
1:39:44structured data
1:39:46sets and then if you compare data League
1:39:49to Data Warehouse if we compare data
1:39:52leak to Data Warehouse here again data
1:39:54leak as you know is simply complimentary
1:39:57to the data warehouse whereas data where
1:40:00if you talk about the data warehousing
1:40:02service it may be soured to data leag
1:40:04data leag is basically used for detail
1:40:06data whereas as we can see data
1:40:08warehouse is basically used for filter
1:40:11summarize and for refinement of data all
1:40:13together data L offers schema on read
1:40:16whereas data house off for schema on
1:40:18right write and data L has one language
1:40:20to process data of any format whereas in
1:40:23data warehouse we can process it using
1:40:25the SQL complying it all together now
1:40:28let's move into the handson here and see
1:40:30how exactly a this is implemented on top
1:40:33of Edo data Factory portal here data
1:40:36warehouse and dat let's understand this
1:40:38by simple use case here and for doing
1:40:41that let's open up our notepad here for
1:40:45example let's say we have multiple data
1:40:47sources for example we have data sources
1:40:49available from our CRM correct for
1:40:51example our main use cases we want to
1:40:53perform the analysis for sales report
1:40:55for sales for any company correct we are
1:40:57here to perform analysis or multiple St
1:40:59data here right or suppose if we have
1:41:01data coming from from CR and we have
1:41:03data from Erp tools as well right and
1:41:06then we have data from the old data set
1:41:09available for example if we have the own
1:41:11CC file also stored here so if you have
1:41:13multiple sources of data coming in then
1:41:16in in if you want to Club it if you want
1:41:18to use this on top of any iCal platform
1:41:22so there are two ways of doing that
1:41:23either we have to connect each and every
1:41:25source with the antical platform and
1:41:28then only we can start working on it
1:41:30correct using the concept of data
1:41:33warehouse what we can do we can connect
1:41:34each and every Source we can store the
1:41:36the data from each of source to a
1:41:38centralized location as in called as a
1:41:41data house right and once the data is
1:41:43available in data house then we can
1:41:46connect that data warehous directly to
1:41:48us through a single collection to our
1:41:50artical platform so that whatever
1:41:51analyst we are trying to perform the
1:41:53entire process can be streamlined as a
1:41:55part of data warehousing
1:41:57service let's erase this
1:42:02up for getting started on data Factory
1:42:05here we can open we can come back to a
1:42:08portal so here we can come back to a
1:42:10portal in case you have not signed up on
1:42:12asure we can open this entire portal
1:42:14here where we can start working on top
1:42:17of data Factory one one by one now for
1:42:20getting started here what we can do is
1:42:21we can simply come back we can simply
1:42:23come back
1:42:25and first of all for having the for
1:42:28having the entire database that we can
1:42:30install and can interpret locally we can
1:42:32use this we can use entire platform
1:42:35called as SS SMS in case you don't have
1:42:38the enre setup done at your end then we
1:42:40can simply go ahead and download the SQL
1:42:42Server management studio in case you
1:42:44don't have the setup done now once we
1:42:46have configured these now what we can do
1:42:47we can come back to our a portal now in
1:42:49this s portal what what we can do we can
1:42:51simply first of all once we are into
1:42:53dashboard here we have to work around
1:42:55with first of all setting up the entire
1:42:57data warehousing service so what we can
1:42:59do is here we can set up the entire dat
1:43:01housing service and for setting it up we
1:43:03can come back to our entire dashboard
1:43:05now for creating a new resource we here
1:43:07we have to click on this option which
1:43:08says create a resource and then we can
1:43:11choose resource type here so from
1:43:13databases here we can choose now yeah
1:43:15these are the most commonly used
1:43:17resources here for example if you want
1:43:18to go for web application for functional
1:43:20applications for SQL databases we can
1:43:23choose accordingly and if you want to
1:43:25start with SQL database we can open up
1:43:27SQL database and then we can set up the
1:43:29entire database platform but again in
1:43:31here our main goal is not to create a
1:43:33SQL database but to set up the entire
1:43:35data with housing first and then we can
1:43:38saish in data services and then get
1:43:41started on top of it right so here what
1:43:43we can do here e either we can go ahead
1:43:45and search for the services using these
1:43:47categories or we can use it directly we
1:43:50can use the the service bar to search
1:43:52for the services directly for example
1:43:55here we can use data
1:43:59L as you can see currently we are
1:44:01planning to use data l so here we can
1:44:03choose data L
1:44:06here now if you want to create it now
1:44:08here we can simply click on create here
1:44:10we can Define the name of the current
1:44:12data storage service that we are going
1:44:13to create we can Define the entire
1:44:18name for example let's say we have we
1:44:20call it as AA app itself here we can
1:44:24choose subscriptions so the currently
1:44:26here either we can go for two type of
1:44:28subscription here we can choose pay as
1:44:29you go or here we can choose the the fre
1:44:31file if in case we have fre file
1:44:33available then we can choose that then
1:44:35here we can choose Resource Group in
1:44:37case we don't have the resource Group
1:44:39created we can use a new Resource Group
1:44:43here let's enable the entire Services
1:44:46there so that we can so we can enable
1:44:48the entire subscription model here
1:44:51so let's enable the IND subscription
1:44:53model here we can verify the
1:44:56code in case we haven't set it up we can
1:44:59simply set it up you by simply
1:45:01specifying our entire
1:45:04details set it up so that we can use it
1:45:07for creation of our enre data data
1:45:09league and data warehousing
1:45:11platform even though once you if we are
1:45:14siging this up for the first time here
1:45:16once we and once we sign it up we will
1:45:19be charge a nominal for a normal fee for
1:45:22making sure the entire account is well
1:45:25authenticated that we can use as a part
1:45:27of signing up for as your platforms and
1:45:29once it is done this will take us back
1:45:31to the D
1:45:34dashboard now once we have enter the
1:45:36account here we can simply click on
1:45:38create resource and here we can choose
1:45:40data L for example let's say we are
1:45:42starting to we are planning to start
1:45:44with the by setting up entire data
1:45:46warehouse first right so ear it was
1:45:48called as data warehouse which was now
1:45:50ch is to data leag itself now we can
1:45:52open up da data leag account all all
1:45:55together so here we can go for data leag
1:45:57generation here we can click on create
1:45:59here we can Define the entire resource
1:46:01name let's say we name it ASA here we
1:46:04can choose the subscription model that
1:46:06we want plan to go for here we can
1:46:08choose Resource Group Resource Group are
1:46:09simply like when we are planning for
1:46:12creation of now when we are creating 10
1:46:14different resources in insure and at the
1:46:16end we want to Simply segregate them
1:46:18based on a certain project for example
1:46:20we want to we want to know which all
1:46:23services have been deployed for which
1:46:25particular project and if you want to
1:46:27see a consolidate billing for that
1:46:29project then we can go for Resource
1:46:31Group for example let say here we are
1:46:32going to create a resource Group for app
1:46:35one we can choose any available Resource
1:46:37Group name that you want to go ahead
1:46:41with all right and here we can go for p
1:46:44ASO model we can choose the encryption
1:46:46if you want to enable encryption we can
1:46:48choose the keys or we can simply choose
1:46:49the default keys from
1:46:52ad suppose let's say we choose a key
1:46:54from the key generation itself we can
1:46:57configure this
1:47:00up it may take a couple of second for
1:47:02this entire data to be
1:47:09created and as you can see it says the
1:47:12deployment in in progressor so it may
1:47:14take a couple of of minutes for this
1:47:15entire data to be processed here step by
1:47:17step and once it is processed we would
1:47:20be able to see this entire in live
1:47:21action and in the meantime in case you
1:47:23want to go ahead and set up the entire
1:47:25SMS that we have discussed we have to
1:47:27have the s SMS so that we can connect
1:47:29our local SQL databases and then we can
1:47:32import directly into the a
1:47:34portal and then we also have to set up
1:47:37entire data Factory for setting up the
1:47:39entire data Factory here we can open up
1:47:40the the data Factory servers available
1:47:42Ino platform here we can click on ADD
1:47:46and here we have to define the data
1:47:48Factory name for example let's say we
1:47:50want to call us as
1:47:53Eda DF as in data Factory here we can
1:47:56choose the current version so basically
1:47:58earlier we have been using V1 but now
1:48:00since last year we have been using V2 so
1:48:03here we can choose the subscription
1:48:04model that we have subscribed for our
1:48:06for we can choose Resource Group that we
1:48:09have already created by the name of Eda
1:48:11app so that if we have 10 different
1:48:13Services if we have 40 different
1:48:15resources being deployed then we can
1:48:17again all of those resources are mapped
1:48:19for single application then we can map
1:48:22it to a single application Al together
1:48:23we can see that and once it is done we
1:48:26can simply choose a location in which we
1:48:28are trying to launch this particular
1:48:30data Factory for again we have to choose
1:48:33a region which is obviously closer to
1:48:35our end users because at the end if we
1:48:38are not using a a location which is not
1:48:40closer to end users then that will
1:48:42create a huge amount of latency so for
1:48:44example if our users base as based our
1:48:47user are based in Singapore so here we
1:48:49can choose region for Singapore Al
1:48:50together if we know that you our users
1:48:52are based not from Singapore suppose
1:48:55from from Central India Australia East
1:48:57Africa north from Europe from USC so
1:48:59here we can choose our regions
1:49:01accordingly for example suppose here we
1:49:02want to start with North Europe so here
1:49:05we can choose North Europe and then if
1:49:07you want to specify a get URL that will
1:49:09be used as a as a main data source here
1:49:12for for this data Factory for example
1:49:14let's say we use our own GitHub URL for
1:49:17this
1:49:18one let's use L into our
1:49:22GitHub and for example suppose here we
1:49:24may have any particular depository here
1:49:27we have any depository that we want to
1:49:28connect here we can simply connect that
1:49:31let's say here we can enter our GitHub
1:49:44URL here if we have multiple
1:49:46repositories here for example here we
1:49:48have repository as 1 we can choose enti
1:49:51one here if we have Branch as we want to
1:49:53go for the for let's suppose for the
1:49:57development Branch so here we have a
1:49:58branch name for development we can
1:50:00choose a branch
1:50:01name so here we have the entire get URL
1:50:04here we have to define the repository
1:50:06name and then we have to define the
1:50:07branch name and then we can choose a
1:50:09root folder for which we are trying to
1:50:11connect to the connect to here so for
1:50:13example if we have Ro folder by the name
1:50:15of index we can Define index and now we
1:50:18can click on create
1:50:24So currently this is getting
1:50:30initialized as you can see our
1:50:32deployment is done data L and currently
1:50:35the deployment is being in progress for
1:50:38data Factory that we currently deployed
1:50:40so
1:50:49far
1:50:50now currently this has been deployed
1:50:52here now if you want to download the
1:50:54entire deployment details we it's a good
1:50:56practice to download the entire
1:50:57deployment details which will contain
1:50:59the entire list of all the piece of of
1:51:01information and then for getting the
1:51:03entire operations detail here we can
1:51:04simply choose the operations detail
1:51:06where we can Define the entire
1:51:08operations name duration and now if you
1:51:10want to configure this we can open up
1:51:12the
1:51:14resource where and this resource here we
1:51:17first of all here if you want to Quick
1:51:18Start if you want to see the entire
1:51:20activity that means what exactly it has
1:51:22been and again what exactly has been the
1:51:24iops the storage usages and the
1:51:27processor usage and we can go ahead and
1:51:28see the entire monitoring
1:51:32part and if you are Conn if you are
1:51:34planning to connect to to this part
1:51:36instance we have to make sure we are
1:51:38defining the current IM user rules and
1:51:41their policies because again if we are
1:51:43using this through some some particular
1:51:45account here then we can end as we
1:51:47Define the rules for them they will not
1:51:48be able to make any changes to this
1:51:50particular server out there that is
1:51:52something that we have to take care
1:51:54of and along with that let's go ahead
1:51:57and create one storage account as well
1:51:59so again for doing that we can use a
1:52:01service bar available on top so here we
1:52:05have to open up our storage account
1:52:07let's open this
1:52:09up now currently as you can see
1:52:11currently by default start up there
1:52:13won't be any storage account created and
1:52:15deployed so here we have to go ahead and
1:52:16create one click on create storage
1:52:19account
1:52:26here we can choose a subscription model
1:52:28here we can choose a resource Group for
1:52:29the same application resource Group that
1:52:31we have currently created here we have
1:52:33to define the entire storage account
1:52:35name let's say we name it ASA itself and
1:52:39then we can choose the location now
1:52:40remember
1:52:41this okay we already have configured
1:52:43that so here okay one
1:52:46two so here we have to choose the same
1:52:49location in which we have deployed the
1:52:51other data Lake and our data wouse our
1:52:54data Factory
1:52:55itself we have deployed this for
1:52:57northern Europe so here we can choose
1:52:59North Europe and based on where our
1:53:02locations are where users are basically
1:53:04located then we can choose the
1:53:06performance to be standard and premium
1:53:09so again in terms of account kind here
1:53:11we can choose storage or we can choose
1:53:14if you're going for general purpose
1:53:15version one or version two so general
1:53:17purpose one again they are depending
1:53:19upon requirement we can choose
1:53:20accordingly if you want a higher
1:53:22performance that means if we are looking
1:53:23to to have a higher workload then we can
1:53:26use Gen 2 that was released early last
1:53:29year and then if you want to replicate
1:53:31the entire data we want to go for Zone
1:53:34redundant we want to go for locally
1:53:35redundant as if we want to maintain
1:53:37local copy or we want to maintain a read
1:53:39access R st that means again based on
1:53:42multiple regions it will be copy but
1:53:45again the copy will have only the read
1:53:46only access that means this will be used
1:53:49as a primary account for storing data
1:53:51for writing data and the replications
1:53:53will be used for just for the read
1:53:55purposes so that the entire situation of
1:53:58Bott leg is also not
1:54:01created and then we can choose assets
1:54:03here to be hot or to be cool depending
1:54:05upon the requirement we can choose it to
1:54:07be hot or cool here just like we in case
1:54:09we have been Avail if we have been
1:54:11familiar with the concept of multiple
1:54:13storage classes in AWS same way we we
1:54:16have standard and then we have
1:54:17infrequent access so here we can chose
1:54:20school if you want to move to some to
1:54:21something like infrequent access which
1:54:23is not used frequently or we can keep it
1:54:25to hard for those standard we can say
1:54:27most frequently used files here so
1:54:29currently we can keep it to host then we
1:54:32have to define the networking point if
1:54:34we are trying to deploy this in our own
1:54:35isolated Network then we can choose
1:54:38public then we can choose a private or
1:54:40public endpoint depending upon the
1:54:42requirement here so if we have if we
1:54:44have a virtual Network created then we
1:54:46can go ahead and use selected networks
1:54:49if you want to deploy this on public
1:54:50endpoints we can go for public endpoints
1:54:52or if again in case you want to deploy
1:54:55this on our private then first of all we
1:54:57have to configure our prior end point
1:54:58and then only we can get started then we
1:55:01can Define the production type here
1:55:03whether we want to go for stop or
1:55:05desktop relite or if sof as in if you
1:55:08want to retrive it we can simply do that
1:55:10and then we have the Gen 2 hierarchal
1:55:12should be disabled as you don't need it
1:55:14as a part of our Handel currently so we
1:55:16can keep it to disable then we can find
1:55:19the tags here
1:55:20tags are simply used for sorting and
1:55:23filtration purposes if you want to sort
1:55:25this out later on if there are multiple
1:55:27services that we deployed and now we
1:55:29want an easier way to host it easily
1:55:32when we can easily do
1:55:39that and in here we can Define tasp if
1:55:42you don't want to use it we can review
1:55:47it we have to wait for this one to be
1:55:50reviewed here once we are done we can
1:55:52click on
1:55:54create let's and let's wait for this one
1:55:56to be
1:56:09created if you are looking to move files
1:56:11here from one from one part to the other
1:56:13again we can easily do that is it access
1:56:16to free to create database for creating
1:56:19database here here we here we can choose
1:56:21the engine for database for creation of
1:56:23multiple databases here we have
1:56:25different services for for doing that
1:56:27for example suppose for creation of
1:56:29databases here we can click on create
1:56:32resource we can choose resource type as
1:56:34databases and here we can choose which
1:56:36part database engine we are going to
1:56:37deploy here just like we have RDS
1:56:39available in adus same way here we can
1:56:42change it
1:56:44up and if we looking to create a
1:56:46complete data pipeline then that's why
1:56:48we use data
1:56:50Factory so for if we are trying to to
1:56:54transfer one data from the other account
1:56:56here we can simply share the across
1:56:57multiple accounts if you're trying to
1:56:59start to transfer data from on premise
1:57:02to SEO or from SE to on premise again
1:57:04for if we looking to transfer from aure
1:57:06to on premise then we can from on a to
1:57:09on premise that me locally then we can
1:57:11use another service called as storage
1:57:13Explorer as a part of AO platform we can
1:57:15use that if we looking to transfer from
1:57:19uh Z platform to on premise we can do
1:57:21that we can take the help of storage
1:57:23Explorer as
1:57:26service and then we can easily use the
1:57:29AO import export to Simply transfer the
1:57:31service from our AO platform to oise
1:57:34just like we have a physical device
1:57:36offered by as a part of snowball in AWS
1:57:39in case we have been working with AWS
1:57:41and we there we have a service called as
1:57:43snowball where it is a simple physical
1:57:45device for if we looking to get
1:57:47connected we can use that we can take
1:57:49take the help of storage gate phase in
1:57:50that so just like storage gate phase
1:57:53here we can take the help of storage
1:57:54Explorer back into Azure for transfering
1:57:58data from in and out of SEO
1:58:02[Music]
1:58:06platforms now let's have an introduction
1:58:08of a database first let's understand why
1:58:10do we need a database so the various
1:58:13reasons a database is important first of
1:58:15all it manages large amounts of data a
1:58:17database stores and manages a large
1:58:18amount of data on a daily basis this
1:58:20would not only be possible using any
1:58:22other tool such as a spreadsheet as they
1:58:25would simply not work second is its
1:58:27accuracy so a database is pretty
1:58:29accurate as it has all sours of building
1:58:31constraints checks Etc this means that
1:58:33the information available in database is
1:58:35guaranteed to be correct in most cases
1:58:38it's easy to update data in a database
1:58:39so in a database it is easy to update
1:58:41data using like various data
1:58:43manipulation languages available one of
1:58:44these languages SQL for the security of
1:58:47data so databases have various methods
1:58:49to ensure security of data there are
1:58:52user logins required before accessing a
1:58:54database and various access specifiers
1:58:57these allow only authorized users to
1:58:59access the database fifth is data
1:59:01Integrity this is ensured in databases
1:59:03by using various constraints for data
1:59:06data Integrity in databases makes sure
1:59:08that the data is accurate and consistent
1:59:10in a database the last is easy to
1:59:12research data it is very easy to access
1:59:15and research data in a database this is
1:59:17done using data Cy language which allow
1:59:20searching of any data in the database
1:59:21and performing computations on it now
1:59:24that you have understood the need of a
1:59:25database let's briefly understand what
1:59:26actually it is so a database is an
1:59:29organized collection of structure
1:59:30information or data typically stored
1:59:32electronically in a computer system a
1:59:34database is usually controlled by
1:59:36database management system together the
1:59:38data and the database management system
1:59:40along with applications that are
1:59:41associated with them are referred to as
1:59:43a database system often shortened to
1:59:46just a database so data within the most
1:59:48common types of database in operation
1:59:50today is typically modeled in rows and
1:59:53columns in a series of tables to make
1:59:55processing and data quering efficient
1:59:57the data can then be easily accessed
1:59:59managed modified updated controlled and
2:00:02organized most databases are structured
2:00:04query language for writing and querying
2:00:06data databases are used to support
2:00:08internal operations of organizations and
2:00:10to underpin online interactions with
2:00:12customers and
2:00:14suppliers databases are used to hold
2:00:16administrative information and more ized
2:00:19data such as engineering data or
2:00:21economic models example includes
2:00:23computerized Library System flight
2:00:24reservation system computerized past
2:00:26inventory system and many content
2:00:28Management systems that store websites
2:00:30as collection of web pages in a database
2:00:33now that you have an understanding of
2:00:34Microsoft aour as well as of a database
2:00:36let's now take a look at different types
2:00:38of databases in Azure first is
2:00:40relational database a relational
2:00:41database is a type of database that
2:00:43stores and provide access to data points
2:00:46that are related to one another
2:00:47relational databases are based on the
2:00:49relational model an intuitive
2:00:51straightforward way of representing data
2:00:53in tables in a relational database each
2:00:56row in the table is a record with a
2:00:57unique ID called the key The Columns of
2:01:00the table hold attributes of the data
2:01:02and each record usually has a value for
2:01:04each attribute making it easy to
2:01:07establish the relationships among data
2:01:08points in a relational database all data
2:01:11is stored and accessed by relations so
2:01:13relations that store data are called
2:01:15base relations and in implementations
2:01:17are called tables other relations do not
2:01:19store data but are computed by applying
2:01:21relational operations to those relations
2:01:24these relations are sometimes called
2:01:26derived relations in implementations
2:01:28these are called views or queries
2:01:30derived relations are convenient in that
2:01:32they act as a single relation even
2:01:34though they may grab information from
2:01:35several relations each relation or table
2:01:38has a primary key this being a
2:01:39consequence of a relation being a set a
2:01:42primary key uniquely specifies a tuple
2:01:44within a table while natural attributes
2:01:46are sometimes good primaries so so this
2:01:49is all about relational database then
2:01:51second we have is non- relational
2:01:52database also known as nosql databases
2:01:55so nosql database or non- relational
2:01:57database provides a mechanism for
2:01:58storage and retrival of data that is
2:02:01modeled in means other than the tabular
2:02:03relations used in relational databases
2:02:05so non-national databases are
2:02:06increasingly used in big data and
2:02:08realtime web applications and
2:02:09non-national databases are also like
2:02:11sometimes called not only SQL to
2:02:13emphasize that they may support SQL like
2:02:16query languages or sit alongside SQL
2:02:19databases so the prominent non
2:02:22relational databases provided by aour is
2:02:23Cosmos database that I will explain you
2:02:26further in this video and third is
2:02:29inmemory database so an inmemory
2:02:31database also like in the short form we
2:02:32say it as IMDb also like a main memory
2:02:35database system or mmdb or memory
2:02:37resident database these are all the
2:02:39names of it so an inmemory database is a
2:02:41database management system that
2:02:43primarily relies on Main memory for
2:02:44computer data storage it is contrasted
2:02:47with database management system that
2:02:48emplo deploy a disk storage mechanism in
2:02:51memory databases are like faster than
2:02:52dis optimized databases because dis
2:02:54access is slower than memory access the
2:02:57internal optimization algorithms are
2:02:59simpler and execute fewer CPU
2:03:01instructions accessing data in memory
2:03:03eliminates seek time when querying the
2:03:05data which provides faster and more
2:03:07predictable performance than disk a
2:03:09potential technical hurdle with inmemory
2:03:11data storage is the volatility of RAM is
2:03:14specifically in the event of a power
2:03:15loss intentional or otherwise data
2:03:18stored in volatile Ram is lost with the
2:03:21introduction of nonvolatile Random
2:03:23Access Memory technology in memory
2:03:25databases will be able to run at full
2:03:27speed and maintain data in the event of
2:03:29power
2:03:29failure let's Now understand the
2:03:32architecture of database Services
2:03:33provided by a so you can see the it
2:03:36looks like a complex architecture but I
2:03:38will explain you in quite easily so the
2:03:40basic fundamental building block that is
2:03:42available in aour is the SQL database so
2:03:45Microsoft offers this SQL server and SQL
2:03:47database on aour in in many ways we can
2:03:49deploy a single database or we can
2:03:51deploy multiple databases as part of a
2:03:54shared elastic pool you can see the
2:03:56elastic pools and single database okay
2:03:59Microsoft introduced a managed instance
2:04:01that is targeted towards on premises
2:04:02customers so if we have some SQ
2:04:05databases within our on premises Data
2:04:06Center and we want to migrate the
2:04:08database into Azure without any complex
2:04:10configuration or ambiguity then we can
2:04:12use a managed instance because this is
2:04:14mainly targeted towards on premises
2:04:16customers who want to lift and share
2:04:18their own on premises database into
2:04:19Azure with the least effort and
2:04:21optimized cost we can also take
2:04:23advantage of Licensing we have within
2:04:26our on premises data center Microsoft
2:04:28will be responsible for maintenance
2:04:30patching and related services but in
2:04:32case if we want to go for the
2:04:34infrastructure as a service for the SQL
2:04:36Server then we can deploy SQL server on
2:04:38the Azure virtual machine if the data
2:04:41has a dependency on the underlying
2:04:42platform and we want to log into the SQL
2:04:44server in that case we can use the SQL
2:04:46server on a virtual machine we can Dey
2:04:48Dey SQL Data Warehouse on the cloud
2:04:50Azure offers many other database
2:04:52services for different types of
2:04:54databases such as MySQL marad DB and
2:04:57also post SQL once we deployed a
2:05:00database into a we need to migrate the
2:05:02data into it or replicate the data into
2:05:04it okay then we have is azure database
2:05:07services for data migration so services
2:05:09that are available in Azure which we can
2:05:11use to migrate the data from our on
2:05:12premises SQL Server into Azure so in
2:05:15that the first one is azure data
2:05:16migration service so it is used to to
2:05:19migrate the data from our existing SQL
2:05:20server and database within the on
2:05:22premises data center into the Azure then
2:05:24we have Azure SQL data synchronization
2:05:27if we want to replicate the data from
2:05:29our on premises database into azour then
2:05:31we can use azour SQL data sync then we
2:05:34have SQL stretch database so it is used
2:05:37to migrate cold data into Azure SQL
2:05:39stretch database is a bit different from
2:05:41other database offerings it works as a
2:05:43hybrid database because it divides the
2:05:45data into different types like hot and
2:05:47cold so hot data will be kept in the on
2:05:49premes Data Center and C data in the
2:05:51Azure then we have is data Factory so
2:05:53Azure data Factory is used for ETL means
2:05:56transformation extraction and loading so
2:05:59using the data Factory we can even
2:06:01extract the data from our on premises
2:06:02data center we can do some conversion
2:06:04and load into the Azure SQL database
2:06:07data Factory is an ETL tool that is
2:06:09offered on the cloud which we can use to
2:06:11connect to different databases and like
2:06:13extract the data or transform it and
2:06:15load into a destination then there's
2:06:17Azure security
2:06:19so all the databases that exist in our
2:06:21need to be secured and also we need to
2:06:23accept connections from known Origins
2:06:25for this purpose all these database
2:06:27Services comes with firewall rules where
2:06:30we can configure from which particular
2:06:32IP address we want to allow connection
2:06:34we can define those firewall rules to
2:06:37limit the number of connections and also
2:06:38reduce the service attack area so now
2:06:41let's talk about the cosmos DV so Cosmos
2:06:43DV is nothing but a SQL data store that
2:06:45is available in aour and it is designed
2:06:47to be globally scalable and also very
2:06:50highly available with extremely low
2:06:51latency Microsoft guarantees latency in
2:06:54terms of reading and wrs with Cosmos DV
2:06:57for example if we have any application
2:06:58such as iot or gaming where we get a lot
2:07:01of data from different users spread
2:07:03across globally then we will go for
2:07:05Cosmos DV because Cosmos DB is designed
2:07:07to be globally scalable and highly
2:07:09available due to which users will like
2:07:11experience low latency finally there are
2:07:14two things and one is we need to secure
2:07:16all the services for that purpose we can
2:07:18integrate all these services with azour
2:07:20active directory and manage the users
2:07:22from Azure active directory also to
2:07:24monitor all these Services we can use
2:07:27the security Center so there is an
2:07:29individual monitoring tool too but AZ
2:07:31security Center will keep on monitoring
2:07:33all these services and provide
2:07:34recommendations if something is wrong I
2:07:36hope the architecture of azure database
2:07:37Services now clear to you now let's uh
2:07:40move forward to briefly understand the
2:07:41database Services provided by Azure so
2:07:44the first is azure SQL database so SQL
2:07:47database is the flagship product for
2:07:48Microsoft in the database area it is a
2:07:50general purpose relational database that
2:07:52supports structures like relation data
2:07:54Json spatial and XML the Azure platform
2:07:56fully manages every azour SQL database
2:07:58and guarantees no data loss and a high
2:08:00percentage of data availability azour
2:08:03automatically handles patching backups
2:08:04replication failure detection underlying
2:08:07potential Hardware software or network
2:08:09failure deploying Buck fixes failovers
2:08:11and like database upgrades and other
2:08:13maintenance tasks so there are three
2:08:15ways we can Implement our SQL database
2:08:18so first is managed instance this is
2:08:20premar targeted towards on premises
2:08:22customers in case if we really have a
2:08:24SQL Server instance in our on premises
2:08:26Data Center and you want to migrate that
2:08:28into Azure with minimum changes to a
2:08:30application and the maximum
2:08:31compatibility then we will go for manage
2:08:34instance second is single database so we
2:08:36can deploy a single database on a its
2:08:38own set of resources managed via L
2:08:41logical server okay then we have his
2:08:42elastic pool we can deploy a pool of
2:08:45databases with a shared set of resources
2:08:47managed bya log local server we can like
2:08:50deploy the SQL database as an
2:08:51infrastructure as a service that means
2:08:53we want to use the SQL server on Azure
2:08:55virtual machine but in the case we are
2:08:56responsible for managing the SQL server
2:08:58on that but in that case we are
2:09:00responsible for managing the SQL server
2:09:02on that particular Azo virtual machine
2:09:04so then we have is the purchasing model
2:09:06so there are two ways we can purchase
2:09:08the SQL server on a j so first is voree
2:09:10purchasing model also as virtual core
2:09:12purchasing model so the vco purchasing
2:09:14model enables us to independently scale
2:09:17compute and Storage resources match on
2:09:20premises performance and optimize price
2:09:22it also allows us to choose a generation
2:09:24Hardware it also allows us to use Azure
2:09:27hybrid benefit for SQL Server to gain
2:09:30cost savings best for the customers who
2:09:32value flexibility control and
2:09:34transparency so second is DD model it is
2:09:37based on a bundled measure or compute
2:09:39storage and input output resources so
2:09:43sizes of the compute are expressed in
2:09:45terms of database transaction units
2:09:47means dtus for single databases and
2:09:50elastic database transaction units for
2:09:51elastic pools this model is best for
2:09:54customers who want simple pre-configured
2:09:56resource options in the second database
2:09:59service is azure Cosmos database so
2:10:02Azure Cosmos database is a no SQL data
2:10:04store it is different from the
2:10:05traditional relational database where we
2:10:08have a table and the table will have a
2:10:09fixed number of columns and each row in
2:10:11the table should ADH to the scheme of
2:10:13the table in the no SQL database you
2:10:16don't Define any schema at all for the
2:10:18table and each item or row within the
2:10:21table can have different values or
2:10:23different schema itself so now let's
2:10:25understand the cosmos database structure
2:10:28first one in the structure is database
2:10:30so we can create one or more Azure
2:10:32Cosmos database under our account a
2:10:34database is analogous to a name space
2:10:37and it is the unit of management for a
2:10:39set of azure Cosmos containers so the
2:10:42second is Cosmos account so the Azure
2:10:44Cosmos account is the basic unit of
2:10:46global distribution and high
2:10:47availability for for globally
2:10:48Distributing our data and throughput
2:10:51across multiple Azure regions we can add
2:10:53or remove Azure regions from our Azure
2:10:55Cosmos at any time I mean Azure Cosmos
2:10:57account at any time so the third is a
2:10:59container so an azour Cosmos container
2:11:02is the unit of scalability for both
2:11:04provision throughput and storage of
2:11:06items a container is horizontally
2:11:08partitioned and then replicated across
2:11:10multiple regions then let's understand
2:11:13the types of consistency under Cosmos DB
2:11:15so Azure Cosmos database approaches the
2:11:17data consistency as a spectrum of
2:11:18choices instead of two extremes so
2:11:21strong compatibility and eventual
2:11:23consistency are at the ends but these
2:11:25are many consistency choices along along
2:11:28the Spectrum so the consistency levels
2:11:30are region agnostic the consistency
2:11:32level of our Azure Cosmos account is
2:11:35guaranteed for all read operations
2:11:37regardless of the region from which the
2:11:39reads and rights are served the number
2:11:40of areas associated with the Azure
2:11:42Cosmos account or whether our account is
2:11:45configured with a single or multiple
2:11:47right regions
2:11:48then there's request unit so we pay for
2:11:51the throughput we provision and the
2:11:53storage we consume on an hourly basis
2:11:55with Azure Cosmos DB remember this DB
2:11:58means database so then there are request
2:12:00units in Cosmos DB means Cosmos database
2:12:03so we pay for the throughput we
2:12:04provision and the storage we consume on
2:12:06an hourly basis with Azure Cosmos GB the
2:12:09cost of all the database operations is
2:12:11normalized by Azure Cosmos DV and is
2:12:13expressed in terms of request units the
2:12:15price to readed a 1 KB item is a one
2:12:18request unit all other database
2:12:19operations are similarly assigned with a
2:12:21cost in terms of research units the
2:12:24number of research units consumed will
2:12:25depend on the type of operations item
2:12:27size data consistency query patters etc
2:12:30for the management and planning of
2:12:32capacity Azure Cosmos database ensures
2:12:34that the number of research units for a
2:12:35given database operations over a given
2:12:37data set is deterministic and the third
2:12:40database service is azure data Factory
2:12:42so Azu data Factory is a data
2:12:43integration service based on the cloud
2:12:44that allows us to create data driven
2:12:46workflows in the cloud for for
2:12:48orchestrating and automating data
2:12:50movement and data transformation data
2:12:52Factory is a perfect ETL tool on cloud
2:12:54data Factory is designed to deliver
2:12:57extraction transformation and loading
2:12:58process within the cloud the ETL process
2:13:01generally involves four steps so the
2:13:03first one is connecting collect we can
2:13:04use the copy activity in a data pipeline
2:13:06to move data from both on premises and
2:13:09Cloud secure data stores so the second
2:13:12is a transform so once the data is
2:13:14present in a centralized data store in
2:13:16the cloud process or trans form the
2:13:18collected data by using compute services
2:13:20such as HD Insight Hadoop spark data
2:13:23leak analytics and machine learning
2:13:24third is published so after the raw data
2:13:26is refined into a business ready
2:13:28consumable form it loads the data into
2:13:30an azour data warehouse azour SQL
2:13:33database and Azure Cosmos database Etc
2:13:35so fourth is Monitor so azour data
2:13:37Factory has built in support for
2:13:39pipeline monitoring via azour monitor
2:13:42API Powershell log analytics and health
2:13:44panels on the azour portal so then there
2:13:46are components of data so data Factory
2:13:48is composed of six key elements all
2:13:51these components work together to
2:13:52provide a data form on which you can
2:13:54form a datadriven workflow with the
2:13:56structure to move and transform the data
2:13:58so first one is pipeline a data Factory
2:14:00can have one or more pipelines it is a
2:14:03logical grouping of activities that
2:14:05perform a unit of work the activities in
2:14:08a pipeline perform the task Al together
2:14:10for example a pipeline can contain a
2:14:12group of activities that inest data from
2:14:14a Azure blob and then runs a hi query
2:14:17and an HD inside cluster to partition
2:14:19the data so second is activity it
2:14:21represents a processing step in a
2:14:23pipeline for example we might use a copy
2:14:25activity to copy data from one data
2:14:27store to another data store then we have
2:14:29a data sets so it represents data
2:14:31structure within the data stores which
2:14:33point to or reference the data or we
2:14:35want to use inov activities as input or
2:14:38output then there are Link services so
2:14:41it is like connection strings which
2:14:43Define the connection information needed
2:14:45for data Factory to connect to external
2:14:47resources
2:14:48a linked service can be a data store and
2:14:50compute resources linked service can be
2:14:53a link to a data store or a compute
2:14:54resource also then we have a triggers so
2:14:58it represents the unit of processing
2:14:59that determines when a pipeline
2:15:00execution needs to be disabled we can
2:15:02also schedule these activities to be
2:15:04performed at some point in time and we
2:15:07can use the trigger to disable an
2:15:09activity then the last one is control
2:15:11flow so it is an orchestration of
2:15:12pipeline activities that include
2:15:14chaining activities in a sequence
2:15:15branching defining parameters for the
2:15:17pipeline label and passing arguments
2:15:20while invoking the pipeline on demand or
2:15:22from a tiger we can use a control flow
2:15:25to sequence certain activities and also
2:15:27Define what parameters need to be passed
2:15:30for each of these activities I hope you
2:15:33have now understood the major Services
2:15:35of azure databases so now let's have a
2:15:38look at some of the use cases for Azure
2:15:39database Services first let's see the
2:15:41use cases for SQL database so the first
2:15:44one is developer or test environment an
2:15:46important use case for replicate getting
2:15:48or migrating data to SQL hosted on Azure
2:15:50is for developer or test environments
2:15:52before deploying to the production
2:15:53environment it is pertinent that the
2:15:56data is tested against developer and
2:15:57test environments so Azure SQL database
2:16:01can act as a target for such
2:16:02environments the life production
2:16:04environment can be replicated to the
2:16:05developer or test environment using a
2:16:07database copy so the second is business
2:16:09continuity one of the most important use
2:16:11cases for SQL on azour is using it as a
2:16:13Dr Target to maintain business
2:16:15continuity azour SQL databases can
2:16:17provide an SLA of up to
2:16:1999.99% by maintaining several copies of
2:16:21the data this provides business
2:16:23continuity as it allows you to restore
2:16:25GE redundant copies of the data or use
2:16:29active Geo redundant copies as failover
2:16:31points in use of outages at data centers
2:16:34or in regions besides SQL databases you
2:16:38can also use availability groups to
2:16:40fulfill business continuity demands not
2:16:41only can you use availability groups in
2:16:43Azure SQL virtual machines but also use
2:16:46Azure SQL virtual machine instances as a
2:16:47target for high availability and
2:16:49disaster recovery and the third one is
2:16:52scaling out readon workloads apart from
2:16:54providing PC or Dr capabilities active
2:16:57Geo replication can also be used to
2:16:59offload readon workload such as
2:17:01reporting jobs to secondary copies you
2:17:03can also extend on premises SQL Server
2:17:05instance using readable always on
2:17:08replicas and the fourth one is backup
2:17:10and G so Azure SQL database are backed
2:17:12up automatically on a regular basis and
2:17:15there are no storage cost for to 200% of
2:17:17the maximum provision database storage
2:17:19you can restore backups to any point in
2:17:22time going back to a pretended period
2:17:24which is determined by the Azure SQL
2:17:26Service Tire in use on premises SQL
2:17:29Server databases and transaction locks
2:17:31can also be bagged up directly to Azure
2:17:33using the backup to URL feature and
2:17:35stored in Azure storage so Azure SQL
2:17:38databases can also be stored on local
2:17:40storage by exporting them to backpack
2:17:43files means backup and the files so
2:17:46fifth one is Advanced analytics so so
2:17:47another important reason for hosting SQL
2:17:49in azour is to make use of azure's
2:17:51advanced gentics platforms such as azour
2:17:53storage blob and Azure data leak store a
2:17:56common scenario with Advanced analytics
2:17:57is when users reference data from
2:17:59various data sources use Azure data L
2:18:02store as the staging area or perform
2:18:05transformation activities using hi or
2:18:07spark and finally load the data into
2:18:08Azure data warehouse for bi and
2:18:10Reporting bi means business intelligence
2:18:13now let's see the use cases for Cosmos
2:18:15database first of all they used in iot
2:18:17and telematics so iot use cases commonly
2:18:19share some patterns in how they ingest
2:18:21process and store data first these
2:18:23systems need to ingest burst of data
2:18:25from device sensors of various locals
2:18:28next these systems process and analyze
2:18:30streaming data to derive real time
2:18:33insights the data is then archive tool
2:18:35Co storage for batch analytics Microsoft
2:18:37Azure offers Rich services that can be
2:18:40applied for iot use cases including
2:18:42Azure Cosmos database Azure event hubs
2:18:45Azure stream analytics Azure
2:18:46notification hub Azure machine learning
2:18:48Azure HD insight and powerbi burst of
2:18:51data can be ingested by Azure event hubs
2:18:53as it offers High throughput data
2:18:55ingestion with low latency data ingested
2:18:57that needs to be processed for realtime
2:18:59Insight can be funneled to aure stream
2:19:01analytics for realtime analytics data
2:19:04can be loaded into Azure Cosmos database
2:19:06for an ad hoc query once the data is
2:19:08loaded into azour Cosmos database the
2:19:10data is ready to be queried in addition
2:19:13new data and changes to existing data
2:19:15can be read on changed feed
2:19:18so change speed is a persistent append
2:19:20Only log that stores changes to Cosmos
2:19:22containers in sequential order then all
2:19:24data or just changes to data in Azure
2:19:26Cosmos database can be used as reference
2:19:29data as part of a realtime analytics in
2:19:31addition data can further be refined and
2:19:34processed by connecting Azure Cosmos
2:19:36database data to HD insight for pig
2:19:38hiive or map reduce jobs refined data is
2:19:41then for a sample of iot solution using
2:19:45Azure Cosmos database event hubs and
2:19:47storm see the HD Insight storm examples
2:19:49repository on GitHub okay then we have
2:19:52is retail and marketing so Azor Cosmo
2:19:55database is used extensively in
2:19:57Microsoft's own e-commerce platforms
2:19:58that runs the Windows store and Xbox
2:20:00Live it is also used in the retail
2:20:02industry for storing catalog data and
2:20:05for event sourcing in order to process
2:20:07pipelines so catalog data storage
2:20:09scenarios involve storage and query a
2:20:11set of attributes for entities such as
2:20:13people places and products some examples
2:20:16of catalog data are user accounts
2:20:18product cataloges iot devices Registries
2:20:21and build of material systems attributes
2:20:24for this data may vary and can change
2:20:26over time to fit application
2:20:28requirements consider an example of a
2:20:30product catalog of an automative part
2:20:32supplier every part may have its own
2:20:34attributes in addition to the common
2:20:36attributes that all parts share
2:20:38furthermore attributes for a specific
2:20:39part can change the following year when
2:20:42a new model is released Azure Cosmos
2:20:44database supports flexible schemas and H
2:20:46High iCal data and thus it is well
2:20:49suited for storing product catalog data
2:20:51Azure Cosmos database is often used for
2:20:53event sourcing to power event driven
2:20:55architectures using its change feed
2:20:57functionality the change feed provides
2:20:59Downstream microservices the ability to
2:21:01reliability and incrementally read
2:21:03inserts and updates made to an Azure
2:21:06Cosmos database this functionality can
2:21:08be leveraged to provide persistent event
2:21:10store as a message broker for State
2:21:13changing events and drive order
2:21:15processing workflow between any micros
2:21:17servic
2:21:18in addition data store in Azure Cosmos
2:21:20database can be integrated with HD
2:21:22insight for big data analytics via
2:21:24Apache spark jobs so the third one is
2:21:27gaming the database tire is a crucial
2:21:29component of gaming applications modern
2:21:31gaming app perform graphical processing
2:21:33on mobile or console clients but rely on
2:21:35the cloud to deliver customized and
2:21:37personalized content like in-game stats
2:21:40social media integration and high school
2:21:42leaderboards games often require single
2:21:44millisecond latencies for reads and WR
2:21:47to provide an engaging in-game
2:21:49experience a game database needs to be
2:21:51fast and be able to handle massive Spice
2:21:53in request rates during new game
2:21:55launches and feature
2:21:57updates so Azure Cosmo database is used
2:21:59by games like The Walking Dead No Man's
2:22:01Land by next games and hello five
2:22:03guardians so Azure Cosmos database
2:22:06provides the number of benefits to game
2:22:07developers like Azure Cosmos DB allows
2:22:10performance to be scaled up or down
2:22:12elastically this allows games to handle
2:22:14updating profiles and stats from dozens
2:22:17millions of simultaneous Gamers by
2:22:19making a single API call then Azure
2:22:21Cosmos DV supports millisecond reads and
2:22:23rights to help avoid any lags during the
2:22:26game play Then Azo Cosmos database
2:22:29automatic indexing allows for filtering
2:22:31against multiple different properties in
2:22:33real time for example locating players
2:22:35by the internal player IDs or their game
2:22:37center Facebook Google IDs or quering
2:22:39based on player membership in a guild
2:22:41this is possible without building
2:22:43complex indexing or shedding
2:22:44infrastructure social features including
2:22:47game that messages player Guild
2:22:48membership challenges completed high
2:22:50score leaderboards and social graphs are
2:22:52easier to implement with a flexible
2:22:54schema so Azor Cosmos database as a
2:22:57managed platform as a service require
2:22:58minimal setup and management work to
2:23:00allow for Rapid iteration and reduce
2:23:02time to market the last one is web and
2:23:05mobile applications so Azure cosos
2:23:06database is commonly used within mobile
2:23:08and web applications and is well suited
2:23:11for modeling social interactions
2:23:13iterating with third party services and
2:23:14for building Rich personal experiences
2:23:17the costos database sdks can be used to
2:23:19build Rich IOS and Android applications
2:23:22using the popular zamarin framework so
2:23:24under web applications first we have the
2:23:27social applications and then we have
2:23:28personalizations so in Social
2:23:30applications a common use for Azure
2:23:32Cosmo database is store and query user
2:23:35generated content means ugc so for web
2:23:38mobile and social media applications
2:23:39some examples of a user generated
2:23:41content are chat sessions tweets blogs
2:23:44posts rating and comments often the ug
2:23:47in social media applications is a bland
2:23:49of free form text properties text and
2:23:51relationships that are not bounded by
2:23:53rigid structure content such as stats
2:23:55comments and posts can be stored in
2:23:58Cosmos DB without requiring
2:23:59Transformations or complex object to
2:24:02relational mapping layers data
2:24:03properties can be added or modified
2:24:05easily to match requirements as
2:24:06developers it iterate over the
2:24:08applications code thus promoting rapid
2:24:10development applications that integrate
2:24:12with third party social network must
2:24:14respond to changing schemas from these
2:24:17networks as data is automatically
2:24:18indexed by default in Cosmos database
2:24:20data is ready to be queried at any time
2:24:23hence these applications have the
2:24:24flexibility to retri projections as
2:24:26their respective needs so the second
2:24:29thing in web mobile applications is
2:24:31personalization so now is mod
2:24:33applications comes with complex views
2:24:34and experiences these are typically
2:24:37Dynamic catering to user preferences or
2:24:39moods and branding needs hence
2:24:41applications need to be able to try
2:24:43personalization settings effectively to
2:24:45render UI elements and experience es
2:24:47quickly Json a format supported by the
2:24:50cosmos DB is an effective format to
2:24:52represent UI layout data as it is not
2:24:54only lightweight but also can be easily
2:24:57interpreted by JavaScript Cosmos R
2:24:59offers turnable consistency levels that
2:25:01allows fast reads with low latency
2:25:04rights hence storing UI layout data
2:25:06including personalized settings as Json
2:25:09documents and Cosmos GB is an effective
2:25:11means to get this data across the wire
2:25:14so these were the use cases for Cosmos
2:25:16DB
2:25:17now that you have a theoretical
2:25:18understanding of azure database Services
2:25:20let's now see a simple deployment of a
2:25:22database service on Microsoft Azure the
2:25:25simply type Microsoft Azure on
2:25:27Google what you can do is you can create
2:25:29a free account on Microsoft Azure you
2:25:32get S 12 months of free services and
2:25:33around 40,000 rupees of free credits
2:25:35also for using the services we can just
2:25:38directly open the console from here I
2:25:41just sign
2:25:43in so for deploying a simple database
2:25:46service so so what we going to deploy
2:25:48today we can like deploy Cosmos database
2:25:50like I have explained you what is cosmos
2:25:51database a new SQL database it is so we
2:25:53can just go to console to the portal I
2:25:56can go to the
2:25:59portal you can create a resource from
2:26:02here can search for
2:26:06Cosmos like J Cosmos yes you can see
2:26:09here like free credits I have a free
2:26:11trial account so it is showing that I
2:26:12have 14,500 three GRS so this is the
2:26:16like credit amount you get for in a free
2:26:18trial okay so you can create a Azure
2:26:20Cosmos DB from here which one you want
2:26:22to create like you can create Pro SQL
2:26:24one so Resource Group can give a new one
2:26:27or we have existing we have a Rec Cosmos
2:26:30one resource Cosmos so if you want to
2:26:32choose the existing one or if you want
2:26:34to choose the new one okay so you can
2:26:36choose the existing one from here or if
2:26:38you want to create a new one then create
2:26:39new One Source One
2:26:41Cosmos I hope that works out yeah then
2:26:45give any unique name for this like uh I
2:26:48will give
2:26:50demoore Cosmos 1 2 3 okay it cannot
2:26:55contain uh UND remember these things
2:26:57okay it do cannot contain this so 1 2 3
2:27:004 I will okay it's not available so I
2:27:03will give five also yeah it's available
2:27:05now and like choose your location
2:27:07whichever location you are located in
2:27:09can use nearby location so mine is Asia
2:27:12Pacific Central India so I've chosen
2:27:14this so free trial account is already
2:27:16there I applied for it then you can just
2:27:19review and
2:27:20create before getting deployed it will
2:27:23show you the review for
2:27:25it so remember that on the basis of the
2:27:28location we have selected the creation
2:27:29time will differ okay so you can review
2:27:32it all the information what you have
2:27:33inserted so now you can
2:27:38create so deployment is in progress it
2:27:41will take a few
2:27:43minutes you can see like how the
2:27:45resource has been created deployment is
2:27:47in progress will soon be created you can
2:27:49check details for it from
2:27:51here let's go to portal
2:27:55again like it is already pinned here or
2:27:57you can search from here okay for Cosmos
2:28:00GV so just click here so it is showing
2:28:03that it is getting created this is the
2:28:05one I have created before only this one
2:28:07is creating this in progress let's
2:28:09refresh one again you can see the
2:28:12processing going on
2:28:15here yeah so your deployment is complete
2:28:18showing you can go to Resource from here
2:28:21also you can just refresh it from
2:28:24here so yeah this is how it's been
2:28:27created you can open the source from
2:28:28here and you can go to activity log or
2:28:31data Explorer you can create a database
2:28:33anything or you can see the consistency
2:28:35of it like default consistency and
2:28:37everything so that's how customers
2:28:39database is been deployed so I hope you
2:28:41have understood this
2:28:45deployment
2:28:49let's look into the family of azure SQL
2:28:52so first one in the family is SQL server
2:28:54on Virtual machines so with this you can
2:28:57lift and shift your SQL Server workloads
2:28:59to the cloud to get the combined
2:29:01performance security and analytics of
2:29:03SQL server with flexibility and hybrid
2:29:05connectivity of azure with 100% code
2:29:08compatibility access the latest SQL
2:29:10Server updates and releases including
2:29:12SQL Server 2019 register your virtual
2:29:15machines with SQ infrastructure as a
2:29:17service agent extension for automated
2:29:20virtual machine management at no
2:29:21additional cost SQL server on Azure
2:29:24virtual machines is part of the Azure
2:29:26SQL family which allows you to migrate
2:29:28existing apps or build new apps on the
2:29:31best cloud destination for a mission
2:29:33critical SQL Server workloads so its
2:29:36features are first of all best TCO that
2:29:38is total cost of ownership with Azure
2:29:40hybrid benefit with Azure SQL Server you
2:29:43can save up to 84% compared to Amazon
2:29:45web services migrating SQ server
2:29:47databases with Azure hybrid benefit and
2:29:49get free extended support for SQL Server
2:29:512008 R2 images in Azure infrastructure
2:29:54as a service activate Azure hybrid
2:29:56benefit when you provision SQL server on
2:29:59Azure virtual machines images from the
2:30:01Azure Marketplace second feature is high
2:30:03performance virtual machines for SQL
2:30:05server on Linux and windows so you can
2:30:08take advantage of SQL Server virtual
2:30:09machines with industry leading
2:30:11performance choose from images with
2:30:13Windows Server redhead Enterprise Linux
2:30:15SU Enterprise Linux server or you been
2:30:18to Linux gain collocated integrated
2:30:21support for your SQL workloads with
2:30:22redhead and suc the third feature is
2:30:25built-in security and manageability so
2:30:28you can ease maintenance with automatic
2:30:30security updates and restore your
2:30:32database to a specific point in time
2:30:34with Azure backup help protect your data
2:30:36address and in motion with the database
2:30:38stated as least vulnerable over the last
2:30:419 years in the cloud with the most
2:30:43global national and Industry
2:30:44certifications the second member in the
2:30:46family is azure SQL managed instance so
2:30:50part of the Azure SQL service portfolio
2:30:52Azure SQL managed instance is the
2:30:54intelligent scalable Cloud database
2:30:56service that combines the broadest SQL
2:30:58Server engine compatibility with all the
2:31:01benefits of a fully managed and everen
2:31:03platform as a service with SQL managed
2:31:06instance confidently modernize your
2:31:08existing apps at scale by combining your
2:31:10experience with familiar tools skills
2:31:12and resources and do more with what you
2:31:15already have Azure Arc enabled SQL
2:31:18manage instance is now in preview you
2:31:21can run the service on premise on any
2:31:23infrastructure of your choice with Azure
2:31:25Cloud benefits like elastic scale
2:31:26unified management and a cloud billing
2:31:28module while staying always current some
2:31:31of its features are always operate on
2:31:33the latest version of SQL so SQL manage
2:31:36instance is built on the SQL Server
2:31:38engine it's overgreen meaning it's
2:31:40always up to date with the latest SQL
2:31:42features and functionality never worry
2:31:45about updates upgrades or end of support
2:31:48again second is fully managed and
2:31:50optimized for DBA productivity so boost
2:31:53productivity and operate more
2:31:54efficiently by letting the service
2:31:56perform timec consuming and complex
2:31:58tasks on your behalf features like
2:32:00built-in High availability disaster
2:32:02recovery and automated backups ensure
2:32:04your data is available when you need it
2:32:06while AI power automatic tuning
2:32:08optimizes performance for you SQL manage
2:32:10instance combines all the best of SQL
2:32:13server with the financial and
2:32:14operational benefits of the platform as
2:32:16a service and the third feature is
2:32:18maintain SQL Server application
2:32:20compatibility so accelerate application
2:32:22modernization with the latest SQL Server
2:32:24capabilities in the cloud SQL managed
2:32:27instance provides an entire SQL Server
2:32:29instance within a managed service so you
2:32:32can continue to use familiar tools and
2:32:34SQL Server features like cross database
2:32:36queries and Link servers SQL managed
2:32:38instance maintains the highest
2:32:40compatibility labels so you can move
2:32:42your on premises workloads without
2:32:44worrying about application comp
2:32:46compatibility of performance changes and
2:32:48the third member in the family is azure
2:32:50SQL Edge so Azure SQL Edge is an
2:32:53optimized relational database engine
2:32:55Geared for iot and iot Edge deployments
2:32:58it provides capabilities to create a
2:33:00high performance data storage and
2:33:02processing layer for iot applications
2:33:04and solutions Azure SQL Edge provides
2:33:07capabilities to stream process and
2:33:09analyze relational and non-relational
2:33:11data such as Json graph and time series
2:33:13data which makes it the right choice for
2:33:15a VAR of modern iot applications Azure
2:33:18SQL Edge is built on the latest version
2:33:21of the SQL Server database engine which
2:33:24provides industry-leading performance
2:33:26security and quering processing
2:33:28capabilities since Azure SQL Edge is
2:33:30built on the same engine as SQL server
2:33:32and Azure SQL it provides the same
2:33:34transact SQL programming surface area
2:33:37that makes development of applications
2:33:38or Solutions easier and faster and makes
2:33:41application probability between iot Edge
2:33:43devices data centers and Cloud straight
2:33:45forward
2:33:46so there are two different deployment
2:33:48models in Azure SQL H so the first one
2:33:52is connected deployment through Azure
2:33:54iot Edge azour SQL Edge is available on
2:33:57the Azure Marketplace and can be
2:33:59deployed as a module for Azure iot Edge
2:34:02second is disconnected deployment so
2:34:04Azure SQL Edge container images can be
2:34:06pulled from Docker Hub and deployed
2:34:08either as a standalone doer container or
2:34:10a kubernetes cluster so some of the
2:34:12features of azour SQL EDR built in data
2:34:14streaming and time SE within database
2:34:17machine learning and graph features for
2:34:18low latency analytics then data
2:34:20processing at the edge for online
2:34:22offline and hybrid environments to
2:34:24overcome latency and bandwidth
2:34:26constraints next is deploy an update
2:34:28from the azuro portal or enterprise
2:34:30portal for consistent security and trunk
2:34:33key management last one is simplified
2:34:35pricing with no upfront cost and
2:34:37subscription offers as low as us $60 per
2:34:41year per device so the fourth and the
2:34:44major family member of azure SQL family
2:34:46is azure SQL database which we're going
2:34:49to briefly understand further first
2:34:51let's understand why one need an Azure
2:34:53SQL database extensively I will tell you
2:34:56the top five benefits that companies are
2:34:58realizing with SQL database so the first
2:35:01one is scalability and Beyond flexible
2:35:03service plans for SQL database meet the
2:35:06need for both big and small business
2:35:08users SQL is no longer Way Out Of Reach
2:35:12for smaller operations because the
2:35:13pricing structure allows users to pay as
2:35:16little as
2:35:17$4.99 per database per month with a
2:35:20maximum storage set at 150 GB per
2:35:23database that's a lot of space for very
2:35:26small cost second is high speed and
2:35:29minimal downtime so high availability
2:35:31architecture mean High speeed
2:35:33connectivity and data retrival as well
2:35:35as low downtime at your organization
2:35:38there's nothing worse than stopping
2:35:40business because your technology can't
2:35:42keep up and that is no longer a problem
2:35:44with SQL database secondly companies can
2:35:47add application instances as needed
2:35:49through sheding for example shedding is
2:35:52a type of database partitioning that
2:35:54separates very large databases into
2:35:56smaller faster more easily managed Parts
2:35:59called Data shards not only can you spin
2:36:02nodes up and down on demand you can
2:36:04leverage a federation infrastructure to
2:36:06scale more easily without affecting
2:36:08other areas of the server SQL azur
2:36:11Federation data migration visard can
2:36:14further automate this process which
2:36:16impacts your organization and the
2:36:17employees much less lastly there are
2:36:20multiple levels of implementation that
2:36:22you can benefit from if you just need a
2:36:24website and a database you can hitch a
2:36:27SQL Azure instance to an Azure website
2:36:29and you are done if you need a
2:36:31full-blown virtual machine now or even
2:36:34down the road you can get that as well
2:36:36you can even use a locally deployed
2:36:38instance of SQL server in the virtual
2:36:40machine instead of SQL Azure these
2:36:43implementation options help make your
2:36:45comp company more adaptable to the
2:36:48inevitable changes it under goes on a
2:36:50regular basis with SQL Azure you are not
2:36:53stuck you are a foundation that
2:36:55encourages growth while working with it
2:36:57third is improved usability so SQL
2:37:00developers are familiar with all things
2:37:03SQL and SQL database can be updated with
2:37:05SQL CMD or the SQL Server management
2:37:08Studio better yet there is no coding
2:37:10required using a standard SQL it's much
2:37:13easier to manage database systems
2:37:16without having to write or update a huge
2:37:18amount of code fourth time is on your
2:37:21side with no administrative duties on
2:37:23your physical location employees can
2:37:25take time for strategic work to advance
2:37:28grow all around business success when
2:37:31your database is hosted in the cloud you
2:37:33don't have to deal with setting up SQL
2:37:35Server appropriating databases and
2:37:37dealing with physical machine
2:37:39maintenance and upkeep all of this
2:37:41results in better alignment of your
2:37:43organization and ultimately more time on
2:37:45your site Fifth and the last one is easy
2:37:48to use migration tools ramp up time with
2:37:51SQL database is now easier than ever and
2:37:53free SQL data synchronization allows you
2:37:56to either synchronize your SQL Server
2:37:59stored data or migrate that data without
2:38:01having to worry about the cost to
2:38:03migrate by syncing gigabyte size tables
2:38:06now that you know why we need aour SQL
2:38:08database let's reply understand what
2:38:10actually it is the basic fundamental
2:38:12building block that is available in a is
2:38:15the SQL data datase Azure SQL database
2:38:17is fully managed platform as a service
2:38:19database engine that handles most of the
2:38:21database management functions such as
2:38:23upgrading patching backups and
2:38:25monitoring without user involvement
2:38:28Azure SQL database is always running on
2:38:30the latest stable version of the SQL
2:38:33Server database engine and ped OS with
2:38:3599.99% availability platform as a
2:38:38service capabilities that are built into
2:38:40Azure SQL database enable you to focus
2:38:42on the domain specific database
2:38:44Administration and optim optimization
2:38:46activities that are critical for your
2:38:47business with Azure SQL database you can
2:38:50create a highly available and high
2:38:52performance data storage layer for the
2:38:54applications and Solutions in azour SQL
2:38:57database can be the right choice for a
2:38:59variety of modern Cloud applications
2:39:01because it enables you to process both
2:39:03relational data and non-relational
2:39:04structures such as graphs Json spatial
2:39:07and XML Azure SQL database is based on
2:39:10the latest stable version of the
2:39:12Microsoft SQL Server database engine you
2:39:15can use Advanced query processing
2:39:16features such as high performance
2:39:18inmemory Technologies and intelligent
2:39:20query processing in fact the newest
2:39:23capabilities of SQL Server are released
2:39:25first to SQL database and then to SQL
2:39:28Server itself you get the newest SQL
2:39:30Server capabilities with no overhead for
2:39:33patching or upgrading tested across
2:39:35millions of databases SQL database
2:39:38enables you to easily Define and scale
2:39:40performance within two different
2:39:42purchasing models and that we will
2:39:44discuss further so Microsoft handles all
2:39:46patching and updating of the SQL and
2:39:48operating system code you don't have to
2:39:50manage the underlying infrastructure so
2:39:52now let's understand the deployment
2:39:54models Azure SQL databas provides the
2:39:56following deployment options for a
2:39:58database so the first one is managed
2:40:00instance this is primarily targeted
2:40:02towards on premises customers in case if
2:40:05we already have a SQL Server instance to
2:40:08on premises Data Center and you want to
2:40:10migrate that into Azure with minimum
2:40:12changes to our application and the
2:40:13maximum compatibility then new will go
2:40:16to the manage instance second is single
2:40:19database so single database represents a
2:40:21fully managed isolated database you
2:40:24might use this option if you have modern
2:40:26Cloud applications and microservices
2:40:28that need single reliable data source a
2:40:30single database is similar to a
2:40:32contained database in the SQL Server
2:40:34database engine last one is elastic pool
2:40:37so elastic pool is a collection of
2:40:39single databases with a shred set of
2:40:41resources such as CPU or memory single
2:40:43databases can be moved into and out of
2:40:45an elastic pool now let's understand the
2:40:48purchasing models so SQL database offers
2:40:50the following purchasing models you can
2:40:52see here so the first one is vcore based
2:40:54purchasing model which is new and it
2:40:56offers a totally different approach to
2:40:58sizing your database it is easier to
2:41:00translate local workloads to a Vore
2:41:02based model because the components are
2:41:04what we are used to the vcore based
2:41:07model lets you choose the number of vour
2:41:09the amount of memory and the amount and
2:41:12speed of storage the Vore based
2:41:14purchasing model also allows you to use
2:41:17Azure hybrid benefit for SQL Server to
2:41:19gain cost savings then the next one is
2:41:22DTU based purchasing model so the D2
2:41:24based purchasing model offers a bland of
2:41:26compute memory and input output
2:41:28resources in three service tires to
2:41:31support light to heavy database
2:41:32workloads compute sizes within each TI
2:41:35provide a different mix of these
2:41:37resources to which you can add
2:41:39additional storage resources as you can
2:41:42see from the following diagram the DTU
2:41:44model offers a pre-configured and
2:41:46predefined amount of compute resources
2:41:49vcore is all about independent
2:41:51scalability where you can look into a
2:41:54specific area such as the CPU core count
2:41:56and memory resources something that you
2:41:59cannot control at the same granular
2:42:00level when using the DTU based model so
2:42:03the third one is the serverless model
2:42:05which automatically scales compute based
2:42:08on workload demand and builds for the
2:42:10amount of compute used per second the
2:42:12serverless compute Tire also
2:42:14automatically pauses data bases during
2:42:16inactive periods when only storage is
2:42:18built and automatically resumes
2:42:20databases when activity returns now
2:42:23let's see the service tries for Azure
2:42:25SQL database so the first one is general
2:42:27purpose or standard model it is based on
2:42:30a separation of computing and storage
2:42:32service this architecture model depends
2:42:35on the high availability and reliability
2:42:37of azure premium storage that
2:42:40transparently copies database files and
2:42:42guarantees for zero data loss if underly
2:42:45infrastructure failure happens second is
2:42:48business critical of premium service Tri
2:42:50model it is based on a cluster of
2:42:53database engine processes both the SQL
2:42:55database engine process and underlying
2:42:58MDF or ldf files are placed on the same
2:43:00node with locally attached SSD storage
2:43:03providing low latency to a workload High
2:43:06availability is implemented using
2:43:08technology similar to SQL Server always
2:43:11on availability groups the third one is
2:43:14hyperscale Service Tire model it is the
2:43:17newest Service Tire in the vord based
2:43:18purchasing model this tire is a highly
2:43:21scalable storage and compute Performance
2:43:24Tire that leverages the Azure
2:43:25architecture to scale out the storage
2:43:27and compute resources for an Azure SQL
2:43:30database beyond the limits available for
2:43:32the general purpose and business
2:43:34critical service tires now let's
2:43:36understand the self-contained services
2:43:37in azour SQL database so Azure SQL
2:43:40database is a database as a platform
2:43:41service designed for applications that
2:43:44will use database as selfcontain service
2:43:46databases can be grouped together to
2:43:49simplify management options or share the
2:43:51resources there are different options
2:43:53that can be used to bound databases in
2:43:56the group so the first one is databases
2:43:58in logical server so logical server is a
2:44:01default container for Azure SQL database
2:44:04logical server enables you to perform
2:44:05administrative tasks across multiple
2:44:07databases including a specifying Regions
2:44:10login information firewall rules
2:44:12auditing thread detection and failover
2:44:14groups all databases within the server
2:44:17are self-contained with independent
2:44:19service tries that can be specified per
2:44:22each database each database can be
2:44:24independently scaled up or down by
2:44:26changing performance tries on the
2:44:28database which will not affect other
2:44:30databases databases cannot share
2:44:33resources and each database has
2:44:34guaranteed and predictable performance
2:44:36defined by its own service style some
2:44:39server level specific features such as
2:44:42cross database quering linked servers
2:44:44SQL agent service broker or CLR are not
2:44:47supported in Azure SQL database placed
2:44:49in logical servers second one is
2:44:51databases in elastic pool so databases
2:44:54need to share resources can be stored in
2:44:57elastic pools instead of The Logical
2:44:59server all databases within the elastic
2:45:02pool share the same resources associated
2:45:04with the elastic pool label currently
2:45:06there are three service ties in the
2:45:07elastic pools basic standard and premium
2:45:10databases within the elastic pools
2:45:12cannot have different service tires
2:45:14because they share resources that are
2:45:16assigned to the entire pool resources
2:45:18usage in one database might affect
2:45:20others however you can specify Reserve
2:45:22performance for the database in the pool
2:45:24that will guarantee a minimal amount of
2:45:26resources that the database can have
2:45:29this model is a good choice for
2:45:30databases that have performance Peaks or
2:45:32heavy usage in different time periods
2:45:34because the amount of resources
2:45:36associated with the pool can be assigned
2:45:38to the databases that need them while
2:45:40the others are inactive elastic pool
2:45:43model is designed for resource sharing
2:45:45and it is still does not support server
2:45:48level features such as SQL agents
2:45:50service broker Etc these are other
2:45:53mechanisms that can be used as a
2:45:55replacement of these features such as
2:45:57elastic jobs and elastic queries instead
2:46:00of some server label features so now
2:46:03let's look at some of the key features
2:46:04of azure SQL database the first one is
2:46:08extensive monitoring and alerting
2:46:09capabilities so Azure SQL database
2:46:12provides Advanced monitoring and trouble
2:46:14shooting features that help you get
2:46:16deeper insights into workload
2:46:18characteristics these features and tools
2:46:20include the buil-in monitoring
2:46:22capabilities provided by the latest
2:46:24version of the SQL Server database
2:46:26engine they enable you to find realtime
2:46:28performance insights also platform as a
2:46:31service monitoring capabilities provided
2:46:33by AO that enable you to Monitor and
2:46:35troubleshoot a large number of database
2:46:37instances query store a buil-in SQL
2:46:40Server monitoring feature records the
2:46:42performance of your queries in real time
2:46:45and enables you to identify the
2:46:47potential performance issues and the top
2:46:49resource consumers automatic tuning and
2:46:52recommendation provides advice regarding
2:46:54the queries with the regress performance
2:46:56and missing or duplicated indexes
2:46:59automatic tning in SQL database enables
2:47:01you to either manually apply the script
2:47:04that can fix the shoes or let SQL
2:47:06database apply the fix SQL database can
2:47:09also test and verify that the fix
2:47:11provides some benefit and retain or
2:47:13reward the change depending on the
2:47:15outcome in addition to query store and
2:47:17automatic twinning capabilities you can
2:47:20use extended DMVs and XC event to
2:47:23monitor the workload performance Azure
2:47:26provides buil-in performance monitoring
2:47:27and alerting tools combined with
2:47:29performance ratings that enable you to
2:47:31monitor the status of thousands of
2:47:33databases using these tools you can
2:47:35quickly asset the impact of scaling up
2:47:37or down based on your current or
2:47:39projected performance needs additionally
2:47:42SQL database can emit metrics and resour
2:47:44Source logs for easier monitoring you
2:47:47can configure SQL database to store
2:47:49resource usage workers and sessions and
2:47:52connectivity into one of these Azure
2:47:54resources so these resources are first
2:47:57one is azure storage so for achieving
2:48:00vast amounts of telemetry for a small
2:48:02price so the second one is azure event
2:48:05hubs for integrating SQL database
2:48:08Telemetry with a custom monitoring
2:48:09solution for hot pipelines and the third
2:48:12one is azure monitor logs for a built in
2:48:15monitoring solution with reporting
2:48:16alerting and mitigating capabilities so
2:48:19the second feature is availability
2:48:21capabilities so aure SQL database
2:48:23enables your business to continue
2:48:25operating during disruptions in a
2:48:27traditional SQL Server environment you
2:48:29generally have at least two machines
2:48:31locally set up these machines have
2:48:33synchronously maintained copies of the
2:48:34data to protect against a failure of a
2:48:37single machine or component this
2:48:39environment provides High availability
2:48:41but it doesn't protect against a natural
2:48:43disaster destroying your your data
2:48:45center Disaster Recovery assumes that a
2:48:47catastrophic event is geographically
2:48:49localized enough to have another machine
2:48:51or set of machines with a copy of your
2:48:53data for far away in SQL Server you can
2:48:56use always on availability groups
2:48:59running in asynchronous mode to get the
2:49:01capability people often don't want to
2:49:04wait for replication to happen that far
2:49:06away from committing a transaction so
2:49:09there's potential for data loss when you
2:49:10do unplanned failovers so databases in
2:49:13the premium and business this critical
2:49:15service tires already do something
2:49:17similar to the synchronization of an
2:49:19availability group databases in lower
2:49:21service tires provideed tendency through
2:49:23storage by using a different but
2:49:25equivalent mechanism like buil-in logic
2:49:28helps protect against a single machine
2:49:29failure the active Geo replication
2:49:32feature gives you the ability to protect
2:49:34against disaster where a whole region is
2:49:36destroyed Azure availability zones tries
2:49:39to protect against the outage of a
2:49:41single Data Center building within a
2:49:43single region it helps helps you protect
2:49:45against the loss of power or network to
2:49:47a building in SQL database you place the
2:49:50different replicas in different
2:49:51availability zones in fact the service
2:49:54label management of azour powered by a
2:49:56Global Network of Microsoft manag data
2:49:58centers helps keep your app running 24/7
2:50:01the Azure platform fully manages every
2:50:03database and it guarantees no data loss
2:50:06and a high percentage of data
2:50:08availability Azure automatically handles
2:50:10patching backups replication failure
2:50:13detection under potential Hardware
2:50:15software or network failures deploying
2:50:18bug fixes failovers database upgrades
2:50:21and other maintenance tasks standard
2:50:23availability is achieved by a separation
2:50:25of compute and storage layers premium
2:50:27availability is achieved by integrating
2:50:29compute and storage on a single node for
2:50:32performance and then implementing
2:50:33technology similar to always on
2:50:35availability groups in addition SQL
2:50:38database provides built-in business
2:50:40continuity and Global scalability
2:50:41features these include automatic backups
2:50:45so SQL database automatically performs
2:50:47full differential and transaction log
2:50:49backups of database to enable you to
2:50:52restore to any point in time for single
2:50:54databases and pool databases you can
2:50:57configure SQL database to store full
2:50:59database backups to Azure storage for
2:51:01long-term backup retention for managed
2:51:03instances you can also perform copy only
2:51:06backups for long-term backup retention
2:51:08second is point in time restor so all
2:51:12SQL database deployment options Support
2:51:14Recovery to any point in time within the
2:51:16automatic backup retention period for
2:51:18any database third one is active Geo
2:51:21replication the single database and pool
2:51:23databases option allow you to configure
2:51:25up to four readable secondary databases
2:51:28in either the same or globally
2:51:29distributed AO data centers for example
2:51:32if you have a service as a platform
2:51:34application with a catalog database that
2:51:36has a high volume of concurrent read
2:51:38only transactions use active GE
2:51:41application to enable global read scale
2:51:43this removes bottleneck on the primary
2:51:45data due to read workloads for managed
2:51:48instances use autof fail groups fourth
2:51:51is autof fail groups all SQL database
2:51:53deployment options allow you to use
2:51:54failover groups to enable High
2:51:56availability and load balancing at
2:51:58global scale this includes transparent
2:52:00GE application and failover of large
2:52:02sets of databases elastic pools and
2:52:04managed instances failover groups enable
2:52:06the creation of globally distributed
2:52:08Service as a platform applications with
2:52:10minimal Administration overhead this
2:52:12leaves all the complex monitoring
2:52:14routing and failover orchestrations to
2:52:16SQL database Fifth and the last one is
2:52:19Zone indendent databases SQL database
2:52:22allows you to provision premium or
2:52:23business critical databases or elastic
2:52:25pools across multiple availability zones
2:52:28because these databases and elastic
2:52:30pools have multiple rendent replicas for
2:52:32high availability placing these replicas
2:52:34into multiple availability zones
2:52:36provides higher resilience this includes
2:52:38the ability to recover automatically
2:52:41from the data center scale features
2:52:42without data loss so the next feature is
2:52:45built-in intelligence with SQL database
2:52:48you get buil-in intelligence that helps
2:52:50you dramatically reduce the costs of
2:52:52running and managing databases and that
2:52:54maximizes both performance and security
2:52:57of your application running millions of
2:52:59customers workloads around the clock SQL
2:53:01database collects and processes a
2:53:03massive amount of telemetry data while
2:53:06also fully respecting customer privacy
2:53:08various algorithms continuously evaluate
2:53:10the Telemetry data so that service can
2:53:12learn and adapt with your applic ation
2:53:15so in its process of work first step is
2:53:17automatic performance monitoring and
2:53:19tuning so SQL database provides detailed
2:53:21insight into the queries that you need
2:53:23to monitor SQL database learns about
2:53:26your database patterns and enables you
2:53:27to adapt your database schema to your
2:53:30workload SQL database provides
2:53:32Performance Tuning recommendations where
2:53:34you can review tuning actions and apply
2:53:36them however constantly monitoring a
2:53:38database is hard and tedious task
2:53:40especially when you are dealing with
2:53:42many databases intelligent insights does
2:53:45this job for you by automatically
2:53:46monitoring SQL database performance at
2:53:48scale it informs you of performance
2:53:51degradation issues it identifies the
2:53:53root cause of each issue and it provides
2:53:55performance Improvement recommendations
2:53:57when possible managing huge number of
2:53:59databases might be impossible to do
2:54:01efficiently even with all available
2:54:03tools and reports that SQL database and
2:54:05Azure provide instead of monitoring and
2:54:08tuning your database manually you might
2:54:10consider delegating some of the
2:54:12monitoring and tuning actions to SQL
2:54:13database by using automatic tuning SQL
2:54:16database automatically applies
2:54:18recommendation tests and verifies each
2:54:20of its tuning actions to ensure the
2:54:22performance keeps improving this way SQL
2:54:24database automatically adapts to your
2:54:26workload in a controlled and Safe Way
2:54:28automatic tning means that the
2:54:30performance of a database is carefully
2:54:32monitored and compared before and after
2:54:34every Twining action if the performance
2:54:37doesn't improve the Twining action is
2:54:39diverted many of our partners that run
2:54:41Service as a platform multitined apps on
2:54:43top of SQL database are relying on
2:54:46automatic Performance Tuning to make
2:54:48sure their applications always have
2:54:50stable and predictable performance for
2:54:52them this feature tremendously reduces
2:54:55the risk of having a performance
2:54:56incident in the middle of the night in
2:54:59addition because part of their customer
2:55:01base also uses SQL Server they are using
2:55:03the same indexing recommendations
2:55:05provided by SQL database to help their
2:55:07SQL Server customers two automatic tning
2:55:10expects are available in SQL database so
2:55:12first one is automatic index management
2:55:15which identifies indexes that should be
2:55:16added in your database and indexes that
2:55:18should be removed second one is
2:55:20automatic plan correction which
2:55:21identifies problematic plans and fixes
2:55:23SQL plan performance problems so the
2:55:25next step in the process of work is
2:55:27adaptive query processing you can use
2:55:29adaptive query processing including
2:55:31interl execution for multi statement
2:55:33table valued functions batch mode memory
2:55:35Grant feedback and batch mode adictive
2:55:37joints each of these adaptive query
2:55:39processing features apply similar learn
2:55:41and adapt techniques helping further
2:55:43address performance issues related to
2:55:45historically inable optimization
2:55:48problems so the next feature is Advanced
2:55:50security and compliance SQL database
2:55:53provides a range of built-in security
2:55:55and compliance features to help your
2:55:56application meet various security and
2:55:58compliance requirement note this that
2:56:00Microsoft has a certified Azure SQL
2:56:02database against number of compliance
2:56:03standards for more information see the
2:56:05Microsoft Azure trust Center where you
2:56:07can find the most current list of SQL
2:56:09database compliance certifications so
2:56:11built-in security and compliance
2:56:13features so the there are certain
2:56:14built-in security and compliance
2:56:16features so the first one is Advanced
2:56:18threat protection Azure Defender for SQL
2:56:20is a unified package for advanced SQL
2:56:23security capabilities it includes
2:56:25functionality for managing your database
2:56:27vulnerabilities and detecting anomalous
2:56:30activities that might indicate a threat
2:56:32to your database it provides a single
2:56:34location for enabling and managing these
2:56:36capabilities so there are two kinds of
2:56:39assessment in advanced protection the
2:56:41first one is vulnerability assessment
2:56:43this service can discover track and help
2:56:45you remediate potential database
2:56:47vulnerabilities it provides visibility
2:56:49into your Security State and includes
2:56:52actionable steps to resolve security
2:56:54issues and enhance your database
2:56:56fortifications second is threat
2:56:57protection so this feature detect anous
2:57:00activities that indicate unusual and
2:57:02potentially harmful attempts to access
2:57:04or exploit your database it continuously
2:57:06monitors your database for suspicious
2:57:08activities and provides immediate
2:57:09security alerts on potential
2:57:11vulnerabilities SQL injection attacks
2:57:13and enous database access patterns
2:57:16threat detection alert provide details
2:57:18of the suspicious activity and recommend
2:57:21action on how to investigate and
2:57:22mitigate the threat so under this the
2:57:25second sub feature is auditing for
2:57:26compliance and security so auditing
2:57:28tracks database events and writes them
2:57:31to an audit log in your Azure storage
2:57:33account auditing can help you maintain
2:57:35Regulatory Compliance understand
2:57:36database activity and gain insight into
2:57:39discrepancies and anomalies that might
2:57:41indicate business concerns or suspected
2:57:43security violations the third one is
2:57:45data encryption SQL database helps
2:57:47secure your data by providing encryption
2:57:50for data at rest it uses transparent
2:57:52data encryption for data in use it uses
2:57:55always encrypted fourth feature is data
2:57:57Discovery and classification data
2:57:59Discovery and classification provides
2:58:00capabilities built into Azure SQL
2:58:03database for discovering classifying
2:58:05labeling and protecting the sensitive
2:58:07data in your databases it provides
2:58:09visibility into your database
2:58:11classification State and tracks the
2:58:12access to sensitive data within the
2:58:14database and Beyond its borders and the
2:58:17last one is azure active directory
2:58:19integration and multiactor
2:58:20authentication so SQL database enables
2:58:23you to centrally manage identities of
2:58:25database user and other Microsoft
2:58:27services with azured active directory
2:58:28integration this capability simplifies
2:58:31permission management and enhances
2:58:33security Azure active directory supports
2:58:35multiactor authentication to increase
2:58:37data and application security while
2:58:39supporting a single signin process the
2:58:41last major feature is easy to use tools
2:58:44SQL database makes building and
2:58:46maintaining applications easier and more
2:58:48productive SQL database allows you to
2:58:50focus on what you do best building great
2:58:52applications so you can manage and
2:58:54develop an SQL datab by using tools and
2:58:56skills you already have so the first
2:58:58tool is the Azure portal a web- based
2:59:01application for managing all Azure
2:59:03Services second is azure data Studio a
2:59:05crossplatform database tool that runs on
2:59:07Windows Mac OS and Linux third is SQL
2:59:10Server management Studio a free
2:59:12downloadable client application
2:59:14for managing any SQL infrastructure from
2:59:16SQL Server to SQL database fourth is SQL
2:59:20Server data Tools in Visual Studio a
2:59:22free downloadable client application for
2:59:24developing SQL Server relational
2:59:25databases databases in aure SQL database
2:59:29integration service packages analysis
2:59:31service data models and Reporting
2:59:33Services reports and the last tool is
2:59:35Visual Studio code a free downloadable
2:59:38open-source code editor for Windows Mac
2:59:40OS and Linux it support extensions
2:59:43including the mssql extension for
2:59:45querying Microsoft SQL Server Azure SQL
2:59:48database and Azure SS analytics now
2:59:51let's look at some of the use cases for
2:59:53Azure SQL database so the first one is
2:59:55developer or test environment it's an
2:59:58important use case for replicating or
3:00:00migrating data to SQL hosted on Azure is
3:00:02for developer or test environments
3:00:05before deploying to the production
3:00:06environment it is pertinent that the
3:00:08data is tested against developer or test
3:00:10environments Azure SQL databases can act
3:00:13as a Target for just such environments
3:00:15the life production environment can be
3:00:17replicated to the developer or test
3:00:19environment using a database copy second
3:00:21is business continuity one of the most
3:00:23important use cases for SQL on aure is
3:00:25using it as a Dr Target to maintain
3:00:28business continuity Azure SQL databases
3:00:30can provide an SLA of up to
3:00:3399.99% by maintaining several copies of
3:00:36the data this provides business
3:00:38continuity as it allows you to restore
3:00:40Geor redundant copies of the data or use
3:00:43active Geo redundant copies as failover
3:00:45points in case of outages at data
3:00:48centers or in regions besides Azure SQL
3:00:51databases you can also use availability
3:00:53groups to fulfill business continuity
3:00:55demands not only can you use
3:00:57availability groups in a SQL virtual
3:00:59machines but also use Azure SQL virtual
3:01:02machine instances as a target for high
3:01:04availability and Disaster Recovery third
3:01:06is scaling our readon workloads apart
3:01:10from providing BC or Dr capabilities
3:01:12active gec ation can also be used to
3:01:15offload readon workload such as
3:01:17reporting jobs to secondary copies you
3:01:19can also extend on premises SQL Server
3:01:22instances using readable always on
3:01:24replicas fourth is backup and restore
3:01:27Azure SQL databases are backed up
3:01:29automatically on a regular basis and
3:01:31there are no storage costs for up to
3:01:34200% of the maximum provision database
3:01:36storage you can restore backups to any
3:01:38point in time going back to the
3:01:40retention period which is determined by
3:01:42the Azure SQL service Tri in use on
3:01:45premises SQL Server databases and
3:01:47transaction logs can also be bagged up
3:01:49directly to Azure using the backup to
3:01:51URL feature and stored in Azure storage
3:01:54Azure SQL databases can also be stored
3:01:57on local storage by exporting them to
3:01:59backpack files means BC PSC files and
3:02:02the last use case is Advanced analytics
3:02:04another important reason for hosting SQL
3:02:06in Azure is to make use of azure's
3:02:09advanced analytics platforms such as
3:02:10Azure storage blob and Azure data L
3:02:13store common scenario with Advanced
3:02:15analytics is when users reference data
3:02:17from various data sources use Azure data
3:02:20Lake store as the stacking area perform
3:02:22transformation activities using Hive or
3:02:25spark and finally load the data into
3:02:27Azure data warehouse for business
3:02:29intelligence and Reporting now that you
3:02:31have a theoretical understanding of
3:02:33azure SQL database let's now see a
3:02:35deployment of a SQL database service on
3:02:38Microsoft Azure we will also connect
3:02:40this database with SQL management server
3:02:42as well as with Azure data studio so
3:02:45let's move ahead can just simply go to
3:02:47the Azure
3:02:48portal if you don't have an account on
3:02:51Azure what you can do is you can create
3:02:52a free account like can start free here
3:02:54and you will get a 12 months free
3:02:56services access as well as if you are
3:02:58from India the currency will be around
3:03:00like 14,500 INR means Indian rupees free
3:03:03currency we will get for free use that's
3:03:06the amount you will get so I already
3:03:07have an account so I will just log in
3:03:09from
3:03:10here no I don't want to I will just go
3:03:12to Microsoft here this is the portal so
3:03:16I don't want to buy see these are the
3:03:17credits it is showing you get 14,500
3:03:20credits I have used some of the credits
3:03:22and this much I left so what you can do
3:03:24is you can go to SQL databases right now
3:03:27because is I have a pin here but what
3:03:29you can do is you can go to SQL
3:03:30databases from here like here or you can
3:03:33search it from here SQL database yeah
3:03:36now let's create the database from
3:03:38here so free trial is here we have
3:03:41chosen this then we can choose the
3:03:43resource Group can create one group so
3:03:46let's create with a new one so let's
3:03:48give the name as resource SQL demo okay
3:03:52so a new resource has been created then
3:03:55enter database name can give some name
3:03:57to like demo
3:03:59SQL so yeah that's okay so give the
3:04:02server name also so create a new server
3:04:05a server name so let's give it like
3:04:08remember these names SQL demo server I
3:04:10have given so the specified server name
3:04:12is already use okay we can give like SQL
3:04:15demo server 662 but you have to remember
3:04:19this don't forget this like login
3:04:21information okay so you can give some
3:04:23server admin login so Azure admin will
3:04:26be okay I think yeah now let's move ahe
3:04:29to give a password don't forget these
3:04:31information we will require these
3:04:33information later on in this process and
3:04:36choose your region also so my region is
3:04:38Central India so it will be Asia Pacific
3:04:41Central India just
3:04:44yeah so okay the next process if you
3:04:47want to use SQL elastic pool right now
3:04:49we don't require so we will choose no if
3:04:51you want to then you can choose yes then
3:04:53we can configure the database and
3:04:54storage also so these are the plans like
3:04:56I have explained you basic standard
3:04:58premium the service tires are there
3:05:00similarly purchasing models are there
3:05:01voco purchasing models are there general
3:05:03purpose hypers scale business critical
3:05:04the also have explained you so like
3:05:07whichever you choose on the basis of
3:05:08that it shows the price so right now
3:05:10it's costing 2949 but we don't require
3:05:12that much bigger so you can just you can
3:05:14choose it from here also like not
3:05:16available because it's the free tire
3:05:18account that's why we don't have a
3:05:19premium version so can use the standard
3:05:21one so in this standard one if you
3:05:23decrease the storage then it will also
3:05:25decrease the price not in this manner
3:05:28like from you can go to basic here you
3:05:30get 2 GB and the price decreased you can
3:05:33choose for 1 GB also but price will
3:05:34remain the same so you can go to 2 gb so
3:05:36we will choose the basic one because we
3:05:38don't require that much of configuration
3:05:39so just apply from here just review and
3:05:42create let simple review plan so yeah
3:05:45create so it's getting created the
3:05:48process it takes certain time so that's
3:05:51why taking a Time few minutes it
3:05:55take so you can see the deployment is
3:05:57complete so you can go to Resource SQL
3:05:59demo so we have came to Resource here
3:06:02you see that server is created but you
3:06:04have to refresh once again just a second
3:06:06where is the database let's refresh
3:06:09again I don't know let's get created
3:06:12usually yeah so you can see have to
3:06:14refresh it so it will takes a little
3:06:16time so database is also created now so
3:06:19what we can do from here is first step
3:06:21is you have to create the firewall so
3:06:23you can go to database so you have to
3:06:26set the firewall set firewall servers
3:06:29you can go from here set the fireable so
3:06:31you have to choose the like all these
3:06:33client IP is given from start IP and end
3:06:35IP so you can just you have to add the
3:06:37client IP so it has been added so you
3:06:40have to give it from 0 to 255 complete
3:06:42you have to
3:06:44so
3:06:45save and just go back then you can see
3:06:48different cool features are given on the
3:06:49left side of this dashboard so you go to
3:06:52the query editor this is here you can
3:06:53like connect your database to the
3:06:55browser and you can configure and then
3:06:57there is computer and storage if you
3:06:58have like you can yeah not a problem so
3:07:01computer and storage is given if you
3:07:03want to like change your plan purchasing
3:07:04model and everything then you can change
3:07:06it from here again like all the things
3:07:08are given here you can again change it
3:07:10it's not a problem then there are like
3:07:12connection Str no problem then you can
3:07:14go to connection strings where you can
3:07:15connect your and make connection to the
3:07:17strings like there's ado.net or if you
3:07:20are working on Java then jdbc is there
3:07:22obbc is there PHP go everything is there
3:07:25it's quite cool feature then there are
3:07:26synchronous to other database where you
3:07:28can synchronize to other databases also
3:07:30then also you can like add Azure search
3:07:32also then there are different security
3:07:34features are there advanc security
3:07:36features are there for data also
3:07:37Advanced Data security features are
3:07:39given here now what we can do major
3:07:41thing we have to do is we have to
3:07:42connect it to the management SQL Server
3:07:44so if you don't have a management SQL
3:07:45Server you can download it from here you
3:07:46can just give management SQL server and
3:07:50just go here and download it from here
3:07:52it's given you can just click here and
3:07:54it will get downloaded it will ask for a
3:07:56restart it will restart the computer and
3:07:57all the settings will be saved in your
3:07:59computer then you can start using SQL
3:08:01Server I already have downloaded it so I
3:08:03already have downloaded it so we can
3:08:05just go here have it print here so
3:08:07Microsoft SQL Server management so yeah
3:08:10you have to give your server name so
3:08:12let's see the server name first here it
3:08:15is no from database only so server name
3:08:19is given so you can just copy it from
3:08:20here paste it here you can just select
3:08:23the SQL Server authentication here and
3:08:25give you a login also so remember the
3:08:28login I told you remember the login
3:08:30password so my is your admin password
3:08:34can just connect it from here so yeah we
3:08:37got it here so we can just use the
3:08:39databases can see the demo SQL we can
3:08:41find the our database here here and
3:08:43there different tables and everything is
3:08:44there so we can just start the new query
3:08:47from here can create the table so let's
3:08:50create the table create table with the
3:08:52name Person
3:09:02persons so yeah it will be enough let's
3:09:04now execute it you can see command
3:09:07completed successfully we can now insert
3:09:09values into it so like before that we
3:09:11can view in tables also the table where
3:09:13we created the name of persons expending
3:09:16so yeah you can see persons is created
3:09:18so let's now
3:09:20insert values you can give like give my
3:09:23name so
3:09:2826 from person let's execute now oh
3:09:33sorry it's a little mistake let's
3:09:35execute now so yeah you can see we have
3:09:37got the table contents everything is
3:09:38there now we can also see for Azure data
3:09:41studio also so we will just just go to
3:09:43aor Studio it's like you have a visual
3:09:46studio code also similarly there's aor
3:09:49Studio can start the new connection then
3:09:51we have to give the server name here
3:09:53just like we have the server name there
3:09:55so what was the server name let's copy
3:09:58it from here so yeah just give here and
3:10:01then we will choose the SQL login
3:10:02username Azure admin the password yeah
3:10:06we can just remember the password not a
3:10:08problem then database we have to select
3:10:10so what was the name of the database
3:10:13demo SQL that's why it's showing here
3:10:15yeah so we connected from
3:10:17here yeah we are here like different
3:10:20features are given like new query new
3:10:22notebook just like you have in Visual
3:10:24Studio code in that way V is given and
3:10:26all the other like it has a very good UI
3:10:28also here we go to the database person
3:10:30it is showing just maximize it okay yeah
3:10:34so it's open now just like you can give
3:10:36it here Al all the details it is showing
3:10:38like it's been executed and details are
3:10:40shown here this is how it looks like
3:10:43also one cool thing about AA studio is
3:10:46you can go here and you can see like it
3:10:48actually remembers all the databases
3:10:52like this one it remembers this one I
3:10:53have closed actually this database I
3:10:55have closed I have created it some time
3:10:57back but it still remembers databases if
3:10:59you give the Azure like server ID now so
3:11:02like it will remember all the databases
3:11:04like from where initially we have
3:11:05selected now from there on so this is
3:11:07how it looks like let's go back to Azure
3:11:10one we have understood how to do it in
3:11:12Azure dat studio also and how to connect
3:11:14it with the management SQL Server also
3:11:16so there are certain other features also
3:11:18you can see here for performance
3:11:19overview like how it's working and to
3:11:22track the performance and everything
3:11:24also the auditing and things are given
3:11:26here so just like these are the things
3:11:27but I have mainly told you about the
3:11:29purchasing models and deployment models
3:11:31and how service tires are there and how
3:11:33we can create the database connected
3:11:34with management SQL Server as well as a
3:11:37studio I think it might be a little
3:11:39hectic but I have explained you in a
3:11:41little simpler manner so please try it
3:11:43with your hands-on experience like take
3:11:45your hands-on experience also by trying
3:11:47it on deploying Azure SQL database on
3:11:49this Azure
3:11:50[Music]
3:11:56portal Azure data Lake storage so what
3:12:00is azure data Lake storage Azure data
3:12:03Lake storage is a repository that stores
3:12:06large amount of raw data in its natural
3:12:08format until it is needed for analytics
3:12:11application
3:12:13so why is it named as data link well
3:12:16James Dixon the chief technology officer
3:12:18of pentaho is the person who has
3:12:21generally been credited with the coining
3:12:23of the term data link according to him
3:12:26he described a data M that is a subset
3:12:29of a data warehouse as a keing to a
3:12:32bottle of water which is cleansed
3:12:35packaged and structured for easy
3:12:37consumption while a data lake is more
3:12:41likely a body of water in its natural
3:12:43States data flows from the streams that
3:12:47is the source of the system to the lake
3:12:49users have access to Lake to examine
3:12:53take samples or dive in so a data lake
3:12:57is a centralized repository designed to
3:12:59store process and secure large amount of
3:13:03structured semi-structured and
3:13:05unstructured data it can store data in
3:13:08its native format and process any
3:13:11variety of it ignoring the size
3:13:15limits next is how to create a data Lake
3:13:19storage for this we'll have a practical
3:13:22demo to have a better understanding so
3:13:25the first step to work on any Azure
3:13:28Services is to first sign in so first we
3:13:31need to sign in make sure you do have an
3:13:34Azure account so that you can have the
3:13:36access to different Azure services so
3:13:39let's sign
3:13:41in
3:13:48stay signed in so once you sign in you
3:13:52enter to this dashboard so you can see
3:13:54here this is a dashboard of your account
3:13:57and you can see different kinds of azure
3:13:59services like storage accounts monitors
3:14:02virtual machine Resource Group SQL
3:14:05database SQL manage instances and also
3:14:08you can see your subscriptions and
3:14:10different types of Resource Group you
3:14:11have created
3:14:13and all other stuffs so let's quickly
3:14:16get started first we need to go to
3:14:18create a
3:14:19resource so once you come here you can
3:14:22see popular isure services so like
3:14:26virtual machine cuberty services Cosmos
3:14:28DB and rest other so for us we need to
3:14:32go to storage
3:14:34account once you click here so you enter
3:14:37to create a storage account here we need
3:14:40to create a resource Cod so first of all
3:14:43what do you mean by Resource Group a
3:14:46resource Group is a container that holds
3:14:48related resources for an Azure solution
3:14:51it can include all these resources for
3:14:53the solution or only those resources
3:14:56that you want to manage as a grou so
3:14:59let's create a new Resource Group let's
3:15:01name it
3:15:03as demo data link one okay and once you
3:15:11create a resource Group so you come to a
3:15:13storage account make so a storage
3:15:16account contains all of your assure
3:15:18storage data objects including blobs
3:15:21file shs cues tables and disks so let's
3:15:25name it as marshmallow 1 2 3 so once you
3:15:30have named your storage account we go to
3:15:32the Advan section here we will directly
3:15:36move to data Lake storage generation 2
3:15:40so what do you mean by data leg storage
3:15:43generation 2 Data leg storage generation
3:15:452 is designed to deal with this variety
3:15:48and volume of data at exhibit scale
3:15:51while securely handling hundreds of
3:15:54gigabytes of through output with this
3:15:57you can use data leg storage generation
3:16:00to on the basis of both real time and
3:16:03bat solution so here we will enable the
3:16:07hierarchical name space once we enable
3:16:10it we'll click on riew plus create so it
3:16:14will take few minutes to validate all
3:16:15your details once validation is passed
3:16:19you just click on create so it's now
3:16:22getting ready to deploy so here you can
3:16:25see your marshmallow 1 2 3 is created
3:16:28and now it is getting ready for
3:16:30deployment here you can see deployment
3:16:33is in progress once it gets ready so we
3:16:37will be working on it here you can see
3:16:40your deployment is complete now what you
3:16:43going to do is we'll move on to go to
3:16:47resources now here you can see your
3:16:51details your resource gr details the
3:16:55whole storage Account Details basically
3:16:58and even you can see the properties
3:17:00enabled with it so here we have data L
3:17:04Storage file service Q service table
3:17:06service networking security so after
3:17:11this we will quickly go to
3:17:14containers and we'll create a new
3:17:17container for our storage account so
3:17:21let's click on container and we give it
3:17:24a name to it make sure you give a name
3:17:28it should be in lower case because they
3:17:30don't accept the upper
3:17:31case so let's keep it demo and create so
3:17:38yeah you can see your container has been
3:17:42created let's click on it so here you
3:17:45can see there are no results as we
3:17:47haven't added any stuff so let's
3:17:50upload for this you can either upload it
3:17:55from Azure portal or else you can also
3:17:59go to storage Explorer so let's see how
3:18:03do we do on the Azure portal so once you
3:18:06come here you can just click on select a
3:18:09file let's take any
3:18:12picture let's see we'll upload a picture
3:18:16so here you can see AWS
3:18:193.png let's upload it so here you can
3:18:22see AWS PNG has been uploaded in your
3:18:26container same as we can do on storage
3:18:29Explorer to for that you need to
3:18:32download a storage Explorer in my case I
3:18:35have already downloaded the storage
3:18:37Explorer so let's quickly go into it so
3:18:42here also make sure you are signed in so
3:18:45that you can see all your containers and
3:18:48Resource Group which you have created so
3:18:51here you can see my storage account that
3:18:54isow 123 which we have created recently
3:18:58under this we'll go and find our
3:19:02container yeah so here we go to blob
3:19:06containers and we can see a demo so here
3:19:09you can see the image file which we had
3:19:11uploaded through our a portal now we'll
3:19:15upload a file and a folder both let's
3:19:19see so let's upload a file first so you
3:19:22click on upload upload files and then
3:19:26you select a file let's take any file
3:19:32let's take a picture again so AWS 4
3:19:37let's take this an image so now here you
3:19:40can see your image is transferring from
3:19:44your path to our demo yeah so here your
3:19:49image file is uploaded so as we said we
3:19:52can upload any kind of datas maybe
3:19:55structured or unstructured so let's
3:19:58check out by uploading a folder so here
3:20:02you go with the same process and you you
3:20:05select your folder let's take any
3:20:10folders let's see if we do have any
3:20:12folder okay I'll just take any one of my
3:20:15folders and just click and upload so
3:20:19your folder is also being uploaded over
3:20:22here and once your folder is uploaded
3:20:25now you can see the inside resources
3:20:28into it so here there were different
3:20:31files text image all of them are there
3:20:35you can access it from here itself and
3:20:38you can see other operations as well if
3:20:40you want to download any of of the file
3:20:42or folder you can do it or if you want
3:20:44to open it let's open it any of the uh
3:20:48files see let's see if it is getting
3:20:52open or not okay yes so here you can see
3:20:56this file has been open now if you want
3:21:01to download this file let's download
3:21:06sequence so we'll download it and just
3:21:10put it in downloads let's see if it is
3:21:13downloaded or no apply to apply so yeah
3:21:17it tells your download is completed
3:21:21let's see if you find it or no we go to
3:21:23our file and downloads and here see now
3:21:28I had downloaded this is the file which
3:21:31I had downloaded through our storage
3:21:33Explorer so this is how you manage your
3:21:35data toward your data leg storage so
3:21:38apart from this you can also give
3:21:40permissions to different users users
3:21:42like whichever file if you want to give
3:21:44an access to a particular user then you
3:21:47can also give permission to them that
3:21:49they can either read write and access
3:21:52the whole file or a folder so for that
3:21:55you just need to click on a file or a
3:21:57folder and right click and you just see
3:22:01here manage Access Control list so once
3:22:04you come here so you can here you can
3:22:06see there are different owners super
3:22:09user owner and all so here I can add an
3:22:13owner like for this file who can just
3:22:17get an access to it so you can find out
3:22:21any name if you find out any relatable
3:22:23person or a user then you can give an
3:22:26access to that currently I don't have
3:22:28anyone so I won't be able to show you
3:22:31that but yes this is how you add or give
3:22:34permission to different users you can
3:22:37also do it with the folder or else you
3:22:39can also do it with the whole container
3:22:42as well if you want to share your
3:22:43containers with different people or
3:22:46different users then you can easily
3:22:48share them so here also you just go
3:22:51right click and just come to manage
3:22:53access control and just give them the
3:22:57access click on ADD and find out the
3:23:01person whom you want to give the access
3:23:03to and just after that once you give
3:23:07them the permission here you can see if
3:23:09you want to permit them for only read or
3:23:13only write or if you want to give read
3:23:16and write or all of the three so you can
3:23:19just give them the permissions
3:23:21accordingly and click on okay and then
3:23:24the particular user gets the access to
3:23:26all your files and folders so this is
3:23:29how we create and work with azo data
3:23:32Lake
3:23:33storage now let us see the comparison
3:23:36between Azure blob storage and the data
3:23:39L
3:23:40Storage so here are some of the
3:23:43comparisons between Azure data L Storage
3:23:45and Azure blob storage Azure data L
3:23:48Storage is a technic of planning and
3:23:51control of the time whereas Azure block
3:23:53storage is an object stored with a flat
3:23:57name space Azure data lake is an
3:24:00optimized storage for big data analytics
3:24:02workload whereas AZ your blob storage is
3:24:06basically a general purpose Object Store
3:24:08for a wide variety of storage scenarios
3:24:11which also include big data analytics in
3:24:15AO data Lake storage the apis are over
3:24:18https only whereas in Blob storage the
3:24:21rest API is over the HTTP as well as the
3:24:25https in AO data Lake storage there is
3:24:29no limits on the account but in aure
3:24:32Blob storage there are specific limits
3:24:34for container sizes and the files in the
3:24:37block so these are the major points
3:24:39which differentiate Azure block storage
3:24:42with aure data Lake storage at last we
3:24:46come to the use cases there are many use
3:24:49cases of data Lake storage out of which
3:24:52we'll discuss about four so at first
3:24:56business intelligence on data Lake
3:24:58storage so data Lake storage
3:25:00dramatically improves the speed for ad
3:25:03hoc queries dashboards and remotes you
3:25:06can run existing bi tools on lower cost
3:25:10data Lakes without compromising
3:25:12performance or data quality it also
3:25:15avoids costly delays adding new data
3:25:18sources and the reports at second we see
3:25:22cloud data Lake migration here we can
3:25:26optionally deploy new applications to
3:25:28the cloud using data Lake storage such
3:25:30as S3 or ADLs you can also migrate from
3:25:34older onri data Lake environments that
3:25:37are expensive and difficult to maintain
3:25:39while ensuring agility and
3:25:42flexibility next data science on the
3:25:45data L Storage here you can accelerate
3:25:49data science on data Lake storage with
3:25:52simplified data exploration and feature
3:25:54engineering dramatically it improves
3:25:57performance making data scientists and
3:25:59Engineers more efficient resulting in
3:26:02high quality analytical models at last
3:26:06data architecture modernization so here
3:26:08you can avoid Reliance on propriatary
3:26:12data warehouse infrastructure and the
3:26:14need to manage the cubes extract and
3:26:16agregation tables you can run
3:26:19operational data warehouse queries on
3:26:21low cost data laks offloading the data
3:26:24warehouse at your own
3:26:28[Music]
3:26:32paas so powerbi is a ba tools okay the
3:26:36business intelligent tools wased on the
3:26:39cloud machines cloud services it's
3:26:42maintained by Microsoft corporations
3:26:44it's a Microsoft tools right Ms tools
3:26:46Microsoft tools it's entirely free of
3:26:49cost we don't need to pay any licensing
3:26:50cost or anything for this powerbi it's a
3:26:53free it's available under Microsoft
3:26:56stores so this powerb is mainly for
3:26:58non-technical people suppose if you're
3:27:00business user if you have data analyst
3:27:02if you have a business analyst right so
3:27:05if if you want to perform some of the
3:27:07aggregate level data I want to see the
3:27:09data from summary level or I want to
3:27:11perform some of the data analysis or I
3:27:14want to uh visualize some of the data in
3:27:16the graphical formats or I want to share
3:27:18the data to the different peoples for
3:27:20different servers different place right
3:27:22so for that the power ba is very well
3:27:24suited for non-technical users business
3:27:27user okay so these are the things we can
3:27:29do it on the
3:27:31powerb okay next we'll talk about Azure
3:27:34in EML so how do we use this um Azure uh
3:27:38in in machine learning right so when
3:27:41where we have a lot of services
3:27:43available in Azu like we can integrate
3:27:45it so basically if you want to apply
3:27:47some filter or if you want to apply some
3:27:48selective filter independently we can do
3:27:50it under Azure platform I'll just
3:27:53quickly walk you through the Azure
3:27:56portal okay so this is the my cloud
3:28:00platform like portal uh aure portals
3:28:03right so as I told you we have a
3:28:07different services available so if you
3:28:09go check right so we have a SQL database
3:28:13Cosmo DB for no SQL database right and
3:28:16storage account like as your active
3:28:18directory for this security things right
3:28:20and then
3:28:21monitor cost management so there are
3:28:24different services are available so
3:28:26basically before going to create a
3:28:28service rate so we need to have the free
3:28:30account so basically this is a one month
3:28:32free account you can able to create it
3:28:35so like this first of all you have to go
3:28:37have the subscriptions so the
3:28:39subscription means same as like what
3:28:41kind of uh subscription currently we're
3:28:43having as of now I don't any
3:28:45subscriptions I'm going to create it
3:28:47like uh this is for the 12 months or
3:28:51subscription free after that we have to
3:28:53go and purchase it either pay as it go
3:28:57let it come so that's the first one is a
3:28:59free trial like after that we have to go
3:29:01P go right as I told you P go means like
3:29:05depending on whatever we using right we
3:29:07are paying for it that's it okay we not
3:29:10once you are application once once you
3:29:12done with the work we can uh stop the
3:29:14services right so it means we're not
3:29:16paying for any anything to the it's not
3:29:19going to consume the cost okay and then
3:29:22for the student also we have the 12
3:29:24months free thing we can use it so first
3:29:27of you have to go to get the
3:29:28subscription once the subscription
3:29:29created we have to go and have the
3:29:31resource Group monitor a lot of things
3:29:34right so for different category wise we
3:29:36talk about right we have a storage we
3:29:38have a networking uh we have a what what
3:29:41is called uh database right so like this
3:29:44we have machine learning also we have a
3:29:45some of services suppose if you want to
3:29:47have the phase API phase deduction
3:29:49computer visions and then uh like for
3:29:53cognitive search right like this for
3:29:55analytical for Azu datab bricks so for
3:29:58that everything for analytical kind of
3:29:59things there are lot of things right
3:30:01powerb Integrations so we talk about
3:30:03powerb integration right so for that we
3:30:05have a separate analytical service okay
3:30:08so we have a compute service so comput
3:30:10service we initially right so how do you
3:30:12want to execute it okay how do we most
3:30:13we want it and what are availability
3:30:16which zone you want to store it what is
3:30:17the disk storage and virtual machines
3:30:21virtual storage like any image app
3:30:24services so there are compute services
3:30:27and then we have a container we have a
3:30:29database level Services we have a devops
3:30:32kind of things Services everything we
3:30:35are having services on the cloud
3:30:38[Music]
3:30:40one
3:30:43before knowing what is azure data brakes
3:30:46we must know what data brakes actually
3:30:49is well according to the definition data
3:30:52brakes is a web-based platform for
3:30:54working with Apache spark that provides
3:30:57automated cluster management and IPython
3:30:59style notebooks so basically data breaks
3:31:03developed by the creators of Apaches
3:31:05spark is nothing but a web-based
3:31:08platform which is also One-Stop product
3:31:10for all data requirements like storage
3:31:13and Analysis it was originally founded
3:31:16to provide an alternative to the map
3:31:19reduce system and provide a just in time
3:31:22cloud-based platform for big data
3:31:24processing clients it can derive
3:31:27insights using spark xql provide active
3:31:31connections to visualization tools such
3:31:33as powerbi click View and tab View and
3:31:37also build predictive models using
3:31:39sparkml datab BRS also can create
3:31:42interactive displays text and code
3:31:45tangibly so in short it is an
3:31:48alternative to a map redu system data
3:31:51brakes is now integrated with Microsoft
3:31:53Azure Google Cloud platform and Amazon
3:31:57making it easy for the business to
3:31:59manage a colossal amount of data and
3:32:02Carry Out machine learning tasks as we
3:32:05got to know that data braks is
3:32:06integrated with all the three
3:32:08cloud-based platforms today we'll
3:32:11discuss on one of those that is azure
3:32:13data brakes so let us understand what is
3:32:16azure data breakes Azure data bricks
3:32:19lakeh house is a platform that offers a
3:32:22uniform collection of tools for building
3:32:24deploying sharing and supporting
3:32:27Enterprise grade Data Solutions to a
3:32:29scale it integrates with cloud storage
3:32:32and Security in your cloud account and
3:32:35manages and deploy Cloud infrastructure
3:32:37on your behalf Azure data break supports
3:32:40python Scala R Java and SQL as well as
3:32:45data science Frameworks and libraries
3:32:48including tensor flow py torch and pych
3:32:51loarn now that we have come to know what
3:32:54is azure data break let us understand
3:32:57why should we use Azure data brakes well
3:33:01our customers use Azure data brakes to
3:33:03process store clean share analyze model
3:33:07and monetize their data sets with
3:33:10solution from powerbi to machine
3:33:12learning you can use Azure data brakes
3:33:15platform to build many different
3:33:16applications spanning data personas
3:33:20customers who fully embrace the lak
3:33:22house take full advantage of the unified
3:33:25platform to build and deploy the data
3:33:28engineering workflows machine learning
3:33:30models and analytics dashboard that
3:33:33power Innovations and insights across an
3:33:37organization the Azure data breakes
3:33:39workspace provid provides user
3:33:41interfaces for many code data task
3:33:44including tools which we'll be
3:33:46discussing one by one so first is
3:33:48optimized Spark engine it is a simple
3:33:51data processing on autoscaling
3:33:54infrastructure powered by highly
3:33:56optimized Apache spark for up to 50x
3:33:59Performance Gaines next is machine
3:34:01learning
3:34:02runtime it is a oneclick access to
3:34:05preconfigured machine learning
3:34:07environment for augumented machine
3:34:10learning with state-ofthe-art and
3:34:12popular Frameworks such as py toch
3:34:15tensor flow and psychic learn we also
3:34:18have mlflow that is used to track and
3:34:22share experiments reproduce runs and
3:34:24manage model collaboratively from a
3:34:28central repository well here you can use
3:34:31your preferred language including python
3:34:34Scala R spark SQL and net whether you
3:34:38use serverless or provis visioned
3:34:41compute resources so that you can
3:34:43quickly access and explore data find and
3:34:46share new insights and build models
3:34:49collaboratively with the language and
3:34:51tools of your choice in Azure data
3:34:55breakes you have Enterprise gr security
3:34:57which is an effortless native security
3:35:00that protects your data where it lives
3:35:02and creates complaint private and
3:35:05isolated analytics workspace across
3:35:08thousands of users and database
3:35:11apart from that it is also production
3:35:14ready that means you can run and scale
3:35:16your most Mission critical data
3:35:19workloads with confidence on a trusted
3:35:22data platform with ecosystem integration
3:35:25for cicd and monitoring you also have
3:35:29collaborative notebooks through which
3:35:32you can quickly access and explore data
3:35:35find and share new insights and build
3:35:37models collaboratively with the
3:35:39languages and tools of your choice and
3:35:42the most important thing it has Delta
3:35:45Lake that brings data reliability and
3:35:47scalability to your existing data lake
3:35:50with an open-source transactional
3:35:52storage layer designed for full data
3:35:55life cycle it has a native integration
3:35:58with Azure services that means you can
3:36:01complete your endtoend analytics and
3:36:03machine learning solution with deep
3:36:06integration with Azure services such as
3:36:08Azure data Factory azure data L Storage
3:36:12aure machine learning and
3:36:14powerbi last but not the least it has an
3:36:17interactive workspace that means you can
3:36:20enable seamless collaborations between
3:36:23data scientists data engineers and
3:36:26business analysts so these were the key
3:36:28features that makes Azure data brakes
3:36:31unique till now we have got an idea
3:36:34about the data brakes and Azure data
3:36:36brakes and why are we using Azure data
3:36:39braks now let let us deep dive in by
3:36:42understanding how does Azure data brakes
3:36:45actually work as I said Azure data
3:36:48brakes is structured to enable secure
3:36:50cross functional team collaboration
3:36:53while keeping a significant amount of
3:36:55backend Services managed by Azure data
3:36:58braks so you can stay focused on your
3:37:01data science and data analytics and data
3:37:04engineering task it operates out of a
3:37:07control plane and a data plane Although
3:37:10our architectures can vary depending on
3:37:12the custom configuration such as when
3:37:15you have deployed a Azure data break
3:37:17workspace to your own virtual Network
3:37:20which is also known as vnet injection
3:37:23now let us consider this architecture it
3:37:26is a common structure and a data flow of
3:37:29an Azure data braks here it consists of
3:37:32control plane and data plane so let us
3:37:35understand what is control plane the
3:37:37control plane includes the backend
3:37:39services that Azure data brakes manage
3:37:42in its own Azure account notebook
3:37:45commands and many other workspace
3:37:47configurations are stored in the control
3:37:50plane and encrypted at the rest if we
3:37:53talk about data plane it is managed by a
3:37:56your Azure account and it is where your
3:37:58data resides this is also where data is
3:38:02processed you can use Azure data break
3:38:04connectors so that your clusters can
3:38:07connect to external data sources outside
3:38:09your aure account to ingest data or for
3:38:13storage you can also ingest data for
3:38:16external streaming data sources such as
3:38:18even datas streaming datas iot data or
3:38:22many more well your data is stored at
3:38:26rest in your Azure account in the data
3:38:29plane and in your own data sources not
3:38:32the control plane so you maintain
3:38:34control and ownership of your data if we
3:38:38talk about the job results then it
3:38:40resides in storage in your account
3:38:42itself interactive notebook results are
3:38:45stored in the combination of control
3:38:47plane that is partial result for
3:38:50presentation in the UI and your Azure
3:38:53storage if you want interactive notebook
3:38:56results stored in only your cloud
3:38:58account storage then you can ask data
3:39:01break representative to enable
3:39:03interactive notebook result in the
3:39:05customer account for your workspace note
3:39:08that some metadata about results such as
3:39:12chart columns names continues to be
3:39:14stored in control plane itself so this
3:39:17is the basic architecture of your data
3:39:20break to know it more in simplified
3:39:23manner let us have a quick Hands-On on
3:39:26Azure data brakes so here we'll be
3:39:29integrating Azure data brakes with the
3:39:32Azure blob storage that is a service
3:39:35provided by Microsoft Azure so as you
3:39:38can see this is a small workflow of how
3:39:41we'll be working on the demo so as we
3:39:44know Azure blob storage and Azure data
3:39:47brakes are both Services provided by
3:39:49Microsoft Azure now these are two
3:39:52separate services but as long as you are
3:39:55using them in same Resource Group you
3:39:58can integrate these two Services well
3:40:00now you must be wondering why do we need
3:40:02to combine these so let's say that
3:40:05Microsoft Azure provides a multitude of
3:40:08services it is often beneficial to
3:40:11combine multiple Services together to
3:40:13approach your use case so if you combine
3:40:16multiple Services then we don't need to
3:40:19engage your local hardware in anything
3:40:21like well for example currently I'm
3:40:24using a laptop maybe my laptop can be a
3:40:26lower configuration or it might not have
3:40:29enough space to process a huge amount of
3:40:32data in that case I would want the cloud
3:40:35service to handle all my use cases and
3:40:38my big data storage that I have
3:40:40so in this workflow as we said that we
3:40:43will integrate Azure data brakes and
3:40:46Azure blob storage that means we're
3:40:48going to combine them so here what you
3:40:50do basically is you interact with the
3:40:52coding notebook which is nothing but the
3:40:55IPython or Jupiter notebook that your
3:40:59Azure data brakes will create then what
3:41:01you do is then you type some coding
3:41:04commands in your coding notebook and
3:41:07then these commands are sent to the data
3:41:10Brak service and then what happens is
3:41:13the datab break service receives the
3:41:16commands from the coding notebook and it
3:41:19sends those commands to your Azure
3:41:21cluster so whatever cluster you have
3:41:23created after creating your data brakes
3:41:26so whatever commands that we are writing
3:41:28in the coding notebook is sent through
3:41:30data brakes to your clusters that you
3:41:33have created now depending on the
3:41:36authentication provided to the cluster
3:41:38with regards to your blob storage
3:41:40account the authentication commands are
3:41:44then sent to the blob storage account
3:41:47saying that the data is fetched from the
3:41:49desired directory then it is brought
3:41:53back inside the cluster and that data
3:41:56that has received by the cluster then is
3:42:00processed and whatever output you get
3:42:03out of after the process is you can see
3:42:07it on your coding notebook now whatever
3:42:10output you receive from the coding
3:42:12notebook you can also store that
3:42:14particular output that you have got from
3:42:17the coding notebook back to your blob
3:42:19storage account or space so all of this
3:42:23is integrated easily and it is handled
3:42:28in a very simpler manner well now that
3:42:31we have understood the architecture let
3:42:33us now know how to implement it with a
3:42:37Hands-On for this we need to quickly
3:42:39sign in to our Azure portal so once we
3:42:42sign in we land in our dashboard of the
3:42:45Microsoft Azure as you can see here the
3:42:47all the Azure services are mentioned
3:42:49here and once you create your resource
3:42:52Group so it gets highlighted also now
3:42:55let us quickly open our Azure data braks
3:42:58and let us create our Azure data braks
3:43:01so for this once we come here we need to
3:43:03just simply
3:43:06create and here they ask for basic
3:43:09details
3:43:10like your resource Group then workspace
3:43:13name region so let us fill up one by one
3:43:16so here as we don't have any Resource
3:43:18Group so we'll be creating a new
3:43:20Resource Group as we working on a demo
3:43:23so let's name it as Azure data brakes
3:43:28demo and
3:43:30let's create it now that we have created
3:43:33the resource Group now they ask for the
3:43:35workspace name so we'll give it a unique
3:43:38name you can keep it any so let's name
3:43:42it as aure data braks hands on and then
3:43:49you can choose your region so here you
3:43:52find uh different types of region where
3:43:54you can like work or create your Azure
3:43:57data bricks so it is up to you whatever
3:44:00region you choose for me I'll be keeping
3:44:03West us as it is after that we come to
3:44:06the pricing tier so here there are three
3:44:09types standard premium and trial so as
3:44:12we are working on the like we are just
3:44:14having a practical knowledge we will be
3:44:17using the trial version and we'll just
3:44:19review and create now here you can just
3:44:22check on to your whatever details you
3:44:24had filled previously whether it is
3:44:26correct or no so once it is validated we
3:44:29can just create this data brakes now as
3:44:32you can see they are initializing the
3:44:35deployment so it may take a while to
3:44:37create a datab break
3:44:40all right so here you see the deployment
3:44:42is in
3:44:43progress so once your deployment is
3:44:46complete you can go to your uh resource
3:44:49and here you land up to your datab
3:44:52brakes page now here what we need to do
3:44:55is launch our workspace so it will move
3:44:59us to our main Azure portal so it will
3:45:02sign us in the Azure data breakes so
3:45:05this is the dashboard of the data braks
3:45:07now here you have different options like
3:45:10notebook and um data import partner
3:45:13connect transform data and many other
3:45:15things here you can set up your
3:45:18workspace like create a cluster import
3:45:21data build a data pipeline as well now
3:45:23one thing I should tell you that data
3:45:25bricks works on dbfs now what do you
3:45:28mean by dbfs is well dbfs is nothing but
3:45:32datab brakes file system which is a
3:45:34distributed file system mounted on Azure
3:45:36databas workpace and are available on
3:45:39the Azure databas cluster so it allows
3:45:42you to interact with the object storage
3:45:45using directory and file semantics
3:45:48instead of cloud specific API commands
3:45:50and it helps you out to mount Cloud
3:45:52object storage location so that you can
3:45:55map your storage credentials to the path
3:45:57in the GE datab brakes workspace it also
3:46:00simplifies the process of persisting
3:46:02files to object storage allowing the
3:46:05virtual machines and attacked volume
3:46:07storage to be safely deleted on the
3:46:09cluster termination well these are the
3:46:12things which you can like implement it
3:46:14while you're working with the Azure data
3:46:17breakes now let's get back to our data
3:46:19brakes now that we have come to this
3:46:21page so the first thing to be done over
3:46:23here is to create a cluster now we go to
3:46:26create a cluster now here we go to
3:46:28create compute now here as you can see
3:46:32we have the cluster name so here we can
3:46:36edit the cluster name based on your
3:46:38requirement so let us keep it as a data
3:46:41brakes cluster and after this we have
3:46:44the policy so here policy is nothing but
3:46:47a cluster policy defines the limit on
3:46:50the attributes available during the
3:46:52cluster creation so here we have
3:46:55different types of policies that is
3:46:57unrestricted personal compute power user
3:47:00compute shared compute so we will keep
3:47:04it unrestricted for time being now there
3:47:08are two types of cluster mode one is
3:47:10multi node and single node so in multi
3:47:14node we can specify the minimum number
3:47:16of workers and the maximum number of
3:47:18workers so here minimum numbers can be
3:47:21two or you can specify it upon this is
3:47:24basically a standard limit for minimum
3:47:27and maximum whereas you can change it
3:47:29accordingly based on your requirement so
3:47:32now as of now as we are only practicing
3:47:34so we will just disable this Autos
3:47:37scaling and we can specify are number of
3:47:40workers over here so we can keep only
3:47:43one worker as it is only for practicing
3:47:46whereas we can also like change the
3:47:49timings for termination if the
3:47:51particular cluster is inactive so you
3:47:54can specify that much of time to it
3:47:57apart from that we have our access mode
3:48:01that is nothing but there are three
3:48:03types of access mode that is single user
3:48:05shared user and no isolation shared so
3:48:09so here we'll keep it as it is and your
3:48:12single user access is nothing but our
3:48:14subscription after that we come to our
3:48:16performance so here we need to specify
3:48:19our runtime version so here in our case
3:48:22it is runtime 11.3 LTS color
3:48:262.12 and our work type can be standard
3:48:30whereas there are other versions as well
3:48:33but as of now we don't require much of
3:48:36it so we'll be going with the standard
3:48:39vers version itself and whereas we have
3:48:41already specified our workers and
3:48:44termination time is also been mentioned
3:48:46now here towards your right you can see
3:48:49the whole summary of your cluster what
3:48:51have you been uh like creating so once
3:48:55you review it you can just create this
3:48:57particular cluster now as you can see it
3:49:00is loading it takes a while to like
3:49:03create a cluster here in this section
3:49:07you see the status of your cluster
3:49:10So currently it is in a pending mode
3:49:12like it has been creating so we'll wait
3:49:16for a while so now as you can see our uh
3:49:19cluster has been created and it is in
3:49:21the running mode now after this what we
3:49:25need to do is go back to our main
3:49:27dashboard and we need to create a new
3:49:30notebook so we'll create a new notebook
3:49:32over here and here we have already
3:49:36specified the cluster so we had created
3:49:39created now so it has automatically
3:49:41taken and uh here we need to specify our
3:49:44default language so you can choose any
3:49:47of them so here there there are four
3:49:49types of languages which you can choose
3:49:52here I would be taking Scala for now and
3:49:55we can give a simple name to this
3:49:57notebook that is it can be anything of
3:50:00your choice so let us name it as data
3:50:04brakes notebook and let's create so this
3:50:08will start very quickly it doesn't take
3:50:10that much of time now here it is like
3:50:14you need to like run your command over
3:50:16here so you just need to type down your
3:50:18command and just hit enter and it starts
3:50:21running so now in this we will know how
3:50:25to upload a file through uh C Azure
3:50:29storage service that is our Azure blob
3:50:32storage so it's basically we need to
3:50:35First integrate the Azure blob storage
3:50:38so let us know how so before this we
3:50:42need to go back to our aure portal and
3:50:45here we need to First create our storage
3:50:47account so here as you can see we have
3:50:49our storage accounts and here we'll
3:50:53create our blob storage so let's create
3:50:57so now here for creating storage account
3:51:00you need to give your details over here
3:51:03so as here we had already created the
3:51:08resource Group previous L while creating
3:51:10a data break so we will be selecting the
3:51:12same and after that we will give a
3:51:15storage account name so let us name it
3:51:18as data braks storage account and in
3:51:23region we need to provide the specific
3:51:26region whichever you like to choose so
3:51:29as previous I had choosen West us so
3:51:32I'll be choosing that as
3:51:34well and here now when we come to the
3:51:38performance so here we can choose any of
3:51:41those um any of the two options given
3:51:44below so here we have standard and
3:51:46premium So currently we'll be going with
3:51:48the Standard
3:51:49Version and when we talk about redund
3:51:52dency so here we have two types of
3:51:55redundancy as you can see locally
3:51:58redundant storage here it means the
3:52:00lowcost option with the basic protection
3:52:02against server rack and drive failur
3:52:05recommended for non-critical scenarios
3:52:08whereas Geo redundant storage is like
3:52:11for intermediate option with failover
3:52:13capabilities in secondary region
3:52:15recommended for backup scenarios so
3:52:18locally means it happens within the
3:52:20region not across the whole world so
3:52:24here we will be choosing the locally
3:52:26redundant storage now we have specified
3:52:29everything so now let us review so here
3:52:33before like creating you need to check
3:52:37all your details once you review it just
3:52:40create and your storage account is in
3:52:45initialization stage so once it gets
3:52:48deployed we will start working on that
3:52:51so now our deployment is complete so we
3:52:53can go to our resource so here all of
3:52:56the permissions are automatically
3:52:58managed within the same Resource Group
3:53:01so we don't have to worry about any
3:53:03permissions
3:53:04requirement so now that we have created
3:53:07our storage account now here we need to
3:53:10create our container so we'll go to our
3:53:15containers and we'll so here we can give
3:53:19it a name to our container it can be
3:53:22anything name it as storage account one
3:53:27and let's just create so as you can see
3:53:30our container has been created now we
3:53:32will quickly go onto this container and
3:53:35here we don't have anything in this
3:53:37container so the container is empty now
3:53:40we need to upload some files in this
3:53:42particular container so we'll just
3:53:44quickly go to upload and here we will go
3:53:49to select file so here what does it do
3:53:51it will get connected to my Windows File
3:53:54you can take any files over here so as
3:53:56of now I will just take a
3:54:00normal Excel a CSV file and we'll just
3:54:05simply upload it well you can upload
3:54:08more than one files if you want to like
3:54:10upload it in the containers you can also
3:54:13have a larger files but it may cost
3:54:16according to the given size now we have
3:54:20a file in place and we also have created
3:54:24a notebooks so now we have to integrate
3:54:28The Blob storage with the data brid so
3:54:32for that we need to run a code now here
3:54:35the code looks a bit complex so that
3:54:38here so here first we need to create a
3:54:41token so that we can get access to the
3:54:44files so we'll just copy this whole code
3:54:50so that there's nothing to memorize this
3:54:53can be provided to you while you are
3:54:56practicing so now we'll just copy paste
3:54:59the whole command so here now as you can
3:55:05see the container name so here we need
3:55:08to spe specify the container name and
3:55:10the storage account that we have created
3:55:12so we'll quickly go back to our aure
3:55:14portal and we will fill up the details
3:55:17so as they asked for the container name
3:55:20so here we need to specify our container
3:55:22name and the storage account which we
3:55:24have created so let's quickly go back to
3:55:27our a portal and here as you can see
3:55:30your container name
3:55:33is given so we'll just quickly copy
3:55:37that go let's go back and we'll just
3:55:43copy this name and we'll paste it here
3:55:49now same thing to be done with our
3:55:51storage account so we'll go back to our
3:55:54storage account and here we'll see this
3:55:57is our storage account name so we'll
3:56:00just copy and we'll just paste it over
3:56:04here now for the SOS token we need to go
3:56:09back to our containers and here we will
3:56:14go to the Shar access signature so here
3:56:18we will be getting our SAS SAS tokens so
3:56:23SAS token can be generated for a limited
3:56:26amount like for a given period of time
3:56:28so you need to specify like when you
3:56:31create your SAS token so the time it has
3:56:34been started it would be valid from the
3:56:37time it has been started till its time
3:56:39of expiry so we'll just allow the
3:56:44services containers and the objects and
3:56:48let the time be as it is as it is been
3:56:51specified and now let's
3:56:54generate so now here as you can see you
3:56:57can find your SAS token now we need to
3:57:00just simply copy this and go back to our
3:57:04data brakes and we'll paste it over here
3:57:08so we have just added our SAS token let
3:57:11us just verify whether it has been
3:57:14correctly copied or not so I feel
3:57:17everything looks perfect rest other
3:57:19things remains to be same so here what
3:57:22happens you keep on creating new
3:57:24variables so here we have created a URL
3:57:28so by appending the container and the
3:57:31storage account so once we like we
3:57:35specify our URL and our configuration
3:57:37then it comes the D PS so here we
3:57:40specify our source Mount point and our
3:57:45maps to that particular token so this is
3:57:49basically where the data braks helps you
3:57:52like get your sources from your
3:57:54different like different services so
3:57:57here the data brakes plays the role
3:57:59where it provides you the data which
3:58:02from where you want to extract from and
3:58:05it shows it over here and now you just
3:58:08just hit shift plus enter so now it is
3:58:13like running the
3:58:15command so here it what does it do as we
3:58:18had discussed in our architecture so it
3:58:20communicates first with the data brakes
3:58:22and then it communicates with our
3:58:23cluster and get all the sources and then
3:58:27shows its result on the cluster itself
3:58:30so as you can see your command has been
3:58:36run okay so it shows some error over
3:58:39here so let's quickly solve that error
3:58:42all right now let's quickly run it
3:58:46again so it may take a while so here as
3:58:50I said it will first connect with the
3:58:53data brakes and then communicate with
3:58:55that and then it will communicate with
3:58:59your cluster and get all the resources
3:59:01from there so here as you can see they
3:59:04have specified your container name
3:59:06storage name and your token which has
3:59:09been
3:59:10specified now what we need to do is
3:59:14check whether our file has been take
3:59:17extracted from our storage account or no
3:59:21so here we'll spec first we'll specify
3:59:23our variable so let's
3:59:26specify with
3:59:29p. read and now we're going to give the
3:59:34format of the file so as we had uh
3:59:37uploaded the CSV file so we'll just
3:59:40specify it over here
3:59:43CSV and then we can give them some of
3:59:46the options that um they should show so
3:59:50let's give it an option like we can ask
3:59:53them to show the header and their value
3:59:58should be
4:00:00true then we can also specify the infers
4:00:03scamma so we'll just add on that as well
4:00:07and for efficient uh data requirement we
4:00:10will just provide them with a mode that
4:00:13can be fail fast mode and then we upload
4:00:17the like mentioned the file name so here
4:00:21we paste this Mount slash staging that
4:00:24is our Mount point and we'll paste it
4:00:28over here and then we will specify our
4:00:34file name that is let's go to our
4:00:37containers
4:00:39let's check what's the name python 1.
4:00:44CSV let's copy the name and we paste it
4:00:49so now let's just shift
4:00:53enter uh so it shows some error let's
4:00:56see what is it okay so here we had
4:01:00specified a wrong command so let's just
4:01:04resolve it and
4:01:07let's run it
4:01:09again all right so let us see how it
4:01:13looks like so now just do
4:01:17TF dot show and
4:01:22let's specify some
4:01:25amount and just shift
4:01:28enter so as you can see so they have
4:01:31mentioned here the decimal and uh
4:01:35description the percentage your headers
4:01:37have been been provided and the number
4:01:41of the number of rows that we required
4:01:43they have specified that so this is how
4:01:45we integrate the two services and we
4:01:49process data through data breaks now let
4:01:52us look onto some of the popular use
4:01:54cases that makes a your data Brakes in a
4:01:57huge demand well data brakes isn't a
4:02:01catchall solution for every business
4:02:03scenario so there are the best use cases
4:02:06for Azure data braks first is database
4:02:09and Mainframe modernization data storage
4:02:13collection and processing are incredibly
4:02:16important in modern businesses well if
4:02:19you're looking to modernize your data
4:02:21legs or looking into Mainframe
4:02:24modernization applications then Azure
4:02:27data brakes has all the Integrations you
4:02:29need next is machine learning production
4:02:33pipeline here using the underlying power
4:02:36of ml flow data bricks is a good choice
4:02:40if you need to get machine learning
4:02:42applications into production getting
4:02:45data signs out of the development and
4:02:47into production is a common problem and
4:02:50Azure data breaks can help you
4:02:52streamline that workflow if you talk
4:02:54about big data processing then Azure
4:02:57data brakes is one of the most coste
4:02:59effective options for big data
4:03:01processing in terms of performance
4:03:04versus cost it offers higher efficiency
4:03:07if your business needs the best
4:03:09performance for one demand data
4:03:11processing then data brakes will likely
4:03:14be your best choice next is business
4:03:17intelligence integration integrating
4:03:19business intelligence tools means you
4:03:22can open your data Lake to analysts and
4:03:25engineer more easily there's no need for
4:03:28creation of new pipelines when analysts
4:03:31need access to new data the data can be
4:03:34shared through SQL analytics powerbi and
4:03:37tablet you if this is a bottleneck for
4:03:40your business then data bricks will help
4:03:43you enable your business intelligence
4:03:45team now these were the four popular use
4:03:48cases by Azure data braks if your
4:03:51business fits in one of these use cases
4:03:54then it might be the solution for
4:03:56[Music]
4:04:01you let us start with the US primary
4:04:04election use case first in this use case
4:04:07we will be discussing about the 2016
4:04:10primary elections in the primary
4:04:12elections the contenders from each party
4:04:14compete against each other to represent
4:04:16his or her own political party in the
4:04:18final elections there are two major
4:04:20political parties in the US the
4:04:22Democrats and Republicans from the
4:04:24Democrats the contenders were Hillary
4:04:26Clinton and Bernie Sanders and out of
4:04:28them Hillary Clinton won the primary
4:04:30elections and from the Republicans the
4:04:32contenders were Donald Trump Ted Cruz
4:04:34and a few others as you already know
4:04:36Donald Trump was the winner from the
4:04:38Republicans so now let us assume that
4:04:41you are an analyst already and you have
4:04:43been hired by Donald Trump and he tells
4:04:45you that I want to know what were the
4:04:47different reasons because of which
4:04:49Hillary Clinton won and I want to carry
4:04:52out my upcoming campaigns based on that
4:04:54so I can win the favor of the people
4:04:56that voted for her so that was the
4:04:58entire agenda so this is the task that
4:05:02has been given to you as a data analyst
4:05:04so what is the first thing that you will
4:05:06need to do the first thing you'll do is
4:05:09that you'll ask for data and you have
4:05:11got two data sets with you so let us
4:05:14take a look at what this data sets
4:05:15contains so this is our first data set
4:05:18which is the US primary election data
4:05:20set so these are the different fields
4:05:21present in our data set so the first
4:05:23field is state so we've got the list of
4:05:25the state of Alabama the state
4:05:27abbreviation for Alabama is Al we've got
4:05:30the different counties in Alabama like
4:05:32aruga Baldwin Barber bib Blount bulock
4:05:34Butler Etc and then we've got fips now
4:05:37fips are federal information processing
4:05:39standards code so this is basically
4:05:41means zip code then we've got the party
4:05:45to which we would be analyzing the
4:05:46Democrats only because we want to know
4:05:50what was the reason for Hillary
4:05:51Clinton's win so we will be analyzing
4:05:54the Democrats only and then we've got
4:05:56the candidate and since I told you there
4:05:58were two candidates Bernie Sanders and
4:06:00Hillary Clinton so we've got the name of
4:06:02the candidate here and the number of the
4:06:04votes each candidate got so Bernie
4:06:06Sanders got 54 44 in aruga county and
4:06:09Hillary Clinton got to 2387 and this
4:06:12field over here represents the fraction
4:06:14of the votes so if you add these two
4:06:16together you will get a one so this
4:06:19basically represents the percentage of
4:06:21vote each of the candidates got so let's
4:06:24take a look at our second data set now
4:06:27so this data set is the US County
4:06:29demographic features data set so the
4:06:31first we will have again fips in the
4:06:34area name aruga County Baldwin and
4:06:36different other count counties in
4:06:38Alabama and other states also the state
4:06:40abbreviation so here it is only showing
4:06:43Alabama and the fields that you see here
4:06:45are actually the different features you
4:06:47won't know what this exactly contains
4:06:49because it is written in a coded form
4:06:52but let let me give you an example what
4:06:54this data set contains let me um tell
4:06:57you that I'm just showing you a few rows
4:06:59of the data set this is not the entire
4:07:01data set so this contains different
4:07:03fields like population in 2014 in 2010
4:07:06the sex ratio how many females males and
4:07:09then based on some ethnicity how many
4:07:11Asians how many Hispanic how many black
4:07:14American people how many um black
4:07:17African people and then there is also
4:07:19based on the age groups how many infants
4:07:22uh how many senior citizens how many
4:07:24adults so there are a lot of fields in
4:07:27our data set and this will help us to
4:07:29analyze and actually find out what led
4:07:30to the winning of Hillary Clinton so now
4:07:33you have seen our data set you have to
4:07:35understand your data set you have have
4:07:37to figure out what are the different
4:07:40features or what are the different
4:07:42columns that you are going to use and
4:07:44you have to think of a strategy or think
4:07:46of how you're going to carry out this
4:07:48analysis so this is the entire solution
4:07:51strategy so the first thing you will do
4:07:53is that you need a data set and you've
4:07:55got two data sets with you the second
4:07:58thing that you'll need to do is to store
4:08:00that data into hdfs now hdfs is how to
4:08:03distributed file system so you need to
4:08:05store the data the next step is to
4:08:07process that data using spark components
4:08:09and we will be using spark SQL spark M
4:08:12lib Etc so the next task is to transform
4:08:16that data using spark SQL transforming
4:08:18here means filtering out the data and
4:08:20the rows and columns that you might need
4:08:22in order to implement or in order to
4:08:24process this the next step is clustering
4:08:27this data using spark M lib and for
4:08:29clustering uh our data we will be using
4:08:32K means and the final step is to
4:08:33visualize the result using Zeppelin now
4:08:36visualizing this step is also very
4:08:38important because without the
4:08:39visualization you won't be able to
4:08:41identify what were the major reasons and
4:08:43you won't be able to gain proper
4:08:45insights from your
4:08:46data now don't be scared if you're not
4:08:49familiar with terms like spark SQL spark
4:08:51AMG K means clustering you will be
4:08:54learning all of these in today's session
4:08:56so this is our entire strategy this is
4:08:59what we're going to do today this is how
4:09:01we're going to implement this use case
4:09:03and find out why Hillary Clinton won so
4:09:06now let me give give you a visualization
4:09:09of the results so I'll just show the
4:09:11analysis that I have performed and I'll
4:09:14show you how it
4:09:15looks so this is zeppelin which is in my
4:09:18master node in my Hadoop cluster and
4:09:21this is where we're going to visualize
4:09:23our data so there's a lot of code don't
4:09:26be scared this is just Scala code with
4:09:28spark SQL and at the end you will be
4:09:30learning how to write this
4:09:32code so I'm just jumping onto the
4:09:35visualization part so this this is the
4:09:38first visualization that we've got and
4:09:40we've analyzed it according to different
4:09:42ethnicities of people for example in our
4:09:45xaxis we have foreign born persons and
4:09:47in y axis we're seeing that among the
4:09:49foreign born people what is the
4:09:51popularity of Hillary Clinton among the
4:09:53Asians and the circles represent the
4:09:55highest values the bigger circle is the
4:09:57bigger counts so we have made a few more
4:10:02visualizations so now we've got a line
4:10:04graph that compares the votes of Hillary
4:10:06Clinton and burn Bernie Sanders together
4:10:09again we have got an area graph also
4:10:12that compares Bernie Sanders and Hillary
4:10:14Clinton votes and hence we have a lot
4:10:17more
4:10:18visualization we have uh got our bar
4:10:20charts and everything finally we also
4:10:23have uh State and County wise
4:10:25distribution of votes so these are the
4:10:28visualizations that will help you derive
4:10:30a conclusion to derive an answer
4:10:33whatever answer that Donald Trump wants
4:10:36from you and don't worry you'll be
4:10:38learning how to do that I'll explain
4:10:40each and every detail of how I've made
4:10:42these
4:10:42visualizations so let's get started with
4:10:44Hadoop and Spark we will start um with
4:10:48an introduction to Hadoop and Spark so
4:10:53now let's take a look at what is Hadoop
4:10:55and what is spark so Hadoop is a
4:10:58framework where you can store large
4:10:59clusters of data in a distributed Manner
4:11:01and then process them parallell then
4:11:04Hadoop has got two components for
4:11:06storage it has hdf f s which stands for
4:11:08Hadoop distributed file system and it
4:11:10allows to dump any kind of data across
4:11:13the Hadoop cluster and it'll be stored
4:11:15in a distributed manner in commodity
4:11:17hardware for processing you've got yarn
4:11:20which stands for yet another resource
4:11:22negotiator and this is the processing
4:11:24unit of Hadoop which allows parallel
4:11:26processing of the distributed data
4:11:28across your Hadoop cluster in
4:11:30hdfs then we've got spark so Apache
4:11:34spark is one of the most popular
4:11:35projects by Apache and this this is an
4:11:37open-source cluster Computing framework
4:11:40for real-time processing where on the
4:11:43other hand Hadoop is used for batch
4:11:45processing spark is used for real-time
4:11:47processing because with spark the
4:11:49processing happens in memory and it
4:11:51provides you with an interface for
4:11:53programming entire clusters with
4:11:54implicit data parallelism and fault
4:11:57tolerance so what is data parallelism
4:12:00data parallelism is a form of
4:12:02parallelization across multiple
4:12:04processes in parallel Computing
4:12:06environments a lot of parallel words in
4:12:08that
4:12:09sentence um so let me tell you simply
4:12:11that it basically means Distributing
4:12:13your data across nodes which operate on
4:12:16the data parallel and it works on fault
4:12:18tolerant systems like hdfs and S3 and is
4:12:22built on top of yarn because with yarn
4:12:24you can combine different tools like
4:12:26Apache spark for better processing of
4:12:28your
4:12:29data and if you see the topology of
4:12:31Hadoop and Spark both of them have the
4:12:33same topology which is a Master Slave
4:12:35topology so in Hadoop if you consider in
4:12:37terms of hdfs the master node as known
4:12:40as the name node and the working node or
4:12:42the slave nodes are known as data node
4:12:44and in spark the master is known as
4:12:47master and slave are known as workers so
4:12:49this is these are basically demons so
4:12:52this is a brief introduction to Hadoop
4:12:54and Spark and now let's take a look at
4:12:56spark complimenting Hadoop there's
4:12:59always been a debate about what to
4:13:00choose Hado spark but let me tell you
4:13:03that there is a stubborn misconception
4:13:04that Apache spark is an alternative to
4:13:06had
4:13:07and that is likely to bring an end to
4:13:09the era for Hadoop it is very difficult
4:13:11to say Hadoop versus spark because the
4:13:13two framers are not mutually exclusive
4:13:16but they are better when they are paired
4:13:18with each other so let's see the
4:13:20different challenges that uh we address
4:13:22when we are using spark and Hadoop
4:13:25together you can see the first point
4:13:28that spark processes data 100 times
4:13:30faster than map ruce so it gives us the
4:13:33results faster and it performs faster
4:13:35analytics the next point is spark
4:13:38applications can run on yarn leveraging
4:13:40Hadoop cluster and you know that Hadoop
4:13:42cluster is usually set up on commodity
4:13:44Hardware so we are getting better
4:13:46processing but we are using very lowcost
4:13:49hardware and this will help us cut our
4:13:51cost a lot so hence also the cheap uh
4:13:54cost optimization the third point is
4:13:56that Apache spark can use hdfs as
4:13:59storage so you don't need a different
4:14:01storage space for Apache spark it can
4:14:03operate on hdfs itself so you don't have
4:14:06to copy the same file again and if you
4:14:08want to process it with spark uh so
4:14:11hence you can avoid duplication of files
4:14:13so Hadoop forms a very strong foundation
4:14:16for any of the future Big Data
4:14:17initiatives and Spark uh is one of those
4:14:20big data
4:14:22initiatives it's got enhanced features
4:14:24like in memory processing machine
4:14:26learning capabilities and you can use it
4:14:28with Hadoop and Hadoop uses commodity
4:14:30Hardware which can give you better
4:14:33processing with minimum cost these are
4:14:37are the benefits that you get when you
4:14:38combine spark and hadu together in order
4:14:40to analyze Big Data let's see some of
4:14:43the big data use cases so the first big
4:14:45data use case is web detailing the
4:14:48recommendation engines uh whenever you
4:14:50go out on Amazon or any other online
4:14:53shopping site in order to buy something
4:14:55you will see some recommended items
4:14:57popping below your screen or to the side
4:14:59of your screen and that is all generated
4:15:01using big data analytics and AD
4:15:04targeting if you go to Facebook you see
4:15:06a lot of different items asking you to
4:15:07buy them and when you got uh search
4:15:09quality abuse and click fraud detection
4:15:13you can use big data analytics and
4:15:15Telecommunications also in order to find
4:15:17out the customer churn prevention the
4:15:20network performance optimization
4:15:22analyzing uh Network to predict failure
4:15:25and you can prevent loss before the
4:15:27error or before the fault actually
4:15:29occurs it's also widely used by
4:15:31governments for fraud detection and
4:15:33cyber security in order to introduce
4:15:35different welfare schemes uh justice it
4:15:37has been widely used by Healthcare and
4:15:40Life Sciences for health information
4:15:42exchange Gene sequencing serialization
4:15:45healthc Care Service quality
4:15:46improvements and Drug safety now let me
4:15:49tell you that with big data analytics it
4:15:51has been very easy in order to diagnose
4:15:53a particular disease and find out the
4:15:55Cure also so these are some more big
4:15:57data use cases it is also used in Banks
4:16:01and financial services for modeling true
4:16:03risk fraud detection credit card scoring
4:16:06analysis and and uh many more it could
4:16:08be used in Retail transportation
4:16:10services hotels and food delivery
4:16:12services and actually every field you
4:16:14name no matter whatever business you
4:16:16have if you're able to use Big Data
4:16:18efficiently your company will grow and
4:16:20you will be gaining different insights
4:16:22by using big data analytics and hence
4:16:24improve your business even more nowadays
4:16:28everyone is using uh big data and you
4:16:31you've seen different fields and
4:16:32everything is different from each other
4:16:34but everyone is using big data analytics
4:16:37and Big Data analysis can be done with
4:16:39tools like Hado and Spark Etc so this is
4:16:43why big data analytics is very much in
4:16:45demand today and why it is very
4:16:46important for you to learn how to
4:16:47perform big data analytics with tools
4:16:49like this so now let's take a look at a
4:16:51big data use solution architecture as a
4:16:53whole you're dealing with big data now
4:16:57the first thing that you need to do is
4:16:58you need to dump all those that data
4:17:00into hdfs and store it in a distributed
4:17:03way and the next thing is to process
4:17:05that data so that you can gain insights
4:17:08and we'll be using yarn because yarn can
4:17:10allow us to integrate different tools
4:17:12together which will help us to process
4:17:14the Big Data these are the tools that
4:17:16you can integrate with yarn you can
4:17:17choose either Apache Hive Apache spark
4:17:20map reduce Apache CFA in order to
4:17:22analyze big data and Apache spark is one
4:17:25of the most popular and most widely used
4:17:27tools with yarn in order to process big
4:17:29data so this is the an entire solution
4:17:32as a whole now so let's take a look at
4:17:35Apache sparkk
4:17:37Apache spark is an open source cluster
4:17:39Computing framework for real-time
4:17:41processing and it has been the thriving
4:17:43open- Source community and is most
4:17:45active Apache project uh at this moment
4:17:48and Spark components are what make
4:17:50Apache spark fast and reliable and a lot
4:17:52of spark components were built to
4:17:54resolve the issues that cropped up while
4:17:56using Hado map reduce so Apache spark
4:18:02has got the following components has got
4:18:04the spark core engine now the core core
4:18:07engine is for the entire spark
4:18:08Frameworks uh every component is based
4:18:11on and it is placed in the core engine
4:18:14so at first we've got uh spark SQL so
4:18:17spark SQL is a spark module for
4:18:19structured data processing and you can
4:18:21run a modified Hive queries on existing
4:18:23hadb deployments and then we've got
4:18:25spark
4:18:26streaming now spark streaming is the
4:18:28component of spark which is used to
4:18:30process real-time streaming data and is
4:18:32useful addition to the core spark API
4:18:35because it enables hive throughput fault
4:18:37tolerance stream processing of live data
4:18:40streams and then we've got spark mli uh
4:18:44this is the machine learning library for
4:18:46sparc and we'll be using spark MMA in uh
4:18:50to implement machine learning in our use
4:18:52cases too and then we've got graphx
4:18:54which is the graph computation engine
4:18:56and this is the spot API for graphs and
4:19:00graph parallel computation it has got a
4:19:03set of fundamental operators like
4:19:04subgraph joint purchases Etc then um
4:19:09you've got uh spark R so this is the
4:19:12package for R language to enable our
4:19:15users to leverage spark power from our
4:19:17shell so the people who have already
4:19:20been working on R are comfortable with
4:19:23it and they can use R shell directly at
4:19:25the same time and they can use spark
4:19:28using this particular component which is
4:19:30spark R you can write all your code in
4:19:32the r shell and Spark will process it
4:19:34for you now let's take a deeper look at
4:19:36a real istic people and all these
4:19:38important components so we've got spark
4:19:41core and Spark core is the basic engine
4:19:44for large scale parallel and distributed
4:19:47data processing the core is the
4:19:49distributed execution engine and Java
4:19:52Scala and python apis offer a platform
4:19:54for distributed edl development and
4:19:57further additional libraries which are
4:19:59built on top of the core allow uh for
4:20:02diverse streaming SQL and machine
4:20:05learning it's it's also responsible for
4:20:07scheduling Distributing and monitoring
4:20:09jobs in a cluster and also interacting
4:20:12with storage systems let's take a look
4:20:15at the spark architecture so Apache
4:20:17spark has a well-defined and layered
4:20:19architecture where all the spark
4:20:21components and layers are Loosely
4:20:23coupled and integrated with various
4:20:25extensions and libraries first let's
4:20:27talk about the driver program this is
4:20:30the spark driver which contains the
4:20:32driver program and Spark context uh this
4:20:35is the Central Point and entry point of
4:20:38the spark shell and the driver program
4:20:40runs the main function of the
4:20:42application and this is the place where
4:20:43Spark context is
4:20:46created well what is spark context spark
4:20:49context represents the connection to the
4:20:51entire spark cluster and it can be used
4:20:54to create resilient distributed data
4:20:56sets accumulators and broadcast
4:20:59variables on that cluster and you should
4:21:01know that only one spark context may be
4:21:04active per Java virtual machine and you
4:21:07must stop any active spark context
4:21:10before creating a new one let's talk
4:21:12about the driver program that runs on
4:21:14the master knob of the spark cluster it
4:21:17schedules the job execution and
4:21:18negotiates with the cluster manager this
4:21:21is the cluster manager over here and the
4:21:24cluster manager is an external service
4:21:26that is responsible for acquiring
4:21:28resources on that spark cluster and
4:21:30allocating them to a spark job then in
4:21:35the worker node we have got the
4:21:36executors the executor is a distributed
4:21:39agent that is responsible for the
4:21:41execution of tasks and Every Spark
4:21:44application has its own executor process
4:21:47executors usually run for their entire
4:21:50lifetime of the spark application and
4:21:52this phenomenon is also known as static
4:21:55allocation of executors but you can also
4:21:57opt for dynamic uh locations of
4:22:01executors where you can add or remove
4:22:03spark executors dynamically to match
4:22:06with the overall workflow okay so now
4:22:08let me tell you what actually happens
4:22:10when the spark job is submitted when a
4:22:12client submits a spark user application
4:22:14code the driver implicitly converts the
4:22:16code containing Transformations and
4:22:18actions into a logical directed ayic
4:22:21graph or dag and at this stage the
4:22:24driver program also performs certain
4:22:26kinds of optimizations like pipelining
4:22:29Transformations and then converts The
4:22:31Logical dag into a physical execution of
4:22:34a plan with a set of stages and after
4:22:38creating a physical execution plan it
4:22:40creates more physical execution units
4:22:44that are referred to as tasks under each
4:22:46stage and these tasks are bundled to be
4:22:49sent to the spark cluster so the driver
4:22:52program then talks to the cluster
4:22:54manager and negotiates for resources and
4:22:56the cluster manager then launches the
4:22:59executors on the worker nodes on behalf
4:23:01of the driver and at this point the
4:23:03driver sends tasks to the cluster
4:23:05manager based on the day of replacement
4:23:07and before the executors begin execution
4:23:10they first register themselves with the
4:23:11driver program so that the driver has
4:23:14got a holistic view of all the
4:23:17executors now the executors will execute
4:23:19the various tasks that are assigned to
4:23:21them by the driver program and at any
4:23:22point of time when the spark application
4:23:25is running the driver program will keep
4:23:27the on monitoring the set of executors
4:23:29that are running the spark application
4:23:31code and this driver program here also
4:23:34schedules future tasks based on data
4:23:37replacement by tracking the location of
4:23:39the cache data so I hope you have
4:23:41understood the architecture of spark any
4:23:44doubts all right no doubts now let's
4:23:47take a look at spark SQL and its
4:23:49architecture so spark SQL is the new
4:23:51module in spark and it integrates
4:23:53relational processing with Spark's
4:23:55functional programming API and it
4:23:57supports querying of data either by a
4:23:59SQL or via Hive query language so for
4:24:02those of you who have been familiar with
4:24:05rdbms uh so spark SQL will be a very
4:24:08easy transition from your earlier tools
4:24:11because you can extend the boundaries of
4:24:13traditional relational data processing
4:24:15with spark SQL and it also provides s
4:24:18support for various data sources and
4:24:21makes it possible to read SQL queries
4:24:23with code transformation and that is why
4:24:25spark SQL has become a very powerful
4:24:28tool this is the architecture of spark
4:24:30SQL so let's talk about each of these
4:24:32components one by one the first we have
4:24:35got the data source our API so this is
4:24:37the universal API for loading and
4:24:39storing structured data and it is built
4:24:42on support for Hive Avro Json jdbc CVS
4:24:47parkette Etc so it also supports the
4:24:50third-party integration through spark
4:24:52packages then you've got the data frame
4:24:54API dataframe API is the distributive
4:24:58collection of data that is organized uh
4:25:00into named columns and is similar to
4:25:03relational table in SQL that is used for
4:25:05storing data in tables so it is the
4:25:08domain specific language applicable to
4:25:10or
4:25:11DSL applicable on structured and
4:25:14semi-structured data so it processes
4:25:16data from kilobytes to pedabytes on a
4:25:19single node cluster to a multi- noode
4:25:21cluster and it provides different apis
4:25:23for python Java Scala and our
4:25:26programming so I hope you have
4:25:28understood all the architecture of spark
4:25:30SQL we will be using spark SQL in order
4:25:32to solve our use cases so these are the
4:25:34different commands to start start the
4:25:36spark Damons these are very similar to
4:25:38had of commands to start The hdfs Damons
4:25:41so you can see to start all the spark
4:25:44Damons uh so the spark Damons are master
4:25:46and worker and you can use this command
4:25:48to check if all the Damons are running
4:25:50on your machine you can use JPS like
4:25:53Hadoop and then in order to start the
4:25:55spark shell you can use this and you can
4:25:58go ahead and try this out so this is
4:26:00very similar to the hadu part that I
4:26:01just showed you earlier so I'm not going
4:26:03to do it again and then we've seen a
4:26:06Pache spark also so now let's take a
4:26:08look at K means and Zeppelin K means is
4:26:10the clustering method and Zeppelin is
4:26:12what we're going to use in order to
4:26:14visualize our data so let's talk about
4:26:17the K mean clustering now K means is one
4:26:20of the most simplest UNS supervised
4:26:22learning algorithms that uh solves the
4:26:25well-known clustering problem so the
4:26:27procedure of K means follows a simple
4:26:30and easy way to classify a data set to a
4:26:33certain number of clusters which is
4:26:35fixed prior to performing the clustering
4:26:37method so the main idea is Define K
4:26:40centroids one for each cluster and the
4:26:42centroids should be placed in a very
4:26:46cunning way because of different
4:26:48location re causes different results so
4:26:52here let's take an example so let's say
4:26:54that we want to Cluster total population
4:26:56of a certain location and so we want to
4:27:00Cluster them into four different uh
4:27:03clusters namely group one two and three
4:27:05and four so the main thing that we
4:27:06should keep in mind is that the objects
4:27:08in group one should be as similar as
4:27:11possible but there should be as much
4:27:13difference between an object in group
4:27:14one and group two it means that the
4:27:16points that are lying in the same group
4:27:19should have similar characteristics and
4:27:21it should be different from the points
4:27:23that are lying in a different cluster
4:27:25and the attributes of the objects are
4:27:27allowed to determine which object should
4:27:30be grouped
4:27:31together for example let us uh take in
4:27:36the same sample that we're using in the
4:27:37US County so let's consider the second
4:27:40data set we have used there are a lot of
4:27:42features that I already told you like
4:27:44there are age groups and they are
4:27:46categorized by professions and they also
4:27:48categorized by the ethnicity and uh so
4:27:52this is the thing that we are talking
4:27:54about so these are the attributes that
4:27:57will allow us to Cluster our data so
4:28:00this is K means
4:28:02clustering here is one more example let
4:28:04us consider a comparison income and
4:28:06balance so in my x- Axis I've got the
4:28:08gross monthly income and in the Y AIS I
4:28:11have the current balance I want to
4:28:13Cluster my data according to these two
4:28:17attributes here if you see this is my
4:28:20first cluster and this is my second
4:28:23cluster so this uh is the cluster that
4:28:26indicates the people who have high
4:28:28income and low balance in uh the account
4:28:31and they spent a lot and this cluster
4:28:33comprises of the people who have got a
4:28:35low income but a high balance and they
4:28:38are safe you can see that all the points
4:28:41that are lying here have got similar
4:28:43characteristics that they have got low
4:28:45income and high balance and here are the
4:28:47people who share the same
4:28:49characteristics where they have uh got
4:28:51low balance and high income and there
4:28:55are a few outliers here and there but
4:28:57they don't uh form a cluster so this is
4:29:01an example of K means clustering and
4:29:02we'll be using that in order to solve
4:29:04our problems so does anybody have any
4:29:07questions so here is one more example
4:29:10and one more problem for you so you guys
4:29:12will tell me now so the problem is that
4:29:15I want to set up schools in my city and
4:29:17these are the points which indicate
4:29:20where each student lives so my question
4:29:23to you is where should I be building my
4:29:25school if I have students living um
4:29:28around the city in these particular
4:29:30locations and in order to find that out
4:29:33we will do K means clustering and we'll
4:29:35find out the center point right so if
4:29:38you can cluster and make groups of all
4:29:40these locations and set up schools at
4:29:42the center point of each cluster that
4:29:44would be Optima isn't it because that is
4:29:47how the students have to travel less it
4:29:50will be close to everyone's house and
4:29:52there it is so we have formed three
4:29:55clusters so you can see the brown dots
4:29:57are one cluster and the blue dots are
4:29:59one cluster and the red dots are one
4:30:01cluster and we have uh set up schools in
4:30:04the center points of each cluster so
4:30:07here is one here is one and here is yet
4:30:10another one so this is where I need to
4:30:12set my schools up so that my students do
4:30:15not have to travel that much so that was
4:30:17all about c means and now let's talk
4:30:19about Apache Zeppelin this is a web page
4:30:22notebook which brings in data ingestion
4:30:24data exploration visualization sharing
4:30:27and collaboration features to Hadoop and
4:30:29Spark so remember when I showed you my
4:30:32Zeppelin notebook you can see that we
4:30:34have written the code code there we have
4:30:37even run SQL codes there and we have
4:30:40more visualizations by executing code
4:30:42there so this is how interactive
4:30:45Zeppelin is and it supports many many
4:30:47interpreters and it is a very powerful
4:30:50visualization tool that can use uh that
4:30:54goes very well with Linux systems and it
4:30:56supports a lot of language interpreters
4:30:59it supports R python and a lot of other
4:31:03interpreters so now let's move on on to
4:31:05the solution of the use case so this is
4:31:07what you've been waiting for first we
4:31:09will solve our us County solution so the
4:31:11first thing we will do is we will store
4:31:13the data into hdfs and then we will
4:31:16analyze the data by using Scala spark
4:31:18SQL and Spark ml Li and then uh finally
4:31:22we'll find out the results and visualize
4:31:24them using Zeppelin so this was the
4:31:27entire us election solution strategy
4:31:29that I told you and I don't think I
4:31:31should repeat it again but if you want
4:31:32me I can uh should I repeat
4:31:36all right so most of the people are
4:31:38saying no so I will go right through
4:31:39this one again so let me just go to my
4:31:41VM and execute this for you so this is
4:31:45my Zeppelin and I opened my notebook
4:31:47here and let us go to my us election
4:31:49notebook and this is the code so first
4:31:52of all what I'm going to do is that I am
4:31:54importing certain packages because I'll
4:31:56be using certain functions that are in
4:31:58those packages so I've imported spark
4:32:00SQL packages and I have also imported
4:32:03spark ml lib packages because I'll be
4:32:06using K means clustering so Vector
4:32:09assembler enables me certain machine
4:32:11learning functions so over here I have
4:32:13the vector assembler package that gives
4:32:15me certain machine learning functions
4:32:16that I'm going to use I've also imported
4:32:19K means package because I'll be using K
4:32:20means clustering then the first thing
4:32:22that you need to do is that you need to
4:32:24start the SQL context so I have started
4:32:27my spark SQL context here and the next
4:32:29thing that you need to do is that you
4:32:30need to define a schema because when you
4:32:33want to dump our data set or we want to
4:32:35to dump our data it should be in a
4:32:37particular format and we have to tell
4:32:38spark in which format it should be so
4:32:41we're defining a schema here so let me
4:32:44take you uh through the code so I'm
4:32:47storing schema in a variable called
4:32:49schema and we have to define the schema
4:32:52in a proper structure so we're going to
4:32:54start with struct type and since you
4:32:55know that our data set has got different
4:32:57fields as columns we're going to Define
4:33:00this as an array of fields then this is
4:33:02an array instruct so we are defining the
4:33:05different Fields now so we'll start with
4:33:07the first field by defining it as struct
4:33:09field inside the braces which should
4:33:10mention what would should be the name of
4:33:13that particular field so I've named it
4:33:15as state it should be a string type and
4:33:18true that means it is a string type the
4:33:21next we've got fips which is of string
4:33:23type now I know that fips is a number
4:33:25but since we are not going to do any
4:33:26kind of numeric operation on fips uh
4:33:29we're going to let it stay as a string
4:33:31then we've got party as a string type
4:33:33candidate as a string type and then
4:33:34votes as integer type because we're
4:33:36going to count the number of votes and
4:33:38there is going to be certain numeric
4:33:40operations that we are going to perform
4:33:42that will help us to analyze our data
4:33:45then we've got a fraction votes which
4:33:47you know is a decimal type so we have to
4:33:49keep it as double type the next thing
4:33:51you need to do is that spark needs to
4:33:53read the data set from the
4:33:55hdfs so for that you have to use the
4:33:58command spark read option header true
4:34:01header true means that you have
4:34:03mentioned and you have told spark that
4:34:05my data set already contains column
4:34:07headers because State as ABR they are
4:34:10nothing but they are column headers so
4:34:12you don't have to explicitly Define the
4:34:14column headers uh for it neither will
4:34:17spark choose any random row as a column
4:34:19header so it will choose only the column
4:34:21headers uh your data set has then you
4:34:24have to mention the schema that you have
4:34:26defined so I have defined it in my
4:34:28variable schema so that's why I have
4:34:30mentioned it in my file should be in CSV
4:34:33format and then I have mentioned the
4:34:35path of the file in my hdfs this is the
4:34:38path and I store this entire data set in
4:34:41my variable
4:34:42DF now what I am going to do is that I'm
4:34:46going to divide up certain rows from my
4:34:48data set because you know that my data
4:34:50set contains both the Republican and
4:34:52Democrat data and I just want the
4:34:54Democrat data right because we're going
4:34:55to analyze the Hillary Clinton and
4:34:57Bernie Sanders part okay so this is how
4:34:59you divide your data set so the first
4:35:01thing that we have done is that we have
4:35:03created one more variable called DF far
4:35:05and we have replied A filter where party
4:35:07is equal to Republican and then we are
4:35:10storing the Democrat Party data into DFD
4:35:13so we're going to use the DFD from
4:35:16onwards and dfr the Republican data is
4:35:19going to be your assignment for the next
4:35:21class now I am going to analyze the
4:35:23Democrat data and then after this class
4:35:26is over I want you guys to take the
4:35:28Republican data this data set is already
4:35:31available in your elements and you've
4:35:33got the VMS also with everything
4:35:35everything installed so please when you
4:35:37are at home when you have free time just
4:35:39analyze the Republican data and tell me
4:35:41uh what were the reasons that Donald
4:35:43Trump want I want you to do all that
4:35:45analysis and come up with that in the
4:35:47next class and we'll discuss about it
4:35:49and whatever results and conclusions
4:35:52that you have made after analyzing the
4:35:54Republican data and that way you'll also
4:35:56learn even more and it will also be
4:35:59practice for you after today's class so
4:36:01all right so we are going to take DFD
4:36:04now in the first thing that we will do
4:36:06is that we will create a table View and
4:36:07I'm going to name the table view as
4:36:09election and let me just show you what
4:36:11it looks like and what it has so this is
4:36:14the command that I have run in Zeppelin
4:36:16so this is SQL code that I have run in
4:36:20Zeppelin and you can see that I have got
4:36:23States state abbr and I have only got
4:36:26the Democrat data all right let's go
4:36:29back all right so after creating the
4:36:32table view now all of the Democrat data
4:36:34is in in my election table so now what
4:36:37I'm going to do is that I'm creating a
4:36:39temporary variable and I'm running spark
4:36:41SQL code so what I'm actually doing by
4:36:43writing this code the motive of writing
4:36:45this SQL code or the SQL query is that I
4:36:48want to refine my data even more so what
4:36:50I'm trying to analyze here is how a
4:36:52particular candidate actually won I
4:36:53don't have to do anything with the
4:36:54losing data because you know that each
4:36:56of fips contain one of the losing
4:36:58candidate members and one of the winning
4:37:00candidate members it contains the data
4:37:02of the winning candidate and the losing
4:37:03candidate also because my data set
4:37:05contains both of the data of Bernie
4:37:07Sanders and Hillary Clinton in some
4:37:09parts Bernie Sanders won and in some
4:37:11counties Hillary Clinton won so I just
4:37:14want to find out uh that who are the
4:37:16winners in a particular County okay so
4:37:19I'm going to refine that data and for
4:37:21that I'm using this query so I'm going
4:37:24to uh select all from election and then
4:37:26I'm going to perform an inner join with
4:37:29their query so this is one more query
4:37:31inside this query and let me tell you
4:37:33what I'm actually doing so first of all
4:37:36what we have done is that we have
4:37:37selected fips as B you know that now you
4:37:40have got two entries for each fips so
4:37:43each fips actually appears twice in the
4:37:45data set so I've named it as B and now
4:37:47we are counting the maximum fraction
4:37:49votes so you know that in each FIP we
4:37:52have the maximum fraction vote and then
4:37:54we can find the winner by actually
4:37:55seeing who has got the maximum fraction
4:37:57votes then we have named it as a the
4:38:00maximum fraction votes column is named
4:38:02as a and we are grouping by Fifth
4:38:06so now each of my fips will be selected
4:38:08which has a maximum fraction vote and I
4:38:11have uh two columns for that fips which
4:38:13is one1 and
4:38:15one1 so the only rle will be selected
4:38:18which has the maximum fraction votes now
4:38:21I'll have the winner data and I've named
4:38:23this entire table inside this query as
4:38:26group TT and then I'm validating it as
4:38:29where election. fips the main table
4:38:31view. fips should be equal to the B
4:38:34column that we have created in group TT
4:38:36table and election. fraction votes
4:38:38should be equal to group tt. a so any
4:38:42doubts on this query and about how I
4:38:44have written this all right so now what
4:38:47we're going to do is that whatever data
4:38:49that we've got here I'm storing that in
4:38:51election one let me just show you what
4:38:54is in election one now so this is my
4:38:57election table only and uh you can see
4:38:59that I've got two fips so 1067
4:39:031067 now let me show you election one so
4:39:06there now I can see that I don't have
4:39:09repetition of fips I have only one entry
4:39:13for fips and that is the row which tells
4:39:16me who won in that county or in that
4:39:19particular FIP or in the FIP associated
4:39:21with a particular County you can see for
4:39:23Bullock it was Hillary Clinton for kahun
4:39:25it was Hillary Clinton Cherokee also
4:39:27Hillary Clinton and then state house
4:39:29district 19 is Bernie Sanders so Alaska
4:39:32is mainly Bernie Sanders so this is what
4:39:36we've done now and then you can see that
4:39:38we have also got additional columns as b
4:39:42and a so a tells you the maximum
4:39:44fraction votes and B tells you the
4:39:47fips so the data in fips and the data in
4:39:50B are the same and data in fraction
4:39:52votes and data in a is the same right
4:39:55what I'm going to do now is since my
4:39:57columns are repeating and they have the
4:39:59same value I Don't Want A and B now
4:40:02right so what I'm going to do is I'm
4:40:04going to filter out the columns I don't
4:40:06need and in this case I don't want b and
4:40:08a and what I'm going to do is I'm going
4:40:10to make a temporary variable again so
4:40:12I'm using the temporary variable to
4:40:14store some data temporarily so I'm
4:40:16writing to the spark SQL code uh to
4:40:19select Only The Columns that I want I
4:40:21want the state state abbreviation County
4:40:24fips Party candidate votes fashion votes
4:40:26from election one I'm storing everything
4:40:28in D winner I've created this new
4:40:30variable and whatever there was in temp
4:40:33I'm assigning it to D winner and now
4:40:35I've uh got only the winner data so I
4:40:38have got all the counties and I've got
4:40:40uh who won in that particular County and
4:40:41by how much in the fraction of votes
4:40:43what I'm just doing till now is that I'm
4:40:46just refining our data set so that it
4:40:48will be easy for us to make some
4:40:49conclusions or gain some insights from
4:40:52that data right and also let me tell you
4:40:54that it's not always necessary that you
4:40:56are filing your data set in the exact
4:40:59way that I'm doing it if you have
4:41:00something in mind after you've seen your
4:41:02data and understand your data and you
4:41:03found out what actually you need to do
4:41:06you can carry out different steps to do
4:41:08that also this is just one way of doing
4:41:09it and this is my way of doing it so I'm
4:41:11just telling you and then we are
4:41:14creating a table for D winner and we are
4:41:16going to name it as Democrat so let me
4:41:18go again and let me show you what the
4:41:20Democrat table view looks like you can
4:41:23press shift enter so there you uh have
4:41:27we have column A and B that we had in
4:41:29election
4:41:32one and so I have just got uh winner
4:41:37data so now let us go back and find out
4:41:40what we're going to find is that I want
4:41:43to find out that uh which of the
4:41:46candidates won my state and then
4:41:48whatever date and whatever result I'll
4:41:50get will be stored in the temporary
4:41:52variable when I'm assigning everything
4:41:55uh that will be stored in the temporary
4:41:57variable to a new variable called dstate
4:41:59and then similarly I'm going to create a
4:42:02table view for dstate which is State let
4:42:06me show you what my state table view
4:42:08actually contains so there it is so I've
4:42:12got State Connecticut Hillary Clinton W
4:42:14155 counties Florida Hillary Clinton won
4:42:1658 counties so this is what we've come
4:42:18up to for our first data set so now
4:42:22let's see what we can do with our second
4:42:23data set that contains all the different
4:42:25demographic
4:42:26features uh first thing again you have
4:42:29to define a schema and this time I'm
4:42:31naming that schema uh schema one say
4:42:34since you know that we have got almost
4:42:3754 columns so I have to Define all those
4:42:4054 columns
4:42:41also so you remember what th each of
4:42:44those columns contains so this is
4:42:47exactly what I have done and I don't
4:42:49need to go through every line but I like
4:42:51I already told you how to define a
4:42:53schema you can have the code in your LMS
4:42:56so you can take a look at it so the next
4:42:58thing we're doing again we have to read
4:43:00our data set and I'm storing my data set
4:43:02into a new variable called df1 and this
4:43:05is the path in my htfs where my data set
4:43:07was and then I have created a table view
4:43:09for my data set which is called
4:43:13facts now let me show you what facts
4:43:19contain as you can see that it contains
4:43:23abbreviation state abbreviation
4:43:24population 2014 so instead of uh using
4:43:30the code now or the encoded form that
4:43:33was actually there in my data set I have
4:43:35given a varied metan name that would
4:43:38describe what it contains right so
4:43:40instead of PST 214 I've got population
4:43:432014 so does that make sense right and
4:43:47contains all the 54 demographic features
4:43:49or different features that was in my
4:43:51data
4:43:53set white alone not Hispanic or Latino
4:43:57living in the same house one year and
4:43:59over foreign born persons language or
4:44:01other than English spoken at home High
4:44:03School gradu or higher uh so it contains
4:44:06basically all the different features or
4:44:08all the different columns that actually
4:44:09was in my data set and that I have
4:44:11defined in my schema so this is what
4:44:14facts I have so now what I'm going to do
4:44:16is that I'm not going to analyze my
4:44:18whole data based on all this different
4:44:20features I'm going to choose some
4:44:22specific features in order to analyze it
4:44:26uh I'm going to take just a few add one
4:44:29so these are the different features that
4:44:32I'm going to use I'm going to use fips
4:44:34I'm going to use state I'm going to use
4:44:35state abbreviation then area name
4:44:38candidate and people who are over 65
4:44:40years senior citizens a female people
4:44:43white Alone um black African alone I'm
4:44:47choosing Asian alone Hispanic or Latino
4:44:50basically what I'm trying to do is I'm
4:44:51trying to check what is the popularity
4:44:53of Hillary Clinton among the foreign
4:44:55people or people from different
4:44:57ethnicities so I'm choosing white people
4:44:59black people and Hispanic people so I'm
4:45:01just trying to analyze it okay and you
4:45:04know that I have stored this in a
4:45:05temporary variable again and then
4:45:07whatever result I'll get by running this
4:45:09spark SQL code I'll store it in a
4:45:12different variable called DFX and then
4:45:14I'll store it and then I'll make a table
4:45:16view for DF facts such as winter
4:45:20facts so let me show you what winter
4:45:23facts look like so it's winter facts
4:45:26you've got uh fips the state is Alabama
4:45:28state abbreviation is Al for
4:45:32Alabama um the area name is our tuga
4:45:37County and the winner was Hillary
4:45:39Clinton and the people over 65 years in
4:45:42that particular county is 13.8% female
4:45:45percentage is 51.4 white alone 77.9 and
4:45:49so these show you the data
4:45:52so uh black white or African is 18% and
4:45:57then I've got the different fields that
4:45:59I have selected Asian alone Hispanic or
4:46:02Latino foreign born so so I have chosen
4:46:0514 features to analyze it from all right
4:46:08so now what I'm doing again is that I'm
4:46:10going to divide the Hillary Clinton data
4:46:11and the Bernie Sanders data so that we
4:46:13can analyze only why Hillary Clinton won
4:46:16or why Bernie Sanders won in some
4:46:18particular counties so we are planning
4:46:21to filter the same way we divided
4:46:23Democrats and Republican data from our
4:46:25initial primary result data set so this
4:46:29is what you have done so you know that
4:46:31is stored in DF fact so we are putting
4:46:33the filter in DFX where uh candidate is
4:46:37equal to Hillary Clinton so that will be
4:46:39stored in HC and the data of Bernie
4:46:41Sanders will be stored in BS so after
4:46:44that what we are doing is that we are
4:46:46doing a one hot encoding so we'll add
4:46:49two more columns in our data set a WBS
4:46:51and
4:46:52whc in this case we are going to do one
4:46:56hot encoding and what we're going to do
4:46:59is that we are going to include or we
4:47:02are going to attach two more columns in
4:47:04Winter facts as whhc and WBS so it'll
4:47:10just contain either one or zero and so
4:47:14you can edit it in that way whichever
4:47:15County so if you're considering a county
4:47:18let's say aruga County say if Hillary
4:47:20Clinton is the winner it will have a one
4:47:22in whc and in WBS it will have zero
4:47:27similarly in which Count's Bernie
4:47:28Sanders one so Bernie Sanders will have
4:47:30one so WBS will have one and whc will
4:47:33have have a
4:47:35zero and then we are creating different
4:47:38views for both of these two together so
4:47:40this will only tell me wherever whc is
4:47:43one that means this will only show me
4:47:45the counties where Hillary Clinton won
4:47:47this will only show me the counties
4:47:48where Bernie Sanders won and we are
4:47:51creating a view for both of these so for
4:47:53Bernie Sanders the view is WBS and for
4:47:56Hillary Clinton it's whc then finally we
4:47:58are merging both of them together using
4:48:00Union so select all from whc Union all
4:48:03select from WBS and finally you have
4:48:05stored it in result and we have created
4:48:07a table view known as result so let me
4:48:10show you what this result
4:48:12contains uh so there it is for UGA it
4:48:15was Hillary Clinton so we've got the
4:48:17Bernie Sanders
4:48:19data over here at the bottom and I've
4:48:22got all the different fields also from
4:48:24my uh second data
4:48:28set the different features that I chose
4:48:31from my second data set uh to to analyze
4:48:34it so now comes the actual analyzing
4:48:36part this is where we're going to
4:48:37perform K means but first we have to
4:48:40define the feature columns actually you
4:48:42have to Define what is the input that
4:48:45you're going to feed so that you get an
4:48:48output so this is actually the input
4:48:50that you are going to feed the to the
4:48:52machine so that the machine learning
4:48:54goes on and finally it gives you some
4:48:56kind of result right so this is where
4:48:58I'm defining again I'm using an array to
4:49:00Define all the different fields from my
4:49:03data sets I'm using person 65 years and
4:49:05older female person percentage white
4:49:07alone or black or African uh American
4:49:10alone Asian alone Hispanic or Latino
4:49:13foreign born persons language other than
4:49:17English spoken at home bachelor degree
4:49:19or higher veterans home ownership rate
4:49:21median household uh income persons below
4:49:24poverty level and population per square
4:49:26mile whc and WBS and then I'm going to
4:49:30use the vector assembler so this is what
4:49:32enables different machines learning
4:49:34algorithms where we are using K means so
4:49:38my input column is features calls so
4:49:41this is going to be the input in my
4:49:43output column and will be called
4:49:46features so whatever result that I'm
4:49:49going to get is features and we have to
4:49:51transform the results so this is the
4:49:54final table view that we have created
4:49:56and you know what transforming means and
4:49:57transforming again means so in our
4:50:00strategy we already saw that we have to
4:50:01transform the data first so my updated
4:50:04data set was results so I'm going to
4:50:06transform result and put columns as
4:50:08going to be these which is feature
4:50:11columns and output uh table view will be
4:50:14called features and then we're going to
4:50:16perform the K means clustering and we're
4:50:18going to store it in a variable called K
4:50:20means so we're using different functions
4:50:22from spark
4:50:23M library and we have chosen spark with
4:50:27uh clustering K means and you know that
4:50:30in K means we already defined that how
4:50:32many clusters do we need and we need
4:50:35four so we have selected four clusters
4:50:37and then we are going to set feature
4:50:39columns as features and then set
4:50:41prediction column as
4:50:43predictions so after that we're going to
4:50:46make a model and we have defined our
4:50:48input and output columns in row so we're
4:50:50going to use uh kefit row and whatever
4:50:53predictions we will get we're going to
4:50:55store it in a model and then we um are
4:50:59going to do this and that we are going
4:51:01to print the cluster centers for each
4:51:04cluster so let me show you what my
4:51:06cluster centers are so after we run this
4:51:10code you can see that these are the
4:51:12different cluster centers know so just
4:51:13what I can make you understand about
4:51:15what we're going to do after K means
4:51:17clustering and how to analyze it so the
4:51:19numbers are present they are placed very
4:51:22haphazardly so what I have done is that
4:51:25I've picked out each of the cluster
4:51:26Center points and then I have made a new
4:51:29table yes so this is it so you know that
4:51:32we have four clusters we have got the
4:51:34zero cluster first cluster second
4:51:36cluster and uh uh third uh so 0 two 3
4:51:42okay so four clusters and we have found
4:51:44out this uh cluster centers according to
4:51:47different features that we fed into my K
4:51:50means algorithm so what we observed here
4:51:54in whc and WBS is that the winning
4:51:56percentage or the winning chances of
4:51:58Hillary Clinton was 0.9 whereas winning
4:52:00chances for Burnie Sanders was
4:52:020.1 and and uh
4:52:06then if You observe the differences in
4:52:09the cluster centers for each feature
4:52:11here you can see that there is not much
4:52:13difference not even here so it's uh 50
4:52:1749 49 51 and then uh it's well again it
4:52:22is not much of a difference but if you
4:52:24see here that it's nine and it's uh
4:52:26going to 16 so you can do a more
4:52:29detailed analysis on black or
4:52:32African-American so if you want want to
4:52:33know the real support of black or
4:52:36African-American and you want to see uh
4:52:40what was their voting pattern or how
4:52:43popular was Hillary Clinton among them
4:52:46so maybe this could be a good field to
4:52:48analyze because you see the variations
4:52:50in the number similarly you can check
4:52:52out other features and you can check out
4:52:55uh here at 16 89 and 36 so maybe again
4:52:58Hispanic or Latino field and you should
4:53:01uh do some more analysis on it and even
4:53:03here you can see in veterans there's
4:53:0447,800 whereas we've got 182,000 all so
4:53:10there is uh also a lot of
4:53:12difference um then here is only 20 2759
4:53:17and we've got uh in the 10,000s we've
4:53:20got numbers and even 100,000s here so
4:53:23this is how we can identify that which
4:53:25are the fields or which we can find the
4:53:26main reasons of the main points where
4:53:28you should make your analysis so let's
4:53:30go back to our Zeppelin notebook and
4:53:33here it is is so now what we're going to
4:53:36do is that we're going to visualize the
4:53:37result first so we are counting from
4:53:41predictions so you can see that in
4:53:43cluster ones the prediction means
4:53:45prediction contains my cluster since you
4:53:47know that I've have stored my clusters
4:53:49my cluster information and prediction
4:53:51this is my output after K means uh so
4:53:55I've got uh these this many clusters so
4:53:58this is the count of my counties or a
4:54:00count of different various that lie in
4:54:02my particular cluster you can see that
4:54:05in cluster one I have got 1917 and
4:54:08cluster 2 I've got 751 so maybe I should
4:54:11pay more attention on analyzing cluster
4:54:13one right uh so that's why I've SE
4:54:16selected cluster one here and we're
4:54:18making different predictions so you can
4:54:20see that in the x axis I have got foure
4:54:23inborn people and in y AIS I have
4:54:26chosen uh language other than English
4:54:29spoken at home and then we are grouping
4:54:32it by candidate so you can see the
4:54:34lighter blue is for Bernie Sanders and
4:54:36the more dark blue is for Hillary
4:54:38Clinton so all this light blue is for
4:54:40Bernie Sanders and you can see that as
4:54:43the number of foreign people increases
4:54:46you can only see Hillary Clinton in the
4:54:48scattered plot here so there might be a
4:54:50few outliers like back here in the size
4:54:52of defined according to black or
4:54:54africanamerican alone you so you
4:54:57remember that this was the feature where
4:54:59we find a lot of variations in the
4:55:01numbers so that's why we grouped it
4:55:02according to that
4:55:04and you can see the bigger the circle
4:55:06represents the more black or
4:55:08African-American alone and that's what
4:55:11what the conclusion we can find out from
4:55:12this scatter plot and we can see that as
4:55:14the number of foreign people increases
4:55:16the popularity of Hillary Clinton is uh
4:55:20more in larger groups of foreign people
4:55:23you can also choose different parameters
4:55:25out of all the different features that
4:55:26you have chosen so remember uh that we
4:55:29have also seen the variation in veterans
4:55:31so let's choose veterans in a y axis so
4:55:33let's also change xaxis and let me just
4:55:35use white alone here so you can see here
4:55:38that
4:55:39uh uh there is the x axis that has white
4:55:45alone and this is the Veterans so you
4:55:47can see that Hillary Clinton is popular
4:55:49among veterans also in a smaller group
4:55:52of veterans since we have decided the
4:55:54size in black or africanamerican alone
4:55:57so the size um also represents some
4:56:01values she is popular among the
4:56:03africanamerican veterans and then uh as
4:56:07you go ahead and as the count increases
4:56:09you can see actually since it's a
4:56:11scatter plot and it almost represents
4:56:14that uh this is a point as the number of
4:56:17people increases or as the number of
4:56:20white people increases the votes are
4:56:21equally kind of distributed between
4:56:24Bernie Sanders and Hillary Clinton
4:56:26because there are a lot of points in
4:56:27this scatter plot over here and you can
4:56:30go ahead and drag and drop different
4:56:31features and you can make different
4:56:33visualizations on that now what we've
4:56:36done is that we know that there are 1917
4:56:39counties in my cluster one so I'm am
4:56:41going to do is that I'm going to see
4:56:44that among these
4:56:451917 how many were in favor for Hillary
4:56:48Clinton and how many were in favor for
4:56:50Bernie
4:56:51Sanders
4:56:53um so in cluster number one you can see
4:56:56clearly Hillary Clinton is the winner
4:56:58and Bernie Sanders only has got 764
4:57:01whereas she got 11
4:57:04153 similarly in cluster 3 again Hillary
4:57:07Clinton is the winner with nine and
4:57:08Bernie Sanders uh with
4:57:12one then it uh two she's also got 388
4:57:17and Bernie Sanders was 363 so this was
4:57:20very close call and again in zero you've
4:57:22got 119 and
4:57:2530 and then we went ahead and create a
4:57:27line chart also of the word distribution
4:57:29for Hillary Clinton and Bernie Sanders
4:57:31so in Keys we have selected prediction
4:57:34the values here are whc and WBS the sum
4:57:37that we have got over here so definitely
4:57:40Bernie Sanders is lagging behind so even
4:57:42though you don't have that table for you
4:57:44you can also find it out according to
4:57:47this line chart so you can see that in
4:57:49cluster zero even again Hillary Clinton
4:57:51was ahead of Bernie Sanders in cluster 2
4:57:55there was a very neck to neck
4:57:56competition and you can see it in the
4:57:58graph year so this res represents
4:58:01cluster 2 and so you can see you have a
4:58:04neck to neck competition and again in
4:58:07cluster
4:58:08three uh they have got neck to neck
4:58:10competition so this describes the
4:58:13distribution of votes for Hillary
4:58:14Clinton and Bernie Sanders and
4:58:16definitely Hillary Clinton uh knows
4:58:20ahead and that's why of course she won
4:58:22the primary elections so again you can
4:58:24go ahead and we have created the same
4:58:27graph it's uh only just area graph
4:58:29instead of a line graph the key here are
4:58:33are State and candidates so I've got
4:58:35States and candidates over here and the
4:58:38values is counties once uh if you just
4:58:41hover onto this bar chart you can see
4:58:43that in Connecticut Bernie Sanders won
4:58:46115 counties in Connecticut Hillary
4:58:49Clinton won 55 only so in Florida
4:58:52Hillary Clinton is 58 and in Florida
4:58:54Bernie Sanders is nine and here you can
4:58:56see in uh the main Bernie Sanders won
4:58:59462 so Bernie Sanders got a majority of
4:59:02votes for
4:59:04Maine so you can also classify it
4:59:07statewise you can find out which uh are
4:59:10the states and as Donald Trump now you
4:59:13will know that which are the states that
4:59:15you can Target right so you know that in
4:59:17Maine a lot of people voted for Bernie
4:59:19Sanders and maybe Hillary uh Clinton is
4:59:23not popular so you can go ahead and lead
4:59:25out so as Donald Trump's party member
4:59:28you can just advise him to go in Maine
4:59:31and carry out different campaigns
4:59:33because uh Hillary Clinton is not so
4:59:36popular there so maybe it would be a
4:59:38little easier to get votes from the
4:59:40people in Maine so this is what you can
4:59:43make a conclusion from it might not be
4:59:45very accurate but this would be very
4:59:46close the thing is that you can make
4:59:48different charts you can make bar charts
4:59:50you can make pie charts so whatever
4:59:52counties won have made in the bar chart
4:59:54so they're here is in a pie chart it
4:59:56looks better but it's not maybe as
4:59:58insightful um I just placed it uh so
5:00:01that I can show you you can make pie
5:00:02chart s also so these are the insights
5:00:05that you can make after analyzing your
5:00:06us County data and this is what you can
5:00:08tell Donald Trump these are the
5:00:10different suggestions that you can
5:00:11actually go and tell Donald Trump uh
5:00:14that she is popular among the foreign
5:00:16people and the people who speak
5:00:18different languages she is popular among
5:00:20the Hispanic people in then in Maine she
5:00:23lost a lot of counties she almost lost
5:00:25all of the counties in Maine so these
5:00:27are different insights that you uh have
5:00:30got and then you can tell your Superior
5:00:33or your employer who has hired you to do
5:00:36that um so this is what you can present
5:00:38right so this is for a very beginner's
5:00:41level and there are some more analytics
5:00:43that you need to do I just showed you a
5:00:45few options you can go ahead and try
5:00:47more in the Democrats section also and
5:00:49you remember that uh you have to do it
5:00:52for the Republican party also now let me
5:00:55see what youve learned today so if you
5:00:57have any questions right now you can
5:00:59just go ahead and ask me so does anyone
5:01:01have any questions
5:01:04so now we will move on and find out the
5:01:05solution for the instant cab use case
5:01:08you remember that we have got the Uber
5:01:10data set which contains the pickup time
5:01:12and the location by two columns latitude
5:01:14and longitude and we have uh also got
5:01:17the license number for a uh particular
5:01:21Uber driver and what we have to do is
5:01:24that we have to find the Beehive
5:01:26locations uh that is the point where we
5:01:29will find the maximum pickups and then
5:01:30we will also have to find out what is
5:01:33the peak hour of the day so this was the
5:01:35entire strategy so we've got the Uber
5:01:38pickup data set and then we store the
5:01:40data into hdfs we will transform the
5:01:43data set and make predictions by using K
5:01:45means clustering on the latitude and
5:01:47longitude and find out the b Point uh or
5:01:50beehive point so now let me open my
5:01:53other notebook the Uber notebook so
5:01:55again the first thing that you have to
5:01:56do is copy the Uber data set into your
5:01:59hdfs now we've done that before
5:02:01explaining to you the US County analysis
5:02:04so again the code is kind of the same
5:02:05the first thing is that again we are
5:02:07importing some spark SQL packages and
5:02:09some spark ml lib packages because we
5:02:12are going to use K means clustering and
5:02:15you can see Vector assembler here again
5:02:16spark ml clustering K means and other
5:02:19spark SQL packages so then we have to
5:02:22start our SQL context and we're doing it
5:02:25same way than the first thing again we
5:02:27have to define a schema now I don't have
5:02:29many fields I've got only four Fields if
5:02:31I remember so the first uh field was the
5:02:34date and time stamp that defines the
5:02:36time we're defining it as DT and next
5:02:39field is the latitude the longitude and
5:02:41base then I'm going to read my data set
5:02:44this is the path in my htfs where my
5:02:46Uber data set is there so I Define
5:02:50schema as schema here the header is true
5:02:53because again my data set contains
5:02:55column headers and I'm going to store in
5:02:57DF so feature calls uh here is going to
5:03:01be latitude longitude because I'm going
5:03:03to find out the Beehive point the point
5:03:07where I will get my maximum kick up from
5:03:09so again I have set the input calls as
5:03:14feature calls and output calls as
5:03:16features so I'm using the assembler to
5:03:19transform my data set and then again I'm
5:03:20using K means and we use the same elbow
5:03:23method we found out that we should make
5:03:24eight clusters for this data set okay
5:03:28and then we are selecting the prediction
5:03:29column and the output column and as
5:03:33predictions and then we have printed the
5:03:35cluster centers for each cluster so
5:03:38definitely whatever result we are going
5:03:40to find the cluster centers will tell me
5:03:42the exact location so this cluster uh
5:03:45centers that we will find after c means
5:03:47is actually the Beehive points this will
5:03:50be the point where I will find maximum
5:03:52pickups
5:03:54right so here I have printed my cluster
5:03:57centers and this defines the latitude
5:03:59and longitude and this is going to be my
5:04:02location where I'm going to find the
5:04:03maximum pickups and I got eight results
5:04:06uh like that because I got eight
5:04:07clusters and uh define the eight centers
5:04:10for different clusters so this is
5:04:13exactly like the K School problem that I
5:04:15explained to you in K means this is
5:04:16exactly what happens just as we found
5:04:18out the center of each cluster and that
5:04:20is where we are replacing the school or
5:04:22building the new school so similarly
5:04:24this is going to be my beehive point and
5:04:27this is where I will place my maximum
5:04:29number of cabs okay so we found out the
5:04:32be Hive points the next thing we will
5:04:34need to do is we need to find the peak
5:04:35hours because I also need to know at
5:04:38what time should I place my cabs in the
5:04:40location so what we're doing now we are
5:04:43taking a new variable called q and we
5:04:45are selecting hour from the timestamp
5:04:47column and then the Alias name should be
5:04:50our and we're getting it from our
5:04:52prediction or from the result that we
5:04:53got after my K means clustering so now
5:04:56we are grouping it and it will have the
5:04:58different hours of the day and then it
5:04:59will just show me the pickups at the
5:05:02different hours of the day in the
5:05:04location that we found out are the
5:05:07Beehive points and then we're going to
5:05:10count uh how many pickups we are going
5:05:13to get from that place right so we're
5:05:15ordering it by descending so the smaller
5:05:17pickup count will be the the first and
5:05:19then the larger will be at the bottom
5:05:21similarly again we are creating new
5:05:23variable called T and we're going to do
5:05:25the same thing so here what we're doing
5:05:28is we are selecting the time hour of the
5:05:30date the latitude longitude prediction
5:05:32and and we filter by hour which is not
5:05:34null so we're filtering out the null
5:05:37values from here so now we have created
5:05:39a table view for categories so let me
5:05:41show you what the categories contain
5:05:44okay let me just go down so I've done
5:05:46some few operations here so let's scroll
5:05:49back up and I'll show you and again we
5:05:51have created table views for T and Q
5:05:54also which is again T and Q all right
5:05:56and then I have made some visualizations
5:05:59for each so then we uh have created a
5:06:03value P where hour is not null so again
5:06:05we have filtered out the null hours and
5:06:07we have created a new view called P so
5:06:11here is my hours this is my count and in
5:06:15the x-axis that show how many pickups
5:06:16were there and this contains different
5:06:18hours of the day and then I have grouped
5:06:20it by prediction so the size is
5:06:23according to the count so you can see
5:06:25that the bigger the circle means more
5:06:27pickups so you can find out the biggest
5:06:30circle and you know that you can find
5:06:31the biggest Circle as you go along the
5:06:33x-axis because this is where the count
5:06:36increases so you can find out the
5:06:38biggest circle would be here and it lies
5:06:40in my fourth cluster and you can see
5:06:43that there are 800 or 8,915 pickups at
5:06:46the 17th hour of the day which is around
5:06:495:00 p.m. and so you know that the
5:06:50maximum pickups are around 4:00 or 5:00
5:06:53and this lies all in my fourth cluster
5:06:57and so it means my peak hours are around
5:06:594 or 5:00 in the evening right so this
5:07:01is what Insight we have gained and you
5:07:04can tell instant cab CEO that I have
5:07:06found out that your cabs should be ready
5:07:08around four or five because that's the
5:07:10time when uh people go home from offices
5:07:14or they're going out for dinner or
5:07:16something and this is what another table
5:07:18view looks like which is T so here we
5:07:21have latitude and longitude and this is
5:07:23where we are finding the Beehive
5:07:25locations so I have uh got this the
5:07:29distribution in a scatter plot again and
5:07:31you can see see that we have got uh very
5:07:33dense points over here it means that
5:07:36these represent the Beehive points so
5:07:39what you can do is that you can just put
5:07:40the US map and scale it according to
5:07:42this scale over here and then you can
5:07:44exactly find out what is the exact
5:07:46location where you need to put your cabs
5:07:49around the 17th hour or the 16th hour of
5:07:52the day all
5:07:54right and you know that we had a lot of
5:07:56rows but the results are only limited by
5:07:5910,000 if it's around 10,000 rows but we
5:08:02obviously had a lot more and you can
5:08:04check in different uh clusters so now we
5:08:08are
5:08:09analyzing
5:08:11uh cluster zero so here if you see this
5:08:15point over here this lies in cluster 4
5:08:20this lies in cluster five and this lies
5:08:22in cluster zero so you can analyze each
5:08:24cluster also so here I have just laid
5:08:28out the latitude and longitude for my uh
5:08:30zeroth cluster so you can see here where
5:08:33prediction is equal to zero and I've
5:08:35selected this from the table view of T
5:08:38so here you can find out the exact
5:08:39latitude and longitude and here the
5:08:42latitude is 4.72 two and the longitude
5:08:45is
5:09:01-73.995411 that tells you what is the
5:09:03count of pickups at each hour of the
5:09:05days starting from 0 to 23 there are 24
5:09:09slices in this circle so you can see uh
5:09:12that these few slices are the bigger
5:09:14chunks and this is the 19th hour of the
5:09:16day which is around 7:00 6:00 5:00 4:00
5:09:203:00 and so on so you can see the
5:09:23midnight maybe nobody travels uh so
5:09:26maybe your cabs could rest or you don't
5:09:29have to place any more cabs during this
5:09:31part of the the day um these are the
5:09:34insights that you gain so any questions
5:09:36on that I think after doing the US
5:09:39County election this was pretty easy to
5:09:40do and this is also uh pretty easy to
5:09:44understand and the results which were
5:09:46also much more clear
5:09:48[Music]
5:09:53correct what is actually a Hadoop
5:09:56ecosystem okay the very first thing is
5:09:59Hadoop ecosystem is not one tool it's
5:10:01not a programming language or it's not a
5:10:03single framework it is a group of tools
5:10:06that are there which are used together
5:10:08by various companies in various domains
5:10:10for different tasks okay Hadoop alone
5:10:14cannot provide all the facilities or
5:10:17services that are required to process
5:10:19the Big Data okay so like for example
5:10:22Hadoop can store Big Data Hadoop can
5:10:24process Big Data up to a certain limit
5:10:26however there are much more other
5:10:28requirements that are there for example
5:10:30we would like to create recommendation
5:10:33engines over big data we would like to
5:10:35run clustering algorithms over big data
5:10:38we would like to get the realtime
5:10:39insights using big data itself because
5:10:41Hadoop is a batch processing framework
5:10:43right so if I want a real time Insight I
5:10:46would need another tool that can run
5:10:48over htfs that can utilize and leverage
5:10:50htfs right the basic thing that you need
5:10:53to understand here is one single tool
5:10:55like Hadoop is not going to solve all
5:10:57your problems you'll have to use various
5:11:00other tools over Hadoop or with Hadoop
5:11:02to get rid or get the solution of every
5:11:05problem that you have okay but before
5:11:09that before doing so it is important
5:11:11that you know what are the different
5:11:12tools that are there which can work with
5:11:14Hadoop and in today's session we'll
5:11:16exactly do that we'll try and find out
5:11:18what are the various tools that are
5:11:20there which can be used with Hadoop and
5:11:22what functions they can perform in their
5:11:25own domains we now move on to the next
5:11:30slide okay the very first tool that
5:11:32we'll understand is
5:11:34hdfs now as you know hdfs is nothing but
5:11:38Hadoop distributed file system it is the
5:11:40storage unit of Hadoop sdfs is entirely
5:11:43the Hadoop cluster which is formed by
5:11:45data nodes data nodes are nothing but
5:11:47commodity Hardwares which are cheap
5:11:49Hardwares which can be clustered
5:11:52together using the Hardo framework and
5:11:54then entire file system that gets
5:11:56created on which you can store big data
5:11:58is called Hadoop distributed file system
5:12:00using hdfs you can store any kind of
5:12:03data be it structured be it unstructured
5:12:05or be it
5:12:06semi-structured okay now once you store
5:12:09the data in hdfs you can view the entire
5:12:12data as a single unit as well hdfs
5:12:15stores data across various nodes that
5:12:17these nodes are nothing but the data
5:12:19nodes and it also maintains the log
5:12:21files of what data is stored at which
5:12:23position so basically hdfs has got two
5:12:26components one is the name node and
5:12:28other is the data node data name node is
5:12:30the one which manages the entire cluster
5:12:33which manages the entire set of data
5:12:35nodes and keeps the information keeps
5:12:37the metadata of the data that is stored
5:12:39in these data nodes data nodes on the
5:12:42other hand are the slave machines the
5:12:44commodity Hardwares which actually
5:12:45stores the data so sdfs is the one which
5:12:49solves the primary problem of storing
5:12:51big data so it's time we move on and
5:12:54explore the next tool in Ado ecosystem
5:12:57which
5:12:58is Yan now we'll explore
5:13:02Yan now as the name suggest it is
5:13:05nothing but a resource
5:13:07negotiator okay the main purpose of yan
5:13:10is to allocate resources to run
5:13:12particular task over the Hadoop cluster
5:13:15so Yan has basically two components one
5:13:17is the resource manager and the other
5:13:19one is the node
5:13:20manager as soon as a client submits a
5:13:23job these resources are nothing but the
5:13:25containers in which the jobs can be
5:13:27executed okay node manager is the one
5:13:30which finally executes the a job within
5:13:32these containers and manages the entire
5:13:35thing on the data nodes okay the
5:13:38resource manager is the master demon and
5:13:40the node manager is the slave
5:13:42demon apart from that I would like to
5:13:44tell you one important thing about yan
5:13:46yan was introduced in Hadoop 2.0 which
5:13:49enabled various ecosystem tools to
5:13:51connect with Hadoop distributed file
5:13:53system and leverage Big Data okay so
5:13:56we'll understand this better when we
5:13:59come to the map reduced slide okay okay
5:14:02we'll move on to the next
5:14:05slide so we now come to map ruce which
5:14:07is the processing unit of Hado so once
5:14:10the data is stored on Hado distributed
5:14:12file system the next task is to process
5:14:15that data for doing so one can use a
5:14:17Pache map produce in Hado 1.x map
5:14:20produce was the only framework that can
5:14:22be used to process the distributed data
5:14:24that is present on
5:14:26hdfs however soan was the next layer
5:14:28over htfs and map ruce now connected
5:14:31with Yan to allocate resources for
5:14:34executing the map reduce task similarly
5:14:36many other ecosystem tools or databases
5:14:39now can connect with Yan and leverage
5:14:41hdfs okay so it happened after Yan so
5:14:45essentially it has got two functions one
5:14:47is the map and the reduce map function
5:14:49is used for filtering grouping and
5:14:50sorting kind of functions and the result
5:14:52of the map function is then aggregated
5:14:54in the reduce phase and the entire
5:14:56summarized result is dumped on the hdfs
5:14:59itself so this is how Hadoop map
5:15:01produced works now it's time we move on
5:15:03to the next
5:15:06slide and we'll explore what is Apache
5:15:11Pig Apache Pig was a tool that was
5:15:13developed at Kahu so it is nothing but a
5:15:16data processing tool that runs over
5:15:18Hadoop or you can say that it sits on
5:15:20top of the Hadoop Apache Pig has its own
5:15:22language that is called Pig Latin which
5:15:24is nothing but a dataflow language or
5:15:26you can call it as instructional
5:15:28language for example if you want to load
5:15:30a data you have a command like like load
5:15:32this data from path and then you can
5:15:34dump that data or perform various
5:15:36functions like filter or grub okay so
5:15:39using Pig Latin the life of the
5:15:41developers became very easy they need
5:15:43not write the entire full map Produce
5:15:45job for executing some processing over
5:15:48the big data for people who cannot write
5:15:52a map produce program or or were not
5:15:54comfortable with map ruce for them pck
5:15:57Latin came as a blessing or for them P
5:16:01came as a Blessing by using p the task
5:16:04for them became very easy and they were
5:16:06able to leverage Big Data it is said
5:16:08that approximately one line of pig latin
5:16:11is equals to 100 lines of map produced
5:16:13so just think how much time are you
5:16:15saving there okay you can perform all
5:16:18the ETL operations that you would like
5:16:20to execute over big data using Peg so I
5:16:23hope this gives you a clear picture how
5:16:26Apache Peg came into picture what is the
5:16:28importance of Apache Peg okay then we'll
5:16:32move on and explore the next tool in the
5:16:35chain that is Apache
5:16:38Hive Apache Hive is one of the most
5:16:41important tool that is there in the
5:16:42Hadoop ecosystem Apache Hive was
5:16:45developed at Facebook now the idea
5:16:47behind aache hiive was the time when it
5:16:50was developed the relational databases
5:16:52were flourishing these were the
5:16:54databases that were used by most of the
5:16:56organizations most of the companies
5:16:58nobody knew about no SQL or system like
5:17:02Hado even Facebook had its website on
5:17:04MySQL so the workforce there mostly was
5:17:08working on myql or SQL like queries or
5:17:11PL SQL
5:17:13plsql so for Facebook it was a problem
5:17:15because the workforce were skilled in
5:17:18SQL however for writing a map produced
5:17:21program you had to know some other
5:17:23programming language so what Facebook
5:17:26did is uh Facebook came up with tool
5:17:28that is called Hive using which you can
5:17:30write SQL like queries that is called
5:17:32high query language and execute the same
5:17:37task over the Hadoop cluster and
5:17:38leverage Big Data just like Pig using
5:17:41Hive you can write simple SQL like
5:17:43queries and the task that you were
5:17:45executing using map produce now can be
5:17:47executing using Hive without getting
5:17:49into the complexities of map ruce okay
5:17:52even using Hive you can connect from
5:17:54client applications like Java as well if
5:17:56you have that
5:17:57requirement okay so Hive is one
5:18:00important to tool that is used by a lot
5:18:03of people out there who do not want to
5:18:05get into writing the map produce program
5:18:08so let's move on to the next tool that
5:18:12is mahot and Spark mlip mahot is a
5:18:16machine learning library written in Java
5:18:19it can be used for creating
5:18:21recommendation engines or uh clusters of
5:18:24data or classify your data into various
5:18:26grps okay so all those algorithms that
5:18:29are there in machine learning can be
5:18:30implemented over big data using mahot
5:18:34okay so it provides you a command line
5:18:36interface to achieve the same task you
5:18:38would have heard about an analysis that
5:18:41is called Market Basket analysis which
5:18:43can be easily executed using Maho over
5:18:46big data the various other things like
5:18:48recommendation engine as I mentioned
5:18:50which you would have seen in many
5:18:52e-commerce websites like Amazon flip
5:18:54cart or many more okay so all those
5:18:57things can be done using
5:19:00Mah now now we come on to spark which is
5:19:03a leading tool in the Hadoop ecosystem
5:19:06map ruce and Hadoop together can only be
5:19:08used for batch processing that means
5:19:10you're not getting the results in real
5:19:12time but out there there is a
5:19:15requirement for realtime analytics as
5:19:17well which cannot be done using Hadoop
5:19:19map reduce right in that case Sparks
5:19:21come into the picture which can run
5:19:24Standalone as well as it can run over
5:19:26the Hado cluster and leverage the same
5:19:28big data to provide you realtime
5:19:30insights right as well as Apache spark
5:19:34is almost 100 times faster than Apache
5:19:37map ruce okay so I hope this excites you
5:19:41right so we'll move on to the next
5:19:45slide and explore Apache Edge
5:19:48Bas so what is Apache hedge Bas Apache
5:19:51Edge Bas is a nosql database that runs
5:19:54over aoop Apache hpas can be used for
5:19:57storing any kind of data that is there
5:19:59okay it could be any structured or
5:20:01unstructured data and P Edge Bas has
5:20:05been modeled after Google big table and
5:20:08can be utilize to store any big data
5:20:10that is there in Hadoop file system with
5:20:13Apache hpas you have an advantage that
5:20:16is you can use edpas as a backend for a
5:20:19website or web application to query in
5:20:21real time which cannot be done with
5:20:23tools like pck Hive or map ruce or even
5:20:27hdfs and hence it is a very important
5:20:30addition to hadopi
5:20:31ecosystem okay so guys are you clear
5:20:34with this you can also write a Java
5:20:36application and connect with Edge base
5:20:38using the rest apis Thrift apis or AO
5:20:42apis
5:20:44okay now we'll go through another tool
5:20:47that is called Apache drill Apache drill
5:20:50is again an open source application
5:20:52which works well with any distributed
5:20:54environment that is out there it can
5:20:56work with any nosql database or a flat
5:20:58file system okay the advantage with
5:21:01Apache drill is that it can connect with
5:21:03various nosql databases or a flat file
5:21:07system or a simple file itself at the
5:21:10same time so if you have data stored in
5:21:12various sources like let's say you have
5:21:15a data stored in Hadoop distributed file
5:21:17system you have a data stored in Edge
5:21:18Bas you have a data stored in mongodb
5:21:21every one of them has their own syntax
5:21:23to execute queries on them to retrieve
5:21:26the same set of Records however using a
5:21:29pacher drill you can connect to all
5:21:31these databases at a single time execute
5:21:34one query and extract the results from
5:21:36all the three databases and use it for
5:21:39your application okay Apache drill is
5:21:42able to do that because it follows the
5:21:44an SQL which enables you to write a
5:21:47query that can execute or that can be
5:21:49understood by all the three
5:21:52databases we'll move on to the next
5:21:55slide and now we'll explore
5:21:58Uzi Apache Uzi is nothing but a
5:22:01scheduler in the Hadoop ecosystem now
5:22:04what does it mean let's say you have a
5:22:06map ruce task that needs to be executed
5:22:08every hour now in that case instead of
5:22:10manually triggering it what you can do
5:22:12is you can define a workflow in Uzi and
5:22:15schedule your task to be executed after
5:22:18every 1 hour okay when you see you're
5:22:21doing two things one is you're defining
5:22:23a workflow that could be one task or it
5:22:26could be a combination of tasks that are
5:22:28executed by various tools like map
5:22:30produce hi Pig Etc in a sequence as well
5:22:33as you're defining the frequency in
5:22:35which the workflow needs to be executed
5:22:38okay so life becomes very easy you need
5:22:40not go and execute or trigger your job
5:22:43every time that need it needs to be done
5:22:46Uzi can do it for you along with that
5:22:48Uzi coordinator is another component
5:22:50that is present in Uzi which ensures
5:22:53that the job or the workflow is only
5:22:55executed when the data is available so
5:22:57at times if the data is coming from an
5:22:59external Source automat automatically it
5:23:01will ensure that as soon as the data is
5:23:04in the system then only the workflow of
5:23:06the job is executed so it is an event
5:23:08based execution that can be triggered
5:23:11using Uzi
5:23:13okay let's move on and we come to
5:23:17Flume Flume is again one of the most
5:23:20widely used tools which is used for data
5:23:22ingestion into hdfs okay so using Flume
5:23:26you can ingest any kind of data it could
5:23:28be structured it could be
5:23:29semi-structured into the the Hado
5:23:31distributed file system and perform
5:23:33various processing after that floom
5:23:36gives you the capability of extracting
5:23:38data out of social media like Twitter
5:23:40Facebook or you can also extract data
5:23:43from servers where logs are getting
5:23:45generated on a regular interval so Flume
5:23:47can be utilized to extract data from
5:23:50there and move into the hdfs okay
5:23:53similarly there could be many other use
5:23:55cases like getting email messages or
5:23:56network traffic
5:23:58Etc the next tool in the as scoop scoop
5:24:02is again used for data inje however
5:24:05scoop is used between relational
5:24:07database and the hdfs so using scope you
5:24:11can move your data from your relational
5:24:12database into hdfs and vice versa that
5:24:15means you can also move data out of hdfs
5:24:18into an rdbms so it mostly deals with
5:24:21structured
5:24:22data okay so if you compare Flume with
5:24:25scoop Flume is mostly used for moving
5:24:28data into the hdfs and deals with
5:24:31streaming data most of the time however
5:24:33scoop works with structured data and it
5:24:36can move data in and out of hdfs unlike
5:24:39flu so let's move on and we come to the
5:24:43next tool that is solar and
5:24:45Lucine okay so solar Lucine is again an
5:24:49Apache project which has been developed
5:24:51in Java Lucine in itself is a Java
5:24:53Library which is for developing search
5:24:56engine and indexers okay Apache solar is
5:25:00an application that is built using aacha
5:25:03loine so if you want to develop a search
5:25:05engine or you want to implement search
5:25:07onto your website which works very fast
5:25:10using indexing you can always use Apache
5:25:13solar to do that okay so this is the
5:25:15main purpose of Apache solar which is an
5:25:18application which is developed using
5:25:20Apache
5:25:21Lucin we come to
5:25:24zookeeper as the name suggests the job
5:25:27of Zookeeper is to ensure coordination
5:25:29between various tools that are there in
5:25:31the Hadoop ecosystem okay so the main
5:25:34purpose of Zookeeper is to ensure that
5:25:37each and every tool is able to
5:25:38communicate with each other without any
5:25:40Interruption so that the entire
5:25:42ecosystems works together in achieving a
5:25:45particular task okay it performs
5:25:47synchronization it performs
5:25:49configuration management grouping and
5:25:51naming of all these things okay it also
5:25:54manages all the services that are
5:25:56running in the Hadoop cluster okay so
5:25:59zookeeper is very important component of
5:26:01Hado cluster if zookeeper is not there
5:26:04your services your demons your tools
5:26:07will not be able to interact with each
5:26:09other or communicate with each other and
5:26:12hence you'll get a broken system if
5:26:14zookeeper
5:26:15fails now we come on to the final tool
5:26:18that is a Apache Amari in the Hadoop
5:26:21ecosystem that we are going to discuss
5:26:23today so apach Amar is a cluster manager
5:26:27okay what does a cluster manager mean or
5:26:31what what does a cluster manager do
5:26:33cluster manager manages the Hadoop
5:26:35cluster okay using Apachi mbari you can
5:26:38provision manage and monitor their P
5:26:40Hadoop clusters okay it makes very easy
5:26:43for you to set up a Hadoop cluster and
5:26:46then configure all the services that
5:26:48needs to run over the Hado cluster it
5:26:50could be a PES spark service it could be
5:26:52a hue service it could be any other
5:26:54service that you need over the cluster
5:26:56and it can be done very easily using
5:26:58Apache ambari so this was basically
5:27:01developed by hoton works a similar tool
5:27:04is developed by claura as well which is
5:27:06called claura manager okay however Cloud
5:27:10manager is not an open-source tool like
5:27:13Apache ambari so Apache ambari was
5:27:15developed by Harden Works however it was
5:27:17given to Apache later on a similar tool
5:27:20which is a propriety tool developed by
5:27:22Cloud named as Cloud manager using which
5:27:25you can deploy the Cloudera clusters
5:27:27however it is a paid service okay using
5:27:30using aache Amari you can also monitor
5:27:32health and status of your Hado cluster
5:27:34okay now you know the importance of
5:27:36aache Amari and how it can be used to
5:27:39make your life
5:27:40[Music]
5:27:45easy the prerequisites to install Hadoop
5:27:48in Windows operating system are Java so
5:27:52we all know that Hadoop supports only
5:27:54Java version 8 so firstly we need to
5:27:57download Java 8 version followed by that
5:28:00a latest Hardo version which we need for
5:28:02our operating system then the
5:28:04configuration files so these were the
5:28:07prerequisites now let's quickly go ahead
5:28:09and download Java 8 version into our
5:28:11local system and also hadu so you can
5:28:14see that this particular web page
5:28:16belongs to Oracle and here you'll be
5:28:19getting your Java development kit number
5:28:21eight so these are the various versions
5:28:23available for Java 8 for Linux as well
5:28:26as Windows so we need a jdk which is
5:28:29compatible with Windows
5:28:31so here you can see that Windows x64 jdk
5:28:35version which will support Windows so
5:28:37this particular link will redirect you
5:28:39and download jdk8 for you into your
5:28:42local system once you click on it it
5:28:44will ask you to accept the license terms
5:28:47from Oracle now you can just click on
5:28:50download followed by this you will be
5:28:52redirected into a login page where you
5:28:54need to create your own account with
5:28:57Oracle so that you can download this jdk
5:29:00don't worry this account is free of cost
5:29:03so you can see the jdk is getting
5:29:04downloaded here so as the jdk is getting
5:29:08downloaded we shall now move ahead and
5:29:10download Hardo for our local system so
5:29:12this particular web page belongs to
5:29:14Apache organization where we can
5:29:16download Hadoop for free so these are
5:29:18the various versions available for
5:29:20Hadoop which are 2.10 3.1.3 3.2.1 and
5:29:24many more so we shall select the latest
5:29:27version of Hardo but while you selecting
5:29:30the latest version of Hardo please make
5:29:32sure that you're not actually
5:29:33downloading the exact latest version of
5:29:35Hardo here you can see we have three
5:29:38different versions 3.1.3 3.2.1 3.1.2 as
5:29:43you can see 3.2.1 is the latest version
5:29:47we have to select the version which is
5:29:49earlier to it which is
5:29:513.1.3 because this particular version
5:29:54will be the stable version now we shall
5:29:56move ahead and select binary once you
5:29:59select binary
5:30:00you will be redirected into a new web
5:30:02page where you will have a mirror link
5:30:05select that mirror link and your Hardo
5:30:06will be downloaded for your local system
5:30:09as you can see Hardo 3.1.3 tar.gz is
5:30:12getting
5:30:13downloaded now here you can see I have
5:30:15successfully downloaded her version
5:30:183.1.3 T file as well as jtk 8 and those
5:30:22two files have successfully moved into
5:30:24my C drive now let's install Java first
5:30:29now make sure you that you create a new
5:30:31folder for Java so select change and
5:30:35here select Windows C drive then select
5:30:38make new folder now rename this new
5:30:41folder as Java click okay and now select
5:30:45next you can see the installation
5:30:48procedure has now been
5:30:51started you can see Java development kit
5:30:548 has been successfully installed now we
5:30:57shall enter into program files and move
5:31:00jdk into Java file because sometimes
5:31:03there will be an error while we set
5:31:06environment variables for Java so you
5:31:09can see inside program files we have
5:31:11another folder called Java so inside
5:31:13Java there you have our jdk so now what
5:31:17I'll be doing is just moving this jdk
5:31:19into Java file which we have created in
5:31:22C drive this
5:31:25one now you can just delete this Java
5:31:28file from your program files so that you
5:31:30don't have to mess with duplication of
5:31:32java file now you have your Java and jdk
5:31:36in one single file which is Java that is
5:31:39you have created in Windows C drive now
5:31:41we shall move ahead and set the
5:31:43environment variables for Java so click
5:31:46windows and then enter into settings and
5:31:49inside the settings select system and
5:31:51inside system just type in environment
5:31:54variables and there you go select the
5:31:56edit the system environment variables
5:31:58option and you have this dialogue box
5:32:01here select environment variables and
5:32:03inside the environment variables you
5:32:05need to set the Java home as well as
5:32:07path for Java now select new and here
5:32:12just type in Java home and here let us
5:32:16add the location of jdk bin so here we
5:32:19will add in the variable value that is
5:32:21the jdk bin location so our jdk bin
5:32:25location is in the C drive and inside
5:32:27the C drive we have the Java folder and
5:32:29inside jav Java folder we have a jdk
5:32:321.8.0 and inside jdk file we have the
5:32:35bin location so this will be the home
5:32:38location for Java select okay and then
5:32:42now move into the next dialogue box
5:32:45which is the system variables and inside
5:32:47that select path and select edit here
5:32:50Create A New Path variable which will be
5:32:53the jdk path the same location that is
5:32:55the bin of jdk Select okay and now
5:32:59select okay again and now okay and close
5:33:03it now Java has been successfully
5:33:06installed into our local system now
5:33:08let's check Java is functional or not we
5:33:10can do that by selecting Windows R and
5:33:13inside windows R just type in CMD so
5:33:16that you can open your command prompt
5:33:18here just type in Java C if you see the
5:33:22set of files popping up into your
5:33:23terminal then it means that Java is
5:33:26working properly so you can see Java is
5:33:28working just fine now let us check the
5:33:30version of our Java installed into our
5:33:32local system so this can be checked by
5:33:34typing in Java space hyphen version so
5:33:38you can see we have 1.8 version which is
5:33:41running in our local system now that we
5:33:43have successfully installed Java into
5:33:45our local system let us now move ahead
5:33:47and install Hado into our local system
5:33:50you can see that we have downloaded the
5:33:51tab version of Hado so for that we need
5:33:54to extract it
5:33:57first now you can see that the process
5:33:59of extraction has been completely
5:34:01finished that is 100% but you have three
5:34:03errors you can ignore these errors now
5:34:06just close the extracting process then
5:34:09you have your Hadoop
5:34:12file now let us rename a Hadoop
5:34:153.1.3 as just Hadoop to reduce the
5:34:18confusion now that we have successfully
5:34:20extracted Hadoop let's set environment
5:34:22variables for Hadoop but before that
5:34:25let's set the configuration of Hadoop
5:34:28you can select Hadoop and inside that
5:34:30you have a file called Etc and inside
5:34:32Etc you have another folder with the
5:34:35name hadu and inside that you have a set
5:34:38of folders so out of these all folders
5:34:41We have four important folders they are
5:34:45core site. XML then htfs site. XML
5:34:50followed by htfs we have another one
5:34:52which is map site. XML and lastly the
5:34:56Yan site. XML file so we need to edit
5:34:59all these four different files and once
5:35:02after we edit these four files we need
5:35:04to edit one last file which is the
5:35:07Hadoop EnV Windows command PR file so
5:35:11here you're just going to add in the
5:35:13Java home location now let's quickly
5:35:16edit all those four files so we have
5:35:19successfully opened our four important
5:35:21files which are cor site. XML map reduce
5:35:24site. XML Yan site. XML htfs site. XML
5:35:29followed by by the four important files
5:35:31the last file which is the Hado
5:35:33environment. CMD file here we are going
5:35:36to set this Java home location now let's
5:35:38first set the values for corite XML so
5:35:42the values that are changed in corite
5:35:44XML are the properties so inside the
5:35:47configuration I have added one property
5:35:49which is the file location that is fs.
5:35:52default file system and the Local Host
5:35:54location that is 9,000 now let us save
5:35:57this course site. XML similarly we need
5:36:00to also edit map reduce side. XML files
5:36:03here inside this we need to add some
5:36:05properties as you can see we have also
5:36:08edited the configuration files of map
5:36:10redu site. XML let's save it now
5:36:13followed by the map redu site. XML we
5:36:15have Yan site. XML let's edit this
5:36:19also as you can see the Yan site. XML is
5:36:22also been updated no worry about this
5:36:25property file I will link this in the
5:36:27description box below you can have the
5:36:29access to it and you can use the same
5:36:30configuration file and install her tube
5:36:32followed by Yan site. XML we have the
5:36:35last one which is htfs site. XML but
5:36:38before editing this particular file I
5:36:41want you to create a new folder in hero
5:36:43location which is data let's see how to
5:36:46create it so this particular folder is
5:36:49inside C drive this is Hadoop and inside
5:36:53this Hadoop file you need to create a
5:36:54new folder with the name data inside
5:36:58data you need to create two more new
5:37:00files which are data node and name node
5:37:03so the first folder will be name node
5:37:06and now another folder which will be our
5:37:09data
5:37:10node so now let's copy the location of
5:37:13data node and name node so this location
5:37:16is the data node location and followed
5:37:19by that the name node location so this
5:37:21particular location will be the name
5:37:22node location we have the two locations
5:37:24copied onto our clipboard now let's go
5:37:27back to the htfs site dox SML file and
5:37:30edit the configurations
5:37:32here so you can see that we have edited
5:37:35the configuration file of htfs site. XML
5:37:38and inside the configuration we have
5:37:40provided the replication factor which is
5:37:42the first property and we have set the
5:37:44value as one since we're using our local
5:37:46system we might want to save memory so
5:37:48the replication is only one but the
5:37:50default value for the Hardo replication
5:37:52factor is three and followed by the
5:37:54first property the second property which
5:37:56is our name node so we have provided our
5:37:59name node a which is Hardo file and
5:38:01followed by that the data file and
5:38:03inside that we have the name node and
5:38:05similarly the last property which is the
5:38:08data node property so here the value is
5:38:11Hadoop data data node now let's save
5:38:16it now that we have successfully edited
5:38:18all our four important files let's get
5:38:21back to Hadoop env. CMD file and edit
5:38:24the Java home location so for safest
5:38:27side let's get back to environment
5:38:28variables and and get our jdk
5:38:33location so this particular location is
5:38:36the location for Java home we might want
5:38:38to remove the bin over here so only C
5:38:41Java jdk is enough to set the Java home
5:38:44into our Hardo env. CMD file now let's
5:38:48save this particular file and close it
5:38:50so all the important files have been now
5:38:52successfully edited now let's go back to
5:38:54environment variables and set home and
5:38:56path for
5:38:57hadu now select new and write in Hadoop
5:39:04home so this particular location that is
5:39:07C Hadoop Ben is the location for Hadoop
5:39:10home now select okay now let's get back
5:39:13to path and set path for Hadoop files in
5:39:17here let's set up a new path variable
5:39:20that is Hadoop bin and now remember to
5:39:23create another path variable that is
5:39:25your spin so to locate Espin get back
5:39:29back into Hadoop and select Spin and
5:39:32this will be the location or path value
5:39:34for your spin select that and edit a new
5:39:37variable in path and paste it so that's
5:39:40how you set sben and select okay okay
5:39:43and finally another okay and close the
5:39:45system properties now that we have
5:39:47successfully set home and path for
5:39:49Hadoop let's go ahead and fix the
5:39:52configuration files you can see that
5:39:54inside the bin folder of Hardo we are
5:39:57missing some important configuration
5:39:59file to fix this we need a new
5:40:01configuration file which will be
5:40:03available in the description box below
5:40:05you can click on that particular link
5:40:07and the required configuration file will
5:40:09be downloaded into your local system and
5:40:11all you need to do is just replace that
5:40:13particular file with your bin folder in
5:40:15your Hadoop you can see that there is a
5:40:17new file in my Hadoop which is Hadoop
5:40:19configuration Fixx bin 1.rar now all you
5:40:23need to do is just extract this
5:40:25particular
5:40:28folder
5:40:34you can see that the folder has been
5:40:36successfully extracted and all the
5:40:38executable files that you require in
5:40:40your Hadoop have been downloaded
5:40:41successfully now what you need to do is
5:40:44just move this bin into your Hadoop bin
5:40:48so cut this bin and get back to Hadoop
5:40:51and enter bin so just delete this
5:40:54particular bin and replace it with a new
5:40:56one so there you go you have
5:40:58successfully done on it now let's delete
5:41:00the unnecessary files there you go as
5:41:03good as new so you have all the
5:41:05executable files and your hop is been
5:41:07set to check if hero is functioning
5:41:10properly or not let's open CMD and type
5:41:13in
5:41:15hdfs space name node space hyphen
5:41:20format if you see a set of files popping
5:41:23up on your terminal that means you have
5:41:25successfully installed heru you can see
5:41:27that the name note has been successfully
5:41:28getting started
5:41:29now let's open a new terminal and start
5:41:32all the Hadoop demons here you just need
5:41:34to enter your Hadoop location file that
5:41:38is CD space
5:41:41Hadoop now you are inside Hadoop and
5:41:44inside Hadoop enter
5:41:47sbin now you're inside sbin now you need
5:41:50to type in start all. shr start all.
5:41:57CMD and there you go all your demons are
5:42:00getting started so that's how you
5:42:03install hardup into your local Windows
5:42:05operating system with the version
5:42:06Windows
5:42:08[Music]
5:42:1210 so what were the problems associated
5:42:15with the relational database system as I
5:42:17have already mentioned that for a Hado
5:42:19developer actual game starts after the
5:42:21data is being loaded in hdfs and the
5:42:24developers play around this data in
5:42:25order to gain various insights that are
5:42:27hidden in the data stored in H hdfs so
5:42:30for this analysis the data residing in
5:42:32the rdbms needs to be transferred to
5:42:35hdfs and you all know that the task of
5:42:37writing map reduce scod for importing
5:42:39and exporting the data from relational
5:42:41database to hdfs is TDS so this is where
5:42:44Apache scoop comes to rescue and remove
5:42:46the pain of data inje so why do we need
5:42:49scoop it's a known fact that before Big
5:42:51Data came into existence the entire data
5:42:54was stored in relational database
5:42:55servers in the relational database
5:42:57structure but with the advancement of
5:42:59scoop it makes the life of developers
5:43:01Easier by providing CLI for importing
5:43:03and exporting the data and scoop
5:43:05internally converts a command into map
5:43:07ruce task which are then executed over
5:43:10hdfs it uses yarn framework to Import
5:43:13and Export the data which provides fall
5:43:15Tolerance on top of parallelism it also
5:43:17uses yarn framework to Import and Export
5:43:20the data which provides fall Tolerance
5:43:22on top of parallelism not only that it
5:43:24is also very useful for data analysis
5:43:27high in performance and provides command
5:43:29line interface now let's understand what
5:43:31a scoop so before I tell you what a
5:43:34scoop let me tell you how the name scoop
5:43:36came into
5:43:37existence by now I hope that you have
5:43:40got an idea that it is used for data
5:43:42transfer between relational database and
5:43:44hdfs so before I tell you what is scoop
5:43:47let me first tell you how the name came
5:43:48into existence the first two letters in
5:43:51scoop stands for the first two letters
5:43:53in SQL and the last three letters in
5:43:55scope refers to the last three letters
5:43:57in Hardo that is o op so it clearly
5:44:00depicts that it is SQL to Hadoop and
5:44:02Hadoop to SQL that is how the name of
5:44:05scop came into existence so what is scop
5:44:08it is a tool used for data transfer
5:44:10between rdbms like MySQL Oracle SQL Etc
5:44:14and Hadoop like Hive hdfs hbas Etc it is
5:44:18used to import the data from rdbms to
5:44:20Hadoop and Export the data from Hadoop
5:44:22to rdbms simple again scoop is one of
5:44:26the top projects by Apache software
5:44:28Foundation and works brilliantly with
5:44:30relational databases such as Terra dat
5:44:32nza Oracle MySQL Etc it also uses map
5:44:36reduce mechanism for its operations like
5:44:38Import and Export work and work on a
5:44:40parel mechanism as well as fall
5:44:42tolerance as I have already mentioned
5:44:44that it provides command line interface
5:44:46for importing and exporting the data the
5:44:48developers just have to provide the
5:44:50basic information like database
5:44:52authentication Source destination
5:44:54operations Etc and the rest of the work
5:44:56will be done by scoop tool itself sounds
5:44:59much reliable correct now let's move
5:45:02further and talk about some of the
5:45:04amazing features of sco for Big Data
5:45:06developers First full load Apache scope
5:45:09can load whole table by a single command
5:45:12you can also load all the tables from a
5:45:14database using a single command next
5:45:17incremental load scool provides a
5:45:19facility of incremental load where you
5:45:21can load the parts of a table wherever
5:45:23it is updated next parallel Import and
5:45:25Export again as already mentioned scoop
5:45:28user this Yan framework to Import and
5:45:30Export the data which provides fall
5:45:32Tolerance on top of parallelism next
5:45:35compression you can compress your data
5:45:37by using Gip algorithm with compress
5:45:39argument or by specifying compression
5:45:41codec argument next karos security
5:45:44integration so what is karos it's a
5:45:47computer network Authentication Protocol
5:45:50which works on the basis of tickets to
5:45:52allow the nodes that are communicating
5:45:53over a non-secure network to prove their
5:45:56identity to one another in a secure
5:45:59manner next load data directly into Hive
5:46:02and hbas here it is very simple you can
5:46:05load the data directly to Apache high
5:46:07for analysis and you can also dump your
5:46:09data in hbase which is a no SQL database
5:46:13now let's see what's next the
5:46:15architecture is one of the empowering
5:46:17Apache scope with its benefits now as we
5:46:20know the features of Apache scope let's
5:46:22move ahead and try to understand Apache
5:46:24scope's architecture and its working so
5:46:27when we submit our job or a command
5:46:28through through scope it is mapped into
5:46:30map task which brings a chunks of data
5:46:32from hdfs and these chunks are exported
5:46:35to a structured data destination and
5:46:38combining all these exported chunks of
5:46:40data we receive the whole data at the
5:46:42destination which in most of the cases
5:46:44is rdbms server next reduce phase is
5:46:47required in case of aggregations but
5:46:50Apache scope just Imports and Export the
5:46:52data it does not perform any
5:46:54aggregations map job launch multiple
5:46:56mappers depending on the number defined
5:46:58by the user for scoop import each maper
5:47:01task will be assigned with the part of
5:47:03data that is to be imported and scoop
5:47:05distributes the input data among all the
5:47:07mappers equally in order to achieve high
5:47:10performance then each mapper creates a
5:47:13connection with the database using gdbc
5:47:15and fetches the part of the data
5:47:17assigned by the scope and then writes
5:47:19that data to hdfs Hive or hbas based on
5:47:22the arguments provided in the command
5:47:24line interface so this is how scope
5:47:26Import and Export works like the gather
5:47:29the metadata again it submits only map
5:47:31job the reduced phase will never occur
5:47:33here and then it stores the data in hdfs
5:47:36storage coming to scoop export it's the
5:47:39same thing the data will be reversed
5:47:40back to
5:47:42rdbms so here the scoop import tool will
5:47:45import each table of the rdbms in Hardo
5:47:48and each row of the table will be
5:47:49considered as a record in the hdfs and
5:47:52all the records are stored as Text data
5:47:54in the text files or binary data in
5:47:56sequence files on the other hand the
5:47:58scoop export tool will export the Hardo
5:48:01files back to the rdbms tables again the
5:48:03records in the hdfs files will be the
5:48:05rows of a table and those are read and
5:48:08passed into a set of records and
5:48:09delimited with the user specified
5:48:11delimiter so this is all about the scoop
5:48:14architecture and its Import and Export
5:48:16now let's execute some scoop commands
5:48:18and understand how it works so at the
5:48:21first we have scoop import that is it
5:48:23Imports the data from rdbms in hadu the
5:48:26command goes very simple here you have
5:48:28to just provide the connection for MySQL
5:48:31your IP address your database name your
5:48:33table name the username for MySQL user
5:48:36if you have set privileges for password
5:48:38you can specify the password or it is
5:48:40not required and the target directory
5:48:42now let's see how to
5:48:44execute so I'll open my terminal and
5:48:47check whether all my Hardo demons are up
5:48:49and running or
5:48:51not so I can see that all my Hardo
5:48:54demons are up and running so now let's
5:48:57execute scoop help and see whether the
5:48:59scoop has been properly installed or
5:49:01not so these are the available commands
5:49:04in scoop here I will show you the
5:49:06execution and explain you few of these
5:49:08commands now this is also properly being
5:49:11set now I will open another terminal and
5:49:13connect to my SQL database the command
5:49:16goes like
5:49:17this my user is edura so I'm giving it
5:49:20as edureka you can give it as root if
5:49:22your user is root simple so now I'm into
5:49:25my SQL database if you want to create a
5:49:28data database you can create the
5:49:30database by giving this command create
5:49:32database database name Etc as I have
5:49:35already created a database so I'll just
5:49:37specify show databases command to list
5:49:39the database present in the mySQL
5:49:41database so now these are the list of
5:49:44database present here so now I want to
5:49:46use employees database so I'll give use
5:49:49employees database got changed now I
5:49:52want to list the tables present in the
5:49:54employees database so I'll give show
5:49:57tables so these are the 11 tables
5:50:00present in the database employees now
5:50:03let's say I want to use employees table
5:50:05so what will I do I'll just give select
5:50:08star from employees that is table
5:50:11name so it's just a huge amount of data
5:50:14that is present in this
5:50:17database so there are these many rows
5:50:20present in this table now open the other
5:50:23terminal where you have executed this
5:50:25command and here I will show you how to
5:50:27import the data present in in that table
5:50:29to hdfs so how we are going to do that
5:50:31by using scoop import command the
5:50:33command goes like
5:50:36this the IP here is Local Host and you
5:50:39know that employees is my database name
5:50:42and the username will be
5:50:44adura and the table that I have chosen
5:50:47is employees so I have not set any
5:50:49privileges for password so I'm not
5:50:50specifying the password over
5:50:55here so it got executed and you can see
5:50:59the number of job counters the map
5:51:00reduce framework the input records
5:51:02output records Etc so what happens after
5:51:05executing this command the map task will
5:51:07be executed at the back end now let's
5:51:10check the webui of hdfs that is the
5:51:13webui port for hdfs is Local Host
5:51:17570 and let's see where the data got
5:51:20imported one important thing to note I
5:51:23have not specified Target directory so
5:51:25by default the data will be imported to
5:51:28this folder user edure Rea and employees
5:51:32so here you can see the four different
5:51:34part files where our data got imported
5:51:36so you might be thinking y4 correct here
5:51:39I have not specified the number of
5:51:41mappers so by default it takes the
5:51:43number of mappers to three and then
5:51:45gives the output in four different part
5:51:47files so let's open the part file and
5:51:49see how the output will be so here you
5:51:52can see the data is imported from rdbms
5:51:55to hdfs there's lots of data being
5:51:57present over here I'm just scrolling
5:51:59down and it's not coming to an end
5:52:02similarly the output will be same in the
5:52:03other part files as well so this is all
5:52:06about the simple scope import command
5:52:08without specifying the target directory
5:52:10and the number of
5:52:11mappers now let's see how to import the
5:52:13data from rdbms to hdfs by specifying
5:52:16the target directory this part remains
5:52:19out to be the same now I will do one
5:52:21thing I'll specify the number of mappers
5:52:23as one so that your output will be in
5:52:25just one single part file and then I
5:52:28will specify the target directory as
5:52:30well and I will name the target
5:52:32directory as employee 10 enter again it
5:52:36takes a lot of time to execute because
5:52:38there is lot of data present in the
5:52:41database so it got executed and
5:52:43retrieved these many records again you
5:52:46can see the map reduce framework the
5:52:48number of bytes written number of read
5:52:50operations write operations Etc now
5:52:52again let's go to the htfs browser and
5:52:54see the output so I had specified the
5:52:57target directory name as employee 10 so
5:52:59you can see it here and the output is
5:53:02just in one single part file because I
5:53:04have specified the number of mappers as
5:53:05one in this case you can control the
5:53:08number of mapers independently from the
5:53:09number of files present in the
5:53:11directory so the entire records will be
5:53:14present in one single part file so this
5:53:17is the
5:53:19output now let me tell you how to import
5:53:22the command using wear Clause here you
5:53:25can import a subset of the table using
5:53:27the wear clause in scoop import tool it
5:53:29executes a corresponding SQL query in
5:53:31the respective database server and
5:53:33stores a result in a Target directory in
5:53:35hdfs so this is how the command goes the
5:53:39same command as before I'll just change
5:53:41the name of the target directory I'll
5:53:43specify employee 11 and I will increase
5:53:46the number of mappers to three and here
5:53:49I will specify the we
5:53:51Clause so let's give a condition like
5:53:54where the employee number will be
5:53:55greater than 499,000 it should dis the
5:53:58output enter so what you expect your
5:54:01output will be so here it displays the
5:54:03output records of the employees whose
5:54:05employee number is greater than
5:54:0949,000 so here it retrieved these many
5:54:12records which are above the employee
5:54:14number 49,000 now let's check the output
5:54:17so I have specified the target directory
5:54:19name as employee
5:54:22Lev so here you can see the output that
5:54:25it has retrieved the records of employe
5:54:28number which is more than
5:54:3049,000 so I hope you understand how to
5:54:33do this next let's see how to import all
5:54:36the tables from the rdbms database
5:54:38server to the hdfs here each table data
5:54:41is stored in a separate directory and
5:54:43the directory name is same as a table
5:54:45name it is mandatory that every table in
5:54:47that database must have a primary key
5:54:51field the command will be simple like
5:54:53simple import but just that you have to
5:54:55remove the table name and you have to
5:54:58replace the import with import all
5:55:00tables that's
5:55:02all so it will retrieve all the tables
5:55:05present in the
5:55:11employes okay so it imported all the
5:55:14tables from rdbms to htfs now let's
5:55:17check the
5:55:18output again I have not specified the
5:55:21target directory file so by default it
5:55:23will be in user edureka and you can see
5:55:26here it imported all the the tables
5:55:28present in the database to
5:55:31hdfs so these are the various tables
5:55:33present in rdbms and now it is present
5:55:36in htfs so this is how import all tables
5:55:39command works so this was all about
5:55:42executing import command in various ways
5:55:45now let's move further and see how does
5:55:47scop export works it exports the data
5:55:50from hdfs to rdbms correct again the
5:55:53command goes very simple you have to
5:55:55specify the connection your table name
5:55:56username and instead of the target
5:55:59directory you have to specify the export
5:56:01directory path so let's see how it works
5:56:04one important thing to notify the target
5:56:07table must exist in the Target database
5:56:09that is the data is stored as records in
5:56:11hdfs and these records are read and pass
5:56:14and delimited with the user specified
5:56:16delimiter the default operation is to
5:56:18insert all the records from the input
5:56:19files to the database using the insert
5:56:21statement in update mode scoop generates
5:56:24the update statement that replaces the
5:56:26existing record in the database so first
5:56:28we are creating an empty table where we
5:56:30will export our
5:56:31data I'm going to create a table called
5:56:34employee
5:56:36zero the primary key value should never
5:56:39be null so I'm specifying it as not
5:56:45null so here I created an empty table
5:56:47called employee zero and now I'll show
5:56:49you how the scoop export
5:56:52works here instead of import I'll make
5:56:54it as export this is the database name
5:56:58and I have created a table called
5:57:00employee zero and I'm going to specify
5:57:02the path for export
5:57:05directory simple that's
5:57:09all so you can see here it exported all
5:57:12these records into the rdbms so now
5:57:15let's cross check I'm going to give
5:57:18select count star from the table name
5:57:21that I have specified to export the
5:57:23tables so you can see the entire records
5:57:26got exported to the this table so this
5:57:29is how the scoop export
5:57:31works now let's see how to list the
5:57:33database present in the relational
5:57:35database here you need not even specify
5:57:38the database name because you're going
5:57:39to list the database that is present in
5:57:41the relational database system so you're
5:57:43going to specify list
5:57:45databases and execute the
5:57:49command so the databases present in the
5:57:51relational database system is test jdbc
5:57:55test employees and information schema
5:57:57again Let's cross check I'm going to
5:58:00give show databases it retrieve the same
5:58:02database so the result tallies so we can
5:58:06also list the tables present in the
5:58:08database let's see how to list all the
5:58:10tables present in the
5:58:13database again it's very simple you have
5:58:16to just specify the database name like
5:58:19here and give list tables instead of
5:58:25UT so you can see that it listed all all
5:58:27the tables present in the database
5:58:29employees so again let's cross check
5:58:33show
5:58:38tables same thing and now let's see what
5:58:41is Cen in object oriented application
5:58:45every database table has one data access
5:58:47object class that contains getter and
5:58:49seter methods to initialize the objects
5:58:51and coachin generates Dao class
5:58:54automatically and it also generates the
5:58:56Dao class in Java based on the table
5:58:59schema structure so this is a simple
5:59:01command let's see how to execute it here
5:59:04I will give scoop Cod
5:59:07gen and give the connection and the
5:59:09database name as
5:59:11employees and I'm going to specify the
5:59:13table name as well so it is going to
5:59:15create a employee CH file in which the
5:59:18backend code will be
5:59:20generated so I'll copy this path and
5:59:24jump into this
5:59:26directory so so you can see here that it
5:59:28created employees class jar file and the
5:59:31Java object file as well so now let's
5:59:34open the file system and check for the
5:59:36file so what was the name of the file it
5:59:40ends with 539
5:59:43D9 so here is the folder so in this you
5:59:47can see the object file that is being
5:59:52generated so this is the pack and code
5:59:54that is being
5:59:55generated so this is all about about how
5:59:57the scoop Cod gen
5:59:59[Music]
6:00:04Works what is Apache Pig so Apache pig
6:00:09is an abstraction over map reduce it is
6:00:12a tool or platform which is used to
6:00:14analyze larger sets of data representing
6:00:17them as data flows pig is generally used
6:00:20with Hadoop we can perform all the data
6:00:23manipulation operations in her doop
6:00:25using Apache pck to write data analysis
6:00:28programs Pig provides a highlevel
6:00:31language known as Pig Latin this
6:00:33language provides various operators
6:00:35using which programmers can develop
6:00:38their own functions for Reading Writing
6:00:41and processing data to analyze data
6:00:43using Apache Pig programmers need to
6:00:46write scripts using Pig Latin language
6:00:49all these scripts are internally
6:00:50converted into map and redu Tas
6:00:54respectively Apache Pig has a component
6:00:56called Apache shape Pig engine that
6:00:59accepts the pig latin scripts as input
6:01:01and converts those particular scripts
6:01:03into map reduced jobs so this was a
6:01:07basic introduction to Apache Pig now
6:01:10moving ahead we shall understand the
6:01:12different modes in which Apache Peg
6:01:14functions there are two particular modes
6:01:17in which Apache Peg functions those are
6:01:20the local mode and map reduce mode first
6:01:24we will understand what exactly is local
6:01:26mode in local mode Apache pig is
6:01:29designed to execute in a single jvm and
6:01:33is used for development experimenting
6:01:35and prototyping here files are installed
6:01:38and run using Local Host the local mode
6:01:42works on local file system and the input
6:01:45and the output data is stored in the
6:01:47local file system to access the command
6:01:50or the CR shell in local mode you need
6:01:53to execute a command called Pig hyphen X
6:01:56local
6:01:58we shall practically execute this in the
6:02:00demo section now moving ahead we shall
6:02:03discuss about the second type of mode in
6:02:05which Apache pck can be run that is the
6:02:08map reduce mode the map reduce mode is
6:02:11also known as Hadoop mode Apache Pig
6:02:14chooses Hadoop mode as its default mode
6:02:17in this pig renders Pig Latin into map
6:02:20rce shops and executes them on a Hadoop
6:02:23cluster it can be executed against
6:02:26semi-distributed
6:02:27or fully distributed Hadoop
6:02:29installation here the input and the
6:02:32output are present on
6:02:34hdfs the command for executing peg in
6:02:37map reduce mode is Peg or Peg hyphen X
6:02:41map reduce again we shall discuss about
6:02:44this particular command in our demo
6:02:46section where I'll show you both the
6:02:48modes and execute the pck scripts there
6:02:51now followed by the peg modes we shall
6:02:54understand the ways to execute Peg
6:02:56program so basically the big scripts are
6:02:59executed in three particular modes those
6:03:02are interactive mode batch mode and
6:03:05lastly the embedded mode no worry I'll
6:03:08explain to you each of these modes
6:03:11firstly we shall discuss about the
6:03:13interactive mode in this particular mode
6:03:15the pig is executed in a grun shell to
6:03:19invoke grun shell run the pig command
6:03:22once the grun mode executes we can
6:03:24provide big Latin statements and command
6:03:27command interactively at the command
6:03:29line itself so the next mode is the
6:03:32batch mode in this particular mode we
6:03:35can run a script file having a DOT Pig
6:03:38extension these files contain the pig
6:03:41latin commands so basically what we do
6:03:43is we write the pig command or the pig
6:03:47script and store it in a location and
6:03:50using the terminal we will access that
6:03:52particular location and that particular
6:03:54file and run the code present in that
6:03:56file
6:03:57so this is what happens in batch mode so
6:04:00followed by that the last mode is the
6:04:02embedded mode in this particular mode we
6:04:05can Define our own functions these
6:04:08functions can be called as userdefined
6:04:11functions here we use programming
6:04:13languages like Java and python to Define
6:04:16our own user defined functions so these
6:04:18were the three different modes in which
6:04:20we can execute Pig scripts now moving
6:04:23ahead we shall enter into our next topic
6:04:26that is the features of pig so there are
6:04:29five different and important features of
6:04:32pig so first up we shall understand the
6:04:34EAS of programming writing complex Java
6:04:38programs for map reduce is quite tough
6:04:40for non-programmers Peg makes this
6:04:43process very easy in the pig the queries
6:04:46are converted into map reduce processes
6:04:49internally so obviously it has grown the
6:04:52ease of programming followed by that the
6:04:55next important feature as the
6:04:57optimization
6:04:59opportunities so what exactly is
6:05:01optimization it is how the tasks are
6:05:03encoded permits the system to optimize
6:05:06the execution automatically allowing the
6:05:08user to focus on semantics rather than
6:05:11efficiency so followed by the
6:05:13optimization we have the
6:05:15extensibility a user defined function is
6:05:18written in which the user can write
6:05:21their logic to execute over the data
6:05:23sets so this increases the extensibility
6:05:26of the programmer followed by that the
6:05:29next important feature is it is highly
6:05:32flexible Apache pck can easily handle
6:05:35structured as well as unstructured data
6:05:38so Apache Pig tool is considered to be
6:05:40highly flexible irrelevant of the data
6:05:43type followed by that the last and
6:05:46important feature is inbuilt operators
6:05:50Apache pick contains various types of
6:05:52operators such as sort filter joints and
6:05:56Men anymore which are inbuild so here
6:06:00the programmer doesn't have to program
6:06:03these functions
6:06:05externally instead he can just directly
6:06:07get an access to those enbu functions
6:06:10and execute them so these were the five
6:06:12important features of Apache Pig so the
6:06:15next Topic in our today's discussion is
6:06:18Apache Pig installation into our local
6:06:20system now we shall discuss about one of
6:06:23the easiest ways to install Apache pig
6:06:26into our local system
6:06:27so today I'll explain you how to install
6:06:30Apache pck into Windows operating system
6:06:33so to do this the easiest way is to
6:06:36download one of the virtual machines so
6:06:38I would prefer you to download Oracle
6:06:41virtual box for this particular task
6:06:43followed by that we shall also download
6:06:45a quick start VM of cloud AR don't worry
6:06:49about these softwares I'll drop down the
6:06:51link for those softwares in the
6:06:52description box below you can use that
6:06:54and download them so once after the
6:06:57Oracle virtual box is installed into a
6:06:59local system and it is running this is
6:07:02how it looks like now what we want to do
6:07:06is to add a new virtual machine into our
6:07:08virtual box so for that you just need to
6:07:11select import option and it will give
6:07:15you a new dialogue box and from here you
6:07:18must redirect to the location where your
6:07:21virtual box is located so in my system
6:07:24it's located in F drive and inside F
6:07:29Drive CCA Cloud era virtual box cloud
6:07:33era quick start VM so there you go now
6:07:37you just have to select open and before
6:07:40we actually select the button import you
6:07:42might want to select the RAM and
6:07:44increase its size to at least 8 GB so
6:07:48just write in 9,000 MV which is just a
6:07:51little above 8GB now we are good to go
6:07:54just select the option import and your
6:07:56virtual Bo will be
6:07:58imported you can see that the quick
6:08:00start VM is getting
6:08:04imported now you can see that the
6:08:06virtual machine got successfully
6:08:08imported now to start it all you need to
6:08:10do is just click on it and then select
6:08:13the start
6:08:14button so you can see that the cloud
6:08:17error quick start VM version
6:08:205.13 has been successfully imported and
6:08:23booted
6:08:24up here we are we have the the quick
6:08:27start cler of VM welcome note and now if
6:08:30you want to start up with big editor you
6:08:33might want to log into Hue first or your
6:08:36hdfs remember in Cloud error the default
6:08:40username is cloud error and the password
6:08:43also is cloud era now let me tell you in
6:08:47Cloud era the default username and
6:08:50password for everything is cloud era for
6:08:53example let's log into our Hue using the
6:08:56username Cloud error and password Cloud
6:09:02error now you might want to just select
6:09:04remember to remember your
6:09:11password and there you go you have
6:09:13successfully logged into
6:09:16here and now if you want your query
6:09:19editor for pick you can just select the
6:09:22bottom Arrow Mark and there you can
6:09:24select editor and inside editor you have
6:09:27the editor designed for
6:09:31pig now you are in the window where you
6:09:34can write pick scripts and execute them
6:09:37now we will move ahead into our next
6:09:39topic and after finishing the theory
6:09:41part we shall come back into our Cloud
6:09:44era and execute some of the basic
6:09:46operations on Pig terminal so our next
6:09:49concept is understanding the pig
6:09:52architecture so first of all let us go
6:09:54through the diagram of pig AR
6:09:56architecture the following diagram
6:09:59represents the architecture of Apache
6:10:01pig as shown in the figure there are
6:10:04various components in Apache Pig
6:10:06framework let us look at the major
6:10:08components firstly the parser initially
6:10:12the pick scripts are handled by the
6:10:14passer it checks the syntax of the
6:10:17script does type checking and other
6:10:19miscellaneous checks the output of the
6:10:22Passa will be a dag which is directed a
6:10:26cyclic graph which represents the big
6:10:28Latin statements and logical operators
6:10:32in directed aycc graphs The Logical
6:10:35operators of the scripts are represented
6:10:37as the notes and data flows are
6:10:39represented as edges and next the
6:10:43optimizer the logical plan for directed
6:10:46asyc graphs is passed to The Logical
6:10:49Optimizer this carries out the logical
6:10:52optimizations such as projections and P
6:10:54Downs next comes the compiler the
6:10:58compiler is used to compile the
6:11:00optimized logical plan into the series
6:11:02of map reduce shops followed by that we
6:11:06have the execution engine finally the
6:11:09map reduce jobs are submitted to the
6:11:11Hadoop in a sorted order finally these
6:11:15map ruce jobs are executed on Hado
6:11:18producing the desired results so this
6:11:21was a basic explanation based on Apache
6:11:24pck architecture now moving ahead we
6:11:27shall understand the major advantages of
6:11:30Apache P so some of the major advantages
6:11:33of Apache Pig are less code the pick
6:11:36consumes less than a line of code to
6:11:39perform any operation so this reduces
6:11:43the number of lines included in the code
6:11:46followed by that code
6:11:48reusability the big code is so flexible
6:11:50enough to reuse it again you can
6:11:53basically write the piig script into a
6:11:55file and access the file whenever you
6:11:57need it so followed by that the next
6:12:00important advantage of Apache pck is the
6:12:03nested data types the pig provides a
6:12:05useful concept of nesting data types
6:12:08like topple bag and map so these were
6:12:11the few important advantages of Apache
6:12:14Pig now we shall move ahead into the
6:12:16next topic which is about the
6:12:18differences between Apache Pig and
6:12:20Apache map reduce so there are basically
6:12:23four differences between Apache Pig and
6:12:26Apache map reduce so first up we will
6:12:29begin with map reduce in map reduce we
6:12:32have low-level data processing when it
6:12:35comes to Apache Peg it is considered as
6:12:38a high level data processing tool
6:12:40followed by that the next important
6:12:42difference between map reduce and Apache
6:12:44p is that we find complex Java programs
6:12:48when we are using map reduce but when we
6:12:51come into Apache Pig we have SIMPLE
6:12:55programming script
6:12:57which are shorter in code length and
6:12:59easily understandable followed by that
6:13:02the third difference is the data
6:13:04operations are completely tough and
6:13:06complicated in map reduce but in Apache
6:13:10Peg you don't have to worry about those
6:13:12operations because they're already built
6:13:15in next and the last difference between
6:13:19map reduce and Apache pick is Apache map
6:13:22reduce does not allow nested data types
6:13:25whereas Apache Pig allows nested data
6:13:28types so these were the basic
6:13:30differences between Apache map ruce and
6:13:32Apache Pig so the next topic is the pig
6:13:36demo here we will understand all the
6:13:39basic commands in aach Pig and the basic
6:13:42functionalities which are available in
6:13:45pig now without further Ado let's
6:13:48quickly begin with our demo for today's
6:13:50session now we have come back to our
6:13:52Cloud era that we have installed into
6:13:54our local system now let's start pig as
6:13:58we have discussed before Apache Pig can
6:14:00be executed in two modes they are the
6:14:03local mode and the map reduce mode
6:14:05firstly we shall execute an example
6:14:08based on local mode to start pck in
6:14:10local mode you need to type in the
6:14:12command Pig space hyphen X space local
6:14:17firing this command will enable Pig in
6:14:19local mode as you can see the command is
6:14:22getting
6:14:25decrypted now you can see that it has
6:14:28successfully started the gr shell in
6:14:30local mode now what are we going to
6:14:33execute we are going to execute a very
6:14:35simple wordon program for this
6:14:38particular wordon example I have
6:14:41considered a basic text that is the
6:14:43definition of Apache Pig which happens
6:14:45to be Apache pig is a high level
6:14:48platform for creating programs that run
6:14:50on Apache Hadoop the language for this
6:14:53particular platform is called Pig Latin
6:14:55Pig can execute its Hadoop jobs in map
6:14:58reduce Apache test and Apache spark so
6:15:01we will consider this particular
6:15:03paragraph and we will also count the
6:15:05number of words which are included in
6:15:07this particular paragraph now let's
6:15:09quickly go back to our terminal and
6:15:11execute our
6:15:12program so now we have come back to our
6:15:15terminal let's clear it using the
6:15:16command control+ L now let's quickly
6:15:20start up typing our
6:15:22commands so as discussed before we are
6:15:24going to execute this particular program
6:15:26in local mode so here we not loading the
6:15:29data into htfs instead we considering
6:15:31the local
6:15:33location which is the local location
6:15:36that is home Cloud error desktop word
6:15:38count now we're going to run this
6:15:39command and see the output so the data
6:15:41has successfully loaded now we'll
6:15:43execute the next
6:15:45command so now we're going to tokenize
6:15:48each and every single word in this
6:15:50particular paragraph and we'll be
6:15:52considering each and every word as a
6:15:54single word and we are going to to
6:15:56separate each and every word by using a
6:15:58space now let's fire up this command and
6:16:00see the output yeah the command just got
6:16:02executed or the script just got executed
6:16:05now in the next script we are going to
6:16:07group the words according to their
6:16:08occurrence yeah even that is done now
6:16:11the next script would help us to count
6:16:13the number of words which have been
6:16:14repeated so even that is done now the
6:16:17last script would be dumping the output
6:16:19which is stored in p word c which is pig
6:16:22word count now let's enter the command
6:16:24and see the output you can see see some
6:16:26map Ru shops that are getting executed
6:16:28and there you go we have our
6:16:31output so the words are a n on for its
6:16:35days Hado Apache and all those words and
6:16:38along with them we also have the number
6:16:39of repetition of each and every word for
6:16:42example we have Hado which is repeated
6:16:44one time and the word Apache is repeated
6:16:46for four times and so on so with this
6:16:49now let us move ahead and execute some
6:16:51examples based on apache's Hadoop mode
6:16:54or map reduce mode now let's close this
6:16:56terminal and open a new one and let's
6:16:59begin uh executing Hardo in map reduce
6:17:01mode as discussed before apach P can be
6:17:05executed in both local mode and map
6:17:08reduce mode we have already executed
6:17:10some examples based on local mode now we
6:17:13shall execute some examples based on map
6:17:15reduce mode to do so we will first load
6:17:18some local data into
6:17:20htfs so using this command I'll be
6:17:23loading my local data that I expect to
6:17:26tutorial. CSV into my hdfs and I'll name
6:17:29it as Eda in as you can see the command
6:17:32is getting
6:17:33deprecated and the data has been
6:17:35successfully loaded now let us use cat
6:17:38command and see what exactly is present
6:17:40in that particular
6:17:43data as you can see the command got
6:17:46deprecated and the data is a simple CSV
6:17:48file related to students that has ID
6:17:51name department and year now let's start
6:17:55pick in map reduce mode to do so we just
6:17:58need to type in Pig and strike
6:18:02enter there you go we have successfully
6:18:04started Pig in map reduce mode now let's
6:18:07execute some
6:18:13commands so using this command we will
6:18:16be loading edura doin that is the data
6:18:18the CSV file which we have discussed
6:18:19before using pick storage as comma
6:18:22separated file and the schema for the
6:18:24data will be ID as character array name
6:18:27as character array Department as
6:18:29character array and ear as integer array
6:18:33now let's track and enter and see the
6:18:35output there you go the data has been
6:18:37successfully loaded now now let's use
6:18:40dump command to see the data what we
6:18:42have
6:18:47loaded so there you go the data has been
6:18:50successfully dumped so this was the data
6:18:53which we have
6:18:54loaded now let's move further and
6:18:56execute few more
6:19:05commands now we shall use for each
6:19:08command and for each data present in our
6:19:10data file we will generate ID name and
6:19:14Department there you go the command got
6:19:17successfully executed now let's dump the
6:19:19data using dump
6:19:21command so Pig for each is our variable
6:19:25that stores the data now let's type in
6:19:27semicolon and
6:19:35enter so there you go the 4 command has
6:19:38been successfully executed and we have
6:19:41generated ID name and department now we
6:19:44shall execute few more
6:19:50examples now we shall use descending
6:19:52operator to arrange the data in the form
6:19:55of descending order of ID so it's been
6:19:58executed now let's dump the data and see
6:20:00the
6:20:06output there you go we have successfully
6:20:09executed the Dum command and the IDS
6:20:11have been arranged in the descending
6:20:13order now as you can see it now this is
6:20:15how the order by descending function
6:20:17works now let's move ahead and execute
6:20:20one last
6:20:24command
6:20:29as you can see here we are using filter
6:20:31operation we are going to filter the
6:20:33students based on the department where
6:20:35department is equals to csse so
6:20:38executing this command will give us the
6:20:39students which are inside the department
6:20:42C now the command got executed now let's
6:20:45dump the data using dump
6:20:50command so there you go you can see some
6:20:52commands getting
6:20:54deprecated
6:20:56here you can see the map redu shops
6:20:58getting
6:21:00executed so there you go you can finally
6:21:03see the output which has the students
6:21:05that belong to the Cs
6:21:08[Music]
6:21:12Department why exactly we needed Apache
6:21:15Hive it All Began in the early '90s when
6:21:18Facebook started slowly the number of
6:21:20users at Facebook increased that is
6:21:23nearly 1 billion users and along with
6:21:25the users increase the data which is
6:21:27nearly equals to thousands of terabytes
6:21:29of data and nearly one lakh queries then
6:21:32also 500 million photographs uploaded
6:21:35daily and this was a huge amount of data
6:21:38that Facebook had to process and the
6:21:40first thing that everybody had in their
6:21:42mind was to use rdbms and we all know
6:21:45that rdbms couldn't handled such a huge
6:21:48amount of data and neither it was
6:21:50capable enough to process it and the
6:21:52very next big guy who was capable enough
6:21:54to handle all this big data was Hadoop
6:21:58even when Hadoop came into picture it
6:21:59was not too easy to manage all the
6:22:01queries it used to take a lot of time to
6:22:04execute all the queries performed so one
6:22:07common thing that all the Hardo
6:22:09developers had was the SQL so they
6:22:12thought to come up with a new solution
6:22:13that has hadoop's capacity and interface
6:22:16like SQL that is when Hive came into
6:22:19picture so now we understand the exact
6:22:21definition of Apache Hive Apache Hive is
6:22:24a data warehouse soft sofware project
6:22:26built on top of Apache hardup for
6:22:28providing data query and data analysis
6:22:31Hive gives a SQL like interface to query
6:22:33data stored in various databases and
6:22:35file systems that integrate with hadu
6:22:38also Apache Hive has data warehousing
6:22:40software utility it can be used for data
6:22:43analytics it is built for SQL users
6:22:46manages querying of structured data and
6:22:48it simplifies and abstracts the load
6:22:51that is on Hadoop and lastly no need to
6:22:53learn Java and Hadoop a API to handle
6:22:56data using Hive so followed by this we
6:23:00shall understand Apache Hive
6:23:02applications Apache Hive is used in many
6:23:05major applications few of the major
6:23:08applications are as follows Hive is a
6:23:11data warehousing infrastructure for hu
6:23:14the primary responsibility of Hive is to
6:23:16provide data summarization query and
6:23:18data analysis it supports analysis of
6:23:21large data sets in Hero's hdfs as well
6:23:24as on Amazon on S3 file system followed
6:23:27by that we have document indexing with
6:23:30Hive the goal of Hive indexing is to
6:23:33improve the speed of query lookup on
6:23:35certain Columns of a table without an
6:23:37index queries could load an entire table
6:23:40or partition a whole process as rowes
6:23:43this would be Troublesome so with Hive
6:23:46we have solved this problem followed by
6:23:48that predictive modeling the data
6:23:51manager allows you to prepare your data
6:23:54so it can be processed in automated
6:23:56analytics it offers a variety of
6:23:58preparation functionalities including
6:24:00the creation of analytical records and
6:24:02timestamp populations followed by that
6:24:05the next important application of Hive
6:24:08is business
6:24:09intelligence Hive is a data warehousing
6:24:11component of Hadoop and it functions
6:24:14well with structured data enabling ad
6:24:16hoc queries against large transactional
6:24:18data sets hence it happens to be a
6:24:21best-in-class tool available for
6:24:23business intelligence and helps many
6:24:24companies to predict their business
6:24:26requirements with high accuracy last but
6:24:29not the least lock processing Apache
6:24:32Hive is a data warehouse infrastructure
6:24:34built on top of Hardo it allows
6:24:37processing of data with SQL Li queries
6:24:39and is very pluggable so that we can
6:24:41configure it to provide our logs quite
6:24:43easily so these were the few important
6:24:45Hive applications now let us move ahead
6:24:48and understand Apache Hive features the
6:24:51first and the foremost important feature
6:24:53of Apache Hive is see SQL type queries
6:24:57the SQL type queries present on hi will
6:24:59help many of the Hado developers to
6:25:01write queries with ease followed by that
6:25:04the next important feature of Apache
6:25:07Hive is oap based design oap is nothing
6:25:11but online analytical processing this
6:25:14allows users to analyze database
6:25:16information from multiple database
6:25:18systems at one time so using Apache Hive
6:25:22we can achieve o AP with higher accuracy
6:25:25followed by the second feature we have
6:25:27the third feature which says Apache Hive
6:25:29is really fast since we have SQL like
6:25:32interface in Apache Hive using this
6:25:35feature on sdfs will help us writing
6:25:38queries faster and executing them
6:25:40followed by that we believe Apache Hive
6:25:43is highly scalable Hive tables are
6:25:45defined directly in Hardo file system
6:25:48hence Hive is fast and scalable and easy
6:25:51to learn followed by that it is known to
6:25:54be highly extensible
6:25:56Apache Hive uses Hadoop file system and
6:25:59Hadoop file systems or hdfs provides
6:26:01horizontal extensibility and finally the
6:26:05ad hoc wearing using H we can execute ad
6:26:08hoc wearing to analyze and predict data
6:26:11so these were the few important features
6:26:12of Apache Hive let us move on to our
6:26:15next topic where we deal with Apache
6:26:17Hive architecture the following
6:26:19architecture explains the flow of
6:26:21submission of query into hiy the first
6:26:24stage is The Hive client Hive allows
6:26:27writing applications in various
6:26:29languages including Java Python and C++
6:26:33it supports different types of clients
6:26:35such as Thrift server jdbc driver and
6:26:38odbc driver so what exactly is Thrift
6:26:42server it is a cross language service
6:26:44provider platform that serves the
6:26:46request from all these programming
6:26:48languages that supports Thrift followed
6:26:51by that jdbc driver it is used to
6:26:54establish connection between Hive and
6:26:56Java applications the jdbc driver is
6:26:59present in the class or. Apache dohad
6:27:02doh. jdbc dohy driver finally we come to
6:27:07odbc driver so what exactly is odbc
6:27:10driver obbc driver allows the
6:27:13applications that support obbc protocol
6:27:15to connect to Hive followed by that we
6:27:18have the hive Services the following are
6:27:21the services provided by hve they are
6:27:24hve C Li Hive web user interface Hive
6:27:28meta store Hive server Hive driver Hive
6:27:31compiler and lastly The Hive execution
6:27:34engine The Hive CLI or command line
6:27:37interface is a shell where we can
6:27:40execute the hive queries and commands
6:27:42followed by that the hive web UI is just
6:27:46an alternative for Hive CLI it provides
6:27:49a web-based graphical user interface for
6:27:52executing High queries and commands
6:27:54follow by that the hive meta store it is
6:27:57a central respository that stores all
6:28:00the structured information of various
6:28:01tables and partitions in the warehouse
6:28:05it also includes metadata of column and
6:28:07its type information the serializers and
6:28:10D serializers which is used to read and
6:28:13write data and the corresponding hdfs
6:28:15files where the data is stored followed
6:28:18by that the H server it is referred to
6:28:21as Apache th server it accepts the
6:28:24request from different clients and
6:28:26provides to The Hive driver moving on we
6:28:28shall deal with Hive driver The Hive
6:28:31driver receives queries from different
6:28:32sources such as web UI CLI Thrift and
6:28:36jdbc or odbc drivers it transfers the
6:28:39queries to the compiler followed by that
6:28:42we have the hive compiler the purpose of
6:28:45Hive compiler is to pass the query and
6:28:48perform semantic analysis on the
6:28:49different query blocks and expressions
6:28:52it converts hiveql statements into
6:28:55produce jobs finally we have Hive
6:28:58execution engine Hive execution engine
6:29:01is the optimizer that generates the
6:29:03logical plan in the form of dag or
6:29:06directed aycc graph of map redu task and
6:29:09hdfs tasks in the end the execution
6:29:12engine executes the incoming task in the
6:29:14order of their dependencies followed by
6:29:17that we have the map reduce and hdfs map
6:29:20reduce is the processing layer which
6:29:22executes the mapping and reducing jobs
6:29:24on the data provided lastly the sdfs or
6:29:28Hardo distributed file system is the
6:29:30location where the data which we provide
6:29:32is stored so this is the architecture of
6:29:35Apache Hive then moving next we have
6:29:38Apache Hive components so what are the
6:29:41different components which are present
6:29:42in Hive they are first one the Shell
6:29:46Shell is the place where we write our
6:29:48queries and execute them followed by
6:29:51that we have metast store as discussed
6:29:54in the AR Ure the metast store is a
6:29:56place where all the details related to
6:29:58our tables is stored like schema Etc
6:30:02followed by that we have the execution
6:30:04engine so execution engine is the
6:30:06component of Apache Hive which converts
6:30:09the query or the code which we have
6:30:11written into the language which The Hive
6:30:14can understand followed by that driver
6:30:16is the component which executes the code
6:30:19or query in the form of acyclic graphs
6:30:22and lastly the compiler compiler
6:30:25compiles whatever the code we write and
6:30:27executes and provides us the output so
6:30:30these are the major Hive components
6:30:32moving ahead we shall understand Apache
6:30:34Hive installation on Windows operating
6:30:37system so urea is all about providing
6:30:39the technical knowledge in the simplest
6:30:41way as possible and later play around
6:30:43with the Technologies to understand the
6:30:45complicated parts of it so now let's try
6:30:48to install hyve into our local system in
6:30:50the most simplest way as possible to do
6:30:53so we might need the Oracle virtual box
6:30:56which looks like this so once after you
6:30:58download Oracle virtual box and install
6:31:01it into your local system The Next Step
6:31:03would be to download the cloud era
6:31:05quickart VM for your local system the
6:31:07link to this will be provided in the
6:31:09description box below now let's quickly
6:31:11start our Cloud era quick start VM with
6:31:14our Oracle virtual box select import
6:31:17option and now provide the location
6:31:19where your cloud data quick start VM is
6:31:21existing in my local system it is in the
6:31:24local desk Drive
6:31:28F there you go select open and now just
6:31:33make sure your Ram size is more than 8GB
6:31:36just randomly I'm providing 9,000 MB
6:31:39which is just above 8GB so that you have
6:31:41a smooth functionality of cloud erra now
6:31:44select
6:31:45import and there you go you can see that
6:31:48cloud era quick start VM is getting
6:31:53imported
6:32:00now you can see that cloud quickart VM
6:32:02has been successfully imported and it's
6:32:04ready for deployment you can just double
6:32:06click on it and it'll get
6:32:17started you can see that cloud era VM
6:32:20has been successfully imported and it
6:32:22started and also you can see that we
6:32:24have gone live on cloud era you can see
6:32:26all the hu Hado Edge space Impala spark
6:32:29which are pre-installed in Cloud era now
6:32:31our concern would be to start up hve so
6:32:34to start hve you need to start up Hue
6:32:37first so let me remind you one thing in
6:32:40Cloud era every single password and
6:32:43username is cloud era by default so for
6:32:46example we've got H username and
6:32:49password here so the username that is
6:32:51the default username for cloud eras h
6:32:54would be Cloud error and along with that
6:32:57even the password will be Cloud error
6:32:59that is by default so we have got Cloud
6:33:02error and Cloud error as username and
6:33:04password respectively let's just sign in
6:33:07you may select remember option in case
6:33:10if you forget your
6:33:12passwords so now we are getting
6:33:14connected to Hugh and we are live on
6:33:16Hugh
6:33:17now there you go we've got started our
6:33:19Hue so now we'll enter into
6:33:23htfs there we go we have a hive
6:33:27here now that we have successfully
6:33:29installed Hive into our local system let
6:33:32us move further and understand few more
6:33:33Concepts in Hadoop firstly we should
6:33:36deal with the data types the data types
6:33:38are completely similar to any other
6:33:40programming language which we have they
6:33:41are tiny end small end integer big end
6:33:45similarly followed by that we have float
6:33:47and in Side High float is used for
6:33:49single precision and if you want double
6:33:51Precision you can go ahead with double
6:33:54and followed by that we have a string
6:33:56and Boolean which are completely similar
6:33:58to any other programming languages which
6:34:00we use in this daily life followed by
6:34:03that we have hve data models so these
6:34:05are the basic data models which we use
6:34:07in Hive that we basically create
6:34:09databases and store our data in the form
6:34:11of tables and sometimes we also need
6:34:14partitions we will discuss each one of
6:34:16these data models in our demo ahead so
6:34:19we'll first create databases and inside
6:34:21databases we will be creating tables
6:34:23inside which we will will be storing
6:34:25data in the form of rows and columns and
6:34:27along with that partitions partitions uh
6:34:30they are like Advanced way of storing
6:34:33data like if you have just imagine you
6:34:35are in a school say standard one and
6:34:38inside standard one you have sections a
6:34:40b c d so partition is like you're
6:34:42getting partitions for Section a section
6:34:44B section c and section D you're storing
6:34:47different different students in
6:34:48different different sections so that
6:34:50when you're querying for a particular
6:34:51data for example say you're searching
6:34:54for a kid called Sam and you have the
6:34:56section of his class SB so you just
6:34:59don't have to just search for Sam in all
6:35:02the four sections you can just directly
6:35:04go into section B and call in Sam and
6:35:07you'll get access to him that's how
6:35:09partitions work followed by partitions
6:35:11we have buckets so similar to partitions
6:35:14even buckets work in the same way let's
6:35:16understand each one of these in much
6:35:17better way through a practical demo
6:35:20after data models we shall understand
6:35:22about hype operators so what are
6:35:24operators operators are any other
6:35:26operators that we use in normal
6:35:28programming languages such as arithmatic
6:35:30operators logical operators we shall
6:35:32also go through some examples based on
6:35:34arithmatic and logical operators in Hive
6:35:37in the hive demo we will use some
6:35:39arithmatic operations as well as logical
6:35:42operations on the data which we have
6:35:43stored in the form of tables in Hive we
6:35:45shall go through a brief look on that as
6:35:47well so before we get started let's have
6:35:50a brief look on the CSV files that I
6:35:52have created for today's demo these are
6:35:54the small CSV files that I've personally
6:35:55created using msxl and I've saved them
6:35:58as CSV files I've made the CSV files to
6:36:01be smaller because just to make sure the
6:36:04execution time consumed is as less as
6:36:06possible since we using Cloud era the
6:36:08execution time might be a little more so
6:36:11it's better we use smaller CSV files so
6:36:14this is my first CSV file which is
6:36:16employee. CSV which has employ IDs
6:36:18employee name salary and age similarly
6:36:21we have another employee 2. CSV file
6:36:24which has the same details along with
6:36:25one more column that is the country
6:36:27column I've included country because we
6:36:29will be using this country column in
6:36:31Joins that we will be performing in
6:36:33future followed by that we have the
6:36:35department so here we have Department ID
6:36:37and Department name so we have
6:36:40Development Department testing product
6:36:42relationship admin and it support
6:36:44similarly we also have student CSV this
6:36:47is another CSP file that I have created
6:36:49this has ID name course and age of the
6:36:52student followed by that we have another
6:36:54CSV this is student report. CSV which
6:36:57has the reports of a particular student
6:37:00gender ttic City parental education
6:37:02lunch course math score reading score
6:37:05writing score and other so these are the
6:37:07CSV files that we will be using in our
6:37:09demo today so now let's quickly begin
6:37:11with our demo so to start Hive we shall
6:37:14open a terminal so starting or firing up
6:37:17Hive in Cloud arise really simple you
6:37:19just have to type in Hive and enter
6:37:22there you go logging initialized using
6:37:25configuration files and Etc the hi CLI
6:37:28is deprecated and migration to bline is
6:37:31recommended and there you go your hi
6:37:33terminal or CLI has been started so
6:37:36first let's try to create a database to
6:37:39save time I've already created the
6:37:41document which has all the codes that we
6:37:43will be executing today so this is the
6:37:46particular file which I will be using
6:37:47today so don't worry this file will be U
6:37:50Linked In the description box below you
6:37:52can use the same file and try executing
6:37:54the same codes in your personal systems
6:37:56just for practice if you feel so so just
6:37:59to save time I've already created uh the
6:38:01document which has the codes that we are
6:38:03going to execute today so this code or
6:38:05this file will be attached in the
6:38:07description box below you can get access
6:38:09to it and you can also execute the same
6:38:11codes in your own personal system to
6:38:13have a practical experience about this
6:38:15particular Hy tutorial so the first
6:38:17thing that we will be doing today is to
6:38:19create a database so I'm going to create
6:38:22the database using SQL type commands
6:38:25which are create database name of the
6:38:27database which is Eda there you go the
6:38:30database has been successfully
6:38:32created so now you can also use the
6:38:35following command to check if your
6:38:37database has been created or not so show
6:38:39databases will help you to find it so
6:38:42there you go you can see the first
6:38:43database which is a default database
6:38:46which will be pre-existing and followed
6:38:48by that you have our own database which
6:38:49we have created now that is edura so
6:38:52followed by this next we will move ahead
6:38:54and try to create a new table so when
6:38:57you come into tables you need to
6:38:59understand there are two types of tables
6:39:01in Hive they are managed tables or
6:39:04internal tables followed by that
6:39:06external tables so what is the
6:39:08difference between these two tables so
6:39:11internal table or manage table is the
6:39:13default table that will be created
6:39:15whenever you try to create a table in
6:39:17high so for example if you're trying to
6:39:19create a new table say Eda then hi
6:39:24considers that particular table as an
6:39:26internal table by default so when you
6:39:29create an internal table your data is
6:39:32not secured understand this so when you
6:39:34create an internal table your data is
6:39:36not secured in case just imagine you are
6:39:39working with a team and all your team
6:39:42members have access to your hive or hue
6:39:45so the table has been existing in your
6:39:47hve and some random newbie or some
6:39:50random inexperienced guy tries to change
6:39:53few things in your table and
6:39:55accidentally he ends up deleting the
6:39:57table so when you delete the table then
6:40:00if the table was created using an
6:40:02internal table code then your data will
6:40:04be erased so that's the disadvantage of
6:40:07using internal tables but in case if you
6:40:10create an external table even if
6:40:12somebody tries to delete your table the
6:40:15table or the data whatever is there will
6:40:18be deleted from their own local system
6:40:20but not from Hy so that's the best part
6:40:22of using external tables don't worry we
6:40:24will discuss about internal tables and
6:40:27external tables as well so first we'll
6:40:29try to create an internal table so this
6:40:31particular code is based on internal
6:40:33tables so we are using SQL type command
6:40:36here which is create table and the table
6:40:38name is employee and the columns inside
6:40:41our table are ID of the employee name of
6:40:44the employee salary and age so row
6:40:47format has been delimited followed by
6:40:49that since this is a CSV file so the fs
6:40:52will be terminated by comma
6:40:54and don't forget you have to use
6:40:56semicolon unless you use semicolon your
6:40:58code is not complete so let's fire and
6:41:01enter and see if the table gets created
6:41:02or not yeah the table is created
6:41:05successfully now we shall see the table
6:41:08or let's describe the table so
6:41:10describing the table means you can see
6:41:12what are the columns which are present
6:41:13in your table so to describe a table you
6:41:15can use the keyword describe a name of
6:41:17the table which is employee and don't
6:41:20forget semicolon there you go so your
6:41:22table has the column ID name salary age
6:41:26so those are the four columns which you
6:41:27have included in your particular table
6:41:29employee now let's move ahead and see if
6:41:32this particular table is an internal
6:41:34table or manage table or the other type
6:41:36of table which is the external table so
6:41:39to do that we can just write and
6:41:41describe formatted table name and
6:41:45semicolon there might be a small issue
6:41:48here yeah there is a typing mistake that
6:41:51is described missed s so there you go we
6:41:55got it so this particular table is
6:41:58managed table as you can see here now
6:42:01let's move ahead and uh try out external
6:42:04tables let's clear our screen first you
6:42:07can use control+ l to clear your screen
6:42:10there you go we have a clear screen now
6:42:12now let's try to create an external
6:42:14table creating an external table is
6:42:17completely similar to that of internal
6:42:19table but the only difference is that
6:42:21you need to add in a keyword which is
6:42:24external so this particular keyword is
6:42:26used to create an external table now
6:42:29let's fire an enter and see if the table
6:42:30gets created or not you can see the
6:42:33table got created now let's try to
6:42:35describe the table employee 2 don't
6:42:38forget the semicolon I'm saying this
6:42:40again and again because most of the
6:42:42times we miss semicolon and we will get
6:42:45an error so you can see the table got
6:42:46described and we have the following
6:42:48columns inside our table now let's move
6:42:51ahead and see if this particular table
6:42:53is an external table or a manage table
6:42:57to do so you can type in describe
6:42:59formatted the same code what we have
6:43:01used earlier that is described formatted
6:43:04name of the table that is employ to
6:43:07semicolon don't forget there is some
6:43:10issue again I think I've missed
6:43:11something or maybe a typing error yeah
6:43:14this is a typing
6:43:17error yeah there you go the table type
6:43:20is external table so that's how we
6:43:23create an internal table or manage table
6:43:26and external table so now that we have
6:43:29understood how to create a database and
6:43:31table and the two types of tables that
6:43:34are internal table or managed table
6:43:36followed by that the second type of
6:43:38table that is the external table now
6:43:40let's try to create an external table in
6:43:42a particular location so for that you
6:43:45can use the following code but the only
6:43:47difference is you are specifying the
6:43:49location that is user Cloud era urea
6:43:52employee edu EMP is a file that we will
6:43:55be creating in our hyp so let's fire and
6:43:58enter and see it if it's created or not
6:44:00yeah it's successfully created let's go
6:44:03back to Hue and see if the following
6:44:05table is created or not so one thing you
6:44:07have to remember is when you fire in a
6:44:09command or if you try to create a table
6:44:12the first folder that will be created is
6:44:14a warehouse so inside hve you have your
6:44:17warehouse and inside Warehouse you have
6:44:19all the databases that we have created
6:44:21our first database was the Ed Raa
6:44:23database and after that we have created
6:44:26table which is employee and the second
6:44:28table is employee 2 so this is in the
6:44:31particular location which is user Cloud
6:44:33error and the file is employed to let's
6:44:36see that this was the
6:44:42file yeah sometimes H will not show it
6:44:45because of network issues you don't have
6:44:47to worry about it you will get back that
6:44:49data now followed by this let us enter
6:44:52into Hue again
6:44:54so when you come back into Hue if you
6:44:56have to upload a file into Hue you can
6:44:58just select this particular option which
6:45:00is plus so selecting this will give you
6:45:02a dialogue box which will be something
6:45:04like this and here you can just select
6:45:06any of the files which you want to
6:45:08upload into Hue now let me select a
6:45:10student report. CSV and select open so
6:45:14there you go upload is in progress so
6:45:16the data file has been successfully
6:45:17uploaded now if you want to access your
6:45:19data file you can just click on that so
6:45:22there you go you have all your data
6:45:24successfully loaded onto
6:45:26Hue you can also perform queries on this
6:45:29particular data you can just select
6:45:31query and inside that you just need to
6:45:33select editor and you have various
6:45:35editors over here which is Peg Impala
6:45:38Java spark map reduce shell scoop and we
6:45:42also have hi in here so if you just
6:45:44select Hive and there you go you have
6:45:46the editor here you can just type in
6:45:48your commands or queries whatever you
6:45:50have see you have many dictionaries as
6:45:53well you can just select any one of
6:45:54those select and that's how you write
6:45:56queries on the hyp terminal now let's
6:46:00not waste much time here and we have a
6:46:02lot to learn so let's continue with the
6:46:05next topics in our today's session now
6:46:07we shall try to edit the
6:46:10tables now we have created the new table
6:46:13that is employee 3 and we have named The
6:46:16Columns as ID name String salary age and
6:46:19Float now we should try to make some
6:46:22alterations to our table
6:46:24so the first alteration that we will try
6:46:26to make to our table is to rename our
6:46:29table as EMP table you know that our
6:46:32employee table was named as employee 3
6:46:35now we trying to rename it to EMP table
6:46:38so we are using the keyword alter here
6:46:41so just fire and enter and see if this
6:46:43is possible or not yeah it is possible
6:46:45the name has been changed to EMP table
6:46:48now let's try if it's completely changed
6:46:50or clearly changed or not you can just
6:46:52type in describe EMP table semicolon if
6:46:57we get the same column names in our
6:47:01description then it should be changed so
6:47:04there you go we can see the same columns
6:47:06here so we have successfully changed the
6:47:08name to EMP table now we shall also try
6:47:11to add in some more columns to our table
6:47:14which is EMP table so here we'll try to
6:47:17add in a new column that is the surname
6:47:20of string data type so I'm doing that by
6:47:23using the keyword alter followed by that
6:47:26table uh the table name is EMP table and
6:47:29I'm using the keyword add columns and
6:47:32the column name is surname and the data
6:47:34type of that column is string so now
6:47:36let's fire in enter and there you go we
6:47:39have successfully added a new row to our
6:47:41table now let's try to describe our
6:47:43table again and see if the column has
6:47:45been successfully added or not there you
6:47:49go you can see the last row which is the
6:47:51surname that we have added most recently
6:47:54so this is how you can alter the table
6:47:57and you can also change the names of the
6:48:00existing columns let's try to do that
6:48:02one as well now what I'm doing is I'm
6:48:05changing the column name to first name
6:48:09so one of the column name in my table
6:48:11EMP table is the name which gives me the
6:48:14names of the employees so since I added
6:48:17surname I'll change this column name
6:48:20from name to first name so this is the
6:48:24command that I'm using for that
6:48:26operation right now let's fire in enter
6:48:28and see the result yeah the change is
6:48:31been made Let's describe our
6:48:34table don't forget the semicolon there
6:48:38you go you can see that earlier we had
6:48:41name now it's been changed to first name
6:48:44and we also have a surname let's clear
6:48:46our screen so that's all for alterations
6:48:50now we shall move ahead into our next
6:48:52Main topic or the data model which is
6:48:56partitioning so we have dealt with the
6:48:58first two data models that are databases
6:49:02and tables so we have learned how to
6:49:04create a database and we have learned
6:49:07how to create a table we have learned
6:49:09how to create internal or manage table
6:49:12and also we have created external table
6:49:15and also we have learned how to create
6:49:16an external table in a particular
6:49:18location in your hive and load data to
6:49:21your table and also how to alterate your
6:49:24tables the column names the name of your
6:49:27table and how to add or delete new
6:49:30columns to your table so far so good and
6:49:32now we shall continue with the next type
6:49:35of data model that is the
6:49:37partitioning as we have discussed
6:49:38earlier about partitioning it's
6:49:40completely similar to a school or a
6:49:44college just imagine that you are in a
6:49:46college and you are in computer science
6:49:49section so a college has many branches
6:49:53so maybe computer science mechanical and
6:49:56electronics and Communications so
6:49:59imagine your name is Harry so if someone
6:50:02comes to your college and if is looking
6:50:04for Harry so there are many harri's in
6:50:06your school so if the person is asking
6:50:09specifically about you that is Harry
6:50:11from computer science then can you
6:50:14imagine how simple is this query so you
6:50:17don't have to search for electronics and
6:50:19mechanical you just have to come into
6:50:20the class computer science and search
6:50:22for for Harry and there you go you are
6:50:24present so that's how partitions work to
6:50:27execute commands or to execute queries
6:50:29on partition we will create a whole new
6:50:32database here let's start everything
6:50:34from fresh so we'll create a separate
6:50:37database for executing our new data
6:50:39model that is partitioning so I'm
6:50:40creating a new database that is Eda
6:50:43student so there you go the database has
6:50:45been successfully created followed by
6:50:48that let's use this database now to use
6:50:51the database you just need to to add in
6:50:53the keyword use and name of the database
6:50:56so let's fire in enter and now we are
6:50:59currently using Eda student database now
6:51:02let's create a table in urea student
6:51:05database so here I'm creating a normal
6:51:09table that is the manage table so inside
6:51:12my student table I'll be having uh some
6:51:14basic columns such as ID number of the
6:51:17student name of the student what is his
6:51:20age and course so you're not find
6:51:22finding course here because I'm going to
6:51:24partition the table based on course so
6:51:28here you can find the course I'm using
6:51:30the keyword partitioned and on what
6:51:33terms so on the terms of course I'm
6:51:35going to partition students so we have
6:51:38discussed about our students CSV file
6:51:41right so here we have our CSV file and
6:51:44the courses that this particular
6:51:46Institute is offering are Hadoop Java
6:51:49Python and yeah so these are the courses
6:51:52that this particular Institute is
6:51:54offering so I'm going to categorize or
6:51:57I'm going to partition these students
6:51:58based on their courses so this is how
6:52:01I'll be partitioning them using this
6:52:03following code so basically the table
6:52:06has all the columns and I'm going to
6:52:07partition the table using course so
6:52:10let's fire and enter and see the
6:52:11execution of this particular code the
6:52:13partition has been done now all we have
6:52:15to do is try to load in our
6:52:19data before that let's try to describe
6:52:22it let's try to see what are the columns
6:52:25present in our particular table student
6:52:28so as you can see the course column is
6:52:30present don't worry the code looks that
6:52:33we have missed out course but we did not
6:52:34miss the course column it is present in
6:52:37the table the only thing is that we have
6:52:40just partitioned it based on the course
6:52:42that we are going to offer now let's try
6:52:44to categorize the students based on
6:52:46their course so you can do that by using
6:52:49the following code we going to load the
6:52:51data use using the command load data
6:52:54local impath so this particular folder
6:52:57that is the student. CSV is in my local
6:53:00location so that is a home Cloud era
6:53:03desktop student. CSV and I'm loading the
6:53:06data present in this particular location
6:53:09into the student which is present in hi
6:53:11right now so I'm going to partition the
6:53:13student based on their course Hado now
6:53:15let's fire in this command and see the
6:53:17output yeah now you can see some map
6:53:19reduce jobs taking place yeah the data
6:53:22has been successful successfully
6:53:23loaded let's now refresh our Hive you
6:53:28can refresh your hive or hue based on
6:53:31two methods the first one is just
6:53:33clicking refresh button on the URL or
6:53:35you can also select an manual refresh
6:53:38this is the manual refresh and there you
6:53:40go it's done you can see the new
6:53:42database that is the urea student
6:53:44database that we have right now created
6:53:47and inside that you can see the student
6:53:48table that we have created and there you
6:53:50go we have the file of students based on
6:53:55course Hadoop now we will try to add in
6:53:58few more students based on the course
6:54:00Java for that all you need to do is just
6:54:03replace the course name with
6:54:05Java there you go here we had Hado
6:54:08course and now here we have Java course
6:54:11just fire and enter and you can see the
6:54:13output followed by that we also had
6:54:15another course that is python so let's
6:54:18also execute a code for that there you
6:54:20go python so now we have uploaded
6:54:23student details into our Hive and we
6:54:25have also partitioned using one of our
6:54:28data models that as partition into three
6:54:30categories that are based on Hadoop Java
6:54:32and python now let's go back to our Hue
6:54:35and see if the three categories are done
6:54:38or not yeah you need to refresh
6:54:41that there you go you have successfully
6:54:43refreshed still there is no sign of java
6:54:46and python maybe a manual refresh could
6:54:50help yeah the manual refresh has
6:54:52resulted in the two new files which are
6:54:56Java and python so you have all the
6:54:58three partitions here Hado Java and
6:55:00python just enter them and you can see
6:55:02the student details and now that we have
6:55:05understood partitioning sorry I forgot
6:55:08to mention we have two types of
6:55:10partitioning which are Dynamic
6:55:12partitioning and static partitioning so
6:55:15uh the static partitioning is in static
6:55:18or manual partitioning it is required to
6:55:20pass the values of partition columns
6:55:22manually while loading the data into the
6:55:25table hence the data file does not
6:55:27contain partitioned columns you can see
6:55:30that we have sent the partition columns
6:55:32manually for python Java and had but
6:55:36when it comes to Dynamic partitioning
6:55:37you just need to do it once and all the
6:55:40three files will be automatically
6:55:41configured and the files will be created
6:55:44so now what is U Dynamic partitioning so
6:55:47uh Dynamic partitioning the values of
6:55:50partition columns exist within the table
6:55:53so it is not required to pass the values
6:55:55of partition columns manually now what
6:55:58is this no worry we shall execute the
6:56:00code based on Dynamic partitioning and
6:56:02we shall understand this in a much
6:56:03better way now let's clear our screen
6:56:06now let's start fresh again let's try to
6:56:09create a new database
6:56:11for dynamic partitioning and let's start
6:56:14again fresh so here we'll be creating a
6:56:17new database that is edura student 2 so
6:56:21earlier we created Eda student and now
6:56:24we'll be testing our Dynamic
6:56:26partitioning on our new database that is
6:56:28urea student 2 so there you go the
6:56:30database has been successfully created
6:56:33now we should use this particular
6:56:35database currently we were in urea
6:56:38student to one database now we'll enter
6:56:40into student 2 database so we'll use it
6:56:43now now we are in Eda student 2 now
6:56:46before we start up with Dynamic
6:56:48partitioning we have to set High
6:56:50execution to Dynamic part is equals to
6:56:53true because by default the partitions
6:56:56that will be taking place in Hive will
6:56:57be static so we need to convert that
6:57:00into Dynamic Partition by specifying
6:57:02this particular code now we are good to
6:57:05go with Dynamic partitioning along with
6:57:07that we need to execute another command
6:57:09which says partition mode would be
6:57:12non-strict so by default when you are
6:57:15partitioning using the static partition
6:57:18the partition mode will be strict so now
6:57:20you're specifying it to be non-strict
6:57:23now let's execute this so there you go
6:57:25we have executed the two required codes
6:57:27for that now let's create a new table so
6:57:31the name of the table will be edura
6:57:34student that is edu St and this will
6:57:37have the same columns which are the ID
6:57:40of the student name of the student
6:57:42course age
6:57:44Etc now we will try to load in the data
6:57:47from our local path that is home cloud
6:57:49or our desktop student. CSV into the
6:57:51table edu stud so the data has been
6:57:55successfully loaded and the size is 267
6:57:58KB number of files is
6:58:02one now comes the tough part so here we
6:58:05are going to partition so we will be
6:58:07partitioning the table based on the same
6:58:09thing which is the course and we will be
6:58:12separating the data using the comma now
6:58:15let's fire and enter now the table has
6:58:18been separated based on course and now
6:58:22we will be loading the data to this
6:58:24particular table which is the student
6:58:26part so this particular table that we
6:58:29have created based on Dynamic
6:58:30partitioning and we are going to
6:58:32partition the data based on course now
6:58:35it's been created so the student part
6:58:37table has been successfully created now
6:58:39the only part remaining is to load the
6:58:41data to this particular table now we
6:58:44will be writing a code so using that
6:58:47code the map reduce will automatically
6:58:50segregate the data members or the
6:58:53students based on their courses so the
6:58:57guys which are in Hardo will be
6:58:58separated guys in Java will be separated
6:59:01and loaded into different file and
6:59:03similarly with python now let's see um
6:59:05how to do it using the
6:59:08code so there you go we are going to
6:59:11insert into student part partition based
6:59:14on course select ID name course age from
6:59:17the table at your so uh the data will be
6:59:20imported from the table what we have
6:59:22created here that is urea student so
6:59:25this particular location has the
6:59:28student. CSV file now let's fire in
6:59:31enter and see if it's created or not
6:59:34fine you can see some of the map produce
6:59:36jobs are getting executed you can see we
6:59:39have three jobs so first one is getting
6:59:42executed we have three because one is
6:59:44for Hado one is for Java and one for
6:59:49python so this will take a little time
6:59:52so this is the reason why I have chosen
6:59:55smaller CSV file so to save time when
6:59:58you take up the course from at your then
7:00:00you can work on realtime data so that
7:00:03you get hands-on experience from real
7:00:05time and you can get yourself placed in
7:00:06some good companies with the experience
7:00:09what you gain from this particular
7:00:10course so the stages have been
7:00:13successfully finished and the data has
7:00:14been loaded now let's see what are the
7:00:17datas present in the particular table
7:00:19student part there you go you have the
7:00:21output executed here so these are the
7:00:24data members present in the partition
7:00:26student part so these are the data
7:00:28members which are separated based on
7:00:29their courses that is the partition
7:00:31based on their courses that is had Java
7:00:33and python so now that we have
7:00:35understood Dynamic partitioning and um
7:00:38static partitioning we shall move ahead
7:00:41into the last type of data model which
7:00:44is bucketing Once after we finished the
7:00:47bucketing we shall enter into some query
7:00:50functionalities of hive or query
7:00:52operations which can be performed in
7:00:54Hive and followed by that we will also
7:00:56learn some functions which are present
7:00:58in Hive and some of the other things
7:01:01like group buy order bu sord bu and
7:01:04finally we shall wind up the session
7:01:06with joins which are available in Hive
7:01:09for now let's get continued with
7:01:11bucketing the last type of data model
7:01:14present in Hive so for that um let's
7:01:16again start fresh we shall create a new
7:01:18database for that before that let's go
7:01:21back to um H and check if our partition
7:01:24has been made or not let's
7:01:27refresh also let us make a manual
7:01:34refresh so our database Wasa student 2
7:01:39database and inside that we have the
7:01:42table that is student part and there you
7:01:44go you can see the files which are based
7:01:47on the partition so 22 is for a
7:01:50different course 23 is for a different
7:01:52course and 24 is for a different course
7:01:55and this is the default partition which
7:01:58has all the data members as we discussed
7:02:00earlier now let's start with the last
7:02:03data model in Hive that is bucket now we
7:02:06have created a new database that is Eda
7:02:09bucket now we shall also create a new
7:02:12table for that before that we need to
7:02:15start with this particular database so
7:02:16we can use the command use edura bucket
7:02:19now we are in Eda bucket now let's
7:02:22create a new table so the table name
7:02:25will be at youra bucket and it will be
7:02:28containing the ID name salary age of the
7:02:31employees the table is created now let's
7:02:34try to load the data so the data file
7:02:37that we will be using is the same one
7:02:39that is the employee. CSV so the data
7:02:41has been successfully loaded into the
7:02:44location now comes to the major part
7:02:47that is the bucketing part so to start
7:02:50bucketing in high we need to use the
7:02:51command and set hive. info. bucketing is
7:02:54equals to true so that's done now we
7:02:58will cluster or classify the data
7:03:00present in this particular file using
7:03:04this particular code so we will be
7:03:06clustering based on the ID and we will
7:03:09be categorizing them into three
7:03:10different buckets so let's fire in this
7:03:13command and see if it's happens yeah
7:03:15that's successfully done now we will
7:03:17overwrite the data using the following
7:03:20command now we'll be inserting data into
7:03:23this buckets that we have made that is
7:03:25three buckets and we will override the
7:03:27table using this particular code there
7:03:30you go you can see some map reduce jobs
7:03:32to be taken care of
7:03:35now so one MPP and red users are three
7:03:39for now so stage one is getting done so
7:03:43we should be having three task basically
7:03:47so let's see what's the output stage one
7:03:50is finished
7:03:52the process is finished and data has
7:03:54been successfully
7:03:55inserted now let's go back to Hive and
7:03:58check if it's done or not so before that
7:04:01let's do a
7:04:04refresh now a manual refresh would be
7:04:07much better there you go we have our
7:04:09database here which is edura bucket and
7:04:12inside edura bucket we have EMP bucket
7:04:16and that's our data employee. CSV there
7:04:20you go now let's move ahead and
7:04:24understand the basic operations we can
7:04:26perform in Hive so for that let's start
7:04:29fresh again let's create a new database
7:04:32I'm creating a new database for each and
7:04:34every option or each and every operation
7:04:36that I'm performing in this particular
7:04:38tutorial just to make things or keep
7:04:40things in a sorted manner so as you can
7:04:44see here in our particular file
7:04:47system I have separated each and
7:04:50everything like I have sorted everything
7:04:52so for bucketing I've got a separate
7:04:54database and for partitioning I've got a
7:04:57separate database and for understanding
7:04:59how to create database and tables I've
7:05:01got a separate database for that just to
7:05:03keep things arranged and sorted this
7:05:05looks uh in a much better way so now
7:05:08let's discuss about the operations that
7:05:10we could perform in h so I'm creating a
7:05:13new database again for this so the
7:05:15database would be Hive query language
7:05:19now let's use this particular database
7:05:21this creates a habit of learning things
7:05:24in a better way or it's like a revision
7:05:27for the things what you have performed
7:05:28or learned so far as you can see the
7:05:31table is been successfully created now
7:05:33let's try to add in some data into this
7:05:36particular location that is employee
7:05:38data it's been successfully loaded now
7:05:41let's try to see what are the details
7:05:43present in this particular file we can
7:05:45use in the command select star from the
7:05:47table urea employee so there you go
7:05:50these are the details or information
7:05:52present in the table meta employee now
7:05:56we shall see what are the functions that
7:05:58we can perform on this particular file
7:06:00so since we discussed that the
7:06:02mathematical operations and logical
7:06:05operations can be performed on H so
7:06:07let's try to perform an addition
7:06:09operation so I'm selecting the column
7:06:12salary and as we have seen here the
7:06:14salaries are 25 30 40 20,000 rupees for
7:06:19every employee now let me add in 5 ,000
7:06:21more to each and every employee so I'm
7:06:23adding uh the value 5,000 by using the
7:06:26addition operation so let's enter you
7:06:29can see we have added 5,000 so the first
7:06:32element was 25 now it's 30 so similarly
7:06:35all the other employees got 5,000 rupees
7:06:38hike all of a sudden now let's try to
7:06:40remove 1,000 so to do so all you need to
7:06:43do is uh replace the addition operation
7:06:47with a subtraction operation that is
7:06:49minus fire and enter and they go each
7:06:52and every employee lost 1,000 so the
7:06:55initial amount was 25,000 so removing
7:06:571,000 from that will result in 24 so
7:07:00this is considering the first initial
7:07:02values so this is how it's working uh
7:07:04followed by that let's also perform some
7:07:07logical operations let's clear the
7:07:10screen and yeah here I'm fetching for
7:07:12the employees who are having a salary
7:07:15equal to or greater than
7:07:1725,000 so these are the employees which
7:07:19are having the salaries above or equal
7:07:22to
7:07:2325,000 similarly let's execute another
7:07:25one which detects the employees with
7:07:28salaries less than 25,000 so we have got
7:07:31two employees which are having lower
7:07:33salaries which are Amit and chaitanya
7:07:35fine so this is how you perform some
7:07:38operations in height so now let's move
7:07:41ahead and understand the functions which
7:07:43you can perform on height so in the same
7:07:46way let's create a new database again
7:07:49and let's use this particular database
7:07:51that has five
7:07:54functions now let's create a table in
7:07:57this particular database so the table is
7:08:01employee function and it's
7:08:03created now let's try to load in the
7:08:06data yeah the data has been successfully
7:08:09loaded and now let's see if the data is
7:08:12correctly loaded or not yeah the data is
7:08:14loaded correctly now let's try to apply
7:08:18some functions in this particular data
7:08:20so the first first thing or the first
7:08:22function I'm going to apply would be a
7:08:24square root function where I'll be
7:08:26finding out square root of the salaries
7:08:28of the employees so there you go the
7:08:30square root of 25,000 was 58 do decimal
7:08:35numbers so this is how you perform some
7:08:38basic functions on your data now let's
7:08:40try to find out the maximum salary so
7:08:45yeah the job is getting executed you can
7:08:47see some map reduce chops here I think
7:08:50the biggest salary would be from sanun
7:08:54so the maximum salary is
7:08:5640,000 so this is how it
7:08:59works since we are working on cloud eror
7:09:02and the system configuration is limited
7:09:06the execution speed is a bit low but if
7:09:08you're working in real time then this
7:09:11process would take like few seconds and
7:09:13it's
7:09:18done there you go you have the value
7:09:2140,000 as shown here so 40,000 the
7:09:25employee name is sanana is the maximum
7:09:28salary so that's what we got here now
7:09:31let's try to find out the minimum
7:09:36salary so the minimum salary is
7:09:3915,000 and who would that be yeah it's
7:09:43chaitanya with minimum salary
7:09:4615,000 so that's how you do some
7:09:49operations in five let's execute some
7:09:52more operations such as converting the
7:09:54names of the employees to uppercase so
7:09:56you can see the employee names are
7:09:58converted to uppercase here and
7:10:00similarly let's try to convert to lower
7:10:04case so here you can see we have
7:10:06converted them to lower case so this is
7:10:08how you learn technology you need to
7:10:10play with the technology then you'll
7:10:12come to know the advantages and
7:10:14disadvantages so you can learn the
7:10:16possible ways where you can make things
7:10:18work out this is how you do it now now
7:10:21let's move ahead and understand Group by
7:10:23function in five so for that we'll be
7:10:25creating a separate database that is
7:10:28group now we will use this particular
7:10:30database that is group so we'll type in
7:10:33command use group semicolon now we will
7:10:37create a table so the table has been
7:10:40successfully created now we will load
7:10:42data into this particular table now we
7:10:44will use the new CSV file which will be
7:10:47employ 2. CSV now we are using this
7:10:51particular table because we have an
7:10:53additional column in this particular
7:10:55table which is the country column now as
7:10:58discussed before we will be grouping the
7:11:01employees based on Country let's see our
7:11:04data first so we have countries such as
7:11:07USA India UAE so these are the three
7:11:10countries that we are having in our CSV
7:11:12file so we will be categorizing the
7:11:14employees based on their countries so
7:11:17this is the particular command that we
7:11:18will be
7:11:20using
7:11:22so maybe I made an error while creating
7:11:25the table I think I gave a wrong table
7:11:29name here so let's drop our table so by
7:11:33mistake I gave different table that is
7:11:35employee order so to drop a table you
7:11:38just need to use the keyword drop and
7:11:41it's
7:11:41done yeah the keyword table was missing
7:11:45so you need to type in drop table and
7:11:47the table name and the table gets
7:11:49dropped so we were supposed to create a
7:11:51different table that is employee group
7:11:55so now let's create a new table that is
7:11:57employee Group Employee group has been
7:11:59created now let's try to add in data
7:12:02into the employee
7:12:04group so we have used employee 2 here
7:12:07because the employee 2 has another
7:12:09column which is based on Country so the
7:12:11countries that we are having here are
7:12:13India USA and UA so we will be using the
7:12:17group by function here and we will
7:12:19categorize the employees BAS based on
7:12:21their
7:12:23countries so there you go you can see
7:12:25some map reduce jobs getting
7:12:32executed yeah there you go we have
7:12:35categorized the employees based on their
7:12:39countries that as India UAE and
7:12:41USA and the sum of the salary so the
7:12:45guys is working in India and their
7:12:46sumission of the salary is 90,000 and
7:12:49similarly UA is nearly 1 lakh 5,000 and
7:12:53USA is 80,000 now let's also execute a
7:12:57different command based on Group by so
7:12:59here we'll be using Group by function
7:13:02and we will categorize based on the
7:13:04country as well as the summation of the
7:13:07salary which is greater than or equals
7:13:09to 15,000 so it's similar to the
7:13:11previous
7:13:18command so you can see the data got
7:13:20executed and we got the same output now
7:13:23let's move ahead and understand order by
7:13:26and sort by methods so for that we'll
7:13:28create a new database orders now we'll
7:13:31use
7:13:34orders now let's create a new table
7:13:36again so the new table is employee order
7:13:39and the table got created now let's load
7:13:42the data into this particular table by
7:13:44now I think you have some good practice
7:13:46of how to create a database how to
7:13:48create a table and how to load data into
7:13:50that particular table so the data got
7:13:53loaded and now we going to order the
7:13:56data present in this particular table
7:13:58based on the descending order of their
7:13:59salary so you're seeing some map reduce
7:14:02jobs going ahead so here we'll see the
7:14:05employees ordered based on their
7:14:07salaries in descending order so the
7:14:10highest salary will be at the first
7:14:11place and the lowest salary will be at
7:14:13the last
7:14:15place yeah so we have sanjana at the
7:14:18first position with 40,000 as the
7:14:21highest salary and she's working for UAE
7:14:24and we have cha with lowest salary
7:14:2715,000 working for India now let us also
7:14:32execute another command based on uh sort
7:14:35by so first we try to execute a command
7:14:38based on orderby now let's see the same
7:14:40output using sort by so basically both
7:14:43work in the same
7:14:48way so there you go we have sorted the
7:14:51records based on descending order of
7:14:53salary now that we have learned what are
7:14:56the various operations that can be
7:14:58performed in h that are the arithmatic
7:15:00operations logical operations and also
7:15:03some of the functions such as maximum
7:15:06minimum Group by order by sort by so
7:15:10these are the various operations and
7:15:12functions that you can perform And Hive
7:15:14now let's move ahead into the last type
7:15:16of operations that can be performed in
7:15:18Hive those are the joints so for that
7:15:22let's again create a new database so
7:15:25here I'll be creating a new database
7:15:26that is urea join and followed by that
7:15:30let's use this particular database now
7:15:32for that we need to use the keyword use
7:15:34and there you go we are in edura join
7:15:38now let's create a new table for
7:15:40that so the table will be EMP join here
7:15:44you can see that I forgot to mention
7:15:46semicolon so now the table got created
7:15:49now we shall load the dat data into this
7:15:51particular table so now I've created the
7:15:54first table that is employee table and
7:15:56I'm loading the employee data into this
7:15:58particular table now to perform join
7:16:00operations we always need two tables so
7:16:04in this particular database at urea join
7:16:06I've already created the first table
7:16:08that is employe join now let's create
7:16:11second table that is the department
7:16:14table which will be present in the same
7:16:16database so this particular table is a
7:16:20department table which will be having
7:16:21the entities that are Department ID and
7:16:24Department name now let's load the data
7:16:27of Department into this particular table
7:16:31so the data has been loaded so you can
7:16:33see the employee 2. CSV had the columns
7:16:36ID name salary age and Country and
7:16:38similarly the department. CSV has the
7:16:41entities which are Department ID and
7:16:44Department name so the department IDs
7:16:46are present here and the names are
7:16:48development testing product relationship
7:16:50ship and admin and ID support now we
7:16:53have created both the tables and we have
7:16:56created or we have loaded the data also
7:16:59now we have four different joints
7:17:02available in Hive they are in a joint
7:17:04left outter joint right outo joint and
7:17:08full outer joint now let's perform the
7:17:10first type of joint which is the inner
7:17:12joint so in inner joint we are going to
7:17:14select the employee name and employee
7:17:17department and based on the employee ID
7:17:19and Department ID we are going to
7:17:21perform the joint operation that is the
7:17:23first joint in the
7:17:26joint so you can see some jobs getting
7:17:29executed so the map reduce task
7:17:32successfully
7:17:36completed so um the first set of join
7:17:39has been successfully finished and the
7:17:40output has been generated now let's try
7:17:43out the second type of join that is the
7:17:45left outter
7:17:47join so the only difference is that
7:17:49we're using the keyword left outer join
7:17:52now you can see one of the job got
7:18:01started so you can see the output is
7:18:04beenin generated as well of the left out
7:18:06of joint now let's move ahead and
7:18:09understand WR out of joint so for WR out
7:18:12of joint you need to use the keyword WR
7:18:14out to join fire in the command and you
7:18:17can see um the jobs getting
7:18:19executed
7:18:25so you can see the output of right out
7:18:27join has been successfully executed or
7:18:30displayed now let's type in the last uh
7:18:33join operation that is full out join so
7:18:37here I'm using the keyword full outter
7:18:39join fire in the command and you can see
7:18:41it's getting
7:18:45executed so uh the output for full out
7:18:49join has been displayed here so this is
7:18:51how the join operations are executed in
7:18:54Hive so we have learned how to create
7:18:57database how to create table how to load
7:19:00data and the various data models present
7:19:03in Hive that are the tables databases
7:19:06partitions bucketing and after that we
7:19:09have also understood various operations
7:19:11that are the arithmatic operations
7:19:13logical operations and functions that
7:19:15can be performed in Hive such as square
7:19:17root and summation minimum Max maximum
7:19:21and after that other operations such as
7:19:24group bu sort by order by and also the
7:19:27joints that are possible in Hive which
7:19:30are inner joint left outer right outer
7:19:33and full outer so each and every
7:19:36operation that could be possibly
7:19:38executed in hi have been displayed in
7:19:40this particular tutorial and everything
7:19:42is sorted here in the base of databases
7:19:45and you can get all the details about
7:19:48this and you'll also get the code that I
7:19:50have used in the description box below
7:19:53and you can try it out and also if
7:19:55you're looking for an online
7:19:56certification and training based on Big
7:19:58Data Hadoop then you can check out the
7:20:00link in the description box below and
7:20:02during the training you'll get to have
7:20:04realtime hands-on experience with
7:20:06realtime data you'll learn a lot of
7:20:08things in the training and so far so
7:20:11good now we shall also discuss some of
7:20:13the limitations of Hive so Apache Hive
7:20:16limitations so Hive is not capable of
7:20:19handling real time data hi is capable of
7:20:22batch processing if you have to work
7:20:24with realtime data then you have to go
7:20:26with realtime tools such as spk and
7:20:29gafka so it's like how will actually
7:20:32take in the data for example imagine
7:20:35you're working on Twitter and you have
7:20:38one lakh commments on a particular post
7:20:41so if you had to process those one lakh
7:20:42comments you'll have to first load all
7:20:45those commments into Hive then you need
7:20:48to process it so while you're loading
7:20:50the data from Twitter to Hive you may
7:20:53also get a few more comments that you
7:20:55will be missed out so it's not
7:20:58preferable for Real Time Hive is
7:21:00preferable for only batch Moree
7:21:02processing so followed by that it is not
7:21:04designed for online transaction
7:21:06processing so online transaction
7:21:09processing is something which only works
7:21:11in real time so Hive cannot support real
7:21:14time processing so last but not the
7:21:16least High queries contain High latency
7:21:19yeah High queries take a longer time to
7:21:22process as you've seen I've have taken a
7:21:24smaller CSV file and the time consumed
7:21:26to process such a small CSV file was
7:21:29taking so long so yeah High queries
7:21:32contain High latency so these are the
7:21:34few important noticeable limitations of
7:21:39[Music]
7:21:42hi so this project is based on the
7:21:45e-commerce domain so let me give you an
7:21:48introduction to this project in context
7:21:51of one of the biggest names of
7:21:52e-commerce platforms none other than
7:21:55Amazon so if you have ever shopped from
7:21:58Amazon before which I presume you must
7:22:00have you must have seen something like
7:22:03this when you click on a product so
7:22:05you'll view the details of the product
7:22:08something like this and as you know most
7:22:11of the e-commerce organizations do not
7:22:13have any inventory so they tie up with
7:22:16different Merchants similar is the case
7:22:18with Amazon and Amazon provides the
7:22:21merchants or the sellers a platform to
7:22:24get connected to the buyers so when you
7:22:27click on the details of a product you
7:22:30can see that Amazon gives you something
7:22:33like this and you also find something
7:22:36like this there are 28 offers from this
7:22:39price and if you click on it you can see
7:22:41the name of the different sellers who
7:22:43are selling the same product at
7:22:46different prices and the prices they're
7:22:48offering are listed like this but you
7:22:51can see over here that by default Amazon
7:22:54has selected the appario retail private
7:22:57limited for this particular product so
7:23:00how does Amazon do that so it is
7:23:02actually based on a merchant rating
7:23:05system and as a platform as a e-commerce
7:23:08platform you have to ensure that you
7:23:11always display the product from the best
7:23:14Merchant in order to ensure quality
7:23:17because you don't want angry customers
7:23:19right right so it is very important that
7:23:22your customers are satisfied with their
7:23:24product so you have to choose from
7:23:26different Merchants for the same product
7:23:29in order to decide which Merchants
7:23:31product needs to get displayed by
7:23:34default and hence Amazon has a merchant
7:23:37rating system in order to decide that
7:23:39and this is what exactly we're going to
7:23:42build all right so we're going to make a
7:23:45merchant rating system similar to this
7:23:49so here here's the problem statement so
7:23:51there are multiple merchants selling the
7:23:53same type of products as you can see
7:23:56that Merchant 1 and Merchant six are
7:23:58selling the same shirt Merchant three
7:24:01and four are selling the same shoes and
7:24:04similarly five and one are selling the
7:24:06same pants and there are multiple other
7:24:09Merchants who are selling the same kind
7:24:12of products and you have to build a
7:24:15merchant rating system or the company
7:24:18wants to build a Merchant rating system
7:24:21in order to determine which Merchant
7:24:23sells the best product so that their
7:24:26product would be displayed by default
7:24:29and as a big data expert let's just
7:24:32assume that you are hired by the company
7:24:33as a big data expert you are assigned
7:24:36this task so this is now your problem to
7:24:39solve so the first thing your
7:24:41organization will give you before you
7:24:44start to do your work is the data set so
7:24:47this is the data set that you're going
7:24:49to get so this is is the transaction
7:24:51data set and has certain Fields like
7:24:53transaction ID customer ID merchant ID
7:24:57amp when the purchase was made the
7:25:00invoice number the amount and the
7:25:02segment of product that was bought you
7:25:05have another data set which is the
7:25:07merchant data so these are the details
7:25:09about the different sellers or the
7:25:11merchants so you have got your merchant
7:25:14ID their tax registration number the
7:25:16merchant name their mobile number start
7:25:19date email address State country pin
7:25:22code description longitude latitude the
7:25:25location basically so these are all the
7:25:28details about the merchant that you have
7:25:30in your data set so let me just show you
7:25:32the data set so this is the data set
7:25:35that is in your htfs right now so here
7:25:39is the transaction
7:25:41data so here is the transaction data
7:25:43which is a 2GB of file and the merchant
7:25:47data set which is 20 MB because as you
7:25:49know there are many transactions but a
7:25:52limited number of sellers or Merchants
7:25:55so that is why the size of the data set
7:25:57the merchant data set is quite smaller
7:25:59as compared to the transaction one and
7:26:02to tell you we haven't actually used the
7:26:04entire data present in the data set we
7:26:07have just selected a subset or a sub
7:26:10data set you can say because the
7:26:11original data set was quite huge and
7:26:14this was a demo project so just for your
7:26:16understanding we have chosen a sub data
7:26:19set we just took 2 GBS of data out of it
7:26:22all right and this is how it exactly
7:26:24looks like so this is the CSV file of
7:26:26the data set that we have this is the
7:26:28transactions data all right so this is
7:26:31the approach to solve so the first thing
7:26:33we'll do is that we'll segregate the
7:26:36merchants based on the price of their
7:26:38products and their sales so we will be
7:26:42segregating them into four categories so
7:26:46the categories are the merchants who are
7:26:48selling products that are below 5,000 or
7:26:52less than
7:26:53$5,000 there is one more category for
7:26:56merchants who sells products between
7:26:58$5,000 to
7:27:00$10,000 another category of merchants
7:27:03who sells products between $10,000 to
7:27:0620,000 and another category greater than
7:27:1020,000 or more than 20,000 we'll be
7:27:13using a simple logic to approach solving
7:27:15this problem so let's say that if there
7:27:18is a merchant who selling their products
7:27:21at quite a low price and you see if he's
7:27:24not making a good number of sales it
7:27:27means that he is not selling quality
7:27:30products because if a merchant who's
7:27:32selling the product at quite a less
7:27:34price and people aren't still buying
7:27:36from him it clearly indicates that his
7:27:39products are not up to the mark so it's
7:27:42a low rating for a merchant similarly on
7:27:45the other hand if you see a merchant who
7:27:47is selling their product at quite a high
7:27:50price and yet he has a very good number
7:27:52of sales so it clearly indicates that
7:27:55his products must be very good because
7:27:57despite of the higher price people are
7:28:00still buying from him so in that case
7:28:03obviously the rating of the merchant
7:28:05would be also very good right so this is
7:28:08a simple logic that we're going to use
7:28:10in order to rate our merchants or
7:28:12Sellers and you have three options to
7:28:15choose from so you have got aachi Hive
7:28:18which is a great analytical tool we have
7:28:21got Hadoop map produce and Apache Pig
7:28:25and today we will be choosing Hadoop map
7:28:27produce so map produce is the core
7:28:30component of Hadoop that process huge
7:28:32amount of data in parallel by dividing
7:28:35the work into a set of independent tasks
7:28:39map produce is the data processing layer
7:28:41of Hado it is a software framework for
7:28:44easily writing applications that
7:28:46processes the vast amount of structured
7:28:48and un structured data that is stored in
7:28:51your hdfs hdfs is Hadoop distributed
7:28:55file system in Hadoop map reduce works
7:28:58by breaking the data processing into two
7:29:01phases maap phase and the reduce phase
7:29:03and that is how exactly it gets us name
7:29:06map reduce so the map is the first phase
7:29:09of processing where we specify all the
7:29:12complex logic business rules reduce is
7:29:15the second phase of processing where we
7:29:17specify lightwe processing like
7:29:19aggregation or summing up the
7:29:23outputs but the question is why choose
7:29:26map ruce well I'll give you two reasons
7:29:28for it first is the custom input format
7:29:33now input format is something that
7:29:35defines how your input files are going
7:29:38to be split and read so in map ruce you
7:29:41can create your own custom input format
7:29:44instead of using the default input
7:29:46format this actually makes handling of
7:29:48your data quite easier because here we
7:29:51can create our custom input format for
7:29:53transactions and pass it as an
7:29:56argument then we have the distributed
7:29:59cachy so distributed cachy is nothing
7:30:02but it is a facility that is provided by
7:30:05map ruce framework to Cache files files
7:30:08like your text files archives jars Etc
7:30:12that is needed by your application let's
7:30:15understand this with an
7:30:17analogy just think about it that that
7:30:19there are three students sitting on a
7:30:22table solving chemistry problems and
7:30:25they have one periodic table so you keep
7:30:28the periodic table in the middle of the
7:30:29table so the students all the three
7:30:31students can refer from the periodic
7:30:33table to find to see the atomic numbers
7:30:36of different elements and solve their
7:30:38problems right so one periodic table
7:30:41everyone can refer to it and solve their
7:30:43own problem so this is what distributed
7:30:46cachier is so with distributed cachier
7:30:49you can put the data that will be used
7:30:51by your different data noes to refer in
7:30:54order to run map produce jobs so we'll
7:30:57be learning more about how to create
7:30:59your custom input format and how to use
7:31:01the distributor caching in the demo part
7:31:03all right so before that let us just
7:31:06understand how map produce exactly works
7:31:08so this is a sample map Produce job
7:31:11execution with an example so this is
7:31:14your input file this is a text input
7:31:17file so it contains some some words so
7:31:21first what will happen is that the input
7:31:23will get divided into three splits all
7:31:27right so I'm taking one sentence at a
7:31:29time so I have got three splits over
7:31:32here and this will distribute the work
7:31:35among all the different map nodes then
7:31:38we will tokenize the words in each of
7:31:40the mapper and give value one to each of
7:31:44the tokens or words so deer one bear One
7:31:48River one similarly here car 1 car One
7:31:51River one and now a list of key value
7:31:55pair will be created where the key is
7:31:58nothing but the individual word and the
7:32:01value is one so this is the key and this
7:32:03is the value and this will happen on all
7:32:06the three nodes so the mapping process
7:32:10remains same on all the nodes so after
7:32:13the mapper phase a partition process
7:32:15takes place where the sorting and
7:32:18shuffling happen happens here all the
7:32:20topples with the same key are sent to
7:32:23the corresponding reducer so all the be
7:32:26are together cars are together deer and
7:32:29River are together so after the sorting
7:32:32and shuffling phase each reducer will
7:32:34have a unique key and a list of
7:32:37corresponding values to that key for
7:32:40example bear 1 one car one one and one
7:32:43and so on now comes the reducing phase
7:32:47now each reducer counts the value which
7:32:49are present in the list of values so
7:32:52reducer gets a list of values which is
7:32:54one one for the key bear and then it
7:32:57counts the number of ones in the list
7:33:00and gives the final output as bare two
7:33:03similarly for car it's three deer Two
7:33:07River it will count 2 1 so two and
7:33:11finally all the output the key value
7:33:14pairs are then collected and written in
7:33:16the output files so this is your output
7:33:18file so it has just combined the result
7:33:21from different reducers and here is your
7:33:25final output so understood map produced
7:33:28with the classic example of the word
7:33:30count program now this is the generic
7:33:33execution flow of the map produ job so
7:33:37you have your input file over here so
7:33:40the data for map produce task is stored
7:33:42in input files and input files typically
7:33:44lives in the
7:33:46hdfs the format of these files is arbit
7:33:49while line based log files and binary
7:33:51format can also be used and you have a
7:33:54input format now input format defines
7:33:57how this input files are split and read
7:34:00it selects the files or other objects
7:34:02that are used for input an input format
7:34:04creates the input split so it logically
7:34:07represents the data which will be
7:34:09processed by an individual mapper one
7:34:12map task is created for each split and
7:34:15thus the number of map task will be
7:34:17equal to the number of in input splits
7:34:20the splits are then divided into records
7:34:22and each record will be processed by the
7:34:25mapper now let's talk about the mapper
7:34:28so mapper processes each input record
7:34:30from the record reader and generates a
7:34:33key value pair and this key value pair
7:34:36is generated by the mapper is completely
7:34:39different from the input pair the output
7:34:42of the mapper is also known as the
7:34:44intermediate output which is written to
7:34:46the local disk the output of the mapper
7:34:49is not stored on htfs as this is
7:34:51temporary data and writing on htfs will
7:34:54create unnecessary copies so then the
7:34:57mapper output is passed on to the
7:34:59combiner for further process the
7:35:02combiner is also known as the mini
7:35:04reducer so Hadoop map produce combiner
7:35:07performs local aggregation on mappers
7:35:09output which helps to minimize the data
7:35:12transfer between mapper and the reducer
7:35:15once the combiner functionality is
7:35:18executed the output is then passed to
7:35:21the partitioner for further work now
7:35:24partitioner comes on the picture if
7:35:25you're working on more than one reducer
7:35:28and here we have two reducers in the
7:35:30example so if you have one reducer you
7:35:33don't actually need a
7:35:35partitioner so the partitioner takes the
7:35:37output from the combiners and performs
7:35:41partitioning partitioning of output
7:35:43takes place on the basis of the key and
7:35:46then sorted so by hash function a key is
7:35:49used to derive the partition according
7:35:52to the key value and map produce each
7:35:54combiner output is partitioned and a
7:35:57record having the same key value goes to
7:35:59the same partition and then each
7:36:01partition is sent to a reducer so by
7:36:05using partitioner it allows to have an
7:36:07even distribution of the map output over
7:36:10the reducer so after that comes the
7:36:12shuffling and sorting part so now the
7:36:15output is shuffled to the reduce node
7:36:18the shuffle Shing is the physical
7:36:19movement of the data which is done over
7:36:22the network once all the mappers are
7:36:24finished and their output is shuffled on
7:36:27the reducer nodes then this intermediate
7:36:29output is merged and sorted which is
7:36:32then provided as an input to the reduced
7:36:34face now comes the reducer so it takes
7:36:37the set of intermediate key value pairs
7:36:39produced by all the mappers as the input
7:36:43and then runs a reducer function on each
7:36:45of them to generate the output the
7:36:48output the reducer is the final output
7:36:50which is stored in hdfs so if you have
7:36:53multiple reducers the result from
7:36:56different reducers will combine and that
7:36:58is going to be your final output which
7:37:00will be written into the
7:37:02hdfs so this was the map Produce job
7:37:06execution flow so we'll also be using
7:37:08the distributed cache so you'll have
7:37:11different data notes each data node will
7:37:13have their local copy of their data and
7:37:16if each of the data nodes needs to refer
7:37:18something something we will keep that in
7:37:20the distributed cachier and in this case
7:37:22we'll be keeping the merchants file all
7:37:24right so distributed cachier is nothing
7:37:26but think of it as a share drive right
7:37:29so if you have multiple users who wants
7:37:32to have access to one data set so you
7:37:35can just put it up in the share drive
7:37:36and all of your users can share the data
7:37:40use the same data right so this is what
7:37:42distributed cache is so applications
7:37:45specify the files via URLs to cach a
7:37:49via the job con so I'll be telling you
7:37:51about the job conf later in the demo
7:37:53section and the distributed cacher
7:37:55assumes that the file specified via URLs
7:37:58are already present on the file system
7:38:00at the path specified by the URL and are
7:38:04accessible by every machine in the
7:38:05cluster so the framework will copy
7:38:07necessary files to the slave node before
7:38:10any jobs are executed on that node and
7:38:13distributed cache tracks modification
7:38:16timestamps of the cache files so click
7:38:18clearly the cach files should not be
7:38:20modified by the applications or
7:38:22externally while the job is
7:38:25executing so how will it works in our
7:38:28case we'll be storing the data into htfs
7:38:31and we'll be executing map reduce
7:38:34program over that file so we'll store
7:38:36the merchant data in the distributed
7:38:39cache then we'll segregate the
7:38:41transaction data into categories such as
7:38:44less than 5,000 5,000 to $10,000 10,000
7:38:4820,000 and greater than $20,000 with the
7:38:52merchant ID then we'll use the merchant
7:38:55file from the distributed cache and map
7:38:58the merchant ID with the merchant name
7:39:01and at last we'll receive the output as
7:39:03the merchant name with date indicating
7:39:06the number of sales in different
7:39:09categories now let us talk about the
7:39:11code sections so the execution will
7:39:14start from the main method where we'll
7:39:16use the tool Runner so the tool tool
7:39:18Runner can be used to run classes
7:39:20implementing tool interface so it paus
7:39:23the generic Hadoop command line
7:39:25arguments and modifies the configuration
7:39:27of the tool then it will point to the
7:39:30run method which will point to the Run
7:39:33Mr jobs so here we are specifying the
7:39:36driver code so I'll be telling you about
7:39:37the driver code later on so next the
7:39:40execution will move to the mapper class
7:39:42which is the transaction mapper the
7:39:44framework first calls the setup method
7:39:48followed by the map method for each key
7:39:51value pair in the input split so in
7:39:54setup method we are loading the cache
7:39:56file and calling a method where we'll be
7:39:59resolving the merchant name from the
7:40:01merchant ID next in the map method we're
7:40:04creating the object of transaction which
7:40:06we will be using to catch all the fields
7:40:08of transaction first and then using the
7:40:11object of aggregate data we will create
7:40:14the segment of transactions as we
7:40:16discussed before the four segments less
7:40:19than 5,000 10,000 20,000 those segments
7:40:22at last using the merchant ID name map
7:40:24we will resolve or find out the merchant
7:40:27name the output of the map method would
7:40:31be the key which will be the combination
7:40:33of merchant name and date of sale while
7:40:35the value would be in the form of number
7:40:37of sales of different categories next
7:40:40the execution will go to the partitioner
7:40:43code where we'll have the get partition
7:40:46method which will send the records with
7:40:48the same key to the same reducer and at
7:40:51last the reducer code will execute which
7:40:53will aggregate the data with the same
7:40:54key and provide the output so I'll be
7:40:58showing you and explaining you all the
7:40:59codes involved over here in the demo
7:41:02part all right now let us move ahead so
7:41:05first let me take you through this
7:41:07transaction class where we are defining
7:41:10all the variables corresponding to the
7:41:12transaction file as you can see we have
7:41:15got the transaction ID customer ID
7:41:17merchant ID Tim stamp invoice number
7:41:20invoice amount and segment here so these
7:41:23are the fields in the transactions and
7:41:25we have created the variable for the
7:41:27same so next we're creating the getter
7:41:30and Setter methods for each of the
7:41:31variables so as to read the value of the
7:41:34field and we write the value of the
7:41:36field so as you can see here we are
7:41:39defining the method get so get segment
7:41:42we have got get segment here where we
7:41:44are returning the value of the field and
7:41:47the set meth method here the set segment
7:41:50where we're writing the value of this
7:41:52field so similarly we're doing it for
7:41:54all the variables as you can see here so
7:41:58we have the get and set for customer ID
7:42:02so we have the get customer ID method
7:42:04which Returns the value of the field and
7:42:06we have got the set customer ID which
7:42:09writes the value of the field so this
7:42:12similar for all the fields in our
7:42:15transaction data so similarly you can
7:42:17have the aggregate data class which we
7:42:20have used to create the categories of
7:42:22the product so here we have Fields like
7:42:25order below 5,000 order below 10,000
7:42:28order below 20,000 order above 20,000
7:42:31which are nothing but the different
7:42:33categories which we have defined earlier
7:42:35and similar to the transaction class
7:42:37here we are defining the getter and
7:42:39Setter method so we have got get total
7:42:41order method which Returns the value of
7:42:44total order and we have set total order
7:42:46method which is writing the value of
7:42:49this field total order so we have the
7:42:51same for all the different variables
7:42:54that we have defined in the aggregate
7:42:56data class the getter and Setter methods
7:42:59then we've got the aggregate writable
7:43:01class so first here we are creating an
7:43:04object of gson class so gson is
7:43:06basically used to convert Java objects
7:43:09to Json format next we're initializing
7:43:12the aggregate data object now we have
7:43:15two Constructors first one is the basic
7:43:17Constructor which is not taking any
7:43:19argument the second Constructor is
7:43:22taking aggregate data format as an
7:43:24argument and trying to initialize the
7:43:26aggregate data object next we have the
7:43:29getter method for the aggregate data
7:43:31which will return the aggregate data
7:43:33object after that we have the write
7:43:36method which will write the value of
7:43:38corresponding Fields using the getter
7:43:40method of the field like for order get
7:43:43order below 5,000 get order below 10,000
7:43:46order below 20,000 get order above
7:43:4920,000 then you have read Fields method
7:43:52which will'll be calling the seter
7:43:53method of each field to assign the
7:43:55values to the corresponding field and at
7:43:58last we are overwriting to string method
7:44:00which will convert the aggregate data
7:44:02object to Json and then return the Json
7:44:06so I hope you guys are clear with the
7:44:07custom input format so now let us take a
7:44:10look at the main Java file which is the
7:44:13merchant analytics job. Java so the main
7:44:17class is the mer Merchant analytics job
7:44:19class inside which all the jobs will
7:44:22reside so the execution will start from
7:44:25the main method so first let us go to
7:44:27the main
7:44:29method so here we are using toolrunner
7:44:32so toolrunner can be used to run classes
7:44:35implementing the tool interface it
7:44:37passes the generic Hadoop command line
7:44:40arguments and modifies the configuration
7:44:42of the tool so tool runner. run method
7:44:45runs the given tool after parsing the
7:44:47given generate arguments it uses the
7:44:50given configuration or builds one if
7:44:52null it sets the tools configuration
7:44:55with the possibly modified version of
7:44:58the conf here we are passing the
7:45:00configuration object Merchant analytics
7:45:03job object which is the main class and
7:45:06arguments which we will be providing
7:45:08while executing the job so in our case
7:45:11there are three arguments first is the
7:45:15path of the transaction file second is
7:45:17the path of the the merchant file and
7:45:19third is the output directory now we
7:45:22will execute the run method where we are
7:45:25returning the values of the Run Mr jobs
7:45:28method we're also parsing the arguments
7:45:31that is all the three parts that is the
7:45:33transaction Merchant and output
7:45:36directory to the Run Mr jobs method now
7:45:39let us see the Run Mr jobs
7:45:42method so here we have the driver code
7:45:46so first we initialize the conf
7:45:48configuration object and then we will
7:45:49initialize the control job object so
7:45:53Control job Class encapsulates A map
7:45:55Produce job and its dependency it
7:45:58monitors the state of the depending jobs
7:46:00and updates the state of this job and
7:46:02now we are creating the object of job
7:46:04class and we will define the properties
7:46:07of the job so first we have the set
7:46:09output key Class Property where we are
7:46:11defining the output format class of the
7:46:13key which is text class similarly we're
7:46:16defining the set out output value class
7:46:18for output format class of the value
7:46:21that is the aggregate writable class
7:46:24next we have set jar by class which
7:46:26tells the class where all the mapper and
7:46:28reducer code resides which the merchant
7:46:31analytics job in our case now we are
7:46:34specifying the reducer class which is
7:46:36Merchant order reducer and next we are
7:46:39providing the input directory so the set
7:46:41input the recursive method will read all
7:46:44the files from the directories
7:46:45recursively if we are providing the
7:46:47directory path so first we're adding the
7:46:50input path of the transaction file which
7:46:52is present in the argument zero then
7:46:54here we are also specifying the input
7:46:56format of the file and the mapper class
7:46:59that is the transaction mapper next
7:47:01we're talking about all the merchant
7:47:03file from the directory provided in
7:47:05argument one and adding this file to the
7:47:07distributed cache using the job. cache
7:47:11archive method and moving ahead we are
7:47:14setting the output directory path which
7:47:16is provided in argument two and we are
7:47:18also adding the timestamp as the
7:47:21subdirectory and at last we are setting
7:47:23the partitioner class that is the
7:47:25Marchant partitioner and then we are
7:47:28returning zero or one depending on
7:47:30whether the job has been executed
7:47:32successfully or not and next we will
7:47:34take a look at the transaction mapper
7:47:36class which implements the mapper
7:47:39interface so it Maps input key value
7:47:42pairs to a set of intermediate key value
7:47:44pairs so maps are the individual task
7:47:47which trans form input records into an
7:47:49intermediate record the transformed
7:47:52intermediate records need not to be of
7:47:54the same type as the input records the
7:47:56Hadoop map ruce framework spawns one map
7:47:59task for each input split generated by
7:48:02the input format for the job and mapper
7:48:05implementations can access the job conf
7:48:08for the job via the job configurable and
7:48:11initialize themselves the framework
7:48:14first calls the setup method followed by
7:48:16map method for each key value pair in
7:48:19the input split so in the setup method
7:48:22we are loading the merchant file from
7:48:24the cache using the get cache archives
7:48:27method then from each file we are
7:48:29calling the load merchant ID name in
7:48:32Cache so we're calling this method and
7:48:35we are passing the path of the cache
7:48:37files and the configuration objects now
7:48:40in this load merchant ID name in cach a
7:48:42method we are initializing the object of
7:48:45file system using the con and next we're
7:48:48opening the file and then we're reading
7:48:50the data line by line from the file now
7:48:53here first we're removing the codes from
7:48:56the line and then we're splitting the
7:48:58line using the comma and at last we're
7:49:01putting the merchant ID and Merchant
7:49:03name in the merchant ID name map
7:49:06variable so here you can see in the
7:49:08merchant file that we have merchant ID
7:49:11at index zero and Merchant name at index
7:49:142 so this merchant ID name map was will
7:49:17help us in resolving the merchant name
7:49:20from merchant ID and next we're defining
7:49:23exception to notify us if the cacher
7:49:26file is not read and at last we're
7:49:29closing the object of the buffered
7:49:32reader now let's talk about the map
7:49:35function so now the map function will be
7:49:38called so the input format of key is
7:49:41long writable and the value is text
7:49:44we're also creating a context of the
7:49:46mapper frame work where we will be
7:49:48writing our intermediate output now the
7:49:51output format of the key is text and the
7:49:54value is aggregate writable again here
7:49:57we are removing the codes from the line
7:49:59and then we are splitting the line using
7:50:01comma so in the split Arrow we have all
7:50:05the fields of the transaction data
7:50:07stored in the consecutive
7:50:09indexes now we are creating an object of
7:50:12transaction class and setting the values
7:50:14of the field using Setter methods and
7:50:17next we are creating the objects of
7:50:19aggregate data and aggregate writable
7:50:22class then using transactions get
7:50:25invoice amount field we are deciding
7:50:28that in which aggregate data field the
7:50:30transaction would lie we will set the
7:50:33value of corresponding field of the
7:50:34aggregate data object to one and next we
7:50:37have the output key which will contain
7:50:40the merchant name and the date of sale
7:50:42we will set the value of corresponding
7:50:45field of that aggregate data object to
7:50:48one and next we have the output key
7:50:51which will contain the merchant name and
7:50:53the date of the sale so merchant ID name
7:50:56map method will return the merchant name
7:50:59from the merchant ID as we just
7:51:01discussed so we are passing the values
7:51:04as merchant ID and the date at last
7:51:07we'll pass the intermediate result in
7:51:09form of key and value to the
7:51:12context next the result will be sent to
7:51:15the partitioner class that is the
7:51:17merchant
7:51:18partitioner so it's over here so in this
7:51:22class we're overwriting the default get
7:51:24partition method and in this method we
7:51:27are converting the key using hash
7:51:29function and using ABS method to return
7:51:33the absolute value of a number and at
7:51:35last we're using the modular function to
7:51:38get the remainder and now we're dividing
7:51:40it with the number of partitions so
7:51:42which is nothing but the number of
7:51:44reducers and in our case we have
7:51:47specified five reducers so the modulo 5
7:51:50would return the value between zero and
7:51:52four and one more thing is the same key
7:51:56would always have the same hash
7:51:58generated and hence the modular result
7:52:00would be also the same and thus the
7:52:03records with the same key will be sent
7:52:05to the same reducer and based on this
7:52:08records are sent to the reducer so next
7:52:11we have the reducer code and it's over
7:52:15here so it's the merchant order reducer
7:52:19the reducer Clause as defined in the
7:52:21driver code resides in the merchant
7:52:23order reducer Clause so here we have the
7:52:26key input as text and value input as
7:52:29aggregate writable which was written by
7:52:31our mapper class and the output key
7:52:34format is again text and the output
7:52:37value format is aggregate writable so
7:52:40here our execution will move to reduce
7:52:43method where we are passing the input
7:52:45key value and context as argument and
7:52:49here we are again creating the objects
7:52:51of aggregate data and aggregate writable
7:52:54class and next we are taking the input
7:52:57values now here we're calling the setter
7:52:59function of each category getting the
7:53:01earlier value of that category and then
7:53:04adding the new value of the new
7:53:06aggregate data object for that category
7:53:09so it will add the value to the
7:53:11corresponding Fields if there is a
7:53:13record with the same key and at last we
7:53:17writing the key and value in context.
7:53:20write
7:53:21method so I have explained you the code
7:53:24so now let us just go ahead and execute
7:53:33it so first let us move to the project
7:53:40directory so we have the palm. XML file
7:53:44which has all the dependencies that we
7:53:46require in order to run our map Produce
7:53:48job so this is the command so it has my
7:53:52jar file and the path of my jar file
7:53:56then I have got my main class over here
7:53:59which is Merchant analytics job so this
7:54:01is where my main function is and then
7:54:03I'm passing the three parts so first is
7:54:06my transactions. CSV this is my data set
7:54:09this is the path of my data set then my
7:54:11Merchant data this is the path of my
7:54:13Merchant data and finally my output
7:54:16directory which is the result this is
7:54:18the path over here so let us just go
7:54:20ahead and execute this
7:54:33command so the code is run so here are
7:54:37the different parameters on which this
7:54:39map Produce job was run so you can see
7:54:43the details over here so you can see the
7:54:46number of reduce task were five since we
7:54:49had five
7:54:51reducers so you can see all the details
7:54:53here let me just show you the result
7:54:56now so it is in the results directory so
7:55:01there are two directories over here
7:55:03because this was the earlier one that
7:55:05when I had previously executed it so
7:55:07this is the one that we have got right
7:55:10now so let me just show it to you so we
7:55:13have got five part files because there
7:55:15are five reducers so I'm just clicking
7:55:18on one
7:55:19part and you can just click on
7:55:22download all right let me just open
7:55:28it so this is what you get so you get
7:55:31the merchant name and the timestamp over
7:55:34here and also the category or the
7:55:37segregation that we did based on the
7:55:39cost of the orders right so it was order
7:55:42above
7:55:44$20,000 at this time stamp and the total
7:55:46order was one so this is the format of
7:55:49the result so we have got a lot of rows
7:55:53so this is the result so we have just
7:55:56used a few fields or parameters from the
7:55:58merchant file we have just used the
7:55:59merchant name and the ID over here since
7:56:02this is a sample project sample demo
7:56:04project but the scope of this particular
7:56:07project is huge you can use a lot of
7:56:09other parameters that was mentioned
7:56:11there like the location you can analyze
7:56:14it based on locations based on the time
7:56:17time period where the order was placed
7:56:19so you can take in account different
7:56:21fields and improve this or make this
7:56:24analysis even better by
7:56:26yourself so when you're doing this
7:56:28project as a part of your course
7:56:31curriculum so you will be exploring the
7:56:33other fields as well I have just shown
7:56:35you with just using two Fields the
7:56:37merchant name and the
7:56:40[Music]
7:56:44ID what is Kafka in general Kafka is a
7:56:49producer to the consumer based messaging
7:56:51system that has a producer that produces
7:56:53the message and the consumer that
7:56:55consumes the message in between the both
7:56:58we have Brokers that distribute the
7:57:00messages to the consumers and data
7:57:02storage unit which is none other than
7:57:04Apache zuker to understand more about
7:57:07Apache zuker and Kafka you can go
7:57:09through the article Link in the
7:57:11description box below Apache Kafka so
7:57:14basically Apache Kafka is an open source
7:57:17messaging tool developed by LinkedIn to
7:57:19provide low latency and high throughput
7:57:22platform for the real-time data feed it
7:57:25is developed using Scala and Java
7:57:27programming languages so followed by the
7:57:29definition of CFA we shall enter and
7:57:32understand what exactly is a stream in
7:57:35general a stream can be defined as an
7:57:37unbounded and continuous flow of data
7:57:39packets in real time data packets are
7:57:42generated in the form of key value Pairs
7:57:45and these are automatically transferred
7:57:47from the publisher there is no need to
7:57:49place a request for the same the below
7:57:51image depicts the key value pairs that
7:57:53are involved in data stream each and
7:57:56every single key value pair is one
7:57:59single unit of data or it is also called
7:58:01as one single unit of a record so
7:58:04followed by the stream we shall
7:58:06understand what exactly is a CFA stream
7:58:10CFA stream is an API that integrates
7:58:12Kafka cluster to the data processing
7:58:15applications which are either written in
7:58:17Java or Scala this API leverages the
7:58:21data processing capabilities of Kafka
7:58:23and increases data pism Apache Kafka
7:58:27stream can be defined as an open-source
7:58:29client library that is used for building
7:58:31applications and
7:58:33microservices here the input and the
7:58:35output data is stored in the form of
7:58:37Kafka clusters it integrates the
7:58:40intelligibility of Designing and
7:58:42deploying standard applications using
7:58:44the programming languages such as scale
7:58:47and Java with the benefits of Kafka
7:58:49server side cluster technology so this
7:58:53was the basic definition of kafa stream
7:58:56now let us understand Kafka stream API
7:58:58in a much better way through its
7:59:01architecture Apache Kafka streams
7:59:03internally use the producer and consumer
7:59:06libraries it is basically coupled with
7:59:08cfom and the API allows you to leverage
7:59:10the capabilities of Kafka by achieving
7:59:13data pism fall tolerance and many other
7:59:16powerful Fe features the following image
7:59:19depicts the basic architecture of Kafka
7:59:21stream here you can see the Kafka
7:59:24cluster which has the input streams and
7:59:26the output streams together followed by
7:59:29that we have Kafka streaming application
7:59:31or the API which takes care of the
7:59:33queries which are received from numerous
7:59:35applications which are connected to
7:59:37Kafka streaming application followed by
7:59:40this we have numerous components present
7:59:42in kfka stream architecture which are as
7:59:44follows they are input stream output
7:59:47stream instance we have two instances
7:59:51here which are stream instance one and
7:59:53stream instance 2 so inside every
7:59:55instance we have consumers as well as
7:59:58local state and Stream
8:00:01typology So in Kafka stream API input
8:00:04stream and output stream can be one
8:00:06single Kafka cluster followed by that we
8:00:09have a consumer which provides the input
8:00:11and receives the output and inside the
8:00:14stream instance we have stream topology
8:00:17and local state we shall understand
8:00:19about stream topology in a much detailed
8:00:21way in the next slide stream topology is
8:00:23all about the directed aycc graph or the
8:00:26steps in which the particular process is
8:00:28executed followed by that we have a
8:00:30local state local state is none other
8:00:33than a memory location which stores the
8:00:35intermediate data or the result provided
8:00:37by the stream topology these results are
8:00:39produced after applying various
8:00:41Transformations such as map flat map Etc
8:00:45so after the data is processed
8:00:47the tasks are united together and sent
8:00:49back to Output stream so this is how the
8:00:51architecture of cfast stream API works
8:00:54now let us understand more about stream
8:00:57topology so this particular diagram
8:00:59explains the stream topology here you
8:01:02can see the stream processor all the
8:01:04dots which are provided here are none
8:01:06other than stream processors and the
8:01:08line which is connecting them is the
8:01:10stream the stream is none other than the
8:01:13key value pairs of the data or records
8:01:16so basically the input is read from
8:01:18Kafka cluster first followed by that we
8:01:21apply various operators such as filter
8:01:23map join Aggregate and many more and
8:01:26finally we will receive the results
8:01:29which will be sent back to the output
8:01:30Kafka cluster so this is how the stream
8:01:33topology works now let us discuss the
8:01:36important features of Kafka streams that
8:01:38give it an edge over the other similar
8:01:40Technologies so the various important
8:01:43features of Apache Kafka streams API are
8:01:46plastic fa tolerant highly viable
8:01:49Integrated Security Java and Scala
8:01:52language support and exactly once don't
8:01:55worry we shall discuss each one of them
8:01:57in detail firstly we shall discuss about
8:02:01elastic nature Apache Kafka is an open
8:02:04source project that was designed to be
8:02:06highly available and horizontally
8:02:07scalable hence with the support of Kafka
8:02:10kfka streams API has achieved its highly
8:02:13elastic nature and can be easily
8:02:15expandable so this was the first feature
8:02:18followed by that we have the second
8:02:20feature which is about fault tolerance
8:02:23the data logs are initially partitioned
8:02:25and these partitions are shared among
8:02:27all the servers in the cluster that are
8:02:29handling the data and their respective
8:02:31requests thus Kafka achieves fall
8:02:33tolerance by duplicating each partition
8:02:35over a number of servers followed by
8:02:38Fall tolerance we have the next
8:02:40important feature that is highly viable
8:02:43since Kafka clusters are highly
8:02:44available they can be preferred to any
8:02:47sort of use cases regardless of their
8:02:49size they are capable of supporting
8:02:51small scale use cases medium scale use
8:02:54cases also the large scale use cases
8:02:57followed by highly viable feature we
8:03:00have Integrated Security Kafka has three
8:03:04major security components that offer the
8:03:06best-in-class security for the data in
8:03:09its clusters they are mentioned as
8:03:11follows they are encryption of data
8:03:15using SSL or or TLS followed by that
8:03:18authentication of SSL or
8:03:21sasl and finally the authorization of
8:03:25ACLS so followed by the security we have
8:03:28its support for top and programming
8:03:30language the best part of Kafka streams
8:03:33API is that it integrates itself with
8:03:35the most dominant programming languages
8:03:37such as Java and Scala and makes
8:03:41designing and deploying Kafka service
8:03:43side applications with ease followed by
8:03:46that we have exactly once processing
8:03:49semantics usually stream processing is a
8:03:53continuous execution of unbounded series
8:03:55of data or events but in the case of
8:03:59Kafka it is not exactly once means that
8:04:02the user defined statement or logic is
8:04:04executed only once and the updates to
8:04:08the state which are managed by SP or
8:04:10stream processing element are committed
8:04:13only once in a durable backend store so
8:04:16this is how Apache Kafka streaming API
8:04:18is considered to be having exactly once
8:04:21processing semantics so followed by the
8:04:24important features we shall go through a
8:04:26sample program based on Kafka streams
8:04:29API so this particular example can be
8:04:31executed using Java programming language
8:04:34yet there are few prerequisites on this
8:04:36one one needs to have Kafka and zeper
8:04:40installed in the local system and it
8:04:42should be running in the background if
8:04:44you have not installed zookeeper and
8:04:46Kafka in your local system then I have
8:04:48linked the article in the description
8:04:50box below which will explain you about
8:04:52the detail installation procedure of
8:04:53Zookeeper and Kafka in your local system
8:04:57once the Zookeeper and Kafka are
8:04:59installed into your local system you
8:05:00need to fire them up once the Kafka and
8:05:03zookeeper are successfully installed
8:05:05into your local system and they're
8:05:06running in the background you can go to
8:05:08Kafka and Define a producer topic and
8:05:11the consumer once the producer topic and
8:05:13consumer are defined you can come back
8:05:15to Kafka article and execute the
8:05:18following code in any of the Java
8:05:20editors the code will count the number
8:05:22of words that you have provided in your
8:05:24text document and you will receive the
8:05:25output as shown in the article here the
8:05:29text given to the code was welcome to
8:05:31Eda Kafka training and this article is
8:05:34based on Kafka streams these were the
8:05:36two sentences given to the program and
8:05:38the output is as shown below here the
8:05:41word welcome is repeated for once two is
8:05:43repeated for once Eda once once Kafka is
8:05:47repeated for two times and training is
8:05:49once this article is about streams so
8:05:53all these words are repeated for once so
8:05:56this is how exactly you should be
8:05:57receiving the output once after you
8:05:59execute the following code in your Java
8:06:01editor so followed by the example based
8:06:03on Kafka streams we can move ahead and
8:06:06understand the important differences
8:06:07between Kafka and Kafka streams so now
8:06:10the first difference is that in Kafka
8:06:13stream API single Kafka cluster can
8:06:16support as both consumer as well as
8:06:18producer while on the other hand in
8:06:21Kafka we need separate consumer and
8:06:23producer and Kafka considers consumer
8:06:26and producer as separate entities the
8:06:29second difference is that in Kafka API
8:06:32exactly once processing semantics are
8:06:34supported whereas in Kafka it is not by
8:06:38default but you can achieve exactly once
8:06:41processing in Kafka manually the third
8:06:45difference is that Kafka streams API is
8:06:47capable enough to perform complex
8:06:49operations whereas Kafka is designed to
8:06:52perform only simple operations the
8:06:55fourth difference is that Kafka API
8:06:57supports single Kafka cluster on the
8:07:00other hand in Kafka you need two
8:07:02different clusters for producer and
8:07:05consumer followed by that in Kafka API
8:07:09the code length is significantly shorter
8:07:12when you come into CFA the code length
8:07:14involved is highly lengthy
8:07:16the next difference between the both is
8:07:18CFA streams API can support both
8:07:21stateless and stateful networks what are
8:07:24stateless and stateful networks in
8:07:26stateless networks the client provides
8:07:29requests to the server and he gets
8:07:31instantaneous reply from server and here
8:07:34the cookies or the requests which are
8:07:36sent by the client are not stored
8:07:39whereas if you come into stateful
8:07:40Network the client requests the server
8:07:43along with some additional data which is
8:07:45requ required by the server in this case
8:07:48the cookies or the requests which are
8:07:50provided by the client are recorded So
8:07:53cfast stream API is capable to support
8:07:55both stateless and stateful networks but
8:07:58on the other hand Kafka is capable only
8:08:01to support stateless Network protocols
8:08:03followed by that the Kafka streams API
8:08:06can support
8:08:07multitasking whereas Kafka is not
8:08:10capable to support multitasking at a
8:08:13single task level followed by that Kafka
8:08:16stream API does not support batch
8:08:19processing whereas Kafka is capable to
8:08:22support batch processing Kafka stream
8:08:25API is all about real time so it doesn't
8:08:28have to support batch processing so
8:08:30these were the few important differences
8:08:32between Kafka streams API and Kafka now
8:08:36we shall move ahead and wind up the
8:08:38session with our last topic which are
8:08:40the important use cases based on Apache
8:08:42Kafka streams API Apache Kafka streams
8:08:46API is used in multiple use cases some
8:08:49of the major applications where streams
8:08:51API is being used are mentioned as
8:08:53follows firstly the New York Times the
8:08:56New York Times is one of the powerful
8:08:58media in the United States of America
8:09:01they use Apache Kafka and Apache Kafka
8:09:04streams API to store and distribute the
8:09:06realtime news through various
8:09:08applications and systems to their
8:09:10readers followed by the New York Times
8:09:13we have trvago trvago is the global
8:09:16Hotel search platform they use Kafka
8:09:19Kafka connect and Kafka streams to
8:09:22enable their developers to access
8:09:23details of various hotels and provide
8:09:26their users with the best-in-class
8:09:27service at lowest prices and finally
8:09:31Pinterest Pinterest uses Kafka at a
8:09:34longer scale to power the real-time
8:09:36predictive budgeting system of their
8:09:38advertising system with Apache streams
8:09:41API backing them up they have more
8:09:43accurate data than ever
8:09:46[Music]
8:09:51now when we talk about big data right
8:09:55the very basic questions lot of time
8:09:57pops up what are five BS available in
8:10:02Big Data can anybody answer that okay so
8:10:06RI want to answer this okay let me
8:10:08unmute RI RI over to you first is volume
8:10:13a volume size of data how much is
8:10:15growing day by day and next is vary VAR
8:10:19is we have three types of data actually
8:10:21structur unstructured and Serv structur
8:10:23data structured data is nothing but
8:10:25relation database all those things
8:10:27unstructured data is audio images files
8:10:31all this data semc is XML files velocity
8:10:36is how much fast is
8:10:39growing okay very very much good answer
8:10:43when Big Data started IBM gave a
8:10:45definition with just three V the three
8:10:48vs were volume variety velocity so what
8:10:53was in volume volume was when we talk
8:10:56about like in terms of amount of data
8:10:58what we are dealing with right for
8:11:00example today's Facebook is dealing with
8:11:02very huge amount of data right when we
8:11:04talk about variety now is it only
8:11:06Facebook which is generating data no
8:11:09right even Twitter is generating data
8:11:11okay we are talking about just social
8:11:13media sites no we can can we talk about
8:11:15medical dos yes in medical domains also
8:11:17the data is getting generated right lot
8:11:19of big data is getting generated so
8:11:21that's a different variety of data we
8:11:23have structur data unstructured data
8:11:25like we can have videos audio right this
8:11:28is basically going to be called as a
8:11:30variety third component is velocity like
8:11:33I said to you Facebook is just a 10 to
8:11:3512 year old company and imagine the
8:11:37growth they have made it in just 10 to
8:11:4012 years imagine each each user is also
8:11:42doing this activity posting video audio
8:11:45all all kind of charts right so with
8:11:48that how much data they are dealing with
8:11:50so that is you will be calling that
8:11:53velocity with the pace they have grown
8:11:55up from scratch to this level with the
8:11:58speed they have grown up to this level
8:12:01is called velocity now these were the
8:12:03three components which were actually
8:12:04going in the market for long actually if
8:12:06you ask me these three were the major
8:12:08components even go today but slowly they
8:12:11started realizing there should be a
8:12:12fourth category of data which is
8:12:14verocity which also makes sense in Big
8:12:17Data because what basically happens is
8:12:19that the data what we receive to us
8:12:21right the problem with that data is we
8:12:23cannot expect that the data is going to
8:12:25be always a clean data there can be a
8:12:28missing data there can be a corrupted
8:12:30data in Middle how to deal with this
8:12:32scenario how to deal basically with
8:12:35those scenarios because that is a
8:12:36component of now that big data right so
8:12:39basically that corrupted or the bad data
8:12:41what we getting how to missing data what
8:12:43we getting so those Cate also they
8:12:46decided to call it as verocity now
8:12:49people started calling that okay these
8:12:50four are the major component but we we
8:12:53are not yet done now they started adding
8:12:55few more components they started saying
8:12:57that no we are not going to stop with
8:12:59four when you are saying that veracity
8:13:01can be added why not value because the
8:13:03data what I'm getting I I want to know
8:13:06what is the value of data what how much
8:13:08important that data is now somebody said
8:13:10that I want to visualize the data so
8:13:12visualization is should also be one one
8:13:14of the we after somebody started saying
8:13:17that I I want to see the vocabulary of
8:13:19the data or the validity of the data now
8:13:22they started keep on adding their V but
8:13:24majorly if you talk about there are four
8:13:26V which carries some good value okay and
8:13:29usually in any interview they will not
8:13:30expect you to know all the V physically
8:13:32if you know it all good but they will be
8:13:34just expecting you to kind of understand
8:13:36that okay do you know at least four V
8:13:39which are important if you can answer
8:13:41that they will be all good okay so
8:13:43sometime to make it tricky they ask you
8:13:45find me to just see that how good you
8:13:47are in terms of thinking like I
8:13:49generally do that so when when I
8:13:51generally ask questions I generally see
8:13:53that okay that guy must be doing four P
8:13:55let me ask five P let me see that is he
8:13:57able to think little beyond what what he
8:14:00knows already moving for we just talked
8:14:03about something called structured and
8:14:04unstructured data right can anybody
8:14:07explain me the difference between them
8:14:09what basically are this uh structure
8:14:12data and unstructured data I let me add
8:14:14one more component to it semi structure
8:14:16data that third category of data let me
8:14:19add it now can anybody explain this plus
8:14:22give me the difference between them as
8:14:24well can anybody give me an
8:14:27answer so unstructured not easy to save
8:14:30the data into rbms that's the question
8:14:32answer from Nish okay structure data is
8:14:35basically in row and column format easy
8:14:37to read and pass from word okay uh n
8:14:42simar saying structure data rbms data
8:14:44unru is like log audio video sem
8:14:48structured is like XML
8:14:51Json yes so lot of people are giving the
8:14:55right answer here if we talk about
8:14:58basically structur data if you go with
8:15:01basically 1980s and all when this Oracle
8:15:04and IBM and all those companies came up
8:15:06into the market so if we talk about
8:15:081970s 1980 even at that time they used
8:15:12to have data but that data was not huge
8:15:14it was small small data but you will be
8:15:17surprised to hear at that moment it was
8:15:19still a challenge to deal with that data
8:15:21though it was having some sort of
8:15:23pattern it it used to have some sort of
8:15:25pattern and it's a small data now people
8:15:28used to think that how to use it how to
8:15:30manage it how to store it where to store
8:15:33it all those questions were coming in
8:15:34people mind and that is where companies
8:15:37like Oracle IBM and all came up into
8:15:39Market with their rdbm solution now they
8:15:42started delivering this solution that
8:15:45you can now store the data you can now
8:15:47process that data what what sounding
8:15:49like a having some pattern and you will
8:15:51be all good and today I need not tell
8:15:53you that today where these companies are
8:15:55like orle Microsoft you not in fact
8:15:57everybody of you must be willing to work
8:15:59for them if given a chance right so they
8:16:01they are basically now the market G and
8:16:04they have given the solution for for
8:16:05that it was going all good but now
8:16:08slowly what happens the other kind of
8:16:11data started coming so the data what
8:16:14they were dealing with was structured
8:16:16kind of data but with today world right
8:16:19like like Facebook came up in few years
8:16:21back right now as soon as Facebook came
8:16:23up into Market I'm just giving you an
8:16:25example now they you what you do in
8:16:27Facebook on Facebook you either upload
8:16:30video upload audio pictures right so you
8:16:33started dealing with this kind of data
8:16:35now do you think this kind of data can
8:16:37be dealt with my rdbm system answer is
8:16:40no right because now we cannot deal with
8:16:42this kind of data now we cannot call
8:16:45this data as structure data because this
8:16:47kind of data do not even have any sort
8:16:50of packing and that is where we started
8:16:54calling it as unstructured data mean any
8:16:57sort of data which do not have pattern
8:17:00kind of thing like your audios videos we
8:17:03started calling it as unstructured data
8:17:06now the third category of data is semi
8:17:10structured data right so there there are
8:17:12some sort of files for example let's
8:17:14talk about XML data so as soon as you
8:17:17see XML what is it sounding like does it
8:17:20have pattern or
8:17:22not does it have pattern or not XML
8:17:25files can I get the
8:17:27answer XML file Json files do they have
8:17:30pattern yes yes it it has pattern so as
8:17:34soon as I say that it has
8:17:36pattern first answer which must be
8:17:38coming in your mind should be that okay
8:17:40it is a structure data but now as soon
8:17:44as you tell me that it's a structure
8:17:45data my question for you is in that case
8:17:48can you do all the activities what you
8:17:50can do in rdbms to XML data I know you
8:17:53can do today even you can deal with
8:17:56unstructured data as well because they
8:17:57have introduced glob and clo data types
8:18:00as well but is that efficient can you
8:18:03deal can you fire Triggers on that can
8:18:05you do all the things which you can do
8:18:06with your traditional data no right now
8:18:10it started sounding to me that it's a
8:18:12unstructured data now I'm confused
8:18:15whether it's a structur data or
8:18:16unstructured data that's where they
8:18:18created a third category called as sem
8:18:21structure so that they can keep it in
8:18:23the middle the data which is sounding
8:18:25something of the structure type or as a
8:18:27unstructured type they started create
8:18:30they created a new category called as
8:18:32semi structure data everybody clear on
8:18:34this part what basically structure data
8:18:36semi structure data unstructured data
8:18:38okay moving further now I have another
8:18:41question for you how had do
8:18:45from your traditional processing system
8:18:49using
8:18:50rdbms can I get an answer what would you
8:18:53answer this part so these are like warm
8:18:55up kind of question usually in interview
8:18:57they will not start with the most
8:18:58complicated questions with right so
8:19:00these are kind of they kind of warm up
8:19:02they want to see your level of expertise
8:19:04how much you know so basically that is
8:19:06what is happening here so can you answer
8:19:08this part how do differ from traditional
8:19:12processing system using rdbm and friends
8:19:15I have a request rather than raising
8:19:17your hand please type it on chat window
8:19:20because this chat window I want to make
8:19:21it more interactive okay okay I don't
8:19:23know why your name is showing up Jo but
8:19:26let let's take it processing is done the
8:19:29data is done no input output okay Hado
8:19:33can store and process any type of data
8:19:37where R can store only relational data
8:19:41okay I can take this answer distributed
8:19:44storage processing okay good answer was
8:19:48par processing large data distributed
8:19:51know so few people have just started
8:19:53answering what is right don't answer me
8:19:55that I want difference right read the
8:19:58question properly the question St give
8:20:01me the difference do not tell me what
8:20:03this do and that do tell me the clearcut
8:20:05difference what are the differences what
8:20:07you notice with ad
8:20:09system or I can ask you in another way
8:20:12as well can had do replace at the system
8:20:16in future or maybe is it really
8:20:18replacing right now okay this question
8:20:21can be asked in this way as now can you
8:20:24answer me this part so I hope anybody
8:20:27who know how do Basics should be able to
8:20:29answer this easy can I get some answers
8:20:32now who want to come on air also can
8:20:34tell me I can unmute you if anybody want
8:20:36to come on air and
8:20:38answer both are complimentary to to each
8:20:41other okay good anyone who want to come
8:20:45here want to answer it this will give
8:20:47you good confidence also when you will
8:20:49be speaking in interview it cannot
8:20:53replace rdbms but I want reason that
8:20:55that's everybody know it cannot replace
8:20:57rdbms but what's the reason very good
8:21:00very good asset property is not
8:21:03supported or in other words can I say
8:21:06cred operation is not supported create
8:21:09delete update can I do that at Pro level
8:21:13no right so that
8:21:15is the major reason you cannot go with
8:21:18Hado systems like you cannot replace
8:21:21them also when the data is small okay if
8:21:25the data is small and it's a structur
8:21:28data which is going to be more efficient
8:21:30RMS or Hado
8:21:32systems yeah that's basically in the
8:21:34latest version they are supporting it
8:21:35but there are lot of distinctions in it
8:21:37that's all about Hado 3.8 stre which
8:21:41which is yet to come in the market
8:21:42properly so hold on until the time it
8:21:44come up because it it has lot of
8:21:46restriction it have right now lot of if
8:21:49you have already seen it you might be
8:21:50aware of
8:21:52this is used when we are more of right
8:21:56once read multiple times very good it
8:21:59follows warm principle W RM which means
8:22:04write once read multiple times okay so
8:22:09and one more important difference RS is
8:22:12free of C is RBM is free of cost no
8:22:17right so basically it's a license
8:22:19software you need to pay for it right
8:22:22but when it comes to it's completely
8:22:24open source there are companies who are
8:22:27now basically making money with this as
8:22:29this because it's an open source
8:22:31Community righto is an open source
8:22:32Community now if you get stuck where you
8:22:35will get the support there is nowhere
8:22:37you can get the support if you get stuck
8:22:39in rdb system there are companies to
8:22:41support you but what about had system if
8:22:43you get let's say some
8:22:45who will fix it for you right it's an
8:22:47open source so basically anybody can
8:22:49come and fix it but let's say if
8:22:50nobody's fixing your bug then what right
8:22:52so in that case what people are doing is
8:22:56they are taking support okay so a lot of
8:22:59companies are providing support also on
8:23:03this so a lot of companies are providing
8:23:06support in terms of now they making
8:23:09money from it they say that okay if you
8:23:11want to use use it we will give you the
8:23:14appropriate support what is required and
8:23:16you have to pay the money for us so a
8:23:17lot of companies actually came up like
8:23:20like this kind of idea that they are
8:23:22expert in a do and if you require any
8:23:24support we will help you with that so
8:23:26that is one thing which is happening now
8:23:28rdps can only deal with structur data
8:23:31right basically when you talk about
8:23:32rdbms you cannot deal with unstructured
8:23:35kind of data though you can right not
8:23:37deal with it as well because like
8:23:39somebody just argued with me she said
8:23:42that you know in the latest version of
8:23:44high that trying to support even C
8:23:45operation right but it it has lot of
8:23:48restriction it's not efficient at of the
8:23:50moment till the time it do not come out
8:23:53of the beta phase we cannot say anything
8:23:55about that similarly rdbms also started
8:23:58creating clock and block dat C O and B O
8:24:03right where what they say that you can
8:24:05now store unstructured data and can work
8:24:07on it but are there efficient answer is
8:24:10no right so similarly like most of the
8:24:13data what you deal with withm is going
8:24:16to be structured data but when it comes
8:24:18to youro system it can be unstructured
8:24:22sem structure as well as structure data
8:24:26right because if we take an example of
8:24:28he and all they deal with structur data
8:24:31right so that is one thing which is
8:24:34there rbms you work just on a single
8:24:36machine right so let's say you work on a
8:24:38single laptop where you have rbms
8:24:40install and working but when it comes to
8:24:43you are working in a distributed fashion
8:24:46right there will be multiple machines
8:24:48which can be involved in this case right
8:24:51and as I said in our mostly when the
8:24:53data is small your speed will be very
8:24:58fast your computation is going to be
8:25:01very quick at the same time with Hado
8:25:05your computation speed is not going to
8:25:08be that great okay with the small data
8:25:13it's not going to be that great with the
8:25:14par stock pitching into the market
8:25:16that's a different story now they are
8:25:18actually picking up basically because if
8:25:20your memory is good then you can
8:25:22actually make up a good speed but with
8:25:24redition ad when we talk about map
8:25:26reduce the speed are lower I believe if
8:25:29you have already done the classes of
8:25:31Maes you might have already noticed this
8:25:34thing right when you were doing that
8:25:36work count example right I hope
8:25:38everybody must have done V example in M
8:25:41right if you have taken the session
8:25:42right so in that case you might have
8:25:44seen the speed of work example that it
8:25:46is not that fast so that is one thing
8:25:50with Hado system if especially the data
8:25:53is smaller your speed is not
8:25:55comparatively with rdbms it's going to
8:25:58be
8:25:59slower now which brings me to another
8:26:02question can anybody tell me the
8:26:05components of had and their services in
8:26:09fact I'm showing you all the components
8:26:11can you explain these components what
8:26:15are the components available of do and
8:26:18what are the services what they
8:26:22provide can I get an
8:26:24answer very good so basically when we
8:26:27talk about hdfs right it is for your
8:26:30storage side okay very good can I get
8:26:34more
8:26:35answers I believe everybody must be
8:26:37knowing this part uh n is saying storage
8:26:41hdfs is for storage Yan is for proc
8:26:44processing cluster very good so you can
8:26:46say Yan is a cluster resource manager
8:26:50right there are a few more thing can I
8:26:53get more answers Name note manages the
8:26:56cluster data Note stores the data okay I
8:27:00can take this answer but partially they
8:27:03name note manages the cluster can you be
8:27:05little more explicit in this part Yan
8:27:09for resource allocation very good our M
8:27:12saying Yan to run map produce
8:27:15okay uh you can say to schedule map
8:27:18produce that would be better anyone
8:27:20right rather than saying to run map ruce
8:27:22can I say schedule map
8:27:24produce right that would be a more
8:27:26appropriate answer here but a good
8:27:28attempt name note has meta data very
8:27:33good right name node has metad data you
8:27:37can take it of something of this sort so
8:27:41uh in your real term scenario right let
8:27:44me go back you can simply relate it with
8:27:47this your real time life also right
8:27:49let's say in real time scenario what
8:27:51happens let's say you have a boss okay
8:27:54how many people have a kind of a smiling
8:27:58boss good
8:28:00boss smiling BS but a cunning BS very
8:28:03cunning he plays with your
8:28:07emotions when I say emotions I basically
8:28:10mean that basically he kind of is a very
8:28:12clever boss anybody who have it okay sh
8:28:17have it Aran have
8:28:20it is very interesting okay so let's say
8:28:24this boss okay let me draw a smiling
8:28:27boss okay The Smiling boss now usually
8:28:30every boss is smiling right they just
8:28:32keep on smiling but the the point to
8:28:34note here is the smile is cunning smile
8:28:37or what kind of SM now what next now
8:28:40these are people who are working right
8:28:42like like sh said he reports to such
8:28:45kind of boss right so basically VI said
8:28:48he report to such kind of Boss who a
8:28:50very clever boss right now let's say
8:28:54um narim is saying that okay fine his
8:28:57boss is also kind of a very elive boss
8:28:59but you know behind that inoc there
8:29:01might be lot of cleverness iting right
8:29:04behind the SC so let's say these three
8:29:06people are reporting to this clever boss
8:29:10what happens boss get products right
8:29:13boss get project let's get the project
8:29:16now what happens boss will be getting a
8:29:18project these three people are reporting
8:29:19to this to this boss right so what they
8:29:22will do so boss usually will distribute
8:29:24the project right so let's say the
8:29:26project was P he distributed into three
8:29:28parts P1 P2 and the third component He
8:29:32Made It P3 what he's going to do so he's
8:29:35going to keep this P1 or he going to
8:29:37come to sh and say that work on P1
8:29:39project he will come to we and say that
8:29:41work on P2 project he will come to
8:29:43nursing will say that work on Project
8:29:46right he will say that now all the three
8:29:49people are working properly given the
8:29:51project on timeline boss is going to be
8:29:53happy now imagine a scenario that we bch
8:29:58is boss not basically ditch but maybe he
8:30:02said that okay fine my boss was inocent
8:30:04so let me take an advantage of him
8:30:06telling him that you know I have a
8:30:07family mcy I can't work I want I have to
8:30:10take Le something of family mergency
8:30:12kind of thing now bosses in right
8:30:14because boss need to work on these two
8:30:16projects right P1 and P P3 who will
8:30:19deliver P2 project so that is the
8:30:20problem so boss what they do they come
8:30:23up basically with earlier kind of
8:30:26backups and boss is clever remember
8:30:28right so boss is going to call Shri in
8:30:31his cabin and say to Shri sh you know
8:30:34you're doing very very good job right so
8:30:37you're doing very good job I'm thinking
8:30:39to promote you if you keep on working
8:30:41like that you will get promoted very
8:30:42quickly in the market for sure
8:30:44and um so you should take up some senior
8:30:47responsibilities from now right so as
8:30:50soon as she will she will not hear
8:30:52anything she will just hear that you
8:30:53know boss is telling me promotion work
8:30:56okay he will just hear promotion word
8:30:57and he will be very happy in his mindset
8:31:00suddenly his boss will throw a Time B
8:31:02Because boss is clever so boss usually
8:31:04what he will say we just discussed right
8:31:06that you are going to take up some
8:31:08senior responsibilities so can you do
8:31:10one thing now can you basically take the
8:31:13backup of project now what happen in
8:31:16this case immediately as he said that
8:31:18can you take the backup of basically V
8:31:20project now he came back to S and
8:31:23started arguing you know I'm already
8:31:24busy but he said that you have to just
8:31:27take back up I'm not asking you to work
8:31:29right just take back up off work if you
8:31:30did me then only you have to work
8:31:32anywhere you are a senior candidate
8:31:34anywhere boss is work she will not be
8:31:36able to say no right you have to work
8:31:37for him similarly he will go to we he
8:31:39will do the same stuff right because
8:31:41managers usually tell this right
8:31:43everything is confidential so same thing
8:31:45he will do with SH also this is
8:31:47confidential do not think about it
8:31:49outside same way will call Vi tell the
8:31:51same story and he will ask now to take
8:31:54the backup of nursing project right so
8:31:56he's backing up this project similarly
8:31:59nursing if you note that nursing have to
8:32:01back up she project he did the same
8:32:03thing with nursing also and basically
8:32:05now nursing have to back up this
8:32:07basically D now if you not in this
8:32:10situation if we to have emergency leave
8:32:13now B will not face any trouble right
8:32:16now boss will not face any trouble
8:32:18because the work who has to do the extra
8:32:20work in fact who will be sad in this
8:32:23scenario definitely boss is not going to
8:32:25be sad but who is going to be sad in
8:32:27this scenario definitely she is going to
8:32:28be sad right so definitely she is going
8:32:30to face the heat of doing the extra work
8:32:33now similarly why I'm telling you all
8:32:35this because all the components you can
8:32:37relate here so basically the in when we
8:32:40talk about Hado Hado is not much
8:32:42different with this what Hado do is and
8:32:45boss is also keeping one more
8:32:46information right so boss is also
8:32:48keeping information of let's say what
8:32:50all projects he have who is working on
8:32:53what project let's say Shri is working
8:32:55on Project P1 and he have also a backup
8:32:58of P2 so all these details what boss is
8:33:00keeping when we talk about had Hado is
8:33:03kind of doing that exactly the similar
8:33:05kind of stuff but when we talk about
8:33:07boss so first thing is now in had word
8:33:09every human is going to be replaced by
8:33:11machines now the first compon what we
8:33:14were talking about right so if we talk
8:33:16about the first component it was name
8:33:18the boss is representing n second
8:33:23component was Data node these employees
8:33:26who are working basically for that boss
8:33:29you can represent them as data node the
8:33:32node where you are doing all the
8:33:35processing because employee do the work
8:33:37right employee do the work boss only
8:33:39instructor manage right same story here
8:33:41so these are your data third property
8:33:44what is this metadata right even name
8:33:47node is going to keep all the data about
8:33:49data that is called your metad data now
8:33:52what is this this basically secondary
8:33:54name node and all those stuff so
8:33:56basically we require some like this is
8:33:58the backup for the data part right for
8:34:00this data file P1 P2 P3 but what about
8:34:04this boss back so basically we want to
8:34:06create some backups for that so that's
8:34:08the reason we keep like casting name Lo
8:34:10and all those stuff there okay so
8:34:12basically that is the back part that's
8:34:13one of the component now that that is
8:34:15means like let's say in your company
8:34:17they have also back up a boss like in
8:34:20case if this boss leave then I should
8:34:22have all the details so that's that's
8:34:23what the backup me there now there is
8:34:25one more component called as load
8:34:27manager resource manager what are these
8:34:30things now what basically a boss is
8:34:32doing right what is boss doing boss is
8:34:34having the skill set to schedule the job
8:34:36right he only decided where to send what
8:34:39all those details right so that you can
8:34:41call it as a part of your resource maner
8:34:43kind of scheduling the work who is
8:34:45helping you to schedule the work you can
8:34:48call it as a resource manager right boss
8:34:51have that skill set similarly in your
8:34:53Hado what your resource manager is going
8:34:55to shule everything up now the the last
8:34:58part right node manager now do you think
8:35:01you will be able to work on this project
8:35:02without any skill set no right you
8:35:05require some skill set to work on that
8:35:07project right so that skill set you can
8:35:09relate it like the node manager which is
8:35:12managing your own no
8:35:14means you can delete it like your skill
8:35:15set which is helping you to solve a
8:35:17project right so same way the node
8:35:20manager you can say that it is managing
8:35:23the whole node it is kind of helping the
8:35:25node to execute the task that you can
8:35:28call it as node manager so this is how
8:35:32you can relate everything I hope with
8:35:34this example now it will make your life
8:35:36easy to remember all these components
8:35:38because this is a very important
8:35:40questions in basically in interview
8:35:42questions they generally ask question
8:35:44what are the configuration files what
8:35:45what basically are the main components
8:35:47so you should be aware of this and
8:35:49that's the reason I've explained you
8:35:51with this anology so that you get some
8:35:53idea and you can relate to it that if
8:35:55you have to explain in interview you do
8:35:57not remember all this stuff okay let's
8:35:59move
8:36:00further now so can I ask this question
8:36:03now what are the main had configuration
8:36:06files right so basically now we talking
8:36:08about configuration F this is related to
8:36:10mostly Administration interviews now
8:36:13when you talk about H administrator
8:36:15right so there will be few files which
8:36:17you need to configure so I hope there
8:36:19are people in this batch in this session
8:36:22also where people must have done some
8:36:24Hado administrator first right so I'm
8:36:26assuming that you people know that
8:36:27people who do not know that that's let's
8:36:29not worry about this that's the reason
8:36:31I'm answering this directly so there are
8:36:33few important FES here one is had
8:36:36environment. this is where you kind of
8:36:39mention all your environmental variables
8:36:41for example where is your Java home
8:36:43where is the Hado form all those things
8:36:46to Define in Hado
8:36:48.s for.xml so these this file basically
8:36:52Define where your let's say your name Lo
8:36:54is going to run right so you need to
8:36:56tell the address of your name Lo where
8:36:58you want to run that maybe you want to
8:36:59run a some machine at 9,000 F so you
8:37:02will be telling all that in your for
8:37:04site XML when we talk about hdfs XM here
8:37:09we talk about what should be the
8:37:10replication Factor where should
8:37:12physically my data Note should be
8:37:14present where physically my name note
8:37:16should be present all those things to
8:37:18Define in hdfs site. XML Yan site and M
8:37:24site basically defines the map jobs
8:37:27right what kind of cluster you are going
8:37:29to use are you going to use let's say
8:37:31Missour or you going to use Yan or you
8:37:33going to run a local distributed one all
8:37:36those things will be defining here also
8:37:38you will be defining where your resource
8:37:40manager should be running it should be
8:37:42running on this machine know 9 9,1 port
8:37:45or whatever Port you want to Define
8:37:47right so those information you will be
8:37:50defining in J site or map rite. XML not
8:37:56last two files are Masters and slaves
8:37:59file in Master's file we usually mention
8:38:03where my secondary name note would be
8:38:06running when I say secondary name it's
8:38:08it's like a backup not exactly I should
8:38:10call it as a backup but I should call it
8:38:12as a snap snapshot of the name it's
8:38:15something like this like somebody just
8:38:17copying the metadata that's it it's not
8:38:19going to become active as soon as the
8:38:21main Lo is down it is copying the data
8:38:23so that if name Lo is down at least I
8:38:25should have a backup that's it slips
8:38:28very clear with the name where are all
8:38:31my data notes what all machines are
8:38:33going to be my data notes that thing we
8:38:35Define in my slaves machine so people
8:38:38who SL file sorry so basically people
8:38:41who have done this administrator you
8:38:43must have basically played with this
8:38:45file these are the major Hado
8:38:46configuration F there are others as well
8:38:49right like hyp site XML there are others
8:38:52like HB site XML now these are very much
8:38:55kind of tool specific so that's the
8:38:57reason they will not be called as the
8:38:58Maino configuration P when somebody say
8:39:01Maino configuration P your answer would
8:39:04be these seven PES you need to remember
8:39:06these seven PES basically these seven
8:39:08files are the one which you will mention
8:39:11if you are going for administrator
8:39:13interview expect good number of
8:39:16questions in this okay they will ask you
8:39:19to explain each and every file how
8:39:21basically you will be what you do in
8:39:23which file which I just explained you
8:39:24make a note of that and that is
8:39:26definitely one of the favorite questions
8:39:29of the interviewer when you go for ad
8:39:32administrator kind of grow moving
8:39:35further now let's talk about some hdfs
8:39:39questions in terms of hdfs now my
8:39:42question is hdfs stores data using
8:39:45commodity Hardware which has higher
8:39:49chance of failure which is obvious right
8:39:52because my laptop can be one of the data
8:39:54Note right now definitely my data Note
8:39:57can fail at any moment now so how hdfs
8:40:01ensures the fault tolerance capability
8:40:05of the system can anybody answer this
8:40:09very good word can I get more
8:40:12answers
8:40:13I just answered it beforehand itself
8:40:16remember that boss and employe
8:40:19relationship what that boss was doing
8:40:22boss was keeping backup right boss was
8:40:25creating backup similarly Hadoop also
8:40:28create backup right that backup in Hado
8:40:32word is called replication right so
8:40:35basically if you have let's say one
8:40:37block P1 you are creating one backup or
8:40:40two backups of basically that block P1
8:40:42and that is called replication this is
8:40:45how Hado is ensuring that in case of any
8:40:50failure also there should be no mistake
8:40:54okay you should not be using data
8:40:56because in case if one machine fails
8:40:58also it's okay it will all work fine for
8:41:00me so that is what going to happen so
8:41:03very good A lot of people have given me
8:41:05the right answer in this case so block
8:41:08replication is the answer in this case
8:41:11as you can see in this example also like
8:41:13block one is replicated three times if
8:41:15you notice here right so block one is
8:41:18replicated three times similarly if you
8:41:20notice block two is also replicated
8:41:22three times so these are four different
8:41:24machines and I have replicated block one
8:41:27block two block three block four block
8:41:30five okay so this is what is happening
8:41:32edit log and Fs image is used to
8:41:34recreate a image FS image and edit log
8:41:38are two different things these these
8:41:40basically are two different things you
8:41:42cannot create the application if you
8:41:44want to know about that let me just
8:41:45answer you here so what happens what is
8:41:48phasically This Ss image and edit logest
8:41:51now what happens basically you have a
8:41:54name not right you have a name note
8:41:56where you keep the data in name not
8:41:59anybody have answer where you keep the
8:42:01data in name node no not ACC F initially
8:42:05where you keep the data in name node and
8:42:07not SDF I'm asking in name note this can
8:42:10also be an interview question very good
8:42:12answer
8:42:13we keep it in memory why let answer this
8:42:17let's say there is one client came up
8:42:19this is one client this is Cent C2 this
8:42:22is Cent C3 let's say there are multiple
8:42:24clients okay now what is happening in
8:42:27this case is let's say if this was my
8:42:30name node and let's say my data in name
8:42:32node is my metadata is let's say setting
8:42:34in this what would be the problem here
8:42:37let's say client one came and want to
8:42:39access some data what is going to happen
8:42:41this data will be because any processing
8:42:44which need to happen right that happens
8:42:46in memory only right now this data will
8:42:49come to the memory once this memory work
8:42:52will be over then what will happen then
8:42:55basically again it will come back to
8:42:57this right it will remove it from memory
8:42:59now don't you think there's a input
8:43:01output operation happening an input
8:43:04output operation is always expensive
8:43:06right now imagine if there are multiple
8:43:09clients asking at the same time to the
8:43:12name don't you think there will be too
8:43:14many input output operation every time
8:43:16you need to basically bring the block to
8:43:18the memory and then basically do the
8:43:20stuff and give the output now this is
8:43:22something which we want to avoid in
8:43:25order to avoid what they came up with
8:43:27the idea is that whatever you are going
8:43:30whatever metadata you are going to
8:43:32create should directly be created and
8:43:35kept in memory okay that is what the
8:43:38idea they came up they are not going to
8:43:39keep any data in the dis right now
8:43:41should be directly kept in the memory
8:43:43which brings another question for you in
8:43:46a serious question for you now as since
8:43:50you are telling that the data what you
8:43:52can keep in memory but my Ram is
8:43:55volatile when I say volatile I will lose
8:43:58I can lose the data at any moment right
8:44:00that's very obvious I can lose the data
8:44:02at any moment because my Ram is always
8:44:04going to be volatile right now I restart
8:44:07my system my Ram data is gone right so I
8:44:10will lose all the metadata in that case
8:44:12how I I will ensure that I should not
8:44:14lose my metadata now what they started
8:44:16doing is okay fine I will create
8:44:19everything in name note memory only but
8:44:22what I am going to do is whenever any
8:44:26like what I'm going to do add some
8:44:28interval of time at some interval of
8:44:32time I will keep on taking a backup of
8:44:36that metad data in dis okay in this I
8:44:41will keep on taking the backup of that
8:44:44data and whatever backup you are taking
8:44:48in the disk from memory is called your
8:44:51FS image now don't you think this FS
8:44:55image is going to be big right this FS
8:44:58image is going to be big now what
8:45:01usually happens is usually let's say
8:45:03today my FS image is version FS1 now
8:45:07usually the backup what they take is
8:45:09every 24 hour just give me one minute
8:45:13if any pop up kind of thing comes up it
8:45:16actually stop my system now anyway now
8:45:19what basically is going to happen so
8:45:21what happens is every 24 hours we can
8:45:25usually do this okay so basically today
8:45:28we did some backup tomorrow again this
8:45:30metadata is going to do backup now which
8:45:32brings another problem the problem now
8:45:35again would be let's say what about in
8:45:3818th hour or maybe in 23rd hour my
8:45:41machine P or my R in that case I'm going
8:45:44to lose the 23 data right which is again
8:45:47not good what should I do for that now
8:45:49for that what they came up is that let's
8:45:52create whatever activity is happening
8:45:55here I will keep on writing in a small
8:45:58file that will be created for let's say
8:46:0124 hours okay and that file is called as
8:46:05edit log then what going to happen for
8:46:0924 hours whatever activity you are doing
8:46:12will be be getting stored as a edit law
8:46:15Okay now what's going to happen after
8:46:17every 24 hours this ss1 plus edit log is
8:46:21going to be added up and FS2 will be
8:46:24created now in this scenario even if I
8:46:27lose the data in 23rd L my edit log will
8:46:30be having the data and that's how I am
8:46:32ensuring that I'm not losing any data
8:46:37now are you clear about this edit log
8:46:38and Fs image question I usually see that
8:46:41people are very confused with this logic
8:46:44that what is FS image they just kind of
8:46:46mg up and come back and tell that you
8:46:48know I know FS image and edit what what
8:46:50exactly are they I have seen people
8:46:53actually kind of confused with this so I
8:46:55hope you should be very clear now on
8:46:56this in which file can this back of time
8:46:58interval be configured so basically
8:47:00wherever the physical location of main
8:47:02Ro you have configured and where you
8:47:04configured I just told you what in hdfs
8:47:07s. XML right so wherever you have
8:47:09configured that so there will be a name
8:47:11directory in it in that name directory
8:47:14there is another subd directory called
8:47:15as current directory in that current
8:47:17directory there is another directory
8:47:19called as SN and directly which is
8:47:21secondary name directory there we keep
8:47:24this SS image and edit Lo clear about
8:47:27this part so this is where we basically
8:47:29keep that so can you please summarize
8:47:32the answer once sure the same answer of
8:47:34FS image and edit log but correct
8:47:38correct not exactly two files one file
8:47:40will be kind of very big file so that's
8:47:42the reason we are creating a smaller
8:47:44version of that F called so that every
8:47:4724 hours activity we can edit logs are
8:47:50stored in the disk as yes in the name so
8:47:53it's kind of act like a back that's it
8:47:56it's a very good interview question
8:47:58that's the reason as soon as this
8:47:59question came up I thought to answer it
8:48:01up though it's not a part of this SL but
8:48:03this is a very famous inter question
8:48:05that can you explain this suim Ed BL and
8:48:08I can tell you that most of the people
8:48:10fail to explain this here memory is R
8:48:13correct correct so it's just metadata
8:48:16which is store correct correct it is
8:48:18just creating a backup of that metadata
8:48:20for the 24 hours activity that means
8:48:23edit log is getting erased and getting
8:48:25and get new data yes yes every 24 hours
8:48:29it just keep on working and kind of er
8:48:31the data that's what keep on happen okay
8:48:33or it creates a new version it depends
8:48:35how your admin have configured that what
8:48:38if the block data goes more than the mem
8:48:43now in that case there is something
8:48:44called as pill usually it's not it do
8:48:47not go like that but there is some
8:48:49concept called as pill so in that case
8:48:51there will be some input output
8:48:53operation happening you have to deal
8:48:54with that so then you are making your
8:48:56name not slow so you have to make sure
8:48:58if you should have a good configuration
8:48:59but if you do not have it then in that
8:49:01case you have to do input output
8:49:02operation no other option then you have
8:49:04to keep in the dis and then the data
8:49:06will be having input outut the window is
8:49:08of 24 backup can be change yes it can be
8:49:11change to summarize this what we just
8:49:14talked about so in name note uh in name
8:49:17note basically you will be storing all
8:49:20the data but the problem is my Ram is
8:49:22going to be volatile now because of that
8:49:25I want to definitely want to have a
8:49:27backup now we keep a backup in the disk
8:49:30and that whatever backup we keeping we
8:49:32call it as FS image now FS image backup
8:49:35is always taken in 24hour slot now the
8:49:38another problem started with this what
8:49:40happen if I lose the data and 23rd in
8:49:43that case I should again create a
8:49:45smaller version of the file called as
8:49:48edit blog okay that will also be a now
8:49:51these things will be added and will be
8:49:53basically given what can be the Ram size
8:49:56the bigger the better so definitely
8:49:58there is no right answer for it now
8:50:01definitely if you say that 32 GB is good
8:50:03I will say how about 128 GB if you say8
8:50:06GB is good I will say how about 256 GB
8:50:09because that will be better what if I
8:50:11more metadata so we can keep on arguing
8:50:13and keep on increasing right so the more
8:50:15the r better it is for you correct
8:50:18correct with with every change in that's
8:50:20a Lo yes correct that's that's what
8:50:22basically
8:50:23happened now let's move further so I
8:50:26hope now everybody should be clear with
8:50:28this question though it's a separate
8:50:30question but actually it's good that you
8:50:32brought it up because that's one of the
8:50:34very famous interview question so I
8:50:35thought would cover it up now another
8:50:38question what is the problem in having
8:50:41lot of small files in GFS please provide
8:50:44one method to overcome this problem can
8:50:47I get this answer can I get this
8:50:49answer so what are the problems if you
8:50:52will have small files in GFS and also
8:50:55can you give me basically a method to
8:50:56overcome this problem change block size
8:50:59if you change the block size no I I want
8:51:02a better answer I want a better answer
8:51:05name note memory will be overload good
8:51:08good now you are coming to right track
8:51:11three Ram will run out yeah because if
8:51:14you will have small data right if you
8:51:17will have small files definitely your
8:51:19metadata is going to be kind of too much
8:51:21right your metadata entry will be too
8:51:23many and that's how you will be kind of
8:51:25filling up your RAM right of your
8:51:27basically name node we just learned that
8:51:29every metadata is stored in the ram of
8:51:31the name node now what are the solution
8:51:35for it so this is the problem what is
8:51:37the solution for
8:51:39it having larger data clock size so that
8:51:42name node will have reasonable metadata
8:51:44to hand it I can take this answer but
8:51:47I'm expecting a better answer
8:51:48she more map jobs will be used yes
8:51:51that's also one of the problem merge
8:51:53them and save them very good job okay
8:51:56what's your name basically
8:51:59joala I hope that that should not be a
8:52:01real name I don't know why it's sh me
8:52:05joa this is your real name because it's
8:52:08telling me twice joa joa okay so it
8:52:11should be once right it should be once I
8:52:13should call it okay then fine now so I
8:52:17don't know this is occuring two times so
8:52:19that's feeling it real so increasing
8:52:21block size in sdfs merging the file with
8:52:24same and it's easier to read and write
8:52:25the data somebody just answered can you
8:52:28combine
8:52:30everything that is the right answer we
8:52:33can
8:52:34create H file in your windows what you
8:52:39do you create a zip file right or a r
8:52:41file similarly in Hadoop also you can do
8:52:44that you can create a h file which is
8:52:48called as Hadoop art F so you can bring
8:52:52all the small files into one folder
8:52:54together kind of zipping it together now
8:52:57basically with that what's going to
8:52:59happen it's going to just keep only one
8:53:02metadata entry for it the metadata entry
8:53:04is going to be reduced how to do that
8:53:06this is the command archive now hyen
8:53:10archive name whatever archive name you
8:53:12want to give it your input location and
8:53:14output location okay so basically this
8:53:16is how you can deal with the smaller
8:53:21files as well better to create a z this
8:53:24is what you do in the real time also
8:53:25right when you have multiple files of
8:53:27the same type you zip them right just to
8:53:30keep them together so the same thing you
8:53:32will be doing in Ado as well moving
8:53:36further now another question this is
8:53:38also a very interesting question and
8:53:40easy question also Suppose there is a
8:53:42file of size 514 MB stored in hdfs 2.x
8:53:48using default block size configuration
8:53:51and default replication Factor we did an
8:53:54assignment with image files I'm want
8:53:56LinkedIn as okay okay now I got it so
8:54:00using uh default block size
8:54:02configuration and default replication
8:54:04Factor then how many blocks will be
8:54:06created in total and what would be the
8:54:09size of this block okay before you
8:54:11answer this can I get an answer what is
8:54:13the default replication factor and what
8:54:16is the default block size if I'm talking
8:54:18about had
8:54:192.x very good so replication as
8:54:22everybody said 3 m what is the size very
8:54:25good 128 M now it's very easy to answer
8:54:28can everybody answer how to split this
8:54:30514 mbf file how to split this 514 mbf
8:54:35file it's in front of you you can do the
8:54:37calculation and give me the answer as
8:54:39well can I get this answer very good 15
8:54:43block lot of people have given me
8:54:45basically less they said five block but
8:54:48don't you think there will be a
8:54:49replication also of all the block so a
8:54:52lot of people who are giving me this
8:54:54answer of four block is completely wrong
8:54:56and right because there will be a block
8:54:59of 2 MB as well right if you notice
8:55:03what's going to happen this is 128 into
8:55:054 is basically 52 right there will be 2
8:55:08MB block so there are going to be five
8:55:10block because the replic is five now
8:55:13sorry replication factor is three so
8:55:15it's going to be 5 into
8:55:173 okay this is how basically you will be
8:55:20calculating this is very famous
8:55:21interview question moving further how to
8:55:25copy a file into
8:55:28hdfs with a different block size to that
8:55:32of existing block size
8:55:35configuration can I get an answer what
8:55:38basically I'm asking is let say you have
8:55:40a block size of one 20 by the but when
8:55:44you are copying that data right when you
8:55:46are doing let's say sdfs hyp input sdfs
8:55:49Hy input maybe you want to now use the
8:55:52block size of 32 bit not the default of
8:55:55128 bit then what you will do to achieve
8:55:58this yes there's a parameter what what
8:56:00is that
8:56:02parameter what is that parameter block
8:56:05size um no can you see this
8:56:10BFS dot block size okay so what you need
8:56:15to do you need to just Define the
8:56:18basically the bytes what you want to
8:56:19mention so 32 bytes is equivalent to
8:56:21this number okay 32 byes is equivalent
8:56:24to this number so you need to basically
8:56:26Define the bytes what you want to put it
8:56:28up now while doing any command let's say
8:56:31hyen put or maybe hyen copy from local
8:56:35there you can mention this DFS do block
8:56:39right and whatever number of B you want
8:56:42to mention so you can mention that okay
8:56:45if you want to check the block size you
8:56:47have another command called as stat I do
8:56:49FS ion stat and you can see all the
8:56:52statistics related to it so it will tell
8:56:54you how many bites it is to basically
8:56:57distributed and everything up you can
8:56:59basically directly use this F thing and
8:57:01you will get that output okay this is
8:57:04some sometimes useful in projects and
8:57:06that's the reason it's it's a very good
8:57:07interview question as well because a lot
8:57:09of time in the projects you want to you
8:57:11don't want to use the default size you
8:57:13want to change some other to some other
8:57:16number so in that case you will be using
8:57:18this because one way is either you
8:57:20change everything from your
8:57:22configuration FES which is not a good
8:57:24idea to do so better thing is
8:57:25programmatically you deal with it and
8:57:27here you can change it by the usage of
8:57:30DFS do block size okay so it's not block
8:57:34underscore size con I hope you got your
8:57:36answer what's the mistake you were doing
8:57:38it should be block
8:57:40size but you are
8:57:43close now what is a block scanner in
8:57:48hdfs can I get this answer this is a
8:57:51usually a question in your Hadoop
8:57:53Administration this this is basically
8:57:55what your Hadoop administrators do so
8:57:57people who have who have done this Hado
8:57:59Administration classes can you answer
8:58:01this I'm expecting this answer basically
8:58:03from you even others can answer what is
8:58:06a block scanner in
8:58:09hdfs what is a block scanner in hdfs can
8:58:13I get answer nobody's answering this who
8:58:16all have done Administration course or
8:58:19no administrator Hadoop Administration
8:58:21can I get answers who all have done I'm
8:58:23not asking you to answer me this part
8:58:25just to answer me who all have done this
8:58:26Hadoop administrator course initially
8:58:29few people mentioned it that we I have
8:58:31done this administrator course you must
8:58:33have read about block
8:58:35scanner okay let me answer this part
8:58:38usually in a block scanner okay both is
8:58:41answering now to check if the block has
8:58:44any empty space left in the block uh
8:58:48okay one of the answer one of the answer
8:58:50I can take but not exact answer it's not
8:58:52very good uh scan the block and Report
8:58:55the remaining spaces okay okay again I
8:58:59can take partially this
8:59:01answer not just declaiming space but in
8:59:05it ensure the Integrity of your data
8:59:10blocks okay it basically keep on
8:59:13reporting every data Note will keep on
8:59:15reporting to the name one and it will
8:59:17keep on checking the Integrity of the
8:59:20data block let's say if any data block
8:59:21for got kind of corrected right or maybe
8:59:24the replic replica value become low
8:59:27right all those things it keeps on
8:59:29monitoring and try to
8:59:32rectify okay so it will basically keep
8:59:34informing the name but that's the reason
8:59:36this is been usely done by
8:59:37administrators because they keep up
8:59:39monitoring the health of the data no
8:59:42data block name Lo they're also
8:59:44responsible for this work right so this
8:59:46is what they keep on doing in order to
8:59:48make sure do they use block SC to do
8:59:51that okay there's one more way to check
8:59:54the replication Factor anybody know what
8:59:58is
8:59:58that there is one more way to check the
9:00:01replication Factor so this is about
9:00:03block St but there is one more way to
9:00:05check the replication fact environmental
9:00:07F no hardbeat no hardbeat will just tell
9:00:10that data is good or not data node right
9:00:13I'm talking about let's say some file
9:00:15got under replicated in that case how
9:00:18who will kind of inform name not let's
9:00:21say blog scanner is not there there is
9:00:24something called as had load
9:00:28balancer I'm not sure if you have read
9:00:30about that I do load balancer that
9:00:33basically ensures that if if your data
9:00:37blocks are not up right if they're under
9:00:39replicated or not that basically informs
9:00:42that okay this is under replicated let
9:00:44me take the off okay so this is
9:00:47basically the way also to check the
9:00:49under replicated or over replicated
9:00:52blocks can multiple clients write into
9:00:56an hdfs file
9:00:58concurrently can I get this answer if
9:01:01somebody ask you this question can
9:01:03multiple clients write into an hdfs file
9:01:07concr interesting I'm getting one yes
9:01:10one no now two yes one no okay two no
9:01:14two yes lot of yes okay do you think it
9:01:18should be fible to write multiple right
9:01:21I'm not saying reading I'm saying
9:01:24writing notice this part now can you
9:01:28answer don't you think it it it will
9:01:30make my file inconsistent yes single
9:01:33file I'm talking about don't you think
9:01:35it will make my file inconsistent if I
9:01:38do that right it will make my
9:01:41inconsistent so it is not allowed
9:01:44basically it allows only one writing and
9:01:48multiple reading stuff so that's the
9:01:51reason single file it will not allow you
9:01:54to basically keep on writing by at the
9:01:56same time by multiple Cent it will not
9:01:58allow you to basically do that for
9:02:00multiple client at the same time once
9:02:02one client is writing it will be kind of
9:02:04file will be kind of block for other
9:02:06client once the client have written
9:02:09after that only the other clients can
9:02:11right but everybody can read read coner
9:02:15that is one thing which is very
9:02:17important in so writing at the same time
9:02:20is not possible concurrently but reading
9:02:23is possible that's how hdfs is basically
9:02:27created why they have not allowed
9:02:29multiple rights together at the same
9:02:31time because if they do it can make the
9:02:33file inconsistent that will be a big
9:02:36trouble and that's the reason they do
9:02:39not allow you to do for current write
9:02:41because this is distributed system if
9:02:44multiple Cent will write on the same
9:02:45file now maybe somebody can overwrite my
9:02:48change right so that's the reason they
9:02:50will not allow you to do it okay this is
9:02:52by
9:02:54architecture another question what do
9:02:57you mean by high availability of name
9:03:01node and how it is achieved can I get
9:03:05this
9:03:05answer how this is achieved and what do
9:03:09you understand by High availability of
9:03:12the namee in fact I already answered
9:03:14this in Boss example you can answer me
9:03:17this part not
9:03:19replication R awareness name t of passes
9:03:22they be stand by so okay few people are
9:03:25giving right answer now active and
9:03:28passive name what basically happens in
9:03:31this cases so let's say if there are two
9:03:35let me show you the slide itself we have
9:03:36drawn it properly see this there will be
9:03:39two name one will be active name note
9:03:43and one will be passive name note so
9:03:45what happens is let's say this is my
9:03:47active name note which is running okay
9:03:50and what these data notes are there for
9:03:53reporting to the active now we also
9:03:56create a passive name node now this
9:03:58passive name node also these data noes
9:04:01will be reporting this passive name node
9:04:03will not be doing anything but it will
9:04:05just keep on collecting the data from
9:04:07your data nodes okay that is what the
9:04:10role of passive name note would be now
9:04:14as soon as this because this now is
9:04:17reading right so it knows the status of
9:04:19data node it knows that where the blocks
9:04:21are being written everything it has the
9:04:23information now suddenly if this machine
9:04:26is down in that case my passive name
9:04:29note will ensure it will immediately
9:04:32start acting like a backup and that is
9:04:34how it is ensuring High availability you
9:04:38are not going to lose your cluster time
9:04:41so basically the down time will not be
9:04:43there immediately your passive name will
9:04:45start acting like your
9:04:47active Okay so this is how basically
9:04:50Theo is ensuring High availability this
9:04:54is a very famous interview question
9:04:57passive and secondary name note good
9:04:59question difference between passive name
9:05:02note and secondary name note secondary
9:05:05name note we used to use in hadu 1.x now
9:05:09secondary name note what used to to
9:05:11happen was secondary n note you can say
9:05:13it's just like a snapshot of this
9:05:16machine means you're just copying the
9:05:19data copying the data to other machine
9:05:21but if my active name note is down if my
9:05:24main name note is down in that case my
9:05:27secondary name mode will not start
9:05:29acting like a backup it will only keep
9:05:32the data but it will not start acting
9:05:34immediately like a name not it will just
9:05:36keep the data you have to physically
9:05:39manually kind of bring name not up copy
9:05:42the data from the secondary name to
9:05:44primary name not and then start working
9:05:46on it but in passive namee yes so it's
9:05:50kind of an human intervention right
9:05:53manual intervention is required here and
9:05:55there will be a downtime there will get
9:05:57downtime here but when it comes to the
9:06:00active and passive name Lo passive name
9:06:03Lo is going to ensure that it is not
9:06:06only collecting the data metadata but it
9:06:10as soon as active name node is down
9:06:12taing name start acting like of active
9:06:15name clear on this difference now no an
9:06:20you have asked this question right clear
9:06:22about this question answer an the
9:06:24difference should be very clear to you
9:06:27wait so let's move
9:06:30further okay so we have few more
9:06:33questions now uh so this is for map Ru
9:06:35side so let's do one thing friends let's
9:06:37take a five minutes break okay let's
9:06:41take a five minute break and we will
9:06:43start with math produce questions okay
9:06:46uh I'm not sure maybe at Guys somebody
9:06:48will answer you just need a sip of water
9:06:50or so just give me five minutes it will
9:06:52take okay you want 10 minutes this is
9:06:55just basically a two to three hour
9:06:57session so we don't want to make it too
9:06:59much let's make it a seven 7 to 8
9:07:01minutes will that work let's come up
9:07:04with a in Middle kind of solution so
9:07:07let's come back by Maybe by 10 three
9:07:10okay we will
9:07:12start have a water break also we will
9:07:14come back to mauce then we have five we
9:07:17have scoop that so let's come everybody
9:07:20please be back by 103 exactly I'm going
9:07:22to start by
9:07:28103 okay so guys everybody is back now
9:07:32everybody
9:07:33back n you should be able to hear me now
9:07:37okay looks like he's starting okay fine
9:07:40so um now we are going to start uh so
9:07:43basically now we are going to start with
9:07:47math produce topic okay now in yeah I'm
9:07:50not shared I'm not shared I'm just going
9:07:51to share that just give me a second it
9:07:54takes almost few seconds to basically
9:07:56this now I hope it should be showing to
9:07:59you okay meanwhile I have talked with
9:08:03Eda team and they have informed me that
9:08:05you all will be receiving this video
9:08:07recording in a day or two okay since
9:08:10that was a question from lot of people
9:08:11they'll be getting all this video
9:08:13recordings or not so you will be getting
9:08:15this video recording on your email ID in
9:08:17a day or two so all these things will be
9:08:20there with you so I think that will be a
9:08:23quick review for you if you want to take
9:08:25a look at any moment now question for
9:08:27you can you explain me the process of
9:08:31spilling in MA
9:08:34RS this is an interesting question in
9:08:37fact think from a perspective of where
9:08:41mapper keeps the output and from that
9:08:43you can basically make out what is this
9:08:45sping I give you a biggest hint possible
9:08:48here so can you give me this
9:08:50answer can you explain the process of
9:08:54spilling in map
9:08:57reduce spills to Temp folder lfs of when
9:09:02it spill and from where it
9:09:05spill can I get this
9:09:07answer map per F very good
9:09:11what usually happens is the output of
9:09:15your mapper task it goes to your R now
9:09:19what basically going to happen they have
9:09:21kept a specific size of that thing so
9:09:25let's say they keep let's say this 100
9:09:28MB now 100 MB of data will be kept let's
9:09:32say in gra but they have they will be
9:09:34keeping a press so it will slowly keep
9:09:36on filling up slowly keep on filling up
9:09:39then what will happen as soon as it will
9:09:41reach a threshold let's say 80% of that
9:09:43Ram of 100 MB is spit it will start
9:09:46spilling that output to the local disk
9:09:50notice here I'm not saying
9:09:53hdfs I am saying local disk okay to
9:09:57local disk only we will be keeping this
9:10:00data local disk means your C drive D
9:10:03drive wherever you want to keep up so
9:10:05this is how they have designed it so as
9:10:08soon as the mapper out putut in the
9:10:11memory reach to a threshold limit it
9:10:14starts filling that mapper data to your
9:10:18local dis and this phas is called as
9:10:22filling phas in Mist okay so this
9:10:27question is also asked and this shows
9:10:29basically the internal working of your
9:10:31math prod okay so basically this is how
9:10:34internally your map produce work so this
9:10:39is all this is filling the data once
9:10:40it's filling it will again come down and
9:10:42you can see more and more data in it
9:10:47which links me to other question can you
9:10:50explain me the difference between block
9:10:53input splits and record this is a very
9:10:57famous interview question can anybody
9:10:59answer me Qui me what is the difference
9:11:02between blocks input split and
9:11:08Records what is the difference between
9:11:10the
9:11:11three friends this is a very important
9:11:14question if you have done this maass
9:11:15reduce part then you must be knowing
9:11:17this
9:11:19part difference between blocks input
9:11:23split and Records very good ji so J is
9:11:26answering record is a single line of
9:11:30data right very good Canan is saying
9:11:32block is hard cut of data just 128 MB
9:11:36very good right block is equal to 128 MB
9:11:40if record is 130 MB input split will
9:11:43happen very good block is based on block
9:11:46size input split makes sure that the
9:11:48line is not broken so it makes sense
9:11:51record is single line block is set by
9:11:54sdfs input split is logically spit very
9:11:57good very good so what usually happens
9:12:00right so let's say when we talk about
9:12:02blw so let's say you have default space
9:12:04is 128 and so that will be called as a
9:12:07physical block okay when we talk about
9:12:10input split right so let's say if your
9:12:12data is of 130 M now in that case don't
9:12:15you think it makes sense to to have a
9:12:18logical spit of 130 MB here so that will
9:12:20be your input spit and record is when
9:12:24you do Mac produce programming right
9:12:26when you do maap produce programming how
9:12:28your mapper take the data it takes line
9:12:31by line right it takes line by line that
9:12:34line is called record so one line of
9:12:38data which it picks up in the map face
9:12:41is called your record okay very very
9:12:45famous question on this part so you can
9:12:48say block is a physical division by
9:12:50logical division are called your input
9:12:52splits and Records okay because the
9:12:55logical division is what your map
9:12:57produce program do which brings me to
9:13:00another question again relate to map
9:13:02produce what is the role of record
9:13:06reader in Hado map prod
9:13:11what is the role of record reader in
9:13:15hard do map
9:13:17ruce make sure to read the complete
9:13:20record no no how that is how Mapp reads
9:13:25a record good but dber can you give be
9:13:28little more exp you coming close to the
9:13:31answer can I get more answer also D can
9:13:34you just be a little more
9:13:35explicit that is how map read and record
9:13:39we coming close
9:13:41what about
9:13:43others what is record record is single
9:13:47line we have just understood it so what
9:13:50should be record reader paring that
9:13:52single line very good word right so
9:13:55don't you think that single WR what you
9:13:57are reading and how mapper convert your
9:14:00data it converts into key value PA right
9:14:03so it initially takes an input as a key
9:14:05value pair so when that conversion is
9:14:08happening that is done basically by
9:14:10record reader look at this see let's say
9:14:13this is the data it will be getting
9:14:15converted to key values here where key
9:14:17is called your offset and value is first
9:14:20line right or second line or third line
9:14:23right so this is done by record reader
9:14:27okay so this is what record reader do
9:14:31now what is the significance of counters
9:14:34in maap
9:14:36is significance of counters in map
9:14:44is uh okay give statistics of data
9:14:48counters will be done in name not okay
9:14:51counters to validate the data R to
9:14:54calculate Bard good not just bad record
9:14:58people it can do even other things right
9:15:00this bad record is just one example
9:15:02right it just one example you can say so
9:15:06what it do is it helps you to identify
9:15:09the statistics right now basically it
9:15:12gives you the statistics about some
9:15:14operation what you want to do you can
9:15:16print it in the console also I I believe
9:15:19that if you have done this Hado course
9:15:20you must have seen one code for your
9:15:22counters right where you must be doing
9:15:25some sort of operations and you you
9:15:27might be printing it on your console
9:15:29window now how you will be doing all
9:15:32that so what you will do let's say we
9:15:34are applying counters on this example
9:15:36and in this example as I think somebody
9:15:39just mentioned for the kind of reading
9:15:41the bad data this example actually bring
9:15:43up the same thing but it can do even
9:15:46other things I will tell you what other
9:15:48things no but let's take this example
9:15:50let's say we want to find out what all
9:15:52bad data I have in this so let's say it
9:15:55is reading it is reading David all good
9:15:57no problem counter will remain as zero
9:16:00then what it did it basically just gave
9:16:02the value as zero it reach to the second
9:16:04line now this value move now basically
9:16:07this was the back data I read it I
9:16:09passed the statistics Now counter value
9:16:11became one it reach to Jeff if the value
9:16:14stays as one because this is a good data
9:16:17now Sean again the value Remains the
9:16:19Same now as soon as it reach the last
9:16:22line of that data it will again increase
9:16:26this counter and will return that I have
9:16:29two bad line of course but does that
9:16:32mean we can only do this operation on
9:16:35this no maybe I can have an example
9:16:38where let's say I have data of which is
9:16:41let's say of time stamp time stamps are
9:16:44there okay I want to print in this time
9:16:46so time stamp I can convert to date type
9:16:49right in my program I can convert it to
9:16:51date type now when I convert to date
9:16:53type maybe I have months being defined
9:16:55right because in in date I have months
9:16:58now maybe I want to find out I want to
9:17:01calculate the statistics that in among
9:17:04this time stamp how many times January
9:17:06is offering how many times February is
9:17:09offering in how many times March April
9:17:12and all are occuring let's say I want to
9:17:14identify all the statistics I can do
9:17:17that with the help of counters very
9:17:20easily okay so this is the purpose of
9:17:24your
9:17:25counters now moving
9:17:28further why the output of math pass or
9:17:32spilled into the local disk and not in
9:17:36hdss good question now can you answer me
9:17:38this remember we talked about uh we
9:17:42talked about yes we have just discussed
9:17:44that question right so we have just seen
9:17:46spilling we have seen spilling
9:17:49now how can we assess counters there are
9:17:52basically libraries available for that
9:17:53there are classes available get counters
9:17:56the is the member function to get that
9:17:59this is how basically will be assessing
9:18:00it so counter is the class in which we
9:18:02have member function predefined
9:18:04functions using that you can assess
9:18:07that because it's an intermediate output
9:18:11okay intermediate output that's fine but
9:18:14why we are not keeping in sdfs that's
9:18:16that's my question when I can keep
9:18:18intermediate output in
9:18:20sdfs very good very good
9:18:23D very good nimma so basically if you
9:18:26notice what happens if you keep in hdf
9:18:29right remember there's a replication
9:18:31Factor right that replication Factor
9:18:33will do what it will increase the number
9:18:36of output blocks do you think it makes
9:18:38sense to increase the replication for
9:18:40your Mapp output definitely a big no
9:18:43right so that's the reason we will be
9:18:45keeping in your local F system we will
9:18:47not keeping sdfs otherwise sdfs will
9:18:50replicate even mapper output which we
9:18:52don't want to happen to do so that's the
9:18:54reason we will stop that so we will be
9:18:56keeping in the local dis not in hdfs in
9:19:01order to avoid basically the
9:19:04replication which brings me to another
9:19:07question can you define this speculative
9:19:11execution can you define this
9:19:14speculative
9:19:17execution can I get this answer can be
9:19:20defined speculative
9:19:25execution it prioritize only some task
9:19:28uh coming close but not exactly
9:19:31right if a job in a note is taking much
9:19:34time very good very good people so this
9:19:38is what happens in speculative execution
9:19:42let's say if any of your task is running
9:19:46very slow in that case your Speculator
9:19:49there will be a scheduler which will
9:19:51basically start a duplicate task of it
9:19:55it will start running a duplicated task
9:19:58for it to ensure that basically that
9:20:01duplicate task run faster and once it
9:20:03will finish it will kill all the
9:20:05duplicate task so it is just kind of
9:20:08making sure that because because it can
9:20:10happen right maybe your task is waiting
9:20:11due to some resource it got blocked due
9:20:13to any reason so in that case it will
9:20:17start immediately a duplicate task and
9:20:21making sure that your job finished
9:20:24quickly okay so this is the part of your
9:20:28speculative
9:20:31execution which brings me another
9:20:33question question is how will you
9:20:36prevent y very good very good how will
9:20:40you prevent a file from splitting in
9:20:43case you want the whole file to be
9:20:46processed by same
9:20:49Ma how you will prevent a file from
9:20:53splitting in case you want the whole
9:20:56file to be processed by same mapper I
9:21:00want my file to be now basically to be
9:21:02used by the same m not combin not combin
9:21:06theb can I get some more answers it's
9:21:09easy
9:21:10can you see this side so what we can do
9:21:14here is first of all we can increase the
9:21:18minimum number of split size which
9:21:21should now make it larger than your last
9:21:24five this operation is itself good
9:21:27enough to make this case look right
9:21:30because if you increase the size itself
9:21:31we would be all good but there is one
9:21:34more thing which you can do after that
9:21:36is this is Method two basically in
9:21:39method one you can just basically
9:21:41increase the size itself it will be all
9:21:43good or what you can do you can go to
9:21:45your input format CL and in that you can
9:21:49just update this property you can make
9:21:52this is splitable to be returning first
9:21:55usually people prefer method one because
9:21:57that's the easiest method right you need
9:21:59not update the Java code basically to
9:22:01achieve all this so what you will do for
9:22:03this can you please tell a scenario
9:22:06where file splitting is not needed it
9:22:08all depends right so let's say if I know
9:22:11that I have only one data node I mean
9:22:13let's say I have only one data node now
9:22:15in that one data node do you think it
9:22:16makes sense to divide a file multiply
9:22:18together I have only one data Note right
9:22:21do you think it makes sense because
9:22:23again it's in the same machine right so
9:22:25that's the reason there you want to
9:22:27basically keep one block it so that I
9:22:29can execute it so this is these can be
9:22:31few situations where you can decide not
9:22:34to split the file and in that case how
9:22:37you will be doing it these are the two
9:22:39method to achieve okay moving
9:22:43further is it legal to set the number of
9:22:46reducer task to zero that's question
9:22:48number one where the output will be
9:22:51stored in this case is it legal to set
9:22:54the number of reducer task to zero is it
9:22:59legal definitely legal right everybody
9:23:03have your school in school was there any
9:23:06reducer was there any reducer in school
9:23:10no scoop only use mapper no reducer
9:23:15right that's itself a tool right when a
9:23:17tool is not even creating any reducer
9:23:20definitely I will not be using that
9:23:22right so what is going to happen here is
9:23:25let say so what what was the purpose of
9:23:27reducer purpose of reducer is when you
9:23:31want to do some sort of aggregation
9:23:33right maybe you want to Summit in the
9:23:35end and all those category kind of
9:23:37things but it's not Mand that all the
9:23:40problem statement in the word require
9:23:43aggregation right so those problems
9:23:45where you do not require agregation like
9:23:48I said scoop because in scoop what
9:23:50happens you copy the data from rdbms to
9:23:53your hdfs or vice versa now are you
9:23:55doing any sort of aggregation no right
9:23:58you're just copying the file from rdbms
9:24:00to your hdfs system no agregation
9:24:03require so in those cases you will be
9:24:06having reducer as zero output where to
9:24:10store definitely whatever mapper output
9:24:12is coming wherever you're telling it
9:24:14will be stored in that sdfs location
9:24:16right so that how you will be using it
9:24:19so definitely the answer is yes and
9:24:21basically wherever the mapper output you
9:24:23seeing there it will be getting SC what
9:24:27is the role of application master in map
9:24:31reduced job can I get this answer this
9:24:34is a very famous interview question what
9:24:37is the role of app ation master in map
9:24:41reduce job to assign the task okay
9:24:46that's it only to assign the task sets
9:24:49the input split okay it manages the
9:24:53application fired and keep track of sub
9:24:55process that's it nothing else to get
9:24:59the resources needed for the task very
9:25:01good now you are coming here we are
9:25:04coming close to create task until until
9:25:07it yes what basically application Master
9:25:11do first thing is application Master is
9:25:15kind of deciding that how many resources
9:25:19it needs okay and it can basically now
9:25:23inform resource manager that I require
9:25:26this many resource and give me this many
9:25:29resource to execute that then after that
9:25:31resource manager give back the container
9:25:33right if you might have gone through
9:25:35your Yan architecture right there you
9:25:37might have understood all this potions
9:25:39right so basically what happens it first
9:25:42basically find out how many resources
9:25:44are required secondly what it want to do
9:25:48is it basically want to find out once it
9:25:51happens when it gets all the container
9:25:53it kind of basically get them working
9:25:56together it collect the output back and
9:25:59return it to basically the master so it
9:26:01is doing multiple things in MA it's not
9:26:04it's basically not doing just one task
9:26:07it is also checking which is a very
9:26:09important role of it it is checking that
9:26:12how many resources I require and kind of
9:26:15helping the resource manager to take
9:26:18that decision okay so this is what the
9:26:21same thing being explained here the role
9:26:24of your application Bas this brings me
9:26:28to another question what do you mean by
9:26:32overb so when your map ruce job runs I
9:26:35don't I'm not sure whether you have
9:26:36noticed that or not there is something
9:26:39called as Uber mode it it's sometime
9:26:42comes in your conso if you have noticed
9:26:45that so can you tell me what is that
9:26:47Uber mode is there any advantage of
9:26:51searching on the Uber mode what
9:26:53basically Uber mode is going to do runs
9:26:56on application Master Mod very good can
9:26:58I get some more answers more insight on
9:27:01this good can I get some more answers
9:27:05yes so let's say if you have a small job
9:27:09right if you have a small job in that
9:27:12case you you're basically again
9:27:15application Master need to be any way up
9:27:17now application Master will request and
9:27:20then it will allocate container right so
9:27:22container sending and all basically
9:27:24creating container all those things are
9:27:26time consuming if the jobs are small
9:27:30what your uh your application Master can
9:27:33do application Master can start a jvm in
9:27:37itself okay so basically application
9:27:40Master can decide to complete your job
9:27:43because the job is small it may decide
9:27:45to complete the job on its own in that
9:27:48case we call it as Uber mode so Uber
9:27:51mode basically require less it's when
9:27:54you you will be using let's say less
9:27:56number of Ms only 10 Maps you have only
9:27:59one reducer to work on so in those cases
9:28:02you use Uber mode now how to enable all
9:28:05that so basically there's a property
9:28:07which you can set to true and it will
9:28:10basically enable your over mode in which
9:28:13what's going to happen your application
9:28:16Master will start acting like a JM and
9:28:19will finish the job so does it will save
9:28:21you to buy creating container making
9:28:24your performance better so usually what
9:28:26people do is whenever they will be
9:28:29having some sort of uh whenever they
9:28:31will be having small jobs right and they
9:28:33want to improve the performance they
9:28:35usually kind of enable the Uber mode so
9:28:39now the application Master itself start
9:28:41executing the task and that way they
9:28:43improve the performance but when you
9:28:45have a bigger job it will not work in
9:28:48that case you need to keep the over word
9:28:49as false otherwise you will degrade your
9:28:52performance it will not even work okay
9:28:54so this is basically your Uber
9:28:58mode another question how you will
9:29:01enhance the performance of M prod jobs
9:29:05when dealing with too many small FES if
9:29:09you have let's say many many small files
9:29:11in that case very good report yeah
9:29:14that's also one thing if you have let's
9:29:16say many many small files how you can
9:29:19improve the performance of your map Ru
9:29:22job The Miner no coming close but not
9:29:26the right answer too many small files
9:29:29are there then what you will do uh
9:29:32basically there is something called as
9:29:35because none of you give the very right
9:29:37answer there is something something
9:29:39called as combined file input format
9:29:42that's the reason I told you was you are
9:29:43coming close but not exactly coming
9:29:45right right so what this do is it kind
9:29:48of packets all the small files together
9:29:51see this diagram see this part like
9:29:54these are some small files it basically
9:29:56now combined all the files together now
9:29:59because these small files got combined
9:30:02together now my execution time will be
9:30:04FAS this is basically one practical
9:30:07thing which which which is is being
9:30:09represented this performance can you see
9:30:11like the small price was taking this
9:30:13much of time and basically with this
9:30:15when you combine this it actually
9:30:17started taking less time so that's
9:30:19improving the performance of your system
9:30:23so this is what this this is how
9:30:26basically you can improve the
9:30:27performance if you have multiple small
9:30:32FES now let's move to hi hi is very
9:30:36important topic now question for you
9:30:40where the data of high stable J
9:30:44St I know that is going to be in sdfs
9:30:48but where is the location of that where
9:30:50is the location for that but what is
9:30:53that default folder that version very
9:30:55good now you are coming close so by
9:30:58default you keep it in slash user slash
9:31:03hi/ Warehouse so this is the location
9:31:07where by by default all your hi table
9:31:11get store if you want to change this you
9:31:14can go to your high. XML and can update
9:31:17the setting as well another question why
9:31:21hdfs is not used by high meta store for
9:31:27storage what I mean is you might have
9:31:30read in your course that you keep your
9:31:34hi basically your metal store in your
9:31:37rtpms right not in
9:31:39hdfs what is the reason behind that why
9:31:44we not keep all these things in my hdfs
9:31:49why your met store is created in your
9:31:53rdbms why you config configuring your
9:31:56meta store in nbms for he why not in
9:32:00hdfs can I get this
9:32:03answer why you are keeping your met
9:32:06store in your rdb system system and not
9:32:09in your hdfs I can tell you this is the
9:32:12most important question of he and any
9:32:16interview of P you will go expect this
9:32:20question this usually everyone ask for
9:32:24random SS but that you can do in sdfs
9:32:27also if you need DB catalog what is that
9:32:31catalog because it need to use jdbc
9:32:34connection no no that's not the
9:32:36answer okay let me ask you this question
9:32:39okay which is actually the main reason
9:32:41for it uh so let's say if uh what you do
9:32:45is when you create a table in height
9:32:48what happens it creates an entry for
9:32:51that table in metas store table right
9:32:54what it's doing it's inserting the roow
9:32:57level right at a row level it is
9:32:59inserted you created another table in
9:33:02again you are basically inserting some
9:33:04values here right now let's say you
9:33:06deleted some table in hand what happened
9:33:09it did a ro level delete can you do this
9:33:12Ro level delete Ro level insertion and
9:33:14all in your
9:33:16hdfs okay you're saying referring to the
9:33:18same thing that is then you're good can
9:33:21you do this basically this in your sdss
9:33:23no right this itself is a good answer to
9:33:26explain this right so and second thing
9:33:29is definetely inm the thinking time is
9:33:31going to be faster and all those are the
9:33:33other facts but first thing is basically
9:33:35your cred operation cannot be done in
9:33:38CFS that's the reason we will not be
9:33:41able to keep in the CFS forget about
9:33:44other factors right they do not make any
9:33:47sense in fact because my first property
9:33:50itself is failing of basically current
9:33:52operation now let's see some scenario
9:33:56questions usually in high Pig you will
9:33:58find some scenario questions coming up
9:34:00now scenario question is suppose I have
9:34:04installed a Pache Hive on top of my Hado
9:34:09can you please show the last answer sure
9:34:11why not see this answer in the here
9:34:16let's move forward now can can you see
9:34:19this question now suppose I have
9:34:22installed AE Hy on top of my Hado
9:34:25cluster using default meta store
9:34:29configuration then will what will happen
9:34:33if we have multiple clients trying to
9:34:36assess High add same time can I get this
9:34:41suppose I have installed AE hi on top of
9:34:45my hadu cluster by using default
9:34:48metastore configuration then what will
9:34:50happen if we have multiple client trying
9:34:53to assess H at the same time very good
9:34:57usually in height you can only basically
9:35:02assess one by one client so multiple
9:35:05client sess itself is not allowed right
9:35:10very good they given basically the right
9:35:12answer for it so usually this these are
9:35:14the scenarios what it follows so main
9:35:16thing is your multiple client sess in hi
9:35:21is not allowed this is by architecture
9:35:25right because you should maintain the re
9:35:27consistency very good addition right so
9:35:30that's the reason this itself is not
9:35:33going to work out so basically that is
9:35:36what you need to keep in mind what is
9:35:39the difference between external table
9:35:43and manage table in fact manage table
9:35:46you also call it as internal table so
9:35:49can I get an answer what is the
9:35:51difference between external table and
9:35:54manage table can I get this
9:35:57answer external table can be a file okay
9:36:01but I want a difference proper
9:36:03difference external table where the hdfs
9:36:07file won't be deleted if you delete the
9:36:09table and you say sdfs file okay okay I
9:36:12can accept your answer external table is
9:36:16stored in separate location of our but
9:36:18even internal table I can store it at
9:36:20some other location by defining the
9:36:22location ke
9:36:24and external table keeps data when it
9:36:27get deleted very good that is the major
9:36:31factor when you talk about manage table
9:36:35what happens is if you have deleted any
9:36:38of the DAT table what is going to happen
9:36:41it will delete the entry in your metas
9:36:45store at the same time it is also going
9:36:48to delete the data file but in external
9:36:52table if you delete a table it is going
9:36:55to only delete the entry in your meta
9:36:59store not from your main data okay so
9:37:04that is the major difference basically
9:37:07in the minus table and external
9:37:10table another question when should we
9:37:14use sort by instead of order by if you
9:37:18notice these two apis belongs to five
9:37:21and these two basically going to do
9:37:23exactly same thing so when should I use
9:37:26sort by and not order by operation when
9:37:31should I use sort by and not order by
9:37:35operation so basically to answer this
9:37:39not able to catch things uh I did not
9:37:41get this R when we should have only one
9:37:46ma okay R I think you are in fifth
9:37:48module right so that's the reason uh
9:37:51yeah I can understand you have told me
9:37:53the starting itself right that you are
9:37:55still going through this course so if
9:37:57you're not getting it is completely fine
9:38:00just listen to this okay just listen to
9:38:02this conversation once you will go over
9:38:05these course topics in uh basically from
9:38:08wherever you are doing you will be all
9:38:10comfortable with it okay basically now I
9:38:13can understand if you will not get
9:38:14anything because these modules is not
9:38:16being taught to you yet so these are
9:38:18basically the new B which which will be
9:38:21taught later now can I get this answer
9:38:24uh in case of numericals no no that you
9:38:27can use also order by one of it use
9:38:30reducer other use mappers okay okay when
9:38:35you use Group by operation no no that
9:38:38turn way actually what happens is if you
9:38:43have huge data set in that case you
9:38:47should use basically the sort by option
9:38:52it usually do this sorting on multiple
9:38:55reducer while order by do it on one
9:38:59reducer that is basically the major
9:39:02difference so when you have huge data
9:39:05set use sort by by instead of order by
9:39:11okay lot of people remain confused with
9:39:13this that's why this is a very tricky
9:39:15question what people are usually if you
9:39:17ask anyone right if you do not go the
9:39:19answer he will tell you both do the same
9:39:21thing but actually that's not the there
9:39:23is a difference now another question
9:39:26what's the difference between partition
9:39:28and bucket in I think the most easiest
9:39:30question to answer can everybody answer
9:39:32this whoever have done on hi topic
9:39:35what's the difference between partition
9:39:37and bucket
9:39:38this is the most easiest
9:39:40answer can I get this answer difference
9:39:43between partition and bucket simple
9:39:45right partition is basically at the
9:39:48first level right when you split the
9:39:50data into different directory Buffet is
9:39:53like a subpartition of that right so
9:39:56even for that partition itself when you
9:39:58create another subpartitions you can
9:40:00call them as bucket like in this case
9:40:03can you see the first partition is PC
9:40:05Department Civil Department electri
9:40:07electrical department but after that we
9:40:09have also created some subpartitions of
9:40:12it and that is your bucket that's
9:40:15basically the differen another question
9:40:18let's say this is the scenario you are
9:40:21creating a transaction table now this is
9:40:25the table what you have like transation
9:40:27table is the table you have this many
9:40:29columns delimited field by comma now
9:40:33let's say you have inserted 50,000
9:40:35couples in this table now I want to know
9:40:39the total revenue generated for each
9:40:41month but he is taking too much time in
9:40:46processing this query can you tell me
9:40:49what the solution you are going to
9:40:51provide this scenario is actually a very
9:40:55good interview
9:40:56scenario very
9:40:59good can I get more answers very good
9:41:02very good can I get more answers you
9:41:05will be partitioning this table how you
9:41:08will be partitioning you will be
9:41:11partitioning your table with month right
9:41:15so basically if you partition your table
9:41:18you will improve your performance so
9:41:21these are the simple steps you can
9:41:23create a table Partition by month set
9:41:26these properties to truth so that you
9:41:27can enable your partition insert the
9:41:30data and then you can retrieve the data
9:41:33where your month is going to January so
9:41:36while parage after partitioning the
9:41:38table you can improve the performance
9:41:41second can I get an answer of this what
9:41:44is dynamic partitioning and when is it
9:41:48used can I get this answer what is
9:41:51dynamic partitioning and when is it
9:41:55used that can be static partitioning
9:41:57also in right so I want to know what is
9:42:00dynamic partitioning very good partition
9:42:04happens when loading the data into table
9:42:08right now I don't know that if I do a
9:42:11dynamic partitioning where it is which
9:42:13how many partitions also it is going to
9:42:15play so the value of your partition
9:42:18columns will be known only during your
9:42:22run time when you will be creating the
9:42:24partition that is called your Dynamic
9:42:29partitioning okay how high distribute
9:42:33the row into bucket can I get this
9:42:36answer how how hi distributes the rules
9:42:40into bucket very good hash algorithm
9:42:45okay it uses the hash algorithm to
9:42:49understand this part if you look what we
9:42:51are doing here is now no you will be
9:42:54using clustered by but basically in how
9:42:57internally this is that's basically how
9:42:59it is let's say you want to put into two
9:43:01bucket in that case it is going to do
9:43:03modul of okay let's say mod of to this
9:43:06output table came out to be one so it
9:43:08will decide to put in bucket one modul
9:43:10two it is going to become zero it is
9:43:12going to keep in this second bucket so
9:43:15this is how it will be decided okay it
9:43:17will be using basically a hash
9:43:20computation of this so basically it will
9:43:22be using hash function of this so let's
9:43:24say hash value of this value came out to
9:43:26be one hash function of this value came
9:43:28out to be two hash function value this
9:43:30came out to be three then basically it
9:43:32is doing this modular operation and
9:43:35giving the output and that's how will
9:43:37distribute the buting data now which
9:43:42brings another question suppose I have a
9:43:45CSV file which is named as sample. CSV
9:43:50present in tm1 directory with the
9:43:54following entries in that case how you
9:43:58will consume this CV file into where hi
9:44:02Warehouse using buildin surday sday mean
9:44:08ization s des that's serialization der
9:44:12serialization when you convert your data
9:44:15into kind of kind of
9:44:17bbes can I get this answer row format
9:44:20delimited by comma not exactly I'm
9:44:24looking for something else actually this
9:44:26requires some API so let me show you
9:44:28this part see this answer in this case
9:44:32you will say raw format Sur or do Apache
9:44:37do. sur2 doop CSV
9:44:41Sur this is what you need to add okay
9:44:45now you got it what's the mistake we
9:44:47doing so basically this is what you need
9:44:49to add otherwise everything is same you
9:44:51can sa in the TMP folder and all that
9:44:54just that the only difference will come
9:44:56in row form at thir another question I
9:45:00have lot of small CSP files present in
9:45:04input directory as and I want to create
9:45:08a single table height table
9:45:11corresponding to these files the data in
9:45:14these files are in this format now as we
9:45:17know had do performance degrades when we
9:45:20use lot of small files so how you will
9:45:24solve this problem can anybody give me a
9:45:27simple answer of this this should be
9:45:28easy you have multiple small FS now in
9:45:31that case what what should be the
9:45:33solution because definitely my
9:45:35performance is going to degrade what
9:45:37should I do concatenate solve file but
9:45:40can there be another answer don't you
9:45:41think you can use sequence file here
9:45:43sequence file right if yeah one solution
9:45:47can be that is from hdfs Level itself
9:45:50right I'm talking from basically height
9:45:52perspective right don't you think I can
9:45:54convert in a sequence file sequence file
9:45:57will convert everything like 0 1 01 kind
9:45:59of thing so that's what make it better
9:46:02right will will improve my performance
9:46:04so first create a table load the data
9:46:07after that what you do store it as
9:46:11sequence file and then basically load
9:46:13this data from this file what you
9:46:15inserted to this file this will ensure
9:46:18now that your speed will be good why do
9:46:22we need to do in serice why do okay so
9:46:25basically when you do cice right the
9:46:28advantage what you get is so so let's
9:46:30say first thing is compression because
9:46:32you are serializing the data when you
9:46:34serialize the data it makes that trans
9:46:37were also very easy because we have
9:46:39converted like 01 01 01 kind of right so
9:46:43first thing is when you convert to sers
9:46:45you comess the data second thing since
9:46:47you convert to this 01 format now the
9:46:49transfer over the network become lot
9:46:51easier for me clear the that's the
9:46:54reason we basically use this so but
9:46:59remember there is no pre- lunch right
9:47:02there is no nothing called as pre- lunch
9:47:04don't you think this will also have a
9:47:05disadvantage when you you do a de
9:47:07realization again you need to convert it
9:47:10back don't you think it will impact the
9:47:11performance a bit right so that remember
9:47:14there is no fre lunch though it is
9:47:16helping you in this act but at the same
9:47:18time it will demand your performance it
9:47:20will leat up some of your performance
9:47:23now some quick
9:47:24questions can you give me this answer
9:47:26difference between logical and physical
9:47:29plan something difference between
9:47:32logical and physical F I know guys that
9:47:35you got little tired because this is
9:47:37big session but don't worry we we are
9:47:40almost getting done so I want everybody
9:47:42attention to be back now can you tell me
9:47:44the difference between logical and
9:47:46physical plan this must be the first
9:47:49thing what you must have learned in your
9:47:51hadum course when you went to pick
9:47:53topic whenever you are executing
9:47:57statement by statement okay it is just
9:48:00executing the statement nothing in that
9:48:03case first it creates a logical plan
9:48:06means let's say if there is no error
9:48:09right in that case it is just creating a
9:48:10logic but when you do dump right when
9:48:14you do dump then only the execution
9:48:16start right because of lazy valuation
9:48:19then your logical plan kind of get
9:48:21converted to like of physical plan means
9:48:24it start getting executed now let's say
9:48:27if you have given the wrong file part
9:48:30right in logical plan it will not give
9:48:33you any R because there is no syntax
9:48:36only at the time of physical plan it
9:48:38will give you an error So Physical plan
9:48:41is when you are basically executing your
9:48:43map job when this pig is getting
9:48:45converted to map job and by logical plan
9:48:48is at the initial
9:48:50level okay so that is what happen can
9:48:53you tell me what is bad what is
9:48:56bad collection of couples very good
9:48:59right so basically when you say a whole
9:49:02data file itself right collection of
9:49:04couples if you notice so let's say this
9:49:06is a data right if this is a data now if
9:49:09you see this is one tle this is second
9:49:12tle this is third tle so collection of
9:49:15all these things will be called as back
9:49:18okay collection of all these things will
9:49:20be called as back
9:49:22now how hi is only working with like hi
9:49:28is able to deal with only structure data
9:49:31but pig is able to deal with
9:49:33unstructured
9:49:34data pig is how Pig is able to deal with
9:49:37unstructured data can I get this answer
9:49:41how pig is able to deal with
9:49:43unstructured data it is actually
9:49:46happening because of schema less part
9:49:49right schema less part basically if you
9:49:52do not have schema if you do not have
9:49:54schema depend then also P can work what
9:49:57P do is let's say you you do not know
9:50:00like in hi you have column names right
9:50:02let's say column name is age integer all
9:50:06that right so you need to define the
9:50:08column names in P there is nothing like
9:50:10that so let's say if you do not know the
9:50:12column name there's no schema being
9:50:14defined you can Define it like this
9:50:16First Column you can Define by dollar
9:50:18one second column you can Define with
9:50:20dollar two so that's how you can also
9:50:22Define so basically Pig you can Define
9:50:25even your schema less thing okay so you
9:50:30don't have anything it will treat it
9:50:32like null it will start treating if you
9:50:34do not Define data type it will start
9:50:37get is buy so basically Pi kind of
9:50:39converts your values to other way so
9:50:43basically that's the reason you can
9:50:46physically go with pig with unstructured
9:50:50data this is one of the major reason
9:50:53that P can deal even unstructured data
9:50:56by height cannot do because P kind of
9:50:59converts the data or Tre that data inv
9:51:02in different way right like if you don't
9:51:04have column name you can Define as
9:51:07dollar two dollar one all those things
9:51:09as
9:51:10well what are the different execution
9:51:14mode available in pig so there are two
9:51:18modes right one is local mode one is map
9:51:23produce mod yeah sure can do
9:51:26that this is the same thing I mean like
9:51:29if you have no data treat it like B if
9:51:31you don't have column it start reading a
9:51:33dollar one dollar two right before okay
9:51:38this one back got it now great let's
9:51:41move
9:51:42further now so there are two modes
9:51:45available one is map ruce mode one is
9:51:48local mode so when you go with pig in
9:51:51map produce mode so when you just type
9:51:53Pig right it take you to the grun by
9:51:56default it take you to the mapm which
9:51:59basically also states that that
9:52:01basically if you're are going with back
9:52:04mode you are assessing your HD this
9:52:07while if you are using local mode what
9:52:09you need to do you need to go like this
9:52:11Pick hyen X local it will now take you
9:52:17as in the local mode when you say local
9:52:20mode what basically happens here it
9:52:22basically now start assessing the file
9:52:25from your local file system now it is no
9:52:27more assessing sdfs but it is directly
9:52:30assessing the data from your local file
9:52:33system these are the two execution mod
9:52:36platin this is very simple right so
9:52:38basically flatten is the keyword
9:52:40available if you have this kind of data
9:52:42you can flatten it up like everything
9:52:44will come together in the line so
9:52:46plattin is just in API right you can see
9:52:50this plattin so basically it is just
9:52:52basically converting this form of data
9:52:54to this form of data okay so this is
9:52:57basically meant by plattin these are
9:53:00simple questions now can anybody explain
9:53:04me this xB B
9:53:08components can anybody explain me these
9:53:11xbase
9:53:13components anyone who want to talk about
9:53:15it hbed components it's in front of you
9:53:20can anybody explain me these components
9:53:23of H
9:53:24base you can start by one by one this is
9:53:28the last topic so friends I want
9:53:30everybody to be attentive here so hbas
9:53:33keeps the data in a distributed mode
9:53:35right so distributed what where it keeps
9:53:38the data it defines a region where it
9:53:41keeps the data so like this will be your
9:53:44first region this will be like region
9:53:46where you're keeping let's say this
9:53:47column value row value right this is a
9:53:50one region not together just like how
9:53:52you define rack right you can Define
9:53:54region surface right where you're
9:53:57defining basically different different
9:53:59regions together so this is one region
9:54:00server this is one region server right
9:54:03now what happens the master will keep
9:54:05basically your will be called as H
9:54:08Master like active Master just like your
9:54:10name note works right similarly hbas
9:54:13uses this H Master what is this Z
9:54:16zookeeper doing here zookeeper is kind
9:54:18of helping you to execute everything so
9:54:21like in zoo what happens in zoo
9:54:23basically they keep animals right they
9:54:25keep animals they manage multiple
9:54:27different category of animals similarly
9:54:29do people like big data also got so many
9:54:31tools available for you now and
9:54:33basically it helps you to manage
9:54:35everything up so botkeeper maintains lot
9:54:38of things for hbas it kind of helps you
9:54:41to basically kind of see that
9:54:43consistency is maintained it like in XB
9:54:47right when you work basically your in
9:54:49your big data when you're working in
9:54:51Hado now you have name Note data Note
9:54:54all those things available but when
9:54:55you're working with XB you don't have
9:54:57all those things right so in XB you
9:54:59don't have concepts of name Lo your
9:55:01keeping data so for that to do this part
9:55:05like you keep play a major role here so
9:55:08zeper is kind of going to act like a
9:55:10coordinator inside your hbas environment
9:55:14okay it will help you to coordinate all
9:55:16the things because here we don't have
9:55:17name node data nodes and all so and
9:55:20there's no Yar basically here so you can
9:55:22treat it like basically just like how
9:55:24Yan was handling things there it is
9:55:25going to help you as a coordinator it
9:55:28also maintains the directory
9:55:30structur can anybody tell me what is
9:55:33Bloom
9:55:33filter Bloom filter
9:55:38anyone know what is Bloom filter in Bas
9:55:41can I get this answer what is Bloom
9:55:45filter basically it helps to improve the
9:55:49overall throughput of your fluster it
9:55:52helps you to basically improve the
9:55:54performance okay now if you want to
9:55:56search any specific row column cells it
9:55:59also help you to do that so it makes a
9:56:01system very fast so that's that's what
9:56:04you need to just enable this and if if
9:56:06it is enabled it kind of includes the
9:56:08toput of your cluster this is the role
9:56:11of your blue filter in X
9:56:15space coming to next question what is
9:56:18the role of jdbc driver in a scoop setup
9:56:24simple can I get this answer this is
9:56:27basically scoop are important topic in
9:56:29interview HB I would still say that
9:56:31they're not very important these this
9:56:33people do not ask questions onb but
9:56:35definitely this schol topic is very
9:56:39important very good rdbs database I want
9:56:43to connect now basically dbms can be of
9:56:45any type it can be my SQL it can be uh
9:56:49it can be my SQL it can also be your
9:56:51Oracle DB it can be db2 right so jdbc
9:56:55driver will be common and can be used
9:56:58for any of the things so this will
9:57:00basically help you to create a
9:57:02connection with any sort of RPMs system
9:57:06this is a very famous question what's
9:57:09the difference between hyen hyen Target
9:57:12Di and the difference with Warehouse
9:57:15hyen
9:57:17di hyphen hyen Target di with warehous
9:57:24v this is a very famous interview
9:57:26question for
9:57:29scoop because both will basically help
9:57:31you to put the data in some SDF specific
9:57:35location then what's the difference
9:57:37between
9:57:38them very good not soting all table and
9:57:42all is fine that you giving me a use
9:57:44case I want basically a proper solution
9:57:48since what's the difference between
9:57:51them is Define no no
9:57:55R what happens in Target DS in Target
9:58:00you if you're defining Target di you
9:58:03need to give the directory path name
9:58:08okay so you need to give the directory
9:58:10name where you will be keeping data so
9:58:12it is possible let's say in your my SQL
9:58:14or in rdbms let's say you have table
9:58:17called as accounts but now you want to
9:58:21import this table to your hdfs using
9:58:24scoop Now by if you give Target Dr you
9:58:29need to tell the name of the hdfs
9:58:32directory where you will be keeping this
9:58:34so now you can change this name of hbfs
9:58:37from accounts to let's say accounts one
9:58:40you can do all that if you're using
9:58:41Target di you can change the name of
9:58:44this accounts to accounts one and ldfs
9:58:47but if you're using Warehouse V in that
9:58:50case whatever the name will be there in
9:58:53your rdbms same name will be created in
9:58:57your
9:58:58hdms same name will be created your HFS
9:59:01no change with that okay so that is one
9:59:04of the major difference between them so
9:59:07Warehouse V will maintain the same name
9:59:10while with Target di you can keep the
9:59:12same name or different name as well so
9:59:15you are forced to basically give the
9:59:18name of the hdfs folder where you want
9:59:22to import while in W that's a
9:59:26different
9:59:27now can you tell me what this quer is
9:59:30doing
9:59:31here read this query and tell me what
9:59:34this query is doing here
9:59:37incremental data no
9:59:40no importing employer table but can you
9:59:42see this hpe and and we
9:59:45Closs it is filtering right it is only
9:59:48filtering all the employees table where
9:59:52your start date is greater than this
9:59:54value okay this is what this is
9:59:58doing see this this is what this is like
10:00:01now let's say this is a question in a
10:00:03scop import command you have mentioned
10:00:06to run eight parall map produce CL but
10:00:09scope is only running four what can be
10:00:12the
10:00:13reason what can be the reason very good
10:00:17very good yes because maybe a number of
10:00:21qus are not allowing you to run eight
10:00:24par right maybe you have less number of
10:00:27Cs itself in that case will not be able
10:00:29to take you up to eight after first
10:00:31right it will only use less number of C
10:00:34so basically if your ques are less in
10:00:36that case this is bound to happen give a
10:00:39scope command to show all the databases
10:00:42in the mql server can you give a scoop
10:00:44command to show all the databases in my
10:00:47SQL Server it should be simple see this
10:00:51scoop list databases not show databases
10:00:55list databases they okay hyen hyen
10:00:58connect give the connections that's it
10:01:00will list all the
10:01:02databases basically that's it so this
10:01:04will list all the databas
10:01:06okay so those sessions will be very
10:01:09useful for all of you thank you everyone
10:01:11for making it interactive and a nice
10:01:12session I hope you have enjoyed
10:01:14listening to this video please be kind
10:01:17enough to like it and you can comment
10:01:19any of your doubts and queries and we
10:01:22will reply them at the earliest do look
10:01:24out for more videos in our playlist And
10:01:27subscribe to Eddie Raa channel to learn
10:01:29more happy
10:01:34learning