Free YouTube Transcribe

Video transcript

Data Engineer Full Course in 10 Hours [2024] | Data Engineer Course For Beginners | Edureka

edureka! · 99,511 words · 453 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:03[Music]

0:08have you ever played with building

0:09blocks data Engineers are like expert

0:11Builders but instead of blocks they work

0:14with information they put together

0:16structures that make data neat and

0:18organized like creating a super cool

0:20puzzle where everything fits

0:23perfectly a data engineer is someone who

0:26designs and builds system to gather and

0:28store and manage lots of

0:31information on that note hello everyone

0:33welcome to this session to the data

0:35engineer full course join us on the

0:38journey and discover how data Engineers

0:40shape the future on decision making and

0:43Innovation we'll Begin by exploring the

0:46data engineering and the steps to become

0:48data engineer next we'll provide the

0:50overview of Big Data followed by the

0:53guidance on becoming the data engineer

0:55and insights into the associated salary

0:59asure is essential for data engineer

1:01because it offers a range of specialized

1:03tools and service designed to manage

1:06analyze and scale data effectively so as

1:09we will cover Azure topics such as Azure

1:12data Factory Azure database services and

1:14Azure SQL database and aure data link

1:18advancing further will explore Advanced

1:20Data modeling using PBI and aure

1:22including assure data bricks as we

1:25progress to Advanced topics we'll

1:27explore Haro Essentials covering HD

1:30Yan map reduce spark Pig hi and EDP for

1:34its

1:35ecosystem following this will progress

1:38to understanding Kafka streams to

1:40conclude the course we'll discuss the

1:42essential her interview question and

1:44answers to advance your career in data

1:46engineering interviews before we begin

1:49please consider subscribing to our

1:50YouTube channel and hit the Bell icon to

1:53stay updated on the latest tech content

1:55from edura also visit the Eda website

1:58for the data engineer master master

2:00program the link to which is given in

2:02the description box

2:04[Music]

2:08below today I'll share you an engaging

2:11story about two friends Alex and Bob hey

2:15have I told you about the incredible

2:16data engineering journey of a

2:18multinational e-commerce company no I

2:21don't think so what happened Alex well

2:24let me share this story with you Bob

2:26this company had lots of valuable data

2:28but struggled with scattered systems and

2:31outdated databases oh that's a tough

2:34situation what did they do about it Alex

2:37they embarked on a data engineering

2:38initiative B it was quite a journey I

2:41can imagine so what changes did they

2:43make Alex they integrated their data

2:46into a centralized system and automated

2:48data cleaning and transformation it made

2:51a huge difference Bob that sounds so

2:54promising did it impact their operations

2:56Alex definitely Bob they gained faster

2:59access to ACC data and generated

3:01realtime analytical reports impressive

3:04did they address data governance as well

3:06Alex yes they implemented data

3:08governance practices for data quality

3:10privacy and compilance that's

3:13commendable it's a great example of how

3:15data engineering can transform a

3:17business what an inspiring story Alex

3:20yeah I thought you would find it

3:22fascinating Bob data engineering truly

3:24has the power to drive success and

3:25Innovation together yeah thanks for

3:28sharing the story Alex it reinforces the

3:30importance of data Engineering in

3:32today's data driven World hello everyone

3:35this is Saia from edua and in this video

3:37we will be diving into the fascinating

3:39world of data engineering so data

3:42engineering is all about harnessing the

3:44power of data to drive meaningful

3:46insights in today's rapidly evolving

3:48digital landscape organizations are

3:51grappling with massive amounts of data

3:53and that's where data and sharing comes

3:55into play now let's take a quick look at

3:57the agenda for this video we will start

4:00by the introduction of what data

4:01insuring is and why it is important for

4:03so many businesses outside then we will

4:06delve into the key components of data

4:07engineering such as data inje data

4:10integration data transformation and data

4:12storage next we will discuss some of the

4:15responsibilities of data engineers and

4:17the comparison between data engineer

4:19data analyst and data scientist after

4:22that we will delve into the installation

4:24process of popular tools and

4:26Technologies which is used in data

4:28engineering finally we will wrap up with

4:30a glimpse of what data pipelines are so

4:33let's understand what is data

4:35engineering First Data engineering

4:37involves the design development and

4:39maintenance of system and infrastructure

4:41to handle large volumes of data

4:43effectively it focuses on the extraction

4:46transformation loading and storage of

4:48data which ensure its quality and

4:49scalability of analysis now let's

4:52understand what are the key components

4:54of data engineering data engineering

4:56helps to collect data from the disparate

4:58sources and integrate into a unified

5:01format which allows organization to have

5:03a comprehensive view of their data it

5:06also involves processing and

5:08transforming data into a suitable format

5:11which ensures data quality consistency

5:13and applications it includes the

5:16development of data pipelines and

5:18processing system that enable the

5:19efficient processing and Analysis of

5:21data this includes batch processing for

5:24large scale data transformation and

5:26realtime processing for immediate

5:28insights and action

5:30data engineering plays a crucial role in

5:32enabling data driven World by providing

5:35clean integrated and accessible data

5:37data insuring empowers organizations to

5:40make an informed decision based on

5:42accurate insights and Analysis overall

5:45it is used to handle the complexity of

5:46data processing storage and integration

5:49which ensures that the organization can

5:51leverage the full potential of their

5:53data assets for strategic operations and

5:56Innovations now after the thorough

5:59understanding of what data engineering

6:00is and its key features also we will Del

6:03into the importance of data

6:05engineering so here are some of the key

6:08reasons why data engineering should be

6:10important or why data engineering is

6:12important so our first reason is it

6:15focuses on building scalable and

6:18efficient data processing systems by

6:20optimizing data pipelines leveraging

6:23distributed computing Technologies and

6:25employing performance tuning techniques

6:27data Engineers ensured did organization

6:30can handle large volumes of data and

6:33Achieve faster processing times it

6:35provides the foundation for advanced

6:38analytics and machine learning

6:40initiatives by structuring and preparing

6:42data in a switchable format data

6:44Engineers unable data scientists and

6:46analyst to extract valuable insights

6:49build predictive models and develop

6:51machine learning algorithms data

6:54insuring enables organization to process

6:56and analyze data as it arrives this is

6:59crucial in scenarios such as fraud

7:02detection recommendation systems iot

7:05applications and monitoring systems that

7:07require immediate insights and actions

7:10as we already know the importance of

7:12data engineering we will cover the real

7:14word applications based on this so our

7:17first application is e-commerce sites so

7:20data insuring is used to collect and

7:21process large volumes of customer data

7:24transaction data and product data this

7:27enables e-commerce companies to predict

7:29recommendations optimize pricing

7:31strategies and improve Inventory

7:33management our second example is social

7:36media sites like data insuring plays a

7:39crucial role in collecting processing

7:41and analyzing social media data which

7:43allows companies to monitor brand

7:45sentiment track user interactions and

7:48derive insights for targeted marketing

7:51campaigns so our third example is in

7:53finance and banking sector data insuring

7:56is used to handle financial data

7:58including trans action records customer

8:00information and Market data it enables

8:03fraud detection risk assessment

8:05algorithmic trading and personalized

8:08Financial Services also okay so now the

8:11fourth application is in the healthcare

8:13sector data insuring is employed to

8:15manage and analyze patient records

8:17Medical Imaging data and clinical child

8:20data it enables Healthcare Providers and

8:22researchers to gain insights improve

8:24patient outcomes and develop predictive

8:27models so this examples highlight the

8:29diverse application and the importance

8:31of data Engineering in various

8:33Industries demonstrating how it enables

8:35organization to Leverage The Power of

8:37data for operation and Innovation now

8:40the data insuring process involves

8:42several key steps including data

8:44injection transformation storage

8:47processing and integration let's discuss

8:50each step in more detail so our first

8:52step is data inje data inje refers to

8:56the process of collecting and importing

8:58data from various sources into a data

9:00system or data pipeline this can involve

9:03extracting data from databases files

9:06apis streaming platforms or the other

9:08resources the goal is to GA relevant

9:11data and make it available for the

9:12further processing and Analysis Second

9:15Step would be the data transformation

9:18once the data is ingested it often needs

9:20to be transformed into a suitable format

9:22for analysis or storage this data

9:25transformation involves cleaning

9:27validating and restructuring the data to

9:29ensure consistency and usability this

9:32step may include tasks such as data

9:34filtering aggregation normalization data

9:37type conversion or the application of

9:39business rules so our third step would

9:41be data storage so after the data is

9:44transformed it needs to be stored in a

9:46structured manner right so this

9:48typically involves using databases or

9:50data storage system that provide

9:52efficient storage and retrial

9:54capabilities popular choices for data

9:56storage include relational databases

9:58data data warehouses data laks or

10:01distributed file systems the selection

10:03depends on the specific requirements of

10:05the project such as data volume SS

10:07patterns and the analytical records so

10:10our next step would be data processing

10:13data processing involves performing

10:15computations and analysis on the stored

10:17data this STP can include tasks such as

10:20data aggregation data Improvement data

10:22summarization statistical calculation or

10:25machine learning algorithms data

10:27processing can be done through various

10:28to tools such as the SQL queries data

10:31processing engines or the custom scripts

10:34Next Step would be the data integration

10:36so in this step data integration

10:39involves combining data from multiple

10:41sources to create a unified and

10:43comprehensive view this step is crucial

10:45when dealing with heterogeneous data

10:47sources or when different style produced

10:49data this needs to be Consolidated data

10:52integration can be achieved through data

10:54consolidation or by using extract

10:56transform load process to combine and

10:59merge data from various sources next our

11:02last step would be the data governance

11:04so data governance is nothing but a

11:06framework in a set of process which

11:08ensures the effective management and

11:10utilization of data assets within an

11:14organization it involves establishing

11:16policies procedures and guidelines for

11:18data management data quality and data

11:21usage all the steps in data Engineering

11:24Process are typically iterative and may

11:26require continuous monitoring optimizing

11:29ation and maintenance to ensure data

11:31quality reliability and performance data

11:35Engineers play a crucial role in

11:37designing and implementing efficient and

11:39scalable data pipelines to support

11:41datadriven applications and

11:43analytics Now we move ahead with the key

11:46responsibilities of data Engineers which

11:48typically includes designing and

11:51developing scalable and efficient data

11:53pipelines which extract transform and

11:56load data from various sources to

11:59targeted systems next would be the

12:01managing and optimizing data storage

12:03infrastructure for performance

12:05scalability and reliability selecting

12:07and configuring appropriate data storage

12:10technology such as relational databases

12:12data warehouses or no SQL databases to

12:14meet data storage and retrial

12:16requirements data Engineers are also

12:19responsible for applying data

12:20transformation techniques to clean

12:22filter Aggregate and structure data for

12:25analysis or consumption by Downstream

12:27systems they're also responsible for

12:30implementing data processing task using

12:32programming languages or data processing

12:34Frameworks to manipulate and transform

12:37data efficiently now after analyzing the

12:40key responsibilities of data insuring we

12:42will proceed with our next topic and

12:44understand the key differences between

12:46these three roles that is data engineer

12:48data analyst and data scientist these

12:51three are the distinct role within the

12:53field of data science each with its own

12:55set of responsibilities and skill

12:57requirements so our our first profile is

12:59of data engineer so data engineer are

13:02responsible for Designing building and

13:04maintaining the infrastructure and

13:06system that enable data storage

13:08processing and retrial as we all know

13:11they focus on creating and managing the

13:12data pipelines and architecture

13:15necessary for efficient data collection

13:17transformation and storage data

13:19Engineers also work closely with

13:21software engineers and database

13:23administrator to ensure data is

13:25accessible reliable and scalable they

13:28typically work work with tools like

13:29Hardo spark SQL ETL Frameworks and

13:33cloud-based platforms for data

13:34processing and storage then our next

13:37profile would be the data analyst so

13:40data analysts are focused on analyzing

13:41and interpreting data to derive

13:43meaningful insights they work with

13:46structured and unstructured data to

13:47identify patterns strengths and

13:49correlations data analysts are

13:51proficient in statistical analysis data

13:54visualization tools and data quaring

13:56techniques as well they often use tools

13:58tools like SQL Excel table or powerbi to

14:02analyze data and create reports and

14:04dashboards after dat our next profile

14:07would be the data scientist data

14:09scientist possess a blend of skills from

14:11mathematics statistics programming and

14:13domain knowledge they leverage their

14:15expertise to develop and Implement

14:17complex algorithms and models to solve

14:20integrate data patterns or extract

14:22insights from large data sets data

14:25scientists employ techniques like

14:26machine learning predictive modeling and

14:28statistical analysis to build predictive

14:30models uncover patterns and make

14:32predictions also they also collaborate

14:34with stakeholders to Define business

14:36problems and design experiments to

14:38together their data so to summarize this

14:41data Engineers focus on the

14:43infrastructure and data pipelines data

14:46analysts work on analyzing and Reporting

14:48data while data scientists concentrate

14:51on Advanced modeling and extracting

14:54insights so to make all this happen data

14:57Engineers rely on power powerful tools

14:59and Technologies platforms like Apache

15:02spark Apache Kafka SQL and no SQL

15:05databases and cloud services such as AWS

15:08and gcp provides the building blocks for

15:10efficient data engineering this tools

15:12help process vast amounts of data

15:15facilate realtime data streaming and

15:17ensure secure and scalable data storage

15:20now let's understand what are the data

15:23pipelines are so here is the glimpse of

15:25what data pipelines are and why we use

15:27data Pipelines

15:29so data pipelines are the series of

15:31steps that extract transform and load

15:34data from source to destination itself

15:36okay so they enable the efficient and

15:38automated flow of data through different

15:41stages ensuring data quality consistency

15:43and avability

15:45let's understand each step one by one so

15:48the first step is extraction the

15:50extraction phase of ETL involves

15:53retrieving data from various sources

15:55such as databases files apis or

15:57streaming platform platforms the goal is

16:00to extract the relevant data needed for

16:02further processing and Analysis this

16:05process typically includes establishing

16:07connection to the data sources

16:09performing data queries or using data

16:11extraction tools to pull the required

16:14data into the data pipeline Next Step

16:16would be the transformation the

16:18transformation phase of ETL focuses on

16:20cleaning validating and reshaping the

16:23exctic data to ensure its quality and

16:26consistency this step involves applying

16:28various operations and rules to the data

16:30such as data filtering data type

16:32conversions data aggregation data

16:35enrichment or data

16:36normalization the transformation process

16:39aims to make the data suitable for

16:41analysis storage or integration into the

16:43destination system Next Step would be

16:45the loading the loading pH of ETL

16:48involves storing the transform data into

16:50the targeted system such as databases

16:52data warehouses or data Lakes then the

16:55data is loaded in a structured format

16:57that aligns with the schema or format of

17:00the destination system this phase may

17:03include tasks such as data mapping

17:05schema matching data passing or indexing

17:08to optimize data storage and retrieval

17:11after that our next step would be the or

17:13maybe we call this our last step would

17:15be the monitoring and handling

17:17monitoring and handling refer to the

17:19ongoing monitoring management and the

17:21maintenance of data pipelines on okay so

17:24now let's discuss about batch processing

17:27and realtime streaming pipeline lines so

17:29in a batch processing pipeline there is

17:31a delay between time data is collected

17:33and when it is processed this delay can

17:36range from minutes to hours or even days

17:38depending on the schedule intervals on

17:41the other hand realtime streaming

17:43pipeline aim to process data as it

17:45arrives which results in a minimal

17:48latency data is processed analyzed in

17:50near real time or within a very low

17:52delay the second point is batch

17:55processing pipelines are designed to

17:56handle large volumes of data efficiently

17:59they can process and analyze massive

18:01amounts of historical data in a batch

18:03mode on the other hand realtime

18:05streaming pipelines focus on processing

18:07data as it arrives making them more

18:10suitable for handling data streams with

18:12continuous High Velocity data updates

18:15now the third point is batch processing

18:17pipelines typically utilize batch

18:19processing Frameworks like Apache spark

18:22or Hadoop map redu this Frameworks

18:25process data in chunks or batches which

18:27allows for parall processing and

18:29optimize resource utilizations also

18:32while realtime streaming pipelines often

18:34use streaming Frameworks like Apache

18:36Kafka Apache Flink or Apache stom this

18:39Frameworks enable continuous processing

18:41of data streams supporting low latency

18:43operations and realtime

18:45analytics now the fourth point is batch

18:48processing pipelines are commonly used

18:50for tasks that involve historical

18:52analysis generating periodic reports or

18:55data preparation for machine learning

18:56models they are Well Suited for

18:58scenarios where processing time is not

19:00critical but analyzing large volumes of

19:02data is essential on the other hand

19:05realtime streaming pipelines are ideal

19:07for application that require realtime

19:09monitoring immediate response or instant

19:12insights based on live data use cases

19:14include like fraud detection realtime

19:16recommendation systems network

19:18monitoring or iot sensor data analysis

19:21so our last point is batch processing

19:24pipelines often require significant

19:25Computing resources during the

19:27processing phase as they process large

19:29volumes of data in a batch mode whereas

19:32realtime streaming pipelines also

19:34requires Computing resources but are

19:36more focused on low latency processing

19:38and continuous data streams requiring

19:41efficient resources allocation and

19:43management so it's important to note

19:45that there can be the overlap between

19:47batch processing and realtime streaming

19:49pipelines and hybrid architectures

19:51combining both approaches are common so

19:54that wraps up our Deep dive into the

19:56world of data engineering we have

19:57covered the essential aspects of data

19:59pipelines governance and security giving

20:02you a comprehensive understanding of how

20:04data is ingested transformed stored and

20:07[Music]

20:12processed now how to become a data

20:15engineer the road map to become a data

20:18engineer can go like from being

20:21proficient in programming languages to

20:23learning Automation and scripting then

20:26understanding your databases

20:28and mastering data processing techniques

20:31to studying cloud computing and

20:34internalizing

20:35infrastructure now we'll see how to get

20:38started with becoming a data engineer to

20:41get started with the learning process

20:43you can look into our Eda YouTube

20:45channel to start with even with no prior

20:48knowledge of data engineering one can

20:50just go through the videos and get an

20:53understanding of the whole

20:55subject we even have edureka blogs to to

20:58help you with detailed information on

21:00the topics and can give you a clearer

21:03picture apart from these we even have

21:06premium courses which can help you

21:10understand the topic at ease with the

21:13personal trainer and 24 hours access to

21:17Lifetime content you can learn here at

21:19your own pace with a live trainer who's

21:22extremely efficient and knowledgeable

21:24and experienc in the particular field

21:27these courses can even get you

21:28certificates which you can add in your

21:30CV for better opportunity in your job

21:33market so that's it for today I hope

21:36this video helped you and will help you

21:38decide how to become a data

21:42[Music]

21:46engineer now I feel Sor of it's the best

21:49time to tell the story about how data

21:51evolved and how big data came fine RMA

21:54so we'll move forward so sort of what

21:57can you notice here RMA I see how

22:00technology has evolved earlier we had

22:02landline phones but now we have

22:04smartphones we have Android we have IOS

22:06that are making our lives smarter as

22:08well as our phone smarter apart from

22:10that we were also using bulky desktops

22:12for processing MBS of data now if you

22:14can remember we were using floppies and

22:16you know how much data it can store

22:18right then came hard dis for storing TBS

22:20of data and now we can store data on

22:22cloud as well and similarly nowadays

22:25even self-driving cars have come up and

22:28I know you must be thinking why are we

22:29telling that now if you notice due to

22:32this enhancement of Technology we're

22:34generating a lot of data so let's take

22:36the example of your phones have you ever

22:39noticed how much data is generated due

22:41to your fancy smartphones your every

22:44action even one video that you send

22:46through WhatsApp or any other messenger

22:48app that generates data now this is just

22:50an example you have no idea how much

22:53data you're generating because of every

22:55action you do now the deal is this data

22:57is not in a format that our relational

22:59database can handle and apart from that

23:02even the volume of data has also

23:04increased exponentially now I was

23:06talking about self-driving cars so

23:08basically these cars have sensors that

23:10records every minute details like the

23:12size of the obstacle the distance from

23:14the obstacle and many more and then it

23:17decides how to react now you can imagine

23:19how much data is generated for each

23:21kilometer that you drive on that car I

23:24completely agree with you RMA so let's

23:26move forward and focus on ious other

23:28factors behind the evolution of data I

23:31think you guys must have heard about iot

23:33if you can recall in the previous slide

23:35we were discussing about self-driving

23:36cars it is nothing but an example of iot

23:39let me tell you what exactly it is iot

23:42connects your physical device with

23:44internet and makes the device smarter so

23:46nowadays if you have noticed we have

23:48Smart ACS TVs Etc so we'll take the

23:51example of smart air conditioners so

23:53this device actually monitors your body

23:55temperature and the outside temperature

23:57and accordingly decides what should be

23:59the temperature of the room now in order

24:01to do this it has to First accumulate

24:04data from where it can accumulate data

24:06from internet through sensors that are

24:08monitoring your body temperature and the

24:10surroundings so basically from various

24:13sources that you might not even know

24:14about it is actually fetching that data

24:17and accordingly it decides what should

24:19be the temperature of your room now we

24:20can actually see that because of iot we

24:22are generating huge amount of data now

24:25there's one startat also that is there

24:26in front of your screen so if you notice

24:28by 2020 we'll have 50 billion iot

24:31devices so I don't think so I need to

24:34explain much that how iot is generating

24:36huge amount of data so we'll move

24:38forward and focus on one more factor

24:39that is social media now when we talk

24:42about social media I think RMA can

24:44explain this better right RMA yeah s but

24:47I'm pretty sure that even you use it so

24:50let me tell you that social media is

24:52actually one of the most important

24:54factor in the evolution of big data so

24:57nowadays everyone is using Facebook

24:59Instagram YouTube and a lot of other

25:02social media websites so these social

25:04media sites have so much data for

25:07example it will have your personal

25:09details like your name age and apart

25:11from that even each picture that you

25:14like or react to also generates data and

25:16even the Facebook pages that you go

25:18around liking that is also generating

25:20data and nowadays you can see that most

25:23people are sharing videos on Facebook so

25:25that is also generating huge amount of

25:28data and the most challenging part here

25:30is that the data is not present in a

25:32structured Manner and at the same time

25:35it is huge in size isn't that right sort

25:38of can't agree more the point you made

25:40about the form of data is actually one

25:42of the biggest factor for the evolution

25:44of big data so due to all these reasons

25:46that we have discussed have not only

25:48increased the amount of data but it has

25:50also shown us that data is actually

25:52getting generated in various formats for

25:54example data is generated with videos

25:56that is actually unstructured same goes

25:58for images as well so there are numerous

26:00or you can say millions of ways in which

26:02data is getting generated nowadays

26:05absolutely and these are just few

26:07examples that we have given you there

26:09are many other driving factors for the

26:11evolution of data so these are few more

26:14examples because of which data is

26:17evolving and converting to Big Data

26:19we'll discuss about the retail part I'm

26:21pretty sure that all of you must have

26:22visited websites like Amazon flip cart

26:24Etc and rishma I know you visited a lot

26:27of times

26:28yeah I do and suppose RMA wants to buy

26:30shoes so she won't just directly go buy

26:32shoes she'll search for a lot of shoes

26:35so somewhere her search history will be

26:36stored and I know for sure that this

26:39won't be the first time that she's

26:40buying something so there will be her

26:42purchase history as well along with her

26:44personal details and there are numerous

26:46ways in which she might not even know

26:47that she's generating data and obviously

26:49Amazon was not present earlier so at

26:52that time there is no way that such huge

26:53amount of data was generated similarly

26:56the data has evolved due to other

26:57reasons as well like Banking and finance

26:59media and entertainment etc etc so now

27:02the deal is what exactly is Big Data how

27:05do we consider data as big data so let's

27:07move forward and understand what exactly

27:10it

27:11is okay now let us look at the proper

27:14definition of Big Data even though we've

27:16put forward our own definitions already

27:19so sort of why don't you take us through

27:21it yes RMA sure so big data is a term

27:24for collection of data sets so large and

27:26complex that it becomes difficult to

27:28process using onhand database system

27:31tools or traditional data processing

27:34applications okay so what I understand

27:37from this is that our traditional

27:39systems are a problem because they're

27:41too oldfashioned to process this data or

27:44something no RMA the real problem is

27:47there is too much data to process when

27:49the traditional systems were invented in

27:51the beginning we never anticipated that

27:53we would have to deal with such enormous

27:55amount of the data it's like like a

27:57disease infected on you you don't change

27:59your body orientation when you get

28:01infected with a disease right RMA you

28:03cure it with medicines couldn't agree

28:06more sort of now the question is how do

28:09we consider some data as Big Data how do

28:11we classify some data as Big Data how do

28:14we know which kind of data is going to

28:16be hard for us to process well Sor of we

28:19have the five vs to tell us

28:21that so let's take a closer look at what

28:24are those so starting with the first V

28:27it's the volume of data it's

28:29tremendously large so if you look at the

28:31stats here you can see the volume of

28:33data is rising exponentially so now

28:35we're dealing with just 4.4 zettabytes

28:38of data and by 2020 just in three years

28:41is expected that the data will rise up

28:44to 44 zettabytes which is like equal to

28:4644 trillion gigabytes so that's really

28:50really

28:51huge it is because all these humongous

28:54all this humongous data is coming from

28:56multiple sources and that is the second

28:59V which is nothing but variety we deal

29:01with so many different kinds of files at

29:03all once there are MP3 files videos Json

29:06CSV tsv and many more now these are all

29:09structured unstructured and

29:10semi-structured all together now let me

29:12explain you this with the diagram that

29:14is there on your screen so over here we

29:16have audio we have video files we have

29:18PNG files we have Json log files emails

29:22various formats of data now this data is

29:24classified into three forms one is

29:26structured format now in structure

29:27format you have a proper schema for your

29:29data so you know what all columns will

29:31be there and basically you know the

29:33schema about your data so it is

29:35structured it is in a structured format

29:37or you can say in a tabular format now

29:39when we talk about semi-structured files

29:41these are nothing but Json XML and CSV

29:43files where schema is not defined

29:45properly now when I go to unstructured

29:47format we have log files here audio

29:50files videos and images so these are all

29:52considered as unstructured files and Sor

29:56it is also because of of the speed of

29:58accumulation of all this variety of data

30:00alog together which brings us to our

30:02third V which is velocity so if you look

30:05here earlier we were using Mainframe

30:07systems huge computers but less data

30:10because there were less people working

30:11with computers at that time but as

30:14computers evolved and we came to the

30:15client server model the time came for

30:18the web applications and the internet

30:20boomed and as it grew among the masses

30:22the web applications got increased over

30:24the internet and everyone started using

30:26all this applications and not only from

30:29their computers and also from mobile

30:31devices so more users more appliances

30:34more apps and H hands a lot of data and

30:37when you talk about people generating

30:39data or Internet RMA the one kind of

30:42application that strikes first in my

30:43mind is social media so you tell me how

30:46much data you generate alone with your

30:47Instagram posts and stories uh it will

30:50be quite a boast if I only talk about

30:52myself here so let's talk including

30:55every social media user so if you see

30:57the stats in front of your screen you

31:00can see that for every 60 seconds there

31:03are 100,000 tweets actually more than

31:06100,000 tweets generated in Twitter

31:08every minute similarly there are 695,000

31:12status updates on Facebook when you talk

31:15about messaging there are 11 million

31:17messages generated every minute and

31:20similarly there are

31:25698,000 email and that equals to almost

31:301,820 terabytes of data and obviously

31:34the number of mobile users are also

31:36increasing every minute and there are

31:382117 plus new mobile users every 60

31:42seconds gez that's a lot of data I don't

31:45even want to go ahead and calculate the

31:46total it would actually scare me yeah

31:49that's a lot now the bigger problem is

31:52how to extract the useful data from here

31:55and that's when we come to our next week

31:57that is value so over here what happens

32:00first you need to mine the useful

32:02content from your data basically you

32:04need to make sure that you have only

32:05useful fields in your data set after

32:07that you perform certain analytics or

32:09you say you you perform certain analysis

32:11on that data that you have cleaned and

32:13you need to make sure that whatever

32:15analysis you have done it is of some

32:17value that is it will help you in your

32:19business to grow it can basically find

32:21out certain insights which were not

32:23possible earlier so you need to make

32:25sure that whatever big data that has

32:26been generated or whatever data that has

32:28been generated it makes sense it will

32:31actually help your business to grow and

32:32it has some value to it now getting the

32:34value of this data is one big challenge

32:37let me tell you why and that brings us

32:39to our next V which is veracity now this

32:42big data has a lot of

32:43inconsistencies obviously when you're

32:45dumping such huge amount of data some

32:48data packets are bound to lose in the

32:50process now what we need to do we need

32:52to fill up these missing data and then

32:54start mining again and then process it

32:56and then come up with a good Insight if

32:59possible so if you can notice there's a

33:01diagram in front of your screen so over

33:03here we have this field which is not

33:04defined similarly this field and if you

33:06can notice here when we talk about this

33:08minimum value you see the other minimum

33:10values and when you talk about this it

33:12is it is way more than the other fields

33:15present in this particular column

33:16similarly goes for this particular

33:18element as well okay so obviously

33:21processing data like this is one

33:23problematic thing and now I get it why

33:26Big Data is a problem statement well we

33:29have only five vs now but maybe later on

33:31we'll have more so there are good

33:33chances that big data might be even more

33:36big okay so there are a lot of problems

33:39in dealing with big data but there are

33:41always different ways to look at some

33:43things so let us get some positivity in

33:45the environment now and let us

33:48understand how can we use Big Data as an

33:50opportunity yes RMA and I would say the

33:53situation is similar to the proverb when

33:55life throws you lemons make

33:57lemonade yeah so let us go through the

34:00fields where we can use Big Data as a

34:02boon and there are certain unknown

34:05problems solved only because we started

34:07dealing with big data and the Boon that

34:09you're talking about RMA is big data

34:11analytics first thing with big data we

34:13figured out how to store our data cost

34:16effectively we were spending too much

34:18money on storage before until Big Data

34:20came into the picture we never thought

34:22of using commodity Hardware to store and

34:25manage a data which is both reliable and

34:27feasible as compared to the costly

34:29servers now let me give you a few

34:31examples in order to show you how

34:33important big data analytics is nowadays

34:35so when you go to a website like Amazon

34:37or YouTube or Pandora Netflix any other

34:40website so they'll actually provide you

34:42certain fields in which they'll

34:43recommend some products or some videos

34:45or some movies or some songs for you

34:47right so how do you think they do that

34:49so basically whatever data that you are

34:51generating on these kind of websites

34:53they make sure that they analyze it

34:54properly and let me tell you guys that

34:57data is not small it is actually big

34:59data now they analyze that big data and

35:02they make sure that whatever you like or

35:04whatever your preferences are

35:05accordingly they'll generate

35:06recommendations for you and when I go to

35:09YouTube I don't know if you guys have

35:10noticed it but I'm pretty sure you must

35:12have done that so when I go to YouTube

35:14YouTube knows what song or what video

35:16that I want to watch next similarly

35:18Netflix knows what kind of movies are

35:20like and when I go to Amazon it actually

35:23shows me what all products that I would

35:25prefer to buy right so so how do you

35:27think it happens it happens only because

35:28of big data analytics okay so there is

35:31one more example that just popped into

35:33my mind I'll share with you guys so

35:36there was this time when the Hurricane

35:38Sandy was about to hit on New Jersey in

35:41United States so what happened then the

35:44Walmart used big data analytics to

35:47profit from it now I'll tell you how

35:49they did it so what Walmart did is that

35:52they studied the purchase patterns of

35:55different customers when a Hur hurri is

35:57about to strike or any kind of natural

35:58Calamity is about to strike on a

36:00particular area and when they made an

36:02analysis of it so they found out that

36:05people tend to buy emergency stuff like

36:08flashlight life jackets and a little bit

36:11of other stuff and interestingly people

36:13also buy a lot of strawberry poptart

36:17strawberry poptarts are you serious yeah

36:20now I didn't do that analysis so I

36:22Walmart did that and apparently it is

36:25true so what they did is so they stuffed

36:28all their stores with a lot of

36:29strawberry Pop-Tarts and emergency stuff

36:32and obviously it was sold out and they

36:34earned a lot of money during that time

36:37but my question here RMA is people want

36:39to die eating strawberry poptarts like

36:42what was the idea behind strawberry

36:43poptarts I'm pretty unsure about it but

36:45yeah since you have given us a very

36:47interesting example and Walmart did that

36:49analysis we didn't do it so yeah so it

36:51is a very good example in order to

36:53understand how big data analytics can

36:55help your business to grow and find

36:57better insights from the data that you

36:59have yeah and also if you want to know

37:01why strawberry poptarts maybe later on

37:04we can start making an analysis by

37:05gathering some more data also yeah that

37:08can be possible okay so now let's move

37:11ahead and take a look at a case study by

37:14IBM how they have used big data

37:16analytics to profit their company so if

37:19you have noticed that earlier the data

37:21that was collected from The Meters that

37:23you have in your home that measures the

37:24electricity consumed it is actually

37:26sending data after one month but

37:29nowadays what IBM did they came up with

37:31this thing called smart meter and that

37:33smart meter used to collect data after

37:35every 15 minutes so whatever energy that

37:38you have consumed after every 15 minutes

37:40it will send that data and because of it

37:42big data was generated so we have some

37:44stats here which says that we have 96

37:47million reads per day for every million

37:50meters which is pretty huge this data

37:52the amount of data that is generated is

37:54pretty huge now IBM actually realize the

37:57data that they're generating it is very

37:59important for them to gain something

38:01from that data so for that what they

38:03need to for that what they need to do

38:04they need to make sure that they analyze

38:06this data so they realize that big data

38:08analytics can solve a lot of problems

38:11and they can get better business Insight

38:13through that so let us move forward and

38:14see what type of analysis they did on

38:16that data so before analyzing that data

38:19they came to know that energy

38:20utilization and billing was only

38:23increasing now after analyzing Big Data

38:25they came to know that during Peak load

38:27the users require more energy and during

38:30off peak times the users require less

38:32energy so what advantage they must have

38:34got from this analysis one thing that I

38:36can think of right now is they can tell

38:39the industries to use their Machinery

38:41only during the off peak times so that

38:43the load will be pretty much balanced

38:45and you can even say that time of use

38:47pricing encourages cost savy retail like

38:50industrial heavy machines to be used off

38:52peak time so yeah they can save money as

38:55well because off peak times pricing will

38:58be less than the peak time prices right

39:01so this is just one analysis now let us

39:03move forward and see the IBM Suite that

39:05they developed so over here what happens

39:09you first dump all your data that you

39:11get in this data warehouse after that it

39:13is very important to make sure that your

39:15user data is secure then what happens

39:18you need to clean that data as I've told

39:20you earlier as well there might be many

39:21feeds that you don't require so you need

39:23to make sure that you have only useful

39:24material or useful data in your data set

39:27and then you perform certain analysis

39:30and in order to use this Suite that IBM

39:32offered you efficiently you have to take

39:34care of a few things the first thing is

39:36that you have to be able to manage the

39:38smart meter data now there is a lot of

39:41data coming from all this million Smart

39:43Meters so you have to be able to manage

39:45that large volume of data and also be

39:48able to retain it because maybe later on

39:50you might need it for some kind of

39:52regulatory requirements or something and

39:55next thing you should keep in mind is is

39:56to monitor the distribution grid so that

39:59you can improve and optimize the overall

40:01grid reliability so that you can

40:03identify the abnormal conditions which

40:05are causing any kind of problem and then

40:07you also have to take care of optimizing

40:10the unit commitment so by optimizing the

40:12unit commitment the companies can

40:14satisfy their customers even more they

40:16can reduce the power outages that is the

40:19they can reduce the power outages so

40:21that their customers don't get angry

40:23more identify problems and then reduce

40:25it obviously

40:27and then you have also to optimize the

40:29energy trading so it means that you can

40:31advise your customers when they should

40:33use their appliances in order to

40:35maintain that balance in the power load

40:38and then you also have to forecast and

40:40schedule loads so companies must be able

40:43to predict when they can profitably sell

40:45the Excess power and when they need to

40:48hedge the supply and continuing from

40:51this now let's talk about how Encore

40:53have made use of the IBM solution so

40:56Encore is an electric delivery company

40:59and it is the largest electrical

41:01distribution and transmission company in

41:03Texas and it is one of the six largest

41:06in the United States they have more than

41:08three million customers and their

41:10service area covers almost

41:13117,000 square miles and they began the

41:17advanced speeder program in 2008 and

41:20they have deployed almost 3.25 million

41:23meters serving customers of North and

41:27Central Texas so when they were

41:29implementing it they kept three things

41:31in mind the first thing was that it

41:34should be instrumented so this solution

41:36utilizes the smart electricity meters so

41:39that they can accurately measure the

41:41electricity usage of a household in

41:44every 15 minutes because like we

41:46discussed that the smart meters were

41:47sending out data every 15 minutes and it

41:50provided data inputs that is essential

41:52for consumption insights next thing is

41:55that it should be interconnected so now

41:58the customers have access to the

41:59detailed information about the

42:01electricity they are consuming and it

42:03creates a very Enterprise wide view of

42:06all the meter assets and it helped them

42:08to improve the service delivery the next

42:11thing is to make your customers

42:13intelligent now since it is getting

42:15monitored already about how each of the

42:17household or each customer is consuming

42:19the power so now they're able to advise

42:22the customers about maybe to tell them

42:24to wash their clothes at night because

42:27they're using a lot of appliances during

42:28the daytime so maybe they could divide

42:30it up so that they can use some

42:32appliances at off peak hours so that

42:34they can even save more money and this

42:37is beneficial for both of them for both

42:39the customers and the company as well

42:42and they have gained a lot of benefits

42:44by using the IBM solution so what are

42:47the benefits they got is that it enabled

42:50oncore to identify and fix outages

42:53before the customers get inconvenience

42:55that means they were able to ident

42:56identify the problem before it even

42:57occurred and it also improved the

42:59emergency response on events of severe

43:02weather events and views of outages and

43:05it also provides the customers the data

43:07needed to become a active participant in

43:09the power consumption management and it

43:12enabled every individual household to

43:14reduce their electrical consumption by

43:16almost 5 to 10% and this is how oncore

43:20used the IBM solution and made huge

43:23benefits out of it just by using big

43:25data analytics that IBM performed but

43:27let me just interrupt right now so since

43:30RMA told us in the beginning as well

43:32that there are no free lunches in life

43:34right so this is an opportunity but

43:36there are many problems to encase this

43:38opportunity right so let us focus on

43:40those problems one by

43:42one so the first problem is storing

43:44colossal amount of data so let's discuss

43:47few starts that are there in front of

43:49your screen so data generated in the

43:51past 2 years is more than the previous

43:53history in total so guys what are we

43:55doing toop generating so much amount of

43:58data and it said that by 2020 total

44:01Digital Data will grow to 44 Zab bytes

44:04approximately and there's one more stat

44:06that amazes me is about 1.7 MB of new

44:11information will be created every second

44:13for every person by 2020 so storing this

44:16huge data in traditional system is not

44:19possible the reason is obvious the

44:21storage will be limited for one system

44:24for example you have a with a storage

44:26limit of 10 terab but your company is

44:29growing really fast and data is

44:31exponentially increasing now what you'll

44:33do now at one point you'll exhaust all

44:35the storage so investing in huge servers

44:38is definitely not a cost effective

44:41solution so RMA what do you think what

44:43can be the solution to this problem uh

44:45according to me a distributed file

44:48system will be a better way to store

44:50this huge data because with this we'll

44:52be uh saving a lot of money let me tell

44:54you how because due to this distributed

44:57system you can actually store your data

45:00in commodity Hardware instead of

45:02spending money on high-end servers don't

45:04you agree sort of completely now we know

45:07storing is a problem but let me tell you

45:09guys it is just one part of the problem

45:11let's see few more okay so since we saw

45:15that the data is not only huge but it is

45:18present in various formats as well like

45:21unstructured semi-structured and

45:23structured so you not only need to store

45:26this huge data but you also need to make

45:28sure that a system is present to store

45:31this varieties of data generated from

45:33various sources and now let's focus on

45:35the next problem now let's focus on the

45:38diagram so over here you can notice that

45:40the hard disk capacity is increasing but

45:43the disc transfer performance or speed

45:45is not increasing at that rate let me

45:47explain you this with an example if you

45:50have only 100 MVPs input output Channel

45:54and you are processing say 1 ter of data

45:56now how much time will it take maybe

45:59calculate it'll be somewhere around 2.91

46:03hours right so it'll be somewhere around

46:062.91 hours and I have taken an example

46:09of 1 terabytes what if you're processing

46:11some Zab bytes of data so you can

46:13imagine how much time will it take now

46:15what if you have four input output

46:17channels for the same amount of data

46:20then it'll take

46:22approximately 72 hours or converted to

46:25minutes so it be around 43 minutes

46:27approximately right and now imagine

46:30instead of 1 TB you have zabt of data

46:32for me more than storage accessing and

46:34processing speed for huge data is a

46:37bigger problem okay so RMA has a very

46:39good example to discuss yeah so since

46:41you were talking about accessing the

46:43data and you told us already about how

46:46Amazon at different websites and YouTube

46:48they make those recommendations so if

46:51there was no solution for it if it would

46:53take so much time to access the data the

46:55recommendation system won't work at all

46:58and they make a lot of money just for

47:00recommendation system because a lot of

47:02people go there and click over there and

47:04buy that product right so let's consider

47:06that that it is taking like hours or

47:08maybe years of time in order to process

47:10my that big amount of data so let's say

47:13that at one time I purchased an iPhone

47:165s from Amazon and after two years I'm

47:20again browsing onto Amazon and since it

47:23took so much time to access the data and

47:25I already switched over to a new iPhone

47:29and they are recommending me the old

47:31iPhone case for 5S so obviously that

47:33won't work I won't go there and click it

47:35because I've already changed my phone

47:37right so that will be a huge problem for

47:40Amazon the recommendation system won't

47:42work anymore and I know that RMA changes

47:45her phone every year so if she has

47:47bought a phone and people are

47:50recommending if she has bought a phone

47:51now and someone's recommending the case

47:54for that phone after 2 years years

47:56doesn't make sense to me at all yeah

47:58only it will work if I have both the two

48:01phones at the same time but yeah I don't

48:03want to waste money on purchasing new

48:05iPhone case for my old phone so

48:08basically it won't be fair if we don't

48:10discuss the solution to these problems

48:12RMA we can't leave our viewers with just

48:15the problems right it won't be fair what

48:17is the solution had doop Hadoop is a

48:20solution so let's introduce Hadoop now

48:24okay so now what is had

48:26so Hadoop is a framework that allows you

48:29to first store big data in a distributed

48:31environment so that you can process it

48:34parallel there are basically two parts

48:37one is hdfs that is Hadoop distributed

48:39file system for storage it allows you to

48:42store data of various formats across a

48:45cluster and the second part is map

48:47reduce now it is nothing but a

48:49processing unit of Hadoop it allows

48:51parallel processing of data that is

48:54stored across the hdfs now let us dig

48:57deep in hdfs and understand it better

49:01yeah so hdfs creates an abstraction of

49:04resources um let me simplify it for you

49:07so similar to virtualization you can see

49:09hdfs logically as a single unit for

49:11storing big data but actually you're

49:14storing your data across multiple

49:16systems or you can say in a distributed

49:18fashion so here you have a Master Slave

49:21architecture in which the name node is a

49:23master node and the data node are slaves

49:26and the name note contains the metadata

49:28about the data that is stored in the

49:30data nodes like which data block is

49:33stored in which data node where are the

49:36replications of the data block kept and

49:38etc etc so the actual data is stored in

49:42the data nodes and I also want to add

49:44that we actually replicate the data

49:46blocks that is present in the data nodes

49:49and by default the replication factor is

49:51three so it means that there are three

49:53copies of each file so s of can you tell

49:56us why do we need that replication sure

49:59RMA since we are using commodity

50:01Hardwares right and we know failure rate

50:03of these Hardwares are pretty high so if

50:06one of the data notes fail I won't have

50:08that data block and that's the reason we

50:10need to replicate the data block now

50:12this replication Factor depends on your

50:14requirements right now let us understand

50:17how actually Hadoop provided the

50:18solution to the big data problems that

50:20we have discussed so RMA can you

50:23remember what was the first problem yeah

50:25it was storing the big data so how hdfs

50:29solved it let's discuss it so hdfs

50:32provides a distributed way to store Big

50:34Data we've already told you that so your

50:37data is stored in blocks in data nodes

50:39and you then specify the size of each

50:42block so basically if you have a 512 MB

50:45of data and you have configured hdf as

50:48such that it will create 128 megabytes

50:51of data block so hdfs will so hdfs will

50:54divide the data in four blocks because

50:5652 divid by 128 is four and it will

51:00store it across different data nodes and

51:02it will also replicate the data blogs on

51:05the different data nodes so now we are

51:07using commodity hardware and storing is

51:09not a challenge so what are your

51:11thoughts on it sort of I will also add

51:13one thing RMA it also solves the scaling

51:16problem it focuses on horizontal scaling

51:18instead of vertical now you can always

51:20add some extra data nodes to your hdfs

51:22cluster ads and when required instead of

51:25scaling the resources of your data nodes

51:27so you're not actually increasing the

51:29resources of your data nodes you're just

51:30adding few more data nodes when you

51:32require let me summarize it for you so

51:35basically for storing one TB of data I

51:37don't need a one TB system I can instead

51:40do it on multiple 128 GB systems or even

51:43less now RMA what was the second

51:46challenge with big data so the next

51:49problem was storing variety of data and

51:52that problem was also addressed by hdfs

51:55so with htfs you can store all kinds of

51:58data whether it's structured

51:59semi-structured or unstructured it is

52:02because in htfs there is no pre- dumping

52:04scheme of validation so you can just

52:06dump all the kinds of data that you have

52:08in one place and it also follows a write

52:11once and read many model and due to this

52:14you can just write the data once and you

52:16can read it multiple times for finding

52:18out insights and if you can recall the

52:21third challenge was accessing the data

52:23faster and this is one of the major

52:26challenge with big data and in order to

52:29solve it we're moving processing to data

52:31and not data to processing so what it

52:34means sort of just go ahead and explain

52:37it yes RMA I will so over here let me

52:40explain you what do you mean by actually

52:42moving process to data so consider this

52:45as our master and these are our slaves

52:48so the data is stored in these slaves so

52:50what happens one way of processing this

52:52data is what I can do is I can send this

52:54data to my Master node and I can process

52:57it over here but what will happen if all

53:00of my slaves will send the data to my

53:02master node it'll cause Network

53:04congestion plus input output Channel

53:06congestion and at the same time my

53:08master node will take a lot of time in

53:10order to process this huge amount of

53:12data so what I can do I can send this

53:14process to data that means I can send

53:17the logic to all these slaves which

53:19actually contain the data and perform

53:22processing in the slaves itself so after

53:24that what will happen the small chunks

53:26of the result that will come out will be

53:28sent to our name node so in that way

53:30there won't be any network congestion or

53:32input output congestion and it will take

53:34comparatively very less time so this is

53:37what actually means sending process to

53:41[Music]

53:44data so who's a big data engineer now

53:47every data driven business needs to have

53:49a framework in place for the data

53:51science and data analytics Pipeline and

53:54a data engineer is the one who's

53:56responsible for building and maintaining

53:58this framework now these Engineers must

54:00ensure that there is an uninterrupted

54:02flow of data between servers and

54:04applications so in simple words a data

54:07engineer builds tests maintains data

54:10structures and architectures for data

54:12ingestion processing and deployment of

54:15large-scale data intensive applications

54:18now data Engineers work in tandem with

54:20data Architects data analysts and data

54:22scientists so they must all share comp

54:25these insights to other stakeholders in

54:28the company through data visualization

54:30and storytelling but what does a big

54:32data engineer do exactly now the most

54:35crucial part of a big data engineer is

54:37to design develop construct install test

54:40and maintain the complete data

54:42management and processing systems they

54:45are basically the ones who handle the

54:47complete endtoend infrastructure for

54:49data management and processing they

54:51build a pipeline for data collection and

54:54storage and funnel the data to data

54:56analysts and scientists so basically

54:59what they do is they create the

55:00framework to make data consumable for

55:03data scientists and analysts so they can

55:06use the data to derive insights from it

55:09know that the data Engineers are the

55:11Builders of data systems and not those

55:14who mine for insights so the data

55:16engineer Works more behind the scenes

55:19and must be comfortable with other

55:20members of the team producing Business

55:22Solutions from this data now all there

55:25responsibilities revolve around this

55:27they need to take care of a lot of

55:28things while performing these activities

55:31hence one of the most sought after

55:33skills in data engineering is the

55:35ability to design and build data

55:37warehouses this is where all the raw

55:39data is collected stored and retrieve

55:41from without data warehouses all the

55:44tasks that a data scientist does will

55:46become obsolete it is either going to

55:48get too expensive or very very large to

55:51scale now data Engineers should always

55:53keep in mind that the system which he or

55:56she builds needs to be scalable robust

55:59and fault tolerant so that the system

56:02can be scaled up without increasing the

56:04number of data sources and can handle a

56:06huge amount of heterogeneous data

56:09without any failure now imagine a

56:11situation wherein the source of data is

56:13doubled or tripled but the system cannot

56:15scale up will it not cost a lot more

56:18time and resources to build the same

56:20system again which is suitable for this

56:22kind of intake exactly this is why the

56:26Big Data Engineers have a role here next

56:29he or she is the one that handles the

56:32extract transform and load process which

56:35is basically the blueprint for how the

56:38collected raw data is processed and

56:40transformed into Data ready for analysis

56:43now you're going to acquire a lot of

56:45data from different sources how do you

56:48bring them together to one platform ETL

56:51is your answer apart from all this a

56:54data engine engineer should always aim

56:57at deriving insights by acquiring data

57:00from new sources some of the

57:02responsibilities of a data engineer also

57:04include improving data foundational

57:07procedures integrating new data

57:09management Technologies and the software

57:12into existing systems and building data

57:14collection pipelines and finally one of

57:17the major roles of a data engineer is to

57:20include performance tuning and make the

57:22whole system way more efficient which is

57:24pretty self-explanatory if you ask me

57:27now most of us have some idea about who

57:30a big data engineer is but there's still

57:33some confusion about their

57:36responsibilities now this ambiguity

57:38further increases when we gain more

57:41information about the role now let me

57:43help you debunk all your queries about

57:46it so let's talk about some big data

57:48engineer

57:49responsibilities first up we have data

57:52ingestion now this is associated with

57:54the task of getting data out of the

57:56source systems and ingesting it into a

57:58data Lake now a data engineer would need

58:01to know how to efficiently extract the

58:03data from a source including multiple

58:05approaches for both batch and real-time

58:07extraction as well as needing to know

58:10about the incremental data loading

58:12fitting within small Source windows and

58:14parallelization of data loading as well

58:17now another small subtask of data

58:19ingestion is data

58:21synchronization but because it's such a

58:23big issue in the Big Data world we are

58:26going to talk about it now since Hadoop

58:28and other big data platforms don't

58:30support incremental loading of data a

58:33data engineer would need to know how to

58:34deal with detecting changes in the data

58:37source merge and sync change data from

58:40sources into the Big Data environment

58:42next we have data transformation this is

58:45basically the T in the extract transform

58:48and load that we had discussed earlier

58:50it is basically focused on integration

58:52and transformation of data for a

58:54specific use case now a major skill set

58:56here is the knowledge of SQL as it turns

58:59out not much has changed in terms of the

59:01type of data Transformations that people

59:03are doing now compared to purely

59:05relational environments now imagine all

59:08this data that you've acquired from

59:10various sources what would you have to

59:12do to make them all palatable in the

59:14same platform you need to transform that

59:17data and this is what a data engineer

59:19does here and finally we have

59:22performance optimization which is one of

59:24the tougher areas because anyone can

59:26build a slow performing system the

59:28challenge is to build data pipelines

59:30that are both scalable and efficient so

59:33the ability and understanding of how to

59:35optimize the performance of an

59:37individual data Pipeline and the overall

59:39systems are a higher level of data

59:41engineering skill now for example Big

59:44Data platforms continue to be

59:45challenging with regard to query

59:47performance and have added complexity to

59:49a data engineer's job in order to

59:51optimize performance of queries and

59:53creation of reports BS the data engineer

59:56needs to know how to denormalize

59:58partition and index data models he also

1:00:01needs to understand tools and Concepts

1:00:04regarding in-memory models and olap

1:00:06cubes now let's quickly move ahead and

1:00:09look at the required skills to fulfill

1:00:11these

1:00:12responsibilities now we'll be going

1:00:14through these skills in a clockwise

1:00:16order so starting with big data

1:00:18Frameworks now with the rise of big data

1:00:20in the early 21st century a new

1:00:22framework was born and that is Hadoop

1:00:24all thanks to Doug cutting for

1:00:26introducing this framework it not only

1:00:28stores big data in a distributed manner

1:00:31but also processes the data parallell

1:00:33there are several tools in the Hadoop

1:00:35ecosystem which cater differently for

1:00:37different purposes and Professionals for

1:00:39a big data engineer mastering Big Data

1:00:42tools is a must some of the tools which

1:00:45you will need to Master first of all you

1:00:47have hdfs which is the storage part of

1:00:49Hadoop being the foundation of Hadoop

1:00:52knowledge of hdfs is a must to start

1:00:54working working with Hadoop framework

1:00:56next we have yarn which performs

1:00:58resource management by allocating

1:01:00resources to different applications and

1:01:02scheduling jobs Now map ruce is a

1:01:05parallel processing Paradigm which

1:01:08allows data to be processed parallely on

1:01:10top of the hdfs next we have pig and

1:01:13Hive now Hive is a data warehousing tool

1:01:16on top of hdfs which caters to

1:01:18professional from an SQL background to

1:01:21perform analytics on top of hdfs whereas

1:01:23apachi pig is a high level platform

1:01:26which is used for data transformation on

1:01:28top of Hado now Hive is generally used

1:01:31by data analyst for creating reports

1:01:33whereas pig is used by researchers for

1:01:35programming both are pretty easy to

1:01:37learn if you already familiar with SQL

1:01:40next we have Flume and scoop Flume is a

1:01:43tool which is used to import

1:01:45unstructured data to hdfs and scoop is

1:01:48used to Import and Export structured

1:01:50data from our dbms now next we have

1:01:53zookeeper which acts as a coordinator

1:01:56among the distributed Services running

1:01:58in a Hadoop environment it basically

1:02:00helps to configure management and

1:02:02synchronized services and finally we

1:02:05have Uzi which is basically a scheduler

1:02:07which binds multiple logical jobs

1:02:10together and helps in accomplishing a

1:02:12complete task next up we have realtime

1:02:15processing Frameworks now real-time

1:02:17processing with quick actions is the

1:02:20need of R either it is a credit card

1:02:22fraud detection system or a Rec

1:02:24recommendation system now imagine if you

1:02:26wanted a red dress today and Amazon

1:02:29decides to suggest it to you a month

1:02:31later now wouldn't that be completely

1:02:33useless for you in this case you need

1:02:36realtime processing it is very important

1:02:38for a data engineer to have knowledge of

1:02:41realtime processing Frameworks now

1:02:43Apachi spark is one of the distributed

1:02:46real-time processing Frameworks which is

1:02:48used in the industry rigorously it can

1:02:50be easily integrated with Hadoop

1:02:52leveraging hdfs as well next we have

1:02:55dbms now a database management system

1:02:58stores organizes and manages a large

1:03:01amount of information within a single

1:03:03software application now data Engineers

1:03:06need to understand the database

1:03:07management system to manage data

1:03:09efficiently and allow users to perform

1:03:11multiple tasks with ease this will help

1:03:14data engineers in improve data sharing

1:03:16data security data access and better

1:03:19data integration with minimized data

1:03:22inconsistencies these are the fun

1:03:24mentals that data Engineers should know

1:03:27prior to building a scalable robust and

1:03:29fall tolerance system next we have SQL

1:03:32based Technologies now there are various

1:03:34relational databases that are used in

1:03:36the industry such as Oracle DB Microsoft

1:03:39SQL Server Etc now data Engineers must

1:03:43have at least the knowledge of one such

1:03:45database now knowing SQL is also a must

1:03:49this structured query language as SQL is

1:03:52also known as used to structure

1:03:54manipulate and manage data stored in

1:03:57relational databases as data Engineers

1:03:59work closely with

1:04:01rdbms's they need to have a strong

1:04:03command on SQL now next we have no SQL

1:04:07Technologies as the requirements of

1:04:09organizations have grown Beyond

1:04:11structured data no SQL databases have

1:04:15been introduced into this environment it

1:04:17can store large volumes of structured

1:04:19semi-structured or unstructured data

1:04:21with quick iteration and agile structure

1:04:24as per application requirements some of

1:04:26the most prominently used databases are

1:04:29hbas Cassandra and mongodb now hbase is

1:04:33a column oriented nosql database on top

1:04:36of hdfs which is great for scalable and

1:04:39distributed Big Data stores it is also

1:04:41great for applications with optimized

1:04:43read and range based scan and it

1:04:46provides consistency and partitioning

1:04:48out of capap now Cassandra is a highly

1:04:51scalable database with incremental

1:04:53scalability and the best part about

1:04:55Cassandra is the minimal Administration

1:04:58and no single point of failure it's good

1:05:01for applications with fast and random

1:05:03read and writs it provides available and

1:05:06partitioning out of capap and finally we

1:05:09have mongodb which is basically a

1:05:12document oriented nosql database which

1:05:15is a schema free database it gives full

1:05:17index support for high performance and

1:05:20replication for fall tolerance it has a

1:05:23Master Slave sort of architecture and

1:05:25provides CP out of capap it is

1:05:28rigorously used by web applications and

1:05:31semi-structured data handling next we're

1:05:34going to discuss programming and

1:05:35scripting languages so various

1:05:37programming languages can serve for the

1:05:39same purpose so knowledge of one

1:05:41programming language is enough I'm

1:05:43saying this because the flavor of

1:05:44language may change but the logic

1:05:46Remains the Same if you're a beginner

1:05:48you can go ahead with python as it is an

1:05:51easy language to learn due to its syntax

1:05:53and good good Community Support whereas

1:05:55R has a steep learning curve which is

1:05:58developed by statisticians and it is

1:06:00mostly used by analysts and data

1:06:02scientists the next skill we're going to

1:06:05discuss is an important one it is ETL or

1:06:08data warehousing now data warehousing is

1:06:11very important when it comes to managing

1:06:13a huge amount of data coming in from

1:06:15heterogeneous sources where you need to

1:06:17apply extract transform and load now

1:06:20data warehousing is used for analytics

1:06:23and Reporting and is a very very crucial

1:06:25part of every business intelligence

1:06:27solution because this is the part which

1:06:30is going to take you most time now it is

1:06:32very important for a big data engineer

1:06:34to Master One data warehousing or ETL

1:06:36tool after mastering one it becomes

1:06:39pretty easy to learn new tools and as

1:06:41the fundamentals remain the same now

1:06:43Informatica click View and talent are

1:06:46very well-known tools used in the

1:06:48industry Informatica and talent Open

1:06:50studio are data integration tools with

1:06:53ETL architecture

1:06:54the major benefit of talent is its

1:06:57support from the Big Data Frameworks if

1:06:59you're new to data warehousing and ETL

1:07:01tools I would definitely recommend you

1:07:03start with talent because after learning

1:07:06this any data warehousing tools will

1:07:08become a piece of cake in finally we

1:07:10have our operating systems now intimate

1:07:13knowledge of Unix Linux and Solaris is

1:07:17very helpful as many mathematical tools

1:07:19are going to be based off of these

1:07:21systems due to their unique demands for

1:07:24root access to hardware and operating

1:07:26system functionality above and beyond

1:07:28that of Microsoft's Windows or Mac OS

1:07:31now some level of understanding of how

1:07:33to act upon this data is also very

1:07:35valuable for data Engineers for this

1:07:38reason some knowledge of statistical

1:07:39analysis and the basics of data modeling

1:07:42are also hugely valuable knowledge of

1:07:45machine learning in Cloud also will

1:07:46serve as a big plus while machine

1:07:48learning is technically something

1:07:50relegated to a data scientist knowledge

1:07:52in this area is helpful to construct

1:07:54Solutions usable by your cohorts now

1:07:57this knowledge has the added benefit of

1:07:59making you extremely marketable in this

1:08:02space as being able to put on both hats

1:08:05in which case makes you a really

1:08:06formidable

1:08:07[Music]

1:08:13tool here you can see the job

1:08:16distribution per salary range in India

1:08:19for a data engineer as we can see people

1:08:22who get paid more than5 inom are about

1:08:2533% 730 perom are 26% 8 lak 70,000 about

1:08:3220% Which are very high salary brackets

1:08:36apart from that the average salary for

1:08:38data engineer is almost 8 lakhs inim and

1:08:41for a senior data engineer is almost 16

1:08:44lakhs in anim if we look at the same

1:08:47numbers in the US there are 32% of

1:08:50professionals who make more than $90,000

1:08:53a a year and 27% professionals who make

1:08:57$105,000 a year the average salary in

1:09:00the US for a data engineer is way more

1:09:03than

1:09:04$90,000 and for a senior data engineer

1:09:07it is

1:09:09$124,000 per anom now as we have

1:09:12discussed the salary of a big data

1:09:14engineer let's look at a few factors in

1:09:17the form of skills and technology that

1:09:19they know on which their salary depends

1:09:23here we've carefully curated a table

1:09:25which lists out the skills and the

1:09:28average salary which can be encashed

1:09:30through them you can see services such

1:09:33as AWS data analysis data mining

1:09:36warehousing machine learning and even

1:09:39programming languages like Java and R

1:09:43apart from that you can see bi tools and

1:09:46statistical tools like Tablo database

1:09:49architecture ETL and structured query

1:09:52languages now another another influence

1:09:54on the salary is experience because

1:09:57experience is also a very important

1:10:00factor in deciding the Big Data engineer

1:10:03salary the distribution of salary is

1:10:06like so an entry-level data engineer

1:10:09makes about

1:10:11$85,000 a year people who have 5 to8

1:10:14years of work experience bag nearly

1:10:18$113,000 a year and people who are

1:10:20experienced I'm talking like 10 years of

1:10:22Industry experience experience get over

1:10:26$118,000 a year now this salary must be

1:10:30coming from somewhere presenting to you

1:10:33the companies that hire in this job role

1:10:36as you can see there are some very big

1:10:38names like Amazon Google Bosch Microsoft

1:10:42and IBM who hire Big Data Engineers now

1:10:46companies that hire Big Data

1:10:48professionals are companies that are

1:10:50invested in the future and the worldwide

1:10:53big data market revenues for software

1:10:55and services are projected to increase

1:10:58from $42 billion in 2018 to $13 billion

1:11:03in

1:11:042027 attaining a compound annual growth

1:11:07rate of about

1:11:0910.5% that sort of a growth needs some

1:11:12kind of work after going through

1:11:15multiple job descriptions we found that

1:11:17the Big Data engineer salary has many

1:11:21[Music]

1:11:22variables

1:11:27so let's go ahead and also understand

1:11:29the average salary of a data engineer so

1:11:31in the US the average salary of a data

1:11:34engineer is

1:11:36$133,000 but remember guys this is

1:11:39basically an average salary which

1:11:41basically means that there are people

1:11:42who are above this salary or can even be

1:11:44a person who is below this salary okay

1:11:47but later in this session I'm also going

1:11:49to talk about some of the job

1:11:51description which can tell you the kind

1:11:52of salary that you can earn once you

1:11:55start applying for a data engineer

1:11:56profile in India the salary is around

1:11:596.5 lakhs per anom or a 7 lakh perom and

1:12:03the same goes for an Indian jobs as well

1:12:05that the salary can be higher than this

1:12:07average package or it can be lower than

1:12:10the package as well now that you have

1:12:12seen the basic or an average scale of a

1:12:15aor data engineer let's go ahead and see

1:12:18the job description of an aor data

1:12:20engineer all right guys so thinking

1:12:22again about what a data engineer does

1:12:25and aor data engineer is responsible for

1:12:28Designing implementing and maintaining

1:12:30data management and data processing

1:12:32system on the Microsoft Azor Cloud

1:12:35platform now they work with large and

1:12:37complex data sets and are responsible

1:12:39for ensuring that data is stored

1:12:41processed and secure efficiently and

1:12:43effectively now when you basically

1:12:45Define a job role like that you have

1:12:47many jobs description which are floating

1:12:49in the market right now how can you

1:12:52identify which job description you have

1:12:55to apply to now let's go ahead and

1:12:57understand that so talking about the job

1:13:00description which basically exists in

1:13:02the market guys you will see a job

1:13:04description which is an entry-level job

1:13:06description and then on the next slide I

1:13:09will show a job description which is a

1:13:11mid or a senior level job description

1:13:13all right so heading back to the entry

1:13:15level so if you see the first job

1:13:17description which is an aor data

1:13:19engineer the salary is anywhere around 6

1:13:22lakh perom to4 14 lakhs perom and these

1:13:26are the skills that you require now as I

1:13:29mentioned before these are the expected

1:13:31skills required for an aor data engineer

1:13:35now what are the skills which are

1:13:37expected now you are expected to know no

1:13:39SQL or a cosmos database skill now they

1:13:42are expected to know data Lake data

1:13:45factory data warehouse and strong

1:13:47experience in building pipelines in Azor

1:13:49data or in aor data Lake then you should

1:13:52be able to analyze and understand

1:13:54complex data you should be able to

1:13:56understand business requirement and

1:13:58actively provides input from data

1:14:00perspective now at the same time if you

1:14:02look at these skill sets these skill

1:14:04sets are all the scales at which you

1:14:06will basically know after you study for

1:14:09or clear the aor data engineer

1:14:11certification so once you're done with

1:14:13the certification once you are done with

1:14:14the skill sets which is just there in

1:14:16the certification you can easily go

1:14:18ahead and apply for a job that lies in

1:14:20the salary range which is for an zor

1:14:22data engineer

1:14:23now talking about a mid senior level

1:14:25data engineer profile the list is quite

1:14:28long as you can see now you can see that

1:14:31over here apart from all these skill

1:14:32sets a lot of other things are also

1:14:34mentioned here as well for example you

1:14:36should have some four or five plus years

1:14:38of experience in implementing or

1:14:41designing solution using Azor Big Data

1:14:43Technologies then you should have an

1:14:45experience with an Hands-On in Azor data

1:14:47factory aor devops aor data Lake storage

1:14:50Etc now you should have a knowledge of

1:14:53big data pipeline then design and build

1:14:55modern data pipelines and maintain the

1:14:57data warehouse schematics layouts

1:14:59architecture and relational or

1:15:01non-relational database for data access

1:15:03and advanced analytics now you should

1:15:06also know Java jQuery SQL or Scala or

1:15:10any preferred programming language right

1:15:12so over here what we recommend to our

1:15:14Learners is you should go ahead and

1:15:15Learn Python because although they have

1:15:17mentioned only these programming

1:15:19language but companies are very much

1:15:21flexible if the target profile has any

1:15:23programming experience but a strong one

1:15:25in any of the programming languages

1:15:27right next thing that they expect you to

1:15:29know is Advanced skill using one or most

1:15:32common language for example like python

1:15:34batch Etc so this python will basically

1:15:38serve as a dual purpose that is all a

1:15:40scripting language and a programming

1:15:42language as well the next thing that

1:15:44they expect you to know is the ETL

1:15:46process using big data Technologies such

1:15:49as spark Kafka Hadoop and others now the

1:15:52next thing that they they expect you to

1:15:53know is the ETL process using big data

1:15:56technology such as spark Kafka doops and

1:16:00others and then you need to understand

1:16:01the data visualization experience using

1:16:04python python here is a plus guys so you

1:16:07need to understand this and then you

1:16:08actually need to learn tableu or a

1:16:10powerbi any one of the tools or

1:16:13technology is a plus point so you either

1:16:15have to have a skill on tblo or a

1:16:19powerbi so this again is an important

1:16:21skill to have and this again coincide

1:16:24with what a data analyst does right

1:16:27because he is also responsible for data

1:16:29visualization to some extent then you

1:16:31have a solid understanding and

1:16:33experience implementing cloud data

1:16:34platform in Microsoft aor devops so if

1:16:38you're a guy who wants to start off you

1:16:39can start off with the freshia profile

1:16:42and after having experience in the

1:16:43freshia profile and learn all the skill

1:16:45sets like big data zor these are the

1:16:48skill sets that if you gain you can

1:16:50actually apply to the senior or midlevel

1:16:52in particular particular okay so Guys

1:16:54these are the few job description which

1:16:56are related to the data engineer profile

1:16:59now that we are clear with why who and

1:17:01what are the career opportunities and

1:17:03salary of an aor data engineer we shall

1:17:05see the path towards becoming an aor

1:17:08data engineer so first of all you will

1:17:10have to talk about the different aor

1:17:12storage which are out there now you'll

1:17:14have to learn about these like the blob

1:17:16storage table storage file storage and

1:17:19the que storage now you will have to

1:17:21learn about the relational database

1:17:22option options as well which are there

1:17:24in the market such as the SQL database

1:17:27SQL DB Warehouse analysis Services now

1:17:30you will also have to learn about no SQL

1:17:33you will have to learn about big data

1:17:35services in Azor like data Lake

1:17:37analytics data Lake storage now you also

1:17:40have to learn about the data factories

1:17:42aor function stream analytics iot hubs

1:17:45even hubs Etc now apart from that you

1:17:48will also need to learn redist cache and

1:17:50aor search so these are the servic that

1:17:53are basically required for you to

1:17:55understand in order to clear the

1:17:56certification and go ahead and become an

1:17:59aor data engineer now these are also

1:18:01Conde with the job description that we

1:18:03had a look earlier right now all the

1:18:05services which are mentioned there is

1:18:07basically a part of what you have

1:18:08learned in order to correct the

1:18:10certification now apart from this we

1:18:12also recommend going through the open

1:18:14source Services of a doop such as spark

1:18:18hi and the Hadoop itself right now let

1:18:21us go step by step to reach our goal in

1:18:23becoming an aor data engineer so in

1:18:26order to become an aor data engineer you

1:18:28will also need to have a strong

1:18:29foundation in data engineering and cloud

1:18:32computing here are some steps you can

1:18:34take to develop the skills and knowledge

1:18:36needed for a career as an aor data

1:18:38engineer so the first one is learn the

1:18:41basics of data engineering now in order

1:18:44to become an Azor data engineer you

1:18:46should first develop a strong foundation

1:18:48in data engineering Concepts such as

1:18:50data modeling data pipelines data

1:18:52processing and data storage now you can

1:18:55learn these Concepts through online

1:18:57courses or books or by working on

1:18:59practical projects now the second is

1:19:02learning or programming language now as

1:19:05an Reser data engineer you will need to

1:19:07proficient in at least one programming

1:19:09languages as I've already mentioned

1:19:11earlier now python is a very popular

1:19:13choice for data engineering but you can

1:19:15also consider learning languages such as

1:19:17Java C or Scala now the next step is you

1:19:21need to learn an SQL now as a data

1:19:23engineer you will be working with large

1:19:25amount of data you've already known that

1:19:27since the name itself as an aor data

1:19:30engineer so you will need to be a

1:19:31proficient in SQL to extract and

1:19:34transform data now the next thing you

1:19:36need to learn is Azor data storage

1:19:38options now as you know Azor offers a

1:19:41range of data storage options such as

1:19:43Azor blob storage Azor data L Storage

1:19:46Azor Cosmos DB Azor SQL database and

1:19:49Azor signups analytics formally SQL

1:19:52database warehouse now you should

1:19:54familiarize yourself with the features

1:19:56and capabilities of these storage

1:19:58options now the next thing you need to

1:20:00learn is Azor data processing

1:20:02Technologies now Azor provides several

1:20:04Technologies for processing data such as

1:20:06Azor stream analytics Azor data bricks

1:20:09Azor data Factory and Azor HD insights

1:20:12you should learn how to use these

1:20:14Technologies to build data pipelines for

1:20:16ingestion transforming and processing

1:20:19data so the next thing you need to learn

1:20:21is aor data management and and security

1:20:24so you as an Azor data engineer should

1:20:26learn about the Azor tools for managing

1:20:28and securing data such as Azor data

1:20:31catalog aor data link security and aor

1:20:34private link so what is the next step

1:20:37the next step is get certified as you

1:20:40know getting certified is very crucial

1:20:42and it's very important in today's

1:20:45generation because all the organization

1:20:48as I mentioned earlier if I just have to

1:20:49go back in my slide I will show you the

1:20:51job description where they have actually

1:20:53mentioned that you need to actually pass

1:20:56the certification all right this is the

1:20:59certification you required that is a dp2

1:21:01200 and a DP 2011 but as for now you

1:21:04don't need dp2 200 or DP 2001 you only

1:21:08have to give one exam which I'll be

1:21:10further talking about it that is the

1:21:12dp23 all right guys so moving ahead

1:21:15again so you need to earn an aor data

1:21:18engineer associate certification by

1:21:20passing the dp23 exam

1:21:23right now what is the next thing you

1:21:25need to do is you need to gain practical

1:21:27experience now the best way to learn aor

1:21:29data engineering is by working on

1:21:31practical projects now we all know this

1:21:34right now you can find data engineering

1:21:36projects on online platform such as

1:21:38keigle or you can work on projects in

1:21:40your own organization or you can enroll

1:21:43with edura and they provide you a tons

1:21:45and tons of Life projects which are

1:21:47developed by the instructor who have

1:21:50already worked as a data engineer in

1:21:52their their organization now the next

1:21:54thing you need to learn is continue

1:21:56learning now as a data engineer you will

1:21:58need to keep your skills and knowledge

1:22:00up to date as Technologies and best

1:22:02practices evolve right now making sure

1:22:05to stay current by learning about new

1:22:06aor data engineering features and

1:22:08participating in professionals

1:22:10development activities all right guys so

1:22:13now let us understand the things that

1:22:14you need to know about this

1:22:15certification that I've just talked

1:22:17about that is your dp23 so guys if

1:22:20you're applying for an aor data engineer

1:22:22associate

1:22:23there used to be two exam that you have

1:22:25to give which I've just mentioned now on

1:22:27the screen in the previous slide right

1:22:29one is implementing an Azor data

1:22:31solution and the next exam is design an

1:22:34Azor data solution now after clearing

1:22:36both these exam it will give you the aor

1:22:38data engineering associate certification

1:22:41and these two exam basically have the

1:22:42code like I've mentioned dp2 200 and DP

1:22:462011 right but now this exam have been

1:22:49retired on 23rd February 2021 now you

1:22:53guys just have to give one exam to clear

1:22:55the data engineer certification and this

1:22:58exam is

1:22:59dp23 now what they have done is they

1:23:02have Club both these exam and they have

1:23:04now included its syllabus in just one

1:23:06examine they're asking questions from it

1:23:09right so earlier what you have to do is

1:23:11you have to pay for two exams then

1:23:13prepare for two exams and give them and

1:23:15then only you can get the certification

1:23:18but now just by passing one exam you can

1:23:20clear out the data engineer

1:23:22certification ation right now as part of

1:23:24the new exams now the skill sets have

1:23:26been updated as you can see on the

1:23:28screen so basically these are the

1:23:30distribution of topics and they'll be

1:23:32covered in this exam now first of all

1:23:34you will be asked most of the questions

1:23:36on design and Implement data storage now

1:23:39this is going to have 40 to 45% of

1:23:42weightage okay now I'm going to tell you

1:23:45what are the topics that comes under the

1:23:47design and Implement data storage so

1:23:49make sure you write it down okay guys so

1:23:51the first is design a data storage

1:23:53structure then the second is design a

1:23:55partitioning strategy then the next

1:23:57question is design the serving layer

1:24:00then you have the Implement physical

1:24:02data storage structures then Implement

1:24:04logical data structure and lastly

1:24:06implement the serving layer now after

1:24:09this you will have design and develop

1:24:12data processing now this is of 25 to 30%

1:24:15weightage now there are four points

1:24:17coming under this design and develop

1:24:19data processing now the first is you

1:24:20need to learn about the inest and

1:24:21transform data second is design and

1:24:24develop a batch processing solution the

1:24:27third is design and develop a stream

1:24:29processing solution and lastly we have

1:24:31the manage batches and pipelines all

1:24:34right so coming to the third is design

1:24:37and Implement data security which is of

1:24:4010 to 15% so there are two main topics

1:24:43here that is design security for data

1:24:45policies and standards and the second is

1:24:47Implement data security Now talking

1:24:50about the last is you have the Monitor

1:24:52and optimize data storage and data

1:24:54processing which is again of 10 to 15%

1:24:57now even here also we have two main

1:25:00topics that you need to be covered that

1:25:01is the monitor data storage and data

1:25:03processing and the second is optimize

1:25:05and troubleshoot data storage and data

1:25:08processing now if you are with me till

1:25:10at this point you shall have understood

1:25:12what are all the things that you have to

1:25:13learn in order to clear the exams and

1:25:16become an aor data engineer right since

1:25:19we have already discussed what are the

1:25:20path and what are the skills to learn

1:25:22learn to prepare ourself so like I said

1:25:24earlier you now just have to give this

1:25:26dp23 exam to get the Microsoft

1:25:29certification associate data engineering

1:25:31certification okay so just one exam and

1:25:34now you will get the certification for

1:25:36it all right so now let's move on guys

1:25:39and now let's talk about how you guys

1:25:41can get started in clearing this exam

1:25:43and developing those skills and go ahead

1:25:46and apply for the job and become a

1:25:47successful aor data engineer so we have

1:25:51mentioned a lot of things that you have

1:25:52to learn but how exactly you should go

1:25:55forward and start learning these things

1:25:57let's go ahead and clear that out so

1:26:00first of all guys what we can do for you

1:26:02is you can basically refer to a lot of

1:26:04blocks that we have written on or we

1:26:06frequently updat videos on YouTube as

1:26:09well such as this video which has info

1:26:11about the engineering certification or

1:26:13how to become an Azor data engineer now

1:26:16we frequently put more videos for such

1:26:18topics now you can go through them and

1:26:20basically get a jump start into how you

1:26:22you can prepare now my recommendation to

1:26:25you will be to plan out your working

1:26:27plan ask after or before working shift

1:26:30of yours now you should at least spend 3

1:26:32to four hours every day for the next two

1:26:34or 3 months in order to clear this

1:26:37certification exam otherwise if you

1:26:39don't invest this much amount of time

1:26:41guys it is going to be very difficult to

1:26:44crack the exam because there are a lot

1:26:46of things to learn especially if you're

1:26:48not from a data engineer domain now it

1:26:50is going to be a little difficult

1:26:52because you will have to read

1:26:53documentation you'll have to do Hands-On

1:26:56right now at some point you will get

1:26:58stuck and you'll have some issues you

1:27:00have to figure out what went wrong or

1:27:02what's wrong with the hands on I know

1:27:04this because I've also been there guys

1:27:06but getting stuck is the most beautiful

1:27:09part of learning anything because that

1:27:11is where you start the actual research

1:27:13and about how things work right so spend

1:27:17at least 2 to three hours every day

1:27:18either after your workshift or before

1:27:21your workshift to basic learn these

1:27:22Technologies right and now for those

1:27:25people who feel like they do not have

1:27:26the time or they do not want to invest

1:27:28time in researching and they want

1:27:30someone to help them out in getting this

1:27:32exam cleared and become a successful AER

1:27:35data engineer so guys we at a directa

1:27:38also offer a course of Microsoft Azor

1:27:40certification training of the Azor data

1:27:42engineer associate certification course

1:27:45and we also have a master program as

1:27:47well all right which will basically help

1:27:49you in clearing the aor data engineer

1:27:51associat certification right now if you

1:27:54need a helping hand and you need someone

1:27:56or you need to be taught by someone

1:27:57who's already cleared this examination

1:27:59and is already working as a data

1:28:01engineer in the industry then this is

1:28:04the right course for you

1:28:06[Music]

1:28:10guys so why as dat Factory because again

1:28:13we know that we have been generating

1:28:16data at an exponential rate especially

1:28:18since last 5 years and in 2015 we were

1:28:22generating data at the rate of almost

1:28:233.4 exabytes per month and now we are

1:28:27generating at the rate of almost more

1:28:29than 44 exabytes of data per month but

1:28:33that's almost 15 times increase in the

1:28:35amount of of data in just a span of 5

1:28:39years and it is going to be increased by

1:28:41almost 30% more because of the current

1:28:44lockdown fa by all the entire globe and

1:28:46that's why the entire consumption of

1:28:48data is again has been increased

1:28:51tremendously right that is that exactly

1:28:53is what we have the that's why we need

1:28:56to have the most optimized solution for

1:28:58data driven Solutions out there

1:28:59especially on the cloud computing

1:29:01platforms right and mod and modern data

1:29:04handling requires us to move from on

1:29:06promise to database to Cloud database

1:29:09services and that to quickly and then

1:29:11here we have to make sure again and

1:29:13that's why the data needs processing and

1:29:16goes through a series of steps making

1:29:18the process tedious because again when

1:29:20we are trying to load the data if we are

1:29:22trying to migrate other data from our

1:29:24database servers from our on premise

1:29:26storage Services then that has to go

1:29:28through a series of steps and that makes

1:29:30the entire process much more complicated

1:29:32it makes it much more slower as well and

1:29:35plus since it requires a good amount of

1:29:37investment both in times of both in

1:29:39terms of time and money it becomes we

1:29:41can see not feasible at all and data

1:29:44Factory simply help us in automating

1:29:46this entire process and does serve the

1:29:48cost that exactly is why we have data

1:29:51Factory

1:29:53now what exactly is data Factory what

1:29:55exactly is data Factory here so data

1:29:57Factory is basically a cloud-based

1:29:59integration service through which we can

1:30:03Define the entire we can design the data

1:30:05the entire workflow of data in cloud and

1:30:08making sure that we we can use for

1:30:11orchestration and for the automation

1:30:13purposes so for example suppose if we

1:30:15have now if you want we can create a

1:30:17complete pipeline we can Define what is

1:30:19a source and how the pipeline should be

1:30:21structured how the data is coming in

1:30:24where to where it should be stored and

1:30:26that to on a regular manner so we can

1:30:28Define the entire Pipeline and we can

1:30:30automate the entire data movement that

1:30:32means if we make any changes to any

1:30:34particular file or any objects in the

1:30:36oromis that will be automatically

1:30:38replicated or we can say moved into the

1:30:41cloud services as well by the help of

1:30:43data

1:30:44Factory right and using data Factory

1:30:47here we can create and schedule the

1:30:48entire data workflows called as

1:30:50pipelines that can inest data from

1:30:52disparate s data source for example if

1:30:54we have multiple data source defined

1:30:56here we can simply injest data from

1:30:58multiple sources and then we can process

1:31:00by simple a single pipeline it is

1:31:03basically used for processing large

1:31:05amount of large volume of data and these

1:31:09are done by using by integrating it with

1:31:11services that we have SEO SG inside Hado

1:31:14spark SEO data L analytics and is your

1:31:17machine

1:31:18learning so basically if we are looking

1:31:20to make sure that we have we have the

1:31:23optimized data available and optimized

1:31:25stream of data available for analytics

1:31:28and for machine learning then we have to

1:31:30do that by using data Factory to

1:31:31maintain that

1:31:33consistency and if you talk about the

1:31:35entire workflow here first of all the

1:31:37pipelines are all datadriven workflows

1:31:40in SEO data Factory typically performs

1:31:42the following four steps it simply help

1:31:44us in connecting and collecting the

1:31:46entire data set so first of all if we

1:31:48have multiple sources if you have wide

1:31:51variety of sources here we have to make

1:31:53sure we are able to set the source for

1:31:55each and every each and every

1:31:58connector we have to make sure that we

1:32:00are able to connect you to multiple

1:32:03services and then we can through

1:32:05multiple apis or if you have multiple

1:32:07sources of data says for example if we

1:32:09have data from our own CRM coming in if

1:32:11we have data from our own Erp tool from

1:32:13multiple social media analytic PL

1:32:15dashboard here we have to connect and

1:32:16collect all the different data sets then

1:32:19we have to transform and enrich because

1:32:21data always consists of multiple

1:32:23anomalities right they may be multiple

1:32:26missing values they will be some

1:32:28incorrect data formats available they

1:32:30the the data may be incomplete or it may

1:32:33be maybe multiple mistakes in that data

1:32:35corre so you have to make sure that we

1:32:37take care of entire data transformation

1:32:39that means if you want to convert the

1:32:40data format from one to the other like

1:32:42we have the ETL so again we can take

1:32:44care of the transformation and the

1:32:47enrichement of data that means data

1:32:49pre-processing and making sure it is

1:32:51much clear for and it is usable by the

1:32:54antical platforms for which we want to

1:32:55use it then we have to publish it to

1:32:58make it usable and then we have to

1:33:01continuously monitor it so that the

1:33:03entire process can be monitored in case

1:33:05there have been some breakdowns in any

1:33:07of the Clusters from which we are

1:33:09collecting data then that can be

1:33:10monitored and if something can be done

1:33:12it can be it can be processed as well we

1:33:15can do that now let's understand each

1:33:18and every concepts for data Factory here

1:33:20so if you talk about the enti Concepts

1:33:22when first of all we do have pipeline we

1:33:24do have something as pipeline so

1:33:26pipeline is a logical grouping of

1:33:28activities for example Suppose there are

1:33:3110 different sequence like we discussed

1:33:33first of all we have to connect then we

1:33:35can transform then we have to process it

1:33:37then we have to make it available for

1:33:39the analytical platforms so there if

1:33:41there are multiple sequences that needs

1:33:43to be followed again that then that is

1:33:45something that we Define as a part of

1:33:47pipeline alog together and then we have

1:33:49data sets so obviously without data data

1:33:51sets the entire data pipeline is of is

1:33:54of no use so data set simply represent

1:33:56the data structure within the available

1:33:59data store so there have multiple data

1:34:01stores available what exactly those are

1:34:03we are going to discuss step by step and

1:34:06then we have activities for activity

1:34:08simply represents a processing step in a

1:34:10pipeline for example there are 10

1:34:12different steps here for example we have

1:34:14to first of all connect to the source

1:34:16then we have to work on transforming

1:34:18data Exel then we have to work on

1:34:20pre-processing then we have to work on

1:34:22setting up the connection so each now

1:34:24this entire process itself is called as

1:34:27pipeline where we it consists of

1:34:28multiple sequence steps and each and

1:34:31every individual step here each and

1:34:32every individual step these are termed

1:34:34as activities so a pipeline is what we

1:34:36can say pipeline is simply a collection

1:34:38of different activities so we can so

1:34:41that activity is focused on completing

1:34:43one different task and pipeline simply

1:34:46defines the structure or we can say

1:34:48sequence of those STS and it simply

1:34:50makes sure that that sequence is

1:34:51followed whenever the entire pipeline is

1:34:53being

1:34:56implemented so let's erase this

1:34:59up and next we have Link services so it

1:35:03simply the information needed to connect

1:35:04to the external sources like for example

1:35:06we have apis if you have third party

1:35:08vendors and we have third party sources

1:35:10then again we do need an active APS for

1:35:12that and that's why these are all termed

1:35:14as link sources from which we can Source

1:35:16the entire assets here these are all

1:35:18additional links

1:35:20available and as we know data Factory is

1:35:23basically used for on premise itself if

1:35:25we are looking to to connect this if we

1:35:27are looking to connect this to multiple

1:35:29on premise data sets here then that

1:35:30exactly is what we use it for let's

1:35:33understand what exactly is data L

1:35:35service so data L as we know data l so

1:35:38data L as you know is simply an

1:35:40Enterprise wide hyperscale repository so

1:35:43here okay it is simply an Enterprise

1:35:44wide hyperscale repository for big data

1:35:46analytics workload and now it simply

1:35:49holds now it has a capability of

1:35:52petabyte so it can hold data for any

1:35:55size it will allow us to do again using

1:35:57that we can we can do multiple

1:36:00operational and exploratory analytics as

1:36:02well for example if we have multiple

1:36:04sources of data so for example if we

1:36:06have data sources like from on premise

1:36:08we have sensors data like for example if

1:36:10we when we talking about aviation

1:36:12industry then we have multiple we have

1:36:14tons of data sets coming it from

1:36:15different sensors especially for and

1:36:17same way for M for any manaing sectors

1:36:20as well same way if we have any data

1:36:22connected for any websites for example

1:36:24we are talking about any so any uh any

1:36:27stream of data coming in for analytics

1:36:29for any special websites for example

1:36:31suppose we have any big e-commerce

1:36:32solution for Amazon flip cart Airbnb so

1:36:35they have tons of data available on the

1:36:37platforms if we had data coming in from

1:36:39different devices from be it can also be

1:36:42a part of the I complete iot networks

1:36:45they can be data in the format of videos

1:36:47multiple social social media streams

1:36:49coming in for example we have streams of

1:36:51Facebook on Twitter redit on Reddit on

1:36:55multiple platforms so if you have

1:36:56multiple social media streams coming in

1:36:58in the formats of post or if you or

1:37:00let's say if you want to understand the

1:37:02real time I can say if you use case is

1:37:05to work on studying the user sentiments

1:37:09right then that in that case we have to

1:37:10make sure we are we are pitching in the

1:37:13multiple social media streams coming in

1:37:15if we have data in the format of images

1:37:17on application then these all Concepts

1:37:19have to be these all sources needs to be

1:37:21connected as a part of data Factory

1:37:24right and that's why here we can in now

1:37:26once we have these data sources

1:37:28available then we can connect it to ad

1:37:30analytics we can connect this to HG

1:37:32insight to R spark and machine learning

1:37:35purposes now if you have been aware if

1:37:38we have been aware of the fact such as

1:37:41now we can say data warehousing we can

1:37:43say data Lake works like a data

1:37:45warehouse for example if you're working

1:37:47on the realtime analytics right and if

1:37:50you have data sources from multiple

1:37:51platforms so instead of connecting each

1:37:53and every Source manually or we can say

1:37:55one by one to spark we can store the

1:37:57entire data and a send a centralized

1:37:59location and then we can we only need to

1:38:01connect a center location with a single

1:38:04connection to spark that's it so it

1:38:06simply improvises the entire entire

1:38:09performance as well just like we have

1:38:10red shift available in awss same way

1:38:14here we have data Le now let's also

1:38:17understand multiple data L Concepts

1:38:19let's understand multiple data Le

1:38:20Concepts here so data Lake as we know

1:38:23again has multiple components inside it

1:38:25like we have analytics now in data Lake

1:38:28components we have anal we have

1:38:29components for analytics such as a

1:38:31Insight we have AO data L available we

1:38:35can use data store for the complete

1:38:37storage purposes or for analytics we can

1:38:40indicate this with a inside and then we

1:38:42have hard inside and then we have AO

1:38:45data

1:38:46available and then when we are trying

1:38:49again in terms of Analytics we have to

1:38:51make sure we are we do remember these

1:38:53three main key points analytics we can

1:38:55do on data of any size there is no

1:38:58limitation users all users are

1:39:01productive on day one and then we have

1:39:04to make sure that ready it is all scale

1:39:07up exactly as per our Enterprise

1:39:09requirement we have to make sure of that

1:39:11part in terms of type of data stored

1:39:14here it supports all data types we have

1:39:16structured semi-structured and

1:39:18unstructured data types supported we

1:39:21have CSC files we have XML files right

1:39:23we have the emails we have Jon's files

1:39:25right so again these all are a part of

1:39:28sem instructure data when we don't have

1:39:30a direct format available on which we

1:39:32can start and perform the entire sorting

1:39:35and that is example for semi structure

1:39:36data set and then we have unstructured

1:39:39data set like we have images videos

1:39:42audio clips these all are part of semi

1:39:44structured data

1:39:46sets and then if you compare data League

1:39:49to Data Warehouse if we compare data

1:39:52leak to Data Warehouse here again data

1:39:54leak as you know is simply complimentary

1:39:57to the data warehouse whereas data where

1:40:00if you talk about the data warehousing

1:40:02service it may be soured to data leag

1:40:04data leag is basically used for detail

1:40:06data whereas as we can see data

1:40:08warehouse is basically used for filter

1:40:11summarize and for refinement of data all

1:40:13together data L offers schema on read

1:40:16whereas data house off for schema on

1:40:18right write and data L has one language

1:40:20to process data of any format whereas in

1:40:23data warehouse we can process it using

1:40:25the SQL complying it all together now

1:40:28let's move into the handson here and see

1:40:30how exactly a this is implemented on top

1:40:33of Edo data Factory portal here data

1:40:36warehouse and dat let's understand this

1:40:38by simple use case here and for doing

1:40:41that let's open up our notepad here for

1:40:45example let's say we have multiple data

1:40:47sources for example we have data sources

1:40:49available from our CRM correct for

1:40:51example our main use cases we want to

1:40:53perform the analysis for sales report

1:40:55for sales for any company correct we are

1:40:57here to perform analysis or multiple St

1:40:59data here right or suppose if we have

1:41:01data coming from from CR and we have

1:41:03data from Erp tools as well right and

1:41:06then we have data from the old data set

1:41:09available for example if we have the own

1:41:11CC file also stored here so if you have

1:41:13multiple sources of data coming in then

1:41:16in in if you want to Club it if you want

1:41:18to use this on top of any iCal platform

1:41:22so there are two ways of doing that

1:41:23either we have to connect each and every

1:41:25source with the antical platform and

1:41:28then only we can start working on it

1:41:30correct using the concept of data

1:41:33warehouse what we can do we can connect

1:41:34each and every Source we can store the

1:41:36the data from each of source to a

1:41:38centralized location as in called as a

1:41:41data house right and once the data is

1:41:43available in data house then we can

1:41:46connect that data warehous directly to

1:41:48us through a single collection to our

1:41:50artical platform so that whatever

1:41:51analyst we are trying to perform the

1:41:53entire process can be streamlined as a

1:41:55part of data warehousing

1:41:57service let's erase this

1:42:02up for getting started on data Factory

1:42:05here we can open we can come back to a

1:42:08portal so here we can come back to a

1:42:10portal in case you have not signed up on

1:42:12asure we can open this entire portal

1:42:14here where we can start working on top

1:42:17of data Factory one one by one now for

1:42:20getting started here what we can do is

1:42:21we can simply come back we can simply

1:42:23come back

1:42:25and first of all for having the for

1:42:28having the entire database that we can

1:42:30install and can interpret locally we can

1:42:32use this we can use entire platform

1:42:35called as SS SMS in case you don't have

1:42:38the enre setup done at your end then we

1:42:40can simply go ahead and download the SQL

1:42:42Server management studio in case you

1:42:44don't have the setup done now once we

1:42:46have configured these now what we can do

1:42:47we can come back to our a portal now in

1:42:49this s portal what what we can do we can

1:42:51simply first of all once we are into

1:42:53dashboard here we have to work around

1:42:55with first of all setting up the entire

1:42:57data warehousing service so what we can

1:42:59do is here we can set up the entire dat

1:43:01housing service and for setting it up we

1:43:03can come back to our entire dashboard

1:43:05now for creating a new resource we here

1:43:07we have to click on this option which

1:43:08says create a resource and then we can

1:43:11choose resource type here so from

1:43:13databases here we can choose now yeah

1:43:15these are the most commonly used

1:43:17resources here for example if you want

1:43:18to go for web application for functional

1:43:20applications for SQL databases we can

1:43:23choose accordingly and if you want to

1:43:25start with SQL database we can open up

1:43:27SQL database and then we can set up the

1:43:29entire database platform but again in

1:43:31here our main goal is not to create a

1:43:33SQL database but to set up the entire

1:43:35data with housing first and then we can

1:43:38saish in data services and then get

1:43:41started on top of it right so here what

1:43:43we can do here e either we can go ahead

1:43:45and search for the services using these

1:43:47categories or we can use it directly we

1:43:50can use the the service bar to search

1:43:52for the services directly for example

1:43:55here we can use data

1:43:59L as you can see currently we are

1:44:01planning to use data l so here we can

1:44:03choose data L

1:44:06here now if you want to create it now

1:44:08here we can simply click on create here

1:44:10we can Define the name of the current

1:44:12data storage service that we are going

1:44:13to create we can Define the entire

1:44:18name for example let's say we have we

1:44:20call it as AA app itself here we can

1:44:24choose subscriptions so the currently

1:44:26here either we can go for two type of

1:44:28subscription here we can choose pay as

1:44:29you go or here we can choose the the fre

1:44:31file if in case we have fre file

1:44:33available then we can choose that then

1:44:35here we can choose Resource Group in

1:44:37case we don't have the resource Group

1:44:39created we can use a new Resource Group

1:44:43here let's enable the entire Services

1:44:46there so that we can so we can enable

1:44:48the entire subscription model here

1:44:51so let's enable the IND subscription

1:44:53model here we can verify the

1:44:56code in case we haven't set it up we can

1:44:59simply set it up you by simply

1:45:01specifying our entire

1:45:04details set it up so that we can use it

1:45:07for creation of our enre data data

1:45:09league and data warehousing

1:45:11platform even though once you if we are

1:45:14siging this up for the first time here

1:45:16once we and once we sign it up we will

1:45:19be charge a nominal for a normal fee for

1:45:22making sure the entire account is well

1:45:25authenticated that we can use as a part

1:45:27of signing up for as your platforms and

1:45:29once it is done this will take us back

1:45:31to the D

1:45:34dashboard now once we have enter the

1:45:36account here we can simply click on

1:45:38create resource and here we can choose

1:45:40data L for example let's say we are

1:45:42starting to we are planning to start

1:45:44with the by setting up entire data

1:45:46warehouse first right so ear it was

1:45:48called as data warehouse which was now

1:45:50ch is to data leag itself now we can

1:45:52open up da data leag account all all

1:45:55together so here we can go for data leag

1:45:57generation here we can click on create

1:45:59here we can Define the entire resource

1:46:01name let's say we name it ASA here we

1:46:04can choose the subscription model that

1:46:06we want plan to go for here we can

1:46:08choose Resource Group Resource Group are

1:46:09simply like when we are planning for

1:46:12creation of now when we are creating 10

1:46:14different resources in insure and at the

1:46:16end we want to Simply segregate them

1:46:18based on a certain project for example

1:46:20we want to we want to know which all

1:46:23services have been deployed for which

1:46:25particular project and if you want to

1:46:27see a consolidate billing for that

1:46:29project then we can go for Resource

1:46:31Group for example let say here we are

1:46:32going to create a resource Group for app

1:46:35one we can choose any available Resource

1:46:37Group name that you want to go ahead

1:46:41with all right and here we can go for p

1:46:44ASO model we can choose the encryption

1:46:46if you want to enable encryption we can

1:46:48choose the keys or we can simply choose

1:46:49the default keys from

1:46:52ad suppose let's say we choose a key

1:46:54from the key generation itself we can

1:46:57configure this

1:47:00up it may take a couple of second for

1:47:02this entire data to be

1:47:09created and as you can see it says the

1:47:12deployment in in progressor so it may

1:47:14take a couple of of minutes for this

1:47:15entire data to be processed here step by

1:47:17step and once it is processed we would

1:47:20be able to see this entire in live

1:47:21action and in the meantime in case you

1:47:23want to go ahead and set up the entire

1:47:25SMS that we have discussed we have to

1:47:27have the s SMS so that we can connect

1:47:29our local SQL databases and then we can

1:47:32import directly into the a

1:47:34portal and then we also have to set up

1:47:37entire data Factory for setting up the

1:47:39entire data Factory here we can open up

1:47:40the the data Factory servers available

1:47:42Ino platform here we can click on ADD

1:47:46and here we have to define the data

1:47:48Factory name for example let's say we

1:47:50want to call us as

1:47:53Eda DF as in data Factory here we can

1:47:56choose the current version so basically

1:47:58earlier we have been using V1 but now

1:48:00since last year we have been using V2 so

1:48:03here we can choose the subscription

1:48:04model that we have subscribed for our

1:48:06for we can choose Resource Group that we

1:48:09have already created by the name of Eda

1:48:11app so that if we have 10 different

1:48:13Services if we have 40 different

1:48:15resources being deployed then we can

1:48:17again all of those resources are mapped

1:48:19for single application then we can map

1:48:22it to a single application Al together

1:48:23we can see that and once it is done we

1:48:26can simply choose a location in which we

1:48:28are trying to launch this particular

1:48:30data Factory for again we have to choose

1:48:33a region which is obviously closer to

1:48:35our end users because at the end if we

1:48:38are not using a a location which is not

1:48:40closer to end users then that will

1:48:42create a huge amount of latency so for

1:48:44example if our users base as based our

1:48:47user are based in Singapore so here we

1:48:49can choose region for Singapore Al

1:48:50together if we know that you our users

1:48:52are based not from Singapore suppose

1:48:55from from Central India Australia East

1:48:57Africa north from Europe from USC so

1:48:59here we can choose our regions

1:49:01accordingly for example suppose here we

1:49:02want to start with North Europe so here

1:49:05we can choose North Europe and then if

1:49:07you want to specify a get URL that will

1:49:09be used as a as a main data source here

1:49:12for for this data Factory for example

1:49:14let's say we use our own GitHub URL for

1:49:17this

1:49:18one let's use L into our

1:49:22GitHub and for example suppose here we

1:49:24may have any particular depository here

1:49:27we have any depository that we want to

1:49:28connect here we can simply connect that

1:49:31let's say here we can enter our GitHub

1:49:44URL here if we have multiple

1:49:46repositories here for example here we

1:49:48have repository as 1 we can choose enti

1:49:51one here if we have Branch as we want to

1:49:53go for the for let's suppose for the

1:49:57development Branch so here we have a

1:49:58branch name for development we can

1:50:00choose a branch

1:50:01name so here we have the entire get URL

1:50:04here we have to define the repository

1:50:06name and then we have to define the

1:50:07branch name and then we can choose a

1:50:09root folder for which we are trying to

1:50:11connect to the connect to here so for

1:50:13example if we have Ro folder by the name

1:50:15of index we can Define index and now we

1:50:18can click on create

1:50:24So currently this is getting

1:50:30initialized as you can see our

1:50:32deployment is done data L and currently

1:50:35the deployment is being in progress for

1:50:38data Factory that we currently deployed

1:50:40so

1:50:49far

1:50:50now currently this has been deployed

1:50:52here now if you want to download the

1:50:54entire deployment details we it's a good

1:50:56practice to download the entire

1:50:57deployment details which will contain

1:50:59the entire list of all the piece of of

1:51:01information and then for getting the

1:51:03entire operations detail here we can

1:51:04simply choose the operations detail

1:51:06where we can Define the entire

1:51:08operations name duration and now if you

1:51:10want to configure this we can open up

1:51:12the

1:51:14resource where and this resource here we

1:51:17first of all here if you want to Quick

1:51:18Start if you want to see the entire

1:51:20activity that means what exactly it has

1:51:22been and again what exactly has been the

1:51:24iops the storage usages and the

1:51:27processor usage and we can go ahead and

1:51:28see the entire monitoring

1:51:32part and if you are Conn if you are

1:51:34planning to connect to to this part

1:51:36instance we have to make sure we are

1:51:38defining the current IM user rules and

1:51:41their policies because again if we are

1:51:43using this through some some particular

1:51:45account here then we can end as we

1:51:47Define the rules for them they will not

1:51:48be able to make any changes to this

1:51:50particular server out there that is

1:51:52something that we have to take care

1:51:54of and along with that let's go ahead

1:51:57and create one storage account as well

1:51:59so again for doing that we can use a

1:52:01service bar available on top so here we

1:52:05have to open up our storage account

1:52:07let's open this

1:52:09up now currently as you can see

1:52:11currently by default start up there

1:52:13won't be any storage account created and

1:52:15deployed so here we have to go ahead and

1:52:16create one click on create storage

1:52:19account

1:52:26here we can choose a subscription model

1:52:28here we can choose a resource Group for

1:52:29the same application resource Group that

1:52:31we have currently created here we have

1:52:33to define the entire storage account

1:52:35name let's say we name it ASA itself and

1:52:39then we can choose the location now

1:52:40remember

1:52:41this okay we already have configured

1:52:43that so here okay one

1:52:46two so here we have to choose the same

1:52:49location in which we have deployed the

1:52:51other data Lake and our data wouse our

1:52:54data Factory

1:52:55itself we have deployed this for

1:52:57northern Europe so here we can choose

1:52:59North Europe and based on where our

1:53:02locations are where users are basically

1:53:04located then we can choose the

1:53:06performance to be standard and premium

1:53:09so again in terms of account kind here

1:53:11we can choose storage or we can choose

1:53:14if you're going for general purpose

1:53:15version one or version two so general

1:53:17purpose one again they are depending

1:53:19upon requirement we can choose

1:53:20accordingly if you want a higher

1:53:22performance that means if we are looking

1:53:23to to have a higher workload then we can

1:53:26use Gen 2 that was released early last

1:53:29year and then if you want to replicate

1:53:31the entire data we want to go for Zone

1:53:34redundant we want to go for locally

1:53:35redundant as if we want to maintain

1:53:37local copy or we want to maintain a read

1:53:39access R st that means again based on

1:53:42multiple regions it will be copy but

1:53:45again the copy will have only the read

1:53:46only access that means this will be used

1:53:49as a primary account for storing data

1:53:51for writing data and the replications

1:53:53will be used for just for the read

1:53:55purposes so that the entire situation of

1:53:58Bott leg is also not

1:54:01created and then we can choose assets

1:54:03here to be hot or to be cool depending

1:54:05upon the requirement we can choose it to

1:54:07be hot or cool here just like we in case

1:54:09we have been Avail if we have been

1:54:11familiar with the concept of multiple

1:54:13storage classes in AWS same way we we

1:54:16have standard and then we have

1:54:17infrequent access so here we can chose

1:54:20school if you want to move to some to

1:54:21something like infrequent access which

1:54:23is not used frequently or we can keep it

1:54:25to hard for those standard we can say

1:54:27most frequently used files here so

1:54:29currently we can keep it to host then we

1:54:32have to define the networking point if

1:54:34we are trying to deploy this in our own

1:54:35isolated Network then we can choose

1:54:38public then we can choose a private or

1:54:40public endpoint depending upon the

1:54:42requirement here so if we have if we

1:54:44have a virtual Network created then we

1:54:46can go ahead and use selected networks

1:54:49if you want to deploy this on public

1:54:50endpoints we can go for public endpoints

1:54:52or if again in case you want to deploy

1:54:55this on our private then first of all we

1:54:57have to configure our prior end point

1:54:58and then only we can get started then we

1:55:01can Define the production type here

1:55:03whether we want to go for stop or

1:55:05desktop relite or if sof as in if you

1:55:08want to retrive it we can simply do that

1:55:10and then we have the Gen 2 hierarchal

1:55:12should be disabled as you don't need it

1:55:14as a part of our Handel currently so we

1:55:16can keep it to disable then we can find

1:55:19the tags here

1:55:20tags are simply used for sorting and

1:55:23filtration purposes if you want to sort

1:55:25this out later on if there are multiple

1:55:27services that we deployed and now we

1:55:29want an easier way to host it easily

1:55:32when we can easily do

1:55:39that and in here we can Define tasp if

1:55:42you don't want to use it we can review

1:55:47it we have to wait for this one to be

1:55:50reviewed here once we are done we can

1:55:52click on

1:55:54create let's and let's wait for this one

1:55:56to be

1:56:09created if you are looking to move files

1:56:11here from one from one part to the other

1:56:13again we can easily do that is it access

1:56:16to free to create database for creating

1:56:19database here here we here we can choose

1:56:21the engine for database for creation of

1:56:23multiple databases here we have

1:56:25different services for for doing that

1:56:27for example suppose for creation of

1:56:29databases here we can click on create

1:56:32resource we can choose resource type as

1:56:34databases and here we can choose which

1:56:36part database engine we are going to

1:56:37deploy here just like we have RDS

1:56:39available in adus same way here we can

1:56:42change it

1:56:44up and if we looking to create a

1:56:46complete data pipeline then that's why

1:56:48we use data

1:56:50Factory so for if we are trying to to

1:56:54transfer one data from the other account

1:56:56here we can simply share the across

1:56:57multiple accounts if you're trying to

1:56:59start to transfer data from on premise

1:57:02to SEO or from SE to on premise again

1:57:04for if we looking to transfer from aure

1:57:06to on premise then we can from on a to

1:57:09on premise that me locally then we can

1:57:11use another service called as storage

1:57:13Explorer as a part of AO platform we can

1:57:15use that if we looking to transfer from

1:57:19uh Z platform to on premise we can do

1:57:21that we can take the help of storage

1:57:23Explorer as

1:57:26service and then we can easily use the

1:57:29AO import export to Simply transfer the

1:57:31service from our AO platform to oise

1:57:34just like we have a physical device

1:57:36offered by as a part of snowball in AWS

1:57:39in case we have been working with AWS

1:57:41and we there we have a service called as

1:57:43snowball where it is a simple physical

1:57:45device for if we looking to get

1:57:47connected we can use that we can take

1:57:49take the help of storage gate phase in

1:57:50that so just like storage gate phase

1:57:53here we can take the help of storage

1:57:54Explorer back into Azure for transfering

1:57:58data from in and out of SEO

1:58:02[Music]

1:58:06platforms now let's have an introduction

1:58:08of a database first let's understand why

1:58:10do we need a database so the various

1:58:13reasons a database is important first of

1:58:15all it manages large amounts of data a

1:58:17database stores and manages a large

1:58:18amount of data on a daily basis this

1:58:20would not only be possible using any

1:58:22other tool such as a spreadsheet as they

1:58:25would simply not work second is its

1:58:27accuracy so a database is pretty

1:58:29accurate as it has all sours of building

1:58:31constraints checks Etc this means that

1:58:33the information available in database is

1:58:35guaranteed to be correct in most cases

1:58:38it's easy to update data in a database

1:58:39so in a database it is easy to update

1:58:41data using like various data

1:58:43manipulation languages available one of

1:58:44these languages SQL for the security of

1:58:47data so databases have various methods

1:58:49to ensure security of data there are

1:58:52user logins required before accessing a

1:58:54database and various access specifiers

1:58:57these allow only authorized users to

1:58:59access the database fifth is data

1:59:01Integrity this is ensured in databases

1:59:03by using various constraints for data

1:59:06data Integrity in databases makes sure

1:59:08that the data is accurate and consistent

1:59:10in a database the last is easy to

1:59:12research data it is very easy to access

1:59:15and research data in a database this is

1:59:17done using data Cy language which allow

1:59:20searching of any data in the database

1:59:21and performing computations on it now

1:59:24that you have understood the need of a

1:59:25database let's briefly understand what

1:59:26actually it is so a database is an

1:59:29organized collection of structure

1:59:30information or data typically stored

1:59:32electronically in a computer system a

1:59:34database is usually controlled by

1:59:36database management system together the

1:59:38data and the database management system

1:59:40along with applications that are

1:59:41associated with them are referred to as

1:59:43a database system often shortened to

1:59:46just a database so data within the most

1:59:48common types of database in operation

1:59:50today is typically modeled in rows and

1:59:53columns in a series of tables to make

1:59:55processing and data quering efficient

1:59:57the data can then be easily accessed

1:59:59managed modified updated controlled and

2:00:02organized most databases are structured

2:00:04query language for writing and querying

2:00:06data databases are used to support

2:00:08internal operations of organizations and

2:00:10to underpin online interactions with

2:00:12customers and

2:00:14suppliers databases are used to hold

2:00:16administrative information and more ized

2:00:19data such as engineering data or

2:00:21economic models example includes

2:00:23computerized Library System flight

2:00:24reservation system computerized past

2:00:26inventory system and many content

2:00:28Management systems that store websites

2:00:30as collection of web pages in a database

2:00:33now that you have an understanding of

2:00:34Microsoft aour as well as of a database

2:00:36let's now take a look at different types

2:00:38of databases in Azure first is

2:00:40relational database a relational

2:00:41database is a type of database that

2:00:43stores and provide access to data points

2:00:46that are related to one another

2:00:47relational databases are based on the

2:00:49relational model an intuitive

2:00:51straightforward way of representing data

2:00:53in tables in a relational database each

2:00:56row in the table is a record with a

2:00:57unique ID called the key The Columns of

2:01:00the table hold attributes of the data

2:01:02and each record usually has a value for

2:01:04each attribute making it easy to

2:01:07establish the relationships among data

2:01:08points in a relational database all data

2:01:11is stored and accessed by relations so

2:01:13relations that store data are called

2:01:15base relations and in implementations

2:01:17are called tables other relations do not

2:01:19store data but are computed by applying

2:01:21relational operations to those relations

2:01:24these relations are sometimes called

2:01:26derived relations in implementations

2:01:28these are called views or queries

2:01:30derived relations are convenient in that

2:01:32they act as a single relation even

2:01:34though they may grab information from

2:01:35several relations each relation or table

2:01:38has a primary key this being a

2:01:39consequence of a relation being a set a

2:01:42primary key uniquely specifies a tuple

2:01:44within a table while natural attributes

2:01:46are sometimes good primaries so so this

2:01:49is all about relational database then

2:01:51second we have is non- relational

2:01:52database also known as nosql databases

2:01:55so nosql database or non- relational

2:01:57database provides a mechanism for

2:01:58storage and retrival of data that is

2:02:01modeled in means other than the tabular

2:02:03relations used in relational databases

2:02:05so non-national databases are

2:02:06increasingly used in big data and

2:02:08realtime web applications and

2:02:09non-national databases are also like

2:02:11sometimes called not only SQL to

2:02:13emphasize that they may support SQL like

2:02:16query languages or sit alongside SQL

2:02:19databases so the prominent non

2:02:22relational databases provided by aour is

2:02:23Cosmos database that I will explain you

2:02:26further in this video and third is

2:02:29inmemory database so an inmemory

2:02:31database also like in the short form we

2:02:32say it as IMDb also like a main memory

2:02:35database system or mmdb or memory

2:02:37resident database these are all the

2:02:39names of it so an inmemory database is a

2:02:41database management system that

2:02:43primarily relies on Main memory for

2:02:44computer data storage it is contrasted

2:02:47with database management system that

2:02:48emplo deploy a disk storage mechanism in

2:02:51memory databases are like faster than

2:02:52dis optimized databases because dis

2:02:54access is slower than memory access the

2:02:57internal optimization algorithms are

2:02:59simpler and execute fewer CPU

2:03:01instructions accessing data in memory

2:03:03eliminates seek time when querying the

2:03:05data which provides faster and more

2:03:07predictable performance than disk a

2:03:09potential technical hurdle with inmemory

2:03:11data storage is the volatility of RAM is

2:03:14specifically in the event of a power

2:03:15loss intentional or otherwise data

2:03:18stored in volatile Ram is lost with the

2:03:21introduction of nonvolatile Random

2:03:23Access Memory technology in memory

2:03:25databases will be able to run at full

2:03:27speed and maintain data in the event of

2:03:29power

2:03:29failure let's Now understand the

2:03:32architecture of database Services

2:03:33provided by a so you can see the it

2:03:36looks like a complex architecture but I

2:03:38will explain you in quite easily so the

2:03:40basic fundamental building block that is

2:03:42available in aour is the SQL database so

2:03:45Microsoft offers this SQL server and SQL

2:03:47database on aour in in many ways we can

2:03:49deploy a single database or we can

2:03:51deploy multiple databases as part of a

2:03:54shared elastic pool you can see the

2:03:56elastic pools and single database okay

2:03:59Microsoft introduced a managed instance

2:04:01that is targeted towards on premises

2:04:02customers so if we have some SQ

2:04:05databases within our on premises Data

2:04:06Center and we want to migrate the

2:04:08database into Azure without any complex

2:04:10configuration or ambiguity then we can

2:04:12use a managed instance because this is

2:04:14mainly targeted towards on premises

2:04:16customers who want to lift and share

2:04:18their own on premises database into

2:04:19Azure with the least effort and

2:04:21optimized cost we can also take

2:04:23advantage of Licensing we have within

2:04:26our on premises data center Microsoft

2:04:28will be responsible for maintenance

2:04:30patching and related services but in

2:04:32case if we want to go for the

2:04:34infrastructure as a service for the SQL

2:04:36Server then we can deploy SQL server on

2:04:38the Azure virtual machine if the data

2:04:41has a dependency on the underlying

2:04:42platform and we want to log into the SQL

2:04:44server in that case we can use the SQL

2:04:46server on a virtual machine we can Dey

2:04:48Dey SQL Data Warehouse on the cloud

2:04:50Azure offers many other database

2:04:52services for different types of

2:04:54databases such as MySQL marad DB and

2:04:57also post SQL once we deployed a

2:05:00database into a we need to migrate the

2:05:02data into it or replicate the data into

2:05:04it okay then we have is azure database

2:05:07services for data migration so services

2:05:09that are available in Azure which we can

2:05:11use to migrate the data from our on

2:05:12premises SQL Server into Azure so in

2:05:15that the first one is azure data

2:05:16migration service so it is used to to

2:05:19migrate the data from our existing SQL

2:05:20server and database within the on

2:05:22premises data center into the Azure then

2:05:24we have Azure SQL data synchronization

2:05:27if we want to replicate the data from

2:05:29our on premises database into azour then

2:05:31we can use azour SQL data sync then we

2:05:34have SQL stretch database so it is used

2:05:37to migrate cold data into Azure SQL

2:05:39stretch database is a bit different from

2:05:41other database offerings it works as a

2:05:43hybrid database because it divides the

2:05:45data into different types like hot and

2:05:47cold so hot data will be kept in the on

2:05:49premes Data Center and C data in the

2:05:51Azure then we have is data Factory so

2:05:53Azure data Factory is used for ETL means

2:05:56transformation extraction and loading so

2:05:59using the data Factory we can even

2:06:01extract the data from our on premises

2:06:02data center we can do some conversion

2:06:04and load into the Azure SQL database

2:06:07data Factory is an ETL tool that is

2:06:09offered on the cloud which we can use to

2:06:11connect to different databases and like

2:06:13extract the data or transform it and

2:06:15load into a destination then there's

2:06:17Azure security

2:06:19so all the databases that exist in our

2:06:21need to be secured and also we need to

2:06:23accept connections from known Origins

2:06:25for this purpose all these database

2:06:27Services comes with firewall rules where

2:06:30we can configure from which particular

2:06:32IP address we want to allow connection

2:06:34we can define those firewall rules to

2:06:37limit the number of connections and also

2:06:38reduce the service attack area so now

2:06:41let's talk about the cosmos DV so Cosmos

2:06:43DV is nothing but a SQL data store that

2:06:45is available in aour and it is designed

2:06:47to be globally scalable and also very

2:06:50highly available with extremely low

2:06:51latency Microsoft guarantees latency in

2:06:54terms of reading and wrs with Cosmos DV

2:06:57for example if we have any application

2:06:58such as iot or gaming where we get a lot

2:07:01of data from different users spread

2:07:03across globally then we will go for

2:07:05Cosmos DV because Cosmos DB is designed

2:07:07to be globally scalable and highly

2:07:09available due to which users will like

2:07:11experience low latency finally there are

2:07:14two things and one is we need to secure

2:07:16all the services for that purpose we can

2:07:18integrate all these services with azour

2:07:20active directory and manage the users

2:07:22from Azure active directory also to

2:07:24monitor all these Services we can use

2:07:27the security Center so there is an

2:07:29individual monitoring tool too but AZ

2:07:31security Center will keep on monitoring

2:07:33all these services and provide

2:07:34recommendations if something is wrong I

2:07:36hope the architecture of azure database

2:07:37Services now clear to you now let's uh

2:07:40move forward to briefly understand the

2:07:41database Services provided by Azure so

2:07:44the first is azure SQL database so SQL

2:07:47database is the flagship product for

2:07:48Microsoft in the database area it is a

2:07:50general purpose relational database that

2:07:52supports structures like relation data

2:07:54Json spatial and XML the Azure platform

2:07:56fully manages every azour SQL database

2:07:58and guarantees no data loss and a high

2:08:00percentage of data availability azour

2:08:03automatically handles patching backups

2:08:04replication failure detection underlying

2:08:07potential Hardware software or network

2:08:09failure deploying Buck fixes failovers

2:08:11and like database upgrades and other

2:08:13maintenance tasks so there are three

2:08:15ways we can Implement our SQL database

2:08:18so first is managed instance this is

2:08:20premar targeted towards on premises

2:08:22customers in case if we really have a

2:08:24SQL Server instance in our on premises

2:08:26Data Center and you want to migrate that

2:08:28into Azure with minimum changes to a

2:08:30application and the maximum

2:08:31compatibility then we will go for manage

2:08:34instance second is single database so we

2:08:36can deploy a single database on a its

2:08:38own set of resources managed via L

2:08:41logical server okay then we have his

2:08:42elastic pool we can deploy a pool of

2:08:45databases with a shared set of resources

2:08:47managed bya log local server we can like

2:08:50deploy the SQL database as an

2:08:51infrastructure as a service that means

2:08:53we want to use the SQL server on Azure

2:08:55virtual machine but in the case we are

2:08:56responsible for managing the SQL server

2:08:58on that but in that case we are

2:09:00responsible for managing the SQL server

2:09:02on that particular Azo virtual machine

2:09:04so then we have is the purchasing model

2:09:06so there are two ways we can purchase

2:09:08the SQL server on a j so first is voree

2:09:10purchasing model also as virtual core

2:09:12purchasing model so the vco purchasing

2:09:14model enables us to independently scale

2:09:17compute and Storage resources match on

2:09:20premises performance and optimize price

2:09:22it also allows us to choose a generation

2:09:24Hardware it also allows us to use Azure

2:09:27hybrid benefit for SQL Server to gain

2:09:30cost savings best for the customers who

2:09:32value flexibility control and

2:09:34transparency so second is DD model it is

2:09:37based on a bundled measure or compute

2:09:39storage and input output resources so

2:09:43sizes of the compute are expressed in

2:09:45terms of database transaction units

2:09:47means dtus for single databases and

2:09:50elastic database transaction units for

2:09:51elastic pools this model is best for

2:09:54customers who want simple pre-configured

2:09:56resource options in the second database

2:09:59service is azure Cosmos database so

2:10:02Azure Cosmos database is a no SQL data

2:10:04store it is different from the

2:10:05traditional relational database where we

2:10:08have a table and the table will have a

2:10:09fixed number of columns and each row in

2:10:11the table should ADH to the scheme of

2:10:13the table in the no SQL database you

2:10:16don't Define any schema at all for the

2:10:18table and each item or row within the

2:10:21table can have different values or

2:10:23different schema itself so now let's

2:10:25understand the cosmos database structure

2:10:28first one in the structure is database

2:10:30so we can create one or more Azure

2:10:32Cosmos database under our account a

2:10:34database is analogous to a name space

2:10:37and it is the unit of management for a

2:10:39set of azure Cosmos containers so the

2:10:42second is Cosmos account so the Azure

2:10:44Cosmos account is the basic unit of

2:10:46global distribution and high

2:10:47availability for for globally

2:10:48Distributing our data and throughput

2:10:51across multiple Azure regions we can add

2:10:53or remove Azure regions from our Azure

2:10:55Cosmos at any time I mean Azure Cosmos

2:10:57account at any time so the third is a

2:10:59container so an azour Cosmos container

2:11:02is the unit of scalability for both

2:11:04provision throughput and storage of

2:11:06items a container is horizontally

2:11:08partitioned and then replicated across

2:11:10multiple regions then let's understand

2:11:13the types of consistency under Cosmos DB

2:11:15so Azure Cosmos database approaches the

2:11:17data consistency as a spectrum of

2:11:18choices instead of two extremes so

2:11:21strong compatibility and eventual

2:11:23consistency are at the ends but these

2:11:25are many consistency choices along along

2:11:28the Spectrum so the consistency levels

2:11:30are region agnostic the consistency

2:11:32level of our Azure Cosmos account is

2:11:35guaranteed for all read operations

2:11:37regardless of the region from which the

2:11:39reads and rights are served the number

2:11:40of areas associated with the Azure

2:11:42Cosmos account or whether our account is

2:11:45configured with a single or multiple

2:11:47right regions

2:11:48then there's request unit so we pay for

2:11:51the throughput we provision and the

2:11:53storage we consume on an hourly basis

2:11:55with Azure Cosmos DB remember this DB

2:11:58means database so then there are request

2:12:00units in Cosmos DB means Cosmos database

2:12:03so we pay for the throughput we

2:12:04provision and the storage we consume on

2:12:06an hourly basis with Azure Cosmos GB the

2:12:09cost of all the database operations is

2:12:11normalized by Azure Cosmos DV and is

2:12:13expressed in terms of request units the

2:12:15price to readed a 1 KB item is a one

2:12:18request unit all other database

2:12:19operations are similarly assigned with a

2:12:21cost in terms of research units the

2:12:24number of research units consumed will

2:12:25depend on the type of operations item

2:12:27size data consistency query patters etc

2:12:30for the management and planning of

2:12:32capacity Azure Cosmos database ensures

2:12:34that the number of research units for a

2:12:35given database operations over a given

2:12:37data set is deterministic and the third

2:12:40database service is azure data Factory

2:12:42so Azu data Factory is a data

2:12:43integration service based on the cloud

2:12:44that allows us to create data driven

2:12:46workflows in the cloud for for

2:12:48orchestrating and automating data

2:12:50movement and data transformation data

2:12:52Factory is a perfect ETL tool on cloud

2:12:54data Factory is designed to deliver

2:12:57extraction transformation and loading

2:12:58process within the cloud the ETL process

2:13:01generally involves four steps so the

2:13:03first one is connecting collect we can

2:13:04use the copy activity in a data pipeline

2:13:06to move data from both on premises and

2:13:09Cloud secure data stores so the second

2:13:12is a transform so once the data is

2:13:14present in a centralized data store in

2:13:16the cloud process or trans form the

2:13:18collected data by using compute services

2:13:20such as HD Insight Hadoop spark data

2:13:23leak analytics and machine learning

2:13:24third is published so after the raw data

2:13:26is refined into a business ready

2:13:28consumable form it loads the data into

2:13:30an azour data warehouse azour SQL

2:13:33database and Azure Cosmos database Etc

2:13:35so fourth is Monitor so azour data

2:13:37Factory has built in support for

2:13:39pipeline monitoring via azour monitor

2:13:42API Powershell log analytics and health

2:13:44panels on the azour portal so then there

2:13:46are components of data so data Factory

2:13:48is composed of six key elements all

2:13:51these components work together to

2:13:52provide a data form on which you can

2:13:54form a datadriven workflow with the

2:13:56structure to move and transform the data

2:13:58so first one is pipeline a data Factory

2:14:00can have one or more pipelines it is a

2:14:03logical grouping of activities that

2:14:05perform a unit of work the activities in

2:14:08a pipeline perform the task Al together

2:14:10for example a pipeline can contain a

2:14:12group of activities that inest data from

2:14:14a Azure blob and then runs a hi query

2:14:17and an HD inside cluster to partition

2:14:19the data so second is activity it

2:14:21represents a processing step in a

2:14:23pipeline for example we might use a copy

2:14:25activity to copy data from one data

2:14:27store to another data store then we have

2:14:29a data sets so it represents data

2:14:31structure within the data stores which

2:14:33point to or reference the data or we

2:14:35want to use inov activities as input or

2:14:38output then there are Link services so

2:14:41it is like connection strings which

2:14:43Define the connection information needed

2:14:45for data Factory to connect to external

2:14:47resources

2:14:48a linked service can be a data store and

2:14:50compute resources linked service can be

2:14:53a link to a data store or a compute

2:14:54resource also then we have a triggers so

2:14:58it represents the unit of processing

2:14:59that determines when a pipeline

2:15:00execution needs to be disabled we can

2:15:02also schedule these activities to be

2:15:04performed at some point in time and we

2:15:07can use the trigger to disable an

2:15:09activity then the last one is control

2:15:11flow so it is an orchestration of

2:15:12pipeline activities that include

2:15:14chaining activities in a sequence

2:15:15branching defining parameters for the

2:15:17pipeline label and passing arguments

2:15:20while invoking the pipeline on demand or

2:15:22from a tiger we can use a control flow

2:15:25to sequence certain activities and also

2:15:27Define what parameters need to be passed

2:15:30for each of these activities I hope you

2:15:33have now understood the major Services

2:15:35of azure databases so now let's have a

2:15:38look at some of the use cases for Azure

2:15:39database Services first let's see the

2:15:41use cases for SQL database so the first

2:15:44one is developer or test environment an

2:15:46important use case for replicate getting

2:15:48or migrating data to SQL hosted on Azure

2:15:50is for developer or test environments

2:15:52before deploying to the production

2:15:53environment it is pertinent that the

2:15:56data is tested against developer and

2:15:57test environments so Azure SQL database

2:16:01can act as a target for such

2:16:02environments the life production

2:16:04environment can be replicated to the

2:16:05developer or test environment using a

2:16:07database copy so the second is business

2:16:09continuity one of the most important use

2:16:11cases for SQL on azour is using it as a

2:16:13Dr Target to maintain business

2:16:15continuity azour SQL databases can

2:16:17provide an SLA of up to

2:16:1999.99% by maintaining several copies of

2:16:21the data this provides business

2:16:23continuity as it allows you to restore

2:16:25GE redundant copies of the data or use

2:16:29active Geo redundant copies as failover

2:16:31points in use of outages at data centers

2:16:34or in regions besides SQL databases you

2:16:38can also use availability groups to

2:16:40fulfill business continuity demands not

2:16:41only can you use availability groups in

2:16:43Azure SQL virtual machines but also use

2:16:46Azure SQL virtual machine instances as a

2:16:47target for high availability and

2:16:49disaster recovery and the third one is

2:16:52scaling out readon workloads apart from

2:16:54providing PC or Dr capabilities active

2:16:57Geo replication can also be used to

2:16:59offload readon workload such as

2:17:01reporting jobs to secondary copies you

2:17:03can also extend on premises SQL Server

2:17:05instance using readable always on

2:17:08replicas and the fourth one is backup

2:17:10and G so Azure SQL database are backed

2:17:12up automatically on a regular basis and

2:17:15there are no storage cost for to 200% of

2:17:17the maximum provision database storage

2:17:19you can restore backups to any point in

2:17:22time going back to a pretended period

2:17:24which is determined by the Azure SQL

2:17:26Service Tire in use on premises SQL

2:17:29Server databases and transaction locks

2:17:31can also be bagged up directly to Azure

2:17:33using the backup to URL feature and

2:17:35stored in Azure storage so Azure SQL

2:17:38databases can also be stored on local

2:17:40storage by exporting them to backpack

2:17:43files means backup and the files so

2:17:46fifth one is Advanced analytics so so

2:17:47another important reason for hosting SQL

2:17:49in azour is to make use of azure's

2:17:51advanced gentics platforms such as azour

2:17:53storage blob and Azure data leak store a

2:17:56common scenario with Advanced analytics

2:17:57is when users reference data from

2:17:59various data sources use Azure data L

2:18:02store as the staging area or perform

2:18:05transformation activities using hi or

2:18:07spark and finally load the data into

2:18:08Azure data warehouse for bi and

2:18:10Reporting bi means business intelligence

2:18:13now let's see the use cases for Cosmos

2:18:15database first of all they used in iot

2:18:17and telematics so iot use cases commonly

2:18:19share some patterns in how they ingest

2:18:21process and store data first these

2:18:23systems need to ingest burst of data

2:18:25from device sensors of various locals

2:18:28next these systems process and analyze

2:18:30streaming data to derive real time

2:18:33insights the data is then archive tool

2:18:35Co storage for batch analytics Microsoft

2:18:37Azure offers Rich services that can be

2:18:40applied for iot use cases including

2:18:42Azure Cosmos database Azure event hubs

2:18:45Azure stream analytics Azure

2:18:46notification hub Azure machine learning

2:18:48Azure HD insight and powerbi burst of

2:18:51data can be ingested by Azure event hubs

2:18:53as it offers High throughput data

2:18:55ingestion with low latency data ingested

2:18:57that needs to be processed for realtime

2:18:59Insight can be funneled to aure stream

2:19:01analytics for realtime analytics data

2:19:04can be loaded into Azure Cosmos database

2:19:06for an ad hoc query once the data is

2:19:08loaded into azour Cosmos database the

2:19:10data is ready to be queried in addition

2:19:13new data and changes to existing data

2:19:15can be read on changed feed

2:19:18so change speed is a persistent append

2:19:20Only log that stores changes to Cosmos

2:19:22containers in sequential order then all

2:19:24data or just changes to data in Azure

2:19:26Cosmos database can be used as reference

2:19:29data as part of a realtime analytics in

2:19:31addition data can further be refined and

2:19:34processed by connecting Azure Cosmos

2:19:36database data to HD insight for pig

2:19:38hiive or map reduce jobs refined data is

2:19:41then for a sample of iot solution using

2:19:45Azure Cosmos database event hubs and

2:19:47storm see the HD Insight storm examples

2:19:49repository on GitHub okay then we have

2:19:52is retail and marketing so Azor Cosmo

2:19:55database is used extensively in

2:19:57Microsoft's own e-commerce platforms

2:19:58that runs the Windows store and Xbox

2:20:00Live it is also used in the retail

2:20:02industry for storing catalog data and

2:20:05for event sourcing in order to process

2:20:07pipelines so catalog data storage

2:20:09scenarios involve storage and query a

2:20:11set of attributes for entities such as

2:20:13people places and products some examples

2:20:16of catalog data are user accounts

2:20:18product cataloges iot devices Registries

2:20:21and build of material systems attributes

2:20:24for this data may vary and can change

2:20:26over time to fit application

2:20:28requirements consider an example of a

2:20:30product catalog of an automative part

2:20:32supplier every part may have its own

2:20:34attributes in addition to the common

2:20:36attributes that all parts share

2:20:38furthermore attributes for a specific

2:20:39part can change the following year when

2:20:42a new model is released Azure Cosmos

2:20:44database supports flexible schemas and H

2:20:46High iCal data and thus it is well

2:20:49suited for storing product catalog data

2:20:51Azure Cosmos database is often used for

2:20:53event sourcing to power event driven

2:20:55architectures using its change feed

2:20:57functionality the change feed provides

2:20:59Downstream microservices the ability to

2:21:01reliability and incrementally read

2:21:03inserts and updates made to an Azure

2:21:06Cosmos database this functionality can

2:21:08be leveraged to provide persistent event

2:21:10store as a message broker for State

2:21:13changing events and drive order

2:21:15processing workflow between any micros

2:21:17servic

2:21:18in addition data store in Azure Cosmos

2:21:20database can be integrated with HD

2:21:22insight for big data analytics via

2:21:24Apache spark jobs so the third one is

2:21:27gaming the database tire is a crucial

2:21:29component of gaming applications modern

2:21:31gaming app perform graphical processing

2:21:33on mobile or console clients but rely on

2:21:35the cloud to deliver customized and

2:21:37personalized content like in-game stats

2:21:40social media integration and high school

2:21:42leaderboards games often require single

2:21:44millisecond latencies for reads and WR

2:21:47to provide an engaging in-game

2:21:49experience a game database needs to be

2:21:51fast and be able to handle massive Spice

2:21:53in request rates during new game

2:21:55launches and feature

2:21:57updates so Azure Cosmo database is used

2:21:59by games like The Walking Dead No Man's

2:22:01Land by next games and hello five

2:22:03guardians so Azure Cosmos database

2:22:06provides the number of benefits to game

2:22:07developers like Azure Cosmos DB allows

2:22:10performance to be scaled up or down

2:22:12elastically this allows games to handle

2:22:14updating profiles and stats from dozens

2:22:17millions of simultaneous Gamers by

2:22:19making a single API call then Azure

2:22:21Cosmos DV supports millisecond reads and

2:22:23rights to help avoid any lags during the

2:22:26game play Then Azo Cosmos database

2:22:29automatic indexing allows for filtering

2:22:31against multiple different properties in

2:22:33real time for example locating players

2:22:35by the internal player IDs or their game

2:22:37center Facebook Google IDs or quering

2:22:39based on player membership in a guild

2:22:41this is possible without building

2:22:43complex indexing or shedding

2:22:44infrastructure social features including

2:22:47game that messages player Guild

2:22:48membership challenges completed high

2:22:50score leaderboards and social graphs are

2:22:52easier to implement with a flexible

2:22:54schema so Azor Cosmos database as a

2:22:57managed platform as a service require

2:22:58minimal setup and management work to

2:23:00allow for Rapid iteration and reduce

2:23:02time to market the last one is web and

2:23:05mobile applications so Azure cosos

2:23:06database is commonly used within mobile

2:23:08and web applications and is well suited

2:23:11for modeling social interactions

2:23:13iterating with third party services and

2:23:14for building Rich personal experiences

2:23:17the costos database sdks can be used to

2:23:19build Rich IOS and Android applications

2:23:22using the popular zamarin framework so

2:23:24under web applications first we have the

2:23:27social applications and then we have

2:23:28personalizations so in Social

2:23:30applications a common use for Azure

2:23:32Cosmo database is store and query user

2:23:35generated content means ugc so for web

2:23:38mobile and social media applications

2:23:39some examples of a user generated

2:23:41content are chat sessions tweets blogs

2:23:44posts rating and comments often the ug

2:23:47in social media applications is a bland

2:23:49of free form text properties text and

2:23:51relationships that are not bounded by

2:23:53rigid structure content such as stats

2:23:55comments and posts can be stored in

2:23:58Cosmos DB without requiring

2:23:59Transformations or complex object to

2:24:02relational mapping layers data

2:24:03properties can be added or modified

2:24:05easily to match requirements as

2:24:06developers it iterate over the

2:24:08applications code thus promoting rapid

2:24:10development applications that integrate

2:24:12with third party social network must

2:24:14respond to changing schemas from these

2:24:17networks as data is automatically

2:24:18indexed by default in Cosmos database

2:24:20data is ready to be queried at any time

2:24:23hence these applications have the

2:24:24flexibility to retri projections as

2:24:26their respective needs so the second

2:24:29thing in web mobile applications is

2:24:31personalization so now is mod

2:24:33applications comes with complex views

2:24:34and experiences these are typically

2:24:37Dynamic catering to user preferences or

2:24:39moods and branding needs hence

2:24:41applications need to be able to try

2:24:43personalization settings effectively to

2:24:45render UI elements and experience es

2:24:47quickly Json a format supported by the

2:24:50cosmos DB is an effective format to

2:24:52represent UI layout data as it is not

2:24:54only lightweight but also can be easily

2:24:57interpreted by JavaScript Cosmos R

2:24:59offers turnable consistency levels that

2:25:01allows fast reads with low latency

2:25:04rights hence storing UI layout data

2:25:06including personalized settings as Json

2:25:09documents and Cosmos GB is an effective

2:25:11means to get this data across the wire

2:25:14so these were the use cases for Cosmos

2:25:16DB

2:25:17now that you have a theoretical

2:25:18understanding of azure database Services

2:25:20let's now see a simple deployment of a

2:25:22database service on Microsoft Azure the

2:25:25simply type Microsoft Azure on

2:25:27Google what you can do is you can create

2:25:29a free account on Microsoft Azure you

2:25:32get S 12 months of free services and

2:25:33around 40,000 rupees of free credits

2:25:35also for using the services we can just

2:25:38directly open the console from here I

2:25:41just sign

2:25:43in so for deploying a simple database

2:25:46service so so what we going to deploy

2:25:48today we can like deploy Cosmos database

2:25:50like I have explained you what is cosmos

2:25:51database a new SQL database it is so we

2:25:53can just go to console to the portal I

2:25:56can go to the

2:25:59portal you can create a resource from

2:26:02here can search for

2:26:06Cosmos like J Cosmos yes you can see

2:26:09here like free credits I have a free

2:26:11trial account so it is showing that I

2:26:12have 14,500 three GRS so this is the

2:26:16like credit amount you get for in a free

2:26:18trial okay so you can create a Azure

2:26:20Cosmos DB from here which one you want

2:26:22to create like you can create Pro SQL

2:26:24one so Resource Group can give a new one

2:26:27or we have existing we have a Rec Cosmos

2:26:30one resource Cosmos so if you want to

2:26:32choose the existing one or if you want

2:26:34to choose the new one okay so you can

2:26:36choose the existing one from here or if

2:26:38you want to create a new one then create

2:26:39new One Source One

2:26:41Cosmos I hope that works out yeah then

2:26:45give any unique name for this like uh I

2:26:48will give

2:26:50demoore Cosmos 1 2 3 okay it cannot

2:26:55contain uh UND remember these things

2:26:57okay it do cannot contain this so 1 2 3

2:27:004 I will okay it's not available so I

2:27:03will give five also yeah it's available

2:27:05now and like choose your location

2:27:07whichever location you are located in

2:27:09can use nearby location so mine is Asia

2:27:12Pacific Central India so I've chosen

2:27:14this so free trial account is already

2:27:16there I applied for it then you can just

2:27:19review and

2:27:20create before getting deployed it will

2:27:23show you the review for

2:27:25it so remember that on the basis of the

2:27:28location we have selected the creation

2:27:29time will differ okay so you can review

2:27:32it all the information what you have

2:27:33inserted so now you can

2:27:38create so deployment is in progress it

2:27:41will take a few

2:27:43minutes you can see like how the

2:27:45resource has been created deployment is

2:27:47in progress will soon be created you can

2:27:49check details for it from

2:27:51here let's go to portal

2:27:55again like it is already pinned here or

2:27:57you can search from here okay for Cosmos

2:28:00GV so just click here so it is showing

2:28:03that it is getting created this is the

2:28:05one I have created before only this one

2:28:07is creating this in progress let's

2:28:09refresh one again you can see the

2:28:12processing going on

2:28:15here yeah so your deployment is complete

2:28:18showing you can go to Resource from here

2:28:21also you can just refresh it from

2:28:24here so yeah this is how it's been

2:28:27created you can open the source from

2:28:28here and you can go to activity log or

2:28:31data Explorer you can create a database

2:28:33anything or you can see the consistency

2:28:35of it like default consistency and

2:28:37everything so that's how customers

2:28:39database is been deployed so I hope you

2:28:41have understood this

2:28:45deployment

2:28:49let's look into the family of azure SQL

2:28:52so first one in the family is SQL server

2:28:54on Virtual machines so with this you can

2:28:57lift and shift your SQL Server workloads

2:28:59to the cloud to get the combined

2:29:01performance security and analytics of

2:29:03SQL server with flexibility and hybrid

2:29:05connectivity of azure with 100% code

2:29:08compatibility access the latest SQL

2:29:10Server updates and releases including

2:29:12SQL Server 2019 register your virtual

2:29:15machines with SQ infrastructure as a

2:29:17service agent extension for automated

2:29:20virtual machine management at no

2:29:21additional cost SQL server on Azure

2:29:24virtual machines is part of the Azure

2:29:26SQL family which allows you to migrate

2:29:28existing apps or build new apps on the

2:29:31best cloud destination for a mission

2:29:33critical SQL Server workloads so its

2:29:36features are first of all best TCO that

2:29:38is total cost of ownership with Azure

2:29:40hybrid benefit with Azure SQL Server you

2:29:43can save up to 84% compared to Amazon

2:29:45web services migrating SQ server

2:29:47databases with Azure hybrid benefit and

2:29:49get free extended support for SQL Server

2:29:512008 R2 images in Azure infrastructure

2:29:54as a service activate Azure hybrid

2:29:56benefit when you provision SQL server on

2:29:59Azure virtual machines images from the

2:30:01Azure Marketplace second feature is high

2:30:03performance virtual machines for SQL

2:30:05server on Linux and windows so you can

2:30:08take advantage of SQL Server virtual

2:30:09machines with industry leading

2:30:11performance choose from images with

2:30:13Windows Server redhead Enterprise Linux

2:30:15SU Enterprise Linux server or you been

2:30:18to Linux gain collocated integrated

2:30:21support for your SQL workloads with

2:30:22redhead and suc the third feature is

2:30:25built-in security and manageability so

2:30:28you can ease maintenance with automatic

2:30:30security updates and restore your

2:30:32database to a specific point in time

2:30:34with Azure backup help protect your data

2:30:36address and in motion with the database

2:30:38stated as least vulnerable over the last

2:30:419 years in the cloud with the most

2:30:43global national and Industry

2:30:44certifications the second member in the

2:30:46family is azure SQL managed instance so

2:30:50part of the Azure SQL service portfolio

2:30:52Azure SQL managed instance is the

2:30:54intelligent scalable Cloud database

2:30:56service that combines the broadest SQL

2:30:58Server engine compatibility with all the

2:31:01benefits of a fully managed and everen

2:31:03platform as a service with SQL managed

2:31:06instance confidently modernize your

2:31:08existing apps at scale by combining your

2:31:10experience with familiar tools skills

2:31:12and resources and do more with what you

2:31:15already have Azure Arc enabled SQL

2:31:18manage instance is now in preview you

2:31:21can run the service on premise on any

2:31:23infrastructure of your choice with Azure

2:31:25Cloud benefits like elastic scale

2:31:26unified management and a cloud billing

2:31:28module while staying always current some

2:31:31of its features are always operate on

2:31:33the latest version of SQL so SQL manage

2:31:36instance is built on the SQL Server

2:31:38engine it's overgreen meaning it's

2:31:40always up to date with the latest SQL

2:31:42features and functionality never worry

2:31:45about updates upgrades or end of support

2:31:48again second is fully managed and

2:31:50optimized for DBA productivity so boost

2:31:53productivity and operate more

2:31:54efficiently by letting the service

2:31:56perform timec consuming and complex

2:31:58tasks on your behalf features like

2:32:00built-in High availability disaster

2:32:02recovery and automated backups ensure

2:32:04your data is available when you need it

2:32:06while AI power automatic tuning

2:32:08optimizes performance for you SQL manage

2:32:10instance combines all the best of SQL

2:32:13server with the financial and

2:32:14operational benefits of the platform as

2:32:16a service and the third feature is

2:32:18maintain SQL Server application

2:32:20compatibility so accelerate application

2:32:22modernization with the latest SQL Server

2:32:24capabilities in the cloud SQL managed

2:32:27instance provides an entire SQL Server

2:32:29instance within a managed service so you

2:32:32can continue to use familiar tools and

2:32:34SQL Server features like cross database

2:32:36queries and Link servers SQL managed

2:32:38instance maintains the highest

2:32:40compatibility labels so you can move

2:32:42your on premises workloads without

2:32:44worrying about application comp

2:32:46compatibility of performance changes and

2:32:48the third member in the family is azure

2:32:50SQL Edge so Azure SQL Edge is an

2:32:53optimized relational database engine

2:32:55Geared for iot and iot Edge deployments

2:32:58it provides capabilities to create a

2:33:00high performance data storage and

2:33:02processing layer for iot applications

2:33:04and solutions Azure SQL Edge provides

2:33:07capabilities to stream process and

2:33:09analyze relational and non-relational

2:33:11data such as Json graph and time series

2:33:13data which makes it the right choice for

2:33:15a VAR of modern iot applications Azure

2:33:18SQL Edge is built on the latest version

2:33:21of the SQL Server database engine which

2:33:24provides industry-leading performance

2:33:26security and quering processing

2:33:28capabilities since Azure SQL Edge is

2:33:30built on the same engine as SQL server

2:33:32and Azure SQL it provides the same

2:33:34transact SQL programming surface area

2:33:37that makes development of applications

2:33:38or Solutions easier and faster and makes

2:33:41application probability between iot Edge

2:33:43devices data centers and Cloud straight

2:33:45forward

2:33:46so there are two different deployment

2:33:48models in Azure SQL H so the first one

2:33:52is connected deployment through Azure

2:33:54iot Edge azour SQL Edge is available on

2:33:57the Azure Marketplace and can be

2:33:59deployed as a module for Azure iot Edge

2:34:02second is disconnected deployment so

2:34:04Azure SQL Edge container images can be

2:34:06pulled from Docker Hub and deployed

2:34:08either as a standalone doer container or

2:34:10a kubernetes cluster so some of the

2:34:12features of azour SQL EDR built in data

2:34:14streaming and time SE within database

2:34:17machine learning and graph features for

2:34:18low latency analytics then data

2:34:20processing at the edge for online

2:34:22offline and hybrid environments to

2:34:24overcome latency and bandwidth

2:34:26constraints next is deploy an update

2:34:28from the azuro portal or enterprise

2:34:30portal for consistent security and trunk

2:34:33key management last one is simplified

2:34:35pricing with no upfront cost and

2:34:37subscription offers as low as us $60 per

2:34:41year per device so the fourth and the

2:34:44major family member of azure SQL family

2:34:46is azure SQL database which we're going

2:34:49to briefly understand further first

2:34:51let's understand why one need an Azure

2:34:53SQL database extensively I will tell you

2:34:56the top five benefits that companies are

2:34:58realizing with SQL database so the first

2:35:01one is scalability and Beyond flexible

2:35:03service plans for SQL database meet the

2:35:06need for both big and small business

2:35:08users SQL is no longer Way Out Of Reach

2:35:12for smaller operations because the

2:35:13pricing structure allows users to pay as

2:35:16little as

2:35:17$4.99 per database per month with a

2:35:20maximum storage set at 150 GB per

2:35:23database that's a lot of space for very

2:35:26small cost second is high speed and

2:35:29minimal downtime so high availability

2:35:31architecture mean High speeed

2:35:33connectivity and data retrival as well

2:35:35as low downtime at your organization

2:35:38there's nothing worse than stopping

2:35:40business because your technology can't

2:35:42keep up and that is no longer a problem

2:35:44with SQL database secondly companies can

2:35:47add application instances as needed

2:35:49through sheding for example shedding is

2:35:52a type of database partitioning that

2:35:54separates very large databases into

2:35:56smaller faster more easily managed Parts

2:35:59called Data shards not only can you spin

2:36:02nodes up and down on demand you can

2:36:04leverage a federation infrastructure to

2:36:06scale more easily without affecting

2:36:08other areas of the server SQL azur

2:36:11Federation data migration visard can

2:36:14further automate this process which

2:36:16impacts your organization and the

2:36:17employees much less lastly there are

2:36:20multiple levels of implementation that

2:36:22you can benefit from if you just need a

2:36:24website and a database you can hitch a

2:36:27SQL Azure instance to an Azure website

2:36:29and you are done if you need a

2:36:31full-blown virtual machine now or even

2:36:34down the road you can get that as well

2:36:36you can even use a locally deployed

2:36:38instance of SQL server in the virtual

2:36:40machine instead of SQL Azure these

2:36:43implementation options help make your

2:36:45comp company more adaptable to the

2:36:48inevitable changes it under goes on a

2:36:50regular basis with SQL Azure you are not

2:36:53stuck you are a foundation that

2:36:55encourages growth while working with it

2:36:57third is improved usability so SQL

2:37:00developers are familiar with all things

2:37:03SQL and SQL database can be updated with

2:37:05SQL CMD or the SQL Server management

2:37:08Studio better yet there is no coding

2:37:10required using a standard SQL it's much

2:37:13easier to manage database systems

2:37:16without having to write or update a huge

2:37:18amount of code fourth time is on your

2:37:21side with no administrative duties on

2:37:23your physical location employees can

2:37:25take time for strategic work to advance

2:37:28grow all around business success when

2:37:31your database is hosted in the cloud you

2:37:33don't have to deal with setting up SQL

2:37:35Server appropriating databases and

2:37:37dealing with physical machine

2:37:39maintenance and upkeep all of this

2:37:41results in better alignment of your

2:37:43organization and ultimately more time on

2:37:45your site Fifth and the last one is easy

2:37:48to use migration tools ramp up time with

2:37:51SQL database is now easier than ever and

2:37:53free SQL data synchronization allows you

2:37:56to either synchronize your SQL Server

2:37:59stored data or migrate that data without

2:38:01having to worry about the cost to

2:38:03migrate by syncing gigabyte size tables

2:38:06now that you know why we need aour SQL

2:38:08database let's reply understand what

2:38:10actually it is the basic fundamental

2:38:12building block that is available in a is

2:38:15the SQL data datase Azure SQL database

2:38:17is fully managed platform as a service

2:38:19database engine that handles most of the

2:38:21database management functions such as

2:38:23upgrading patching backups and

2:38:25monitoring without user involvement

2:38:28Azure SQL database is always running on

2:38:30the latest stable version of the SQL

2:38:33Server database engine and ped OS with

2:38:3599.99% availability platform as a

2:38:38service capabilities that are built into

2:38:40Azure SQL database enable you to focus

2:38:42on the domain specific database

2:38:44Administration and optim optimization

2:38:46activities that are critical for your

2:38:47business with Azure SQL database you can

2:38:50create a highly available and high

2:38:52performance data storage layer for the

2:38:54applications and Solutions in azour SQL

2:38:57database can be the right choice for a

2:38:59variety of modern Cloud applications

2:39:01because it enables you to process both

2:39:03relational data and non-relational

2:39:04structures such as graphs Json spatial

2:39:07and XML Azure SQL database is based on

2:39:10the latest stable version of the

2:39:12Microsoft SQL Server database engine you

2:39:15can use Advanced query processing

2:39:16features such as high performance

2:39:18inmemory Technologies and intelligent

2:39:20query processing in fact the newest

2:39:23capabilities of SQL Server are released

2:39:25first to SQL database and then to SQL

2:39:28Server itself you get the newest SQL

2:39:30Server capabilities with no overhead for

2:39:33patching or upgrading tested across

2:39:35millions of databases SQL database

2:39:38enables you to easily Define and scale

2:39:40performance within two different

2:39:42purchasing models and that we will

2:39:44discuss further so Microsoft handles all

2:39:46patching and updating of the SQL and

2:39:48operating system code you don't have to

2:39:50manage the underlying infrastructure so

2:39:52now let's understand the deployment

2:39:54models Azure SQL databas provides the

2:39:56following deployment options for a

2:39:58database so the first one is managed

2:40:00instance this is primarily targeted

2:40:02towards on premises customers in case if

2:40:05we already have a SQL Server instance to

2:40:08on premises Data Center and you want to

2:40:10migrate that into Azure with minimum

2:40:12changes to our application and the

2:40:13maximum compatibility then new will go

2:40:16to the manage instance second is single

2:40:19database so single database represents a

2:40:21fully managed isolated database you

2:40:24might use this option if you have modern

2:40:26Cloud applications and microservices

2:40:28that need single reliable data source a

2:40:30single database is similar to a

2:40:32contained database in the SQL Server

2:40:34database engine last one is elastic pool

2:40:37so elastic pool is a collection of

2:40:39single databases with a shred set of

2:40:41resources such as CPU or memory single

2:40:43databases can be moved into and out of

2:40:45an elastic pool now let's understand the

2:40:48purchasing models so SQL database offers

2:40:50the following purchasing models you can

2:40:52see here so the first one is vcore based

2:40:54purchasing model which is new and it

2:40:56offers a totally different approach to

2:40:58sizing your database it is easier to

2:41:00translate local workloads to a Vore

2:41:02based model because the components are

2:41:04what we are used to the vcore based

2:41:07model lets you choose the number of vour

2:41:09the amount of memory and the amount and

2:41:12speed of storage the Vore based

2:41:14purchasing model also allows you to use

2:41:17Azure hybrid benefit for SQL Server to

2:41:19gain cost savings then the next one is

2:41:22DTU based purchasing model so the D2

2:41:24based purchasing model offers a bland of

2:41:26compute memory and input output

2:41:28resources in three service tires to

2:41:31support light to heavy database

2:41:32workloads compute sizes within each TI

2:41:35provide a different mix of these

2:41:37resources to which you can add

2:41:39additional storage resources as you can

2:41:42see from the following diagram the DTU

2:41:44model offers a pre-configured and

2:41:46predefined amount of compute resources

2:41:49vcore is all about independent

2:41:51scalability where you can look into a

2:41:54specific area such as the CPU core count

2:41:56and memory resources something that you

2:41:59cannot control at the same granular

2:42:00level when using the DTU based model so

2:42:03the third one is the serverless model

2:42:05which automatically scales compute based

2:42:08on workload demand and builds for the

2:42:10amount of compute used per second the

2:42:12serverless compute Tire also

2:42:14automatically pauses data bases during

2:42:16inactive periods when only storage is

2:42:18built and automatically resumes

2:42:20databases when activity returns now

2:42:23let's see the service tries for Azure

2:42:25SQL database so the first one is general

2:42:27purpose or standard model it is based on

2:42:30a separation of computing and storage

2:42:32service this architecture model depends

2:42:35on the high availability and reliability

2:42:37of azure premium storage that

2:42:40transparently copies database files and

2:42:42guarantees for zero data loss if underly

2:42:45infrastructure failure happens second is

2:42:48business critical of premium service Tri

2:42:50model it is based on a cluster of

2:42:53database engine processes both the SQL

2:42:55database engine process and underlying

2:42:58MDF or ldf files are placed on the same

2:43:00node with locally attached SSD storage

2:43:03providing low latency to a workload High

2:43:06availability is implemented using

2:43:08technology similar to SQL Server always

2:43:11on availability groups the third one is

2:43:14hyperscale Service Tire model it is the

2:43:17newest Service Tire in the vord based

2:43:18purchasing model this tire is a highly

2:43:21scalable storage and compute Performance

2:43:24Tire that leverages the Azure

2:43:25architecture to scale out the storage

2:43:27and compute resources for an Azure SQL

2:43:30database beyond the limits available for

2:43:32the general purpose and business

2:43:34critical service tires now let's

2:43:36understand the self-contained services

2:43:37in azour SQL database so Azure SQL

2:43:40database is a database as a platform

2:43:41service designed for applications that

2:43:44will use database as selfcontain service

2:43:46databases can be grouped together to

2:43:49simplify management options or share the

2:43:51resources there are different options

2:43:53that can be used to bound databases in

2:43:56the group so the first one is databases

2:43:58in logical server so logical server is a

2:44:01default container for Azure SQL database

2:44:04logical server enables you to perform

2:44:05administrative tasks across multiple

2:44:07databases including a specifying Regions

2:44:10login information firewall rules

2:44:12auditing thread detection and failover

2:44:14groups all databases within the server

2:44:17are self-contained with independent

2:44:19service tries that can be specified per

2:44:22each database each database can be

2:44:24independently scaled up or down by

2:44:26changing performance tries on the

2:44:28database which will not affect other

2:44:30databases databases cannot share

2:44:33resources and each database has

2:44:34guaranteed and predictable performance

2:44:36defined by its own service style some

2:44:39server level specific features such as

2:44:42cross database quering linked servers

2:44:44SQL agent service broker or CLR are not

2:44:47supported in Azure SQL database placed

2:44:49in logical servers second one is

2:44:51databases in elastic pool so databases

2:44:54need to share resources can be stored in

2:44:57elastic pools instead of The Logical

2:44:59server all databases within the elastic

2:45:02pool share the same resources associated

2:45:04with the elastic pool label currently

2:45:06there are three service ties in the

2:45:07elastic pools basic standard and premium

2:45:10databases within the elastic pools

2:45:12cannot have different service tires

2:45:14because they share resources that are

2:45:16assigned to the entire pool resources

2:45:18usage in one database might affect

2:45:20others however you can specify Reserve

2:45:22performance for the database in the pool

2:45:24that will guarantee a minimal amount of

2:45:26resources that the database can have

2:45:29this model is a good choice for

2:45:30databases that have performance Peaks or

2:45:32heavy usage in different time periods

2:45:34because the amount of resources

2:45:36associated with the pool can be assigned

2:45:38to the databases that need them while

2:45:40the others are inactive elastic pool

2:45:43model is designed for resource sharing

2:45:45and it is still does not support server

2:45:48level features such as SQL agents

2:45:50service broker Etc these are other

2:45:53mechanisms that can be used as a

2:45:55replacement of these features such as

2:45:57elastic jobs and elastic queries instead

2:46:00of some server label features so now

2:46:03let's look at some of the key features

2:46:04of azure SQL database the first one is

2:46:08extensive monitoring and alerting

2:46:09capabilities so Azure SQL database

2:46:12provides Advanced monitoring and trouble

2:46:14shooting features that help you get

2:46:16deeper insights into workload

2:46:18characteristics these features and tools

2:46:20include the buil-in monitoring

2:46:22capabilities provided by the latest

2:46:24version of the SQL Server database

2:46:26engine they enable you to find realtime

2:46:28performance insights also platform as a

2:46:31service monitoring capabilities provided

2:46:33by AO that enable you to Monitor and

2:46:35troubleshoot a large number of database

2:46:37instances query store a buil-in SQL

2:46:40Server monitoring feature records the

2:46:42performance of your queries in real time

2:46:45and enables you to identify the

2:46:47potential performance issues and the top

2:46:49resource consumers automatic tuning and

2:46:52recommendation provides advice regarding

2:46:54the queries with the regress performance

2:46:56and missing or duplicated indexes

2:46:59automatic tning in SQL database enables

2:47:01you to either manually apply the script

2:47:04that can fix the shoes or let SQL

2:47:06database apply the fix SQL database can

2:47:09also test and verify that the fix

2:47:11provides some benefit and retain or

2:47:13reward the change depending on the

2:47:15outcome in addition to query store and

2:47:17automatic twinning capabilities you can

2:47:20use extended DMVs and XC event to

2:47:23monitor the workload performance Azure

2:47:26provides buil-in performance monitoring

2:47:27and alerting tools combined with

2:47:29performance ratings that enable you to

2:47:31monitor the status of thousands of

2:47:33databases using these tools you can

2:47:35quickly asset the impact of scaling up

2:47:37or down based on your current or

2:47:39projected performance needs additionally

2:47:42SQL database can emit metrics and resour

2:47:44Source logs for easier monitoring you

2:47:47can configure SQL database to store

2:47:49resource usage workers and sessions and

2:47:52connectivity into one of these Azure

2:47:54resources so these resources are first

2:47:57one is azure storage so for achieving

2:48:00vast amounts of telemetry for a small

2:48:02price so the second one is azure event

2:48:05hubs for integrating SQL database

2:48:08Telemetry with a custom monitoring

2:48:09solution for hot pipelines and the third

2:48:12one is azure monitor logs for a built in

2:48:15monitoring solution with reporting

2:48:16alerting and mitigating capabilities so

2:48:19the second feature is availability

2:48:21capabilities so aure SQL database

2:48:23enables your business to continue

2:48:25operating during disruptions in a

2:48:27traditional SQL Server environment you

2:48:29generally have at least two machines

2:48:31locally set up these machines have

2:48:33synchronously maintained copies of the

2:48:34data to protect against a failure of a

2:48:37single machine or component this

2:48:39environment provides High availability

2:48:41but it doesn't protect against a natural

2:48:43disaster destroying your your data

2:48:45center Disaster Recovery assumes that a

2:48:47catastrophic event is geographically

2:48:49localized enough to have another machine

2:48:51or set of machines with a copy of your

2:48:53data for far away in SQL Server you can

2:48:56use always on availability groups

2:48:59running in asynchronous mode to get the

2:49:01capability people often don't want to

2:49:04wait for replication to happen that far

2:49:06away from committing a transaction so

2:49:09there's potential for data loss when you

2:49:10do unplanned failovers so databases in

2:49:13the premium and business this critical

2:49:15service tires already do something

2:49:17similar to the synchronization of an

2:49:19availability group databases in lower

2:49:21service tires provideed tendency through

2:49:23storage by using a different but

2:49:25equivalent mechanism like buil-in logic

2:49:28helps protect against a single machine

2:49:29failure the active Geo replication

2:49:32feature gives you the ability to protect

2:49:34against disaster where a whole region is

2:49:36destroyed Azure availability zones tries

2:49:39to protect against the outage of a

2:49:41single Data Center building within a

2:49:43single region it helps helps you protect

2:49:45against the loss of power or network to

2:49:47a building in SQL database you place the

2:49:50different replicas in different

2:49:51availability zones in fact the service

2:49:54label management of azour powered by a

2:49:56Global Network of Microsoft manag data

2:49:58centers helps keep your app running 24/7

2:50:01the Azure platform fully manages every

2:50:03database and it guarantees no data loss

2:50:06and a high percentage of data

2:50:08availability Azure automatically handles

2:50:10patching backups replication failure

2:50:13detection under potential Hardware

2:50:15software or network failures deploying

2:50:18bug fixes failovers database upgrades

2:50:21and other maintenance tasks standard

2:50:23availability is achieved by a separation

2:50:25of compute and storage layers premium

2:50:27availability is achieved by integrating

2:50:29compute and storage on a single node for

2:50:32performance and then implementing

2:50:33technology similar to always on

2:50:35availability groups in addition SQL

2:50:38database provides built-in business

2:50:40continuity and Global scalability

2:50:41features these include automatic backups

2:50:45so SQL database automatically performs

2:50:47full differential and transaction log

2:50:49backups of database to enable you to

2:50:52restore to any point in time for single

2:50:54databases and pool databases you can

2:50:57configure SQL database to store full

2:50:59database backups to Azure storage for

2:51:01long-term backup retention for managed

2:51:03instances you can also perform copy only

2:51:06backups for long-term backup retention

2:51:08second is point in time restor so all

2:51:12SQL database deployment options Support

2:51:14Recovery to any point in time within the

2:51:16automatic backup retention period for

2:51:18any database third one is active Geo

2:51:21replication the single database and pool

2:51:23databases option allow you to configure

2:51:25up to four readable secondary databases

2:51:28in either the same or globally

2:51:29distributed AO data centers for example

2:51:32if you have a service as a platform

2:51:34application with a catalog database that

2:51:36has a high volume of concurrent read

2:51:38only transactions use active GE

2:51:41application to enable global read scale

2:51:43this removes bottleneck on the primary

2:51:45data due to read workloads for managed

2:51:48instances use autof fail groups fourth

2:51:51is autof fail groups all SQL database

2:51:53deployment options allow you to use

2:51:54failover groups to enable High

2:51:56availability and load balancing at

2:51:58global scale this includes transparent

2:52:00GE application and failover of large

2:52:02sets of databases elastic pools and

2:52:04managed instances failover groups enable

2:52:06the creation of globally distributed

2:52:08Service as a platform applications with

2:52:10minimal Administration overhead this

2:52:12leaves all the complex monitoring

2:52:14routing and failover orchestrations to

2:52:16SQL database Fifth and the last one is

2:52:19Zone indendent databases SQL database

2:52:22allows you to provision premium or

2:52:23business critical databases or elastic

2:52:25pools across multiple availability zones

2:52:28because these databases and elastic

2:52:30pools have multiple rendent replicas for

2:52:32high availability placing these replicas

2:52:34into multiple availability zones

2:52:36provides higher resilience this includes

2:52:38the ability to recover automatically

2:52:41from the data center scale features

2:52:42without data loss so the next feature is

2:52:45built-in intelligence with SQL database

2:52:48you get buil-in intelligence that helps

2:52:50you dramatically reduce the costs of

2:52:52running and managing databases and that

2:52:54maximizes both performance and security

2:52:57of your application running millions of

2:52:59customers workloads around the clock SQL

2:53:01database collects and processes a

2:53:03massive amount of telemetry data while

2:53:06also fully respecting customer privacy

2:53:08various algorithms continuously evaluate

2:53:10the Telemetry data so that service can

2:53:12learn and adapt with your applic ation

2:53:15so in its process of work first step is

2:53:17automatic performance monitoring and

2:53:19tuning so SQL database provides detailed

2:53:21insight into the queries that you need

2:53:23to monitor SQL database learns about

2:53:26your database patterns and enables you

2:53:27to adapt your database schema to your

2:53:30workload SQL database provides

2:53:32Performance Tuning recommendations where

2:53:34you can review tuning actions and apply

2:53:36them however constantly monitoring a

2:53:38database is hard and tedious task

2:53:40especially when you are dealing with

2:53:42many databases intelligent insights does

2:53:45this job for you by automatically

2:53:46monitoring SQL database performance at

2:53:48scale it informs you of performance

2:53:51degradation issues it identifies the

2:53:53root cause of each issue and it provides

2:53:55performance Improvement recommendations

2:53:57when possible managing huge number of

2:53:59databases might be impossible to do

2:54:01efficiently even with all available

2:54:03tools and reports that SQL database and

2:54:05Azure provide instead of monitoring and

2:54:08tuning your database manually you might

2:54:10consider delegating some of the

2:54:12monitoring and tuning actions to SQL

2:54:13database by using automatic tuning SQL

2:54:16database automatically applies

2:54:18recommendation tests and verifies each

2:54:20of its tuning actions to ensure the

2:54:22performance keeps improving this way SQL

2:54:24database automatically adapts to your

2:54:26workload in a controlled and Safe Way

2:54:28automatic tning means that the

2:54:30performance of a database is carefully

2:54:32monitored and compared before and after

2:54:34every Twining action if the performance

2:54:37doesn't improve the Twining action is

2:54:39diverted many of our partners that run

2:54:41Service as a platform multitined apps on

2:54:43top of SQL database are relying on

2:54:46automatic Performance Tuning to make

2:54:48sure their applications always have

2:54:50stable and predictable performance for

2:54:52them this feature tremendously reduces

2:54:55the risk of having a performance

2:54:56incident in the middle of the night in

2:54:59addition because part of their customer

2:55:01base also uses SQL Server they are using

2:55:03the same indexing recommendations

2:55:05provided by SQL database to help their

2:55:07SQL Server customers two automatic tning

2:55:10expects are available in SQL database so

2:55:12first one is automatic index management

2:55:15which identifies indexes that should be

2:55:16added in your database and indexes that

2:55:18should be removed second one is

2:55:20automatic plan correction which

2:55:21identifies problematic plans and fixes

2:55:23SQL plan performance problems so the

2:55:25next step in the process of work is

2:55:27adaptive query processing you can use

2:55:29adaptive query processing including

2:55:31interl execution for multi statement

2:55:33table valued functions batch mode memory

2:55:35Grant feedback and batch mode adictive

2:55:37joints each of these adaptive query

2:55:39processing features apply similar learn

2:55:41and adapt techniques helping further

2:55:43address performance issues related to

2:55:45historically inable optimization

2:55:48problems so the next feature is Advanced

2:55:50security and compliance SQL database

2:55:53provides a range of built-in security

2:55:55and compliance features to help your

2:55:56application meet various security and

2:55:58compliance requirement note this that

2:56:00Microsoft has a certified Azure SQL

2:56:02database against number of compliance

2:56:03standards for more information see the

2:56:05Microsoft Azure trust Center where you

2:56:07can find the most current list of SQL

2:56:09database compliance certifications so

2:56:11built-in security and compliance

2:56:13features so the there are certain

2:56:14built-in security and compliance

2:56:16features so the first one is Advanced

2:56:18threat protection Azure Defender for SQL

2:56:20is a unified package for advanced SQL

2:56:23security capabilities it includes

2:56:25functionality for managing your database

2:56:27vulnerabilities and detecting anomalous

2:56:30activities that might indicate a threat

2:56:32to your database it provides a single

2:56:34location for enabling and managing these

2:56:36capabilities so there are two kinds of

2:56:39assessment in advanced protection the

2:56:41first one is vulnerability assessment

2:56:43this service can discover track and help

2:56:45you remediate potential database

2:56:47vulnerabilities it provides visibility

2:56:49into your Security State and includes

2:56:52actionable steps to resolve security

2:56:54issues and enhance your database

2:56:56fortifications second is threat

2:56:57protection so this feature detect anous

2:57:00activities that indicate unusual and

2:57:02potentially harmful attempts to access

2:57:04or exploit your database it continuously

2:57:06monitors your database for suspicious

2:57:08activities and provides immediate

2:57:09security alerts on potential

2:57:11vulnerabilities SQL injection attacks

2:57:13and enous database access patterns

2:57:16threat detection alert provide details

2:57:18of the suspicious activity and recommend

2:57:21action on how to investigate and

2:57:22mitigate the threat so under this the

2:57:25second sub feature is auditing for

2:57:26compliance and security so auditing

2:57:28tracks database events and writes them

2:57:31to an audit log in your Azure storage

2:57:33account auditing can help you maintain

2:57:35Regulatory Compliance understand

2:57:36database activity and gain insight into

2:57:39discrepancies and anomalies that might

2:57:41indicate business concerns or suspected

2:57:43security violations the third one is

2:57:45data encryption SQL database helps

2:57:47secure your data by providing encryption

2:57:50for data at rest it uses transparent

2:57:52data encryption for data in use it uses

2:57:55always encrypted fourth feature is data

2:57:57Discovery and classification data

2:57:59Discovery and classification provides

2:58:00capabilities built into Azure SQL

2:58:03database for discovering classifying

2:58:05labeling and protecting the sensitive

2:58:07data in your databases it provides

2:58:09visibility into your database

2:58:11classification State and tracks the

2:58:12access to sensitive data within the

2:58:14database and Beyond its borders and the

2:58:17last one is azure active directory

2:58:19integration and multiactor

2:58:20authentication so SQL database enables

2:58:23you to centrally manage identities of

2:58:25database user and other Microsoft

2:58:27services with azured active directory

2:58:28integration this capability simplifies

2:58:31permission management and enhances

2:58:33security Azure active directory supports

2:58:35multiactor authentication to increase

2:58:37data and application security while

2:58:39supporting a single signin process the

2:58:41last major feature is easy to use tools

2:58:44SQL database makes building and

2:58:46maintaining applications easier and more

2:58:48productive SQL database allows you to

2:58:50focus on what you do best building great

2:58:52applications so you can manage and

2:58:54develop an SQL datab by using tools and

2:58:56skills you already have so the first

2:58:58tool is the Azure portal a web- based

2:59:01application for managing all Azure

2:59:03Services second is azure data Studio a

2:59:05crossplatform database tool that runs on

2:59:07Windows Mac OS and Linux third is SQL

2:59:10Server management Studio a free

2:59:12downloadable client application

2:59:14for managing any SQL infrastructure from

2:59:16SQL Server to SQL database fourth is SQL

2:59:20Server data Tools in Visual Studio a

2:59:22free downloadable client application for

2:59:24developing SQL Server relational

2:59:25databases databases in aure SQL database

2:59:29integration service packages analysis

2:59:31service data models and Reporting

2:59:33Services reports and the last tool is

2:59:35Visual Studio code a free downloadable

2:59:38open-source code editor for Windows Mac

2:59:40OS and Linux it support extensions

2:59:43including the mssql extension for

2:59:45querying Microsoft SQL Server Azure SQL

2:59:48database and Azure SS analytics now

2:59:51let's look at some of the use cases for

2:59:53Azure SQL database so the first one is

2:59:55developer or test environment it's an

2:59:58important use case for replicating or

3:00:00migrating data to SQL hosted on Azure is

3:00:02for developer or test environments

3:00:05before deploying to the production

3:00:06environment it is pertinent that the

3:00:08data is tested against developer or test

3:00:10environments Azure SQL databases can act

3:00:13as a Target for just such environments

3:00:15the life production environment can be

3:00:17replicated to the developer or test

3:00:19environment using a database copy second

3:00:21is business continuity one of the most

3:00:23important use cases for SQL on aure is

3:00:25using it as a Dr Target to maintain

3:00:28business continuity Azure SQL databases

3:00:30can provide an SLA of up to

3:00:3399.99% by maintaining several copies of

3:00:36the data this provides business

3:00:38continuity as it allows you to restore

3:00:40Geor redundant copies of the data or use

3:00:43active Geo redundant copies as failover

3:00:45points in case of outages at data

3:00:48centers or in regions besides Azure SQL

3:00:51databases you can also use availability

3:00:53groups to fulfill business continuity

3:00:55demands not only can you use

3:00:57availability groups in a SQL virtual

3:00:59machines but also use Azure SQL virtual

3:01:02machine instances as a target for high

3:01:04availability and Disaster Recovery third

3:01:06is scaling our readon workloads apart

3:01:10from providing BC or Dr capabilities

3:01:12active gec ation can also be used to

3:01:15offload readon workload such as

3:01:17reporting jobs to secondary copies you

3:01:19can also extend on premises SQL Server

3:01:22instances using readable always on

3:01:24replicas fourth is backup and restore

3:01:27Azure SQL databases are backed up

3:01:29automatically on a regular basis and

3:01:31there are no storage costs for up to

3:01:34200% of the maximum provision database

3:01:36storage you can restore backups to any

3:01:38point in time going back to the

3:01:40retention period which is determined by

3:01:42the Azure SQL service Tri in use on

3:01:45premises SQL Server databases and

3:01:47transaction logs can also be bagged up

3:01:49directly to Azure using the backup to

3:01:51URL feature and stored in Azure storage

3:01:54Azure SQL databases can also be stored

3:01:57on local storage by exporting them to

3:01:59backpack files means BC PSC files and

3:02:02the last use case is Advanced analytics

3:02:04another important reason for hosting SQL

3:02:06in Azure is to make use of azure's

3:02:09advanced analytics platforms such as

3:02:10Azure storage blob and Azure data L

3:02:13store common scenario with Advanced

3:02:15analytics is when users reference data

3:02:17from various data sources use Azure data

3:02:20Lake store as the stacking area perform

3:02:22transformation activities using Hive or

3:02:25spark and finally load the data into

3:02:27Azure data warehouse for business

3:02:29intelligence and Reporting now that you

3:02:31have a theoretical understanding of

3:02:33azure SQL database let's now see a

3:02:35deployment of a SQL database service on

3:02:38Microsoft Azure we will also connect

3:02:40this database with SQL management server

3:02:42as well as with Azure data studio so

3:02:45let's move ahead can just simply go to

3:02:47the Azure

3:02:48portal if you don't have an account on

3:02:51Azure what you can do is you can create

3:02:52a free account like can start free here

3:02:54and you will get a 12 months free

3:02:56services access as well as if you are

3:02:58from India the currency will be around

3:03:00like 14,500 INR means Indian rupees free

3:03:03currency we will get for free use that's

3:03:06the amount you will get so I already

3:03:07have an account so I will just log in

3:03:09from

3:03:10here no I don't want to I will just go

3:03:12to Microsoft here this is the portal so

3:03:16I don't want to buy see these are the

3:03:17credits it is showing you get 14,500

3:03:20credits I have used some of the credits

3:03:22and this much I left so what you can do

3:03:24is you can go to SQL databases right now

3:03:27because is I have a pin here but what

3:03:29you can do is you can go to SQL

3:03:30databases from here like here or you can

3:03:33search it from here SQL database yeah

3:03:36now let's create the database from

3:03:38here so free trial is here we have

3:03:41chosen this then we can choose the

3:03:43resource Group can create one group so

3:03:46let's create with a new one so let's

3:03:48give the name as resource SQL demo okay

3:03:52so a new resource has been created then

3:03:55enter database name can give some name

3:03:57to like demo

3:03:59SQL so yeah that's okay so give the

3:04:02server name also so create a new server

3:04:05a server name so let's give it like

3:04:08remember these names SQL demo server I

3:04:10have given so the specified server name

3:04:12is already use okay we can give like SQL

3:04:15demo server 662 but you have to remember

3:04:19this don't forget this like login

3:04:21information okay so you can give some

3:04:23server admin login so Azure admin will

3:04:26be okay I think yeah now let's move ahe

3:04:29to give a password don't forget these

3:04:31information we will require these

3:04:33information later on in this process and

3:04:36choose your region also so my region is

3:04:38Central India so it will be Asia Pacific

3:04:41Central India just

3:04:44yeah so okay the next process if you

3:04:47want to use SQL elastic pool right now

3:04:49we don't require so we will choose no if

3:04:51you want to then you can choose yes then

3:04:53we can configure the database and

3:04:54storage also so these are the plans like

3:04:56I have explained you basic standard

3:04:58premium the service tires are there

3:05:00similarly purchasing models are there

3:05:01voco purchasing models are there general

3:05:03purpose hypers scale business critical

3:05:04the also have explained you so like

3:05:07whichever you choose on the basis of

3:05:08that it shows the price so right now

3:05:10it's costing 2949 but we don't require

3:05:12that much bigger so you can just you can

3:05:14choose it from here also like not

3:05:16available because it's the free tire

3:05:18account that's why we don't have a

3:05:19premium version so can use the standard

3:05:21one so in this standard one if you

3:05:23decrease the storage then it will also

3:05:25decrease the price not in this manner

3:05:28like from you can go to basic here you

3:05:30get 2 GB and the price decreased you can

3:05:33choose for 1 GB also but price will

3:05:34remain the same so you can go to 2 gb so

3:05:36we will choose the basic one because we

3:05:38don't require that much of configuration

3:05:39so just apply from here just review and

3:05:42create let simple review plan so yeah

3:05:45create so it's getting created the

3:05:48process it takes certain time so that's

3:05:51why taking a Time few minutes it

3:05:55take so you can see the deployment is

3:05:57complete so you can go to Resource SQL

3:05:59demo so we have came to Resource here

3:06:02you see that server is created but you

3:06:04have to refresh once again just a second

3:06:06where is the database let's refresh

3:06:09again I don't know let's get created

3:06:12usually yeah so you can see have to

3:06:14refresh it so it will takes a little

3:06:16time so database is also created now so

3:06:19what we can do from here is first step

3:06:21is you have to create the firewall so

3:06:23you can go to database so you have to

3:06:26set the firewall set firewall servers

3:06:29you can go from here set the fireable so

3:06:31you have to choose the like all these

3:06:33client IP is given from start IP and end

3:06:35IP so you can just you have to add the

3:06:37client IP so it has been added so you

3:06:40have to give it from 0 to 255 complete

3:06:42you have to

3:06:44so

3:06:45save and just go back then you can see

3:06:48different cool features are given on the

3:06:49left side of this dashboard so you go to

3:06:52the query editor this is here you can

3:06:53like connect your database to the

3:06:55browser and you can configure and then

3:06:57there is computer and storage if you

3:06:58have like you can yeah not a problem so

3:07:01computer and storage is given if you

3:07:03want to like change your plan purchasing

3:07:04model and everything then you can change

3:07:06it from here again like all the things

3:07:08are given here you can again change it

3:07:10it's not a problem then there are like

3:07:12connection Str no problem then you can

3:07:14go to connection strings where you can

3:07:15connect your and make connection to the

3:07:17strings like there's ado.net or if you

3:07:20are working on Java then jdbc is there

3:07:22obbc is there PHP go everything is there

3:07:25it's quite cool feature then there are

3:07:26synchronous to other database where you

3:07:28can synchronize to other databases also

3:07:30then also you can like add Azure search

3:07:32also then there are different security

3:07:34features are there advanc security

3:07:36features are there for data also

3:07:37Advanced Data security features are

3:07:39given here now what we can do major

3:07:41thing we have to do is we have to

3:07:42connect it to the management SQL Server

3:07:44so if you don't have a management SQL

3:07:45Server you can download it from here you

3:07:46can just give management SQL server and

3:07:50just go here and download it from here

3:07:52it's given you can just click here and

3:07:54it will get downloaded it will ask for a

3:07:56restart it will restart the computer and

3:07:57all the settings will be saved in your

3:07:59computer then you can start using SQL

3:08:01Server I already have downloaded it so I

3:08:03already have downloaded it so we can

3:08:05just go here have it print here so

3:08:07Microsoft SQL Server management so yeah

3:08:10you have to give your server name so

3:08:12let's see the server name first here it

3:08:15is no from database only so server name

3:08:19is given so you can just copy it from

3:08:20here paste it here you can just select

3:08:23the SQL Server authentication here and

3:08:25give you a login also so remember the

3:08:28login I told you remember the login

3:08:30password so my is your admin password

3:08:34can just connect it from here so yeah we

3:08:37got it here so we can just use the

3:08:39databases can see the demo SQL we can

3:08:41find the our database here here and

3:08:43there different tables and everything is

3:08:44there so we can just start the new query

3:08:47from here can create the table so let's

3:08:50create the table create table with the

3:08:52name Person

3:09:02persons so yeah it will be enough let's

3:09:04now execute it you can see command

3:09:07completed successfully we can now insert

3:09:09values into it so like before that we

3:09:11can view in tables also the table where

3:09:13we created the name of persons expending

3:09:16so yeah you can see persons is created

3:09:18so let's now

3:09:20insert values you can give like give my

3:09:23name so

3:09:2826 from person let's execute now oh

3:09:33sorry it's a little mistake let's

3:09:35execute now so yeah you can see we have

3:09:37got the table contents everything is

3:09:38there now we can also see for Azure data

3:09:41studio also so we will just just go to

3:09:43aor Studio it's like you have a visual

3:09:46studio code also similarly there's aor

3:09:49Studio can start the new connection then

3:09:51we have to give the server name here

3:09:53just like we have the server name there

3:09:55so what was the server name let's copy

3:09:58it from here so yeah just give here and

3:10:01then we will choose the SQL login

3:10:02username Azure admin the password yeah

3:10:06we can just remember the password not a

3:10:08problem then database we have to select

3:10:10so what was the name of the database

3:10:13demo SQL that's why it's showing here

3:10:15yeah so we connected from

3:10:17here yeah we are here like different

3:10:20features are given like new query new

3:10:22notebook just like you have in Visual

3:10:24Studio code in that way V is given and

3:10:26all the other like it has a very good UI

3:10:28also here we go to the database person

3:10:30it is showing just maximize it okay yeah

3:10:34so it's open now just like you can give

3:10:36it here Al all the details it is showing

3:10:38like it's been executed and details are

3:10:40shown here this is how it looks like

3:10:43also one cool thing about AA studio is

3:10:46you can go here and you can see like it

3:10:48actually remembers all the databases

3:10:52like this one it remembers this one I

3:10:53have closed actually this database I

3:10:55have closed I have created it some time

3:10:57back but it still remembers databases if

3:10:59you give the Azure like server ID now so

3:11:02like it will remember all the databases

3:11:04like from where initially we have

3:11:05selected now from there on so this is

3:11:07how it looks like let's go back to Azure

3:11:10one we have understood how to do it in

3:11:12Azure dat studio also and how to connect

3:11:14it with the management SQL Server also

3:11:16so there are certain other features also

3:11:18you can see here for performance

3:11:19overview like how it's working and to

3:11:22track the performance and everything

3:11:24also the auditing and things are given

3:11:26here so just like these are the things

3:11:27but I have mainly told you about the

3:11:29purchasing models and deployment models

3:11:31and how service tires are there and how

3:11:33we can create the database connected

3:11:34with management SQL Server as well as a

3:11:37studio I think it might be a little

3:11:39hectic but I have explained you in a

3:11:41little simpler manner so please try it

3:11:43with your hands-on experience like take

3:11:45your hands-on experience also by trying

3:11:47it on deploying Azure SQL database on

3:11:49this Azure

3:11:50[Music]

3:11:56portal Azure data Lake storage so what

3:12:00is azure data Lake storage Azure data

3:12:03Lake storage is a repository that stores

3:12:06large amount of raw data in its natural

3:12:08format until it is needed for analytics

3:12:11application

3:12:13so why is it named as data link well

3:12:16James Dixon the chief technology officer

3:12:18of pentaho is the person who has

3:12:21generally been credited with the coining

3:12:23of the term data link according to him

3:12:26he described a data M that is a subset

3:12:29of a data warehouse as a keing to a

3:12:32bottle of water which is cleansed

3:12:35packaged and structured for easy

3:12:37consumption while a data lake is more

3:12:41likely a body of water in its natural

3:12:43States data flows from the streams that

3:12:47is the source of the system to the lake

3:12:49users have access to Lake to examine

3:12:53take samples or dive in so a data lake

3:12:57is a centralized repository designed to

3:12:59store process and secure large amount of

3:13:03structured semi-structured and

3:13:05unstructured data it can store data in

3:13:08its native format and process any

3:13:11variety of it ignoring the size

3:13:15limits next is how to create a data Lake

3:13:19storage for this we'll have a practical

3:13:22demo to have a better understanding so

3:13:25the first step to work on any Azure

3:13:28Services is to first sign in so first we

3:13:31need to sign in make sure you do have an

3:13:34Azure account so that you can have the

3:13:36access to different Azure services so

3:13:39let's sign

3:13:41in

3:13:48stay signed in so once you sign in you

3:13:52enter to this dashboard so you can see

3:13:54here this is a dashboard of your account

3:13:57and you can see different kinds of azure

3:13:59services like storage accounts monitors

3:14:02virtual machine Resource Group SQL

3:14:05database SQL manage instances and also

3:14:08you can see your subscriptions and

3:14:10different types of Resource Group you

3:14:11have created

3:14:13and all other stuffs so let's quickly

3:14:16get started first we need to go to

3:14:18create a

3:14:19resource so once you come here you can

3:14:22see popular isure services so like

3:14:26virtual machine cuberty services Cosmos

3:14:28DB and rest other so for us we need to

3:14:32go to storage

3:14:34account once you click here so you enter

3:14:37to create a storage account here we need

3:14:40to create a resource Cod so first of all

3:14:43what do you mean by Resource Group a

3:14:46resource Group is a container that holds

3:14:48related resources for an Azure solution

3:14:51it can include all these resources for

3:14:53the solution or only those resources

3:14:56that you want to manage as a grou so

3:14:59let's create a new Resource Group let's

3:15:01name it

3:15:03as demo data link one okay and once you

3:15:11create a resource Group so you come to a

3:15:13storage account make so a storage

3:15:16account contains all of your assure

3:15:18storage data objects including blobs

3:15:21file shs cues tables and disks so let's

3:15:25name it as marshmallow 1 2 3 so once you

3:15:30have named your storage account we go to

3:15:32the Advan section here we will directly

3:15:36move to data Lake storage generation 2

3:15:40so what do you mean by data leg storage

3:15:43generation 2 Data leg storage generation

3:15:452 is designed to deal with this variety

3:15:48and volume of data at exhibit scale

3:15:51while securely handling hundreds of

3:15:54gigabytes of through output with this

3:15:57you can use data leg storage generation

3:16:00to on the basis of both real time and

3:16:03bat solution so here we will enable the

3:16:07hierarchical name space once we enable

3:16:10it we'll click on riew plus create so it

3:16:14will take few minutes to validate all

3:16:15your details once validation is passed

3:16:19you just click on create so it's now

3:16:22getting ready to deploy so here you can

3:16:25see your marshmallow 1 2 3 is created

3:16:28and now it is getting ready for

3:16:30deployment here you can see deployment

3:16:33is in progress once it gets ready so we

3:16:37will be working on it here you can see

3:16:40your deployment is complete now what you

3:16:43going to do is we'll move on to go to

3:16:47resources now here you can see your

3:16:51details your resource gr details the

3:16:55whole storage Account Details basically

3:16:58and even you can see the properties

3:17:00enabled with it so here we have data L

3:17:04Storage file service Q service table

3:17:06service networking security so after

3:17:11this we will quickly go to

3:17:14containers and we'll create a new

3:17:17container for our storage account so

3:17:21let's click on container and we give it

3:17:24a name to it make sure you give a name

3:17:28it should be in lower case because they

3:17:30don't accept the upper

3:17:31case so let's keep it demo and create so

3:17:38yeah you can see your container has been

3:17:42created let's click on it so here you

3:17:45can see there are no results as we

3:17:47haven't added any stuff so let's

3:17:50upload for this you can either upload it

3:17:55from Azure portal or else you can also

3:17:59go to storage Explorer so let's see how

3:18:03do we do on the Azure portal so once you

3:18:06come here you can just click on select a

3:18:09file let's take any

3:18:12picture let's see we'll upload a picture

3:18:16so here you can see AWS

3:18:193.png let's upload it so here you can

3:18:22see AWS PNG has been uploaded in your

3:18:26container same as we can do on storage

3:18:29Explorer to for that you need to

3:18:32download a storage Explorer in my case I

3:18:35have already downloaded the storage

3:18:37Explorer so let's quickly go into it so

3:18:42here also make sure you are signed in so

3:18:45that you can see all your containers and

3:18:48Resource Group which you have created so

3:18:51here you can see my storage account that

3:18:54isow 123 which we have created recently

3:18:58under this we'll go and find our

3:19:02container yeah so here we go to blob

3:19:06containers and we can see a demo so here

3:19:09you can see the image file which we had

3:19:11uploaded through our a portal now we'll

3:19:15upload a file and a folder both let's

3:19:19see so let's upload a file first so you

3:19:22click on upload upload files and then

3:19:26you select a file let's take any file

3:19:32let's take a picture again so AWS 4

3:19:37let's take this an image so now here you

3:19:40can see your image is transferring from

3:19:44your path to our demo yeah so here your

3:19:49image file is uploaded so as we said we

3:19:52can upload any kind of datas maybe

3:19:55structured or unstructured so let's

3:19:58check out by uploading a folder so here

3:20:02you go with the same process and you you

3:20:05select your folder let's take any

3:20:10folders let's see if we do have any

3:20:12folder okay I'll just take any one of my

3:20:15folders and just click and upload so

3:20:19your folder is also being uploaded over

3:20:22here and once your folder is uploaded

3:20:25now you can see the inside resources

3:20:28into it so here there were different

3:20:31files text image all of them are there

3:20:35you can access it from here itself and

3:20:38you can see other operations as well if

3:20:40you want to download any of of the file

3:20:42or folder you can do it or if you want

3:20:44to open it let's open it any of the uh

3:20:48files see let's see if it is getting

3:20:52open or not okay yes so here you can see

3:20:56this file has been open now if you want

3:21:01to download this file let's download

3:21:06sequence so we'll download it and just

3:21:10put it in downloads let's see if it is

3:21:13downloaded or no apply to apply so yeah

3:21:17it tells your download is completed

3:21:21let's see if you find it or no we go to

3:21:23our file and downloads and here see now

3:21:28I had downloaded this is the file which

3:21:31I had downloaded through our storage

3:21:33Explorer so this is how you manage your

3:21:35data toward your data leg storage so

3:21:38apart from this you can also give

3:21:40permissions to different users users

3:21:42like whichever file if you want to give

3:21:44an access to a particular user then you

3:21:47can also give permission to them that

3:21:49they can either read write and access

3:21:52the whole file or a folder so for that

3:21:55you just need to click on a file or a

3:21:57folder and right click and you just see

3:22:01here manage Access Control list so once

3:22:04you come here so you can here you can

3:22:06see there are different owners super

3:22:09user owner and all so here I can add an

3:22:13owner like for this file who can just

3:22:17get an access to it so you can find out

3:22:21any name if you find out any relatable

3:22:23person or a user then you can give an

3:22:26access to that currently I don't have

3:22:28anyone so I won't be able to show you

3:22:31that but yes this is how you add or give

3:22:34permission to different users you can

3:22:37also do it with the folder or else you

3:22:39can also do it with the whole container

3:22:42as well if you want to share your

3:22:43containers with different people or

3:22:46different users then you can easily

3:22:48share them so here also you just go

3:22:51right click and just come to manage

3:22:53access control and just give them the

3:22:57access click on ADD and find out the

3:23:01person whom you want to give the access

3:23:03to and just after that once you give

3:23:07them the permission here you can see if

3:23:09you want to permit them for only read or

3:23:13only write or if you want to give read

3:23:16and write or all of the three so you can

3:23:19just give them the permissions

3:23:21accordingly and click on okay and then

3:23:24the particular user gets the access to

3:23:26all your files and folders so this is

3:23:29how we create and work with azo data

3:23:32Lake

3:23:33storage now let us see the comparison

3:23:36between Azure blob storage and the data

3:23:39L

3:23:40Storage so here are some of the

3:23:43comparisons between Azure data L Storage

3:23:45and Azure blob storage Azure data L

3:23:48Storage is a technic of planning and

3:23:51control of the time whereas Azure block

3:23:53storage is an object stored with a flat

3:23:57name space Azure data lake is an

3:24:00optimized storage for big data analytics

3:24:02workload whereas AZ your blob storage is

3:24:06basically a general purpose Object Store

3:24:08for a wide variety of storage scenarios

3:24:11which also include big data analytics in

3:24:15AO data Lake storage the apis are over

3:24:18https only whereas in Blob storage the

3:24:21rest API is over the HTTP as well as the

3:24:25https in AO data Lake storage there is

3:24:29no limits on the account but in aure

3:24:32Blob storage there are specific limits

3:24:34for container sizes and the files in the

3:24:37block so these are the major points

3:24:39which differentiate Azure block storage

3:24:42with aure data Lake storage at last we

3:24:46come to the use cases there are many use

3:24:49cases of data Lake storage out of which

3:24:52we'll discuss about four so at first

3:24:56business intelligence on data Lake

3:24:58storage so data Lake storage

3:25:00dramatically improves the speed for ad

3:25:03hoc queries dashboards and remotes you

3:25:06can run existing bi tools on lower cost

3:25:10data Lakes without compromising

3:25:12performance or data quality it also

3:25:15avoids costly delays adding new data

3:25:18sources and the reports at second we see

3:25:22cloud data Lake migration here we can

3:25:26optionally deploy new applications to

3:25:28the cloud using data Lake storage such

3:25:30as S3 or ADLs you can also migrate from

3:25:34older onri data Lake environments that

3:25:37are expensive and difficult to maintain

3:25:39while ensuring agility and

3:25:42flexibility next data science on the

3:25:45data L Storage here you can accelerate

3:25:49data science on data Lake storage with

3:25:52simplified data exploration and feature

3:25:54engineering dramatically it improves

3:25:57performance making data scientists and

3:25:59Engineers more efficient resulting in

3:26:02high quality analytical models at last

3:26:06data architecture modernization so here

3:26:08you can avoid Reliance on propriatary

3:26:12data warehouse infrastructure and the

3:26:14need to manage the cubes extract and

3:26:16agregation tables you can run

3:26:19operational data warehouse queries on

3:26:21low cost data laks offloading the data

3:26:24warehouse at your own

3:26:28[Music]

3:26:32paas so powerbi is a ba tools okay the

3:26:36business intelligent tools wased on the

3:26:39cloud machines cloud services it's

3:26:42maintained by Microsoft corporations

3:26:44it's a Microsoft tools right Ms tools

3:26:46Microsoft tools it's entirely free of

3:26:49cost we don't need to pay any licensing

3:26:50cost or anything for this powerbi it's a

3:26:53free it's available under Microsoft

3:26:56stores so this powerb is mainly for

3:26:58non-technical people suppose if you're

3:27:00business user if you have data analyst

3:27:02if you have a business analyst right so

3:27:05if if you want to perform some of the

3:27:07aggregate level data I want to see the

3:27:09data from summary level or I want to

3:27:11perform some of the data analysis or I

3:27:14want to uh visualize some of the data in

3:27:16the graphical formats or I want to share

3:27:18the data to the different peoples for

3:27:20different servers different place right

3:27:22so for that the power ba is very well

3:27:24suited for non-technical users business

3:27:27user okay so these are the things we can

3:27:29do it on the

3:27:31powerb okay next we'll talk about Azure

3:27:34in EML so how do we use this um Azure uh

3:27:38in in machine learning right so when

3:27:41where we have a lot of services

3:27:43available in Azu like we can integrate

3:27:45it so basically if you want to apply

3:27:47some filter or if you want to apply some

3:27:48selective filter independently we can do

3:27:50it under Azure platform I'll just

3:27:53quickly walk you through the Azure

3:27:56portal okay so this is the my cloud

3:28:00platform like portal uh aure portals

3:28:03right so as I told you we have a

3:28:07different services available so if you

3:28:09go check right so we have a SQL database

3:28:13Cosmo DB for no SQL database right and

3:28:16storage account like as your active

3:28:18directory for this security things right

3:28:20and then

3:28:21monitor cost management so there are

3:28:24different services are available so

3:28:26basically before going to create a

3:28:28service rate so we need to have the free

3:28:30account so basically this is a one month

3:28:32free account you can able to create it

3:28:35so like this first of all you have to go

3:28:37have the subscriptions so the

3:28:39subscription means same as like what

3:28:41kind of uh subscription currently we're

3:28:43having as of now I don't any

3:28:45subscriptions I'm going to create it

3:28:47like uh this is for the 12 months or

3:28:51subscription free after that we have to

3:28:53go and purchase it either pay as it go

3:28:57let it come so that's the first one is a

3:28:59free trial like after that we have to go

3:29:01P go right as I told you P go means like

3:29:05depending on whatever we using right we

3:29:07are paying for it that's it okay we not

3:29:10once you are application once once you

3:29:12done with the work we can uh stop the

3:29:14services right so it means we're not

3:29:16paying for any anything to the it's not

3:29:19going to consume the cost okay and then

3:29:22for the student also we have the 12

3:29:24months free thing we can use it so first

3:29:27of you have to go to get the

3:29:28subscription once the subscription

3:29:29created we have to go and have the

3:29:31resource Group monitor a lot of things

3:29:34right so for different category wise we

3:29:36talk about right we have a storage we

3:29:38have a networking uh we have a what what

3:29:41is called uh database right so like this

3:29:44we have machine learning also we have a

3:29:45some of services suppose if you want to

3:29:47have the phase API phase deduction

3:29:49computer visions and then uh like for

3:29:53cognitive search right like this for

3:29:55analytical for Azu datab bricks so for

3:29:58that everything for analytical kind of

3:29:59things there are lot of things right

3:30:01powerb Integrations so we talk about

3:30:03powerb integration right so for that we

3:30:05have a separate analytical service okay

3:30:08so we have a compute service so comput

3:30:10service we initially right so how do you

3:30:12want to execute it okay how do we most

3:30:13we want it and what are availability

3:30:16which zone you want to store it what is

3:30:17the disk storage and virtual machines

3:30:21virtual storage like any image app

3:30:24services so there are compute services

3:30:27and then we have a container we have a

3:30:29database level Services we have a devops

3:30:32kind of things Services everything we

3:30:35are having services on the cloud

3:30:38[Music]

3:30:40one

3:30:43before knowing what is azure data brakes

3:30:46we must know what data brakes actually

3:30:49is well according to the definition data

3:30:52brakes is a web-based platform for

3:30:54working with Apache spark that provides

3:30:57automated cluster management and IPython

3:30:59style notebooks so basically data breaks

3:31:03developed by the creators of Apaches

3:31:05spark is nothing but a web-based

3:31:08platform which is also One-Stop product

3:31:10for all data requirements like storage

3:31:13and Analysis it was originally founded

3:31:16to provide an alternative to the map

3:31:19reduce system and provide a just in time

3:31:22cloud-based platform for big data

3:31:24processing clients it can derive

3:31:27insights using spark xql provide active

3:31:31connections to visualization tools such

3:31:33as powerbi click View and tab View and

3:31:37also build predictive models using

3:31:39sparkml datab BRS also can create

3:31:42interactive displays text and code

3:31:45tangibly so in short it is an

3:31:48alternative to a map redu system data

3:31:51brakes is now integrated with Microsoft

3:31:53Azure Google Cloud platform and Amazon

3:31:57making it easy for the business to

3:31:59manage a colossal amount of data and

3:32:02Carry Out machine learning tasks as we

3:32:05got to know that data braks is

3:32:06integrated with all the three

3:32:08cloud-based platforms today we'll

3:32:11discuss on one of those that is azure

3:32:13data brakes so let us understand what is

3:32:16azure data breakes Azure data bricks

3:32:19lakeh house is a platform that offers a

3:32:22uniform collection of tools for building

3:32:24deploying sharing and supporting

3:32:27Enterprise grade Data Solutions to a

3:32:29scale it integrates with cloud storage

3:32:32and Security in your cloud account and

3:32:35manages and deploy Cloud infrastructure

3:32:37on your behalf Azure data break supports

3:32:40python Scala R Java and SQL as well as

3:32:45data science Frameworks and libraries

3:32:48including tensor flow py torch and pych

3:32:51loarn now that we have come to know what

3:32:54is azure data break let us understand

3:32:57why should we use Azure data brakes well

3:33:01our customers use Azure data brakes to

3:33:03process store clean share analyze model

3:33:07and monetize their data sets with

3:33:10solution from powerbi to machine

3:33:12learning you can use Azure data brakes

3:33:15platform to build many different

3:33:16applications spanning data personas

3:33:20customers who fully embrace the lak

3:33:22house take full advantage of the unified

3:33:25platform to build and deploy the data

3:33:28engineering workflows machine learning

3:33:30models and analytics dashboard that

3:33:33power Innovations and insights across an

3:33:37organization the Azure data breakes

3:33:39workspace provid provides user

3:33:41interfaces for many code data task

3:33:44including tools which we'll be

3:33:46discussing one by one so first is

3:33:48optimized Spark engine it is a simple

3:33:51data processing on autoscaling

3:33:54infrastructure powered by highly

3:33:56optimized Apache spark for up to 50x

3:33:59Performance Gaines next is machine

3:34:01learning

3:34:02runtime it is a oneclick access to

3:34:05preconfigured machine learning

3:34:07environment for augumented machine

3:34:10learning with state-ofthe-art and

3:34:12popular Frameworks such as py toch

3:34:15tensor flow and psychic learn we also

3:34:18have mlflow that is used to track and

3:34:22share experiments reproduce runs and

3:34:24manage model collaboratively from a

3:34:28central repository well here you can use

3:34:31your preferred language including python

3:34:34Scala R spark SQL and net whether you

3:34:38use serverless or provis visioned

3:34:41compute resources so that you can

3:34:43quickly access and explore data find and

3:34:46share new insights and build models

3:34:49collaboratively with the language and

3:34:51tools of your choice in Azure data

3:34:55breakes you have Enterprise gr security

3:34:57which is an effortless native security

3:35:00that protects your data where it lives

3:35:02and creates complaint private and

3:35:05isolated analytics workspace across

3:35:08thousands of users and database

3:35:11apart from that it is also production

3:35:14ready that means you can run and scale

3:35:16your most Mission critical data

3:35:19workloads with confidence on a trusted

3:35:22data platform with ecosystem integration

3:35:25for cicd and monitoring you also have

3:35:29collaborative notebooks through which

3:35:32you can quickly access and explore data

3:35:35find and share new insights and build

3:35:37models collaboratively with the

3:35:39languages and tools of your choice and

3:35:42the most important thing it has Delta

3:35:45Lake that brings data reliability and

3:35:47scalability to your existing data lake

3:35:50with an open-source transactional

3:35:52storage layer designed for full data

3:35:55life cycle it has a native integration

3:35:58with Azure services that means you can

3:36:01complete your endtoend analytics and

3:36:03machine learning solution with deep

3:36:06integration with Azure services such as

3:36:08Azure data Factory azure data L Storage

3:36:12aure machine learning and

3:36:14powerbi last but not the least it has an

3:36:17interactive workspace that means you can

3:36:20enable seamless collaborations between

3:36:23data scientists data engineers and

3:36:26business analysts so these were the key

3:36:28features that makes Azure data brakes

3:36:31unique till now we have got an idea

3:36:34about the data brakes and Azure data

3:36:36brakes and why are we using Azure data

3:36:39braks now let let us deep dive in by

3:36:42understanding how does Azure data brakes

3:36:45actually work as I said Azure data

3:36:48brakes is structured to enable secure

3:36:50cross functional team collaboration

3:36:53while keeping a significant amount of

3:36:55backend Services managed by Azure data

3:36:58braks so you can stay focused on your

3:37:01data science and data analytics and data

3:37:04engineering task it operates out of a

3:37:07control plane and a data plane Although

3:37:10our architectures can vary depending on

3:37:12the custom configuration such as when

3:37:15you have deployed a Azure data break

3:37:17workspace to your own virtual Network

3:37:20which is also known as vnet injection

3:37:23now let us consider this architecture it

3:37:26is a common structure and a data flow of

3:37:29an Azure data braks here it consists of

3:37:32control plane and data plane so let us

3:37:35understand what is control plane the

3:37:37control plane includes the backend

3:37:39services that Azure data brakes manage

3:37:42in its own Azure account notebook

3:37:45commands and many other workspace

3:37:47configurations are stored in the control

3:37:50plane and encrypted at the rest if we

3:37:53talk about data plane it is managed by a

3:37:56your Azure account and it is where your

3:37:58data resides this is also where data is

3:38:02processed you can use Azure data break

3:38:04connectors so that your clusters can

3:38:07connect to external data sources outside

3:38:09your aure account to ingest data or for

3:38:13storage you can also ingest data for

3:38:16external streaming data sources such as

3:38:18even datas streaming datas iot data or

3:38:22many more well your data is stored at

3:38:26rest in your Azure account in the data

3:38:29plane and in your own data sources not

3:38:32the control plane so you maintain

3:38:34control and ownership of your data if we

3:38:38talk about the job results then it

3:38:40resides in storage in your account

3:38:42itself interactive notebook results are

3:38:45stored in the combination of control

3:38:47plane that is partial result for

3:38:50presentation in the UI and your Azure

3:38:53storage if you want interactive notebook

3:38:56results stored in only your cloud

3:38:58account storage then you can ask data

3:39:01break representative to enable

3:39:03interactive notebook result in the

3:39:05customer account for your workspace note

3:39:08that some metadata about results such as

3:39:12chart columns names continues to be

3:39:14stored in control plane itself so this

3:39:17is the basic architecture of your data

3:39:20break to know it more in simplified

3:39:23manner let us have a quick Hands-On on

3:39:26Azure data brakes so here we'll be

3:39:29integrating Azure data brakes with the

3:39:32Azure blob storage that is a service

3:39:35provided by Microsoft Azure so as you

3:39:38can see this is a small workflow of how

3:39:41we'll be working on the demo so as we

3:39:44know Azure blob storage and Azure data

3:39:47brakes are both Services provided by

3:39:49Microsoft Azure now these are two

3:39:52separate services but as long as you are

3:39:55using them in same Resource Group you

3:39:58can integrate these two Services well

3:40:00now you must be wondering why do we need

3:40:02to combine these so let's say that

3:40:05Microsoft Azure provides a multitude of

3:40:08services it is often beneficial to

3:40:11combine multiple Services together to

3:40:13approach your use case so if you combine

3:40:16multiple Services then we don't need to

3:40:19engage your local hardware in anything

3:40:21like well for example currently I'm

3:40:24using a laptop maybe my laptop can be a

3:40:26lower configuration or it might not have

3:40:29enough space to process a huge amount of

3:40:32data in that case I would want the cloud

3:40:35service to handle all my use cases and

3:40:38my big data storage that I have

3:40:40so in this workflow as we said that we

3:40:43will integrate Azure data brakes and

3:40:46Azure blob storage that means we're

3:40:48going to combine them so here what you

3:40:50do basically is you interact with the

3:40:52coding notebook which is nothing but the

3:40:55IPython or Jupiter notebook that your

3:40:59Azure data brakes will create then what

3:41:01you do is then you type some coding

3:41:04commands in your coding notebook and

3:41:07then these commands are sent to the data

3:41:10Brak service and then what happens is

3:41:13the datab break service receives the

3:41:16commands from the coding notebook and it

3:41:19sends those commands to your Azure

3:41:21cluster so whatever cluster you have

3:41:23created after creating your data brakes

3:41:26so whatever commands that we are writing

3:41:28in the coding notebook is sent through

3:41:30data brakes to your clusters that you

3:41:33have created now depending on the

3:41:36authentication provided to the cluster

3:41:38with regards to your blob storage

3:41:40account the authentication commands are

3:41:44then sent to the blob storage account

3:41:47saying that the data is fetched from the

3:41:49desired directory then it is brought

3:41:53back inside the cluster and that data

3:41:56that has received by the cluster then is

3:42:00processed and whatever output you get

3:42:03out of after the process is you can see

3:42:07it on your coding notebook now whatever

3:42:10output you receive from the coding

3:42:12notebook you can also store that

3:42:14particular output that you have got from

3:42:17the coding notebook back to your blob

3:42:19storage account or space so all of this

3:42:23is integrated easily and it is handled

3:42:28in a very simpler manner well now that

3:42:31we have understood the architecture let

3:42:33us now know how to implement it with a

3:42:37Hands-On for this we need to quickly

3:42:39sign in to our Azure portal so once we

3:42:42sign in we land in our dashboard of the

3:42:45Microsoft Azure as you can see here the

3:42:47all the Azure services are mentioned

3:42:49here and once you create your resource

3:42:52Group so it gets highlighted also now

3:42:55let us quickly open our Azure data braks

3:42:58and let us create our Azure data braks

3:43:01so for this once we come here we need to

3:43:03just simply

3:43:06create and here they ask for basic

3:43:09details

3:43:10like your resource Group then workspace

3:43:13name region so let us fill up one by one

3:43:16so here as we don't have any Resource

3:43:18Group so we'll be creating a new

3:43:20Resource Group as we working on a demo

3:43:23so let's name it as Azure data brakes

3:43:28demo and

3:43:30let's create it now that we have created

3:43:33the resource Group now they ask for the

3:43:35workspace name so we'll give it a unique

3:43:38name you can keep it any so let's name

3:43:42it as aure data braks hands on and then

3:43:49you can choose your region so here you

3:43:52find uh different types of region where

3:43:54you can like work or create your Azure

3:43:57data bricks so it is up to you whatever

3:44:00region you choose for me I'll be keeping

3:44:03West us as it is after that we come to

3:44:06the pricing tier so here there are three

3:44:09types standard premium and trial so as

3:44:12we are working on the like we are just

3:44:14having a practical knowledge we will be

3:44:17using the trial version and we'll just

3:44:19review and create now here you can just

3:44:22check on to your whatever details you

3:44:24had filled previously whether it is

3:44:26correct or no so once it is validated we

3:44:29can just create this data brakes now as

3:44:32you can see they are initializing the

3:44:35deployment so it may take a while to

3:44:37create a datab break

3:44:40all right so here you see the deployment

3:44:42is in

3:44:43progress so once your deployment is

3:44:46complete you can go to your uh resource

3:44:49and here you land up to your datab

3:44:52brakes page now here what we need to do

3:44:55is launch our workspace so it will move

3:44:59us to our main Azure portal so it will

3:45:02sign us in the Azure data breakes so

3:45:05this is the dashboard of the data braks

3:45:07now here you have different options like

3:45:10notebook and um data import partner

3:45:13connect transform data and many other

3:45:15things here you can set up your

3:45:18workspace like create a cluster import

3:45:21data build a data pipeline as well now

3:45:23one thing I should tell you that data

3:45:25bricks works on dbfs now what do you

3:45:28mean by dbfs is well dbfs is nothing but

3:45:32datab brakes file system which is a

3:45:34distributed file system mounted on Azure

3:45:36databas workpace and are available on

3:45:39the Azure databas cluster so it allows

3:45:42you to interact with the object storage

3:45:45using directory and file semantics

3:45:48instead of cloud specific API commands

3:45:50and it helps you out to mount Cloud

3:45:52object storage location so that you can

3:45:55map your storage credentials to the path

3:45:57in the GE datab brakes workspace it also

3:46:00simplifies the process of persisting

3:46:02files to object storage allowing the

3:46:05virtual machines and attacked volume

3:46:07storage to be safely deleted on the

3:46:09cluster termination well these are the

3:46:12things which you can like implement it

3:46:14while you're working with the Azure data

3:46:17breakes now let's get back to our data

3:46:19brakes now that we have come to this

3:46:21page so the first thing to be done over

3:46:23here is to create a cluster now we go to

3:46:26create a cluster now here we go to

3:46:28create compute now here as you can see

3:46:32we have the cluster name so here we can

3:46:36edit the cluster name based on your

3:46:38requirement so let us keep it as a data

3:46:41brakes cluster and after this we have

3:46:44the policy so here policy is nothing but

3:46:47a cluster policy defines the limit on

3:46:50the attributes available during the

3:46:52cluster creation so here we have

3:46:55different types of policies that is

3:46:57unrestricted personal compute power user

3:47:00compute shared compute so we will keep

3:47:04it unrestricted for time being now there

3:47:08are two types of cluster mode one is

3:47:10multi node and single node so in multi

3:47:14node we can specify the minimum number

3:47:16of workers and the maximum number of

3:47:18workers so here minimum numbers can be

3:47:21two or you can specify it upon this is

3:47:24basically a standard limit for minimum

3:47:27and maximum whereas you can change it

3:47:29accordingly based on your requirement so

3:47:32now as of now as we are only practicing

3:47:34so we will just disable this Autos

3:47:37scaling and we can specify are number of

3:47:40workers over here so we can keep only

3:47:43one worker as it is only for practicing

3:47:46whereas we can also like change the

3:47:49timings for termination if the

3:47:51particular cluster is inactive so you

3:47:54can specify that much of time to it

3:47:57apart from that we have our access mode

3:48:01that is nothing but there are three

3:48:03types of access mode that is single user

3:48:05shared user and no isolation shared so

3:48:09so here we'll keep it as it is and your

3:48:12single user access is nothing but our

3:48:14subscription after that we come to our

3:48:16performance so here we need to specify

3:48:19our runtime version so here in our case

3:48:22it is runtime 11.3 LTS color

3:48:262.12 and our work type can be standard

3:48:30whereas there are other versions as well

3:48:33but as of now we don't require much of

3:48:36it so we'll be going with the standard

3:48:39vers version itself and whereas we have

3:48:41already specified our workers and

3:48:44termination time is also been mentioned

3:48:46now here towards your right you can see

3:48:49the whole summary of your cluster what

3:48:51have you been uh like creating so once

3:48:55you review it you can just create this

3:48:57particular cluster now as you can see it

3:49:00is loading it takes a while to like

3:49:03create a cluster here in this section

3:49:07you see the status of your cluster

3:49:10So currently it is in a pending mode

3:49:12like it has been creating so we'll wait

3:49:16for a while so now as you can see our uh

3:49:19cluster has been created and it is in

3:49:21the running mode now after this what we

3:49:25need to do is go back to our main

3:49:27dashboard and we need to create a new

3:49:30notebook so we'll create a new notebook

3:49:32over here and here we have already

3:49:36specified the cluster so we had created

3:49:39created now so it has automatically

3:49:41taken and uh here we need to specify our

3:49:44default language so you can choose any

3:49:47of them so here there there are four

3:49:49types of languages which you can choose

3:49:52here I would be taking Scala for now and

3:49:55we can give a simple name to this

3:49:57notebook that is it can be anything of

3:50:00your choice so let us name it as data

3:50:04brakes notebook and let's create so this

3:50:08will start very quickly it doesn't take

3:50:10that much of time now here it is like

3:50:14you need to like run your command over

3:50:16here so you just need to type down your

3:50:18command and just hit enter and it starts

3:50:21running so now in this we will know how

3:50:25to upload a file through uh C Azure

3:50:29storage service that is our Azure blob

3:50:32storage so it's basically we need to

3:50:35First integrate the Azure blob storage

3:50:38so let us know how so before this we

3:50:42need to go back to our aure portal and

3:50:45here we need to First create our storage

3:50:47account so here as you can see we have

3:50:49our storage accounts and here we'll

3:50:53create our blob storage so let's create

3:50:57so now here for creating storage account

3:51:00you need to give your details over here

3:51:03so as here we had already created the

3:51:08resource Group previous L while creating

3:51:10a data break so we will be selecting the

3:51:12same and after that we will give a

3:51:15storage account name so let us name it

3:51:18as data braks storage account and in

3:51:23region we need to provide the specific

3:51:26region whichever you like to choose so

3:51:29as previous I had choosen West us so

3:51:32I'll be choosing that as

3:51:34well and here now when we come to the

3:51:38performance so here we can choose any of

3:51:41those um any of the two options given

3:51:44below so here we have standard and

3:51:46premium So currently we'll be going with

3:51:48the Standard

3:51:49Version and when we talk about redund

3:51:52dency so here we have two types of

3:51:55redundancy as you can see locally

3:51:58redundant storage here it means the

3:52:00lowcost option with the basic protection

3:52:02against server rack and drive failur

3:52:05recommended for non-critical scenarios

3:52:08whereas Geo redundant storage is like

3:52:11for intermediate option with failover

3:52:13capabilities in secondary region

3:52:15recommended for backup scenarios so

3:52:18locally means it happens within the

3:52:20region not across the whole world so

3:52:24here we will be choosing the locally

3:52:26redundant storage now we have specified

3:52:29everything so now let us review so here

3:52:33before like creating you need to check

3:52:37all your details once you review it just

3:52:40create and your storage account is in

3:52:45initialization stage so once it gets

3:52:48deployed we will start working on that

3:52:51so now our deployment is complete so we

3:52:53can go to our resource so here all of

3:52:56the permissions are automatically

3:52:58managed within the same Resource Group

3:53:01so we don't have to worry about any

3:53:03permissions

3:53:04requirement so now that we have created

3:53:07our storage account now here we need to

3:53:10create our container so we'll go to our

3:53:15containers and we'll so here we can give

3:53:19it a name to our container it can be

3:53:22anything name it as storage account one

3:53:27and let's just create so as you can see

3:53:30our container has been created now we

3:53:32will quickly go onto this container and

3:53:35here we don't have anything in this

3:53:37container so the container is empty now

3:53:40we need to upload some files in this

3:53:42particular container so we'll just

3:53:44quickly go to upload and here we will go

3:53:49to select file so here what does it do

3:53:51it will get connected to my Windows File

3:53:54you can take any files over here so as

3:53:56of now I will just take a

3:54:00normal Excel a CSV file and we'll just

3:54:05simply upload it well you can upload

3:54:08more than one files if you want to like

3:54:10upload it in the containers you can also

3:54:13have a larger files but it may cost

3:54:16according to the given size now we have

3:54:20a file in place and we also have created

3:54:24a notebooks so now we have to integrate

3:54:28The Blob storage with the data brid so

3:54:32for that we need to run a code now here

3:54:35the code looks a bit complex so that

3:54:38here so here first we need to create a

3:54:41token so that we can get access to the

3:54:44files so we'll just copy this whole code

3:54:50so that there's nothing to memorize this

3:54:53can be provided to you while you are

3:54:56practicing so now we'll just copy paste

3:54:59the whole command so here now as you can

3:55:05see the container name so here we need

3:55:08to spe specify the container name and

3:55:10the storage account that we have created

3:55:12so we'll quickly go back to our aure

3:55:14portal and we will fill up the details

3:55:17so as they asked for the container name

3:55:20so here we need to specify our container

3:55:22name and the storage account which we

3:55:24have created so let's quickly go back to

3:55:27our a portal and here as you can see

3:55:30your container name

3:55:33is given so we'll just quickly copy

3:55:37that go let's go back and we'll just

3:55:43copy this name and we'll paste it here

3:55:49now same thing to be done with our

3:55:51storage account so we'll go back to our

3:55:54storage account and here we'll see this

3:55:57is our storage account name so we'll

3:56:00just copy and we'll just paste it over

3:56:04here now for the SOS token we need to go

3:56:09back to our containers and here we will

3:56:14go to the Shar access signature so here

3:56:18we will be getting our SAS SAS tokens so

3:56:23SAS token can be generated for a limited

3:56:26amount like for a given period of time

3:56:28so you need to specify like when you

3:56:31create your SAS token so the time it has

3:56:34been started it would be valid from the

3:56:37time it has been started till its time

3:56:39of expiry so we'll just allow the

3:56:44services containers and the objects and

3:56:48let the time be as it is as it is been

3:56:51specified and now let's

3:56:54generate so now here as you can see you

3:56:57can find your SAS token now we need to

3:57:00just simply copy this and go back to our

3:57:04data brakes and we'll paste it over here

3:57:08so we have just added our SAS token let

3:57:11us just verify whether it has been

3:57:14correctly copied or not so I feel

3:57:17everything looks perfect rest other

3:57:19things remains to be same so here what

3:57:22happens you keep on creating new

3:57:24variables so here we have created a URL

3:57:28so by appending the container and the

3:57:31storage account so once we like we

3:57:35specify our URL and our configuration

3:57:37then it comes the D PS so here we

3:57:40specify our source Mount point and our

3:57:45maps to that particular token so this is

3:57:49basically where the data braks helps you

3:57:52like get your sources from your

3:57:54different like different services so

3:57:57here the data brakes plays the role

3:57:59where it provides you the data which

3:58:02from where you want to extract from and

3:58:05it shows it over here and now you just

3:58:08just hit shift plus enter so now it is

3:58:13like running the

3:58:15command so here it what does it do as we

3:58:18had discussed in our architecture so it

3:58:20communicates first with the data brakes

3:58:22and then it communicates with our

3:58:23cluster and get all the sources and then

3:58:27shows its result on the cluster itself

3:58:30so as you can see your command has been

3:58:36run okay so it shows some error over

3:58:39here so let's quickly solve that error

3:58:42all right now let's quickly run it

3:58:46again so it may take a while so here as

3:58:50I said it will first connect with the

3:58:53data brakes and then communicate with

3:58:55that and then it will communicate with

3:58:59your cluster and get all the resources

3:59:01from there so here as you can see they

3:59:04have specified your container name

3:59:06storage name and your token which has

3:59:09been

3:59:10specified now what we need to do is

3:59:14check whether our file has been take

3:59:17extracted from our storage account or no

3:59:21so here we'll spec first we'll specify

3:59:23our variable so let's

3:59:26specify with

3:59:29p. read and now we're going to give the

3:59:34format of the file so as we had uh

3:59:37uploaded the CSV file so we'll just

3:59:40specify it over here

3:59:43CSV and then we can give them some of

3:59:46the options that um they should show so

3:59:50let's give it an option like we can ask

3:59:53them to show the header and their value

3:59:58should be

4:00:00true then we can also specify the infers

4:00:03scamma so we'll just add on that as well

4:00:07and for efficient uh data requirement we

4:00:10will just provide them with a mode that

4:00:13can be fail fast mode and then we upload

4:00:17the like mentioned the file name so here

4:00:21we paste this Mount slash staging that

4:00:24is our Mount point and we'll paste it

4:00:28over here and then we will specify our

4:00:34file name that is let's go to our

4:00:37containers

4:00:39let's check what's the name python 1.

4:00:44CSV let's copy the name and we paste it

4:00:49so now let's just shift

4:00:53enter uh so it shows some error let's

4:00:56see what is it okay so here we had

4:01:00specified a wrong command so let's just

4:01:04resolve it and

4:01:07let's run it

4:01:09again all right so let us see how it

4:01:13looks like so now just do

4:01:17TF dot show and

4:01:22let's specify some

4:01:25amount and just shift

4:01:28enter so as you can see so they have

4:01:31mentioned here the decimal and uh

4:01:35description the percentage your headers

4:01:37have been been provided and the number

4:01:41of the number of rows that we required

4:01:43they have specified that so this is how

4:01:45we integrate the two services and we

4:01:49process data through data breaks now let

4:01:52us look onto some of the popular use

4:01:54cases that makes a your data Brakes in a

4:01:57huge demand well data brakes isn't a

4:02:01catchall solution for every business

4:02:03scenario so there are the best use cases

4:02:06for Azure data braks first is database

4:02:09and Mainframe modernization data storage

4:02:13collection and processing are incredibly

4:02:16important in modern businesses well if

4:02:19you're looking to modernize your data

4:02:21legs or looking into Mainframe

4:02:24modernization applications then Azure

4:02:27data brakes has all the Integrations you

4:02:29need next is machine learning production

4:02:33pipeline here using the underlying power

4:02:36of ml flow data bricks is a good choice

4:02:40if you need to get machine learning

4:02:42applications into production getting

4:02:45data signs out of the development and

4:02:47into production is a common problem and

4:02:50Azure data breaks can help you

4:02:52streamline that workflow if you talk

4:02:54about big data processing then Azure

4:02:57data brakes is one of the most coste

4:02:59effective options for big data

4:03:01processing in terms of performance

4:03:04versus cost it offers higher efficiency

4:03:07if your business needs the best

4:03:09performance for one demand data

4:03:11processing then data brakes will likely

4:03:14be your best choice next is business

4:03:17intelligence integration integrating

4:03:19business intelligence tools means you

4:03:22can open your data Lake to analysts and

4:03:25engineer more easily there's no need for

4:03:28creation of new pipelines when analysts

4:03:31need access to new data the data can be

4:03:34shared through SQL analytics powerbi and

4:03:37tablet you if this is a bottleneck for

4:03:40your business then data bricks will help

4:03:43you enable your business intelligence

4:03:45team now these were the four popular use

4:03:48cases by Azure data braks if your

4:03:51business fits in one of these use cases

4:03:54then it might be the solution for

4:03:56[Music]

4:04:01you let us start with the US primary

4:04:04election use case first in this use case

4:04:07we will be discussing about the 2016

4:04:10primary elections in the primary

4:04:12elections the contenders from each party

4:04:14compete against each other to represent

4:04:16his or her own political party in the

4:04:18final elections there are two major

4:04:20political parties in the US the

4:04:22Democrats and Republicans from the

4:04:24Democrats the contenders were Hillary

4:04:26Clinton and Bernie Sanders and out of

4:04:28them Hillary Clinton won the primary

4:04:30elections and from the Republicans the

4:04:32contenders were Donald Trump Ted Cruz

4:04:34and a few others as you already know

4:04:36Donald Trump was the winner from the

4:04:38Republicans so now let us assume that

4:04:41you are an analyst already and you have

4:04:43been hired by Donald Trump and he tells

4:04:45you that I want to know what were the

4:04:47different reasons because of which

4:04:49Hillary Clinton won and I want to carry

4:04:52out my upcoming campaigns based on that

4:04:54so I can win the favor of the people

4:04:56that voted for her so that was the

4:04:58entire agenda so this is the task that

4:05:02has been given to you as a data analyst

4:05:04so what is the first thing that you will

4:05:06need to do the first thing you'll do is

4:05:09that you'll ask for data and you have

4:05:11got two data sets with you so let us

4:05:14take a look at what this data sets

4:05:15contains so this is our first data set

4:05:18which is the US primary election data

4:05:20set so these are the different fields

4:05:21present in our data set so the first

4:05:23field is state so we've got the list of

4:05:25the state of Alabama the state

4:05:27abbreviation for Alabama is Al we've got

4:05:30the different counties in Alabama like

4:05:32aruga Baldwin Barber bib Blount bulock

4:05:34Butler Etc and then we've got fips now

4:05:37fips are federal information processing

4:05:39standards code so this is basically

4:05:41means zip code then we've got the party

4:05:45to which we would be analyzing the

4:05:46Democrats only because we want to know

4:05:50what was the reason for Hillary

4:05:51Clinton's win so we will be analyzing

4:05:54the Democrats only and then we've got

4:05:56the candidate and since I told you there

4:05:58were two candidates Bernie Sanders and

4:06:00Hillary Clinton so we've got the name of

4:06:02the candidate here and the number of the

4:06:04votes each candidate got so Bernie

4:06:06Sanders got 54 44 in aruga county and

4:06:09Hillary Clinton got to 2387 and this

4:06:12field over here represents the fraction

4:06:14of the votes so if you add these two

4:06:16together you will get a one so this

4:06:19basically represents the percentage of

4:06:21vote each of the candidates got so let's

4:06:24take a look at our second data set now

4:06:27so this data set is the US County

4:06:29demographic features data set so the

4:06:31first we will have again fips in the

4:06:34area name aruga County Baldwin and

4:06:36different other count counties in

4:06:38Alabama and other states also the state

4:06:40abbreviation so here it is only showing

4:06:43Alabama and the fields that you see here

4:06:45are actually the different features you

4:06:47won't know what this exactly contains

4:06:49because it is written in a coded form

4:06:52but let let me give you an example what

4:06:54this data set contains let me um tell

4:06:57you that I'm just showing you a few rows

4:06:59of the data set this is not the entire

4:07:01data set so this contains different

4:07:03fields like population in 2014 in 2010

4:07:06the sex ratio how many females males and

4:07:09then based on some ethnicity how many

4:07:11Asians how many Hispanic how many black

4:07:14American people how many um black

4:07:17African people and then there is also

4:07:19based on the age groups how many infants

4:07:22uh how many senior citizens how many

4:07:24adults so there are a lot of fields in

4:07:27our data set and this will help us to

4:07:29analyze and actually find out what led

4:07:30to the winning of Hillary Clinton so now

4:07:33you have seen our data set you have to

4:07:35understand your data set you have have

4:07:37to figure out what are the different

4:07:40features or what are the different

4:07:42columns that you are going to use and

4:07:44you have to think of a strategy or think

4:07:46of how you're going to carry out this

4:07:48analysis so this is the entire solution

4:07:51strategy so the first thing you will do

4:07:53is that you need a data set and you've

4:07:55got two data sets with you the second

4:07:58thing that you'll need to do is to store

4:08:00that data into hdfs now hdfs is how to

4:08:03distributed file system so you need to

4:08:05store the data the next step is to

4:08:07process that data using spark components

4:08:09and we will be using spark SQL spark M

4:08:12lib Etc so the next task is to transform

4:08:16that data using spark SQL transforming

4:08:18here means filtering out the data and

4:08:20the rows and columns that you might need

4:08:22in order to implement or in order to

4:08:24process this the next step is clustering

4:08:27this data using spark M lib and for

4:08:29clustering uh our data we will be using

4:08:32K means and the final step is to

4:08:33visualize the result using Zeppelin now

4:08:36visualizing this step is also very

4:08:38important because without the

4:08:39visualization you won't be able to

4:08:41identify what were the major reasons and

4:08:43you won't be able to gain proper

4:08:45insights from your

4:08:46data now don't be scared if you're not

4:08:49familiar with terms like spark SQL spark

4:08:51AMG K means clustering you will be

4:08:54learning all of these in today's session

4:08:56so this is our entire strategy this is

4:08:59what we're going to do today this is how

4:09:01we're going to implement this use case

4:09:03and find out why Hillary Clinton won so

4:09:06now let me give give you a visualization

4:09:09of the results so I'll just show the

4:09:11analysis that I have performed and I'll

4:09:14show you how it

4:09:15looks so this is zeppelin which is in my

4:09:18master node in my Hadoop cluster and

4:09:21this is where we're going to visualize

4:09:23our data so there's a lot of code don't

4:09:26be scared this is just Scala code with

4:09:28spark SQL and at the end you will be

4:09:30learning how to write this

4:09:32code so I'm just jumping onto the

4:09:35visualization part so this this is the

4:09:38first visualization that we've got and

4:09:40we've analyzed it according to different

4:09:42ethnicities of people for example in our

4:09:45xaxis we have foreign born persons and

4:09:47in y axis we're seeing that among the

4:09:49foreign born people what is the

4:09:51popularity of Hillary Clinton among the

4:09:53Asians and the circles represent the

4:09:55highest values the bigger circle is the

4:09:57bigger counts so we have made a few more

4:10:02visualizations so now we've got a line

4:10:04graph that compares the votes of Hillary

4:10:06Clinton and burn Bernie Sanders together

4:10:09again we have got an area graph also

4:10:12that compares Bernie Sanders and Hillary

4:10:14Clinton votes and hence we have a lot

4:10:17more

4:10:18visualization we have uh got our bar

4:10:20charts and everything finally we also

4:10:23have uh State and County wise

4:10:25distribution of votes so these are the

4:10:28visualizations that will help you derive

4:10:30a conclusion to derive an answer

4:10:33whatever answer that Donald Trump wants

4:10:36from you and don't worry you'll be

4:10:38learning how to do that I'll explain

4:10:40each and every detail of how I've made

4:10:42these

4:10:42visualizations so let's get started with

4:10:44Hadoop and Spark we will start um with

4:10:48an introduction to Hadoop and Spark so

4:10:53now let's take a look at what is Hadoop

4:10:55and what is spark so Hadoop is a

4:10:58framework where you can store large

4:10:59clusters of data in a distributed Manner

4:11:01and then process them parallell then

4:11:04Hadoop has got two components for

4:11:06storage it has hdf f s which stands for

4:11:08Hadoop distributed file system and it

4:11:10allows to dump any kind of data across

4:11:13the Hadoop cluster and it'll be stored

4:11:15in a distributed manner in commodity

4:11:17hardware for processing you've got yarn

4:11:20which stands for yet another resource

4:11:22negotiator and this is the processing

4:11:24unit of Hadoop which allows parallel

4:11:26processing of the distributed data

4:11:28across your Hadoop cluster in

4:11:30hdfs then we've got spark so Apache

4:11:34spark is one of the most popular

4:11:35projects by Apache and this this is an

4:11:37open-source cluster Computing framework

4:11:40for real-time processing where on the

4:11:43other hand Hadoop is used for batch

4:11:45processing spark is used for real-time

4:11:47processing because with spark the

4:11:49processing happens in memory and it

4:11:51provides you with an interface for

4:11:53programming entire clusters with

4:11:54implicit data parallelism and fault

4:11:57tolerance so what is data parallelism

4:12:00data parallelism is a form of

4:12:02parallelization across multiple

4:12:04processes in parallel Computing

4:12:06environments a lot of parallel words in

4:12:08that

4:12:09sentence um so let me tell you simply

4:12:11that it basically means Distributing

4:12:13your data across nodes which operate on

4:12:16the data parallel and it works on fault

4:12:18tolerant systems like hdfs and S3 and is

4:12:22built on top of yarn because with yarn

4:12:24you can combine different tools like

4:12:26Apache spark for better processing of

4:12:28your

4:12:29data and if you see the topology of

4:12:31Hadoop and Spark both of them have the

4:12:33same topology which is a Master Slave

4:12:35topology so in Hadoop if you consider in

4:12:37terms of hdfs the master node as known

4:12:40as the name node and the working node or

4:12:42the slave nodes are known as data node

4:12:44and in spark the master is known as

4:12:47master and slave are known as workers so

4:12:49this is these are basically demons so

4:12:52this is a brief introduction to Hadoop

4:12:54and Spark and now let's take a look at

4:12:56spark complimenting Hadoop there's

4:12:59always been a debate about what to

4:13:00choose Hado spark but let me tell you

4:13:03that there is a stubborn misconception

4:13:04that Apache spark is an alternative to

4:13:06had

4:13:07and that is likely to bring an end to

4:13:09the era for Hadoop it is very difficult

4:13:11to say Hadoop versus spark because the

4:13:13two framers are not mutually exclusive

4:13:16but they are better when they are paired

4:13:18with each other so let's see the

4:13:20different challenges that uh we address

4:13:22when we are using spark and Hadoop

4:13:25together you can see the first point

4:13:28that spark processes data 100 times

4:13:30faster than map ruce so it gives us the

4:13:33results faster and it performs faster

4:13:35analytics the next point is spark

4:13:38applications can run on yarn leveraging

4:13:40Hadoop cluster and you know that Hadoop

4:13:42cluster is usually set up on commodity

4:13:44Hardware so we are getting better

4:13:46processing but we are using very lowcost

4:13:49hardware and this will help us cut our

4:13:51cost a lot so hence also the cheap uh

4:13:54cost optimization the third point is

4:13:56that Apache spark can use hdfs as

4:13:59storage so you don't need a different

4:14:01storage space for Apache spark it can

4:14:03operate on hdfs itself so you don't have

4:14:06to copy the same file again and if you

4:14:08want to process it with spark uh so

4:14:11hence you can avoid duplication of files

4:14:13so Hadoop forms a very strong foundation

4:14:16for any of the future Big Data

4:14:17initiatives and Spark uh is one of those

4:14:20big data

4:14:22initiatives it's got enhanced features

4:14:24like in memory processing machine

4:14:26learning capabilities and you can use it

4:14:28with Hadoop and Hadoop uses commodity

4:14:30Hardware which can give you better

4:14:33processing with minimum cost these are

4:14:37are the benefits that you get when you

4:14:38combine spark and hadu together in order

4:14:40to analyze Big Data let's see some of

4:14:43the big data use cases so the first big

4:14:45data use case is web detailing the

4:14:48recommendation engines uh whenever you

4:14:50go out on Amazon or any other online

4:14:53shopping site in order to buy something

4:14:55you will see some recommended items

4:14:57popping below your screen or to the side

4:14:59of your screen and that is all generated

4:15:01using big data analytics and AD

4:15:04targeting if you go to Facebook you see

4:15:06a lot of different items asking you to

4:15:07buy them and when you got uh search

4:15:09quality abuse and click fraud detection

4:15:13you can use big data analytics and

4:15:15Telecommunications also in order to find

4:15:17out the customer churn prevention the

4:15:20network performance optimization

4:15:22analyzing uh Network to predict failure

4:15:25and you can prevent loss before the

4:15:27error or before the fault actually

4:15:29occurs it's also widely used by

4:15:31governments for fraud detection and

4:15:33cyber security in order to introduce

4:15:35different welfare schemes uh justice it

4:15:37has been widely used by Healthcare and

4:15:40Life Sciences for health information

4:15:42exchange Gene sequencing serialization

4:15:45healthc Care Service quality

4:15:46improvements and Drug safety now let me

4:15:49tell you that with big data analytics it

4:15:51has been very easy in order to diagnose

4:15:53a particular disease and find out the

4:15:55Cure also so these are some more big

4:15:57data use cases it is also used in Banks

4:16:01and financial services for modeling true

4:16:03risk fraud detection credit card scoring

4:16:06analysis and and uh many more it could

4:16:08be used in Retail transportation

4:16:10services hotels and food delivery

4:16:12services and actually every field you

4:16:14name no matter whatever business you

4:16:16have if you're able to use Big Data

4:16:18efficiently your company will grow and

4:16:20you will be gaining different insights

4:16:22by using big data analytics and hence

4:16:24improve your business even more nowadays

4:16:28everyone is using uh big data and you

4:16:31you've seen different fields and

4:16:32everything is different from each other

4:16:34but everyone is using big data analytics

4:16:37and Big Data analysis can be done with

4:16:39tools like Hado and Spark Etc so this is

4:16:43why big data analytics is very much in

4:16:45demand today and why it is very

4:16:46important for you to learn how to

4:16:47perform big data analytics with tools

4:16:49like this so now let's take a look at a

4:16:51big data use solution architecture as a

4:16:53whole you're dealing with big data now

4:16:57the first thing that you need to do is

4:16:58you need to dump all those that data

4:17:00into hdfs and store it in a distributed

4:17:03way and the next thing is to process

4:17:05that data so that you can gain insights

4:17:08and we'll be using yarn because yarn can

4:17:10allow us to integrate different tools

4:17:12together which will help us to process

4:17:14the Big Data these are the tools that

4:17:16you can integrate with yarn you can

4:17:17choose either Apache Hive Apache spark

4:17:20map reduce Apache CFA in order to

4:17:22analyze big data and Apache spark is one

4:17:25of the most popular and most widely used

4:17:27tools with yarn in order to process big

4:17:29data so this is the an entire solution

4:17:32as a whole now so let's take a look at

4:17:35Apache sparkk

4:17:37Apache spark is an open source cluster

4:17:39Computing framework for real-time

4:17:41processing and it has been the thriving

4:17:43open- Source community and is most

4:17:45active Apache project uh at this moment

4:17:48and Spark components are what make

4:17:50Apache spark fast and reliable and a lot

4:17:52of spark components were built to

4:17:54resolve the issues that cropped up while

4:17:56using Hado map reduce so Apache spark

4:18:02has got the following components has got

4:18:04the spark core engine now the core core

4:18:07engine is for the entire spark

4:18:08Frameworks uh every component is based

4:18:11on and it is placed in the core engine

4:18:14so at first we've got uh spark SQL so

4:18:17spark SQL is a spark module for

4:18:19structured data processing and you can

4:18:21run a modified Hive queries on existing

4:18:23hadb deployments and then we've got

4:18:25spark

4:18:26streaming now spark streaming is the

4:18:28component of spark which is used to

4:18:30process real-time streaming data and is

4:18:32useful addition to the core spark API

4:18:35because it enables hive throughput fault

4:18:37tolerance stream processing of live data

4:18:40streams and then we've got spark mli uh

4:18:44this is the machine learning library for

4:18:46sparc and we'll be using spark MMA in uh

4:18:50to implement machine learning in our use

4:18:52cases too and then we've got graphx

4:18:54which is the graph computation engine

4:18:56and this is the spot API for graphs and

4:19:00graph parallel computation it has got a

4:19:03set of fundamental operators like

4:19:04subgraph joint purchases Etc then um

4:19:09you've got uh spark R so this is the

4:19:12package for R language to enable our

4:19:15users to leverage spark power from our

4:19:17shell so the people who have already

4:19:20been working on R are comfortable with

4:19:23it and they can use R shell directly at

4:19:25the same time and they can use spark

4:19:28using this particular component which is

4:19:30spark R you can write all your code in

4:19:32the r shell and Spark will process it

4:19:34for you now let's take a deeper look at

4:19:36a real istic people and all these

4:19:38important components so we've got spark

4:19:41core and Spark core is the basic engine

4:19:44for large scale parallel and distributed

4:19:47data processing the core is the

4:19:49distributed execution engine and Java

4:19:52Scala and python apis offer a platform

4:19:54for distributed edl development and

4:19:57further additional libraries which are

4:19:59built on top of the core allow uh for

4:20:02diverse streaming SQL and machine

4:20:05learning it's it's also responsible for

4:20:07scheduling Distributing and monitoring

4:20:09jobs in a cluster and also interacting

4:20:12with storage systems let's take a look

4:20:15at the spark architecture so Apache

4:20:17spark has a well-defined and layered

4:20:19architecture where all the spark

4:20:21components and layers are Loosely

4:20:23coupled and integrated with various

4:20:25extensions and libraries first let's

4:20:27talk about the driver program this is

4:20:30the spark driver which contains the

4:20:32driver program and Spark context uh this

4:20:35is the Central Point and entry point of

4:20:38the spark shell and the driver program

4:20:40runs the main function of the

4:20:42application and this is the place where

4:20:43Spark context is

4:20:46created well what is spark context spark

4:20:49context represents the connection to the

4:20:51entire spark cluster and it can be used

4:20:54to create resilient distributed data

4:20:56sets accumulators and broadcast

4:20:59variables on that cluster and you should

4:21:01know that only one spark context may be

4:21:04active per Java virtual machine and you

4:21:07must stop any active spark context

4:21:10before creating a new one let's talk

4:21:12about the driver program that runs on

4:21:14the master knob of the spark cluster it

4:21:17schedules the job execution and

4:21:18negotiates with the cluster manager this

4:21:21is the cluster manager over here and the

4:21:24cluster manager is an external service

4:21:26that is responsible for acquiring

4:21:28resources on that spark cluster and

4:21:30allocating them to a spark job then in

4:21:35the worker node we have got the

4:21:36executors the executor is a distributed

4:21:39agent that is responsible for the

4:21:41execution of tasks and Every Spark

4:21:44application has its own executor process

4:21:47executors usually run for their entire

4:21:50lifetime of the spark application and

4:21:52this phenomenon is also known as static

4:21:55allocation of executors but you can also

4:21:57opt for dynamic uh locations of

4:22:01executors where you can add or remove

4:22:03spark executors dynamically to match

4:22:06with the overall workflow okay so now

4:22:08let me tell you what actually happens

4:22:10when the spark job is submitted when a

4:22:12client submits a spark user application

4:22:14code the driver implicitly converts the

4:22:16code containing Transformations and

4:22:18actions into a logical directed ayic

4:22:21graph or dag and at this stage the

4:22:24driver program also performs certain

4:22:26kinds of optimizations like pipelining

4:22:29Transformations and then converts The

4:22:31Logical dag into a physical execution of

4:22:34a plan with a set of stages and after

4:22:38creating a physical execution plan it

4:22:40creates more physical execution units

4:22:44that are referred to as tasks under each

4:22:46stage and these tasks are bundled to be

4:22:49sent to the spark cluster so the driver

4:22:52program then talks to the cluster

4:22:54manager and negotiates for resources and

4:22:56the cluster manager then launches the

4:22:59executors on the worker nodes on behalf

4:23:01of the driver and at this point the

4:23:03driver sends tasks to the cluster

4:23:05manager based on the day of replacement

4:23:07and before the executors begin execution

4:23:10they first register themselves with the

4:23:11driver program so that the driver has

4:23:14got a holistic view of all the

4:23:17executors now the executors will execute

4:23:19the various tasks that are assigned to

4:23:21them by the driver program and at any

4:23:22point of time when the spark application

4:23:25is running the driver program will keep

4:23:27the on monitoring the set of executors

4:23:29that are running the spark application

4:23:31code and this driver program here also

4:23:34schedules future tasks based on data

4:23:37replacement by tracking the location of

4:23:39the cache data so I hope you have

4:23:41understood the architecture of spark any

4:23:44doubts all right no doubts now let's

4:23:47take a look at spark SQL and its

4:23:49architecture so spark SQL is the new

4:23:51module in spark and it integrates

4:23:53relational processing with Spark's

4:23:55functional programming API and it

4:23:57supports querying of data either by a

4:23:59SQL or via Hive query language so for

4:24:02those of you who have been familiar with

4:24:05rdbms uh so spark SQL will be a very

4:24:08easy transition from your earlier tools

4:24:11because you can extend the boundaries of

4:24:13traditional relational data processing

4:24:15with spark SQL and it also provides s

4:24:18support for various data sources and

4:24:21makes it possible to read SQL queries

4:24:23with code transformation and that is why

4:24:25spark SQL has become a very powerful

4:24:28tool this is the architecture of spark

4:24:30SQL so let's talk about each of these

4:24:32components one by one the first we have

4:24:35got the data source our API so this is

4:24:37the universal API for loading and

4:24:39storing structured data and it is built

4:24:42on support for Hive Avro Json jdbc CVS

4:24:47parkette Etc so it also supports the

4:24:50third-party integration through spark

4:24:52packages then you've got the data frame

4:24:54API dataframe API is the distributive

4:24:58collection of data that is organized uh

4:25:00into named columns and is similar to

4:25:03relational table in SQL that is used for

4:25:05storing data in tables so it is the

4:25:08domain specific language applicable to

4:25:10or

4:25:11DSL applicable on structured and

4:25:14semi-structured data so it processes

4:25:16data from kilobytes to pedabytes on a

4:25:19single node cluster to a multi- noode

4:25:21cluster and it provides different apis

4:25:23for python Java Scala and our

4:25:26programming so I hope you have

4:25:28understood all the architecture of spark

4:25:30SQL we will be using spark SQL in order

4:25:32to solve our use cases so these are the

4:25:34different commands to start start the

4:25:36spark Damons these are very similar to

4:25:38had of commands to start The hdfs Damons

4:25:41so you can see to start all the spark

4:25:44Damons uh so the spark Damons are master

4:25:46and worker and you can use this command

4:25:48to check if all the Damons are running

4:25:50on your machine you can use JPS like

4:25:53Hadoop and then in order to start the

4:25:55spark shell you can use this and you can

4:25:58go ahead and try this out so this is

4:26:00very similar to the hadu part that I

4:26:01just showed you earlier so I'm not going

4:26:03to do it again and then we've seen a

4:26:06Pache spark also so now let's take a

4:26:08look at K means and Zeppelin K means is

4:26:10the clustering method and Zeppelin is

4:26:12what we're going to use in order to

4:26:14visualize our data so let's talk about

4:26:17the K mean clustering now K means is one

4:26:20of the most simplest UNS supervised

4:26:22learning algorithms that uh solves the

4:26:25well-known clustering problem so the

4:26:27procedure of K means follows a simple

4:26:30and easy way to classify a data set to a

4:26:33certain number of clusters which is

4:26:35fixed prior to performing the clustering

4:26:37method so the main idea is Define K

4:26:40centroids one for each cluster and the

4:26:42centroids should be placed in a very

4:26:46cunning way because of different

4:26:48location re causes different results so

4:26:52here let's take an example so let's say

4:26:54that we want to Cluster total population

4:26:56of a certain location and so we want to

4:27:00Cluster them into four different uh

4:27:03clusters namely group one two and three

4:27:05and four so the main thing that we

4:27:06should keep in mind is that the objects

4:27:08in group one should be as similar as

4:27:11possible but there should be as much

4:27:13difference between an object in group

4:27:14one and group two it means that the

4:27:16points that are lying in the same group

4:27:19should have similar characteristics and

4:27:21it should be different from the points

4:27:23that are lying in a different cluster

4:27:25and the attributes of the objects are

4:27:27allowed to determine which object should

4:27:30be grouped

4:27:31together for example let us uh take in

4:27:36the same sample that we're using in the

4:27:37US County so let's consider the second

4:27:40data set we have used there are a lot of

4:27:42features that I already told you like

4:27:44there are age groups and they are

4:27:46categorized by professions and they also

4:27:48categorized by the ethnicity and uh so

4:27:52this is the thing that we are talking

4:27:54about so these are the attributes that

4:27:57will allow us to Cluster our data so

4:28:00this is K means

4:28:02clustering here is one more example let

4:28:04us consider a comparison income and

4:28:06balance so in my x- Axis I've got the

4:28:08gross monthly income and in the Y AIS I

4:28:11have the current balance I want to

4:28:13Cluster my data according to these two

4:28:17attributes here if you see this is my

4:28:20first cluster and this is my second

4:28:23cluster so this uh is the cluster that

4:28:26indicates the people who have high

4:28:28income and low balance in uh the account

4:28:31and they spent a lot and this cluster

4:28:33comprises of the people who have got a

4:28:35low income but a high balance and they

4:28:38are safe you can see that all the points

4:28:41that are lying here have got similar

4:28:43characteristics that they have got low

4:28:45income and high balance and here are the

4:28:47people who share the same

4:28:49characteristics where they have uh got

4:28:51low balance and high income and there

4:28:55are a few outliers here and there but

4:28:57they don't uh form a cluster so this is

4:29:01an example of K means clustering and

4:29:02we'll be using that in order to solve

4:29:04our problems so does anybody have any

4:29:07questions so here is one more example

4:29:10and one more problem for you so you guys

4:29:12will tell me now so the problem is that

4:29:15I want to set up schools in my city and

4:29:17these are the points which indicate

4:29:20where each student lives so my question

4:29:23to you is where should I be building my

4:29:25school if I have students living um

4:29:28around the city in these particular

4:29:30locations and in order to find that out

4:29:33we will do K means clustering and we'll

4:29:35find out the center point right so if

4:29:38you can cluster and make groups of all

4:29:40these locations and set up schools at

4:29:42the center point of each cluster that

4:29:44would be Optima isn't it because that is

4:29:47how the students have to travel less it

4:29:50will be close to everyone's house and

4:29:52there it is so we have formed three

4:29:55clusters so you can see the brown dots

4:29:57are one cluster and the blue dots are

4:29:59one cluster and the red dots are one

4:30:01cluster and we have uh set up schools in

4:30:04the center points of each cluster so

4:30:07here is one here is one and here is yet

4:30:10another one so this is where I need to

4:30:12set my schools up so that my students do

4:30:15not have to travel that much so that was

4:30:17all about c means and now let's talk

4:30:19about Apache Zeppelin this is a web page

4:30:22notebook which brings in data ingestion

4:30:24data exploration visualization sharing

4:30:27and collaboration features to Hadoop and

4:30:29Spark so remember when I showed you my

4:30:32Zeppelin notebook you can see that we

4:30:34have written the code code there we have

4:30:37even run SQL codes there and we have

4:30:40more visualizations by executing code

4:30:42there so this is how interactive

4:30:45Zeppelin is and it supports many many

4:30:47interpreters and it is a very powerful

4:30:50visualization tool that can use uh that

4:30:54goes very well with Linux systems and it

4:30:56supports a lot of language interpreters

4:30:59it supports R python and a lot of other

4:31:03interpreters so now let's move on on to

4:31:05the solution of the use case so this is

4:31:07what you've been waiting for first we

4:31:09will solve our us County solution so the

4:31:11first thing we will do is we will store

4:31:13the data into hdfs and then we will

4:31:16analyze the data by using Scala spark

4:31:18SQL and Spark ml Li and then uh finally

4:31:22we'll find out the results and visualize

4:31:24them using Zeppelin so this was the

4:31:27entire us election solution strategy

4:31:29that I told you and I don't think I

4:31:31should repeat it again but if you want

4:31:32me I can uh should I repeat

4:31:36all right so most of the people are

4:31:38saying no so I will go right through

4:31:39this one again so let me just go to my

4:31:41VM and execute this for you so this is

4:31:45my Zeppelin and I opened my notebook

4:31:47here and let us go to my us election

4:31:49notebook and this is the code so first

4:31:52of all what I'm going to do is that I am

4:31:54importing certain packages because I'll

4:31:56be using certain functions that are in

4:31:58those packages so I've imported spark

4:32:00SQL packages and I have also imported

4:32:03spark ml lib packages because I'll be

4:32:06using K means clustering so Vector

4:32:09assembler enables me certain machine

4:32:11learning functions so over here I have

4:32:13the vector assembler package that gives

4:32:15me certain machine learning functions

4:32:16that I'm going to use I've also imported

4:32:19K means package because I'll be using K

4:32:20means clustering then the first thing

4:32:22that you need to do is that you need to

4:32:24start the SQL context so I have started

4:32:27my spark SQL context here and the next

4:32:29thing that you need to do is that you

4:32:30need to define a schema because when you

4:32:33want to dump our data set or we want to

4:32:35to dump our data it should be in a

4:32:37particular format and we have to tell

4:32:38spark in which format it should be so

4:32:41we're defining a schema here so let me

4:32:44take you uh through the code so I'm

4:32:47storing schema in a variable called

4:32:49schema and we have to define the schema

4:32:52in a proper structure so we're going to

4:32:54start with struct type and since you

4:32:55know that our data set has got different

4:32:57fields as columns we're going to Define

4:33:00this as an array of fields then this is

4:33:02an array instruct so we are defining the

4:33:05different Fields now so we'll start with

4:33:07the first field by defining it as struct

4:33:09field inside the braces which should

4:33:10mention what would should be the name of

4:33:13that particular field so I've named it

4:33:15as state it should be a string type and

4:33:18true that means it is a string type the

4:33:21next we've got fips which is of string

4:33:23type now I know that fips is a number

4:33:25but since we are not going to do any

4:33:26kind of numeric operation on fips uh

4:33:29we're going to let it stay as a string

4:33:31then we've got party as a string type

4:33:33candidate as a string type and then

4:33:34votes as integer type because we're

4:33:36going to count the number of votes and

4:33:38there is going to be certain numeric

4:33:40operations that we are going to perform

4:33:42that will help us to analyze our data

4:33:45then we've got a fraction votes which

4:33:47you know is a decimal type so we have to

4:33:49keep it as double type the next thing

4:33:51you need to do is that spark needs to

4:33:53read the data set from the

4:33:55hdfs so for that you have to use the

4:33:58command spark read option header true

4:34:01header true means that you have

4:34:03mentioned and you have told spark that

4:34:05my data set already contains column

4:34:07headers because State as ABR they are

4:34:10nothing but they are column headers so

4:34:12you don't have to explicitly Define the

4:34:14column headers uh for it neither will

4:34:17spark choose any random row as a column

4:34:19header so it will choose only the column

4:34:21headers uh your data set has then you

4:34:24have to mention the schema that you have

4:34:26defined so I have defined it in my

4:34:28variable schema so that's why I have

4:34:30mentioned it in my file should be in CSV

4:34:33format and then I have mentioned the

4:34:35path of the file in my hdfs this is the

4:34:38path and I store this entire data set in

4:34:41my variable

4:34:42DF now what I am going to do is that I'm

4:34:46going to divide up certain rows from my

4:34:48data set because you know that my data

4:34:50set contains both the Republican and

4:34:52Democrat data and I just want the

4:34:54Democrat data right because we're going

4:34:55to analyze the Hillary Clinton and

4:34:57Bernie Sanders part okay so this is how

4:34:59you divide your data set so the first

4:35:01thing that we have done is that we have

4:35:03created one more variable called DF far

4:35:05and we have replied A filter where party

4:35:07is equal to Republican and then we are

4:35:10storing the Democrat Party data into DFD

4:35:13so we're going to use the DFD from

4:35:16onwards and dfr the Republican data is

4:35:19going to be your assignment for the next

4:35:21class now I am going to analyze the

4:35:23Democrat data and then after this class

4:35:26is over I want you guys to take the

4:35:28Republican data this data set is already

4:35:31available in your elements and you've

4:35:33got the VMS also with everything

4:35:35everything installed so please when you

4:35:37are at home when you have free time just

4:35:39analyze the Republican data and tell me

4:35:41uh what were the reasons that Donald

4:35:43Trump want I want you to do all that

4:35:45analysis and come up with that in the

4:35:47next class and we'll discuss about it

4:35:49and whatever results and conclusions

4:35:52that you have made after analyzing the

4:35:54Republican data and that way you'll also

4:35:56learn even more and it will also be

4:35:59practice for you after today's class so

4:36:01all right so we are going to take DFD

4:36:04now in the first thing that we will do

4:36:06is that we will create a table View and

4:36:07I'm going to name the table view as

4:36:09election and let me just show you what

4:36:11it looks like and what it has so this is

4:36:14the command that I have run in Zeppelin

4:36:16so this is SQL code that I have run in

4:36:20Zeppelin and you can see that I have got

4:36:23States state abbr and I have only got

4:36:26the Democrat data all right let's go

4:36:29back all right so after creating the

4:36:32table view now all of the Democrat data

4:36:34is in in my election table so now what

4:36:37I'm going to do is that I'm creating a

4:36:39temporary variable and I'm running spark

4:36:41SQL code so what I'm actually doing by

4:36:43writing this code the motive of writing

4:36:45this SQL code or the SQL query is that I

4:36:48want to refine my data even more so what

4:36:50I'm trying to analyze here is how a

4:36:52particular candidate actually won I

4:36:53don't have to do anything with the

4:36:54losing data because you know that each

4:36:56of fips contain one of the losing

4:36:58candidate members and one of the winning

4:37:00candidate members it contains the data

4:37:02of the winning candidate and the losing

4:37:03candidate also because my data set

4:37:05contains both of the data of Bernie

4:37:07Sanders and Hillary Clinton in some

4:37:09parts Bernie Sanders won and in some

4:37:11counties Hillary Clinton won so I just

4:37:14want to find out uh that who are the

4:37:16winners in a particular County okay so

4:37:19I'm going to refine that data and for

4:37:21that I'm using this query so I'm going

4:37:24to uh select all from election and then

4:37:26I'm going to perform an inner join with

4:37:29their query so this is one more query

4:37:31inside this query and let me tell you

4:37:33what I'm actually doing so first of all

4:37:36what we have done is that we have

4:37:37selected fips as B you know that now you

4:37:40have got two entries for each fips so

4:37:43each fips actually appears twice in the

4:37:45data set so I've named it as B and now

4:37:47we are counting the maximum fraction

4:37:49votes so you know that in each FIP we

4:37:52have the maximum fraction vote and then

4:37:54we can find the winner by actually

4:37:55seeing who has got the maximum fraction

4:37:57votes then we have named it as a the

4:38:00maximum fraction votes column is named

4:38:02as a and we are grouping by Fifth

4:38:06so now each of my fips will be selected

4:38:08which has a maximum fraction vote and I

4:38:11have uh two columns for that fips which

4:38:13is one1 and

4:38:15one1 so the only rle will be selected

4:38:18which has the maximum fraction votes now

4:38:21I'll have the winner data and I've named

4:38:23this entire table inside this query as

4:38:26group TT and then I'm validating it as

4:38:29where election. fips the main table

4:38:31view. fips should be equal to the B

4:38:34column that we have created in group TT

4:38:36table and election. fraction votes

4:38:38should be equal to group tt. a so any

4:38:42doubts on this query and about how I

4:38:44have written this all right so now what

4:38:47we're going to do is that whatever data

4:38:49that we've got here I'm storing that in

4:38:51election one let me just show you what

4:38:54is in election one now so this is my

4:38:57election table only and uh you can see

4:38:59that I've got two fips so 1067

4:39:031067 now let me show you election one so

4:39:06there now I can see that I don't have

4:39:09repetition of fips I have only one entry

4:39:13for fips and that is the row which tells

4:39:16me who won in that county or in that

4:39:19particular FIP or in the FIP associated

4:39:21with a particular County you can see for

4:39:23Bullock it was Hillary Clinton for kahun

4:39:25it was Hillary Clinton Cherokee also

4:39:27Hillary Clinton and then state house

4:39:29district 19 is Bernie Sanders so Alaska

4:39:32is mainly Bernie Sanders so this is what

4:39:36we've done now and then you can see that

4:39:38we have also got additional columns as b

4:39:42and a so a tells you the maximum

4:39:44fraction votes and B tells you the

4:39:47fips so the data in fips and the data in

4:39:50B are the same and data in fraction

4:39:52votes and data in a is the same right

4:39:55what I'm going to do now is since my

4:39:57columns are repeating and they have the

4:39:59same value I Don't Want A and B now

4:40:02right so what I'm going to do is I'm

4:40:04going to filter out the columns I don't

4:40:06need and in this case I don't want b and

4:40:08a and what I'm going to do is I'm going

4:40:10to make a temporary variable again so

4:40:12I'm using the temporary variable to

4:40:14store some data temporarily so I'm

4:40:16writing to the spark SQL code uh to

4:40:19select Only The Columns that I want I

4:40:21want the state state abbreviation County

4:40:24fips Party candidate votes fashion votes

4:40:26from election one I'm storing everything

4:40:28in D winner I've created this new

4:40:30variable and whatever there was in temp

4:40:33I'm assigning it to D winner and now

4:40:35I've uh got only the winner data so I

4:40:38have got all the counties and I've got

4:40:40uh who won in that particular County and

4:40:41by how much in the fraction of votes

4:40:43what I'm just doing till now is that I'm

4:40:46just refining our data set so that it

4:40:48will be easy for us to make some

4:40:49conclusions or gain some insights from

4:40:52that data right and also let me tell you

4:40:54that it's not always necessary that you

4:40:56are filing your data set in the exact

4:40:59way that I'm doing it if you have

4:41:00something in mind after you've seen your

4:41:02data and understand your data and you

4:41:03found out what actually you need to do

4:41:06you can carry out different steps to do

4:41:08that also this is just one way of doing

4:41:09it and this is my way of doing it so I'm

4:41:11just telling you and then we are

4:41:14creating a table for D winner and we are

4:41:16going to name it as Democrat so let me

4:41:18go again and let me show you what the

4:41:20Democrat table view looks like you can

4:41:23press shift enter so there you uh have

4:41:27we have column A and B that we had in

4:41:29election

4:41:32one and so I have just got uh winner

4:41:37data so now let us go back and find out

4:41:40what we're going to find is that I want

4:41:43to find out that uh which of the

4:41:46candidates won my state and then

4:41:48whatever date and whatever result I'll

4:41:50get will be stored in the temporary

4:41:52variable when I'm assigning everything

4:41:55uh that will be stored in the temporary

4:41:57variable to a new variable called dstate

4:41:59and then similarly I'm going to create a

4:42:02table view for dstate which is State let

4:42:06me show you what my state table view

4:42:08actually contains so there it is so I've

4:42:12got State Connecticut Hillary Clinton W

4:42:14155 counties Florida Hillary Clinton won

4:42:1658 counties so this is what we've come

4:42:18up to for our first data set so now

4:42:22let's see what we can do with our second

4:42:23data set that contains all the different

4:42:25demographic

4:42:26features uh first thing again you have

4:42:29to define a schema and this time I'm

4:42:31naming that schema uh schema one say

4:42:34since you know that we have got almost

4:42:3754 columns so I have to Define all those

4:42:4054 columns

4:42:41also so you remember what th each of

4:42:44those columns contains so this is

4:42:47exactly what I have done and I don't

4:42:49need to go through every line but I like

4:42:51I already told you how to define a

4:42:53schema you can have the code in your LMS

4:42:56so you can take a look at it so the next

4:42:58thing we're doing again we have to read

4:43:00our data set and I'm storing my data set

4:43:02into a new variable called df1 and this

4:43:05is the path in my htfs where my data set

4:43:07was and then I have created a table view

4:43:09for my data set which is called

4:43:13facts now let me show you what facts

4:43:19contain as you can see that it contains

4:43:23abbreviation state abbreviation

4:43:24population 2014 so instead of uh using

4:43:30the code now or the encoded form that

4:43:33was actually there in my data set I have

4:43:35given a varied metan name that would

4:43:38describe what it contains right so

4:43:40instead of PST 214 I've got population

4:43:432014 so does that make sense right and

4:43:47contains all the 54 demographic features

4:43:49or different features that was in my

4:43:51data

4:43:53set white alone not Hispanic or Latino

4:43:57living in the same house one year and

4:43:59over foreign born persons language or

4:44:01other than English spoken at home High

4:44:03School gradu or higher uh so it contains

4:44:06basically all the different features or

4:44:08all the different columns that actually

4:44:09was in my data set and that I have

4:44:11defined in my schema so this is what

4:44:14facts I have so now what I'm going to do

4:44:16is that I'm not going to analyze my

4:44:18whole data based on all this different

4:44:20features I'm going to choose some

4:44:22specific features in order to analyze it

4:44:26uh I'm going to take just a few add one

4:44:29so these are the different features that

4:44:32I'm going to use I'm going to use fips

4:44:34I'm going to use state I'm going to use

4:44:35state abbreviation then area name

4:44:38candidate and people who are over 65

4:44:40years senior citizens a female people

4:44:43white Alone um black African alone I'm

4:44:47choosing Asian alone Hispanic or Latino

4:44:50basically what I'm trying to do is I'm

4:44:51trying to check what is the popularity

4:44:53of Hillary Clinton among the foreign

4:44:55people or people from different

4:44:57ethnicities so I'm choosing white people

4:44:59black people and Hispanic people so I'm

4:45:01just trying to analyze it okay and you

4:45:04know that I have stored this in a

4:45:05temporary variable again and then

4:45:07whatever result I'll get by running this

4:45:09spark SQL code I'll store it in a

4:45:12different variable called DFX and then

4:45:14I'll store it and then I'll make a table

4:45:16view for DF facts such as winter

4:45:20facts so let me show you what winter

4:45:23facts look like so it's winter facts

4:45:26you've got uh fips the state is Alabama

4:45:28state abbreviation is Al for

4:45:32Alabama um the area name is our tuga

4:45:37County and the winner was Hillary

4:45:39Clinton and the people over 65 years in

4:45:42that particular county is 13.8% female

4:45:45percentage is 51.4 white alone 77.9 and

4:45:49so these show you the data

4:45:52so uh black white or African is 18% and

4:45:57then I've got the different fields that

4:45:59I have selected Asian alone Hispanic or

4:46:02Latino foreign born so so I have chosen

4:46:0514 features to analyze it from all right

4:46:08so now what I'm doing again is that I'm

4:46:10going to divide the Hillary Clinton data

4:46:11and the Bernie Sanders data so that we

4:46:13can analyze only why Hillary Clinton won

4:46:16or why Bernie Sanders won in some

4:46:18particular counties so we are planning

4:46:21to filter the same way we divided

4:46:23Democrats and Republican data from our

4:46:25initial primary result data set so this

4:46:29is what you have done so you know that

4:46:31is stored in DF fact so we are putting

4:46:33the filter in DFX where uh candidate is

4:46:37equal to Hillary Clinton so that will be

4:46:39stored in HC and the data of Bernie

4:46:41Sanders will be stored in BS so after

4:46:44that what we are doing is that we are

4:46:46doing a one hot encoding so we'll add

4:46:49two more columns in our data set a WBS

4:46:51and

4:46:52whc in this case we are going to do one

4:46:56hot encoding and what we're going to do

4:46:59is that we are going to include or we

4:47:02are going to attach two more columns in

4:47:04Winter facts as whhc and WBS so it'll

4:47:10just contain either one or zero and so

4:47:14you can edit it in that way whichever

4:47:15County so if you're considering a county

4:47:18let's say aruga County say if Hillary

4:47:20Clinton is the winner it will have a one

4:47:22in whc and in WBS it will have zero

4:47:27similarly in which Count's Bernie

4:47:28Sanders one so Bernie Sanders will have

4:47:30one so WBS will have one and whc will

4:47:33have have a

4:47:35zero and then we are creating different

4:47:38views for both of these two together so

4:47:40this will only tell me wherever whc is

4:47:43one that means this will only show me

4:47:45the counties where Hillary Clinton won

4:47:47this will only show me the counties

4:47:48where Bernie Sanders won and we are

4:47:51creating a view for both of these so for

4:47:53Bernie Sanders the view is WBS and for

4:47:56Hillary Clinton it's whc then finally we

4:47:58are merging both of them together using

4:48:00Union so select all from whc Union all

4:48:03select from WBS and finally you have

4:48:05stored it in result and we have created

4:48:07a table view known as result so let me

4:48:10show you what this result

4:48:12contains uh so there it is for UGA it

4:48:15was Hillary Clinton so we've got the

4:48:17Bernie Sanders

4:48:19data over here at the bottom and I've

4:48:22got all the different fields also from

4:48:24my uh second data

4:48:28set the different features that I chose

4:48:31from my second data set uh to to analyze

4:48:34it so now comes the actual analyzing

4:48:36part this is where we're going to

4:48:37perform K means but first we have to

4:48:40define the feature columns actually you

4:48:42have to Define what is the input that

4:48:45you're going to feed so that you get an

4:48:48output so this is actually the input

4:48:50that you are going to feed the to the

4:48:52machine so that the machine learning

4:48:54goes on and finally it gives you some

4:48:56kind of result right so this is where

4:48:58I'm defining again I'm using an array to

4:49:00Define all the different fields from my

4:49:03data sets I'm using person 65 years and

4:49:05older female person percentage white

4:49:07alone or black or African uh American

4:49:10alone Asian alone Hispanic or Latino

4:49:13foreign born persons language other than

4:49:17English spoken at home bachelor degree

4:49:19or higher veterans home ownership rate

4:49:21median household uh income persons below

4:49:24poverty level and population per square

4:49:26mile whc and WBS and then I'm going to

4:49:30use the vector assembler so this is what

4:49:32enables different machines learning

4:49:34algorithms where we are using K means so

4:49:38my input column is features calls so

4:49:41this is going to be the input in my

4:49:43output column and will be called

4:49:46features so whatever result that I'm

4:49:49going to get is features and we have to

4:49:51transform the results so this is the

4:49:54final table view that we have created

4:49:56and you know what transforming means and

4:49:57transforming again means so in our

4:50:00strategy we already saw that we have to

4:50:01transform the data first so my updated

4:50:04data set was results so I'm going to

4:50:06transform result and put columns as

4:50:08going to be these which is feature

4:50:11columns and output uh table view will be

4:50:14called features and then we're going to

4:50:16perform the K means clustering and we're

4:50:18going to store it in a variable called K

4:50:20means so we're using different functions

4:50:22from spark

4:50:23M library and we have chosen spark with

4:50:27uh clustering K means and you know that

4:50:30in K means we already defined that how

4:50:32many clusters do we need and we need

4:50:35four so we have selected four clusters

4:50:37and then we are going to set feature

4:50:39columns as features and then set

4:50:41prediction column as

4:50:43predictions so after that we're going to

4:50:46make a model and we have defined our

4:50:48input and output columns in row so we're

4:50:50going to use uh kefit row and whatever

4:50:53predictions we will get we're going to

4:50:55store it in a model and then we um are

4:50:59going to do this and that we are going

4:51:01to print the cluster centers for each

4:51:04cluster so let me show you what my

4:51:06cluster centers are so after we run this

4:51:10code you can see that these are the

4:51:12different cluster centers know so just

4:51:13what I can make you understand about

4:51:15what we're going to do after K means

4:51:17clustering and how to analyze it so the

4:51:19numbers are present they are placed very

4:51:22haphazardly so what I have done is that

4:51:25I've picked out each of the cluster

4:51:26Center points and then I have made a new

4:51:29table yes so this is it so you know that

4:51:32we have four clusters we have got the

4:51:34zero cluster first cluster second

4:51:36cluster and uh uh third uh so 0 two 3

4:51:42okay so four clusters and we have found

4:51:44out this uh cluster centers according to

4:51:47different features that we fed into my K

4:51:50means algorithm so what we observed here

4:51:54in whc and WBS is that the winning

4:51:56percentage or the winning chances of

4:51:58Hillary Clinton was 0.9 whereas winning

4:52:00chances for Burnie Sanders was

4:52:020.1 and and uh

4:52:06then if You observe the differences in

4:52:09the cluster centers for each feature

4:52:11here you can see that there is not much

4:52:13difference not even here so it's uh 50

4:52:1749 49 51 and then uh it's well again it

4:52:22is not much of a difference but if you

4:52:24see here that it's nine and it's uh

4:52:26going to 16 so you can do a more

4:52:29detailed analysis on black or

4:52:32African-American so if you want want to

4:52:33know the real support of black or

4:52:36African-American and you want to see uh

4:52:40what was their voting pattern or how

4:52:43popular was Hillary Clinton among them

4:52:46so maybe this could be a good field to

4:52:48analyze because you see the variations

4:52:50in the number similarly you can check

4:52:52out other features and you can check out

4:52:55uh here at 16 89 and 36 so maybe again

4:52:58Hispanic or Latino field and you should

4:53:01uh do some more analysis on it and even

4:53:03here you can see in veterans there's

4:53:0447,800 whereas we've got 182,000 all so

4:53:10there is uh also a lot of

4:53:12difference um then here is only 20 2759

4:53:17and we've got uh in the 10,000s we've

4:53:20got numbers and even 100,000s here so

4:53:23this is how we can identify that which

4:53:25are the fields or which we can find the

4:53:26main reasons of the main points where

4:53:28you should make your analysis so let's

4:53:30go back to our Zeppelin notebook and

4:53:33here it is is so now what we're going to

4:53:36do is that we're going to visualize the

4:53:37result first so we are counting from

4:53:41predictions so you can see that in

4:53:43cluster ones the prediction means

4:53:45prediction contains my cluster since you

4:53:47know that I've have stored my clusters

4:53:49my cluster information and prediction

4:53:51this is my output after K means uh so

4:53:55I've got uh these this many clusters so

4:53:58this is the count of my counties or a

4:54:00count of different various that lie in

4:54:02my particular cluster you can see that

4:54:05in cluster one I have got 1917 and

4:54:08cluster 2 I've got 751 so maybe I should

4:54:11pay more attention on analyzing cluster

4:54:13one right uh so that's why I've SE

4:54:16selected cluster one here and we're

4:54:18making different predictions so you can

4:54:20see that in the x axis I have got foure

4:54:23inborn people and in y AIS I have

4:54:26chosen uh language other than English

4:54:29spoken at home and then we are grouping

4:54:32it by candidate so you can see the

4:54:34lighter blue is for Bernie Sanders and

4:54:36the more dark blue is for Hillary

4:54:38Clinton so all this light blue is for

4:54:40Bernie Sanders and you can see that as

4:54:43the number of foreign people increases

4:54:46you can only see Hillary Clinton in the

4:54:48scattered plot here so there might be a

4:54:50few outliers like back here in the size

4:54:52of defined according to black or

4:54:54africanamerican alone you so you

4:54:57remember that this was the feature where

4:54:59we find a lot of variations in the

4:55:01numbers so that's why we grouped it

4:55:02according to that

4:55:04and you can see the bigger the circle

4:55:06represents the more black or

4:55:08African-American alone and that's what

4:55:11what the conclusion we can find out from

4:55:12this scatter plot and we can see that as

4:55:14the number of foreign people increases

4:55:16the popularity of Hillary Clinton is uh

4:55:20more in larger groups of foreign people

4:55:23you can also choose different parameters

4:55:25out of all the different features that

4:55:26you have chosen so remember uh that we

4:55:29have also seen the variation in veterans

4:55:31so let's choose veterans in a y axis so

4:55:33let's also change xaxis and let me just

4:55:35use white alone here so you can see here

4:55:38that

4:55:39uh uh there is the x axis that has white

4:55:45alone and this is the Veterans so you

4:55:47can see that Hillary Clinton is popular

4:55:49among veterans also in a smaller group

4:55:52of veterans since we have decided the

4:55:54size in black or africanamerican alone

4:55:57so the size um also represents some

4:56:01values she is popular among the

4:56:03africanamerican veterans and then uh as

4:56:07you go ahead and as the count increases

4:56:09you can see actually since it's a

4:56:11scatter plot and it almost represents

4:56:14that uh this is a point as the number of

4:56:17people increases or as the number of

4:56:20white people increases the votes are

4:56:21equally kind of distributed between

4:56:24Bernie Sanders and Hillary Clinton

4:56:26because there are a lot of points in

4:56:27this scatter plot over here and you can

4:56:30go ahead and drag and drop different

4:56:31features and you can make different

4:56:33visualizations on that now what we've

4:56:36done is that we know that there are 1917

4:56:39counties in my cluster one so I'm am

4:56:41going to do is that I'm going to see

4:56:44that among these

4:56:451917 how many were in favor for Hillary

4:56:48Clinton and how many were in favor for

4:56:50Bernie

4:56:51Sanders

4:56:53um so in cluster number one you can see

4:56:56clearly Hillary Clinton is the winner

4:56:58and Bernie Sanders only has got 764

4:57:01whereas she got 11

4:57:04153 similarly in cluster 3 again Hillary

4:57:07Clinton is the winner with nine and

4:57:08Bernie Sanders uh with

4:57:12one then it uh two she's also got 388

4:57:17and Bernie Sanders was 363 so this was

4:57:20very close call and again in zero you've

4:57:22got 119 and

4:57:2530 and then we went ahead and create a

4:57:27line chart also of the word distribution

4:57:29for Hillary Clinton and Bernie Sanders

4:57:31so in Keys we have selected prediction

4:57:34the values here are whc and WBS the sum

4:57:37that we have got over here so definitely

4:57:40Bernie Sanders is lagging behind so even

4:57:42though you don't have that table for you

4:57:44you can also find it out according to

4:57:47this line chart so you can see that in

4:57:49cluster zero even again Hillary Clinton

4:57:51was ahead of Bernie Sanders in cluster 2

4:57:55there was a very neck to neck

4:57:56competition and you can see it in the

4:57:58graph year so this res represents

4:58:01cluster 2 and so you can see you have a

4:58:04neck to neck competition and again in

4:58:07cluster

4:58:08three uh they have got neck to neck

4:58:10competition so this describes the

4:58:13distribution of votes for Hillary

4:58:14Clinton and Bernie Sanders and

4:58:16definitely Hillary Clinton uh knows

4:58:20ahead and that's why of course she won

4:58:22the primary elections so again you can

4:58:24go ahead and we have created the same

4:58:27graph it's uh only just area graph

4:58:29instead of a line graph the key here are

4:58:33are State and candidates so I've got

4:58:35States and candidates over here and the

4:58:38values is counties once uh if you just

4:58:41hover onto this bar chart you can see

4:58:43that in Connecticut Bernie Sanders won

4:58:46115 counties in Connecticut Hillary

4:58:49Clinton won 55 only so in Florida

4:58:52Hillary Clinton is 58 and in Florida

4:58:54Bernie Sanders is nine and here you can

4:58:56see in uh the main Bernie Sanders won

4:58:59462 so Bernie Sanders got a majority of

4:59:02votes for

4:59:04Maine so you can also classify it

4:59:07statewise you can find out which uh are

4:59:10the states and as Donald Trump now you

4:59:13will know that which are the states that

4:59:15you can Target right so you know that in

4:59:17Maine a lot of people voted for Bernie

4:59:19Sanders and maybe Hillary uh Clinton is

4:59:23not popular so you can go ahead and lead

4:59:25out so as Donald Trump's party member

4:59:28you can just advise him to go in Maine

4:59:31and carry out different campaigns

4:59:33because uh Hillary Clinton is not so

4:59:36popular there so maybe it would be a

4:59:38little easier to get votes from the

4:59:40people in Maine so this is what you can

4:59:43make a conclusion from it might not be

4:59:45very accurate but this would be very

4:59:46close the thing is that you can make

4:59:48different charts you can make bar charts

4:59:50you can make pie charts so whatever

4:59:52counties won have made in the bar chart

4:59:54so they're here is in a pie chart it

4:59:56looks better but it's not maybe as

4:59:58insightful um I just placed it uh so

5:00:01that I can show you you can make pie

5:00:02chart s also so these are the insights

5:00:05that you can make after analyzing your

5:00:06us County data and this is what you can

5:00:08tell Donald Trump these are the

5:00:10different suggestions that you can

5:00:11actually go and tell Donald Trump uh

5:00:14that she is popular among the foreign

5:00:16people and the people who speak

5:00:18different languages she is popular among

5:00:20the Hispanic people in then in Maine she

5:00:23lost a lot of counties she almost lost

5:00:25all of the counties in Maine so these

5:00:27are different insights that you uh have

5:00:30got and then you can tell your Superior

5:00:33or your employer who has hired you to do

5:00:36that um so this is what you can present

5:00:38right so this is for a very beginner's

5:00:41level and there are some more analytics

5:00:43that you need to do I just showed you a

5:00:45few options you can go ahead and try

5:00:47more in the Democrats section also and

5:00:49you remember that uh you have to do it

5:00:52for the Republican party also now let me

5:00:55see what youve learned today so if you

5:00:57have any questions right now you can

5:00:59just go ahead and ask me so does anyone

5:01:01have any questions

5:01:04so now we will move on and find out the

5:01:05solution for the instant cab use case

5:01:08you remember that we have got the Uber

5:01:10data set which contains the pickup time

5:01:12and the location by two columns latitude

5:01:14and longitude and we have uh also got

5:01:17the license number for a uh particular

5:01:21Uber driver and what we have to do is

5:01:24that we have to find the Beehive

5:01:26locations uh that is the point where we

5:01:29will find the maximum pickups and then

5:01:30we will also have to find out what is

5:01:33the peak hour of the day so this was the

5:01:35entire strategy so we've got the Uber

5:01:38pickup data set and then we store the

5:01:40data into hdfs we will transform the

5:01:43data set and make predictions by using K

5:01:45means clustering on the latitude and

5:01:47longitude and find out the b Point uh or

5:01:50beehive point so now let me open my

5:01:53other notebook the Uber notebook so

5:01:55again the first thing that you have to

5:01:56do is copy the Uber data set into your

5:01:59hdfs now we've done that before

5:02:01explaining to you the US County analysis

5:02:04so again the code is kind of the same

5:02:05the first thing is that again we are

5:02:07importing some spark SQL packages and

5:02:09some spark ml lib packages because we

5:02:12are going to use K means clustering and

5:02:15you can see Vector assembler here again

5:02:16spark ml clustering K means and other

5:02:19spark SQL packages so then we have to

5:02:22start our SQL context and we're doing it

5:02:25same way than the first thing again we

5:02:27have to define a schema now I don't have

5:02:29many fields I've got only four Fields if

5:02:31I remember so the first uh field was the

5:02:34date and time stamp that defines the

5:02:36time we're defining it as DT and next

5:02:39field is the latitude the longitude and

5:02:41base then I'm going to read my data set

5:02:44this is the path in my htfs where my

5:02:46Uber data set is there so I Define

5:02:50schema as schema here the header is true

5:02:53because again my data set contains

5:02:55column headers and I'm going to store in

5:02:57DF so feature calls uh here is going to

5:03:01be latitude longitude because I'm going

5:03:03to find out the Beehive point the point

5:03:07where I will get my maximum kick up from

5:03:09so again I have set the input calls as

5:03:14feature calls and output calls as

5:03:16features so I'm using the assembler to

5:03:19transform my data set and then again I'm

5:03:20using K means and we use the same elbow

5:03:23method we found out that we should make

5:03:24eight clusters for this data set okay

5:03:28and then we are selecting the prediction

5:03:29column and the output column and as

5:03:33predictions and then we have printed the

5:03:35cluster centers for each cluster so

5:03:38definitely whatever result we are going

5:03:40to find the cluster centers will tell me

5:03:42the exact location so this cluster uh

5:03:45centers that we will find after c means

5:03:47is actually the Beehive points this will

5:03:50be the point where I will find maximum

5:03:52pickups

5:03:54right so here I have printed my cluster

5:03:57centers and this defines the latitude

5:03:59and longitude and this is going to be my

5:04:02location where I'm going to find the

5:04:03maximum pickups and I got eight results

5:04:06uh like that because I got eight

5:04:07clusters and uh define the eight centers

5:04:10for different clusters so this is

5:04:13exactly like the K School problem that I

5:04:15explained to you in K means this is

5:04:16exactly what happens just as we found

5:04:18out the center of each cluster and that

5:04:20is where we are replacing the school or

5:04:22building the new school so similarly

5:04:24this is going to be my beehive point and

5:04:27this is where I will place my maximum

5:04:29number of cabs okay so we found out the

5:04:32be Hive points the next thing we will

5:04:34need to do is we need to find the peak

5:04:35hours because I also need to know at

5:04:38what time should I place my cabs in the

5:04:40location so what we're doing now we are

5:04:43taking a new variable called q and we

5:04:45are selecting hour from the timestamp

5:04:47column and then the Alias name should be

5:04:50our and we're getting it from our

5:04:52prediction or from the result that we

5:04:53got after my K means clustering so now

5:04:56we are grouping it and it will have the

5:04:58different hours of the day and then it

5:04:59will just show me the pickups at the

5:05:02different hours of the day in the

5:05:04location that we found out are the

5:05:07Beehive points and then we're going to

5:05:10count uh how many pickups we are going

5:05:13to get from that place right so we're

5:05:15ordering it by descending so the smaller

5:05:17pickup count will be the the first and

5:05:19then the larger will be at the bottom

5:05:21similarly again we are creating new

5:05:23variable called T and we're going to do

5:05:25the same thing so here what we're doing

5:05:28is we are selecting the time hour of the

5:05:30date the latitude longitude prediction

5:05:32and and we filter by hour which is not

5:05:34null so we're filtering out the null

5:05:37values from here so now we have created

5:05:39a table view for categories so let me

5:05:41show you what the categories contain

5:05:44okay let me just go down so I've done

5:05:46some few operations here so let's scroll

5:05:49back up and I'll show you and again we

5:05:51have created table views for T and Q

5:05:54also which is again T and Q all right

5:05:56and then I have made some visualizations

5:05:59for each so then we uh have created a

5:06:03value P where hour is not null so again

5:06:05we have filtered out the null hours and

5:06:07we have created a new view called P so

5:06:11here is my hours this is my count and in

5:06:15the x-axis that show how many pickups

5:06:16were there and this contains different

5:06:18hours of the day and then I have grouped

5:06:20it by prediction so the size is

5:06:23according to the count so you can see

5:06:25that the bigger the circle means more

5:06:27pickups so you can find out the biggest

5:06:30circle and you know that you can find

5:06:31the biggest Circle as you go along the

5:06:33x-axis because this is where the count

5:06:36increases so you can find out the

5:06:38biggest circle would be here and it lies

5:06:40in my fourth cluster and you can see

5:06:43that there are 800 or 8,915 pickups at

5:06:46the 17th hour of the day which is around

5:06:495:00 p.m. and so you know that the

5:06:50maximum pickups are around 4:00 or 5:00

5:06:53and this lies all in my fourth cluster

5:06:57and so it means my peak hours are around

5:06:594 or 5:00 in the evening right so this

5:07:01is what Insight we have gained and you

5:07:04can tell instant cab CEO that I have

5:07:06found out that your cabs should be ready

5:07:08around four or five because that's the

5:07:10time when uh people go home from offices

5:07:14or they're going out for dinner or

5:07:16something and this is what another table

5:07:18view looks like which is T so here we

5:07:21have latitude and longitude and this is

5:07:23where we are finding the Beehive

5:07:25locations so I have uh got this the

5:07:29distribution in a scatter plot again and

5:07:31you can see see that we have got uh very

5:07:33dense points over here it means that

5:07:36these represent the Beehive points so

5:07:39what you can do is that you can just put

5:07:40the US map and scale it according to

5:07:42this scale over here and then you can

5:07:44exactly find out what is the exact

5:07:46location where you need to put your cabs

5:07:49around the 17th hour or the 16th hour of

5:07:52the day all

5:07:54right and you know that we had a lot of

5:07:56rows but the results are only limited by

5:07:5910,000 if it's around 10,000 rows but we

5:08:02obviously had a lot more and you can

5:08:04check in different uh clusters so now we

5:08:08are

5:08:09analyzing

5:08:11uh cluster zero so here if you see this

5:08:15point over here this lies in cluster 4

5:08:20this lies in cluster five and this lies

5:08:22in cluster zero so you can analyze each

5:08:24cluster also so here I have just laid

5:08:28out the latitude and longitude for my uh

5:08:30zeroth cluster so you can see here where

5:08:33prediction is equal to zero and I've

5:08:35selected this from the table view of T

5:08:38so here you can find out the exact

5:08:39latitude and longitude and here the

5:08:42latitude is 4.72 two and the longitude

5:08:45is

5:09:01-73.995411 that tells you what is the

5:09:03count of pickups at each hour of the

5:09:05days starting from 0 to 23 there are 24

5:09:09slices in this circle so you can see uh

5:09:12that these few slices are the bigger

5:09:14chunks and this is the 19th hour of the

5:09:16day which is around 7:00 6:00 5:00 4:00

5:09:203:00 and so on so you can see the

5:09:23midnight maybe nobody travels uh so

5:09:26maybe your cabs could rest or you don't

5:09:29have to place any more cabs during this

5:09:31part of the the day um these are the

5:09:34insights that you gain so any questions

5:09:36on that I think after doing the US

5:09:39County election this was pretty easy to

5:09:40do and this is also uh pretty easy to

5:09:44understand and the results which were

5:09:46also much more clear

5:09:48[Music]

5:09:53correct what is actually a Hadoop

5:09:56ecosystem okay the very first thing is

5:09:59Hadoop ecosystem is not one tool it's

5:10:01not a programming language or it's not a

5:10:03single framework it is a group of tools

5:10:06that are there which are used together

5:10:08by various companies in various domains

5:10:10for different tasks okay Hadoop alone

5:10:14cannot provide all the facilities or

5:10:17services that are required to process

5:10:19the Big Data okay so like for example

5:10:22Hadoop can store Big Data Hadoop can

5:10:24process Big Data up to a certain limit

5:10:26however there are much more other

5:10:28requirements that are there for example

5:10:30we would like to create recommendation

5:10:33engines over big data we would like to

5:10:35run clustering algorithms over big data

5:10:38we would like to get the realtime

5:10:39insights using big data itself because

5:10:41Hadoop is a batch processing framework

5:10:43right so if I want a real time Insight I

5:10:46would need another tool that can run

5:10:48over htfs that can utilize and leverage

5:10:50htfs right the basic thing that you need

5:10:53to understand here is one single tool

5:10:55like Hadoop is not going to solve all

5:10:57your problems you'll have to use various

5:11:00other tools over Hadoop or with Hadoop

5:11:02to get rid or get the solution of every

5:11:05problem that you have okay but before

5:11:09that before doing so it is important

5:11:11that you know what are the different

5:11:12tools that are there which can work with

5:11:14Hadoop and in today's session we'll

5:11:16exactly do that we'll try and find out

5:11:18what are the various tools that are

5:11:20there which can be used with Hadoop and

5:11:22what functions they can perform in their

5:11:25own domains we now move on to the next

5:11:30slide okay the very first tool that

5:11:32we'll understand is

5:11:34hdfs now as you know hdfs is nothing but

5:11:38Hadoop distributed file system it is the

5:11:40storage unit of Hadoop sdfs is entirely

5:11:43the Hadoop cluster which is formed by

5:11:45data nodes data nodes are nothing but

5:11:47commodity Hardwares which are cheap

5:11:49Hardwares which can be clustered

5:11:52together using the Hardo framework and

5:11:54then entire file system that gets

5:11:56created on which you can store big data

5:11:58is called Hadoop distributed file system

5:12:00using hdfs you can store any kind of

5:12:03data be it structured be it unstructured

5:12:05or be it

5:12:06semi-structured okay now once you store

5:12:09the data in hdfs you can view the entire

5:12:12data as a single unit as well hdfs

5:12:15stores data across various nodes that

5:12:17these nodes are nothing but the data

5:12:19nodes and it also maintains the log

5:12:21files of what data is stored at which

5:12:23position so basically hdfs has got two

5:12:26components one is the name node and

5:12:28other is the data node data name node is

5:12:30the one which manages the entire cluster

5:12:33which manages the entire set of data

5:12:35nodes and keeps the information keeps

5:12:37the metadata of the data that is stored

5:12:39in these data nodes data nodes on the

5:12:42other hand are the slave machines the

5:12:44commodity Hardwares which actually

5:12:45stores the data so sdfs is the one which

5:12:49solves the primary problem of storing

5:12:51big data so it's time we move on and

5:12:54explore the next tool in Ado ecosystem

5:12:57which

5:12:58is Yan now we'll explore

5:13:02Yan now as the name suggest it is

5:13:05nothing but a resource

5:13:07negotiator okay the main purpose of yan

5:13:10is to allocate resources to run

5:13:12particular task over the Hadoop cluster

5:13:15so Yan has basically two components one

5:13:17is the resource manager and the other

5:13:19one is the node

5:13:20manager as soon as a client submits a

5:13:23job these resources are nothing but the

5:13:25containers in which the jobs can be

5:13:27executed okay node manager is the one

5:13:30which finally executes the a job within

5:13:32these containers and manages the entire

5:13:35thing on the data nodes okay the

5:13:38resource manager is the master demon and

5:13:40the node manager is the slave

5:13:42demon apart from that I would like to

5:13:44tell you one important thing about yan

5:13:46yan was introduced in Hadoop 2.0 which

5:13:49enabled various ecosystem tools to

5:13:51connect with Hadoop distributed file

5:13:53system and leverage Big Data okay so

5:13:56we'll understand this better when we

5:13:59come to the map reduced slide okay okay

5:14:02we'll move on to the next

5:14:05slide so we now come to map ruce which

5:14:07is the processing unit of Hado so once

5:14:10the data is stored on Hado distributed

5:14:12file system the next task is to process

5:14:15that data for doing so one can use a

5:14:17Pache map produce in Hado 1.x map

5:14:20produce was the only framework that can

5:14:22be used to process the distributed data

5:14:24that is present on

5:14:26hdfs however soan was the next layer

5:14:28over htfs and map ruce now connected

5:14:31with Yan to allocate resources for

5:14:34executing the map reduce task similarly

5:14:36many other ecosystem tools or databases

5:14:39now can connect with Yan and leverage

5:14:41hdfs okay so it happened after Yan so

5:14:45essentially it has got two functions one

5:14:47is the map and the reduce map function

5:14:49is used for filtering grouping and

5:14:50sorting kind of functions and the result

5:14:52of the map function is then aggregated

5:14:54in the reduce phase and the entire

5:14:56summarized result is dumped on the hdfs

5:14:59itself so this is how Hadoop map

5:15:01produced works now it's time we move on

5:15:03to the next

5:15:06slide and we'll explore what is Apache

5:15:11Pig Apache Pig was a tool that was

5:15:13developed at Kahu so it is nothing but a

5:15:16data processing tool that runs over

5:15:18Hadoop or you can say that it sits on

5:15:20top of the Hadoop Apache Pig has its own

5:15:22language that is called Pig Latin which

5:15:24is nothing but a dataflow language or

5:15:26you can call it as instructional

5:15:28language for example if you want to load

5:15:30a data you have a command like like load

5:15:32this data from path and then you can

5:15:34dump that data or perform various

5:15:36functions like filter or grub okay so

5:15:39using Pig Latin the life of the

5:15:41developers became very easy they need

5:15:43not write the entire full map Produce

5:15:45job for executing some processing over

5:15:48the big data for people who cannot write

5:15:52a map produce program or or were not

5:15:54comfortable with map ruce for them pck

5:15:57Latin came as a blessing or for them P

5:16:01came as a Blessing by using p the task

5:16:04for them became very easy and they were

5:16:06able to leverage Big Data it is said

5:16:08that approximately one line of pig latin

5:16:11is equals to 100 lines of map produced

5:16:13so just think how much time are you

5:16:15saving there okay you can perform all

5:16:18the ETL operations that you would like

5:16:20to execute over big data using Peg so I

5:16:23hope this gives you a clear picture how

5:16:26Apache Peg came into picture what is the

5:16:28importance of Apache Peg okay then we'll

5:16:32move on and explore the next tool in the

5:16:35chain that is Apache

5:16:38Hive Apache Hive is one of the most

5:16:41important tool that is there in the

5:16:42Hadoop ecosystem Apache Hive was

5:16:45developed at Facebook now the idea

5:16:47behind aache hiive was the time when it

5:16:50was developed the relational databases

5:16:52were flourishing these were the

5:16:54databases that were used by most of the

5:16:56organizations most of the companies

5:16:58nobody knew about no SQL or system like

5:17:02Hado even Facebook had its website on

5:17:04MySQL so the workforce there mostly was

5:17:08working on myql or SQL like queries or

5:17:11PL SQL

5:17:13plsql so for Facebook it was a problem

5:17:15because the workforce were skilled in

5:17:18SQL however for writing a map produced

5:17:21program you had to know some other

5:17:23programming language so what Facebook

5:17:26did is uh Facebook came up with tool

5:17:28that is called Hive using which you can

5:17:30write SQL like queries that is called

5:17:32high query language and execute the same

5:17:37task over the Hadoop cluster and

5:17:38leverage Big Data just like Pig using

5:17:41Hive you can write simple SQL like

5:17:43queries and the task that you were

5:17:45executing using map produce now can be

5:17:47executing using Hive without getting

5:17:49into the complexities of map ruce okay

5:17:52even using Hive you can connect from

5:17:54client applications like Java as well if

5:17:56you have that

5:17:57requirement okay so Hive is one

5:18:00important to tool that is used by a lot

5:18:03of people out there who do not want to

5:18:05get into writing the map produce program

5:18:08so let's move on to the next tool that

5:18:12is mahot and Spark mlip mahot is a

5:18:16machine learning library written in Java

5:18:19it can be used for creating

5:18:21recommendation engines or uh clusters of

5:18:24data or classify your data into various

5:18:26grps okay so all those algorithms that

5:18:29are there in machine learning can be

5:18:30implemented over big data using mahot

5:18:34okay so it provides you a command line

5:18:36interface to achieve the same task you

5:18:38would have heard about an analysis that

5:18:41is called Market Basket analysis which

5:18:43can be easily executed using Maho over

5:18:46big data the various other things like

5:18:48recommendation engine as I mentioned

5:18:50which you would have seen in many

5:18:52e-commerce websites like Amazon flip

5:18:54cart or many more okay so all those

5:18:57things can be done using

5:19:00Mah now now we come on to spark which is

5:19:03a leading tool in the Hadoop ecosystem

5:19:06map ruce and Hadoop together can only be

5:19:08used for batch processing that means

5:19:10you're not getting the results in real

5:19:12time but out there there is a

5:19:15requirement for realtime analytics as

5:19:17well which cannot be done using Hadoop

5:19:19map reduce right in that case Sparks

5:19:21come into the picture which can run

5:19:24Standalone as well as it can run over

5:19:26the Hado cluster and leverage the same

5:19:28big data to provide you realtime

5:19:30insights right as well as Apache spark

5:19:34is almost 100 times faster than Apache

5:19:37map ruce okay so I hope this excites you

5:19:41right so we'll move on to the next

5:19:45slide and explore Apache Edge

5:19:48Bas so what is Apache hedge Bas Apache

5:19:51Edge Bas is a nosql database that runs

5:19:54over aoop Apache hpas can be used for

5:19:57storing any kind of data that is there

5:19:59okay it could be any structured or

5:20:01unstructured data and P Edge Bas has

5:20:05been modeled after Google big table and

5:20:08can be utilize to store any big data

5:20:10that is there in Hadoop file system with

5:20:13Apache hpas you have an advantage that

5:20:16is you can use edpas as a backend for a

5:20:19website or web application to query in

5:20:21real time which cannot be done with

5:20:23tools like pck Hive or map ruce or even

5:20:27hdfs and hence it is a very important

5:20:30addition to hadopi

5:20:31ecosystem okay so guys are you clear

5:20:34with this you can also write a Java

5:20:36application and connect with Edge base

5:20:38using the rest apis Thrift apis or AO

5:20:42apis

5:20:44okay now we'll go through another tool

5:20:47that is called Apache drill Apache drill

5:20:50is again an open source application

5:20:52which works well with any distributed

5:20:54environment that is out there it can

5:20:56work with any nosql database or a flat

5:20:58file system okay the advantage with

5:21:01Apache drill is that it can connect with

5:21:03various nosql databases or a flat file

5:21:07system or a simple file itself at the

5:21:10same time so if you have data stored in

5:21:12various sources like let's say you have

5:21:15a data stored in Hadoop distributed file

5:21:17system you have a data stored in Edge

5:21:18Bas you have a data stored in mongodb

5:21:21every one of them has their own syntax

5:21:23to execute queries on them to retrieve

5:21:26the same set of Records however using a

5:21:29pacher drill you can connect to all

5:21:31these databases at a single time execute

5:21:34one query and extract the results from

5:21:36all the three databases and use it for

5:21:39your application okay Apache drill is

5:21:42able to do that because it follows the

5:21:44an SQL which enables you to write a

5:21:47query that can execute or that can be

5:21:49understood by all the three

5:21:52databases we'll move on to the next

5:21:55slide and now we'll explore

5:21:58Uzi Apache Uzi is nothing but a

5:22:01scheduler in the Hadoop ecosystem now

5:22:04what does it mean let's say you have a

5:22:06map ruce task that needs to be executed

5:22:08every hour now in that case instead of

5:22:10manually triggering it what you can do

5:22:12is you can define a workflow in Uzi and

5:22:15schedule your task to be executed after

5:22:18every 1 hour okay when you see you're

5:22:21doing two things one is you're defining

5:22:23a workflow that could be one task or it

5:22:26could be a combination of tasks that are

5:22:28executed by various tools like map

5:22:30produce hi Pig Etc in a sequence as well

5:22:33as you're defining the frequency in

5:22:35which the workflow needs to be executed

5:22:38okay so life becomes very easy you need

5:22:40not go and execute or trigger your job

5:22:43every time that need it needs to be done

5:22:46Uzi can do it for you along with that

5:22:48Uzi coordinator is another component

5:22:50that is present in Uzi which ensures

5:22:53that the job or the workflow is only

5:22:55executed when the data is available so

5:22:57at times if the data is coming from an

5:22:59external Source automat automatically it

5:23:01will ensure that as soon as the data is

5:23:04in the system then only the workflow of

5:23:06the job is executed so it is an event

5:23:08based execution that can be triggered

5:23:11using Uzi

5:23:13okay let's move on and we come to

5:23:17Flume Flume is again one of the most

5:23:20widely used tools which is used for data

5:23:22ingestion into hdfs okay so using Flume

5:23:26you can ingest any kind of data it could

5:23:28be structured it could be

5:23:29semi-structured into the the Hado

5:23:31distributed file system and perform

5:23:33various processing after that floom

5:23:36gives you the capability of extracting

5:23:38data out of social media like Twitter

5:23:40Facebook or you can also extract data

5:23:43from servers where logs are getting

5:23:45generated on a regular interval so Flume

5:23:47can be utilized to extract data from

5:23:50there and move into the hdfs okay

5:23:53similarly there could be many other use

5:23:55cases like getting email messages or

5:23:56network traffic

5:23:58Etc the next tool in the as scoop scoop

5:24:02is again used for data inje however

5:24:05scoop is used between relational

5:24:07database and the hdfs so using scope you

5:24:11can move your data from your relational

5:24:12database into hdfs and vice versa that

5:24:15means you can also move data out of hdfs

5:24:18into an rdbms so it mostly deals with

5:24:21structured

5:24:22data okay so if you compare Flume with

5:24:25scoop Flume is mostly used for moving

5:24:28data into the hdfs and deals with

5:24:31streaming data most of the time however

5:24:33scoop works with structured data and it

5:24:36can move data in and out of hdfs unlike

5:24:39flu so let's move on and we come to the

5:24:43next tool that is solar and

5:24:45Lucine okay so solar Lucine is again an

5:24:49Apache project which has been developed

5:24:51in Java Lucine in itself is a Java

5:24:53Library which is for developing search

5:24:56engine and indexers okay Apache solar is

5:25:00an application that is built using aacha

5:25:03loine so if you want to develop a search

5:25:05engine or you want to implement search

5:25:07onto your website which works very fast

5:25:10using indexing you can always use Apache

5:25:13solar to do that okay so this is the

5:25:15main purpose of Apache solar which is an

5:25:18application which is developed using

5:25:20Apache

5:25:21Lucin we come to

5:25:24zookeeper as the name suggests the job

5:25:27of Zookeeper is to ensure coordination

5:25:29between various tools that are there in

5:25:31the Hadoop ecosystem okay so the main

5:25:34purpose of Zookeeper is to ensure that

5:25:37each and every tool is able to

5:25:38communicate with each other without any

5:25:40Interruption so that the entire

5:25:42ecosystems works together in achieving a

5:25:45particular task okay it performs

5:25:47synchronization it performs

5:25:49configuration management grouping and

5:25:51naming of all these things okay it also

5:25:54manages all the services that are

5:25:56running in the Hadoop cluster okay so

5:25:59zookeeper is very important component of

5:26:01Hado cluster if zookeeper is not there

5:26:04your services your demons your tools

5:26:07will not be able to interact with each

5:26:09other or communicate with each other and

5:26:12hence you'll get a broken system if

5:26:14zookeeper

5:26:15fails now we come on to the final tool

5:26:18that is a Apache Amari in the Hadoop

5:26:21ecosystem that we are going to discuss

5:26:23today so apach Amar is a cluster manager

5:26:27okay what does a cluster manager mean or

5:26:31what what does a cluster manager do

5:26:33cluster manager manages the Hadoop

5:26:35cluster okay using Apachi mbari you can

5:26:38provision manage and monitor their P

5:26:40Hadoop clusters okay it makes very easy

5:26:43for you to set up a Hadoop cluster and

5:26:46then configure all the services that

5:26:48needs to run over the Hado cluster it

5:26:50could be a PES spark service it could be

5:26:52a hue service it could be any other

5:26:54service that you need over the cluster

5:26:56and it can be done very easily using

5:26:58Apache ambari so this was basically

5:27:01developed by hoton works a similar tool

5:27:04is developed by claura as well which is

5:27:06called claura manager okay however Cloud

5:27:10manager is not an open-source tool like

5:27:13Apache ambari so Apache ambari was

5:27:15developed by Harden Works however it was

5:27:17given to Apache later on a similar tool

5:27:20which is a propriety tool developed by

5:27:22Cloud named as Cloud manager using which

5:27:25you can deploy the Cloudera clusters

5:27:27however it is a paid service okay using

5:27:30using aache Amari you can also monitor

5:27:32health and status of your Hado cluster

5:27:34okay now you know the importance of

5:27:36aache Amari and how it can be used to

5:27:39make your life

5:27:40[Music]

5:27:45easy the prerequisites to install Hadoop

5:27:48in Windows operating system are Java so

5:27:52we all know that Hadoop supports only

5:27:54Java version 8 so firstly we need to

5:27:57download Java 8 version followed by that

5:28:00a latest Hardo version which we need for

5:28:02our operating system then the

5:28:04configuration files so these were the

5:28:07prerequisites now let's quickly go ahead

5:28:09and download Java 8 version into our

5:28:11local system and also hadu so you can

5:28:14see that this particular web page

5:28:16belongs to Oracle and here you'll be

5:28:19getting your Java development kit number

5:28:21eight so these are the various versions

5:28:23available for Java 8 for Linux as well

5:28:26as Windows so we need a jdk which is

5:28:29compatible with Windows

5:28:31so here you can see that Windows x64 jdk

5:28:35version which will support Windows so

5:28:37this particular link will redirect you

5:28:39and download jdk8 for you into your

5:28:42local system once you click on it it

5:28:44will ask you to accept the license terms

5:28:47from Oracle now you can just click on

5:28:50download followed by this you will be

5:28:52redirected into a login page where you

5:28:54need to create your own account with

5:28:57Oracle so that you can download this jdk

5:29:00don't worry this account is free of cost

5:29:03so you can see the jdk is getting

5:29:04downloaded here so as the jdk is getting

5:29:08downloaded we shall now move ahead and

5:29:10download Hardo for our local system so

5:29:12this particular web page belongs to

5:29:14Apache organization where we can

5:29:16download Hadoop for free so these are

5:29:18the various versions available for

5:29:20Hadoop which are 2.10 3.1.3 3.2.1 and

5:29:24many more so we shall select the latest

5:29:27version of Hardo but while you selecting

5:29:30the latest version of Hardo please make

5:29:32sure that you're not actually

5:29:33downloading the exact latest version of

5:29:35Hardo here you can see we have three

5:29:38different versions 3.1.3 3.2.1 3.1.2 as

5:29:43you can see 3.2.1 is the latest version

5:29:47we have to select the version which is

5:29:49earlier to it which is

5:29:513.1.3 because this particular version

5:29:54will be the stable version now we shall

5:29:56move ahead and select binary once you

5:29:59select binary

5:30:00you will be redirected into a new web

5:30:02page where you will have a mirror link

5:30:05select that mirror link and your Hardo

5:30:06will be downloaded for your local system

5:30:09as you can see Hardo 3.1.3 tar.gz is

5:30:12getting

5:30:13downloaded now here you can see I have

5:30:15successfully downloaded her version

5:30:183.1.3 T file as well as jtk 8 and those

5:30:22two files have successfully moved into

5:30:24my C drive now let's install Java first

5:30:29now make sure you that you create a new

5:30:31folder for Java so select change and

5:30:35here select Windows C drive then select

5:30:38make new folder now rename this new

5:30:41folder as Java click okay and now select

5:30:45next you can see the installation

5:30:48procedure has now been

5:30:51started you can see Java development kit

5:30:548 has been successfully installed now we

5:30:57shall enter into program files and move

5:31:00jdk into Java file because sometimes

5:31:03there will be an error while we set

5:31:06environment variables for Java so you

5:31:09can see inside program files we have

5:31:11another folder called Java so inside

5:31:13Java there you have our jdk so now what

5:31:17I'll be doing is just moving this jdk

5:31:19into Java file which we have created in

5:31:22C drive this

5:31:25one now you can just delete this Java

5:31:28file from your program files so that you

5:31:30don't have to mess with duplication of

5:31:32java file now you have your Java and jdk

5:31:36in one single file which is Java that is

5:31:39you have created in Windows C drive now

5:31:41we shall move ahead and set the

5:31:43environment variables for Java so click

5:31:46windows and then enter into settings and

5:31:49inside the settings select system and

5:31:51inside system just type in environment

5:31:54variables and there you go select the

5:31:56edit the system environment variables

5:31:58option and you have this dialogue box

5:32:01here select environment variables and

5:32:03inside the environment variables you

5:32:05need to set the Java home as well as

5:32:07path for Java now select new and here

5:32:12just type in Java home and here let us

5:32:16add the location of jdk bin so here we

5:32:19will add in the variable value that is

5:32:21the jdk bin location so our jdk bin

5:32:25location is in the C drive and inside

5:32:27the C drive we have the Java folder and

5:32:29inside jav Java folder we have a jdk

5:32:321.8.0 and inside jdk file we have the

5:32:35bin location so this will be the home

5:32:38location for Java select okay and then

5:32:42now move into the next dialogue box

5:32:45which is the system variables and inside

5:32:47that select path and select edit here

5:32:50Create A New Path variable which will be

5:32:53the jdk path the same location that is

5:32:55the bin of jdk Select okay and now

5:32:59select okay again and now okay and close

5:33:03it now Java has been successfully

5:33:06installed into our local system now

5:33:08let's check Java is functional or not we

5:33:10can do that by selecting Windows R and

5:33:13inside windows R just type in CMD so

5:33:16that you can open your command prompt

5:33:18here just type in Java C if you see the

5:33:22set of files popping up into your

5:33:23terminal then it means that Java is

5:33:26working properly so you can see Java is

5:33:28working just fine now let us check the

5:33:30version of our Java installed into our

5:33:32local system so this can be checked by

5:33:34typing in Java space hyphen version so

5:33:38you can see we have 1.8 version which is

5:33:41running in our local system now that we

5:33:43have successfully installed Java into

5:33:45our local system let us now move ahead

5:33:47and install Hado into our local system

5:33:50you can see that we have downloaded the

5:33:51tab version of Hado so for that we need

5:33:54to extract it

5:33:57first now you can see that the process

5:33:59of extraction has been completely

5:34:01finished that is 100% but you have three

5:34:03errors you can ignore these errors now

5:34:06just close the extracting process then

5:34:09you have your Hadoop

5:34:12file now let us rename a Hadoop

5:34:153.1.3 as just Hadoop to reduce the

5:34:18confusion now that we have successfully

5:34:20extracted Hadoop let's set environment

5:34:22variables for Hadoop but before that

5:34:25let's set the configuration of Hadoop

5:34:28you can select Hadoop and inside that

5:34:30you have a file called Etc and inside

5:34:32Etc you have another folder with the

5:34:35name hadu and inside that you have a set

5:34:38of folders so out of these all folders

5:34:41We have four important folders they are

5:34:45core site. XML then htfs site. XML

5:34:50followed by htfs we have another one

5:34:52which is map site. XML and lastly the

5:34:56Yan site. XML file so we need to edit

5:34:59all these four different files and once

5:35:02after we edit these four files we need

5:35:04to edit one last file which is the

5:35:07Hadoop EnV Windows command PR file so

5:35:11here you're just going to add in the

5:35:13Java home location now let's quickly

5:35:16edit all those four files so we have

5:35:19successfully opened our four important

5:35:21files which are cor site. XML map reduce

5:35:24site. XML Yan site. XML htfs site. XML

5:35:29followed by by the four important files

5:35:31the last file which is the Hado

5:35:33environment. CMD file here we are going

5:35:36to set this Java home location now let's

5:35:38first set the values for corite XML so

5:35:42the values that are changed in corite

5:35:44XML are the properties so inside the

5:35:47configuration I have added one property

5:35:49which is the file location that is fs.

5:35:52default file system and the Local Host

5:35:54location that is 9,000 now let us save

5:35:57this course site. XML similarly we need

5:36:00to also edit map reduce side. XML files

5:36:03here inside this we need to add some

5:36:05properties as you can see we have also

5:36:08edited the configuration files of map

5:36:10redu site. XML let's save it now

5:36:13followed by the map redu site. XML we

5:36:15have Yan site. XML let's edit this

5:36:19also as you can see the Yan site. XML is

5:36:22also been updated no worry about this

5:36:25property file I will link this in the

5:36:27description box below you can have the

5:36:29access to it and you can use the same

5:36:30configuration file and install her tube

5:36:32followed by Yan site. XML we have the

5:36:35last one which is htfs site. XML but

5:36:38before editing this particular file I

5:36:41want you to create a new folder in hero

5:36:43location which is data let's see how to

5:36:46create it so this particular folder is

5:36:49inside C drive this is Hadoop and inside

5:36:53this Hadoop file you need to create a

5:36:54new folder with the name data inside

5:36:58data you need to create two more new

5:37:00files which are data node and name node

5:37:03so the first folder will be name node

5:37:06and now another folder which will be our

5:37:09data

5:37:10node so now let's copy the location of

5:37:13data node and name node so this location

5:37:16is the data node location and followed

5:37:19by that the name node location so this

5:37:21particular location will be the name

5:37:22node location we have the two locations

5:37:24copied onto our clipboard now let's go

5:37:27back to the htfs site dox SML file and

5:37:30edit the configurations

5:37:32here so you can see that we have edited

5:37:35the configuration file of htfs site. XML

5:37:38and inside the configuration we have

5:37:40provided the replication factor which is

5:37:42the first property and we have set the

5:37:44value as one since we're using our local

5:37:46system we might want to save memory so

5:37:48the replication is only one but the

5:37:50default value for the Hardo replication

5:37:52factor is three and followed by the

5:37:54first property the second property which

5:37:56is our name node so we have provided our

5:37:59name node a which is Hardo file and

5:38:01followed by that the data file and

5:38:03inside that we have the name node and

5:38:05similarly the last property which is the

5:38:08data node property so here the value is

5:38:11Hadoop data data node now let's save

5:38:16it now that we have successfully edited

5:38:18all our four important files let's get

5:38:21back to Hadoop env. CMD file and edit

5:38:24the Java home location so for safest

5:38:27side let's get back to environment

5:38:28variables and and get our jdk

5:38:33location so this particular location is

5:38:36the location for Java home we might want

5:38:38to remove the bin over here so only C

5:38:41Java jdk is enough to set the Java home

5:38:44into our Hardo env. CMD file now let's

5:38:48save this particular file and close it

5:38:50so all the important files have been now

5:38:52successfully edited now let's go back to

5:38:54environment variables and set home and

5:38:56path for

5:38:57hadu now select new and write in Hadoop

5:39:04home so this particular location that is

5:39:07C Hadoop Ben is the location for Hadoop

5:39:10home now select okay now let's get back

5:39:13to path and set path for Hadoop files in

5:39:17here let's set up a new path variable

5:39:20that is Hadoop bin and now remember to

5:39:23create another path variable that is

5:39:25your spin so to locate Espin get back

5:39:29back into Hadoop and select Spin and

5:39:32this will be the location or path value

5:39:34for your spin select that and edit a new

5:39:37variable in path and paste it so that's

5:39:40how you set sben and select okay okay

5:39:43and finally another okay and close the

5:39:45system properties now that we have

5:39:47successfully set home and path for

5:39:49Hadoop let's go ahead and fix the

5:39:52configuration files you can see that

5:39:54inside the bin folder of Hardo we are

5:39:57missing some important configuration

5:39:59file to fix this we need a new

5:40:01configuration file which will be

5:40:03available in the description box below

5:40:05you can click on that particular link

5:40:07and the required configuration file will

5:40:09be downloaded into your local system and

5:40:11all you need to do is just replace that

5:40:13particular file with your bin folder in

5:40:15your Hadoop you can see that there is a

5:40:17new file in my Hadoop which is Hadoop

5:40:19configuration Fixx bin 1.rar now all you

5:40:23need to do is just extract this

5:40:25particular

5:40:28folder

5:40:34you can see that the folder has been

5:40:36successfully extracted and all the

5:40:38executable files that you require in

5:40:40your Hadoop have been downloaded

5:40:41successfully now what you need to do is

5:40:44just move this bin into your Hadoop bin

5:40:48so cut this bin and get back to Hadoop

5:40:51and enter bin so just delete this

5:40:54particular bin and replace it with a new

5:40:56one so there you go you have

5:40:58successfully done on it now let's delete

5:41:00the unnecessary files there you go as

5:41:03good as new so you have all the

5:41:05executable files and your hop is been

5:41:07set to check if hero is functioning

5:41:10properly or not let's open CMD and type

5:41:13in

5:41:15hdfs space name node space hyphen

5:41:20format if you see a set of files popping

5:41:23up on your terminal that means you have

5:41:25successfully installed heru you can see

5:41:27that the name note has been successfully

5:41:28getting started

5:41:29now let's open a new terminal and start

5:41:32all the Hadoop demons here you just need

5:41:34to enter your Hadoop location file that

5:41:38is CD space

5:41:41Hadoop now you are inside Hadoop and

5:41:44inside Hadoop enter

5:41:47sbin now you're inside sbin now you need

5:41:50to type in start all. shr start all.

5:41:57CMD and there you go all your demons are

5:42:00getting started so that's how you

5:42:03install hardup into your local Windows

5:42:05operating system with the version

5:42:06Windows

5:42:08[Music]

5:42:1210 so what were the problems associated

5:42:15with the relational database system as I

5:42:17have already mentioned that for a Hado

5:42:19developer actual game starts after the

5:42:21data is being loaded in hdfs and the

5:42:24developers play around this data in

5:42:25order to gain various insights that are

5:42:27hidden in the data stored in H hdfs so

5:42:30for this analysis the data residing in

5:42:32the rdbms needs to be transferred to

5:42:35hdfs and you all know that the task of

5:42:37writing map reduce scod for importing

5:42:39and exporting the data from relational

5:42:41database to hdfs is TDS so this is where

5:42:44Apache scoop comes to rescue and remove

5:42:46the pain of data inje so why do we need

5:42:49scoop it's a known fact that before Big

5:42:51Data came into existence the entire data

5:42:54was stored in relational database

5:42:55servers in the relational database

5:42:57structure but with the advancement of

5:42:59scoop it makes the life of developers

5:43:01Easier by providing CLI for importing

5:43:03and exporting the data and scoop

5:43:05internally converts a command into map

5:43:07ruce task which are then executed over

5:43:10hdfs it uses yarn framework to Import

5:43:13and Export the data which provides fall

5:43:15Tolerance on top of parallelism it also

5:43:17uses yarn framework to Import and Export

5:43:20the data which provides fall Tolerance

5:43:22on top of parallelism not only that it

5:43:24is also very useful for data analysis

5:43:27high in performance and provides command

5:43:29line interface now let's understand what

5:43:31a scoop so before I tell you what a

5:43:34scoop let me tell you how the name scoop

5:43:36came into

5:43:37existence by now I hope that you have

5:43:40got an idea that it is used for data

5:43:42transfer between relational database and

5:43:44hdfs so before I tell you what is scoop

5:43:47let me first tell you how the name came

5:43:48into existence the first two letters in

5:43:51scoop stands for the first two letters

5:43:53in SQL and the last three letters in

5:43:55scope refers to the last three letters

5:43:57in Hardo that is o op so it clearly

5:44:00depicts that it is SQL to Hadoop and

5:44:02Hadoop to SQL that is how the name of

5:44:05scop came into existence so what is scop

5:44:08it is a tool used for data transfer

5:44:10between rdbms like MySQL Oracle SQL Etc

5:44:14and Hadoop like Hive hdfs hbas Etc it is

5:44:18used to import the data from rdbms to

5:44:20Hadoop and Export the data from Hadoop

5:44:22to rdbms simple again scoop is one of

5:44:26the top projects by Apache software

5:44:28Foundation and works brilliantly with

5:44:30relational databases such as Terra dat

5:44:32nza Oracle MySQL Etc it also uses map

5:44:36reduce mechanism for its operations like

5:44:38Import and Export work and work on a

5:44:40parel mechanism as well as fall

5:44:42tolerance as I have already mentioned

5:44:44that it provides command line interface

5:44:46for importing and exporting the data the

5:44:48developers just have to provide the

5:44:50basic information like database

5:44:52authentication Source destination

5:44:54operations Etc and the rest of the work

5:44:56will be done by scoop tool itself sounds

5:44:59much reliable correct now let's move

5:45:02further and talk about some of the

5:45:04amazing features of sco for Big Data

5:45:06developers First full load Apache scope

5:45:09can load whole table by a single command

5:45:12you can also load all the tables from a

5:45:14database using a single command next

5:45:17incremental load scool provides a

5:45:19facility of incremental load where you

5:45:21can load the parts of a table wherever

5:45:23it is updated next parallel Import and

5:45:25Export again as already mentioned scoop

5:45:28user this Yan framework to Import and

5:45:30Export the data which provides fall

5:45:32Tolerance on top of parallelism next

5:45:35compression you can compress your data

5:45:37by using Gip algorithm with compress

5:45:39argument or by specifying compression

5:45:41codec argument next karos security

5:45:44integration so what is karos it's a

5:45:47computer network Authentication Protocol

5:45:50which works on the basis of tickets to

5:45:52allow the nodes that are communicating

5:45:53over a non-secure network to prove their

5:45:56identity to one another in a secure

5:45:59manner next load data directly into Hive

5:46:02and hbas here it is very simple you can

5:46:05load the data directly to Apache high

5:46:07for analysis and you can also dump your

5:46:09data in hbase which is a no SQL database

5:46:13now let's see what's next the

5:46:15architecture is one of the empowering

5:46:17Apache scope with its benefits now as we

5:46:20know the features of Apache scope let's

5:46:22move ahead and try to understand Apache

5:46:24scope's architecture and its working so

5:46:27when we submit our job or a command

5:46:28through through scope it is mapped into

5:46:30map task which brings a chunks of data

5:46:32from hdfs and these chunks are exported

5:46:35to a structured data destination and

5:46:38combining all these exported chunks of

5:46:40data we receive the whole data at the

5:46:42destination which in most of the cases

5:46:44is rdbms server next reduce phase is

5:46:47required in case of aggregations but

5:46:50Apache scope just Imports and Export the

5:46:52data it does not perform any

5:46:54aggregations map job launch multiple

5:46:56mappers depending on the number defined

5:46:58by the user for scoop import each maper

5:47:01task will be assigned with the part of

5:47:03data that is to be imported and scoop

5:47:05distributes the input data among all the

5:47:07mappers equally in order to achieve high

5:47:10performance then each mapper creates a

5:47:13connection with the database using gdbc

5:47:15and fetches the part of the data

5:47:17assigned by the scope and then writes

5:47:19that data to hdfs Hive or hbas based on

5:47:22the arguments provided in the command

5:47:24line interface so this is how scope

5:47:26Import and Export works like the gather

5:47:29the metadata again it submits only map

5:47:31job the reduced phase will never occur

5:47:33here and then it stores the data in hdfs

5:47:36storage coming to scoop export it's the

5:47:39same thing the data will be reversed

5:47:40back to

5:47:42rdbms so here the scoop import tool will

5:47:45import each table of the rdbms in Hardo

5:47:48and each row of the table will be

5:47:49considered as a record in the hdfs and

5:47:52all the records are stored as Text data

5:47:54in the text files or binary data in

5:47:56sequence files on the other hand the

5:47:58scoop export tool will export the Hardo

5:48:01files back to the rdbms tables again the

5:48:03records in the hdfs files will be the

5:48:05rows of a table and those are read and

5:48:08passed into a set of records and

5:48:09delimited with the user specified

5:48:11delimiter so this is all about the scoop

5:48:14architecture and its Import and Export

5:48:16now let's execute some scoop commands

5:48:18and understand how it works so at the

5:48:21first we have scoop import that is it

5:48:23Imports the data from rdbms in hadu the

5:48:26command goes very simple here you have

5:48:28to just provide the connection for MySQL

5:48:31your IP address your database name your

5:48:33table name the username for MySQL user

5:48:36if you have set privileges for password

5:48:38you can specify the password or it is

5:48:40not required and the target directory

5:48:42now let's see how to

5:48:44execute so I'll open my terminal and

5:48:47check whether all my Hardo demons are up

5:48:49and running or

5:48:51not so I can see that all my Hardo

5:48:54demons are up and running so now let's

5:48:57execute scoop help and see whether the

5:48:59scoop has been properly installed or

5:49:01not so these are the available commands

5:49:04in scoop here I will show you the

5:49:06execution and explain you few of these

5:49:08commands now this is also properly being

5:49:11set now I will open another terminal and

5:49:13connect to my SQL database the command

5:49:16goes like

5:49:17this my user is edura so I'm giving it

5:49:20as edureka you can give it as root if

5:49:22your user is root simple so now I'm into

5:49:25my SQL database if you want to create a

5:49:28data database you can create the

5:49:30database by giving this command create

5:49:32database database name Etc as I have

5:49:35already created a database so I'll just

5:49:37specify show databases command to list

5:49:39the database present in the mySQL

5:49:41database so now these are the list of

5:49:44database present here so now I want to

5:49:46use employees database so I'll give use

5:49:49employees database got changed now I

5:49:52want to list the tables present in the

5:49:54employees database so I'll give show

5:49:57tables so these are the 11 tables

5:50:00present in the database employees now

5:50:03let's say I want to use employees table

5:50:05so what will I do I'll just give select

5:50:08star from employees that is table

5:50:11name so it's just a huge amount of data

5:50:14that is present in this

5:50:17database so there are these many rows

5:50:20present in this table now open the other

5:50:23terminal where you have executed this

5:50:25command and here I will show you how to

5:50:27import the data present in in that table

5:50:29to hdfs so how we are going to do that

5:50:31by using scoop import command the

5:50:33command goes like

5:50:36this the IP here is Local Host and you

5:50:39know that employees is my database name

5:50:42and the username will be

5:50:44adura and the table that I have chosen

5:50:47is employees so I have not set any

5:50:49privileges for password so I'm not

5:50:50specifying the password over

5:50:55here so it got executed and you can see

5:50:59the number of job counters the map

5:51:00reduce framework the input records

5:51:02output records Etc so what happens after

5:51:05executing this command the map task will

5:51:07be executed at the back end now let's

5:51:10check the webui of hdfs that is the

5:51:13webui port for hdfs is Local Host

5:51:17570 and let's see where the data got

5:51:20imported one important thing to note I

5:51:23have not specified Target directory so

5:51:25by default the data will be imported to

5:51:28this folder user edure Rea and employees

5:51:32so here you can see the four different

5:51:34part files where our data got imported

5:51:36so you might be thinking y4 correct here

5:51:39I have not specified the number of

5:51:41mappers so by default it takes the

5:51:43number of mappers to three and then

5:51:45gives the output in four different part

5:51:47files so let's open the part file and

5:51:49see how the output will be so here you

5:51:52can see the data is imported from rdbms

5:51:55to hdfs there's lots of data being

5:51:57present over here I'm just scrolling

5:51:59down and it's not coming to an end

5:52:02similarly the output will be same in the

5:52:03other part files as well so this is all

5:52:06about the simple scope import command

5:52:08without specifying the target directory

5:52:10and the number of

5:52:11mappers now let's see how to import the

5:52:13data from rdbms to hdfs by specifying

5:52:16the target directory this part remains

5:52:19out to be the same now I will do one

5:52:21thing I'll specify the number of mappers

5:52:23as one so that your output will be in

5:52:25just one single part file and then I

5:52:28will specify the target directory as

5:52:30well and I will name the target

5:52:32directory as employee 10 enter again it

5:52:36takes a lot of time to execute because

5:52:38there is lot of data present in the

5:52:41database so it got executed and

5:52:43retrieved these many records again you

5:52:46can see the map reduce framework the

5:52:48number of bytes written number of read

5:52:50operations write operations Etc now

5:52:52again let's go to the htfs browser and

5:52:54see the output so I had specified the

5:52:57target directory name as employee 10 so

5:52:59you can see it here and the output is

5:53:02just in one single part file because I

5:53:04have specified the number of mappers as

5:53:05one in this case you can control the

5:53:08number of mapers independently from the

5:53:09number of files present in the

5:53:11directory so the entire records will be

5:53:14present in one single part file so this

5:53:17is the

5:53:19output now let me tell you how to import

5:53:22the command using wear Clause here you

5:53:25can import a subset of the table using

5:53:27the wear clause in scoop import tool it

5:53:29executes a corresponding SQL query in

5:53:31the respective database server and

5:53:33stores a result in a Target directory in

5:53:35hdfs so this is how the command goes the

5:53:39same command as before I'll just change

5:53:41the name of the target directory I'll

5:53:43specify employee 11 and I will increase

5:53:46the number of mappers to three and here

5:53:49I will specify the we

5:53:51Clause so let's give a condition like

5:53:54where the employee number will be

5:53:55greater than 499,000 it should dis the

5:53:58output enter so what you expect your

5:54:01output will be so here it displays the

5:54:03output records of the employees whose

5:54:05employee number is greater than

5:54:0949,000 so here it retrieved these many

5:54:12records which are above the employee

5:54:14number 49,000 now let's check the output

5:54:17so I have specified the target directory

5:54:19name as employee

5:54:22Lev so here you can see the output that

5:54:25it has retrieved the records of employe

5:54:28number which is more than

5:54:3049,000 so I hope you understand how to

5:54:33do this next let's see how to import all

5:54:36the tables from the rdbms database

5:54:38server to the hdfs here each table data

5:54:41is stored in a separate directory and

5:54:43the directory name is same as a table

5:54:45name it is mandatory that every table in

5:54:47that database must have a primary key

5:54:51field the command will be simple like

5:54:53simple import but just that you have to

5:54:55remove the table name and you have to

5:54:58replace the import with import all

5:55:00tables that's

5:55:02all so it will retrieve all the tables

5:55:05present in the

5:55:11employes okay so it imported all the

5:55:14tables from rdbms to htfs now let's

5:55:17check the

5:55:18output again I have not specified the

5:55:21target directory file so by default it

5:55:23will be in user edureka and you can see

5:55:26here it imported all the the tables

5:55:28present in the database to

5:55:31hdfs so these are the various tables

5:55:33present in rdbms and now it is present

5:55:36in htfs so this is how import all tables

5:55:39command works so this was all about

5:55:42executing import command in various ways

5:55:45now let's move further and see how does

5:55:47scop export works it exports the data

5:55:50from hdfs to rdbms correct again the

5:55:53command goes very simple you have to

5:55:55specify the connection your table name

5:55:56username and instead of the target

5:55:59directory you have to specify the export

5:56:01directory path so let's see how it works

5:56:04one important thing to notify the target

5:56:07table must exist in the Target database

5:56:09that is the data is stored as records in

5:56:11hdfs and these records are read and pass

5:56:14and delimited with the user specified

5:56:16delimiter the default operation is to

5:56:18insert all the records from the input

5:56:19files to the database using the insert

5:56:21statement in update mode scoop generates

5:56:24the update statement that replaces the

5:56:26existing record in the database so first

5:56:28we are creating an empty table where we

5:56:30will export our

5:56:31data I'm going to create a table called

5:56:34employee

5:56:36zero the primary key value should never

5:56:39be null so I'm specifying it as not

5:56:45null so here I created an empty table

5:56:47called employee zero and now I'll show

5:56:49you how the scoop export

5:56:52works here instead of import I'll make

5:56:54it as export this is the database name

5:56:58and I have created a table called

5:57:00employee zero and I'm going to specify

5:57:02the path for export

5:57:05directory simple that's

5:57:09all so you can see here it exported all

5:57:12these records into the rdbms so now

5:57:15let's cross check I'm going to give

5:57:18select count star from the table name

5:57:21that I have specified to export the

5:57:23tables so you can see the entire records

5:57:26got exported to the this table so this

5:57:29is how the scoop export

5:57:31works now let's see how to list the

5:57:33database present in the relational

5:57:35database here you need not even specify

5:57:38the database name because you're going

5:57:39to list the database that is present in

5:57:41the relational database system so you're

5:57:43going to specify list

5:57:45databases and execute the

5:57:49command so the databases present in the

5:57:51relational database system is test jdbc

5:57:55test employees and information schema

5:57:57again Let's cross check I'm going to

5:58:00give show databases it retrieve the same

5:58:02database so the result tallies so we can

5:58:06also list the tables present in the

5:58:08database let's see how to list all the

5:58:10tables present in the

5:58:13database again it's very simple you have

5:58:16to just specify the database name like

5:58:19here and give list tables instead of

5:58:25UT so you can see that it listed all all

5:58:27the tables present in the database

5:58:29employees so again let's cross check

5:58:33show

5:58:38tables same thing and now let's see what

5:58:41is Cen in object oriented application

5:58:45every database table has one data access

5:58:47object class that contains getter and

5:58:49seter methods to initialize the objects

5:58:51and coachin generates Dao class

5:58:54automatically and it also generates the

5:58:56Dao class in Java based on the table

5:58:59schema structure so this is a simple

5:59:01command let's see how to execute it here

5:59:04I will give scoop Cod

5:59:07gen and give the connection and the

5:59:09database name as

5:59:11employees and I'm going to specify the

5:59:13table name as well so it is going to

5:59:15create a employee CH file in which the

5:59:18backend code will be

5:59:20generated so I'll copy this path and

5:59:24jump into this

5:59:26directory so so you can see here that it

5:59:28created employees class jar file and the

5:59:31Java object file as well so now let's

5:59:34open the file system and check for the

5:59:36file so what was the name of the file it

5:59:40ends with 539

5:59:43D9 so here is the folder so in this you

5:59:47can see the object file that is being

5:59:52generated so this is the pack and code

5:59:54that is being

5:59:55generated so this is all about about how

5:59:57the scoop Cod gen

5:59:59[Music]

6:00:04Works what is Apache Pig so Apache pig

6:00:09is an abstraction over map reduce it is

6:00:12a tool or platform which is used to

6:00:14analyze larger sets of data representing

6:00:17them as data flows pig is generally used

6:00:20with Hadoop we can perform all the data

6:00:23manipulation operations in her doop

6:00:25using Apache pck to write data analysis

6:00:28programs Pig provides a highlevel

6:00:31language known as Pig Latin this

6:00:33language provides various operators

6:00:35using which programmers can develop

6:00:38their own functions for Reading Writing

6:00:41and processing data to analyze data

6:00:43using Apache Pig programmers need to

6:00:46write scripts using Pig Latin language

6:00:49all these scripts are internally

6:00:50converted into map and redu Tas

6:00:54respectively Apache Pig has a component

6:00:56called Apache shape Pig engine that

6:00:59accepts the pig latin scripts as input

6:01:01and converts those particular scripts

6:01:03into map reduced jobs so this was a

6:01:07basic introduction to Apache Pig now

6:01:10moving ahead we shall understand the

6:01:12different modes in which Apache Peg

6:01:14functions there are two particular modes

6:01:17in which Apache Peg functions those are

6:01:20the local mode and map reduce mode first

6:01:24we will understand what exactly is local

6:01:26mode in local mode Apache pig is

6:01:29designed to execute in a single jvm and

6:01:33is used for development experimenting

6:01:35and prototyping here files are installed

6:01:38and run using Local Host the local mode

6:01:42works on local file system and the input

6:01:45and the output data is stored in the

6:01:47local file system to access the command

6:01:50or the CR shell in local mode you need

6:01:53to execute a command called Pig hyphen X

6:01:56local

6:01:58we shall practically execute this in the

6:02:00demo section now moving ahead we shall

6:02:03discuss about the second type of mode in

6:02:05which Apache pck can be run that is the

6:02:08map reduce mode the map reduce mode is

6:02:11also known as Hadoop mode Apache Pig

6:02:14chooses Hadoop mode as its default mode

6:02:17in this pig renders Pig Latin into map

6:02:20rce shops and executes them on a Hadoop

6:02:23cluster it can be executed against

6:02:26semi-distributed

6:02:27or fully distributed Hadoop

6:02:29installation here the input and the

6:02:32output are present on

6:02:34hdfs the command for executing peg in

6:02:37map reduce mode is Peg or Peg hyphen X

6:02:41map reduce again we shall discuss about

6:02:44this particular command in our demo

6:02:46section where I'll show you both the

6:02:48modes and execute the pck scripts there

6:02:51now followed by the peg modes we shall

6:02:54understand the ways to execute Peg

6:02:56program so basically the big scripts are

6:02:59executed in three particular modes those

6:03:02are interactive mode batch mode and

6:03:05lastly the embedded mode no worry I'll

6:03:08explain to you each of these modes

6:03:11firstly we shall discuss about the

6:03:13interactive mode in this particular mode

6:03:15the pig is executed in a grun shell to

6:03:19invoke grun shell run the pig command

6:03:22once the grun mode executes we can

6:03:24provide big Latin statements and command

6:03:27command interactively at the command

6:03:29line itself so the next mode is the

6:03:32batch mode in this particular mode we

6:03:35can run a script file having a DOT Pig

6:03:38extension these files contain the pig

6:03:41latin commands so basically what we do

6:03:43is we write the pig command or the pig

6:03:47script and store it in a location and

6:03:50using the terminal we will access that

6:03:52particular location and that particular

6:03:54file and run the code present in that

6:03:56file

6:03:57so this is what happens in batch mode so

6:04:00followed by that the last mode is the

6:04:02embedded mode in this particular mode we

6:04:05can Define our own functions these

6:04:08functions can be called as userdefined

6:04:11functions here we use programming

6:04:13languages like Java and python to Define

6:04:16our own user defined functions so these

6:04:18were the three different modes in which

6:04:20we can execute Pig scripts now moving

6:04:23ahead we shall enter into our next topic

6:04:26that is the features of pig so there are

6:04:29five different and important features of

6:04:32pig so first up we shall understand the

6:04:34EAS of programming writing complex Java

6:04:38programs for map reduce is quite tough

6:04:40for non-programmers Peg makes this

6:04:43process very easy in the pig the queries

6:04:46are converted into map reduce processes

6:04:49internally so obviously it has grown the

6:04:52ease of programming followed by that the

6:04:55next important feature as the

6:04:57optimization

6:04:59opportunities so what exactly is

6:05:01optimization it is how the tasks are

6:05:03encoded permits the system to optimize

6:05:06the execution automatically allowing the

6:05:08user to focus on semantics rather than

6:05:11efficiency so followed by the

6:05:13optimization we have the

6:05:15extensibility a user defined function is

6:05:18written in which the user can write

6:05:21their logic to execute over the data

6:05:23sets so this increases the extensibility

6:05:26of the programmer followed by that the

6:05:29next important feature is it is highly

6:05:32flexible Apache pck can easily handle

6:05:35structured as well as unstructured data

6:05:38so Apache Pig tool is considered to be

6:05:40highly flexible irrelevant of the data

6:05:43type followed by that the last and

6:05:46important feature is inbuilt operators

6:05:50Apache pick contains various types of

6:05:52operators such as sort filter joints and

6:05:56Men anymore which are inbuild so here

6:06:00the programmer doesn't have to program

6:06:03these functions

6:06:05externally instead he can just directly

6:06:07get an access to those enbu functions

6:06:10and execute them so these were the five

6:06:12important features of Apache Pig so the

6:06:15next Topic in our today's discussion is

6:06:18Apache Pig installation into our local

6:06:20system now we shall discuss about one of

6:06:23the easiest ways to install Apache pig

6:06:26into our local system

6:06:27so today I'll explain you how to install

6:06:30Apache pck into Windows operating system

6:06:33so to do this the easiest way is to

6:06:36download one of the virtual machines so

6:06:38I would prefer you to download Oracle

6:06:41virtual box for this particular task

6:06:43followed by that we shall also download

6:06:45a quick start VM of cloud AR don't worry

6:06:49about these softwares I'll drop down the

6:06:51link for those softwares in the

6:06:52description box below you can use that

6:06:54and download them so once after the

6:06:57Oracle virtual box is installed into a

6:06:59local system and it is running this is

6:07:02how it looks like now what we want to do

6:07:06is to add a new virtual machine into our

6:07:08virtual box so for that you just need to

6:07:11select import option and it will give

6:07:15you a new dialogue box and from here you

6:07:18must redirect to the location where your

6:07:21virtual box is located so in my system

6:07:24it's located in F drive and inside F

6:07:29Drive CCA Cloud era virtual box cloud

6:07:33era quick start VM so there you go now

6:07:37you just have to select open and before

6:07:40we actually select the button import you

6:07:42might want to select the RAM and

6:07:44increase its size to at least 8 GB so

6:07:48just write in 9,000 MV which is just a

6:07:51little above 8GB now we are good to go

6:07:54just select the option import and your

6:07:56virtual Bo will be

6:07:58imported you can see that the quick

6:08:00start VM is getting

6:08:04imported now you can see that the

6:08:06virtual machine got successfully

6:08:08imported now to start it all you need to

6:08:10do is just click on it and then select

6:08:13the start

6:08:14button so you can see that the cloud

6:08:17error quick start VM version

6:08:205.13 has been successfully imported and

6:08:23booted

6:08:24up here we are we have the the quick

6:08:27start cler of VM welcome note and now if

6:08:30you want to start up with big editor you

6:08:33might want to log into Hue first or your

6:08:36hdfs remember in Cloud error the default

6:08:40username is cloud error and the password

6:08:43also is cloud era now let me tell you in

6:08:47Cloud era the default username and

6:08:50password for everything is cloud era for

6:08:53example let's log into our Hue using the

6:08:56username Cloud error and password Cloud

6:09:02error now you might want to just select

6:09:04remember to remember your

6:09:11password and there you go you have

6:09:13successfully logged into

6:09:16here and now if you want your query

6:09:19editor for pick you can just select the

6:09:22bottom Arrow Mark and there you can

6:09:24select editor and inside editor you have

6:09:27the editor designed for

6:09:31pig now you are in the window where you

6:09:34can write pick scripts and execute them

6:09:37now we will move ahead into our next

6:09:39topic and after finishing the theory

6:09:41part we shall come back into our Cloud

6:09:44era and execute some of the basic

6:09:46operations on Pig terminal so our next

6:09:49concept is understanding the pig

6:09:52architecture so first of all let us go

6:09:54through the diagram of pig AR

6:09:56architecture the following diagram

6:09:59represents the architecture of Apache

6:10:01pig as shown in the figure there are

6:10:04various components in Apache Pig

6:10:06framework let us look at the major

6:10:08components firstly the parser initially

6:10:12the pick scripts are handled by the

6:10:14passer it checks the syntax of the

6:10:17script does type checking and other

6:10:19miscellaneous checks the output of the

6:10:22Passa will be a dag which is directed a

6:10:26cyclic graph which represents the big

6:10:28Latin statements and logical operators

6:10:32in directed aycc graphs The Logical

6:10:35operators of the scripts are represented

6:10:37as the notes and data flows are

6:10:39represented as edges and next the

6:10:43optimizer the logical plan for directed

6:10:46asyc graphs is passed to The Logical

6:10:49Optimizer this carries out the logical

6:10:52optimizations such as projections and P

6:10:54Downs next comes the compiler the

6:10:58compiler is used to compile the

6:11:00optimized logical plan into the series

6:11:02of map reduce shops followed by that we

6:11:06have the execution engine finally the

6:11:09map reduce jobs are submitted to the

6:11:11Hadoop in a sorted order finally these

6:11:15map ruce jobs are executed on Hado

6:11:18producing the desired results so this

6:11:21was a basic explanation based on Apache

6:11:24pck architecture now moving ahead we

6:11:27shall understand the major advantages of

6:11:30Apache P so some of the major advantages

6:11:33of Apache Pig are less code the pick

6:11:36consumes less than a line of code to

6:11:39perform any operation so this reduces

6:11:43the number of lines included in the code

6:11:46followed by that code

6:11:48reusability the big code is so flexible

6:11:50enough to reuse it again you can

6:11:53basically write the piig script into a

6:11:55file and access the file whenever you

6:11:57need it so followed by that the next

6:12:00important advantage of Apache pck is the

6:12:03nested data types the pig provides a

6:12:05useful concept of nesting data types

6:12:08like topple bag and map so these were

6:12:11the few important advantages of Apache

6:12:14Pig now we shall move ahead into the

6:12:16next topic which is about the

6:12:18differences between Apache Pig and

6:12:20Apache map reduce so there are basically

6:12:23four differences between Apache Pig and

6:12:26Apache map reduce so first up we will

6:12:29begin with map reduce in map reduce we

6:12:32have low-level data processing when it

6:12:35comes to Apache Peg it is considered as

6:12:38a high level data processing tool

6:12:40followed by that the next important

6:12:42difference between map reduce and Apache

6:12:44p is that we find complex Java programs

6:12:48when we are using map reduce but when we

6:12:51come into Apache Pig we have SIMPLE

6:12:55programming script

6:12:57which are shorter in code length and

6:12:59easily understandable followed by that

6:13:02the third difference is the data

6:13:04operations are completely tough and

6:13:06complicated in map reduce but in Apache

6:13:10Peg you don't have to worry about those

6:13:12operations because they're already built

6:13:15in next and the last difference between

6:13:19map reduce and Apache pick is Apache map

6:13:22reduce does not allow nested data types

6:13:25whereas Apache Pig allows nested data

6:13:28types so these were the basic

6:13:30differences between Apache map ruce and

6:13:32Apache Pig so the next topic is the pig

6:13:36demo here we will understand all the

6:13:39basic commands in aach Pig and the basic

6:13:42functionalities which are available in

6:13:45pig now without further Ado let's

6:13:48quickly begin with our demo for today's

6:13:50session now we have come back to our

6:13:52Cloud era that we have installed into

6:13:54our local system now let's start pig as

6:13:58we have discussed before Apache Pig can

6:14:00be executed in two modes they are the

6:14:03local mode and the map reduce mode

6:14:05firstly we shall execute an example

6:14:08based on local mode to start pck in

6:14:10local mode you need to type in the

6:14:12command Pig space hyphen X space local

6:14:17firing this command will enable Pig in

6:14:19local mode as you can see the command is

6:14:22getting

6:14:25decrypted now you can see that it has

6:14:28successfully started the gr shell in

6:14:30local mode now what are we going to

6:14:33execute we are going to execute a very

6:14:35simple wordon program for this

6:14:38particular wordon example I have

6:14:41considered a basic text that is the

6:14:43definition of Apache Pig which happens

6:14:45to be Apache pig is a high level

6:14:48platform for creating programs that run

6:14:50on Apache Hadoop the language for this

6:14:53particular platform is called Pig Latin

6:14:55Pig can execute its Hadoop jobs in map

6:14:58reduce Apache test and Apache spark so

6:15:01we will consider this particular

6:15:03paragraph and we will also count the

6:15:05number of words which are included in

6:15:07this particular paragraph now let's

6:15:09quickly go back to our terminal and

6:15:11execute our

6:15:12program so now we have come back to our

6:15:15terminal let's clear it using the

6:15:16command control+ L now let's quickly

6:15:20start up typing our

6:15:22commands so as discussed before we are

6:15:24going to execute this particular program

6:15:26in local mode so here we not loading the

6:15:29data into htfs instead we considering

6:15:31the local

6:15:33location which is the local location

6:15:36that is home Cloud error desktop word

6:15:38count now we're going to run this

6:15:39command and see the output so the data

6:15:41has successfully loaded now we'll

6:15:43execute the next

6:15:45command so now we're going to tokenize

6:15:48each and every single word in this

6:15:50particular paragraph and we'll be

6:15:52considering each and every word as a

6:15:54single word and we are going to to

6:15:56separate each and every word by using a

6:15:58space now let's fire up this command and

6:16:00see the output yeah the command just got

6:16:02executed or the script just got executed

6:16:05now in the next script we are going to

6:16:07group the words according to their

6:16:08occurrence yeah even that is done now

6:16:11the next script would help us to count

6:16:13the number of words which have been

6:16:14repeated so even that is done now the

6:16:17last script would be dumping the output

6:16:19which is stored in p word c which is pig

6:16:22word count now let's enter the command

6:16:24and see the output you can see see some

6:16:26map Ru shops that are getting executed

6:16:28and there you go we have our

6:16:31output so the words are a n on for its

6:16:35days Hado Apache and all those words and

6:16:38along with them we also have the number

6:16:39of repetition of each and every word for

6:16:42example we have Hado which is repeated

6:16:44one time and the word Apache is repeated

6:16:46for four times and so on so with this

6:16:49now let us move ahead and execute some

6:16:51examples based on apache's Hadoop mode

6:16:54or map reduce mode now let's close this

6:16:56terminal and open a new one and let's

6:16:59begin uh executing Hardo in map reduce

6:17:01mode as discussed before apach P can be

6:17:05executed in both local mode and map

6:17:08reduce mode we have already executed

6:17:10some examples based on local mode now we

6:17:13shall execute some examples based on map

6:17:15reduce mode to do so we will first load

6:17:18some local data into

6:17:20htfs so using this command I'll be

6:17:23loading my local data that I expect to

6:17:26tutorial. CSV into my hdfs and I'll name

6:17:29it as Eda in as you can see the command

6:17:32is getting

6:17:33deprecated and the data has been

6:17:35successfully loaded now let us use cat

6:17:38command and see what exactly is present

6:17:40in that particular

6:17:43data as you can see the command got

6:17:46deprecated and the data is a simple CSV

6:17:48file related to students that has ID

6:17:51name department and year now let's start

6:17:55pick in map reduce mode to do so we just

6:17:58need to type in Pig and strike

6:18:02enter there you go we have successfully

6:18:04started Pig in map reduce mode now let's

6:18:07execute some

6:18:13commands so using this command we will

6:18:16be loading edura doin that is the data

6:18:18the CSV file which we have discussed

6:18:19before using pick storage as comma

6:18:22separated file and the schema for the

6:18:24data will be ID as character array name

6:18:27as character array Department as

6:18:29character array and ear as integer array

6:18:33now let's track and enter and see the

6:18:35output there you go the data has been

6:18:37successfully loaded now now let's use

6:18:40dump command to see the data what we

6:18:42have

6:18:47loaded so there you go the data has been

6:18:50successfully dumped so this was the data

6:18:53which we have

6:18:54loaded now let's move further and

6:18:56execute few more

6:19:05commands now we shall use for each

6:19:08command and for each data present in our

6:19:10data file we will generate ID name and

6:19:14Department there you go the command got

6:19:17successfully executed now let's dump the

6:19:19data using dump

6:19:21command so Pig for each is our variable

6:19:25that stores the data now let's type in

6:19:27semicolon and

6:19:35enter so there you go the 4 command has

6:19:38been successfully executed and we have

6:19:41generated ID name and department now we

6:19:44shall execute few more

6:19:50examples now we shall use descending

6:19:52operator to arrange the data in the form

6:19:55of descending order of ID so it's been

6:19:58executed now let's dump the data and see

6:20:00the

6:20:06output there you go we have successfully

6:20:09executed the Dum command and the IDS

6:20:11have been arranged in the descending

6:20:13order now as you can see it now this is

6:20:15how the order by descending function

6:20:17works now let's move ahead and execute

6:20:20one last

6:20:24command

6:20:29as you can see here we are using filter

6:20:31operation we are going to filter the

6:20:33students based on the department where

6:20:35department is equals to csse so

6:20:38executing this command will give us the

6:20:39students which are inside the department

6:20:42C now the command got executed now let's

6:20:45dump the data using dump

6:20:50command so there you go you can see some

6:20:52commands getting

6:20:54deprecated

6:20:56here you can see the map redu shops

6:20:58getting

6:21:00executed so there you go you can finally

6:21:03see the output which has the students

6:21:05that belong to the Cs

6:21:08[Music]

6:21:12Department why exactly we needed Apache

6:21:15Hive it All Began in the early '90s when

6:21:18Facebook started slowly the number of

6:21:20users at Facebook increased that is

6:21:23nearly 1 billion users and along with

6:21:25the users increase the data which is

6:21:27nearly equals to thousands of terabytes

6:21:29of data and nearly one lakh queries then

6:21:32also 500 million photographs uploaded

6:21:35daily and this was a huge amount of data

6:21:38that Facebook had to process and the

6:21:40first thing that everybody had in their

6:21:42mind was to use rdbms and we all know

6:21:45that rdbms couldn't handled such a huge

6:21:48amount of data and neither it was

6:21:50capable enough to process it and the

6:21:52very next big guy who was capable enough

6:21:54to handle all this big data was Hadoop

6:21:58even when Hadoop came into picture it

6:21:59was not too easy to manage all the

6:22:01queries it used to take a lot of time to

6:22:04execute all the queries performed so one

6:22:07common thing that all the Hardo

6:22:09developers had was the SQL so they

6:22:12thought to come up with a new solution

6:22:13that has hadoop's capacity and interface

6:22:16like SQL that is when Hive came into

6:22:19picture so now we understand the exact

6:22:21definition of Apache Hive Apache Hive is

6:22:24a data warehouse soft sofware project

6:22:26built on top of Apache hardup for

6:22:28providing data query and data analysis

6:22:31Hive gives a SQL like interface to query

6:22:33data stored in various databases and

6:22:35file systems that integrate with hadu

6:22:38also Apache Hive has data warehousing

6:22:40software utility it can be used for data

6:22:43analytics it is built for SQL users

6:22:46manages querying of structured data and

6:22:48it simplifies and abstracts the load

6:22:51that is on Hadoop and lastly no need to

6:22:53learn Java and Hadoop a API to handle

6:22:56data using Hive so followed by this we

6:23:00shall understand Apache Hive

6:23:02applications Apache Hive is used in many

6:23:05major applications few of the major

6:23:08applications are as follows Hive is a

6:23:11data warehousing infrastructure for hu

6:23:14the primary responsibility of Hive is to

6:23:16provide data summarization query and

6:23:18data analysis it supports analysis of

6:23:21large data sets in Hero's hdfs as well

6:23:24as on Amazon on S3 file system followed

6:23:27by that we have document indexing with

6:23:30Hive the goal of Hive indexing is to

6:23:33improve the speed of query lookup on

6:23:35certain Columns of a table without an

6:23:37index queries could load an entire table

6:23:40or partition a whole process as rowes

6:23:43this would be Troublesome so with Hive

6:23:46we have solved this problem followed by

6:23:48that predictive modeling the data

6:23:51manager allows you to prepare your data

6:23:54so it can be processed in automated

6:23:56analytics it offers a variety of

6:23:58preparation functionalities including

6:24:00the creation of analytical records and

6:24:02timestamp populations followed by that

6:24:05the next important application of Hive

6:24:08is business

6:24:09intelligence Hive is a data warehousing

6:24:11component of Hadoop and it functions

6:24:14well with structured data enabling ad

6:24:16hoc queries against large transactional

6:24:18data sets hence it happens to be a

6:24:21best-in-class tool available for

6:24:23business intelligence and helps many

6:24:24companies to predict their business

6:24:26requirements with high accuracy last but

6:24:29not the least lock processing Apache

6:24:32Hive is a data warehouse infrastructure

6:24:34built on top of Hardo it allows

6:24:37processing of data with SQL Li queries

6:24:39and is very pluggable so that we can

6:24:41configure it to provide our logs quite

6:24:43easily so these were the few important

6:24:45Hive applications now let us move ahead

6:24:48and understand Apache Hive features the

6:24:51first and the foremost important feature

6:24:53of Apache Hive is see SQL type queries

6:24:57the SQL type queries present on hi will

6:24:59help many of the Hado developers to

6:25:01write queries with ease followed by that

6:25:04the next important feature of Apache

6:25:07Hive is oap based design oap is nothing

6:25:11but online analytical processing this

6:25:14allows users to analyze database

6:25:16information from multiple database

6:25:18systems at one time so using Apache Hive

6:25:22we can achieve o AP with higher accuracy

6:25:25followed by the second feature we have

6:25:27the third feature which says Apache Hive

6:25:29is really fast since we have SQL like

6:25:32interface in Apache Hive using this

6:25:35feature on sdfs will help us writing

6:25:38queries faster and executing them

6:25:40followed by that we believe Apache Hive

6:25:43is highly scalable Hive tables are

6:25:45defined directly in Hardo file system

6:25:48hence Hive is fast and scalable and easy

6:25:51to learn followed by that it is known to

6:25:54be highly extensible

6:25:56Apache Hive uses Hadoop file system and

6:25:59Hadoop file systems or hdfs provides

6:26:01horizontal extensibility and finally the

6:26:05ad hoc wearing using H we can execute ad

6:26:08hoc wearing to analyze and predict data

6:26:11so these were the few important features

6:26:12of Apache Hive let us move on to our

6:26:15next topic where we deal with Apache

6:26:17Hive architecture the following

6:26:19architecture explains the flow of

6:26:21submission of query into hiy the first

6:26:24stage is The Hive client Hive allows

6:26:27writing applications in various

6:26:29languages including Java Python and C++

6:26:33it supports different types of clients

6:26:35such as Thrift server jdbc driver and

6:26:38odbc driver so what exactly is Thrift

6:26:42server it is a cross language service

6:26:44provider platform that serves the

6:26:46request from all these programming

6:26:48languages that supports Thrift followed

6:26:51by that jdbc driver it is used to

6:26:54establish connection between Hive and

6:26:56Java applications the jdbc driver is

6:26:59present in the class or. Apache dohad

6:27:02doh. jdbc dohy driver finally we come to

6:27:07odbc driver so what exactly is odbc

6:27:10driver obbc driver allows the

6:27:13applications that support obbc protocol

6:27:15to connect to Hive followed by that we

6:27:18have the hive Services the following are

6:27:21the services provided by hve they are

6:27:24hve C Li Hive web user interface Hive

6:27:28meta store Hive server Hive driver Hive

6:27:31compiler and lastly The Hive execution

6:27:34engine The Hive CLI or command line

6:27:37interface is a shell where we can

6:27:40execute the hive queries and commands

6:27:42followed by that the hive web UI is just

6:27:46an alternative for Hive CLI it provides

6:27:49a web-based graphical user interface for

6:27:52executing High queries and commands

6:27:54follow by that the hive meta store it is

6:27:57a central respository that stores all

6:28:00the structured information of various

6:28:01tables and partitions in the warehouse

6:28:05it also includes metadata of column and

6:28:07its type information the serializers and

6:28:10D serializers which is used to read and

6:28:13write data and the corresponding hdfs

6:28:15files where the data is stored followed

6:28:18by that the H server it is referred to

6:28:21as Apache th server it accepts the

6:28:24request from different clients and

6:28:26provides to The Hive driver moving on we

6:28:28shall deal with Hive driver The Hive

6:28:31driver receives queries from different

6:28:32sources such as web UI CLI Thrift and

6:28:36jdbc or odbc drivers it transfers the

6:28:39queries to the compiler followed by that

6:28:42we have the hive compiler the purpose of

6:28:45Hive compiler is to pass the query and

6:28:48perform semantic analysis on the

6:28:49different query blocks and expressions

6:28:52it converts hiveql statements into

6:28:55produce jobs finally we have Hive

6:28:58execution engine Hive execution engine

6:29:01is the optimizer that generates the

6:29:03logical plan in the form of dag or

6:29:06directed aycc graph of map redu task and

6:29:09hdfs tasks in the end the execution

6:29:12engine executes the incoming task in the

6:29:14order of their dependencies followed by

6:29:17that we have the map reduce and hdfs map

6:29:20reduce is the processing layer which

6:29:22executes the mapping and reducing jobs

6:29:24on the data provided lastly the sdfs or

6:29:28Hardo distributed file system is the

6:29:30location where the data which we provide

6:29:32is stored so this is the architecture of

6:29:35Apache Hive then moving next we have

6:29:38Apache Hive components so what are the

6:29:41different components which are present

6:29:42in Hive they are first one the Shell

6:29:46Shell is the place where we write our

6:29:48queries and execute them followed by

6:29:51that we have metast store as discussed

6:29:54in the AR Ure the metast store is a

6:29:56place where all the details related to

6:29:58our tables is stored like schema Etc

6:30:02followed by that we have the execution

6:30:04engine so execution engine is the

6:30:06component of Apache Hive which converts

6:30:09the query or the code which we have

6:30:11written into the language which The Hive

6:30:14can understand followed by that driver

6:30:16is the component which executes the code

6:30:19or query in the form of acyclic graphs

6:30:22and lastly the compiler compiler

6:30:25compiles whatever the code we write and

6:30:27executes and provides us the output so

6:30:30these are the major Hive components

6:30:32moving ahead we shall understand Apache

6:30:34Hive installation on Windows operating

6:30:37system so urea is all about providing

6:30:39the technical knowledge in the simplest

6:30:41way as possible and later play around

6:30:43with the Technologies to understand the

6:30:45complicated parts of it so now let's try

6:30:48to install hyve into our local system in

6:30:50the most simplest way as possible to do

6:30:53so we might need the Oracle virtual box

6:30:56which looks like this so once after you

6:30:58download Oracle virtual box and install

6:31:01it into your local system The Next Step

6:31:03would be to download the cloud era

6:31:05quickart VM for your local system the

6:31:07link to this will be provided in the

6:31:09description box below now let's quickly

6:31:11start our Cloud era quick start VM with

6:31:14our Oracle virtual box select import

6:31:17option and now provide the location

6:31:19where your cloud data quick start VM is

6:31:21existing in my local system it is in the

6:31:24local desk Drive

6:31:28F there you go select open and now just

6:31:33make sure your Ram size is more than 8GB

6:31:36just randomly I'm providing 9,000 MB

6:31:39which is just above 8GB so that you have

6:31:41a smooth functionality of cloud erra now

6:31:44select

6:31:45import and there you go you can see that

6:31:48cloud era quick start VM is getting

6:31:53imported

6:32:00now you can see that cloud quickart VM

6:32:02has been successfully imported and it's

6:32:04ready for deployment you can just double

6:32:06click on it and it'll get

6:32:17started you can see that cloud era VM

6:32:20has been successfully imported and it

6:32:22started and also you can see that we

6:32:24have gone live on cloud era you can see

6:32:26all the hu Hado Edge space Impala spark

6:32:29which are pre-installed in Cloud era now

6:32:31our concern would be to start up hve so

6:32:34to start hve you need to start up Hue

6:32:37first so let me remind you one thing in

6:32:40Cloud era every single password and

6:32:43username is cloud era by default so for

6:32:46example we've got H username and

6:32:49password here so the username that is

6:32:51the default username for cloud eras h

6:32:54would be Cloud error and along with that

6:32:57even the password will be Cloud error

6:32:59that is by default so we have got Cloud

6:33:02error and Cloud error as username and

6:33:04password respectively let's just sign in

6:33:07you may select remember option in case

6:33:10if you forget your

6:33:12passwords so now we are getting

6:33:14connected to Hugh and we are live on

6:33:16Hugh

6:33:17now there you go we've got started our

6:33:19Hue so now we'll enter into

6:33:23htfs there we go we have a hive

6:33:27here now that we have successfully

6:33:29installed Hive into our local system let

6:33:32us move further and understand few more

6:33:33Concepts in Hadoop firstly we should

6:33:36deal with the data types the data types

6:33:38are completely similar to any other

6:33:40programming language which we have they

6:33:41are tiny end small end integer big end

6:33:45similarly followed by that we have float

6:33:47and in Side High float is used for

6:33:49single precision and if you want double

6:33:51Precision you can go ahead with double

6:33:54and followed by that we have a string

6:33:56and Boolean which are completely similar

6:33:58to any other programming languages which

6:34:00we use in this daily life followed by

6:34:03that we have hve data models so these

6:34:05are the basic data models which we use

6:34:07in Hive that we basically create

6:34:09databases and store our data in the form

6:34:11of tables and sometimes we also need

6:34:14partitions we will discuss each one of

6:34:16these data models in our demo ahead so

6:34:19we'll first create databases and inside

6:34:21databases we will be creating tables

6:34:23inside which we will will be storing

6:34:25data in the form of rows and columns and

6:34:27along with that partitions partitions uh

6:34:30they are like Advanced way of storing

6:34:33data like if you have just imagine you

6:34:35are in a school say standard one and

6:34:38inside standard one you have sections a

6:34:40b c d so partition is like you're

6:34:42getting partitions for Section a section

6:34:44B section c and section D you're storing

6:34:47different different students in

6:34:48different different sections so that

6:34:50when you're querying for a particular

6:34:51data for example say you're searching

6:34:54for a kid called Sam and you have the

6:34:56section of his class SB so you just

6:34:59don't have to just search for Sam in all

6:35:02the four sections you can just directly

6:35:04go into section B and call in Sam and

6:35:07you'll get access to him that's how

6:35:09partitions work followed by partitions

6:35:11we have buckets so similar to partitions

6:35:14even buckets work in the same way let's

6:35:16understand each one of these in much

6:35:17better way through a practical demo

6:35:20after data models we shall understand

6:35:22about hype operators so what are

6:35:24operators operators are any other

6:35:26operators that we use in normal

6:35:28programming languages such as arithmatic

6:35:30operators logical operators we shall

6:35:32also go through some examples based on

6:35:34arithmatic and logical operators in Hive

6:35:37in the hive demo we will use some

6:35:39arithmatic operations as well as logical

6:35:42operations on the data which we have

6:35:43stored in the form of tables in Hive we

6:35:45shall go through a brief look on that as

6:35:47well so before we get started let's have

6:35:50a brief look on the CSV files that I

6:35:52have created for today's demo these are

6:35:54the small CSV files that I've personally

6:35:55created using msxl and I've saved them

6:35:58as CSV files I've made the CSV files to

6:36:01be smaller because just to make sure the

6:36:04execution time consumed is as less as

6:36:06possible since we using Cloud era the

6:36:08execution time might be a little more so

6:36:11it's better we use smaller CSV files so

6:36:14this is my first CSV file which is

6:36:16employee. CSV which has employ IDs

6:36:18employee name salary and age similarly

6:36:21we have another employee 2. CSV file

6:36:24which has the same details along with

6:36:25one more column that is the country

6:36:27column I've included country because we

6:36:29will be using this country column in

6:36:31Joins that we will be performing in

6:36:33future followed by that we have the

6:36:35department so here we have Department ID

6:36:37and Department name so we have

6:36:40Development Department testing product

6:36:42relationship admin and it support

6:36:44similarly we also have student CSV this

6:36:47is another CSP file that I have created

6:36:49this has ID name course and age of the

6:36:52student followed by that we have another

6:36:54CSV this is student report. CSV which

6:36:57has the reports of a particular student

6:37:00gender ttic City parental education

6:37:02lunch course math score reading score

6:37:05writing score and other so these are the

6:37:07CSV files that we will be using in our

6:37:09demo today so now let's quickly begin

6:37:11with our demo so to start Hive we shall

6:37:14open a terminal so starting or firing up

6:37:17Hive in Cloud arise really simple you

6:37:19just have to type in Hive and enter

6:37:22there you go logging initialized using

6:37:25configuration files and Etc the hi CLI

6:37:28is deprecated and migration to bline is

6:37:31recommended and there you go your hi

6:37:33terminal or CLI has been started so

6:37:36first let's try to create a database to

6:37:39save time I've already created the

6:37:41document which has all the codes that we

6:37:43will be executing today so this is the

6:37:46particular file which I will be using

6:37:47today so don't worry this file will be U

6:37:50Linked In the description box below you

6:37:52can use the same file and try executing

6:37:54the same codes in your personal systems

6:37:56just for practice if you feel so so just

6:37:59to save time I've already created uh the

6:38:01document which has the codes that we are

6:38:03going to execute today so this code or

6:38:05this file will be attached in the

6:38:07description box below you can get access

6:38:09to it and you can also execute the same

6:38:11codes in your own personal system to

6:38:13have a practical experience about this

6:38:15particular Hy tutorial so the first

6:38:17thing that we will be doing today is to

6:38:19create a database so I'm going to create

6:38:22the database using SQL type commands

6:38:25which are create database name of the

6:38:27database which is Eda there you go the

6:38:30database has been successfully

6:38:32created so now you can also use the

6:38:35following command to check if your

6:38:37database has been created or not so show

6:38:39databases will help you to find it so

6:38:42there you go you can see the first

6:38:43database which is a default database

6:38:46which will be pre-existing and followed

6:38:48by that you have our own database which

6:38:49we have created now that is edura so

6:38:52followed by this next we will move ahead

6:38:54and try to create a new table so when

6:38:57you come into tables you need to

6:38:59understand there are two types of tables

6:39:01in Hive they are managed tables or

6:39:04internal tables followed by that

6:39:06external tables so what is the

6:39:08difference between these two tables so

6:39:11internal table or manage table is the

6:39:13default table that will be created

6:39:15whenever you try to create a table in

6:39:17high so for example if you're trying to

6:39:19create a new table say Eda then hi

6:39:24considers that particular table as an

6:39:26internal table by default so when you

6:39:29create an internal table your data is

6:39:32not secured understand this so when you

6:39:34create an internal table your data is

6:39:36not secured in case just imagine you are

6:39:39working with a team and all your team

6:39:42members have access to your hive or hue

6:39:45so the table has been existing in your

6:39:47hve and some random newbie or some

6:39:50random inexperienced guy tries to change

6:39:53few things in your table and

6:39:55accidentally he ends up deleting the

6:39:57table so when you delete the table then

6:40:00if the table was created using an

6:40:02internal table code then your data will

6:40:04be erased so that's the disadvantage of

6:40:07using internal tables but in case if you

6:40:10create an external table even if

6:40:12somebody tries to delete your table the

6:40:15table or the data whatever is there will

6:40:18be deleted from their own local system

6:40:20but not from Hy so that's the best part

6:40:22of using external tables don't worry we

6:40:24will discuss about internal tables and

6:40:27external tables as well so first we'll

6:40:29try to create an internal table so this

6:40:31particular code is based on internal

6:40:33tables so we are using SQL type command

6:40:36here which is create table and the table

6:40:38name is employee and the columns inside

6:40:41our table are ID of the employee name of

6:40:44the employee salary and age so row

6:40:47format has been delimited followed by

6:40:49that since this is a CSV file so the fs

6:40:52will be terminated by comma

6:40:54and don't forget you have to use

6:40:56semicolon unless you use semicolon your

6:40:58code is not complete so let's fire and

6:41:01enter and see if the table gets created

6:41:02or not yeah the table is created

6:41:05successfully now we shall see the table

6:41:08or let's describe the table so

6:41:10describing the table means you can see

6:41:12what are the columns which are present

6:41:13in your table so to describe a table you

6:41:15can use the keyword describe a name of

6:41:17the table which is employee and don't

6:41:20forget semicolon there you go so your

6:41:22table has the column ID name salary age

6:41:26so those are the four columns which you

6:41:27have included in your particular table

6:41:29employee now let's move ahead and see if

6:41:32this particular table is an internal

6:41:34table or manage table or the other type

6:41:36of table which is the external table so

6:41:39to do that we can just write and

6:41:41describe formatted table name and

6:41:45semicolon there might be a small issue

6:41:48here yeah there is a typing mistake that

6:41:51is described missed s so there you go we

6:41:55got it so this particular table is

6:41:58managed table as you can see here now

6:42:01let's move ahead and uh try out external

6:42:04tables let's clear our screen first you

6:42:07can use control+ l to clear your screen

6:42:10there you go we have a clear screen now

6:42:12now let's try to create an external

6:42:14table creating an external table is

6:42:17completely similar to that of internal

6:42:19table but the only difference is that

6:42:21you need to add in a keyword which is

6:42:24external so this particular keyword is

6:42:26used to create an external table now

6:42:29let's fire an enter and see if the table

6:42:30gets created or not you can see the

6:42:33table got created now let's try to

6:42:35describe the table employee 2 don't

6:42:38forget the semicolon I'm saying this

6:42:40again and again because most of the

6:42:42times we miss semicolon and we will get

6:42:45an error so you can see the table got

6:42:46described and we have the following

6:42:48columns inside our table now let's move

6:42:51ahead and see if this particular table

6:42:53is an external table or a manage table

6:42:57to do so you can type in describe

6:42:59formatted the same code what we have

6:43:01used earlier that is described formatted

6:43:04name of the table that is employ to

6:43:07semicolon don't forget there is some

6:43:10issue again I think I've missed

6:43:11something or maybe a typing error yeah

6:43:14this is a typing

6:43:17error yeah there you go the table type

6:43:20is external table so that's how we

6:43:23create an internal table or manage table

6:43:26and external table so now that we have

6:43:29understood how to create a database and

6:43:31table and the two types of tables that

6:43:34are internal table or managed table

6:43:36followed by that the second type of

6:43:38table that is the external table now

6:43:40let's try to create an external table in

6:43:42a particular location so for that you

6:43:45can use the following code but the only

6:43:47difference is you are specifying the

6:43:49location that is user Cloud era urea

6:43:52employee edu EMP is a file that we will

6:43:55be creating in our hyp so let's fire and

6:43:58enter and see it if it's created or not

6:44:00yeah it's successfully created let's go

6:44:03back to Hue and see if the following

6:44:05table is created or not so one thing you

6:44:07have to remember is when you fire in a

6:44:09command or if you try to create a table

6:44:12the first folder that will be created is

6:44:14a warehouse so inside hve you have your

6:44:17warehouse and inside Warehouse you have

6:44:19all the databases that we have created

6:44:21our first database was the Ed Raa

6:44:23database and after that we have created

6:44:26table which is employee and the second

6:44:28table is employee 2 so this is in the

6:44:31particular location which is user Cloud

6:44:33error and the file is employed to let's

6:44:36see that this was the

6:44:42file yeah sometimes H will not show it

6:44:45because of network issues you don't have

6:44:47to worry about it you will get back that

6:44:49data now followed by this let us enter

6:44:52into Hue again

6:44:54so when you come back into Hue if you

6:44:56have to upload a file into Hue you can

6:44:58just select this particular option which

6:45:00is plus so selecting this will give you

6:45:02a dialogue box which will be something

6:45:04like this and here you can just select

6:45:06any of the files which you want to

6:45:08upload into Hue now let me select a

6:45:10student report. CSV and select open so

6:45:14there you go upload is in progress so

6:45:16the data file has been successfully

6:45:17uploaded now if you want to access your

6:45:19data file you can just click on that so

6:45:22there you go you have all your data

6:45:24successfully loaded onto

6:45:26Hue you can also perform queries on this

6:45:29particular data you can just select

6:45:31query and inside that you just need to

6:45:33select editor and you have various

6:45:35editors over here which is Peg Impala

6:45:38Java spark map reduce shell scoop and we

6:45:42also have hi in here so if you just

6:45:44select Hive and there you go you have

6:45:46the editor here you can just type in

6:45:48your commands or queries whatever you

6:45:50have see you have many dictionaries as

6:45:53well you can just select any one of

6:45:54those select and that's how you write

6:45:56queries on the hyp terminal now let's

6:46:00not waste much time here and we have a

6:46:02lot to learn so let's continue with the

6:46:05next topics in our today's session now

6:46:07we shall try to edit the

6:46:10tables now we have created the new table

6:46:13that is employee 3 and we have named The

6:46:16Columns as ID name String salary age and

6:46:19Float now we should try to make some

6:46:22alterations to our table

6:46:24so the first alteration that we will try

6:46:26to make to our table is to rename our

6:46:29table as EMP table you know that our

6:46:32employee table was named as employee 3

6:46:35now we trying to rename it to EMP table

6:46:38so we are using the keyword alter here

6:46:41so just fire and enter and see if this

6:46:43is possible or not yeah it is possible

6:46:45the name has been changed to EMP table

6:46:48now let's try if it's completely changed

6:46:50or clearly changed or not you can just

6:46:52type in describe EMP table semicolon if

6:46:57we get the same column names in our

6:47:01description then it should be changed so

6:47:04there you go we can see the same columns

6:47:06here so we have successfully changed the

6:47:08name to EMP table now we shall also try

6:47:11to add in some more columns to our table

6:47:14which is EMP table so here we'll try to

6:47:17add in a new column that is the surname

6:47:20of string data type so I'm doing that by

6:47:23using the keyword alter followed by that

6:47:26table uh the table name is EMP table and

6:47:29I'm using the keyword add columns and

6:47:32the column name is surname and the data

6:47:34type of that column is string so now

6:47:36let's fire in enter and there you go we

6:47:39have successfully added a new row to our

6:47:41table now let's try to describe our

6:47:43table again and see if the column has

6:47:45been successfully added or not there you

6:47:49go you can see the last row which is the

6:47:51surname that we have added most recently

6:47:54so this is how you can alter the table

6:47:57and you can also change the names of the

6:48:00existing columns let's try to do that

6:48:02one as well now what I'm doing is I'm

6:48:05changing the column name to first name

6:48:09so one of the column name in my table

6:48:11EMP table is the name which gives me the

6:48:14names of the employees so since I added

6:48:17surname I'll change this column name

6:48:20from name to first name so this is the

6:48:24command that I'm using for that

6:48:26operation right now let's fire in enter

6:48:28and see the result yeah the change is

6:48:31been made Let's describe our

6:48:34table don't forget the semicolon there

6:48:38you go you can see that earlier we had

6:48:41name now it's been changed to first name

6:48:44and we also have a surname let's clear

6:48:46our screen so that's all for alterations

6:48:50now we shall move ahead into our next

6:48:52Main topic or the data model which is

6:48:56partitioning so we have dealt with the

6:48:58first two data models that are databases

6:49:02and tables so we have learned how to

6:49:04create a database and we have learned

6:49:07how to create a table we have learned

6:49:09how to create internal or manage table

6:49:12and also we have created external table

6:49:15and also we have learned how to create

6:49:16an external table in a particular

6:49:18location in your hive and load data to

6:49:21your table and also how to alterate your

6:49:24tables the column names the name of your

6:49:27table and how to add or delete new

6:49:30columns to your table so far so good and

6:49:32now we shall continue with the next type

6:49:35of data model that is the

6:49:37partitioning as we have discussed

6:49:38earlier about partitioning it's

6:49:40completely similar to a school or a

6:49:44college just imagine that you are in a

6:49:46college and you are in computer science

6:49:49section so a college has many branches

6:49:53so maybe computer science mechanical and

6:49:56electronics and Communications so

6:49:59imagine your name is Harry so if someone

6:50:02comes to your college and if is looking

6:50:04for Harry so there are many harri's in

6:50:06your school so if the person is asking

6:50:09specifically about you that is Harry

6:50:11from computer science then can you

6:50:14imagine how simple is this query so you

6:50:17don't have to search for electronics and

6:50:19mechanical you just have to come into

6:50:20the class computer science and search

6:50:22for for Harry and there you go you are

6:50:24present so that's how partitions work to

6:50:27execute commands or to execute queries

6:50:29on partition we will create a whole new

6:50:32database here let's start everything

6:50:34from fresh so we'll create a separate

6:50:37database for executing our new data

6:50:39model that is partitioning so I'm

6:50:40creating a new database that is Eda

6:50:43student so there you go the database has

6:50:45been successfully created followed by

6:50:48that let's use this database now to use

6:50:51the database you just need to to add in

6:50:53the keyword use and name of the database

6:50:56so let's fire in enter and now we are

6:50:59currently using Eda student database now

6:51:02let's create a table in urea student

6:51:05database so here I'm creating a normal

6:51:09table that is the manage table so inside

6:51:12my student table I'll be having uh some

6:51:14basic columns such as ID number of the

6:51:17student name of the student what is his

6:51:20age and course so you're not find

6:51:22finding course here because I'm going to

6:51:24partition the table based on course so

6:51:28here you can find the course I'm using

6:51:30the keyword partitioned and on what

6:51:33terms so on the terms of course I'm

6:51:35going to partition students so we have

6:51:38discussed about our students CSV file

6:51:41right so here we have our CSV file and

6:51:44the courses that this particular

6:51:46Institute is offering are Hadoop Java

6:51:49Python and yeah so these are the courses

6:51:52that this particular Institute is

6:51:54offering so I'm going to categorize or

6:51:57I'm going to partition these students

6:51:58based on their courses so this is how

6:52:01I'll be partitioning them using this

6:52:03following code so basically the table

6:52:06has all the columns and I'm going to

6:52:07partition the table using course so

6:52:10let's fire and enter and see the

6:52:11execution of this particular code the

6:52:13partition has been done now all we have

6:52:15to do is try to load in our

6:52:19data before that let's try to describe

6:52:22it let's try to see what are the columns

6:52:25present in our particular table student

6:52:28so as you can see the course column is

6:52:30present don't worry the code looks that

6:52:33we have missed out course but we did not

6:52:34miss the course column it is present in

6:52:37the table the only thing is that we have

6:52:40just partitioned it based on the course

6:52:42that we are going to offer now let's try

6:52:44to categorize the students based on

6:52:46their course so you can do that by using

6:52:49the following code we going to load the

6:52:51data use using the command load data

6:52:54local impath so this particular folder

6:52:57that is the student. CSV is in my local

6:53:00location so that is a home Cloud era

6:53:03desktop student. CSV and I'm loading the

6:53:06data present in this particular location

6:53:09into the student which is present in hi

6:53:11right now so I'm going to partition the

6:53:13student based on their course Hado now

6:53:15let's fire in this command and see the

6:53:17output yeah now you can see some map

6:53:19reduce jobs taking place yeah the data

6:53:22has been successful successfully

6:53:23loaded let's now refresh our Hive you

6:53:28can refresh your hive or hue based on

6:53:31two methods the first one is just

6:53:33clicking refresh button on the URL or

6:53:35you can also select an manual refresh

6:53:38this is the manual refresh and there you

6:53:40go it's done you can see the new

6:53:42database that is the urea student

6:53:44database that we have right now created

6:53:47and inside that you can see the student

6:53:48table that we have created and there you

6:53:50go we have the file of students based on

6:53:55course Hadoop now we will try to add in

6:53:58few more students based on the course

6:54:00Java for that all you need to do is just

6:54:03replace the course name with

6:54:05Java there you go here we had Hado

6:54:08course and now here we have Java course

6:54:11just fire and enter and you can see the

6:54:13output followed by that we also had

6:54:15another course that is python so let's

6:54:18also execute a code for that there you

6:54:20go python so now we have uploaded

6:54:23student details into our Hive and we

6:54:25have also partitioned using one of our

6:54:28data models that as partition into three

6:54:30categories that are based on Hadoop Java

6:54:32and python now let's go back to our Hue

6:54:35and see if the three categories are done

6:54:38or not yeah you need to refresh

6:54:41that there you go you have successfully

6:54:43refreshed still there is no sign of java

6:54:46and python maybe a manual refresh could

6:54:50help yeah the manual refresh has

6:54:52resulted in the two new files which are

6:54:56Java and python so you have all the

6:54:58three partitions here Hado Java and

6:55:00python just enter them and you can see

6:55:02the student details and now that we have

6:55:05understood partitioning sorry I forgot

6:55:08to mention we have two types of

6:55:10partitioning which are Dynamic

6:55:12partitioning and static partitioning so

6:55:15uh the static partitioning is in static

6:55:18or manual partitioning it is required to

6:55:20pass the values of partition columns

6:55:22manually while loading the data into the

6:55:25table hence the data file does not

6:55:27contain partitioned columns you can see

6:55:30that we have sent the partition columns

6:55:32manually for python Java and had but

6:55:36when it comes to Dynamic partitioning

6:55:37you just need to do it once and all the

6:55:40three files will be automatically

6:55:41configured and the files will be created

6:55:44so now what is U Dynamic partitioning so

6:55:47uh Dynamic partitioning the values of

6:55:50partition columns exist within the table

6:55:53so it is not required to pass the values

6:55:55of partition columns manually now what

6:55:58is this no worry we shall execute the

6:56:00code based on Dynamic partitioning and

6:56:02we shall understand this in a much

6:56:03better way now let's clear our screen

6:56:06now let's start fresh again let's try to

6:56:09create a new database

6:56:11for dynamic partitioning and let's start

6:56:14again fresh so here we'll be creating a

6:56:17new database that is edura student 2 so

6:56:21earlier we created Eda student and now

6:56:24we'll be testing our Dynamic

6:56:26partitioning on our new database that is

6:56:28urea student 2 so there you go the

6:56:30database has been successfully created

6:56:33now we should use this particular

6:56:35database currently we were in urea

6:56:38student to one database now we'll enter

6:56:40into student 2 database so we'll use it

6:56:43now now we are in Eda student 2 now

6:56:46before we start up with Dynamic

6:56:48partitioning we have to set High

6:56:50execution to Dynamic part is equals to

6:56:53true because by default the partitions

6:56:56that will be taking place in Hive will

6:56:57be static so we need to convert that

6:57:00into Dynamic Partition by specifying

6:57:02this particular code now we are good to

6:57:05go with Dynamic partitioning along with

6:57:07that we need to execute another command

6:57:09which says partition mode would be

6:57:12non-strict so by default when you are

6:57:15partitioning using the static partition

6:57:18the partition mode will be strict so now

6:57:20you're specifying it to be non-strict

6:57:23now let's execute this so there you go

6:57:25we have executed the two required codes

6:57:27for that now let's create a new table so

6:57:31the name of the table will be edura

6:57:34student that is edu St and this will

6:57:37have the same columns which are the ID

6:57:40of the student name of the student

6:57:42course age

6:57:44Etc now we will try to load in the data

6:57:47from our local path that is home cloud

6:57:49or our desktop student. CSV into the

6:57:51table edu stud so the data has been

6:57:55successfully loaded and the size is 267

6:57:58KB number of files is

6:58:02one now comes the tough part so here we

6:58:05are going to partition so we will be

6:58:07partitioning the table based on the same

6:58:09thing which is the course and we will be

6:58:12separating the data using the comma now

6:58:15let's fire and enter now the table has

6:58:18been separated based on course and now

6:58:22we will be loading the data to this

6:58:24particular table which is the student

6:58:26part so this particular table that we

6:58:29have created based on Dynamic

6:58:30partitioning and we are going to

6:58:32partition the data based on course now

6:58:35it's been created so the student part

6:58:37table has been successfully created now

6:58:39the only part remaining is to load the

6:58:41data to this particular table now we

6:58:44will be writing a code so using that

6:58:47code the map reduce will automatically

6:58:50segregate the data members or the

6:58:53students based on their courses so the

6:58:57guys which are in Hardo will be

6:58:58separated guys in Java will be separated

6:59:01and loaded into different file and

6:59:03similarly with python now let's see um

6:59:05how to do it using the

6:59:08code so there you go we are going to

6:59:11insert into student part partition based

6:59:14on course select ID name course age from

6:59:17the table at your so uh the data will be

6:59:20imported from the table what we have

6:59:22created here that is urea student so

6:59:25this particular location has the

6:59:28student. CSV file now let's fire in

6:59:31enter and see if it's created or not

6:59:34fine you can see some of the map produce

6:59:36jobs are getting executed you can see we

6:59:39have three jobs so first one is getting

6:59:42executed we have three because one is

6:59:44for Hado one is for Java and one for

6:59:49python so this will take a little time

6:59:52so this is the reason why I have chosen

6:59:55smaller CSV file so to save time when

6:59:58you take up the course from at your then

7:00:00you can work on realtime data so that

7:00:03you get hands-on experience from real

7:00:05time and you can get yourself placed in

7:00:06some good companies with the experience

7:00:09what you gain from this particular

7:00:10course so the stages have been

7:00:13successfully finished and the data has

7:00:14been loaded now let's see what are the

7:00:17datas present in the particular table

7:00:19student part there you go you have the

7:00:21output executed here so these are the

7:00:24data members present in the partition

7:00:26student part so these are the data

7:00:28members which are separated based on

7:00:29their courses that is the partition

7:00:31based on their courses that is had Java

7:00:33and python so now that we have

7:00:35understood Dynamic partitioning and um

7:00:38static partitioning we shall move ahead

7:00:41into the last type of data model which

7:00:44is bucketing Once after we finished the

7:00:47bucketing we shall enter into some query

7:00:50functionalities of hive or query

7:00:52operations which can be performed in

7:00:54Hive and followed by that we will also

7:00:56learn some functions which are present

7:00:58in Hive and some of the other things

7:01:01like group buy order bu sord bu and

7:01:04finally we shall wind up the session

7:01:06with joins which are available in Hive

7:01:09for now let's get continued with

7:01:11bucketing the last type of data model

7:01:14present in Hive so for that um let's

7:01:16again start fresh we shall create a new

7:01:18database for that before that let's go

7:01:21back to um H and check if our partition

7:01:24has been made or not let's

7:01:27refresh also let us make a manual

7:01:34refresh so our database Wasa student 2

7:01:39database and inside that we have the

7:01:42table that is student part and there you

7:01:44go you can see the files which are based

7:01:47on the partition so 22 is for a

7:01:50different course 23 is for a different

7:01:52course and 24 is for a different course

7:01:55and this is the default partition which

7:01:58has all the data members as we discussed

7:02:00earlier now let's start with the last

7:02:03data model in Hive that is bucket now we

7:02:06have created a new database that is Eda

7:02:09bucket now we shall also create a new

7:02:12table for that before that we need to

7:02:15start with this particular database so

7:02:16we can use the command use edura bucket

7:02:19now we are in Eda bucket now let's

7:02:22create a new table so the table name

7:02:25will be at youra bucket and it will be

7:02:28containing the ID name salary age of the

7:02:31employees the table is created now let's

7:02:34try to load the data so the data file

7:02:37that we will be using is the same one

7:02:39that is the employee. CSV so the data

7:02:41has been successfully loaded into the

7:02:44location now comes to the major part

7:02:47that is the bucketing part so to start

7:02:50bucketing in high we need to use the

7:02:51command and set hive. info. bucketing is

7:02:54equals to true so that's done now we

7:02:58will cluster or classify the data

7:03:00present in this particular file using

7:03:04this particular code so we will be

7:03:06clustering based on the ID and we will

7:03:09be categorizing them into three

7:03:10different buckets so let's fire in this

7:03:13command and see if it's happens yeah

7:03:15that's successfully done now we will

7:03:17overwrite the data using the following

7:03:20command now we'll be inserting data into

7:03:23this buckets that we have made that is

7:03:25three buckets and we will override the

7:03:27table using this particular code there

7:03:30you go you can see some map reduce jobs

7:03:32to be taken care of

7:03:35now so one MPP and red users are three

7:03:39for now so stage one is getting done so

7:03:43we should be having three task basically

7:03:47so let's see what's the output stage one

7:03:50is finished

7:03:52the process is finished and data has

7:03:54been successfully

7:03:55inserted now let's go back to Hive and

7:03:58check if it's done or not so before that

7:04:01let's do a

7:04:04refresh now a manual refresh would be

7:04:07much better there you go we have our

7:04:09database here which is edura bucket and

7:04:12inside edura bucket we have EMP bucket

7:04:16and that's our data employee. CSV there

7:04:20you go now let's move ahead and

7:04:24understand the basic operations we can

7:04:26perform in Hive so for that let's start

7:04:29fresh again let's create a new database

7:04:32I'm creating a new database for each and

7:04:34every option or each and every operation

7:04:36that I'm performing in this particular

7:04:38tutorial just to make things or keep

7:04:40things in a sorted manner so as you can

7:04:44see here in our particular file

7:04:47system I have separated each and

7:04:50everything like I have sorted everything

7:04:52so for bucketing I've got a separate

7:04:54database and for partitioning I've got a

7:04:57separate database and for understanding

7:04:59how to create database and tables I've

7:05:01got a separate database for that just to

7:05:03keep things arranged and sorted this

7:05:05looks uh in a much better way so now

7:05:08let's discuss about the operations that

7:05:10we could perform in h so I'm creating a

7:05:13new database again for this so the

7:05:15database would be Hive query language

7:05:19now let's use this particular database

7:05:21this creates a habit of learning things

7:05:24in a better way or it's like a revision

7:05:27for the things what you have performed

7:05:28or learned so far as you can see the

7:05:31table is been successfully created now

7:05:33let's try to add in some data into this

7:05:36particular location that is employee

7:05:38data it's been successfully loaded now

7:05:41let's try to see what are the details

7:05:43present in this particular file we can

7:05:45use in the command select star from the

7:05:47table urea employee so there you go

7:05:50these are the details or information

7:05:52present in the table meta employee now

7:05:56we shall see what are the functions that

7:05:58we can perform on this particular file

7:06:00so since we discussed that the

7:06:02mathematical operations and logical

7:06:05operations can be performed on H so

7:06:07let's try to perform an addition

7:06:09operation so I'm selecting the column

7:06:12salary and as we have seen here the

7:06:14salaries are 25 30 40 20,000 rupees for

7:06:19every employee now let me add in 5 ,000

7:06:21more to each and every employee so I'm

7:06:23adding uh the value 5,000 by using the

7:06:26addition operation so let's enter you

7:06:29can see we have added 5,000 so the first

7:06:32element was 25 now it's 30 so similarly

7:06:35all the other employees got 5,000 rupees

7:06:38hike all of a sudden now let's try to

7:06:40remove 1,000 so to do so all you need to

7:06:43do is uh replace the addition operation

7:06:47with a subtraction operation that is

7:06:49minus fire and enter and they go each

7:06:52and every employee lost 1,000 so the

7:06:55initial amount was 25,000 so removing

7:06:571,000 from that will result in 24 so

7:07:00this is considering the first initial

7:07:02values so this is how it's working uh

7:07:04followed by that let's also perform some

7:07:07logical operations let's clear the

7:07:10screen and yeah here I'm fetching for

7:07:12the employees who are having a salary

7:07:15equal to or greater than

7:07:1725,000 so these are the employees which

7:07:19are having the salaries above or equal

7:07:22to

7:07:2325,000 similarly let's execute another

7:07:25one which detects the employees with

7:07:28salaries less than 25,000 so we have got

7:07:31two employees which are having lower

7:07:33salaries which are Amit and chaitanya

7:07:35fine so this is how you perform some

7:07:38operations in height so now let's move

7:07:41ahead and understand the functions which

7:07:43you can perform on height so in the same

7:07:46way let's create a new database again

7:07:49and let's use this particular database

7:07:51that has five

7:07:54functions now let's create a table in

7:07:57this particular database so the table is

7:08:01employee function and it's

7:08:03created now let's try to load in the

7:08:06data yeah the data has been successfully

7:08:09loaded and now let's see if the data is

7:08:12correctly loaded or not yeah the data is

7:08:14loaded correctly now let's try to apply

7:08:18some functions in this particular data

7:08:20so the first first thing or the first

7:08:22function I'm going to apply would be a

7:08:24square root function where I'll be

7:08:26finding out square root of the salaries

7:08:28of the employees so there you go the

7:08:30square root of 25,000 was 58 do decimal

7:08:35numbers so this is how you perform some

7:08:38basic functions on your data now let's

7:08:40try to find out the maximum salary so

7:08:45yeah the job is getting executed you can

7:08:47see some map reduce chops here I think

7:08:50the biggest salary would be from sanun

7:08:54so the maximum salary is

7:08:5640,000 so this is how it

7:08:59works since we are working on cloud eror

7:09:02and the system configuration is limited

7:09:06the execution speed is a bit low but if

7:09:08you're working in real time then this

7:09:11process would take like few seconds and

7:09:13it's

7:09:18done there you go you have the value

7:09:2140,000 as shown here so 40,000 the

7:09:25employee name is sanana is the maximum

7:09:28salary so that's what we got here now

7:09:31let's try to find out the minimum

7:09:36salary so the minimum salary is

7:09:3915,000 and who would that be yeah it's

7:09:43chaitanya with minimum salary

7:09:4615,000 so that's how you do some

7:09:49operations in five let's execute some

7:09:52more operations such as converting the

7:09:54names of the employees to uppercase so

7:09:56you can see the employee names are

7:09:58converted to uppercase here and

7:10:00similarly let's try to convert to lower

7:10:04case so here you can see we have

7:10:06converted them to lower case so this is

7:10:08how you learn technology you need to

7:10:10play with the technology then you'll

7:10:12come to know the advantages and

7:10:14disadvantages so you can learn the

7:10:16possible ways where you can make things

7:10:18work out this is how you do it now now

7:10:21let's move ahead and understand Group by

7:10:23function in five so for that we'll be

7:10:25creating a separate database that is

7:10:28group now we will use this particular

7:10:30database that is group so we'll type in

7:10:33command use group semicolon now we will

7:10:37create a table so the table has been

7:10:40successfully created now we will load

7:10:42data into this particular table now we

7:10:44will use the new CSV file which will be

7:10:47employ 2. CSV now we are using this

7:10:51particular table because we have an

7:10:53additional column in this particular

7:10:55table which is the country column now as

7:10:58discussed before we will be grouping the

7:11:01employees based on Country let's see our

7:11:04data first so we have countries such as

7:11:07USA India UAE so these are the three

7:11:10countries that we are having in our CSV

7:11:12file so we will be categorizing the

7:11:14employees based on their countries so

7:11:17this is the particular command that we

7:11:18will be

7:11:20using

7:11:22so maybe I made an error while creating

7:11:25the table I think I gave a wrong table

7:11:29name here so let's drop our table so by

7:11:33mistake I gave different table that is

7:11:35employee order so to drop a table you

7:11:38just need to use the keyword drop and

7:11:41it's

7:11:41done yeah the keyword table was missing

7:11:45so you need to type in drop table and

7:11:47the table name and the table gets

7:11:49dropped so we were supposed to create a

7:11:51different table that is employee group

7:11:55so now let's create a new table that is

7:11:57employee Group Employee group has been

7:11:59created now let's try to add in data

7:12:02into the employee

7:12:04group so we have used employee 2 here

7:12:07because the employee 2 has another

7:12:09column which is based on Country so the

7:12:11countries that we are having here are

7:12:13India USA and UA so we will be using the

7:12:17group by function here and we will

7:12:19categorize the employees BAS based on

7:12:21their

7:12:23countries so there you go you can see

7:12:25some map reduce jobs getting

7:12:32executed yeah there you go we have

7:12:35categorized the employees based on their

7:12:39countries that as India UAE and

7:12:41USA and the sum of the salary so the

7:12:45guys is working in India and their

7:12:46sumission of the salary is 90,000 and

7:12:49similarly UA is nearly 1 lakh 5,000 and

7:12:53USA is 80,000 now let's also execute a

7:12:57different command based on Group by so

7:12:59here we'll be using Group by function

7:13:02and we will categorize based on the

7:13:04country as well as the summation of the

7:13:07salary which is greater than or equals

7:13:09to 15,000 so it's similar to the

7:13:11previous

7:13:18command so you can see the data got

7:13:20executed and we got the same output now

7:13:23let's move ahead and understand order by

7:13:26and sort by methods so for that we'll

7:13:28create a new database orders now we'll

7:13:31use

7:13:34orders now let's create a new table

7:13:36again so the new table is employee order

7:13:39and the table got created now let's load

7:13:42the data into this particular table by

7:13:44now I think you have some good practice

7:13:46of how to create a database how to

7:13:48create a table and how to load data into

7:13:50that particular table so the data got

7:13:53loaded and now we going to order the

7:13:56data present in this particular table

7:13:58based on the descending order of their

7:13:59salary so you're seeing some map reduce

7:14:02jobs going ahead so here we'll see the

7:14:05employees ordered based on their

7:14:07salaries in descending order so the

7:14:10highest salary will be at the first

7:14:11place and the lowest salary will be at

7:14:13the last

7:14:15place yeah so we have sanjana at the

7:14:18first position with 40,000 as the

7:14:21highest salary and she's working for UAE

7:14:24and we have cha with lowest salary

7:14:2715,000 working for India now let us also

7:14:32execute another command based on uh sort

7:14:35by so first we try to execute a command

7:14:38based on orderby now let's see the same

7:14:40output using sort by so basically both

7:14:43work in the same

7:14:48way so there you go we have sorted the

7:14:51records based on descending order of

7:14:53salary now that we have learned what are

7:14:56the various operations that can be

7:14:58performed in h that are the arithmatic

7:15:00operations logical operations and also

7:15:03some of the functions such as maximum

7:15:06minimum Group by order by sort by so

7:15:10these are the various operations and

7:15:12functions that you can perform And Hive

7:15:14now let's move ahead into the last type

7:15:16of operations that can be performed in

7:15:18Hive those are the joints so for that

7:15:22let's again create a new database so

7:15:25here I'll be creating a new database

7:15:26that is urea join and followed by that

7:15:30let's use this particular database now

7:15:32for that we need to use the keyword use

7:15:34and there you go we are in edura join

7:15:38now let's create a new table for

7:15:40that so the table will be EMP join here

7:15:44you can see that I forgot to mention

7:15:46semicolon so now the table got created

7:15:49now we shall load the dat data into this

7:15:51particular table so now I've created the

7:15:54first table that is employee table and

7:15:56I'm loading the employee data into this

7:15:58particular table now to perform join

7:16:00operations we always need two tables so

7:16:04in this particular database at urea join

7:16:06I've already created the first table

7:16:08that is employe join now let's create

7:16:11second table that is the department

7:16:14table which will be present in the same

7:16:16database so this particular table is a

7:16:20department table which will be having

7:16:21the entities that are Department ID and

7:16:24Department name now let's load the data

7:16:27of Department into this particular table

7:16:31so the data has been loaded so you can

7:16:33see the employee 2. CSV had the columns

7:16:36ID name salary age and Country and

7:16:38similarly the department. CSV has the

7:16:41entities which are Department ID and

7:16:44Department name so the department IDs

7:16:46are present here and the names are

7:16:48development testing product relationship

7:16:50ship and admin and ID support now we

7:16:53have created both the tables and we have

7:16:56created or we have loaded the data also

7:16:59now we have four different joints

7:17:02available in Hive they are in a joint

7:17:04left outter joint right outo joint and

7:17:08full outer joint now let's perform the

7:17:10first type of joint which is the inner

7:17:12joint so in inner joint we are going to

7:17:14select the employee name and employee

7:17:17department and based on the employee ID

7:17:19and Department ID we are going to

7:17:21perform the joint operation that is the

7:17:23first joint in the

7:17:26joint so you can see some jobs getting

7:17:29executed so the map reduce task

7:17:32successfully

7:17:36completed so um the first set of join

7:17:39has been successfully finished and the

7:17:40output has been generated now let's try

7:17:43out the second type of join that is the

7:17:45left outter

7:17:47join so the only difference is that

7:17:49we're using the keyword left outer join

7:17:52now you can see one of the job got

7:18:01started so you can see the output is

7:18:04beenin generated as well of the left out

7:18:06of joint now let's move ahead and

7:18:09understand WR out of joint so for WR out

7:18:12of joint you need to use the keyword WR

7:18:14out to join fire in the command and you

7:18:17can see um the jobs getting

7:18:19executed

7:18:25so you can see the output of right out

7:18:27join has been successfully executed or

7:18:30displayed now let's type in the last uh

7:18:33join operation that is full out join so

7:18:37here I'm using the keyword full outter

7:18:39join fire in the command and you can see

7:18:41it's getting

7:18:45executed so uh the output for full out

7:18:49join has been displayed here so this is

7:18:51how the join operations are executed in

7:18:54Hive so we have learned how to create

7:18:57database how to create table how to load

7:19:00data and the various data models present

7:19:03in Hive that are the tables databases

7:19:06partitions bucketing and after that we

7:19:09have also understood various operations

7:19:11that are the arithmatic operations

7:19:13logical operations and functions that

7:19:15can be performed in Hive such as square

7:19:17root and summation minimum Max maximum

7:19:21and after that other operations such as

7:19:24group bu sort by order by and also the

7:19:27joints that are possible in Hive which

7:19:30are inner joint left outer right outer

7:19:33and full outer so each and every

7:19:36operation that could be possibly

7:19:38executed in hi have been displayed in

7:19:40this particular tutorial and everything

7:19:42is sorted here in the base of databases

7:19:45and you can get all the details about

7:19:48this and you'll also get the code that I

7:19:50have used in the description box below

7:19:53and you can try it out and also if

7:19:55you're looking for an online

7:19:56certification and training based on Big

7:19:58Data Hadoop then you can check out the

7:20:00link in the description box below and

7:20:02during the training you'll get to have

7:20:04realtime hands-on experience with

7:20:06realtime data you'll learn a lot of

7:20:08things in the training and so far so

7:20:11good now we shall also discuss some of

7:20:13the limitations of Hive so Apache Hive

7:20:16limitations so Hive is not capable of

7:20:19handling real time data hi is capable of

7:20:22batch processing if you have to work

7:20:24with realtime data then you have to go

7:20:26with realtime tools such as spk and

7:20:29gafka so it's like how will actually

7:20:32take in the data for example imagine

7:20:35you're working on Twitter and you have

7:20:38one lakh commments on a particular post

7:20:41so if you had to process those one lakh

7:20:42comments you'll have to first load all

7:20:45those commments into Hive then you need

7:20:48to process it so while you're loading

7:20:50the data from Twitter to Hive you may

7:20:53also get a few more comments that you

7:20:55will be missed out so it's not

7:20:58preferable for Real Time Hive is

7:21:00preferable for only batch Moree

7:21:02processing so followed by that it is not

7:21:04designed for online transaction

7:21:06processing so online transaction

7:21:09processing is something which only works

7:21:11in real time so Hive cannot support real

7:21:14time processing so last but not the

7:21:16least High queries contain High latency

7:21:19yeah High queries take a longer time to

7:21:22process as you've seen I've have taken a

7:21:24smaller CSV file and the time consumed

7:21:26to process such a small CSV file was

7:21:29taking so long so yeah High queries

7:21:32contain High latency so these are the

7:21:34few important noticeable limitations of

7:21:39[Music]

7:21:42hi so this project is based on the

7:21:45e-commerce domain so let me give you an

7:21:48introduction to this project in context

7:21:51of one of the biggest names of

7:21:52e-commerce platforms none other than

7:21:55Amazon so if you have ever shopped from

7:21:58Amazon before which I presume you must

7:22:00have you must have seen something like

7:22:03this when you click on a product so

7:22:05you'll view the details of the product

7:22:08something like this and as you know most

7:22:11of the e-commerce organizations do not

7:22:13have any inventory so they tie up with

7:22:16different Merchants similar is the case

7:22:18with Amazon and Amazon provides the

7:22:21merchants or the sellers a platform to

7:22:24get connected to the buyers so when you

7:22:27click on the details of a product you

7:22:30can see that Amazon gives you something

7:22:33like this and you also find something

7:22:36like this there are 28 offers from this

7:22:39price and if you click on it you can see

7:22:41the name of the different sellers who

7:22:43are selling the same product at

7:22:46different prices and the prices they're

7:22:48offering are listed like this but you

7:22:51can see over here that by default Amazon

7:22:54has selected the appario retail private

7:22:57limited for this particular product so

7:23:00how does Amazon do that so it is

7:23:02actually based on a merchant rating

7:23:05system and as a platform as a e-commerce

7:23:08platform you have to ensure that you

7:23:11always display the product from the best

7:23:14Merchant in order to ensure quality

7:23:17because you don't want angry customers

7:23:19right right so it is very important that

7:23:22your customers are satisfied with their

7:23:24product so you have to choose from

7:23:26different Merchants for the same product

7:23:29in order to decide which Merchants

7:23:31product needs to get displayed by

7:23:34default and hence Amazon has a merchant

7:23:37rating system in order to decide that

7:23:39and this is what exactly we're going to

7:23:42build all right so we're going to make a

7:23:45merchant rating system similar to this

7:23:49so here here's the problem statement so

7:23:51there are multiple merchants selling the

7:23:53same type of products as you can see

7:23:56that Merchant 1 and Merchant six are

7:23:58selling the same shirt Merchant three

7:24:01and four are selling the same shoes and

7:24:04similarly five and one are selling the

7:24:06same pants and there are multiple other

7:24:09Merchants who are selling the same kind

7:24:12of products and you have to build a

7:24:15merchant rating system or the company

7:24:18wants to build a Merchant rating system

7:24:21in order to determine which Merchant

7:24:23sells the best product so that their

7:24:26product would be displayed by default

7:24:29and as a big data expert let's just

7:24:32assume that you are hired by the company

7:24:33as a big data expert you are assigned

7:24:36this task so this is now your problem to

7:24:39solve so the first thing your

7:24:41organization will give you before you

7:24:44start to do your work is the data set so

7:24:47this is the data set that you're going

7:24:49to get so this is is the transaction

7:24:51data set and has certain Fields like

7:24:53transaction ID customer ID merchant ID

7:24:57amp when the purchase was made the

7:25:00invoice number the amount and the

7:25:02segment of product that was bought you

7:25:05have another data set which is the

7:25:07merchant data so these are the details

7:25:09about the different sellers or the

7:25:11merchants so you have got your merchant

7:25:14ID their tax registration number the

7:25:16merchant name their mobile number start

7:25:19date email address State country pin

7:25:22code description longitude latitude the

7:25:25location basically so these are all the

7:25:28details about the merchant that you have

7:25:30in your data set so let me just show you

7:25:32the data set so this is the data set

7:25:35that is in your htfs right now so here

7:25:39is the transaction

7:25:41data so here is the transaction data

7:25:43which is a 2GB of file and the merchant

7:25:47data set which is 20 MB because as you

7:25:49know there are many transactions but a

7:25:52limited number of sellers or Merchants

7:25:55so that is why the size of the data set

7:25:57the merchant data set is quite smaller

7:25:59as compared to the transaction one and

7:26:02to tell you we haven't actually used the

7:26:04entire data present in the data set we

7:26:07have just selected a subset or a sub

7:26:10data set you can say because the

7:26:11original data set was quite huge and

7:26:14this was a demo project so just for your

7:26:16understanding we have chosen a sub data

7:26:19set we just took 2 GBS of data out of it

7:26:22all right and this is how it exactly

7:26:24looks like so this is the CSV file of

7:26:26the data set that we have this is the

7:26:28transactions data all right so this is

7:26:31the approach to solve so the first thing

7:26:33we'll do is that we'll segregate the

7:26:36merchants based on the price of their

7:26:38products and their sales so we will be

7:26:42segregating them into four categories so

7:26:46the categories are the merchants who are

7:26:48selling products that are below 5,000 or

7:26:52less than

7:26:53$5,000 there is one more category for

7:26:56merchants who sells products between

7:26:58$5,000 to

7:27:00$10,000 another category of merchants

7:27:03who sells products between $10,000 to

7:27:0620,000 and another category greater than

7:27:1020,000 or more than 20,000 we'll be

7:27:13using a simple logic to approach solving

7:27:15this problem so let's say that if there

7:27:18is a merchant who selling their products

7:27:21at quite a low price and you see if he's

7:27:24not making a good number of sales it

7:27:27means that he is not selling quality

7:27:30products because if a merchant who's

7:27:32selling the product at quite a less

7:27:34price and people aren't still buying

7:27:36from him it clearly indicates that his

7:27:39products are not up to the mark so it's

7:27:42a low rating for a merchant similarly on

7:27:45the other hand if you see a merchant who

7:27:47is selling their product at quite a high

7:27:50price and yet he has a very good number

7:27:52of sales so it clearly indicates that

7:27:55his products must be very good because

7:27:57despite of the higher price people are

7:28:00still buying from him so in that case

7:28:03obviously the rating of the merchant

7:28:05would be also very good right so this is

7:28:08a simple logic that we're going to use

7:28:10in order to rate our merchants or

7:28:12Sellers and you have three options to

7:28:15choose from so you have got aachi Hive

7:28:18which is a great analytical tool we have

7:28:21got Hadoop map produce and Apache Pig

7:28:25and today we will be choosing Hadoop map

7:28:27produce so map produce is the core

7:28:30component of Hadoop that process huge

7:28:32amount of data in parallel by dividing

7:28:35the work into a set of independent tasks

7:28:39map produce is the data processing layer

7:28:41of Hado it is a software framework for

7:28:44easily writing applications that

7:28:46processes the vast amount of structured

7:28:48and un structured data that is stored in

7:28:51your hdfs hdfs is Hadoop distributed

7:28:55file system in Hadoop map reduce works

7:28:58by breaking the data processing into two

7:29:01phases maap phase and the reduce phase

7:29:03and that is how exactly it gets us name

7:29:06map reduce so the map is the first phase

7:29:09of processing where we specify all the

7:29:12complex logic business rules reduce is

7:29:15the second phase of processing where we

7:29:17specify lightwe processing like

7:29:19aggregation or summing up the

7:29:23outputs but the question is why choose

7:29:26map ruce well I'll give you two reasons

7:29:28for it first is the custom input format

7:29:33now input format is something that

7:29:35defines how your input files are going

7:29:38to be split and read so in map ruce you

7:29:41can create your own custom input format

7:29:44instead of using the default input

7:29:46format this actually makes handling of

7:29:48your data quite easier because here we

7:29:51can create our custom input format for

7:29:53transactions and pass it as an

7:29:56argument then we have the distributed

7:29:59cachy so distributed cachy is nothing

7:30:02but it is a facility that is provided by

7:30:05map ruce framework to Cache files files

7:30:08like your text files archives jars Etc

7:30:12that is needed by your application let's

7:30:15understand this with an

7:30:17analogy just think about it that that

7:30:19there are three students sitting on a

7:30:22table solving chemistry problems and

7:30:25they have one periodic table so you keep

7:30:28the periodic table in the middle of the

7:30:29table so the students all the three

7:30:31students can refer from the periodic

7:30:33table to find to see the atomic numbers

7:30:36of different elements and solve their

7:30:38problems right so one periodic table

7:30:41everyone can refer to it and solve their

7:30:43own problem so this is what distributed

7:30:46cachier is so with distributed cachier

7:30:49you can put the data that will be used

7:30:51by your different data noes to refer in

7:30:54order to run map produce jobs so we'll

7:30:57be learning more about how to create

7:30:59your custom input format and how to use

7:31:01the distributor caching in the demo part

7:31:03all right so before that let us just

7:31:06understand how map produce exactly works

7:31:08so this is a sample map Produce job

7:31:11execution with an example so this is

7:31:14your input file this is a text input

7:31:17file so it contains some some words so

7:31:21first what will happen is that the input

7:31:23will get divided into three splits all

7:31:27right so I'm taking one sentence at a

7:31:29time so I have got three splits over

7:31:32here and this will distribute the work

7:31:35among all the different map nodes then

7:31:38we will tokenize the words in each of

7:31:40the mapper and give value one to each of

7:31:44the tokens or words so deer one bear One

7:31:48River one similarly here car 1 car One

7:31:51River one and now a list of key value

7:31:55pair will be created where the key is

7:31:58nothing but the individual word and the

7:32:01value is one so this is the key and this

7:32:03is the value and this will happen on all

7:32:06the three nodes so the mapping process

7:32:10remains same on all the nodes so after

7:32:13the mapper phase a partition process

7:32:15takes place where the sorting and

7:32:18shuffling happen happens here all the

7:32:20topples with the same key are sent to

7:32:23the corresponding reducer so all the be

7:32:26are together cars are together deer and

7:32:29River are together so after the sorting

7:32:32and shuffling phase each reducer will

7:32:34have a unique key and a list of

7:32:37corresponding values to that key for

7:32:40example bear 1 one car one one and one

7:32:43and so on now comes the reducing phase

7:32:47now each reducer counts the value which

7:32:49are present in the list of values so

7:32:52reducer gets a list of values which is

7:32:54one one for the key bear and then it

7:32:57counts the number of ones in the list

7:33:00and gives the final output as bare two

7:33:03similarly for car it's three deer Two

7:33:07River it will count 2 1 so two and

7:33:11finally all the output the key value

7:33:14pairs are then collected and written in

7:33:16the output files so this is your output

7:33:18file so it has just combined the result

7:33:21from different reducers and here is your

7:33:25final output so understood map produced

7:33:28with the classic example of the word

7:33:30count program now this is the generic

7:33:33execution flow of the map produ job so

7:33:37you have your input file over here so

7:33:40the data for map produce task is stored

7:33:42in input files and input files typically

7:33:44lives in the

7:33:46hdfs the format of these files is arbit

7:33:49while line based log files and binary

7:33:51format can also be used and you have a

7:33:54input format now input format defines

7:33:57how this input files are split and read

7:34:00it selects the files or other objects

7:34:02that are used for input an input format

7:34:04creates the input split so it logically

7:34:07represents the data which will be

7:34:09processed by an individual mapper one

7:34:12map task is created for each split and

7:34:15thus the number of map task will be

7:34:17equal to the number of in input splits

7:34:20the splits are then divided into records

7:34:22and each record will be processed by the

7:34:25mapper now let's talk about the mapper

7:34:28so mapper processes each input record

7:34:30from the record reader and generates a

7:34:33key value pair and this key value pair

7:34:36is generated by the mapper is completely

7:34:39different from the input pair the output

7:34:42of the mapper is also known as the

7:34:44intermediate output which is written to

7:34:46the local disk the output of the mapper

7:34:49is not stored on htfs as this is

7:34:51temporary data and writing on htfs will

7:34:54create unnecessary copies so then the

7:34:57mapper output is passed on to the

7:34:59combiner for further process the

7:35:02combiner is also known as the mini

7:35:04reducer so Hadoop map produce combiner

7:35:07performs local aggregation on mappers

7:35:09output which helps to minimize the data

7:35:12transfer between mapper and the reducer

7:35:15once the combiner functionality is

7:35:18executed the output is then passed to

7:35:21the partitioner for further work now

7:35:24partitioner comes on the picture if

7:35:25you're working on more than one reducer

7:35:28and here we have two reducers in the

7:35:30example so if you have one reducer you

7:35:33don't actually need a

7:35:35partitioner so the partitioner takes the

7:35:37output from the combiners and performs

7:35:41partitioning partitioning of output

7:35:43takes place on the basis of the key and

7:35:46then sorted so by hash function a key is

7:35:49used to derive the partition according

7:35:52to the key value and map produce each

7:35:54combiner output is partitioned and a

7:35:57record having the same key value goes to

7:35:59the same partition and then each

7:36:01partition is sent to a reducer so by

7:36:05using partitioner it allows to have an

7:36:07even distribution of the map output over

7:36:10the reducer so after that comes the

7:36:12shuffling and sorting part so now the

7:36:15output is shuffled to the reduce node

7:36:18the shuffle Shing is the physical

7:36:19movement of the data which is done over

7:36:22the network once all the mappers are

7:36:24finished and their output is shuffled on

7:36:27the reducer nodes then this intermediate

7:36:29output is merged and sorted which is

7:36:32then provided as an input to the reduced

7:36:34face now comes the reducer so it takes

7:36:37the set of intermediate key value pairs

7:36:39produced by all the mappers as the input

7:36:43and then runs a reducer function on each

7:36:45of them to generate the output the

7:36:48output the reducer is the final output

7:36:50which is stored in hdfs so if you have

7:36:53multiple reducers the result from

7:36:56different reducers will combine and that

7:36:58is going to be your final output which

7:37:00will be written into the

7:37:02hdfs so this was the map Produce job

7:37:06execution flow so we'll also be using

7:37:08the distributed cache so you'll have

7:37:11different data notes each data node will

7:37:13have their local copy of their data and

7:37:16if each of the data nodes needs to refer

7:37:18something something we will keep that in

7:37:20the distributed cachier and in this case

7:37:22we'll be keeping the merchants file all

7:37:24right so distributed cachier is nothing

7:37:26but think of it as a share drive right

7:37:29so if you have multiple users who wants

7:37:32to have access to one data set so you

7:37:35can just put it up in the share drive

7:37:36and all of your users can share the data

7:37:40use the same data right so this is what

7:37:42distributed cache is so applications

7:37:45specify the files via URLs to cach a

7:37:49via the job con so I'll be telling you

7:37:51about the job conf later in the demo

7:37:53section and the distributed cacher

7:37:55assumes that the file specified via URLs

7:37:58are already present on the file system

7:38:00at the path specified by the URL and are

7:38:04accessible by every machine in the

7:38:05cluster so the framework will copy

7:38:07necessary files to the slave node before

7:38:10any jobs are executed on that node and

7:38:13distributed cache tracks modification

7:38:16timestamps of the cache files so click

7:38:18clearly the cach files should not be

7:38:20modified by the applications or

7:38:22externally while the job is

7:38:25executing so how will it works in our

7:38:28case we'll be storing the data into htfs

7:38:31and we'll be executing map reduce

7:38:34program over that file so we'll store

7:38:36the merchant data in the distributed

7:38:39cache then we'll segregate the

7:38:41transaction data into categories such as

7:38:44less than 5,000 5,000 to $10,000 10,000

7:38:4820,000 and greater than $20,000 with the

7:38:52merchant ID then we'll use the merchant

7:38:55file from the distributed cache and map

7:38:58the merchant ID with the merchant name

7:39:01and at last we'll receive the output as

7:39:03the merchant name with date indicating

7:39:06the number of sales in different

7:39:09categories now let us talk about the

7:39:11code sections so the execution will

7:39:14start from the main method where we'll

7:39:16use the tool Runner so the tool tool

7:39:18Runner can be used to run classes

7:39:20implementing tool interface so it paus

7:39:23the generic Hadoop command line

7:39:25arguments and modifies the configuration

7:39:27of the tool then it will point to the

7:39:30run method which will point to the Run

7:39:33Mr jobs so here we are specifying the

7:39:36driver code so I'll be telling you about

7:39:37the driver code later on so next the

7:39:40execution will move to the mapper class

7:39:42which is the transaction mapper the

7:39:44framework first calls the setup method

7:39:48followed by the map method for each key

7:39:51value pair in the input split so in

7:39:54setup method we are loading the cache

7:39:56file and calling a method where we'll be

7:39:59resolving the merchant name from the

7:40:01merchant ID next in the map method we're

7:40:04creating the object of transaction which

7:40:06we will be using to catch all the fields

7:40:08of transaction first and then using the

7:40:11object of aggregate data we will create

7:40:14the segment of transactions as we

7:40:16discussed before the four segments less

7:40:19than 5,000 10,000 20,000 those segments

7:40:22at last using the merchant ID name map

7:40:24we will resolve or find out the merchant

7:40:27name the output of the map method would

7:40:31be the key which will be the combination

7:40:33of merchant name and date of sale while

7:40:35the value would be in the form of number

7:40:37of sales of different categories next

7:40:40the execution will go to the partitioner

7:40:43code where we'll have the get partition

7:40:46method which will send the records with

7:40:48the same key to the same reducer and at

7:40:51last the reducer code will execute which

7:40:53will aggregate the data with the same

7:40:54key and provide the output so I'll be

7:40:58showing you and explaining you all the

7:40:59codes involved over here in the demo

7:41:02part all right now let us move ahead so

7:41:05first let me take you through this

7:41:07transaction class where we are defining

7:41:10all the variables corresponding to the

7:41:12transaction file as you can see we have

7:41:15got the transaction ID customer ID

7:41:17merchant ID Tim stamp invoice number

7:41:20invoice amount and segment here so these

7:41:23are the fields in the transactions and

7:41:25we have created the variable for the

7:41:27same so next we're creating the getter

7:41:30and Setter methods for each of the

7:41:31variables so as to read the value of the

7:41:34field and we write the value of the

7:41:36field so as you can see here we are

7:41:39defining the method get so get segment

7:41:42we have got get segment here where we

7:41:44are returning the value of the field and

7:41:47the set meth method here the set segment

7:41:50where we're writing the value of this

7:41:52field so similarly we're doing it for

7:41:54all the variables as you can see here so

7:41:58we have the get and set for customer ID

7:42:02so we have the get customer ID method

7:42:04which Returns the value of the field and

7:42:06we have got the set customer ID which

7:42:09writes the value of the field so this

7:42:12similar for all the fields in our

7:42:15transaction data so similarly you can

7:42:17have the aggregate data class which we

7:42:20have used to create the categories of

7:42:22the product so here we have Fields like

7:42:25order below 5,000 order below 10,000

7:42:28order below 20,000 order above 20,000

7:42:31which are nothing but the different

7:42:33categories which we have defined earlier

7:42:35and similar to the transaction class

7:42:37here we are defining the getter and

7:42:39Setter method so we have got get total

7:42:41order method which Returns the value of

7:42:44total order and we have set total order

7:42:46method which is writing the value of

7:42:49this field total order so we have the

7:42:51same for all the different variables

7:42:54that we have defined in the aggregate

7:42:56data class the getter and Setter methods

7:42:59then we've got the aggregate writable

7:43:01class so first here we are creating an

7:43:04object of gson class so gson is

7:43:06basically used to convert Java objects

7:43:09to Json format next we're initializing

7:43:12the aggregate data object now we have

7:43:15two Constructors first one is the basic

7:43:17Constructor which is not taking any

7:43:19argument the second Constructor is

7:43:22taking aggregate data format as an

7:43:24argument and trying to initialize the

7:43:26aggregate data object next we have the

7:43:29getter method for the aggregate data

7:43:31which will return the aggregate data

7:43:33object after that we have the write

7:43:36method which will write the value of

7:43:38corresponding Fields using the getter

7:43:40method of the field like for order get

7:43:43order below 5,000 get order below 10,000

7:43:46order below 20,000 get order above

7:43:4920,000 then you have read Fields method

7:43:52which will'll be calling the seter

7:43:53method of each field to assign the

7:43:55values to the corresponding field and at

7:43:58last we are overwriting to string method

7:44:00which will convert the aggregate data

7:44:02object to Json and then return the Json

7:44:06so I hope you guys are clear with the

7:44:07custom input format so now let us take a

7:44:10look at the main Java file which is the

7:44:13merchant analytics job. Java so the main

7:44:17class is the mer Merchant analytics job

7:44:19class inside which all the jobs will

7:44:22reside so the execution will start from

7:44:25the main method so first let us go to

7:44:27the main

7:44:29method so here we are using toolrunner

7:44:32so toolrunner can be used to run classes

7:44:35implementing the tool interface it

7:44:37passes the generic Hadoop command line

7:44:40arguments and modifies the configuration

7:44:42of the tool so tool runner. run method

7:44:45runs the given tool after parsing the

7:44:47given generate arguments it uses the

7:44:50given configuration or builds one if

7:44:52null it sets the tools configuration

7:44:55with the possibly modified version of

7:44:58the conf here we are passing the

7:45:00configuration object Merchant analytics

7:45:03job object which is the main class and

7:45:06arguments which we will be providing

7:45:08while executing the job so in our case

7:45:11there are three arguments first is the

7:45:15path of the transaction file second is

7:45:17the path of the the merchant file and

7:45:19third is the output directory now we

7:45:22will execute the run method where we are

7:45:25returning the values of the Run Mr jobs

7:45:28method we're also parsing the arguments

7:45:31that is all the three parts that is the

7:45:33transaction Merchant and output

7:45:36directory to the Run Mr jobs method now

7:45:39let us see the Run Mr jobs

7:45:42method so here we have the driver code

7:45:46so first we initialize the conf

7:45:48configuration object and then we will

7:45:49initialize the control job object so

7:45:53Control job Class encapsulates A map

7:45:55Produce job and its dependency it

7:45:58monitors the state of the depending jobs

7:46:00and updates the state of this job and

7:46:02now we are creating the object of job

7:46:04class and we will define the properties

7:46:07of the job so first we have the set

7:46:09output key Class Property where we are

7:46:11defining the output format class of the

7:46:13key which is text class similarly we're

7:46:16defining the set out output value class

7:46:18for output format class of the value

7:46:21that is the aggregate writable class

7:46:24next we have set jar by class which

7:46:26tells the class where all the mapper and

7:46:28reducer code resides which the merchant

7:46:31analytics job in our case now we are

7:46:34specifying the reducer class which is

7:46:36Merchant order reducer and next we are

7:46:39providing the input directory so the set

7:46:41input the recursive method will read all

7:46:44the files from the directories

7:46:45recursively if we are providing the

7:46:47directory path so first we're adding the

7:46:50input path of the transaction file which

7:46:52is present in the argument zero then

7:46:54here we are also specifying the input

7:46:56format of the file and the mapper class

7:46:59that is the transaction mapper next

7:47:01we're talking about all the merchant

7:47:03file from the directory provided in

7:47:05argument one and adding this file to the

7:47:07distributed cache using the job. cache

7:47:11archive method and moving ahead we are

7:47:14setting the output directory path which

7:47:16is provided in argument two and we are

7:47:18also adding the timestamp as the

7:47:21subdirectory and at last we are setting

7:47:23the partitioner class that is the

7:47:25Marchant partitioner and then we are

7:47:28returning zero or one depending on

7:47:30whether the job has been executed

7:47:32successfully or not and next we will

7:47:34take a look at the transaction mapper

7:47:36class which implements the mapper

7:47:39interface so it Maps input key value

7:47:42pairs to a set of intermediate key value

7:47:44pairs so maps are the individual task

7:47:47which trans form input records into an

7:47:49intermediate record the transformed

7:47:52intermediate records need not to be of

7:47:54the same type as the input records the

7:47:56Hadoop map ruce framework spawns one map

7:47:59task for each input split generated by

7:48:02the input format for the job and mapper

7:48:05implementations can access the job conf

7:48:08for the job via the job configurable and

7:48:11initialize themselves the framework

7:48:14first calls the setup method followed by

7:48:16map method for each key value pair in

7:48:19the input split so in the setup method

7:48:22we are loading the merchant file from

7:48:24the cache using the get cache archives

7:48:27method then from each file we are

7:48:29calling the load merchant ID name in

7:48:32Cache so we're calling this method and

7:48:35we are passing the path of the cache

7:48:37files and the configuration objects now

7:48:40in this load merchant ID name in cach a

7:48:42method we are initializing the object of

7:48:45file system using the con and next we're

7:48:48opening the file and then we're reading

7:48:50the data line by line from the file now

7:48:53here first we're removing the codes from

7:48:56the line and then we're splitting the

7:48:58line using the comma and at last we're

7:49:01putting the merchant ID and Merchant

7:49:03name in the merchant ID name map

7:49:06variable so here you can see in the

7:49:08merchant file that we have merchant ID

7:49:11at index zero and Merchant name at index

7:49:142 so this merchant ID name map was will

7:49:17help us in resolving the merchant name

7:49:20from merchant ID and next we're defining

7:49:23exception to notify us if the cacher

7:49:26file is not read and at last we're

7:49:29closing the object of the buffered

7:49:32reader now let's talk about the map

7:49:35function so now the map function will be

7:49:38called so the input format of key is

7:49:41long writable and the value is text

7:49:44we're also creating a context of the

7:49:46mapper frame work where we will be

7:49:48writing our intermediate output now the

7:49:51output format of the key is text and the

7:49:54value is aggregate writable again here

7:49:57we are removing the codes from the line

7:49:59and then we are splitting the line using

7:50:01comma so in the split Arrow we have all

7:50:05the fields of the transaction data

7:50:07stored in the consecutive

7:50:09indexes now we are creating an object of

7:50:12transaction class and setting the values

7:50:14of the field using Setter methods and

7:50:17next we are creating the objects of

7:50:19aggregate data and aggregate writable

7:50:22class then using transactions get

7:50:25invoice amount field we are deciding

7:50:28that in which aggregate data field the

7:50:30transaction would lie we will set the

7:50:33value of corresponding field of the

7:50:34aggregate data object to one and next we

7:50:37have the output key which will contain

7:50:40the merchant name and the date of sale

7:50:42we will set the value of corresponding

7:50:45field of that aggregate data object to

7:50:48one and next we have the output key

7:50:51which will contain the merchant name and

7:50:53the date of the sale so merchant ID name

7:50:56map method will return the merchant name

7:50:59from the merchant ID as we just

7:51:01discussed so we are passing the values

7:51:04as merchant ID and the date at last

7:51:07we'll pass the intermediate result in

7:51:09form of key and value to the

7:51:12context next the result will be sent to

7:51:15the partitioner class that is the

7:51:17merchant

7:51:18partitioner so it's over here so in this

7:51:22class we're overwriting the default get

7:51:24partition method and in this method we

7:51:27are converting the key using hash

7:51:29function and using ABS method to return

7:51:33the absolute value of a number and at

7:51:35last we're using the modular function to

7:51:38get the remainder and now we're dividing

7:51:40it with the number of partitions so

7:51:42which is nothing but the number of

7:51:44reducers and in our case we have

7:51:47specified five reducers so the modulo 5

7:51:50would return the value between zero and

7:51:52four and one more thing is the same key

7:51:56would always have the same hash

7:51:58generated and hence the modular result

7:52:00would be also the same and thus the

7:52:03records with the same key will be sent

7:52:05to the same reducer and based on this

7:52:08records are sent to the reducer so next

7:52:11we have the reducer code and it's over

7:52:15here so it's the merchant order reducer

7:52:19the reducer Clause as defined in the

7:52:21driver code resides in the merchant

7:52:23order reducer Clause so here we have the

7:52:26key input as text and value input as

7:52:29aggregate writable which was written by

7:52:31our mapper class and the output key

7:52:34format is again text and the output

7:52:37value format is aggregate writable so

7:52:40here our execution will move to reduce

7:52:43method where we are passing the input

7:52:45key value and context as argument and

7:52:49here we are again creating the objects

7:52:51of aggregate data and aggregate writable

7:52:54class and next we are taking the input

7:52:57values now here we're calling the setter

7:52:59function of each category getting the

7:53:01earlier value of that category and then

7:53:04adding the new value of the new

7:53:06aggregate data object for that category

7:53:09so it will add the value to the

7:53:11corresponding Fields if there is a

7:53:13record with the same key and at last we

7:53:17writing the key and value in context.

7:53:20write

7:53:21method so I have explained you the code

7:53:24so now let us just go ahead and execute

7:53:33it so first let us move to the project

7:53:40directory so we have the palm. XML file

7:53:44which has all the dependencies that we

7:53:46require in order to run our map Produce

7:53:48job so this is the command so it has my

7:53:52jar file and the path of my jar file

7:53:56then I have got my main class over here

7:53:59which is Merchant analytics job so this

7:54:01is where my main function is and then

7:54:03I'm passing the three parts so first is

7:54:06my transactions. CSV this is my data set

7:54:09this is the path of my data set then my

7:54:11Merchant data this is the path of my

7:54:13Merchant data and finally my output

7:54:16directory which is the result this is

7:54:18the path over here so let us just go

7:54:20ahead and execute this

7:54:33command so the code is run so here are

7:54:37the different parameters on which this

7:54:39map Produce job was run so you can see

7:54:43the details over here so you can see the

7:54:46number of reduce task were five since we

7:54:49had five

7:54:51reducers so you can see all the details

7:54:53here let me just show you the result

7:54:56now so it is in the results directory so

7:55:01there are two directories over here

7:55:03because this was the earlier one that

7:55:05when I had previously executed it so

7:55:07this is the one that we have got right

7:55:10now so let me just show it to you so we

7:55:13have got five part files because there

7:55:15are five reducers so I'm just clicking

7:55:18on one

7:55:19part and you can just click on

7:55:22download all right let me just open

7:55:28it so this is what you get so you get

7:55:31the merchant name and the timestamp over

7:55:34here and also the category or the

7:55:37segregation that we did based on the

7:55:39cost of the orders right so it was order

7:55:42above

7:55:44$20,000 at this time stamp and the total

7:55:46order was one so this is the format of

7:55:49the result so we have got a lot of rows

7:55:53so this is the result so we have just

7:55:56used a few fields or parameters from the

7:55:58merchant file we have just used the

7:55:59merchant name and the ID over here since

7:56:02this is a sample project sample demo

7:56:04project but the scope of this particular

7:56:07project is huge you can use a lot of

7:56:09other parameters that was mentioned

7:56:11there like the location you can analyze

7:56:14it based on locations based on the time

7:56:17time period where the order was placed

7:56:19so you can take in account different

7:56:21fields and improve this or make this

7:56:24analysis even better by

7:56:26yourself so when you're doing this

7:56:28project as a part of your course

7:56:31curriculum so you will be exploring the

7:56:33other fields as well I have just shown

7:56:35you with just using two Fields the

7:56:37merchant name and the

7:56:40[Music]

7:56:44ID what is Kafka in general Kafka is a

7:56:49producer to the consumer based messaging

7:56:51system that has a producer that produces

7:56:53the message and the consumer that

7:56:55consumes the message in between the both

7:56:58we have Brokers that distribute the

7:57:00messages to the consumers and data

7:57:02storage unit which is none other than

7:57:04Apache zuker to understand more about

7:57:07Apache zuker and Kafka you can go

7:57:09through the article Link in the

7:57:11description box below Apache Kafka so

7:57:14basically Apache Kafka is an open source

7:57:17messaging tool developed by LinkedIn to

7:57:19provide low latency and high throughput

7:57:22platform for the real-time data feed it

7:57:25is developed using Scala and Java

7:57:27programming languages so followed by the

7:57:29definition of CFA we shall enter and

7:57:32understand what exactly is a stream in

7:57:35general a stream can be defined as an

7:57:37unbounded and continuous flow of data

7:57:39packets in real time data packets are

7:57:42generated in the form of key value Pairs

7:57:45and these are automatically transferred

7:57:47from the publisher there is no need to

7:57:49place a request for the same the below

7:57:51image depicts the key value pairs that

7:57:53are involved in data stream each and

7:57:56every single key value pair is one

7:57:59single unit of data or it is also called

7:58:01as one single unit of a record so

7:58:04followed by the stream we shall

7:58:06understand what exactly is a CFA stream

7:58:10CFA stream is an API that integrates

7:58:12Kafka cluster to the data processing

7:58:15applications which are either written in

7:58:17Java or Scala this API leverages the

7:58:21data processing capabilities of Kafka

7:58:23and increases data pism Apache Kafka

7:58:27stream can be defined as an open-source

7:58:29client library that is used for building

7:58:31applications and

7:58:33microservices here the input and the

7:58:35output data is stored in the form of

7:58:37Kafka clusters it integrates the

7:58:40intelligibility of Designing and

7:58:42deploying standard applications using

7:58:44the programming languages such as scale

7:58:47and Java with the benefits of Kafka

7:58:49server side cluster technology so this

7:58:53was the basic definition of kafa stream

7:58:56now let us understand Kafka stream API

7:58:58in a much better way through its

7:59:01architecture Apache Kafka streams

7:59:03internally use the producer and consumer

7:59:06libraries it is basically coupled with

7:59:08cfom and the API allows you to leverage

7:59:10the capabilities of Kafka by achieving

7:59:13data pism fall tolerance and many other

7:59:16powerful Fe features the following image

7:59:19depicts the basic architecture of Kafka

7:59:21stream here you can see the Kafka

7:59:24cluster which has the input streams and

7:59:26the output streams together followed by

7:59:29that we have Kafka streaming application

7:59:31or the API which takes care of the

7:59:33queries which are received from numerous

7:59:35applications which are connected to

7:59:37Kafka streaming application followed by

7:59:40this we have numerous components present

7:59:42in kfka stream architecture which are as

7:59:44follows they are input stream output

7:59:47stream instance we have two instances

7:59:51here which are stream instance one and

7:59:53stream instance 2 so inside every

7:59:55instance we have consumers as well as

7:59:58local state and Stream

8:00:01typology So in Kafka stream API input

8:00:04stream and output stream can be one

8:00:06single Kafka cluster followed by that we

8:00:09have a consumer which provides the input

8:00:11and receives the output and inside the

8:00:14stream instance we have stream topology

8:00:17and local state we shall understand

8:00:19about stream topology in a much detailed

8:00:21way in the next slide stream topology is

8:00:23all about the directed aycc graph or the

8:00:26steps in which the particular process is

8:00:28executed followed by that we have a

8:00:30local state local state is none other

8:00:33than a memory location which stores the

8:00:35intermediate data or the result provided

8:00:37by the stream topology these results are

8:00:39produced after applying various

8:00:41Transformations such as map flat map Etc

8:00:45so after the data is processed

8:00:47the tasks are united together and sent

8:00:49back to Output stream so this is how the

8:00:51architecture of cfast stream API works

8:00:54now let us understand more about stream

8:00:57topology so this particular diagram

8:00:59explains the stream topology here you

8:01:02can see the stream processor all the

8:01:04dots which are provided here are none

8:01:06other than stream processors and the

8:01:08line which is connecting them is the

8:01:10stream the stream is none other than the

8:01:13key value pairs of the data or records

8:01:16so basically the input is read from

8:01:18Kafka cluster first followed by that we

8:01:21apply various operators such as filter

8:01:23map join Aggregate and many more and

8:01:26finally we will receive the results

8:01:29which will be sent back to the output

8:01:30Kafka cluster so this is how the stream

8:01:33topology works now let us discuss the

8:01:36important features of Kafka streams that

8:01:38give it an edge over the other similar

8:01:40Technologies so the various important

8:01:43features of Apache Kafka streams API are

8:01:46plastic fa tolerant highly viable

8:01:49Integrated Security Java and Scala

8:01:52language support and exactly once don't

8:01:55worry we shall discuss each one of them

8:01:57in detail firstly we shall discuss about

8:02:01elastic nature Apache Kafka is an open

8:02:04source project that was designed to be

8:02:06highly available and horizontally

8:02:07scalable hence with the support of Kafka

8:02:10kfka streams API has achieved its highly

8:02:13elastic nature and can be easily

8:02:15expandable so this was the first feature

8:02:18followed by that we have the second

8:02:20feature which is about fault tolerance

8:02:23the data logs are initially partitioned

8:02:25and these partitions are shared among

8:02:27all the servers in the cluster that are

8:02:29handling the data and their respective

8:02:31requests thus Kafka achieves fall

8:02:33tolerance by duplicating each partition

8:02:35over a number of servers followed by

8:02:38Fall tolerance we have the next

8:02:40important feature that is highly viable

8:02:43since Kafka clusters are highly

8:02:44available they can be preferred to any

8:02:47sort of use cases regardless of their

8:02:49size they are capable of supporting

8:02:51small scale use cases medium scale use

8:02:54cases also the large scale use cases

8:02:57followed by highly viable feature we

8:03:00have Integrated Security Kafka has three

8:03:04major security components that offer the

8:03:06best-in-class security for the data in

8:03:09its clusters they are mentioned as

8:03:11follows they are encryption of data

8:03:15using SSL or or TLS followed by that

8:03:18authentication of SSL or

8:03:21sasl and finally the authorization of

8:03:25ACLS so followed by the security we have

8:03:28its support for top and programming

8:03:30language the best part of Kafka streams

8:03:33API is that it integrates itself with

8:03:35the most dominant programming languages

8:03:37such as Java and Scala and makes

8:03:41designing and deploying Kafka service

8:03:43side applications with ease followed by

8:03:46that we have exactly once processing

8:03:49semantics usually stream processing is a

8:03:53continuous execution of unbounded series

8:03:55of data or events but in the case of

8:03:59Kafka it is not exactly once means that

8:04:02the user defined statement or logic is

8:04:04executed only once and the updates to

8:04:08the state which are managed by SP or

8:04:10stream processing element are committed

8:04:13only once in a durable backend store so

8:04:16this is how Apache Kafka streaming API

8:04:18is considered to be having exactly once

8:04:21processing semantics so followed by the

8:04:24important features we shall go through a

8:04:26sample program based on Kafka streams

8:04:29API so this particular example can be

8:04:31executed using Java programming language

8:04:34yet there are few prerequisites on this

8:04:36one one needs to have Kafka and zeper

8:04:40installed in the local system and it

8:04:42should be running in the background if

8:04:44you have not installed zookeeper and

8:04:46Kafka in your local system then I have

8:04:48linked the article in the description

8:04:50box below which will explain you about

8:04:52the detail installation procedure of

8:04:53Zookeeper and Kafka in your local system

8:04:57once the Zookeeper and Kafka are

8:04:59installed into your local system you

8:05:00need to fire them up once the Kafka and

8:05:03zookeeper are successfully installed

8:05:05into your local system and they're

8:05:06running in the background you can go to

8:05:08Kafka and Define a producer topic and

8:05:11the consumer once the producer topic and

8:05:13consumer are defined you can come back

8:05:15to Kafka article and execute the

8:05:18following code in any of the Java

8:05:20editors the code will count the number

8:05:22of words that you have provided in your

8:05:24text document and you will receive the

8:05:25output as shown in the article here the

8:05:29text given to the code was welcome to

8:05:31Eda Kafka training and this article is

8:05:34based on Kafka streams these were the

8:05:36two sentences given to the program and

8:05:38the output is as shown below here the

8:05:41word welcome is repeated for once two is

8:05:43repeated for once Eda once once Kafka is

8:05:47repeated for two times and training is

8:05:49once this article is about streams so

8:05:53all these words are repeated for once so

8:05:56this is how exactly you should be

8:05:57receiving the output once after you

8:05:59execute the following code in your Java

8:06:01editor so followed by the example based

8:06:03on Kafka streams we can move ahead and

8:06:06understand the important differences

8:06:07between Kafka and Kafka streams so now

8:06:10the first difference is that in Kafka

8:06:13stream API single Kafka cluster can

8:06:16support as both consumer as well as

8:06:18producer while on the other hand in

8:06:21Kafka we need separate consumer and

8:06:23producer and Kafka considers consumer

8:06:26and producer as separate entities the

8:06:29second difference is that in Kafka API

8:06:32exactly once processing semantics are

8:06:34supported whereas in Kafka it is not by

8:06:38default but you can achieve exactly once

8:06:41processing in Kafka manually the third

8:06:45difference is that Kafka streams API is

8:06:47capable enough to perform complex

8:06:49operations whereas Kafka is designed to

8:06:52perform only simple operations the

8:06:55fourth difference is that Kafka API

8:06:57supports single Kafka cluster on the

8:07:00other hand in Kafka you need two

8:07:02different clusters for producer and

8:07:05consumer followed by that in Kafka API

8:07:09the code length is significantly shorter

8:07:12when you come into CFA the code length

8:07:14involved is highly lengthy

8:07:16the next difference between the both is

8:07:18CFA streams API can support both

8:07:21stateless and stateful networks what are

8:07:24stateless and stateful networks in

8:07:26stateless networks the client provides

8:07:29requests to the server and he gets

8:07:31instantaneous reply from server and here

8:07:34the cookies or the requests which are

8:07:36sent by the client are not stored

8:07:39whereas if you come into stateful

8:07:40Network the client requests the server

8:07:43along with some additional data which is

8:07:45requ required by the server in this case

8:07:48the cookies or the requests which are

8:07:50provided by the client are recorded So

8:07:53cfast stream API is capable to support

8:07:55both stateless and stateful networks but

8:07:58on the other hand Kafka is capable only

8:08:01to support stateless Network protocols

8:08:03followed by that the Kafka streams API

8:08:06can support

8:08:07multitasking whereas Kafka is not

8:08:10capable to support multitasking at a

8:08:13single task level followed by that Kafka

8:08:16stream API does not support batch

8:08:19processing whereas Kafka is capable to

8:08:22support batch processing Kafka stream

8:08:25API is all about real time so it doesn't

8:08:28have to support batch processing so

8:08:30these were the few important differences

8:08:32between Kafka streams API and Kafka now

8:08:36we shall move ahead and wind up the

8:08:38session with our last topic which are

8:08:40the important use cases based on Apache

8:08:42Kafka streams API Apache Kafka streams

8:08:46API is used in multiple use cases some

8:08:49of the major applications where streams

8:08:51API is being used are mentioned as

8:08:53follows firstly the New York Times the

8:08:56New York Times is one of the powerful

8:08:58media in the United States of America

8:09:01they use Apache Kafka and Apache Kafka

8:09:04streams API to store and distribute the

8:09:06realtime news through various

8:09:08applications and systems to their

8:09:10readers followed by the New York Times

8:09:13we have trvago trvago is the global

8:09:16Hotel search platform they use Kafka

8:09:19Kafka connect and Kafka streams to

8:09:22enable their developers to access

8:09:23details of various hotels and provide

8:09:26their users with the best-in-class

8:09:27service at lowest prices and finally

8:09:31Pinterest Pinterest uses Kafka at a

8:09:34longer scale to power the real-time

8:09:36predictive budgeting system of their

8:09:38advertising system with Apache streams

8:09:41API backing them up they have more

8:09:43accurate data than ever

8:09:46[Music]

8:09:51now when we talk about big data right

8:09:55the very basic questions lot of time

8:09:57pops up what are five BS available in

8:10:02Big Data can anybody answer that okay so

8:10:06RI want to answer this okay let me

8:10:08unmute RI RI over to you first is volume

8:10:13a volume size of data how much is

8:10:15growing day by day and next is vary VAR

8:10:19is we have three types of data actually

8:10:21structur unstructured and Serv structur

8:10:23data structured data is nothing but

8:10:25relation database all those things

8:10:27unstructured data is audio images files

8:10:31all this data semc is XML files velocity

8:10:36is how much fast is

8:10:39growing okay very very much good answer

8:10:43when Big Data started IBM gave a

8:10:45definition with just three V the three

8:10:48vs were volume variety velocity so what

8:10:53was in volume volume was when we talk

8:10:56about like in terms of amount of data

8:10:58what we are dealing with right for

8:11:00example today's Facebook is dealing with

8:11:02very huge amount of data right when we

8:11:04talk about variety now is it only

8:11:06Facebook which is generating data no

8:11:09right even Twitter is generating data

8:11:11okay we are talking about just social

8:11:13media sites no we can can we talk about

8:11:15medical dos yes in medical domains also

8:11:17the data is getting generated right lot

8:11:19of big data is getting generated so

8:11:21that's a different variety of data we

8:11:23have structur data unstructured data

8:11:25like we can have videos audio right this

8:11:28is basically going to be called as a

8:11:30variety third component is velocity like

8:11:33I said to you Facebook is just a 10 to

8:11:3512 year old company and imagine the

8:11:37growth they have made it in just 10 to

8:11:4012 years imagine each each user is also

8:11:42doing this activity posting video audio

8:11:45all all kind of charts right so with

8:11:48that how much data they are dealing with

8:11:50so that is you will be calling that

8:11:53velocity with the pace they have grown

8:11:55up from scratch to this level with the

8:11:58speed they have grown up to this level

8:12:01is called velocity now these were the

8:12:03three components which were actually

8:12:04going in the market for long actually if

8:12:06you ask me these three were the major

8:12:08components even go today but slowly they

8:12:11started realizing there should be a

8:12:12fourth category of data which is

8:12:14verocity which also makes sense in Big

8:12:17Data because what basically happens is

8:12:19that the data what we receive to us

8:12:21right the problem with that data is we

8:12:23cannot expect that the data is going to

8:12:25be always a clean data there can be a

8:12:28missing data there can be a corrupted

8:12:30data in Middle how to deal with this

8:12:32scenario how to deal basically with

8:12:35those scenarios because that is a

8:12:36component of now that big data right so

8:12:39basically that corrupted or the bad data

8:12:41what we getting how to missing data what

8:12:43we getting so those Cate also they

8:12:46decided to call it as verocity now

8:12:49people started calling that okay these

8:12:50four are the major component but we we

8:12:53are not yet done now they started adding

8:12:55few more components they started saying

8:12:57that no we are not going to stop with

8:12:59four when you are saying that veracity

8:13:01can be added why not value because the

8:13:03data what I'm getting I I want to know

8:13:06what is the value of data what how much

8:13:08important that data is now somebody said

8:13:10that I want to visualize the data so

8:13:12visualization is should also be one one

8:13:14of the we after somebody started saying

8:13:17that I I want to see the vocabulary of

8:13:19the data or the validity of the data now

8:13:22they started keep on adding their V but

8:13:24majorly if you talk about there are four

8:13:26V which carries some good value okay and

8:13:29usually in any interview they will not

8:13:30expect you to know all the V physically

8:13:32if you know it all good but they will be

8:13:34just expecting you to kind of understand

8:13:36that okay do you know at least four V

8:13:39which are important if you can answer

8:13:41that they will be all good okay so

8:13:43sometime to make it tricky they ask you

8:13:45find me to just see that how good you

8:13:47are in terms of thinking like I

8:13:49generally do that so when when I

8:13:51generally ask questions I generally see

8:13:53that okay that guy must be doing four P

8:13:55let me ask five P let me see that is he

8:13:57able to think little beyond what what he

8:14:00knows already moving for we just talked

8:14:03about something called structured and

8:14:04unstructured data right can anybody

8:14:07explain me the difference between them

8:14:09what basically are this uh structure

8:14:12data and unstructured data I let me add

8:14:14one more component to it semi structure

8:14:16data that third category of data let me

8:14:19add it now can anybody explain this plus

8:14:22give me the difference between them as

8:14:24well can anybody give me an

8:14:27answer so unstructured not easy to save

8:14:30the data into rbms that's the question

8:14:32answer from Nish okay structure data is

8:14:35basically in row and column format easy

8:14:37to read and pass from word okay uh n

8:14:42simar saying structure data rbms data

8:14:44unru is like log audio video sem

8:14:48structured is like XML

8:14:51Json yes so lot of people are giving the

8:14:55right answer here if we talk about

8:14:58basically structur data if you go with

8:15:01basically 1980s and all when this Oracle

8:15:04and IBM and all those companies came up

8:15:06into the market so if we talk about

8:15:081970s 1980 even at that time they used

8:15:12to have data but that data was not huge

8:15:14it was small small data but you will be

8:15:17surprised to hear at that moment it was

8:15:19still a challenge to deal with that data

8:15:21though it was having some sort of

8:15:23pattern it it used to have some sort of

8:15:25pattern and it's a small data now people

8:15:28used to think that how to use it how to

8:15:30manage it how to store it where to store

8:15:33it all those questions were coming in

8:15:34people mind and that is where companies

8:15:37like Oracle IBM and all came up into

8:15:39Market with their rdbm solution now they

8:15:42started delivering this solution that

8:15:45you can now store the data you can now

8:15:47process that data what what sounding

8:15:49like a having some pattern and you will

8:15:51be all good and today I need not tell

8:15:53you that today where these companies are

8:15:55like orle Microsoft you not in fact

8:15:57everybody of you must be willing to work

8:15:59for them if given a chance right so they

8:16:01they are basically now the market G and

8:16:04they have given the solution for for

8:16:05that it was going all good but now

8:16:08slowly what happens the other kind of

8:16:11data started coming so the data what

8:16:14they were dealing with was structured

8:16:16kind of data but with today world right

8:16:19like like Facebook came up in few years

8:16:21back right now as soon as Facebook came

8:16:23up into Market I'm just giving you an

8:16:25example now they you what you do in

8:16:27Facebook on Facebook you either upload

8:16:30video upload audio pictures right so you

8:16:33started dealing with this kind of data

8:16:35now do you think this kind of data can

8:16:37be dealt with my rdbm system answer is

8:16:40no right because now we cannot deal with

8:16:42this kind of data now we cannot call

8:16:45this data as structure data because this

8:16:47kind of data do not even have any sort

8:16:50of packing and that is where we started

8:16:54calling it as unstructured data mean any

8:16:57sort of data which do not have pattern

8:17:00kind of thing like your audios videos we

8:17:03started calling it as unstructured data

8:17:06now the third category of data is semi

8:17:10structured data right so there there are

8:17:12some sort of files for example let's

8:17:14talk about XML data so as soon as you

8:17:17see XML what is it sounding like does it

8:17:20have pattern or

8:17:22not does it have pattern or not XML

8:17:25files can I get the

8:17:27answer XML file Json files do they have

8:17:30pattern yes yes it it has pattern so as

8:17:34soon as I say that it has

8:17:36pattern first answer which must be

8:17:38coming in your mind should be that okay

8:17:40it is a structure data but now as soon

8:17:44as you tell me that it's a structure

8:17:45data my question for you is in that case

8:17:48can you do all the activities what you

8:17:50can do in rdbms to XML data I know you

8:17:53can do today even you can deal with

8:17:56unstructured data as well because they

8:17:57have introduced glob and clo data types

8:18:00as well but is that efficient can you

8:18:03deal can you fire Triggers on that can

8:18:05you do all the things which you can do

8:18:06with your traditional data no right now

8:18:10it started sounding to me that it's a

8:18:12unstructured data now I'm confused

8:18:15whether it's a structur data or

8:18:16unstructured data that's where they

8:18:18created a third category called as sem

8:18:21structure so that they can keep it in

8:18:23the middle the data which is sounding

8:18:25something of the structure type or as a

8:18:27unstructured type they started create

8:18:30they created a new category called as

8:18:32semi structure data everybody clear on

8:18:34this part what basically structure data

8:18:36semi structure data unstructured data

8:18:38okay moving further now I have another

8:18:41question for you how had do

8:18:45from your traditional processing system

8:18:49using

8:18:50rdbms can I get an answer what would you

8:18:53answer this part so these are like warm

8:18:55up kind of question usually in interview

8:18:57they will not start with the most

8:18:58complicated questions with right so

8:19:00these are kind of they kind of warm up

8:19:02they want to see your level of expertise

8:19:04how much you know so basically that is

8:19:06what is happening here so can you answer

8:19:08this part how do differ from traditional

8:19:12processing system using rdbm and friends

8:19:15I have a request rather than raising

8:19:17your hand please type it on chat window

8:19:20because this chat window I want to make

8:19:21it more interactive okay okay I don't

8:19:23know why your name is showing up Jo but

8:19:26let let's take it processing is done the

8:19:29data is done no input output okay Hado

8:19:33can store and process any type of data

8:19:37where R can store only relational data

8:19:41okay I can take this answer distributed

8:19:44storage processing okay good answer was

8:19:48par processing large data distributed

8:19:51know so few people have just started

8:19:53answering what is right don't answer me

8:19:55that I want difference right read the

8:19:58question properly the question St give

8:20:01me the difference do not tell me what

8:20:03this do and that do tell me the clearcut

8:20:05difference what are the differences what

8:20:07you notice with ad

8:20:09system or I can ask you in another way

8:20:12as well can had do replace at the system

8:20:16in future or maybe is it really

8:20:18replacing right now okay this question

8:20:21can be asked in this way as now can you

8:20:24answer me this part so I hope anybody

8:20:27who know how do Basics should be able to

8:20:29answer this easy can I get some answers

8:20:32now who want to come on air also can

8:20:34tell me I can unmute you if anybody want

8:20:36to come on air and

8:20:38answer both are complimentary to to each

8:20:41other okay good anyone who want to come

8:20:45here want to answer it this will give

8:20:47you good confidence also when you will

8:20:49be speaking in interview it cannot

8:20:53replace rdbms but I want reason that

8:20:55that's everybody know it cannot replace

8:20:57rdbms but what's the reason very good

8:21:00very good asset property is not

8:21:03supported or in other words can I say

8:21:06cred operation is not supported create

8:21:09delete update can I do that at Pro level

8:21:13no right so that

8:21:15is the major reason you cannot go with

8:21:18Hado systems like you cannot replace

8:21:21them also when the data is small okay if

8:21:25the data is small and it's a structur

8:21:28data which is going to be more efficient

8:21:30RMS or Hado

8:21:32systems yeah that's basically in the

8:21:34latest version they are supporting it

8:21:35but there are lot of distinctions in it

8:21:37that's all about Hado 3.8 stre which

8:21:41which is yet to come in the market

8:21:42properly so hold on until the time it

8:21:44come up because it it has lot of

8:21:46restriction it have right now lot of if

8:21:49you have already seen it you might be

8:21:50aware of

8:21:52this is used when we are more of right

8:21:56once read multiple times very good it

8:21:59follows warm principle W RM which means

8:22:04write once read multiple times okay so

8:22:09and one more important difference RS is

8:22:12free of C is RBM is free of cost no

8:22:17right so basically it's a license

8:22:19software you need to pay for it right

8:22:22but when it comes to it's completely

8:22:24open source there are companies who are

8:22:27now basically making money with this as

8:22:29this because it's an open source

8:22:31Community righto is an open source

8:22:32Community now if you get stuck where you

8:22:35will get the support there is nowhere

8:22:37you can get the support if you get stuck

8:22:39in rdb system there are companies to

8:22:41support you but what about had system if

8:22:43you get let's say some

8:22:45who will fix it for you right it's an

8:22:47open source so basically anybody can

8:22:49come and fix it but let's say if

8:22:50nobody's fixing your bug then what right

8:22:52so in that case what people are doing is

8:22:56they are taking support okay so a lot of

8:22:59companies are providing support also on

8:23:03this so a lot of companies are providing

8:23:06support in terms of now they making

8:23:09money from it they say that okay if you

8:23:11want to use use it we will give you the

8:23:14appropriate support what is required and

8:23:16you have to pay the money for us so a

8:23:17lot of companies actually came up like

8:23:20like this kind of idea that they are

8:23:22expert in a do and if you require any

8:23:24support we will help you with that so

8:23:26that is one thing which is happening now

8:23:28rdps can only deal with structur data

8:23:31right basically when you talk about

8:23:32rdbms you cannot deal with unstructured

8:23:35kind of data though you can right not

8:23:37deal with it as well because like

8:23:39somebody just argued with me she said

8:23:42that you know in the latest version of

8:23:44high that trying to support even C

8:23:45operation right but it it has lot of

8:23:48restriction it's not efficient at of the

8:23:50moment till the time it do not come out

8:23:53of the beta phase we cannot say anything

8:23:55about that similarly rdbms also started

8:23:58creating clock and block dat C O and B O

8:24:03right where what they say that you can

8:24:05now store unstructured data and can work

8:24:07on it but are there efficient answer is

8:24:10no right so similarly like most of the

8:24:13data what you deal with withm is going

8:24:16to be structured data but when it comes

8:24:18to youro system it can be unstructured

8:24:22sem structure as well as structure data

8:24:26right because if we take an example of

8:24:28he and all they deal with structur data

8:24:31right so that is one thing which is

8:24:34there rbms you work just on a single

8:24:36machine right so let's say you work on a

8:24:38single laptop where you have rbms

8:24:40install and working but when it comes to

8:24:43you are working in a distributed fashion

8:24:46right there will be multiple machines

8:24:48which can be involved in this case right

8:24:51and as I said in our mostly when the

8:24:53data is small your speed will be very

8:24:58fast your computation is going to be

8:25:01very quick at the same time with Hado

8:25:05your computation speed is not going to

8:25:08be that great okay with the small data

8:25:13it's not going to be that great with the

8:25:14par stock pitching into the market

8:25:16that's a different story now they are

8:25:18actually picking up basically because if

8:25:20your memory is good then you can

8:25:22actually make up a good speed but with

8:25:24redition ad when we talk about map

8:25:26reduce the speed are lower I believe if

8:25:29you have already done the classes of

8:25:31Maes you might have already noticed this

8:25:34thing right when you were doing that

8:25:36work count example right I hope

8:25:38everybody must have done V example in M

8:25:41right if you have taken the session

8:25:42right so in that case you might have

8:25:44seen the speed of work example that it

8:25:46is not that fast so that is one thing

8:25:50with Hado system if especially the data

8:25:53is smaller your speed is not

8:25:55comparatively with rdbms it's going to

8:25:58be

8:25:59slower now which brings me to another

8:26:02question can anybody tell me the

8:26:05components of had and their services in

8:26:09fact I'm showing you all the components

8:26:11can you explain these components what

8:26:15are the components available of do and

8:26:18what are the services what they

8:26:22provide can I get an

8:26:24answer very good so basically when we

8:26:27talk about hdfs right it is for your

8:26:30storage side okay very good can I get

8:26:34more

8:26:35answers I believe everybody must be

8:26:37knowing this part uh n is saying storage

8:26:41hdfs is for storage Yan is for proc

8:26:44processing cluster very good so you can

8:26:46say Yan is a cluster resource manager

8:26:50right there are a few more thing can I

8:26:53get more answers Name note manages the

8:26:56cluster data Note stores the data okay I

8:27:00can take this answer but partially they

8:27:03name note manages the cluster can you be

8:27:05little more explicit in this part Yan

8:27:09for resource allocation very good our M

8:27:12saying Yan to run map produce

8:27:15okay uh you can say to schedule map

8:27:18produce that would be better anyone

8:27:20right rather than saying to run map ruce

8:27:22can I say schedule map

8:27:24produce right that would be a more

8:27:26appropriate answer here but a good

8:27:28attempt name note has meta data very

8:27:33good right name node has metad data you

8:27:37can take it of something of this sort so

8:27:41uh in your real term scenario right let

8:27:44me go back you can simply relate it with

8:27:47this your real time life also right

8:27:49let's say in real time scenario what

8:27:51happens let's say you have a boss okay

8:27:54how many people have a kind of a smiling

8:27:58boss good

8:28:00boss smiling BS but a cunning BS very

8:28:03cunning he plays with your

8:28:07emotions when I say emotions I basically

8:28:10mean that basically he kind of is a very

8:28:12clever boss anybody who have it okay sh

8:28:17have it Aran have

8:28:20it is very interesting okay so let's say

8:28:24this boss okay let me draw a smiling

8:28:27boss okay The Smiling boss now usually

8:28:30every boss is smiling right they just

8:28:32keep on smiling but the the point to

8:28:34note here is the smile is cunning smile

8:28:37or what kind of SM now what next now

8:28:40these are people who are working right

8:28:42like like sh said he reports to such

8:28:45kind of boss right so basically VI said

8:28:48he report to such kind of Boss who a

8:28:50very clever boss right now let's say

8:28:54um narim is saying that okay fine his

8:28:57boss is also kind of a very elive boss

8:28:59but you know behind that inoc there

8:29:01might be lot of cleverness iting right

8:29:04behind the SC so let's say these three

8:29:06people are reporting to this clever boss

8:29:10what happens boss get products right

8:29:13boss get project let's get the project

8:29:16now what happens boss will be getting a

8:29:18project these three people are reporting

8:29:19to this to this boss right so what they

8:29:22will do so boss usually will distribute

8:29:24the project right so let's say the

8:29:26project was P he distributed into three

8:29:28parts P1 P2 and the third component He

8:29:32Made It P3 what he's going to do so he's

8:29:35going to keep this P1 or he going to

8:29:37come to sh and say that work on P1

8:29:39project he will come to we and say that

8:29:41work on P2 project he will come to

8:29:43nursing will say that work on Project

8:29:46right he will say that now all the three

8:29:49people are working properly given the

8:29:51project on timeline boss is going to be

8:29:53happy now imagine a scenario that we bch

8:29:58is boss not basically ditch but maybe he

8:30:02said that okay fine my boss was inocent

8:30:04so let me take an advantage of him

8:30:06telling him that you know I have a

8:30:07family mcy I can't work I want I have to

8:30:10take Le something of family mergency

8:30:12kind of thing now bosses in right

8:30:14because boss need to work on these two

8:30:16projects right P1 and P P3 who will

8:30:19deliver P2 project so that is the

8:30:20problem so boss what they do they come

8:30:23up basically with earlier kind of

8:30:26backups and boss is clever remember

8:30:28right so boss is going to call Shri in

8:30:31his cabin and say to Shri sh you know

8:30:34you're doing very very good job right so

8:30:37you're doing very good job I'm thinking

8:30:39to promote you if you keep on working

8:30:41like that you will get promoted very

8:30:42quickly in the market for sure

8:30:44and um so you should take up some senior

8:30:47responsibilities from now right so as

8:30:50soon as she will she will not hear

8:30:52anything she will just hear that you

8:30:53know boss is telling me promotion work

8:30:56okay he will just hear promotion word

8:30:57and he will be very happy in his mindset

8:31:00suddenly his boss will throw a Time B

8:31:02Because boss is clever so boss usually

8:31:04what he will say we just discussed right

8:31:06that you are going to take up some

8:31:08senior responsibilities so can you do

8:31:10one thing now can you basically take the

8:31:13backup of project now what happen in

8:31:16this case immediately as he said that

8:31:18can you take the backup of basically V

8:31:20project now he came back to S and

8:31:23started arguing you know I'm already

8:31:24busy but he said that you have to just

8:31:27take back up I'm not asking you to work

8:31:29right just take back up off work if you

8:31:30did me then only you have to work

8:31:32anywhere you are a senior candidate

8:31:34anywhere boss is work she will not be

8:31:36able to say no right you have to work

8:31:37for him similarly he will go to we he

8:31:39will do the same stuff right because

8:31:41managers usually tell this right

8:31:43everything is confidential so same thing

8:31:45he will do with SH also this is

8:31:47confidential do not think about it

8:31:49outside same way will call Vi tell the

8:31:51same story and he will ask now to take

8:31:54the backup of nursing project right so

8:31:56he's backing up this project similarly

8:31:59nursing if you note that nursing have to

8:32:01back up she project he did the same

8:32:03thing with nursing also and basically

8:32:05now nursing have to back up this

8:32:07basically D now if you not in this

8:32:10situation if we to have emergency leave

8:32:13now B will not face any trouble right

8:32:16now boss will not face any trouble

8:32:18because the work who has to do the extra

8:32:20work in fact who will be sad in this

8:32:23scenario definitely boss is not going to

8:32:25be sad but who is going to be sad in

8:32:27this scenario definitely she is going to

8:32:28be sad right so definitely she is going

8:32:30to face the heat of doing the extra work

8:32:33now similarly why I'm telling you all

8:32:35this because all the components you can

8:32:37relate here so basically the in when we

8:32:40talk about Hado Hado is not much

8:32:42different with this what Hado do is and

8:32:45boss is also keeping one more

8:32:46information right so boss is also

8:32:48keeping information of let's say what

8:32:50all projects he have who is working on

8:32:53what project let's say Shri is working

8:32:55on Project P1 and he have also a backup

8:32:58of P2 so all these details what boss is

8:33:00keeping when we talk about had Hado is

8:33:03kind of doing that exactly the similar

8:33:05kind of stuff but when we talk about

8:33:07boss so first thing is now in had word

8:33:09every human is going to be replaced by

8:33:11machines now the first compon what we

8:33:14were talking about right so if we talk

8:33:16about the first component it was name

8:33:18the boss is representing n second

8:33:23component was Data node these employees

8:33:26who are working basically for that boss

8:33:29you can represent them as data node the

8:33:32node where you are doing all the

8:33:35processing because employee do the work

8:33:37right employee do the work boss only

8:33:39instructor manage right same story here

8:33:41so these are your data third property

8:33:44what is this metadata right even name

8:33:47node is going to keep all the data about

8:33:49data that is called your metad data now

8:33:52what is this this basically secondary

8:33:54name node and all those stuff so

8:33:56basically we require some like this is

8:33:58the backup for the data part right for

8:34:00this data file P1 P2 P3 but what about

8:34:04this boss back so basically we want to

8:34:06create some backups for that so that's

8:34:08the reason we keep like casting name Lo

8:34:10and all those stuff there okay so

8:34:12basically that is the back part that's

8:34:13one of the component now that that is

8:34:15means like let's say in your company

8:34:17they have also back up a boss like in

8:34:20case if this boss leave then I should

8:34:22have all the details so that's that's

8:34:23what the backup me there now there is

8:34:25one more component called as load

8:34:27manager resource manager what are these

8:34:30things now what basically a boss is

8:34:32doing right what is boss doing boss is

8:34:34having the skill set to schedule the job

8:34:36right he only decided where to send what

8:34:39all those details right so that you can

8:34:41call it as a part of your resource maner

8:34:43kind of scheduling the work who is

8:34:45helping you to schedule the work you can

8:34:48call it as a resource manager right boss

8:34:51have that skill set similarly in your

8:34:53Hado what your resource manager is going

8:34:55to shule everything up now the the last

8:34:58part right node manager now do you think

8:35:01you will be able to work on this project

8:35:02without any skill set no right you

8:35:05require some skill set to work on that

8:35:07project right so that skill set you can

8:35:09relate it like the node manager which is

8:35:12managing your own no

8:35:14means you can delete it like your skill

8:35:15set which is helping you to solve a

8:35:17project right so same way the node

8:35:20manager you can say that it is managing

8:35:23the whole node it is kind of helping the

8:35:25node to execute the task that you can

8:35:28call it as node manager so this is how

8:35:32you can relate everything I hope with

8:35:34this example now it will make your life

8:35:36easy to remember all these components

8:35:38because this is a very important

8:35:40questions in basically in interview

8:35:42questions they generally ask question

8:35:44what are the configuration files what

8:35:45what basically are the main components

8:35:47so you should be aware of this and

8:35:49that's the reason I've explained you

8:35:51with this anology so that you get some

8:35:53idea and you can relate to it that if

8:35:55you have to explain in interview you do

8:35:57not remember all this stuff okay let's

8:35:59move

8:36:00further now so can I ask this question

8:36:03now what are the main had configuration

8:36:06files right so basically now we talking

8:36:08about configuration F this is related to

8:36:10mostly Administration interviews now

8:36:13when you talk about H administrator

8:36:15right so there will be few files which

8:36:17you need to configure so I hope there

8:36:19are people in this batch in this session

8:36:22also where people must have done some

8:36:24Hado administrator first right so I'm

8:36:26assuming that you people know that

8:36:27people who do not know that that's let's

8:36:29not worry about this that's the reason

8:36:31I'm answering this directly so there are

8:36:33few important FES here one is had

8:36:36environment. this is where you kind of

8:36:39mention all your environmental variables

8:36:41for example where is your Java home

8:36:43where is the Hado form all those things

8:36:46to Define in Hado

8:36:48.s for.xml so these this file basically

8:36:52Define where your let's say your name Lo

8:36:54is going to run right so you need to

8:36:56tell the address of your name Lo where

8:36:58you want to run that maybe you want to

8:36:59run a some machine at 9,000 F so you

8:37:02will be telling all that in your for

8:37:04site XML when we talk about hdfs XM here

8:37:09we talk about what should be the

8:37:10replication Factor where should

8:37:12physically my data Note should be

8:37:14present where physically my name note

8:37:16should be present all those things to

8:37:18Define in hdfs site. XML Yan site and M

8:37:24site basically defines the map jobs

8:37:27right what kind of cluster you are going

8:37:29to use are you going to use let's say

8:37:31Missour or you going to use Yan or you

8:37:33going to run a local distributed one all

8:37:36those things will be defining here also

8:37:38you will be defining where your resource

8:37:40manager should be running it should be

8:37:42running on this machine know 9 9,1 port

8:37:45or whatever Port you want to Define

8:37:47right so those information you will be

8:37:50defining in J site or map rite. XML not

8:37:56last two files are Masters and slaves

8:37:59file in Master's file we usually mention

8:38:03where my secondary name note would be

8:38:06running when I say secondary name it's

8:38:08it's like a backup not exactly I should

8:38:10call it as a backup but I should call it

8:38:12as a snap snapshot of the name it's

8:38:15something like this like somebody just

8:38:17copying the metadata that's it it's not

8:38:19going to become active as soon as the

8:38:21main Lo is down it is copying the data

8:38:23so that if name Lo is down at least I

8:38:25should have a backup that's it slips

8:38:28very clear with the name where are all

8:38:31my data notes what all machines are

8:38:33going to be my data notes that thing we

8:38:35Define in my slaves machine so people

8:38:38who SL file sorry so basically people

8:38:41who have done this administrator you

8:38:43must have basically played with this

8:38:45file these are the major Hado

8:38:46configuration F there are others as well

8:38:49right like hyp site XML there are others

8:38:52like HB site XML now these are very much

8:38:55kind of tool specific so that's the

8:38:57reason they will not be called as the

8:38:58Maino configuration P when somebody say

8:39:01Maino configuration P your answer would

8:39:04be these seven PES you need to remember

8:39:06these seven PES basically these seven

8:39:08files are the one which you will mention

8:39:11if you are going for administrator

8:39:13interview expect good number of

8:39:16questions in this okay they will ask you

8:39:19to explain each and every file how

8:39:21basically you will be what you do in

8:39:23which file which I just explained you

8:39:24make a note of that and that is

8:39:26definitely one of the favorite questions

8:39:29of the interviewer when you go for ad

8:39:32administrator kind of grow moving

8:39:35further now let's talk about some hdfs

8:39:39questions in terms of hdfs now my

8:39:42question is hdfs stores data using

8:39:45commodity Hardware which has higher

8:39:49chance of failure which is obvious right

8:39:52because my laptop can be one of the data

8:39:54Note right now definitely my data Note

8:39:57can fail at any moment now so how hdfs

8:40:01ensures the fault tolerance capability

8:40:05of the system can anybody answer this

8:40:09very good word can I get more

8:40:12answers

8:40:13I just answered it beforehand itself

8:40:16remember that boss and employe

8:40:19relationship what that boss was doing

8:40:22boss was keeping backup right boss was

8:40:25creating backup similarly Hadoop also

8:40:28create backup right that backup in Hado

8:40:32word is called replication right so

8:40:35basically if you have let's say one

8:40:37block P1 you are creating one backup or

8:40:40two backups of basically that block P1

8:40:42and that is called replication this is

8:40:45how Hado is ensuring that in case of any

8:40:50failure also there should be no mistake

8:40:54okay you should not be using data

8:40:56because in case if one machine fails

8:40:58also it's okay it will all work fine for

8:41:00me so that is what going to happen so

8:41:03very good A lot of people have given me

8:41:05the right answer in this case so block

8:41:08replication is the answer in this case

8:41:11as you can see in this example also like

8:41:13block one is replicated three times if

8:41:15you notice here right so block one is

8:41:18replicated three times similarly if you

8:41:20notice block two is also replicated

8:41:22three times so these are four different

8:41:24machines and I have replicated block one

8:41:27block two block three block four block

8:41:30five okay so this is what is happening

8:41:32edit log and Fs image is used to

8:41:34recreate a image FS image and edit log

8:41:38are two different things these these

8:41:40basically are two different things you

8:41:42cannot create the application if you

8:41:44want to know about that let me just

8:41:45answer you here so what happens what is

8:41:48phasically This Ss image and edit logest

8:41:51now what happens basically you have a

8:41:54name not right you have a name note

8:41:56where you keep the data in name not

8:41:59anybody have answer where you keep the

8:42:01data in name node no not ACC F initially

8:42:05where you keep the data in name node and

8:42:07not SDF I'm asking in name note this can

8:42:10also be an interview question very good

8:42:12answer

8:42:13we keep it in memory why let answer this

8:42:17let's say there is one client came up

8:42:19this is one client this is Cent C2 this

8:42:22is Cent C3 let's say there are multiple

8:42:24clients okay now what is happening in

8:42:27this case is let's say if this was my

8:42:30name node and let's say my data in name

8:42:32node is my metadata is let's say setting

8:42:34in this what would be the problem here

8:42:37let's say client one came and want to

8:42:39access some data what is going to happen

8:42:41this data will be because any processing

8:42:44which need to happen right that happens

8:42:46in memory only right now this data will

8:42:49come to the memory once this memory work

8:42:52will be over then what will happen then

8:42:55basically again it will come back to

8:42:57this right it will remove it from memory

8:42:59now don't you think there's a input

8:43:01output operation happening an input

8:43:04output operation is always expensive

8:43:06right now imagine if there are multiple

8:43:09clients asking at the same time to the

8:43:12name don't you think there will be too

8:43:14many input output operation every time

8:43:16you need to basically bring the block to

8:43:18the memory and then basically do the

8:43:20stuff and give the output now this is

8:43:22something which we want to avoid in

8:43:25order to avoid what they came up with

8:43:27the idea is that whatever you are going

8:43:30whatever metadata you are going to

8:43:32create should directly be created and

8:43:35kept in memory okay that is what the

8:43:38idea they came up they are not going to

8:43:39keep any data in the dis right now

8:43:41should be directly kept in the memory

8:43:43which brings another question for you in

8:43:46a serious question for you now as since

8:43:50you are telling that the data what you

8:43:52can keep in memory but my Ram is

8:43:55volatile when I say volatile I will lose

8:43:58I can lose the data at any moment right

8:44:00that's very obvious I can lose the data

8:44:02at any moment because my Ram is always

8:44:04going to be volatile right now I restart

8:44:07my system my Ram data is gone right so I

8:44:10will lose all the metadata in that case

8:44:12how I I will ensure that I should not

8:44:14lose my metadata now what they started

8:44:16doing is okay fine I will create

8:44:19everything in name note memory only but

8:44:22what I am going to do is whenever any

8:44:26like what I'm going to do add some

8:44:28interval of time at some interval of

8:44:32time I will keep on taking a backup of

8:44:36that metad data in dis okay in this I

8:44:41will keep on taking the backup of that

8:44:44data and whatever backup you are taking

8:44:48in the disk from memory is called your

8:44:51FS image now don't you think this FS

8:44:55image is going to be big right this FS

8:44:58image is going to be big now what

8:45:01usually happens is usually let's say

8:45:03today my FS image is version FS1 now

8:45:07usually the backup what they take is

8:45:09every 24 hour just give me one minute

8:45:13if any pop up kind of thing comes up it

8:45:16actually stop my system now anyway now

8:45:19what basically is going to happen so

8:45:21what happens is every 24 hours we can

8:45:25usually do this okay so basically today

8:45:28we did some backup tomorrow again this

8:45:30metadata is going to do backup now which

8:45:32brings another problem the problem now

8:45:35again would be let's say what about in

8:45:3818th hour or maybe in 23rd hour my

8:45:41machine P or my R in that case I'm going

8:45:44to lose the 23 data right which is again

8:45:47not good what should I do for that now

8:45:49for that what they came up is that let's

8:45:52create whatever activity is happening

8:45:55here I will keep on writing in a small

8:45:58file that will be created for let's say

8:46:0124 hours okay and that file is called as

8:46:05edit log then what going to happen for

8:46:0924 hours whatever activity you are doing

8:46:12will be be getting stored as a edit law

8:46:15Okay now what's going to happen after

8:46:17every 24 hours this ss1 plus edit log is

8:46:21going to be added up and FS2 will be

8:46:24created now in this scenario even if I

8:46:27lose the data in 23rd L my edit log will

8:46:30be having the data and that's how I am

8:46:32ensuring that I'm not losing any data

8:46:37now are you clear about this edit log

8:46:38and Fs image question I usually see that

8:46:41people are very confused with this logic

8:46:44that what is FS image they just kind of

8:46:46mg up and come back and tell that you

8:46:48know I know FS image and edit what what

8:46:50exactly are they I have seen people

8:46:53actually kind of confused with this so I

8:46:55hope you should be very clear now on

8:46:56this in which file can this back of time

8:46:58interval be configured so basically

8:47:00wherever the physical location of main

8:47:02Ro you have configured and where you

8:47:04configured I just told you what in hdfs

8:47:07s. XML right so wherever you have

8:47:09configured that so there will be a name

8:47:11directory in it in that name directory

8:47:14there is another subd directory called

8:47:15as current directory in that current

8:47:17directory there is another directory

8:47:19called as SN and directly which is

8:47:21secondary name directory there we keep

8:47:24this SS image and edit Lo clear about

8:47:27this part so this is where we basically

8:47:29keep that so can you please summarize

8:47:32the answer once sure the same answer of

8:47:34FS image and edit log but correct

8:47:38correct not exactly two files one file

8:47:40will be kind of very big file so that's

8:47:42the reason we are creating a smaller

8:47:44version of that F called so that every

8:47:4724 hours activity we can edit logs are

8:47:50stored in the disk as yes in the name so

8:47:53it's kind of act like a back that's it

8:47:56it's a very good interview question

8:47:58that's the reason as soon as this

8:47:59question came up I thought to answer it

8:48:01up though it's not a part of this SL but

8:48:03this is a very famous inter question

8:48:05that can you explain this suim Ed BL and

8:48:08I can tell you that most of the people

8:48:10fail to explain this here memory is R

8:48:13correct correct so it's just metadata

8:48:16which is store correct correct it is

8:48:18just creating a backup of that metadata

8:48:20for the 24 hours activity that means

8:48:23edit log is getting erased and getting

8:48:25and get new data yes yes every 24 hours

8:48:29it just keep on working and kind of er

8:48:31the data that's what keep on happen okay

8:48:33or it creates a new version it depends

8:48:35how your admin have configured that what

8:48:38if the block data goes more than the mem

8:48:43now in that case there is something

8:48:44called as pill usually it's not it do

8:48:47not go like that but there is some

8:48:49concept called as pill so in that case

8:48:51there will be some input output

8:48:53operation happening you have to deal

8:48:54with that so then you are making your

8:48:56name not slow so you have to make sure

8:48:58if you should have a good configuration

8:48:59but if you do not have it then in that

8:49:01case you have to do input output

8:49:02operation no other option then you have

8:49:04to keep in the dis and then the data

8:49:06will be having input outut the window is

8:49:08of 24 backup can be change yes it can be

8:49:11change to summarize this what we just

8:49:14talked about so in name note uh in name

8:49:17note basically you will be storing all

8:49:20the data but the problem is my Ram is

8:49:22going to be volatile now because of that

8:49:25I want to definitely want to have a

8:49:27backup now we keep a backup in the disk

8:49:30and that whatever backup we keeping we

8:49:32call it as FS image now FS image backup

8:49:35is always taken in 24hour slot now the

8:49:38another problem started with this what

8:49:40happen if I lose the data and 23rd in

8:49:43that case I should again create a

8:49:45smaller version of the file called as

8:49:48edit blog okay that will also be a now

8:49:51these things will be added and will be

8:49:53basically given what can be the Ram size

8:49:56the bigger the better so definitely

8:49:58there is no right answer for it now

8:50:01definitely if you say that 32 GB is good

8:50:03I will say how about 128 GB if you say8

8:50:06GB is good I will say how about 256 GB

8:50:09because that will be better what if I

8:50:11more metadata so we can keep on arguing

8:50:13and keep on increasing right so the more

8:50:15the r better it is for you correct

8:50:18correct with with every change in that's

8:50:20a Lo yes correct that's that's what

8:50:22basically

8:50:23happened now let's move further so I

8:50:26hope now everybody should be clear with

8:50:28this question though it's a separate

8:50:30question but actually it's good that you

8:50:32brought it up because that's one of the

8:50:34very famous interview question so I

8:50:35thought would cover it up now another

8:50:38question what is the problem in having

8:50:41lot of small files in GFS please provide

8:50:44one method to overcome this problem can

8:50:47I get this answer can I get this

8:50:49answer so what are the problems if you

8:50:52will have small files in GFS and also

8:50:55can you give me basically a method to

8:50:56overcome this problem change block size

8:50:59if you change the block size no I I want

8:51:02a better answer I want a better answer

8:51:05name note memory will be overload good

8:51:08good now you are coming to right track

8:51:11three Ram will run out yeah because if

8:51:14you will have small data right if you

8:51:17will have small files definitely your

8:51:19metadata is going to be kind of too much

8:51:21right your metadata entry will be too

8:51:23many and that's how you will be kind of

8:51:25filling up your RAM right of your

8:51:27basically name node we just learned that

8:51:29every metadata is stored in the ram of

8:51:31the name node now what are the solution

8:51:35for it so this is the problem what is

8:51:37the solution for

8:51:39it having larger data clock size so that

8:51:42name node will have reasonable metadata

8:51:44to hand it I can take this answer but

8:51:47I'm expecting a better answer

8:51:48she more map jobs will be used yes

8:51:51that's also one of the problem merge

8:51:53them and save them very good job okay

8:51:56what's your name basically

8:51:59joala I hope that that should not be a

8:52:01real name I don't know why it's sh me

8:52:05joa this is your real name because it's

8:52:08telling me twice joa joa okay so it

8:52:11should be once right it should be once I

8:52:13should call it okay then fine now so I

8:52:17don't know this is occuring two times so

8:52:19that's feeling it real so increasing

8:52:21block size in sdfs merging the file with

8:52:24same and it's easier to read and write

8:52:25the data somebody just answered can you

8:52:28combine

8:52:30everything that is the right answer we

8:52:33can

8:52:34create H file in your windows what you

8:52:39do you create a zip file right or a r

8:52:41file similarly in Hadoop also you can do

8:52:44that you can create a h file which is

8:52:48called as Hadoop art F so you can bring

8:52:52all the small files into one folder

8:52:54together kind of zipping it together now

8:52:57basically with that what's going to

8:52:59happen it's going to just keep only one

8:53:02metadata entry for it the metadata entry

8:53:04is going to be reduced how to do that

8:53:06this is the command archive now hyen

8:53:10archive name whatever archive name you

8:53:12want to give it your input location and

8:53:14output location okay so basically this

8:53:16is how you can deal with the smaller

8:53:21files as well better to create a z this

8:53:24is what you do in the real time also

8:53:25right when you have multiple files of

8:53:27the same type you zip them right just to

8:53:30keep them together so the same thing you

8:53:32will be doing in Ado as well moving

8:53:36further now another question this is

8:53:38also a very interesting question and

8:53:40easy question also Suppose there is a

8:53:42file of size 514 MB stored in hdfs 2.x

8:53:48using default block size configuration

8:53:51and default replication Factor we did an

8:53:54assignment with image files I'm want

8:53:56LinkedIn as okay okay now I got it so

8:54:00using uh default block size

8:54:02configuration and default replication

8:54:04Factor then how many blocks will be

8:54:06created in total and what would be the

8:54:09size of this block okay before you

8:54:11answer this can I get an answer what is

8:54:13the default replication factor and what

8:54:16is the default block size if I'm talking

8:54:18about had

8:54:192.x very good so replication as

8:54:22everybody said 3 m what is the size very

8:54:25good 128 M now it's very easy to answer

8:54:28can everybody answer how to split this

8:54:30514 mbf file how to split this 514 mbf

8:54:35file it's in front of you you can do the

8:54:37calculation and give me the answer as

8:54:39well can I get this answer very good 15

8:54:43block lot of people have given me

8:54:45basically less they said five block but

8:54:48don't you think there will be a

8:54:49replication also of all the block so a

8:54:52lot of people who are giving me this

8:54:54answer of four block is completely wrong

8:54:56and right because there will be a block

8:54:59of 2 MB as well right if you notice

8:55:03what's going to happen this is 128 into

8:55:054 is basically 52 right there will be 2

8:55:08MB block so there are going to be five

8:55:10block because the replic is five now

8:55:13sorry replication factor is three so

8:55:15it's going to be 5 into

8:55:173 okay this is how basically you will be

8:55:20calculating this is very famous

8:55:21interview question moving further how to

8:55:25copy a file into

8:55:28hdfs with a different block size to that

8:55:32of existing block size

8:55:35configuration can I get an answer what

8:55:38basically I'm asking is let say you have

8:55:40a block size of one 20 by the but when

8:55:44you are copying that data right when you

8:55:46are doing let's say sdfs hyp input sdfs

8:55:49Hy input maybe you want to now use the

8:55:52block size of 32 bit not the default of

8:55:55128 bit then what you will do to achieve

8:55:58this yes there's a parameter what what

8:56:00is that

8:56:02parameter what is that parameter block

8:56:05size um no can you see this

8:56:10BFS dot block size okay so what you need

8:56:15to do you need to just Define the

8:56:18basically the bytes what you want to

8:56:19mention so 32 bytes is equivalent to

8:56:21this number okay 32 byes is equivalent

8:56:24to this number so you need to basically

8:56:26Define the bytes what you want to put it

8:56:28up now while doing any command let's say

8:56:31hyen put or maybe hyen copy from local

8:56:35there you can mention this DFS do block

8:56:39right and whatever number of B you want

8:56:42to mention so you can mention that okay

8:56:45if you want to check the block size you

8:56:47have another command called as stat I do

8:56:49FS ion stat and you can see all the

8:56:52statistics related to it so it will tell

8:56:54you how many bites it is to basically

8:56:57distributed and everything up you can

8:56:59basically directly use this F thing and

8:57:01you will get that output okay this is

8:57:04some sometimes useful in projects and

8:57:06that's the reason it's it's a very good

8:57:07interview question as well because a lot

8:57:09of time in the projects you want to you

8:57:11don't want to use the default size you

8:57:13want to change some other to some other

8:57:16number so in that case you will be using

8:57:18this because one way is either you

8:57:20change everything from your

8:57:22configuration FES which is not a good

8:57:24idea to do so better thing is

8:57:25programmatically you deal with it and

8:57:27here you can change it by the usage of

8:57:30DFS do block size okay so it's not block

8:57:34underscore size con I hope you got your

8:57:36answer what's the mistake you were doing

8:57:38it should be block

8:57:40size but you are

8:57:43close now what is a block scanner in

8:57:48hdfs can I get this answer this is a

8:57:51usually a question in your Hadoop

8:57:53Administration this this is basically

8:57:55what your Hadoop administrators do so

8:57:57people who have who have done this Hado

8:57:59Administration classes can you answer

8:58:01this I'm expecting this answer basically

8:58:03from you even others can answer what is

8:58:06a block scanner in

8:58:09hdfs what is a block scanner in hdfs can

8:58:13I get answer nobody's answering this who

8:58:16all have done Administration course or

8:58:19no administrator Hadoop Administration

8:58:21can I get answers who all have done I'm

8:58:23not asking you to answer me this part

8:58:25just to answer me who all have done this

8:58:26Hadoop administrator course initially

8:58:29few people mentioned it that we I have

8:58:31done this administrator course you must

8:58:33have read about block

8:58:35scanner okay let me answer this part

8:58:38usually in a block scanner okay both is

8:58:41answering now to check if the block has

8:58:44any empty space left in the block uh

8:58:48okay one of the answer one of the answer

8:58:50I can take but not exact answer it's not

8:58:52very good uh scan the block and Report

8:58:55the remaining spaces okay okay again I

8:58:59can take partially this

8:59:01answer not just declaiming space but in

8:59:05it ensure the Integrity of your data

8:59:10blocks okay it basically keep on

8:59:13reporting every data Note will keep on

8:59:15reporting to the name one and it will

8:59:17keep on checking the Integrity of the

8:59:20data block let's say if any data block

8:59:21for got kind of corrected right or maybe

8:59:24the replic replica value become low

8:59:27right all those things it keeps on

8:59:29monitoring and try to

8:59:32rectify okay so it will basically keep

8:59:34informing the name but that's the reason

8:59:36this is been usely done by

8:59:37administrators because they keep up

8:59:39monitoring the health of the data no

8:59:42data block name Lo they're also

8:59:44responsible for this work right so this

8:59:46is what they keep on doing in order to

8:59:48make sure do they use block SC to do

8:59:51that okay there's one more way to check

8:59:54the replication Factor anybody know what

8:59:58is

8:59:58that there is one more way to check the

9:00:01replication Factor so this is about

9:00:03block St but there is one more way to

9:00:05check the replication fact environmental

9:00:07F no hardbeat no hardbeat will just tell

9:00:10that data is good or not data node right

9:00:13I'm talking about let's say some file

9:00:15got under replicated in that case how

9:00:18who will kind of inform name not let's

9:00:21say blog scanner is not there there is

9:00:24something called as had load

9:00:28balancer I'm not sure if you have read

9:00:30about that I do load balancer that

9:00:33basically ensures that if if your data

9:00:37blocks are not up right if they're under

9:00:39replicated or not that basically informs

9:00:42that okay this is under replicated let

9:00:44me take the off okay so this is

9:00:47basically the way also to check the

9:00:49under replicated or over replicated

9:00:52blocks can multiple clients write into

9:00:56an hdfs file

9:00:58concurrently can I get this answer if

9:01:01somebody ask you this question can

9:01:03multiple clients write into an hdfs file

9:01:07concr interesting I'm getting one yes

9:01:10one no now two yes one no okay two no

9:01:14two yes lot of yes okay do you think it

9:01:18should be fible to write multiple right

9:01:21I'm not saying reading I'm saying

9:01:24writing notice this part now can you

9:01:28answer don't you think it it it will

9:01:30make my file inconsistent yes single

9:01:33file I'm talking about don't you think

9:01:35it will make my file inconsistent if I

9:01:38do that right it will make my

9:01:41inconsistent so it is not allowed

9:01:44basically it allows only one writing and

9:01:48multiple reading stuff so that's the

9:01:51reason single file it will not allow you

9:01:54to basically keep on writing by at the

9:01:56same time by multiple Cent it will not

9:01:58allow you to basically do that for

9:02:00multiple client at the same time once

9:02:02one client is writing it will be kind of

9:02:04file will be kind of block for other

9:02:06client once the client have written

9:02:09after that only the other clients can

9:02:11right but everybody can read read coner

9:02:15that is one thing which is very

9:02:17important in so writing at the same time

9:02:20is not possible concurrently but reading

9:02:23is possible that's how hdfs is basically

9:02:27created why they have not allowed

9:02:29multiple rights together at the same

9:02:31time because if they do it can make the

9:02:33file inconsistent that will be a big

9:02:36trouble and that's the reason they do

9:02:39not allow you to do for current write

9:02:41because this is distributed system if

9:02:44multiple Cent will write on the same

9:02:45file now maybe somebody can overwrite my

9:02:48change right so that's the reason they

9:02:50will not allow you to do it okay this is

9:02:52by

9:02:54architecture another question what do

9:02:57you mean by high availability of name

9:03:01node and how it is achieved can I get

9:03:05this

9:03:05answer how this is achieved and what do

9:03:09you understand by High availability of

9:03:12the namee in fact I already answered

9:03:14this in Boss example you can answer me

9:03:17this part not

9:03:19replication R awareness name t of passes

9:03:22they be stand by so okay few people are

9:03:25giving right answer now active and

9:03:28passive name what basically happens in

9:03:31this cases so let's say if there are two

9:03:35let me show you the slide itself we have

9:03:36drawn it properly see this there will be

9:03:39two name one will be active name note

9:03:43and one will be passive name note so

9:03:45what happens is let's say this is my

9:03:47active name note which is running okay

9:03:50and what these data notes are there for

9:03:53reporting to the active now we also

9:03:56create a passive name node now this

9:03:58passive name node also these data noes

9:04:01will be reporting this passive name node

9:04:03will not be doing anything but it will

9:04:05just keep on collecting the data from

9:04:07your data nodes okay that is what the

9:04:10role of passive name note would be now

9:04:14as soon as this because this now is

9:04:17reading right so it knows the status of

9:04:19data node it knows that where the blocks

9:04:21are being written everything it has the

9:04:23information now suddenly if this machine

9:04:26is down in that case my passive name

9:04:29note will ensure it will immediately

9:04:32start acting like a backup and that is

9:04:34how it is ensuring High availability you

9:04:38are not going to lose your cluster time

9:04:41so basically the down time will not be

9:04:43there immediately your passive name will

9:04:45start acting like your

9:04:47active Okay so this is how basically

9:04:50Theo is ensuring High availability this

9:04:54is a very famous interview question

9:04:57passive and secondary name note good

9:04:59question difference between passive name

9:05:02note and secondary name note secondary

9:05:05name note we used to use in hadu 1.x now

9:05:09secondary name note what used to to

9:05:11happen was secondary n note you can say

9:05:13it's just like a snapshot of this

9:05:16machine means you're just copying the

9:05:19data copying the data to other machine

9:05:21but if my active name note is down if my

9:05:24main name note is down in that case my

9:05:27secondary name mode will not start

9:05:29acting like a backup it will only keep

9:05:32the data but it will not start acting

9:05:34immediately like a name not it will just

9:05:36keep the data you have to physically

9:05:39manually kind of bring name not up copy

9:05:42the data from the secondary name to

9:05:44primary name not and then start working

9:05:46on it but in passive namee yes so it's

9:05:50kind of an human intervention right

9:05:53manual intervention is required here and

9:05:55there will be a downtime there will get

9:05:57downtime here but when it comes to the

9:06:00active and passive name Lo passive name

9:06:03Lo is going to ensure that it is not

9:06:06only collecting the data metadata but it

9:06:10as soon as active name node is down

9:06:12taing name start acting like of active

9:06:15name clear on this difference now no an

9:06:20you have asked this question right clear

9:06:22about this question answer an the

9:06:24difference should be very clear to you

9:06:27wait so let's move

9:06:30further okay so we have few more

9:06:33questions now uh so this is for map Ru

9:06:35side so let's do one thing friends let's

9:06:37take a five minutes break okay let's

9:06:41take a five minute break and we will

9:06:43start with math produce questions okay

9:06:46uh I'm not sure maybe at Guys somebody

9:06:48will answer you just need a sip of water

9:06:50or so just give me five minutes it will

9:06:52take okay you want 10 minutes this is

9:06:55just basically a two to three hour

9:06:57session so we don't want to make it too

9:06:59much let's make it a seven 7 to 8

9:07:01minutes will that work let's come up

9:07:04with a in Middle kind of solution so

9:07:07let's come back by Maybe by 10 three

9:07:10okay we will

9:07:12start have a water break also we will

9:07:14come back to mauce then we have five we

9:07:17have scoop that so let's come everybody

9:07:20please be back by 103 exactly I'm going

9:07:22to start by

9:07:28103 okay so guys everybody is back now

9:07:32everybody

9:07:33back n you should be able to hear me now

9:07:37okay looks like he's starting okay fine

9:07:40so um now we are going to start uh so

9:07:43basically now we are going to start with

9:07:47math produce topic okay now in yeah I'm

9:07:50not shared I'm not shared I'm just going

9:07:51to share that just give me a second it

9:07:54takes almost few seconds to basically

9:07:56this now I hope it should be showing to

9:07:59you okay meanwhile I have talked with

9:08:03Eda team and they have informed me that

9:08:05you all will be receiving this video

9:08:07recording in a day or two okay since

9:08:10that was a question from lot of people

9:08:11they'll be getting all this video

9:08:13recordings or not so you will be getting

9:08:15this video recording on your email ID in

9:08:17a day or two so all these things will be

9:08:20there with you so I think that will be a

9:08:23quick review for you if you want to take

9:08:25a look at any moment now question for

9:08:27you can you explain me the process of

9:08:31spilling in MA

9:08:34RS this is an interesting question in

9:08:37fact think from a perspective of where

9:08:41mapper keeps the output and from that

9:08:43you can basically make out what is this

9:08:45sping I give you a biggest hint possible

9:08:48here so can you give me this

9:08:50answer can you explain the process of

9:08:54spilling in map

9:08:57reduce spills to Temp folder lfs of when

9:09:02it spill and from where it

9:09:05spill can I get this

9:09:07answer map per F very good

9:09:11what usually happens is the output of

9:09:15your mapper task it goes to your R now

9:09:19what basically going to happen they have

9:09:21kept a specific size of that thing so

9:09:25let's say they keep let's say this 100

9:09:28MB now 100 MB of data will be kept let's

9:09:32say in gra but they have they will be

9:09:34keeping a press so it will slowly keep

9:09:36on filling up slowly keep on filling up

9:09:39then what will happen as soon as it will

9:09:41reach a threshold let's say 80% of that

9:09:43Ram of 100 MB is spit it will start

9:09:46spilling that output to the local disk

9:09:50notice here I'm not saying

9:09:53hdfs I am saying local disk okay to

9:09:57local disk only we will be keeping this

9:10:00data local disk means your C drive D

9:10:03drive wherever you want to keep up so

9:10:05this is how they have designed it so as

9:10:08soon as the mapper out putut in the

9:10:11memory reach to a threshold limit it

9:10:14starts filling that mapper data to your

9:10:18local dis and this phas is called as

9:10:22filling phas in Mist okay so this

9:10:27question is also asked and this shows

9:10:29basically the internal working of your

9:10:31math prod okay so basically this is how

9:10:34internally your map produce work so this

9:10:39is all this is filling the data once

9:10:40it's filling it will again come down and

9:10:42you can see more and more data in it

9:10:47which links me to other question can you

9:10:50explain me the difference between block

9:10:53input splits and record this is a very

9:10:57famous interview question can anybody

9:10:59answer me Qui me what is the difference

9:11:02between blocks input split and

9:11:08Records what is the difference between

9:11:10the

9:11:11three friends this is a very important

9:11:14question if you have done this maass

9:11:15reduce part then you must be knowing

9:11:17this

9:11:19part difference between blocks input

9:11:23split and Records very good ji so J is

9:11:26answering record is a single line of

9:11:30data right very good Canan is saying

9:11:32block is hard cut of data just 128 MB

9:11:36very good right block is equal to 128 MB

9:11:40if record is 130 MB input split will

9:11:43happen very good block is based on block

9:11:46size input split makes sure that the

9:11:48line is not broken so it makes sense

9:11:51record is single line block is set by

9:11:54sdfs input split is logically spit very

9:11:57good very good so what usually happens

9:12:00right so let's say when we talk about

9:12:02blw so let's say you have default space

9:12:04is 128 and so that will be called as a

9:12:07physical block okay when we talk about

9:12:10input split right so let's say if your

9:12:12data is of 130 M now in that case don't

9:12:15you think it makes sense to to have a

9:12:18logical spit of 130 MB here so that will

9:12:20be your input spit and record is when

9:12:24you do Mac produce programming right

9:12:26when you do maap produce programming how

9:12:28your mapper take the data it takes line

9:12:31by line right it takes line by line that

9:12:34line is called record so one line of

9:12:38data which it picks up in the map face

9:12:41is called your record okay very very

9:12:45famous question on this part so you can

9:12:48say block is a physical division by

9:12:50logical division are called your input

9:12:52splits and Records okay because the

9:12:55logical division is what your map

9:12:57produce program do which brings me to

9:13:00another question again relate to map

9:13:02produce what is the role of record

9:13:06reader in Hado map prod

9:13:11what is the role of record reader in

9:13:15hard do map

9:13:17ruce make sure to read the complete

9:13:20record no no how that is how Mapp reads

9:13:25a record good but dber can you give be

9:13:28little more exp you coming close to the

9:13:31answer can I get more answer also D can

9:13:34you just be a little more

9:13:35explicit that is how map read and record

9:13:39we coming close

9:13:41what about

9:13:43others what is record record is single

9:13:47line we have just understood it so what

9:13:50should be record reader paring that

9:13:52single line very good word right so

9:13:55don't you think that single WR what you

9:13:57are reading and how mapper convert your

9:14:00data it converts into key value PA right

9:14:03so it initially takes an input as a key

9:14:05value pair so when that conversion is

9:14:08happening that is done basically by

9:14:10record reader look at this see let's say

9:14:13this is the data it will be getting

9:14:15converted to key values here where key

9:14:17is called your offset and value is first

9:14:20line right or second line or third line

9:14:23right so this is done by record reader

9:14:27okay so this is what record reader do

9:14:31now what is the significance of counters

9:14:34in maap

9:14:36is significance of counters in map

9:14:44is uh okay give statistics of data

9:14:48counters will be done in name not okay

9:14:51counters to validate the data R to

9:14:54calculate Bard good not just bad record

9:14:58people it can do even other things right

9:15:00this bad record is just one example

9:15:02right it just one example you can say so

9:15:06what it do is it helps you to identify

9:15:09the statistics right now basically it

9:15:12gives you the statistics about some

9:15:14operation what you want to do you can

9:15:16print it in the console also I I believe

9:15:19that if you have done this Hado course

9:15:20you must have seen one code for your

9:15:22counters right where you must be doing

9:15:25some sort of operations and you you

9:15:27might be printing it on your console

9:15:29window now how you will be doing all

9:15:32that so what you will do let's say we

9:15:34are applying counters on this example

9:15:36and in this example as I think somebody

9:15:39just mentioned for the kind of reading

9:15:41the bad data this example actually bring

9:15:43up the same thing but it can do even

9:15:46other things I will tell you what other

9:15:48things no but let's take this example

9:15:50let's say we want to find out what all

9:15:52bad data I have in this so let's say it

9:15:55is reading it is reading David all good

9:15:57no problem counter will remain as zero

9:16:00then what it did it basically just gave

9:16:02the value as zero it reach to the second

9:16:04line now this value move now basically

9:16:07this was the back data I read it I

9:16:09passed the statistics Now counter value

9:16:11became one it reach to Jeff if the value

9:16:14stays as one because this is a good data

9:16:17now Sean again the value Remains the

9:16:19Same now as soon as it reach the last

9:16:22line of that data it will again increase

9:16:26this counter and will return that I have

9:16:29two bad line of course but does that

9:16:32mean we can only do this operation on

9:16:35this no maybe I can have an example

9:16:38where let's say I have data of which is

9:16:41let's say of time stamp time stamps are

9:16:44there okay I want to print in this time

9:16:46so time stamp I can convert to date type

9:16:49right in my program I can convert it to

9:16:51date type now when I convert to date

9:16:53type maybe I have months being defined

9:16:55right because in in date I have months

9:16:58now maybe I want to find out I want to

9:17:01calculate the statistics that in among

9:17:04this time stamp how many times January

9:17:06is offering how many times February is

9:17:09offering in how many times March April

9:17:12and all are occuring let's say I want to

9:17:14identify all the statistics I can do

9:17:17that with the help of counters very

9:17:20easily okay so this is the purpose of

9:17:24your

9:17:25counters now moving

9:17:28further why the output of math pass or

9:17:32spilled into the local disk and not in

9:17:36hdss good question now can you answer me

9:17:38this remember we talked about uh we

9:17:42talked about yes we have just discussed

9:17:44that question right so we have just seen

9:17:46spilling we have seen spilling

9:17:49now how can we assess counters there are

9:17:52basically libraries available for that

9:17:53there are classes available get counters

9:17:56the is the member function to get that

9:17:59this is how basically will be assessing

9:18:00it so counter is the class in which we

9:18:02have member function predefined

9:18:04functions using that you can assess

9:18:07that because it's an intermediate output

9:18:11okay intermediate output that's fine but

9:18:14why we are not keeping in sdfs that's

9:18:16that's my question when I can keep

9:18:18intermediate output in

9:18:20sdfs very good very good

9:18:23D very good nimma so basically if you

9:18:26notice what happens if you keep in hdf

9:18:29right remember there's a replication

9:18:31Factor right that replication Factor

9:18:33will do what it will increase the number

9:18:36of output blocks do you think it makes

9:18:38sense to increase the replication for

9:18:40your Mapp output definitely a big no

9:18:43right so that's the reason we will be

9:18:45keeping in your local F system we will

9:18:47not keeping sdfs otherwise sdfs will

9:18:50replicate even mapper output which we

9:18:52don't want to happen to do so that's the

9:18:54reason we will stop that so we will be

9:18:56keeping in the local dis not in hdfs in

9:19:01order to avoid basically the

9:19:04replication which brings me to another

9:19:07question can you define this speculative

9:19:11execution can you define this

9:19:14speculative

9:19:17execution can I get this answer can be

9:19:20defined speculative

9:19:25execution it prioritize only some task

9:19:28uh coming close but not exactly

9:19:31right if a job in a note is taking much

9:19:34time very good very good people so this

9:19:38is what happens in speculative execution

9:19:42let's say if any of your task is running

9:19:46very slow in that case your Speculator

9:19:49there will be a scheduler which will

9:19:51basically start a duplicate task of it

9:19:55it will start running a duplicated task

9:19:58for it to ensure that basically that

9:20:01duplicate task run faster and once it

9:20:03will finish it will kill all the

9:20:05duplicate task so it is just kind of

9:20:08making sure that because because it can

9:20:10happen right maybe your task is waiting

9:20:11due to some resource it got blocked due

9:20:13to any reason so in that case it will

9:20:17start immediately a duplicate task and

9:20:21making sure that your job finished

9:20:24quickly okay so this is the part of your

9:20:28speculative

9:20:31execution which brings me another

9:20:33question question is how will you

9:20:36prevent y very good very good how will

9:20:40you prevent a file from splitting in

9:20:43case you want the whole file to be

9:20:46processed by same

9:20:49Ma how you will prevent a file from

9:20:53splitting in case you want the whole

9:20:56file to be processed by same mapper I

9:21:00want my file to be now basically to be

9:21:02used by the same m not combin not combin

9:21:06theb can I get some more answers it's

9:21:09easy

9:21:10can you see this side so what we can do

9:21:14here is first of all we can increase the

9:21:18minimum number of split size which

9:21:21should now make it larger than your last

9:21:24five this operation is itself good

9:21:27enough to make this case look right

9:21:30because if you increase the size itself

9:21:31we would be all good but there is one

9:21:34more thing which you can do after that

9:21:36is this is Method two basically in

9:21:39method one you can just basically

9:21:41increase the size itself it will be all

9:21:43good or what you can do you can go to

9:21:45your input format CL and in that you can

9:21:49just update this property you can make

9:21:52this is splitable to be returning first

9:21:55usually people prefer method one because

9:21:57that's the easiest method right you need

9:21:59not update the Java code basically to

9:22:01achieve all this so what you will do for

9:22:03this can you please tell a scenario

9:22:06where file splitting is not needed it

9:22:08all depends right so let's say if I know

9:22:11that I have only one data node I mean

9:22:13let's say I have only one data node now

9:22:15in that one data node do you think it

9:22:16makes sense to divide a file multiply

9:22:18together I have only one data Note right

9:22:21do you think it makes sense because

9:22:23again it's in the same machine right so

9:22:25that's the reason there you want to

9:22:27basically keep one block it so that I

9:22:29can execute it so this is these can be

9:22:31few situations where you can decide not

9:22:34to split the file and in that case how

9:22:37you will be doing it these are the two

9:22:39method to achieve okay moving

9:22:43further is it legal to set the number of

9:22:46reducer task to zero that's question

9:22:48number one where the output will be

9:22:51stored in this case is it legal to set

9:22:54the number of reducer task to zero is it

9:22:59legal definitely legal right everybody

9:23:03have your school in school was there any

9:23:06reducer was there any reducer in school

9:23:10no scoop only use mapper no reducer

9:23:15right that's itself a tool right when a

9:23:17tool is not even creating any reducer

9:23:20definitely I will not be using that

9:23:22right so what is going to happen here is

9:23:25let say so what what was the purpose of

9:23:27reducer purpose of reducer is when you

9:23:31want to do some sort of aggregation

9:23:33right maybe you want to Summit in the

9:23:35end and all those category kind of

9:23:37things but it's not Mand that all the

9:23:40problem statement in the word require

9:23:43aggregation right so those problems

9:23:45where you do not require agregation like

9:23:48I said scoop because in scoop what

9:23:50happens you copy the data from rdbms to

9:23:53your hdfs or vice versa now are you

9:23:55doing any sort of aggregation no right

9:23:58you're just copying the file from rdbms

9:24:00to your hdfs system no agregation

9:24:03require so in those cases you will be

9:24:06having reducer as zero output where to

9:24:10store definitely whatever mapper output

9:24:12is coming wherever you're telling it

9:24:14will be stored in that sdfs location

9:24:16right so that how you will be using it

9:24:19so definitely the answer is yes and

9:24:21basically wherever the mapper output you

9:24:23seeing there it will be getting SC what

9:24:27is the role of application master in map

9:24:31reduced job can I get this answer this

9:24:34is a very famous interview question what

9:24:37is the role of app ation master in map

9:24:41reduce job to assign the task okay

9:24:46that's it only to assign the task sets

9:24:49the input split okay it manages the

9:24:53application fired and keep track of sub

9:24:55process that's it nothing else to get

9:24:59the resources needed for the task very

9:25:01good now you are coming here we are

9:25:04coming close to create task until until

9:25:07it yes what basically application Master

9:25:11do first thing is application Master is

9:25:15kind of deciding that how many resources

9:25:19it needs okay and it can basically now

9:25:23inform resource manager that I require

9:25:26this many resource and give me this many

9:25:29resource to execute that then after that

9:25:31resource manager give back the container

9:25:33right if you might have gone through

9:25:35your Yan architecture right there you

9:25:37might have understood all this potions

9:25:39right so basically what happens it first

9:25:42basically find out how many resources

9:25:44are required secondly what it want to do

9:25:48is it basically want to find out once it

9:25:51happens when it gets all the container

9:25:53it kind of basically get them working

9:25:56together it collect the output back and

9:25:59return it to basically the master so it

9:26:01is doing multiple things in MA it's not

9:26:04it's basically not doing just one task

9:26:07it is also checking which is a very

9:26:09important role of it it is checking that

9:26:12how many resources I require and kind of

9:26:15helping the resource manager to take

9:26:18that decision okay so this is what the

9:26:21same thing being explained here the role

9:26:24of your application Bas this brings me

9:26:28to another question what do you mean by

9:26:32overb so when your map ruce job runs I

9:26:35don't I'm not sure whether you have

9:26:36noticed that or not there is something

9:26:39called as Uber mode it it's sometime

9:26:42comes in your conso if you have noticed

9:26:45that so can you tell me what is that

9:26:47Uber mode is there any advantage of

9:26:51searching on the Uber mode what

9:26:53basically Uber mode is going to do runs

9:26:56on application Master Mod very good can

9:26:58I get some more answers more insight on

9:27:01this good can I get some more answers

9:27:05yes so let's say if you have a small job

9:27:09right if you have a small job in that

9:27:12case you you're basically again

9:27:15application Master need to be any way up

9:27:17now application Master will request and

9:27:20then it will allocate container right so

9:27:22container sending and all basically

9:27:24creating container all those things are

9:27:26time consuming if the jobs are small

9:27:30what your uh your application Master can

9:27:33do application Master can start a jvm in

9:27:37itself okay so basically application

9:27:40Master can decide to complete your job

9:27:43because the job is small it may decide

9:27:45to complete the job on its own in that

9:27:48case we call it as Uber mode so Uber

9:27:51mode basically require less it's when

9:27:54you you will be using let's say less

9:27:56number of Ms only 10 Maps you have only

9:27:59one reducer to work on so in those cases

9:28:02you use Uber mode now how to enable all

9:28:05that so basically there's a property

9:28:07which you can set to true and it will

9:28:10basically enable your over mode in which

9:28:13what's going to happen your application

9:28:16Master will start acting like a JM and

9:28:19will finish the job so does it will save

9:28:21you to buy creating container making

9:28:24your performance better so usually what

9:28:26people do is whenever they will be

9:28:29having some sort of uh whenever they

9:28:31will be having small jobs right and they

9:28:33want to improve the performance they

9:28:35usually kind of enable the Uber mode so

9:28:39now the application Master itself start

9:28:41executing the task and that way they

9:28:43improve the performance but when you

9:28:45have a bigger job it will not work in

9:28:48that case you need to keep the over word

9:28:49as false otherwise you will degrade your

9:28:52performance it will not even work okay

9:28:54so this is basically your Uber

9:28:58mode another question how you will

9:29:01enhance the performance of M prod jobs

9:29:05when dealing with too many small FES if

9:29:09you have let's say many many small files

9:29:11in that case very good report yeah

9:29:14that's also one thing if you have let's

9:29:16say many many small files how you can

9:29:19improve the performance of your map Ru

9:29:22job The Miner no coming close but not

9:29:26the right answer too many small files

9:29:29are there then what you will do uh

9:29:32basically there is something called as

9:29:35because none of you give the very right

9:29:37answer there is something something

9:29:39called as combined file input format

9:29:42that's the reason I told you was you are

9:29:43coming close but not exactly coming

9:29:45right right so what this do is it kind

9:29:48of packets all the small files together

9:29:51see this diagram see this part like

9:29:54these are some small files it basically

9:29:56now combined all the files together now

9:29:59because these small files got combined

9:30:02together now my execution time will be

9:30:04FAS this is basically one practical

9:30:07thing which which which is is being

9:30:09represented this performance can you see

9:30:11like the small price was taking this

9:30:13much of time and basically with this

9:30:15when you combine this it actually

9:30:17started taking less time so that's

9:30:19improving the performance of your system

9:30:23so this is what this this is how

9:30:26basically you can improve the

9:30:27performance if you have multiple small

9:30:32FES now let's move to hi hi is very

9:30:36important topic now question for you

9:30:40where the data of high stable J

9:30:44St I know that is going to be in sdfs

9:30:48but where is the location of that where

9:30:50is the location for that but what is

9:30:53that default folder that version very

9:30:55good now you are coming close so by

9:30:58default you keep it in slash user slash

9:31:03hi/ Warehouse so this is the location

9:31:07where by by default all your hi table

9:31:11get store if you want to change this you

9:31:14can go to your high. XML and can update

9:31:17the setting as well another question why

9:31:21hdfs is not used by high meta store for

9:31:27storage what I mean is you might have

9:31:30read in your course that you keep your

9:31:34hi basically your metal store in your

9:31:37rtpms right not in

9:31:39hdfs what is the reason behind that why

9:31:44we not keep all these things in my hdfs

9:31:49why your met store is created in your

9:31:53rdbms why you config configuring your

9:31:56meta store in nbms for he why not in

9:32:00hdfs can I get this

9:32:03answer why you are keeping your met

9:32:06store in your rdb system system and not

9:32:09in your hdfs I can tell you this is the

9:32:12most important question of he and any

9:32:16interview of P you will go expect this

9:32:20question this usually everyone ask for

9:32:24random SS but that you can do in sdfs

9:32:27also if you need DB catalog what is that

9:32:31catalog because it need to use jdbc

9:32:34connection no no that's not the

9:32:36answer okay let me ask you this question

9:32:39okay which is actually the main reason

9:32:41for it uh so let's say if uh what you do

9:32:45is when you create a table in height

9:32:48what happens it creates an entry for

9:32:51that table in metas store table right

9:32:54what it's doing it's inserting the roow

9:32:57level right at a row level it is

9:32:59inserted you created another table in

9:33:02again you are basically inserting some

9:33:04values here right now let's say you

9:33:06deleted some table in hand what happened

9:33:09it did a ro level delete can you do this

9:33:12Ro level delete Ro level insertion and

9:33:14all in your

9:33:16hdfs okay you're saying referring to the

9:33:18same thing that is then you're good can

9:33:21you do this basically this in your sdss

9:33:23no right this itself is a good answer to

9:33:26explain this right so and second thing

9:33:29is definetely inm the thinking time is

9:33:31going to be faster and all those are the

9:33:33other facts but first thing is basically

9:33:35your cred operation cannot be done in

9:33:38CFS that's the reason we will not be

9:33:41able to keep in the CFS forget about

9:33:44other factors right they do not make any

9:33:47sense in fact because my first property

9:33:50itself is failing of basically current

9:33:52operation now let's see some scenario

9:33:56questions usually in high Pig you will

9:33:58find some scenario questions coming up

9:34:00now scenario question is suppose I have

9:34:04installed a Pache Hive on top of my Hado

9:34:09can you please show the last answer sure

9:34:11why not see this answer in the here

9:34:16let's move forward now can can you see

9:34:19this question now suppose I have

9:34:22installed AE Hy on top of my Hado

9:34:25cluster using default meta store

9:34:29configuration then will what will happen

9:34:33if we have multiple clients trying to

9:34:36assess High add same time can I get this

9:34:41suppose I have installed AE hi on top of

9:34:45my hadu cluster by using default

9:34:48metastore configuration then what will

9:34:50happen if we have multiple client trying

9:34:53to assess H at the same time very good

9:34:57usually in height you can only basically

9:35:02assess one by one client so multiple

9:35:05client sess itself is not allowed right

9:35:10very good they given basically the right

9:35:12answer for it so usually this these are

9:35:14the scenarios what it follows so main

9:35:16thing is your multiple client sess in hi

9:35:21is not allowed this is by architecture

9:35:25right because you should maintain the re

9:35:27consistency very good addition right so

9:35:30that's the reason this itself is not

9:35:33going to work out so basically that is

9:35:36what you need to keep in mind what is

9:35:39the difference between external table

9:35:43and manage table in fact manage table

9:35:46you also call it as internal table so

9:35:49can I get an answer what is the

9:35:51difference between external table and

9:35:54manage table can I get this

9:35:57answer external table can be a file okay

9:36:01but I want a difference proper

9:36:03difference external table where the hdfs

9:36:07file won't be deleted if you delete the

9:36:09table and you say sdfs file okay okay I

9:36:12can accept your answer external table is

9:36:16stored in separate location of our but

9:36:18even internal table I can store it at

9:36:20some other location by defining the

9:36:22location ke

9:36:24and external table keeps data when it

9:36:27get deleted very good that is the major

9:36:31factor when you talk about manage table

9:36:35what happens is if you have deleted any

9:36:38of the DAT table what is going to happen

9:36:41it will delete the entry in your metas

9:36:45store at the same time it is also going

9:36:48to delete the data file but in external

9:36:52table if you delete a table it is going

9:36:55to only delete the entry in your meta

9:36:59store not from your main data okay so

9:37:04that is the major difference basically

9:37:07in the minus table and external

9:37:10table another question when should we

9:37:14use sort by instead of order by if you

9:37:18notice these two apis belongs to five

9:37:21and these two basically going to do

9:37:23exactly same thing so when should I use

9:37:26sort by and not order by operation when

9:37:31should I use sort by and not order by

9:37:35operation so basically to answer this

9:37:39not able to catch things uh I did not

9:37:41get this R when we should have only one

9:37:46ma okay R I think you are in fifth

9:37:48module right so that's the reason uh

9:37:51yeah I can understand you have told me

9:37:53the starting itself right that you are

9:37:55still going through this course so if

9:37:57you're not getting it is completely fine

9:38:00just listen to this okay just listen to

9:38:02this conversation once you will go over

9:38:05these course topics in uh basically from

9:38:08wherever you are doing you will be all

9:38:10comfortable with it okay basically now I

9:38:13can understand if you will not get

9:38:14anything because these modules is not

9:38:16being taught to you yet so these are

9:38:18basically the new B which which will be

9:38:21taught later now can I get this answer

9:38:24uh in case of numericals no no that you

9:38:27can use also order by one of it use

9:38:30reducer other use mappers okay okay when

9:38:35you use Group by operation no no that

9:38:38turn way actually what happens is if you

9:38:43have huge data set in that case you

9:38:47should use basically the sort by option

9:38:52it usually do this sorting on multiple

9:38:55reducer while order by do it on one

9:38:59reducer that is basically the major

9:39:02difference so when you have huge data

9:39:05set use sort by by instead of order by

9:39:11okay lot of people remain confused with

9:39:13this that's why this is a very tricky

9:39:15question what people are usually if you

9:39:17ask anyone right if you do not go the

9:39:19answer he will tell you both do the same

9:39:21thing but actually that's not the there

9:39:23is a difference now another question

9:39:26what's the difference between partition

9:39:28and bucket in I think the most easiest

9:39:30question to answer can everybody answer

9:39:32this whoever have done on hi topic

9:39:35what's the difference between partition

9:39:37and bucket

9:39:38this is the most easiest

9:39:40answer can I get this answer difference

9:39:43between partition and bucket simple

9:39:45right partition is basically at the

9:39:48first level right when you split the

9:39:50data into different directory Buffet is

9:39:53like a subpartition of that right so

9:39:56even for that partition itself when you

9:39:58create another subpartitions you can

9:40:00call them as bucket like in this case

9:40:03can you see the first partition is PC

9:40:05Department Civil Department electri

9:40:07electrical department but after that we

9:40:09have also created some subpartitions of

9:40:12it and that is your bucket that's

9:40:15basically the differen another question

9:40:18let's say this is the scenario you are

9:40:21creating a transaction table now this is

9:40:25the table what you have like transation

9:40:27table is the table you have this many

9:40:29columns delimited field by comma now

9:40:33let's say you have inserted 50,000

9:40:35couples in this table now I want to know

9:40:39the total revenue generated for each

9:40:41month but he is taking too much time in

9:40:46processing this query can you tell me

9:40:49what the solution you are going to

9:40:51provide this scenario is actually a very

9:40:55good interview

9:40:56scenario very

9:40:59good can I get more answers very good

9:41:02very good can I get more answers you

9:41:05will be partitioning this table how you

9:41:08will be partitioning you will be

9:41:11partitioning your table with month right

9:41:15so basically if you partition your table

9:41:18you will improve your performance so

9:41:21these are the simple steps you can

9:41:23create a table Partition by month set

9:41:26these properties to truth so that you

9:41:27can enable your partition insert the

9:41:30data and then you can retrieve the data

9:41:33where your month is going to January so

9:41:36while parage after partitioning the

9:41:38table you can improve the performance

9:41:41second can I get an answer of this what

9:41:44is dynamic partitioning and when is it

9:41:48used can I get this answer what is

9:41:51dynamic partitioning and when is it

9:41:55used that can be static partitioning

9:41:57also in right so I want to know what is

9:42:00dynamic partitioning very good partition

9:42:04happens when loading the data into table

9:42:08right now I don't know that if I do a

9:42:11dynamic partitioning where it is which

9:42:13how many partitions also it is going to

9:42:15play so the value of your partition

9:42:18columns will be known only during your

9:42:22run time when you will be creating the

9:42:24partition that is called your Dynamic

9:42:29partitioning okay how high distribute

9:42:33the row into bucket can I get this

9:42:36answer how how hi distributes the rules

9:42:40into bucket very good hash algorithm

9:42:45okay it uses the hash algorithm to

9:42:49understand this part if you look what we

9:42:51are doing here is now no you will be

9:42:54using clustered by but basically in how

9:42:57internally this is that's basically how

9:42:59it is let's say you want to put into two

9:43:01bucket in that case it is going to do

9:43:03modul of okay let's say mod of to this

9:43:06output table came out to be one so it

9:43:08will decide to put in bucket one modul

9:43:10two it is going to become zero it is

9:43:12going to keep in this second bucket so

9:43:15this is how it will be decided okay it

9:43:17will be using basically a hash

9:43:20computation of this so basically it will

9:43:22be using hash function of this so let's

9:43:24say hash value of this value came out to

9:43:26be one hash function of this value came

9:43:28out to be two hash function value this

9:43:30came out to be three then basically it

9:43:32is doing this modular operation and

9:43:35giving the output and that's how will

9:43:37distribute the buting data now which

9:43:42brings another question suppose I have a

9:43:45CSV file which is named as sample. CSV

9:43:50present in tm1 directory with the

9:43:54following entries in that case how you

9:43:58will consume this CV file into where hi

9:44:02Warehouse using buildin surday sday mean

9:44:08ization s des that's serialization der

9:44:12serialization when you convert your data

9:44:15into kind of kind of

9:44:17bbes can I get this answer row format

9:44:20delimited by comma not exactly I'm

9:44:24looking for something else actually this

9:44:26requires some API so let me show you

9:44:28this part see this answer in this case

9:44:32you will say raw format Sur or do Apache

9:44:37do. sur2 doop CSV

9:44:41Sur this is what you need to add okay

9:44:45now you got it what's the mistake we

9:44:47doing so basically this is what you need

9:44:49to add otherwise everything is same you

9:44:51can sa in the TMP folder and all that

9:44:54just that the only difference will come

9:44:56in row form at thir another question I

9:45:00have lot of small CSP files present in

9:45:04input directory as and I want to create

9:45:08a single table height table

9:45:11corresponding to these files the data in

9:45:14these files are in this format now as we

9:45:17know had do performance degrades when we

9:45:20use lot of small files so how you will

9:45:24solve this problem can anybody give me a

9:45:27simple answer of this this should be

9:45:28easy you have multiple small FS now in

9:45:31that case what what should be the

9:45:33solution because definitely my

9:45:35performance is going to degrade what

9:45:37should I do concatenate solve file but

9:45:40can there be another answer don't you

9:45:41think you can use sequence file here

9:45:43sequence file right if yeah one solution

9:45:47can be that is from hdfs Level itself

9:45:50right I'm talking from basically height

9:45:52perspective right don't you think I can

9:45:54convert in a sequence file sequence file

9:45:57will convert everything like 0 1 01 kind

9:45:59of thing so that's what make it better

9:46:02right will will improve my performance

9:46:04so first create a table load the data

9:46:07after that what you do store it as

9:46:11sequence file and then basically load

9:46:13this data from this file what you

9:46:15inserted to this file this will ensure

9:46:18now that your speed will be good why do

9:46:22we need to do in serice why do okay so

9:46:25basically when you do cice right the

9:46:28advantage what you get is so so let's

9:46:30say first thing is compression because

9:46:32you are serializing the data when you

9:46:34serialize the data it makes that trans

9:46:37were also very easy because we have

9:46:39converted like 01 01 01 kind of right so

9:46:43first thing is when you convert to sers

9:46:45you comess the data second thing since

9:46:47you convert to this 01 format now the

9:46:49transfer over the network become lot

9:46:51easier for me clear the that's the

9:46:54reason we basically use this so but

9:46:59remember there is no pre- lunch right

9:47:02there is no nothing called as pre- lunch

9:47:04don't you think this will also have a

9:47:05disadvantage when you you do a de

9:47:07realization again you need to convert it

9:47:10back don't you think it will impact the

9:47:11performance a bit right so that remember

9:47:14there is no fre lunch though it is

9:47:16helping you in this act but at the same

9:47:18time it will demand your performance it

9:47:20will leat up some of your performance

9:47:23now some quick

9:47:24questions can you give me this answer

9:47:26difference between logical and physical

9:47:29plan something difference between

9:47:32logical and physical F I know guys that

9:47:35you got little tired because this is

9:47:37big session but don't worry we we are

9:47:40almost getting done so I want everybody

9:47:42attention to be back now can you tell me

9:47:44the difference between logical and

9:47:46physical plan this must be the first

9:47:49thing what you must have learned in your

9:47:51hadum course when you went to pick

9:47:53topic whenever you are executing

9:47:57statement by statement okay it is just

9:48:00executing the statement nothing in that

9:48:03case first it creates a logical plan

9:48:06means let's say if there is no error

9:48:09right in that case it is just creating a

9:48:10logic but when you do dump right when

9:48:14you do dump then only the execution

9:48:16start right because of lazy valuation

9:48:19then your logical plan kind of get

9:48:21converted to like of physical plan means

9:48:24it start getting executed now let's say

9:48:27if you have given the wrong file part

9:48:30right in logical plan it will not give

9:48:33you any R because there is no syntax

9:48:36only at the time of physical plan it

9:48:38will give you an error So Physical plan

9:48:41is when you are basically executing your

9:48:43map job when this pig is getting

9:48:45converted to map job and by logical plan

9:48:48is at the initial

9:48:50level okay so that is what happen can

9:48:53you tell me what is bad what is

9:48:56bad collection of couples very good

9:48:59right so basically when you say a whole

9:49:02data file itself right collection of

9:49:04couples if you notice so let's say this

9:49:06is a data right if this is a data now if

9:49:09you see this is one tle this is second

9:49:12tle this is third tle so collection of

9:49:15all these things will be called as back

9:49:18okay collection of all these things will

9:49:20be called as back

9:49:22now how hi is only working with like hi

9:49:28is able to deal with only structure data

9:49:31but pig is able to deal with

9:49:33unstructured

9:49:34data pig is how Pig is able to deal with

9:49:37unstructured data can I get this answer

9:49:41how pig is able to deal with

9:49:43unstructured data it is actually

9:49:46happening because of schema less part

9:49:49right schema less part basically if you

9:49:52do not have schema if you do not have

9:49:54schema depend then also P can work what

9:49:57P do is let's say you you do not know

9:50:00like in hi you have column names right

9:50:02let's say column name is age integer all

9:50:06that right so you need to define the

9:50:08column names in P there is nothing like

9:50:10that so let's say if you do not know the

9:50:12column name there's no schema being

9:50:14defined you can Define it like this

9:50:16First Column you can Define by dollar

9:50:18one second column you can Define with

9:50:20dollar two so that's how you can also

9:50:22Define so basically Pig you can Define

9:50:25even your schema less thing okay so you

9:50:30don't have anything it will treat it

9:50:32like null it will start treating if you

9:50:34do not Define data type it will start

9:50:37get is buy so basically Pi kind of

9:50:39converts your values to other way so

9:50:43basically that's the reason you can

9:50:46physically go with pig with unstructured

9:50:50data this is one of the major reason

9:50:53that P can deal even unstructured data

9:50:56by height cannot do because P kind of

9:50:59converts the data or Tre that data inv

9:51:02in different way right like if you don't

9:51:04have column name you can Define as

9:51:07dollar two dollar one all those things

9:51:09as

9:51:10well what are the different execution

9:51:14mode available in pig so there are two

9:51:18modes right one is local mode one is map

9:51:23produce mod yeah sure can do

9:51:26that this is the same thing I mean like

9:51:29if you have no data treat it like B if

9:51:31you don't have column it start reading a

9:51:33dollar one dollar two right before okay

9:51:38this one back got it now great let's

9:51:41move

9:51:42further now so there are two modes

9:51:45available one is map ruce mode one is

9:51:48local mode so when you go with pig in

9:51:51map produce mode so when you just type

9:51:53Pig right it take you to the grun by

9:51:56default it take you to the mapm which

9:51:59basically also states that that

9:52:01basically if you're are going with back

9:52:04mode you are assessing your HD this

9:52:07while if you are using local mode what

9:52:09you need to do you need to go like this

9:52:11Pick hyen X local it will now take you

9:52:17as in the local mode when you say local

9:52:20mode what basically happens here it

9:52:22basically now start assessing the file

9:52:25from your local file system now it is no

9:52:27more assessing sdfs but it is directly

9:52:30assessing the data from your local file

9:52:33system these are the two execution mod

9:52:36platin this is very simple right so

9:52:38basically flatten is the keyword

9:52:40available if you have this kind of data

9:52:42you can flatten it up like everything

9:52:44will come together in the line so

9:52:46plattin is just in API right you can see

9:52:50this plattin so basically it is just

9:52:52basically converting this form of data

9:52:54to this form of data okay so this is

9:52:57basically meant by plattin these are

9:53:00simple questions now can anybody explain

9:53:04me this xB B

9:53:08components can anybody explain me these

9:53:11xbase

9:53:13components anyone who want to talk about

9:53:15it hbed components it's in front of you

9:53:20can anybody explain me these components

9:53:23of H

9:53:24base you can start by one by one this is

9:53:28the last topic so friends I want

9:53:30everybody to be attentive here so hbas

9:53:33keeps the data in a distributed mode

9:53:35right so distributed what where it keeps

9:53:38the data it defines a region where it

9:53:41keeps the data so like this will be your

9:53:44first region this will be like region

9:53:46where you're keeping let's say this

9:53:47column value row value right this is a

9:53:50one region not together just like how

9:53:52you define rack right you can Define

9:53:54region surface right where you're

9:53:57defining basically different different

9:53:59regions together so this is one region

9:54:00server this is one region server right

9:54:03now what happens the master will keep

9:54:05basically your will be called as H

9:54:08Master like active Master just like your

9:54:10name note works right similarly hbas

9:54:13uses this H Master what is this Z

9:54:16zookeeper doing here zookeeper is kind

9:54:18of helping you to execute everything so

9:54:21like in zoo what happens in zoo

9:54:23basically they keep animals right they

9:54:25keep animals they manage multiple

9:54:27different category of animals similarly

9:54:29do people like big data also got so many

9:54:31tools available for you now and

9:54:33basically it helps you to manage

9:54:35everything up so botkeeper maintains lot

9:54:38of things for hbas it kind of helps you

9:54:41to basically kind of see that

9:54:43consistency is maintained it like in XB

9:54:47right when you work basically your in

9:54:49your big data when you're working in

9:54:51Hado now you have name Note data Note

9:54:54all those things available but when

9:54:55you're working with XB you don't have

9:54:57all those things right so in XB you

9:54:59don't have concepts of name Lo your

9:55:01keeping data so for that to do this part

9:55:05like you keep play a major role here so

9:55:08zeper is kind of going to act like a

9:55:10coordinator inside your hbas environment

9:55:14okay it will help you to coordinate all

9:55:16the things because here we don't have

9:55:17name node data nodes and all so and

9:55:20there's no Yar basically here so you can

9:55:22treat it like basically just like how

9:55:24Yan was handling things there it is

9:55:25going to help you as a coordinator it

9:55:28also maintains the directory

9:55:30structur can anybody tell me what is

9:55:33Bloom

9:55:33filter Bloom filter

9:55:38anyone know what is Bloom filter in Bas

9:55:41can I get this answer what is Bloom

9:55:45filter basically it helps to improve the

9:55:49overall throughput of your fluster it

9:55:52helps you to basically improve the

9:55:54performance okay now if you want to

9:55:56search any specific row column cells it

9:55:59also help you to do that so it makes a

9:56:01system very fast so that's that's what

9:56:04you need to just enable this and if if

9:56:06it is enabled it kind of includes the

9:56:08toput of your cluster this is the role

9:56:11of your blue filter in X

9:56:15space coming to next question what is

9:56:18the role of jdbc driver in a scoop setup

9:56:24simple can I get this answer this is

9:56:27basically scoop are important topic in

9:56:29interview HB I would still say that

9:56:31they're not very important these this

9:56:33people do not ask questions onb but

9:56:35definitely this schol topic is very

9:56:39important very good rdbs database I want

9:56:43to connect now basically dbms can be of

9:56:45any type it can be my SQL it can be uh

9:56:49it can be my SQL it can also be your

9:56:51Oracle DB it can be db2 right so jdbc

9:56:55driver will be common and can be used

9:56:58for any of the things so this will

9:57:00basically help you to create a

9:57:02connection with any sort of RPMs system

9:57:06this is a very famous question what's

9:57:09the difference between hyen hyen Target

9:57:12Di and the difference with Warehouse

9:57:15hyen

9:57:17di hyphen hyen Target di with warehous

9:57:24v this is a very famous interview

9:57:26question for

9:57:29scoop because both will basically help

9:57:31you to put the data in some SDF specific

9:57:35location then what's the difference

9:57:37between

9:57:38them very good not soting all table and

9:57:42all is fine that you giving me a use

9:57:44case I want basically a proper solution

9:57:48since what's the difference between

9:57:51them is Define no no

9:57:55R what happens in Target DS in Target

9:58:00you if you're defining Target di you

9:58:03need to give the directory path name

9:58:08okay so you need to give the directory

9:58:10name where you will be keeping data so

9:58:12it is possible let's say in your my SQL

9:58:14or in rdbms let's say you have table

9:58:17called as accounts but now you want to

9:58:21import this table to your hdfs using

9:58:24scoop Now by if you give Target Dr you

9:58:29need to tell the name of the hdfs

9:58:32directory where you will be keeping this

9:58:34so now you can change this name of hbfs

9:58:37from accounts to let's say accounts one

9:58:40you can do all that if you're using

9:58:41Target di you can change the name of

9:58:44this accounts to accounts one and ldfs

9:58:47but if you're using Warehouse V in that

9:58:50case whatever the name will be there in

9:58:53your rdbms same name will be created in

9:58:57your

9:58:58hdms same name will be created your HFS

9:59:01no change with that okay so that is one

9:59:04of the major difference between them so

9:59:07Warehouse V will maintain the same name

9:59:10while with Target di you can keep the

9:59:12same name or different name as well so

9:59:15you are forced to basically give the

9:59:18name of the hdfs folder where you want

9:59:22to import while in W that's a

9:59:26different

9:59:27now can you tell me what this quer is

9:59:30doing

9:59:31here read this query and tell me what

9:59:34this query is doing here

9:59:37incremental data no

9:59:40no importing employer table but can you

9:59:42see this hpe and and we

9:59:45Closs it is filtering right it is only

9:59:48filtering all the employees table where

9:59:52your start date is greater than this

9:59:54value okay this is what this is

9:59:58doing see this this is what this is like

10:00:01now let's say this is a question in a

10:00:03scop import command you have mentioned

10:00:06to run eight parall map produce CL but

10:00:09scope is only running four what can be

10:00:12the

10:00:13reason what can be the reason very good

10:00:17very good yes because maybe a number of

10:00:21qus are not allowing you to run eight

10:00:24par right maybe you have less number of

10:00:27Cs itself in that case will not be able

10:00:29to take you up to eight after first

10:00:31right it will only use less number of C

10:00:34so basically if your ques are less in

10:00:36that case this is bound to happen give a

10:00:39scope command to show all the databases

10:00:42in the mql server can you give a scoop

10:00:44command to show all the databases in my

10:00:47SQL Server it should be simple see this

10:00:51scoop list databases not show databases

10:00:55list databases they okay hyen hyen

10:00:58connect give the connections that's it

10:01:00will list all the

10:01:02databases basically that's it so this

10:01:04will list all the databas

10:01:06okay so those sessions will be very

10:01:09useful for all of you thank you everyone

10:01:11for making it interactive and a nice

10:01:12session I hope you have enjoyed

10:01:14listening to this video please be kind

10:01:17enough to like it and you can comment

10:01:19any of your doubts and queries and we

10:01:22will reply them at the earliest do look

10:01:24out for more videos in our playlist And

10:01:27subscribe to Eddie Raa channel to learn

10:01:29more happy

10:01:34learning

More from edureka!

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.