Full transcript
Introduction
0:09Hello everyone and welcome to the AI
0:11[music] and machine learning full
0:12course. Artificial intelligence and
0:15machine learning are rapidly changing
0:17the way organization [music] analyze
0:18data, automate processes, and build
0:21intelligent applications. [music]
0:23This course provides a complete
0:25introduction to a key ideas, techniques,
0:27and tools
0:28>> [music]
0:28>> used in modern AI and machine learning.
0:31Throughout the course, you will explore
0:33how machines [music] learn from data,
0:35understand the role of algorithms and
0:37models, and see how AI systems are
0:40applied to solve real-world problems.
0:41[music]
0:42The course is designed to build your
0:44understanding step-by-step, making
0:46complex concepts [music] easier to
0:48understand. And by the end of this
0:50course, you will have a strong
0:51foundation in AI and machine learning
0:54and a clear view [music] of how these
0:56technologies are driving innovation
0:57across industries. So, before we begin,
1:00please like, share, and subscribe to
1:01Edureka's YouTube channel and hit the
1:03bell icon to stay updated on the latest
1:05tech content [music] from Edureka. Also,
1:08check out Edureka's postgraduate program
1:10in generative [music] AI and machine
1:12learning in collaboration with Illinois
1:14Tech. It offers a unique opportunity to
1:16explore [music] the cutting-edge world
1:18of generative AI and develop advanced
1:20AI-powered solutions. This program
1:23[music] covers in-demand topics
1:24including machine learning, deep
1:26learning, natural language processing,
1:28[music] prompt engineering, generative
1:30AI, LLMs, RAG, agentic AI, and much
1:33more. Learn from industry [music]
1:35experts through a curriculum built
1:37around real-world hands-on use cases
1:39designed to equip you with a practical
1:41and job-ready skills. So, check out the
1:43course link given in the description box
1:45[music] below. Now, let us get started
What is Artificial Intelligence?
1:48by understanding what artificial
1:49intelligence is.
1:51AI or artificial intelligence is a
1:54branch of computer science where focused
1:56on creating systems that were performed
1:59task and that would be normally required
2:01by human intelligence. These tasks can
2:04range understanding of natural language.
2:07Secondly, recognizing patterns, then
2:10making decisions, and lastly, learning
2:12from the experiences.
2:14AI is a collaboration of ideas, methods,
2:18and knowledge where from the multiple
2:20academic disciplines work on a different
2:22problem-solving and share their
2:24knowledge to a better understanding and
2:26come up with a good solution.
2:28But, it can also be rule-based and
2:31operate under a set of rules and
2:33conditions only.
2:34So, I would like to tell you all guys a
2:37small real-time experience or an
2:39experiment done by Alan Turing to
2:42propagate and to establish the
2:44artificial intelligence.
2:46Alan Turing was the first person to
2:48conduct the sustainable research in the
2:50field that he called machine
2:52intelligence.
2:53The Turing test was conducted to explore
2:55whether machines could exhibit
2:57human-like intelligence.
2:59Proposed by Alan Turing in 1950, it
3:02involves a human evaluator communicating
3:05with both a human and a machine through
3:08a text interface.
3:09If a evaluator cannot distinguish
3:11between the two based on their
3:13responses, the machine is said to have
3:15passed the test. It serves as a
3:17benchmark of the assessing the progress
3:20of AI and discussions of the nature and
3:23intelligence and consciousness.
3:25So, this is to evolve and to establish
3:28that even machines can work as humans.
3:32And that is how it is made to bring up
3:35the machines' knowledge and the humans'
3:37knowledge into machines.
3:39There are two types of artificial
3:40intelligence.
3:42First, let's talk about weak AI. Before
3:44getting into what is weak AI, I would
3:46like to tell you with an example that is
3:49performed in a real world.
3:51I know you're all guys will be knowing
3:53about Alexa, Apple Siri, and also
3:55self-driving vehicles.
3:57So, all of these considered as in weak
4:00AI. It is also known as a narrow AI or
4:04an AI narrow intelligent that is trained
4:07by the AI and focused to perform a
4:09specific task. Weak AI drives most of
4:12the AI that surrounds by us today.
4:14Narrow might be a more apt descriptor
4:17for this type of AI as it is anything
4:19but weak.
4:20Next, let us learn about strong AI.
4:23Strong AI is made up of artificial
4:25general intelligence or artificial super
4:28intelligence.
4:30Simply, it could be told as where a
4:31machine would have a intelligence equal
4:34to humans.
4:35Where humans will track what AI needs to
4:37be done.
4:39It would be able to self-aware with the
4:41consciousness that would be this ability
4:43to solve problems, learn, and also plan
4:45for the future.
4:46Have you ever thought what does AI do at
4:49its core?
4:50I'm here to tell you what. AI is
4:52essential. It works by analyzing a lot
4:55of data
4:57to find patterns and a useful
4:58information.
4:59It even learns from this data to get
5:02better at the task over time. With this
5:04learning, AI can make decisions, predict
5:06future events, and do tasks
5:08automatically that would normally need
5:10human intelligence.
5:11This help business and other
5:13organizations work more efficiently.
5:16So, AI is about making computers smarter
5:19and more helpful in everyday life. So,
5:22let me tell you some use cases that is
5:24happening in daily uses or daily real
5:27life based.
5:28Firstly, I have taken is about
5:30cybersecurity. As it is a very important
5:33and a vital role for many platforms here
5:35after.
5:36Cybersecurity is a critical concern for
5:38individual businesses and also
5:40governments as cyber threats continue to
5:43evolve in a complexity and
5:44sophistication. It would play a very
5:47role for augmenting cybersecurity
5:49defenses.
5:50It is all based on the false positive
5:53effects that is made by the false
5:55information given by any criteria.
5:58Leading the alert of fatigue and reduced
6:01operational efficiencies.
6:03Cyber attacks can exploit
6:05vulnerabilities in AI models
6:08by invading detection and compromising
6:10security defenses.
6:12They have some private data where it is
6:15trained on the basis of incomplete data
6:18sets may produce the outcomes of private
6:21data. This will raise an ethical concern
6:24and also regulatory compliances issues.
6:27Having a very complicated problems, too.
6:29So, the next one will be your
6:31entertainment. As people know,
6:33entertainment is taking a very huge part
6:36in everyone's life.
6:37For example, it would be your social
6:39media, too.
6:40So, AI is transforming the entertainment
6:42industry by revolutionizing content
6:45creation, personalization, and audience
6:47engagement. So, this can be such as
6:50television, gaming, music, and also
6:52digital media platforms.
6:54One significant use of AI in
6:56entertainment is personalized content
6:59recommendation.
7:01So, personalized entertainment
7:02experiences are enhanced through AI
7:04driven. It is recommended through
7:06systems.
7:07They are even having some platforms like
7:10Netflix,
7:11Prime Amazon, and also Spotify
7:13that is making people engaged in a very
7:15hype as of now.
7:17These recommendations improve over time.
7:20So, I would like to conclude by telling
7:22artificial intelligence can do amazing
7:24things like analyzing data and making
7:27task easier.
7:28But, it also brings up huge question
7:31about fairness, jobs, and who controls
7:33it.
7:34We need to be careful in how we develop
7:36and use the AI.
7:37Making sure it helps everyone and
7:39doesn't cause harm at all.
7:41So, this makes easier for people to get
7:43into creativity and make rules and
7:45guidelines.
Types of Artificial Intelligence
7:49>> [music]
7:52>> So, now let's get started with the first
7:53topic, which is history of artificial
7:56intelligence.
7:57The concept of AI goes back to the
7:59classical ages. Under Greek mythology,
8:02the concept of machines and mechanical
8:04men were well thought of. An example is
8:07Talos. Talos was supposedly a giant
8:10animated bronze warrior who was
8:12programmed to guard the island of Crete.
8:15Now, let's get back to the 19th century.
8:17In 1950, Alan Turing proposed the Turing
8:20test. The Turing test basically
8:22determines whether or not a computer can
8:25intelligently think like a human being.
8:27The Turing test was the first serious
8:29proposal in the philosophy of artificial
8:31intelligence.
8:331951 marked the era for game artificial
8:36intelligence. This period was called
8:38game AI because here a lot of computer
8:40scientists developed programs for
8:42checkers and for chess. However, these
8:45programs were later rewritten and redone
8:47in a better way.
8:491956 marked the most important year for
8:52artificial intelligence.
8:54During this year, John McCarthy first
8:56coined the term artificial intelligence.
8:58This was followed by the first AI
9:00laboratory, which was set up in 1959.
9:03MIT AI Lab was the first setup, which
9:06was basically dedicated to the research
9:08of AI.
9:09In 1960, the first robot was introduced
9:12to the General Motors assembly line. In
9:151961, the first AI chatbot called Eliza
9:19was introduced. In 1997, IBM's Deep Blue
9:23beats the world champion Garry Kasparov
9:25in the game of chess.
9:272005 marks for the year when an
9:29autonomous robotic car called Stanley
9:32won the DARPA Grand Challenge.
9:34In 2011, IBM's question-answering
9:37machine Watson defeated the two greatest
9:40Jeopardy champions Brad Rutter and Ken
9:42Jennings. So, that was a brief history
9:45of AI. Now guys, since the emergence of
9:47artificial intelligence in 1950s, we
9:50have seen an exponential growth in its
9:52potential. AI covers domains such as
9:55machine learning, deep learning, neural
9:57networks, natural language processing,
9:59knowledge-based expert systems, and so
10:02on.
10:02Now that you know a brief history of
10:04artificial intelligence, let's move on
10:06and understand what exactly artificial
10:08intelligence is. So, the term artificial
10:11intelligence was first coined by John
10:13McCarthy, like I mentioned earlier. He
10:15defined AI as a science and engineering
10:18of making intelligent machines. In other
10:20words, artificial intelligence can also
10:23be defined as a development of computer
10:25systems that are capable of performing
10:28tasks that require human intelligence
10:30such as decision-making, object
10:32detection, solving complex problems, and
10:34so on. So, like I mentioned, artificial
10:37intelligence helps in decision-making,
10:39solving complex problems, it performs
10:42high-level computations, and also
10:44increases the accuracy of your
10:46predictions. Right? These are the main
10:48features of AI.
10:49So, now let's understand the different
10:51stages of artificial intelligence. So,
10:54basically, when I was doing my research,
10:55I found a lot of videos and a lot of
10:58articles that stated that artificial
11:00general intelligence, artificial narrow
11:03intelligence, and artificial super
11:04intelligence are the different types of
11:07AI. If I have to be more precise with
11:09you, then artificial intelligence has
11:11three different stages. Right? The types
11:13of AI are completely different from the
11:15stages of AI. So, under the stages of
11:17artificial intelligence, we have
11:19artificial narrow intelligence,
11:21artificial general intelligence, and
11:23artificial super intelligence.
11:25So, what is artificial narrow
11:26intelligence? Artificial narrow
11:28intelligence, also known as weak AI, is
11:31a stage of artificial intelligence that
11:33involves machines that can perform only
11:36a narrowly defined set of specific
11:38tasks. Right? At this stage, the
11:40machines don't possess any thinking
11:42ability. They just perform a set of
11:44predefined functions. Examples of weak
11:47AI include Siri, Alexa, AlphaGo, Sophia,
11:50the self-driving cars, and so on. Almost
11:53all the AI-based systems that are built
11:55till this date fall under the category
11:57of weak AI or artificial narrow
11:59intelligence.
12:01Next, we have something known as
12:02artificial general intelligence.
12:04Artificial general intelligence is also
12:06known as strong AI. This stage is the
12:09evolution of artificial intelligence,
12:11wherein machines will possess the
12:13ability to think and make decisions just
12:16like human beings. There are currently
12:18no existing examples of strong AI, but
12:21it's believed that we will soon be able
12:23to create machines that are as smart as
12:25human beings. Strong AI is actually
12:28considered a threat to human existence
12:30by many scientists. This includes
12:32Stephen Hawking. Stephen Hawking quoted
12:35that the development of full artificial
12:37intelligence could spell the end of
12:40human race. Moving on to our last stage,
12:42which is artificial super intelligence.
12:45Artificial super intelligence is that
12:47stage of AI when the capability of
12:49computers will surpass human beings.
12:52Artificial super intelligence is
12:54currently seen as a hypothetical
12:56situation as depicted in movies and
12:58science fiction books. You see a lot of
13:00movies which show that machines are
13:02taking over the world. All of that is
13:04artificial super intelligence. Now, I
13:06believe that machines are not very far
13:08from reaching the stage taking into
13:10consideration our current pace.
13:12However, such systems don't currently
13:14exist, right? We don't have any machine
13:16that is capable of thinking better than
13:19a human being or reasoning in a better
13:21way than a human. Artificial super
13:23intelligence, basically any robot that
13:24is much smarter than humans. Now, moving
13:27on to the different types of artificial
13:29intelligence. Based on the functionality
13:32of AI-based systems, artificial
13:34intelligence can be categorized into
13:36four types. The first type is reactive
13:38machines AI. This type of AI includes
13:41machines that operate solely based on
13:44the present data and take into
13:46consideration only the current
13:48situation.
13:49Reactive AI machines cannot form
13:51inferences from the data to evaluate any
13:54future actions. They can perform a
13:56narrowed range of predefined tasks.
13:59An example of reactive AI is the famous
14:02IBM chess program that beat the world
14:04champion Garry Kasparov.
14:06This is one of the most impressive AI
14:08machines built so far.
14:10Next, we have limited memory AI. Now,
14:13like the name suggests, limited memory
14:15AI can make informed and improved
14:17decisions by studying the past data from
14:20its memory. So, such an AI has a
14:22short-lived or you can say a temporary
14:25memory that can be used to store past
14:27experiences and hence evaluate your
14:29future actions.
14:31Self-driving cars are limited memory AI
14:33that use the data collected in the
14:35recent past to make immediate decisions.
14:38For example, self-driving cars use
14:40sensors to identify civilians that are
14:43crossing the road. They identify any
14:45steep roads or traffic signals and they
14:48use this to make better driving
14:49decisions. This also helps in preventing
14:52any future accidents. Next, we have
14:54something known as theory of mind
14:56artificial intelligence. The theory of
14:58mind AI is a more advanced type of
15:00artificial intelligence. This category
15:03is speculated to play a very important
15:06role in psychology. This type of AI will
15:08mainly focus on emotional intelligence
15:11so that human beliefs and thoughts can
15:13be better comprehended. The theory of
15:15mind AI has not been fully developed
15:18yet, but rigorous research is happening
15:20in this area.
15:21Moving on to our last type of artificial
15:23intelligence is the self-aware
15:25artificial intelligence.
15:27So guys, let us fold hands and pray that
15:29we don't reach the state of AI where
15:32machines have their own consciousness
15:34and become self-aware. This type of AI
15:37is a little far-fetched, but in the
15:38future achieving a stage of
15:40superintelligence might be possible.
15:43Geniuses like Elon Musk and Stephen
15:45Hawking have constantly warned us about
15:47evolution of AI.
15:49So guys, let me know your thoughts in
15:50the comment section. Do you ever think
15:52we'll reach the stage of artificial
15:54superintelligence?
15:56Moving on to the last topic of today's
15:58session is the different domains or the
16:00different branches of artificial
16:02intelligence. So artificial intelligence
16:04can be used to solve real-world problems
16:06by implementing machine learning, deep
16:08learning, natural language processing,
16:10robotics, expert systems, and fuzzy
16:13logic. Now guys, these are the different
16:15domains or you can say the different
16:16branches that AI uses in order to solve
16:19any problem. Recently, AI has also been
16:22used as an application in computer
16:24vision and image processing. Right, for
16:26now let me tell you briefly about each
16:28of these domains. Machine learning is
16:30basically the science of getting
16:32machines to interpret, process, and
16:34analyze data in order to solve
16:36real-world problems. Right, under
16:38machine learning there's supervised,
16:40unsupervised, and reinforcement
16:41learning. If any of you are interested
16:43in learning about these technologies,
16:45I'll leave a link in the description
16:46box. You all can go through that
16:47content. Next, we have deep learning or
16:50neural networks. So deep learning is a
16:52process of implementing neural networks
16:54on high-dimensional data to gain
16:57insights and form solutions. It is
16:59basically the logic behind the face
17:01verification algorithm on Facebook. It
17:04is the logic behind the self-driving
17:06cars, virtual assistants like Siri and
17:08Alexa. Then we have natural language
17:10processing. Natural language processing
17:12refers to the science of drawing
17:14insights from natural human language in
17:16order to communicate with machines and
17:19grow businesses. So, an example of NLP
17:22is Twitter and Amazon. Twitter uses NLP
17:25to filter out terroristic language in
17:27their tweets. Amazon uses NLP to
17:30understand customer reviews and improve
17:32user experience. Then we have robotics.
17:35Robotics is a branch of artificial
17:37intelligence which focuses on the
17:39different branches and applications of
17:41robots. AI robots are artificial agents
17:44which act in the real world environment
17:47to produce results by taking some
17:49accountable actions.
17:51So, I'm sure all of you have heard of
17:52Sophia. Sophia the humanoid is a very
17:55good example of AI in robotics. Then we
17:58have fuzzy logic. So, fuzzy logic is a
18:00computing approach that is based on the
18:02principle of degree of truth instead of
18:05the usual modern logic that we use which
18:08is basically the Boolean logic. Fuzzy
18:10logic is used in medical fields to solve
18:12complex problems which involve decision
18:15making. It is also used in automating
18:18gear systems in your cars and all of
18:20that. Then we have expert systems. An
18:22expert system is an AI-based computer
18:24system that learns and reciprocates the
18:27decision-making ability of a human
18:29expert. Expert systems use if-then logic
18:32notions in order to solve any complex
18:34problem. They do not rely on
18:37conventional procedural programming.
18:39Expert systems are mainly used in
18:41information management. They're seen to
18:43be used in fraud detection, virus
18:46detection, also in managing medical and
18:48hospital records, and so on. So, guys,
18:50to sum it up, these were the different
18:52branches of artificial intelligence.
18:55>> [music]
AI vs Machine Learning vs Deep Learning
19:00>> These are the term which have confused a
19:02lot of people. And if you too are one
19:04among them, let me resolve it for you.
19:07Well, artificial intelligence is a
19:09broader umbrella under which machine
19:11learning and deep learning come. You can
19:13also see in the diagram that even deep
19:15learning is a subset of machine
19:17learning. So, you can say that all three
19:19of them, the AI, the machine learning,
19:21and deep learning, are just the subset
19:24of each other. So, let's move on and
19:26understand how exactly they differ from
19:28each other. So, let's start with
19:30artificial intelligence. The term
19:32artificial intelligence was first coined
19:35in the year 1956.
19:37The concept is pretty old, but it has
19:39gained its popularity recently. But why?
19:42Well, the reason is earlier we had very
19:45small amount of data. The data we had
19:48was not enough to predict the accurate
19:50result. But now, there's a tremendous
19:52increase in the amount of data.
19:54Statistics suggest that by 2020, the
19:57accumulated volume of data will increase
20:00from 4.4 zettabytes to roughly around 44
20:03zettabytes, or 44 trillion GBs of data.
20:06Along with such enormous amount of data,
20:09now we have more advanced algorithm and
20:12high-end computing power and storage
20:14that can deal with such large amount of
20:15data.
20:16As a result, it is expected that 70% of
20:19enterprise will implement AI over the
20:21next 12 months, which is up from 40% in
20:242016 and 51% in 2017.
20:28Just for your understanding, what is AI?
20:31Well, it's nothing but a technique that
20:33enables the machine to act like humans
20:35by replicating the behavior and nature.
20:38With AI, it is possible for machine to
20:40learn from the experience. The machines
20:43adjust their responses based on new
20:45input, thereby performing human-like
20:47tasks.
20:48Artificial intelligence can be trained
20:50to accomplish specific tasks by
20:51processing large amount of data and
20:53recognizing pattern in them.
20:55You can consider that building an
20:57artificial intelligence is like building
20:59a church. The first church took
21:01generations to finish. So, most of the
21:04workers who were working on it never saw
21:06the final outcome. Those working on it
21:08took pride in their crafts, building
21:10bricks and chiseling stone that was
21:12going to be placed into the great
21:13structure. So, as AI researchers, we
21:16should think of ourselves as humble
21:18brick makers whose job is to study how
21:21to build components, example parsers,
21:23planners, or learning algorithm, or
21:25etc., anything that someday someone and
21:27somewhere will integrate into the
21:29intelligent systems. Some of the
21:31examples of artificial intelligence from
21:33our day-to-day life are Apple series,
21:36chess-playing computer, Tesla's
21:38self-driving car, and many more. These
21:40examples are based on deep learning and
21:42natural language processing.
21:44Well, this was about what is AI and how
21:46it gained its height. So, moving on
21:48ahead, let's discuss about machine
21:50learning and see what it is and why it
21:53was even introduced. Well, machine
21:55learning came into existence in the late
21:57'80s and the early '90s. But, what were
21:59the issues with the people which made
22:01the machine learning come into
22:02existence? Let us discuss them one by
22:05one.
22:05In the field of statistics, the problem
22:08was how to efficiently train large
22:10complex model. In the field of computer
22:12science and artificial intelligence, the
22:14problem was how to train more robust
22:16version of AI system. While in the case
22:18of neuroscience, problem faced by the
22:20researchers was how to design
22:22operational model of the brain.
22:24So, these were some of the issues which
22:26had the largest influence and led to the
22:28existence of the machine learning.
22:30Now, this machine learning shifted its
22:32focus from the symbolic approaches it
22:33had inherited from the AI and moved
22:36towards the methods and model it had
22:38bought from statistics and probability
22:40theory.
22:41So, let's proceed and see what exactly
22:43is machine learning. Well, machine
22:45learning is a subset of AI which enables
22:48the computer to act and make data-driven
22:50decisions to carry out a certain task.
22:52These programs or algorithms are
22:54designed in a way that they can learn
22:56and improve over time when exposed to
22:58new data. Let's see an example of
23:00machine learning. Let's say you want to
23:02create a system which tells the expected
23:04weight of a person based on its height.
23:07The first thing you do is you collect
23:08the data. Let's see, this how your data
23:10looks like. Now, each point on the graph
23:13represent one data point. To start with,
23:16we can draw a simple line to predict the
23:18weight based on the height. For example,
23:20a simple line W equal H minus 100, where
23:23W is weight in kg and H is height in cm.
23:27This line can help us to make the
23:28prediction. Our main goal is to reduce
23:31the difference between the estimated
23:32value and the actual value. So, in order
23:35to achieve it, we try to draw a straight
23:37line that fits through all these
23:39different points and minimize the error.
23:41So, our main goal is to minimize the
23:43error and make them as small as
23:45possible. Decreasing the error or the
23:47difference between the actual value and
23:49estimated value increases the
23:50performance of the model. Further on,
23:53the more data points we collect, the
23:55better our model will become. We can
23:56also improve our model by adding more
23:58variables and creating different
24:00prediction lines for them. Once the line
24:02is created, so from the next time if we
24:04feed a new data, for example, height of
24:06a person to the model, it would easily
24:08predict the data for you and it will
24:10tell you what its predicted weight could
24:12be. I hope you got a clear understanding
24:14of machine learning. So, moving on
24:16ahead, let's learn about deep learning.
24:18Now, what is deep learning? You can
24:20consider deep learning model as a rocket
24:22engine and its fuel is its huge amount
24:25of data that we feed to these
24:26algorithms.
24:27The concept of deep learning is not new.
24:30But recently, it's hype has increased
24:32and deep learning is getting more
24:33attention.
24:34This field is a particular kind of
24:36machine learning that is inspired by the
24:38functionality of our brain cells called
24:39neuron, which led to the concept of
24:42artificial neural network.
24:44It simply takes the data connection
24:45between all the artificial neurons and
24:47adjust them according to the data
24:49pattern. More neurons are added if the
24:51size of the data is large. It
24:53automatically features learning at
24:55multiple levels of abstraction, thereby
24:57allowing a system to learn complex
24:59function mapping without depending on
25:01any specific algorithm. You know what?
25:04No one actually knows what happens
25:06inside a neural network and why it works
25:08so well. So, currently you can call it
25:10as a black box. Let [snorts] us discuss
25:12some of the example of deep learning and
25:14understand it in a better way. Let me
25:16start with a simple example and explain
25:18you how things happen at a conceptual
25:21level. Let us try and understand how you
25:23recognize a square from other shapes.
25:26The first thing you do is you check
25:28whether there are four lines associated
25:30with the figure or not. Simple concept,
25:32right? If yes, we further check if they
25:35are connected and closed. Again, if yes,
25:37we finally check whether it is
25:39perpendicular and all its sides are
25:41equal. Correct? If everything fulfills,
25:44yes, it is a square.
25:46Well, it is nothing but a nested
25:47hierarchy of concepts.
25:50What we did here, we took a complex task
25:52of identifying a square in this case and
25:54broke it into simpler task. Now, this
25:56deep learning also does the same thing
25:58but at a larger scale. Let's take an
26:01example of machine which recognizes the
26:03animal. The task of the machine is to
26:05recognize whether the given image is of
26:07a cat or of a dog.
26:09What if we were asked to resolve the
26:10same issue using the concept of machine
26:12learning? What we would do? First, we
26:15would define the features such as check
26:17whether the animal has whiskers or not
26:19or check if the animal has pointed ears
26:21or not or whether its tail is straight
26:23or curved. In short, we will define the
26:25facial features and let the system
26:27identify which features are more
26:29important in classifying a particular
26:31animal. Now, when it comes to deep
26:34learning, it takes this to one step
26:35ahead. Deep learning automatically finds
26:38out the feature which are most important
26:40for classification compared to machine
26:42learning where we had to manually give
26:44out that features.
26:46By now, I guess you have understood that
26:48AI is a bigger picture and machine
26:50learning and deep learning are its
26:51subpart. So, let's move on and focus our
26:53discussion on machine learning and deep
26:55learning.
26:56The easiest way to understand the
26:58difference between the machine learning
26:59and deep learning is to know that deep
27:01learning is machine learning. More
27:03specifically, it is the next evolution
27:05of machine learning. Let's take few
27:07important parameter and compare machine
27:09learning with deep learning. So,
27:11starting with data dependencies. The
27:13most important difference between deep
27:15learning and machine learning is its
27:17performance as the volume of the data
27:19gets increased. From the below graph,
27:21you can see that when the size of the
27:23data is small, deep learning algorithm
27:25doesn't perform that well. But, why?
27:28Well, this is because deep learning
27:30algorithm needs a large amount of data
27:32to understand it perfectly.
27:34On the other hand, the machine learning
27:36algorithm can easily work with smaller
27:38data set. Fine?
27:40Next comes the hardware dependencies.
27:42Deep learning algorithms are heavily
27:44dependent on high-end machines, while
27:46the machine learning algorithm can work
27:48on low-end machines as well.
27:50This is because the requirement of deep
27:52learning algorithm include GPUs, which
27:55is an integral part of its working.
27:57The deep learning algorithm require GPUs
27:59as they do a large amount of matrix
28:01multiplication operations, and these
28:03operations can only be efficiently
28:06optimized using a GPU as it is built for
28:09this purpose only.
28:10Our third parameter will be feature
28:12engineering. Well, feature engineering
28:15is a process of putting the domain
28:17knowledge to reduce the complexity of
28:19the data and make patterns more visible
28:21to learning algorithms.
28:23This process is difficult and expensive
28:25in terms of time and expertise.
28:28In case of machine learning, most of the
28:29features are needed to be identified by
28:31an expert and then hand coded as per the
28:34domain and the data type. For example,
28:37the features can be a pixel value,
28:38shapes, texture, position, orientation,
28:41or anything. Fine? The performance of
28:44most of the machine learning algorithm
28:46depends on how accurately the features
28:48are identified and extracted.
28:50Whereas in case of deep learning
28:52algorithms, it try to learn high-level
28:54features from the data. This is a very
28:56distinctive part of deep learning, which
28:57makes it way ahead of traditional
28:59machine learning.
29:01Deep learning reduces the task of
29:03developing new feature extractor for
29:04every problem. Like in the case of CNN
29:07algorithm, it first try to learn the
29:09low-level features of the image, such as
29:11edges and lines, and then it proceeds to
29:13the parts of faces of people, and then
29:16finally to the high-level representation
29:17of the face. I hope the things are
29:19getting clear to you.
29:21So, let's move on ahead and see the next
29:23parameter. So, our next parameter is
29:25problem-solving approach.
29:27When we are solving a problem using
29:29traditional machine learning algorithm,
29:30it is generally recommended that we
29:33first break down the problem into
29:34different sub parts, solve them
29:36individually, and then finally combine
29:38them to get the desired result. This is
29:41how the machine learning algorithm
29:42handles the problem. On the other hand,
29:45the deep learning algorithm solves the
29:46problem from end to end.
29:48Let's take an example to understand
29:50this.
29:51Suppose you have a task of multiple
29:52object detection, and your task is to
29:54identify what is the object and where it
29:57is present in the image. So, let's see
29:59and compare how will you tackle this
30:01issue using the concept of machine
30:03learning and deep learning.
30:04Starting with machine learning, in a
30:06typical machine learning approach, you
30:08would first divide the problem into two
30:10step. First, object detection and then
30:13object recognition. First of all, you'd
30:16use a bounding box detection algorithm
30:18like GrabCut for example,
30:20to scan through the image and find out
30:22all the possible objects. Now, once the
30:25objects are recognized, you'd use object
30:27recognition algorithm like SVM with HOG,
30:31to recognize relevant objects.
30:33Now, finally when you combine the
30:35result, you would be able to identify
30:37what is the object and where it is
30:38present in the image.
30:40On the other hand, in deep learning
30:42approach, you would do the process from
30:44end to end. For example, in a YOLO net,
30:46which is a type of deep learning
30:48algorithm, you would pass an image and
30:50it would give out the location along
30:52with the name of the object. Now, let's
30:54move on to our fifth comparison
30:56parameter.
30:57It's execution time.
30:59Usually, a deep learning algorithm takes
31:01a long time to train. This is because
31:03there are so many parameter in a deep
31:05learning algorithm that makes the
31:06training longer than usual. The training
31:09might even last for 2 weeks or more than
31:11that if you're training completely from
31:13the scratch. Whereas in the case of
31:15machine learning, it relatively takes
31:17much less time to train, ranging from a
31:19few weeks to few hours.
31:21Now, the execution time is completely
31:23reversed when it comes to the testing of
31:25data. During testing, the deep learning
31:28algorithm takes much less time to run.
31:30Whereas if you compare it with a KNN
31:32algorithm, which is a type of machine
31:33learning algorithm, the test time
31:35increases as the size of the data
31:36increase.
31:38Last but not the least, we have
31:39interpretability as a factor for
31:41comparison of machine learning and deep
31:43learning. This factor is the main reason
31:46why deep learning is still thought 10
31:48times before anyone uses it in the
31:50industry. Let's take an example. Suppose
31:53we use deep learning to give automated
31:56scoring to essays. The performance it
31:58gives in scoring is quite excellent and
32:00is near to the human performance. But
32:02there's an issue with it. It does not
32:04reveal why it has given that score.
32:06Indeed, mathematically it is possible to
32:09find out that which node of a deep
32:11neural network were activated, but we
32:13don't know what the neurons are supposed
32:15to model and what these layers of neuron
32:17were doing collectively.
32:19So, we failed to interpret the result.
32:21On the other hand, machine learning
32:22algorithm like decision tree gives us a
32:25crisp rule for why it chose and what it
32:27chose. So, it is particularly easy to
32:30interpret the reasoning behind it.
32:32Therefore, the algorithms like decision
32:33tree and linear or logistic regression
32:36are primarily used in industry for
32:38interpretability.
Artificial Intelligence with Python
32:45So, why exactly are we using Python for
32:47artificial intelligence? Why aren't we
32:49using any other language? Right? Now,
32:52there are a couple of reasons as to why
32:54Python is so popular when it comes to
32:56AI, machine learning, and deep learning.
32:58The first reason is less coding is
33:00required. Now, artificial intelligence
33:02has a lot of algorithms. If you have to
33:04implement AI in any code or in any
33:07problem, then there are going to be tons
33:09and tons of machine learning algorithms
33:11involved, deep learning algorithms
33:13involved, right? Now, testing all of
33:15these can become a very tiresome task.
33:18That's where Python usually comes in
33:20handy. Now, the language has something
33:23known as check as you code methodology,
33:26which eases the process of testing,
33:28right? You can check your program as you
33:30code it. Basically, as you're typing
33:32each sentence, your errors or your any
33:34sort of mistakes in your code will be
33:36given to you. Right? So, testing becomes
33:38much easier when it comes to Python.
33:40The next important reason why we're
33:42choosing Python is it has support for
33:44pre-built libraries. Right? Python is
33:47very convenient for AI developers
33:50because all of the algorithms, machine
33:52learning algorithms, and deep learning
33:53algorithms are already predefined in
33:56libraries, right? So, you don't have to
33:57actually sit down and code each and
33:59every algorithm. That would take a lot
34:01of time. And that's a very
34:03time-consuming task. And thanks to
34:05Python, you don't have to do that
34:06because they have libraries and packages
34:09that have all the algorithms built in
34:11them, right? So, if you want to run any
34:14algorithm, all you have to do is you
34:15have to call the function and load the
34:17library. That's all. It's as simple as
34:19that. Now, the next reason is ease of
34:21learning. So, guys, Python is actually
34:24the most simplest programming language,
34:26right? If you ask me, I think it is is
34:28the most easiest programming language.
34:30It's very similar to English language,
34:32right? If you read a couple of lines in
34:34Python, you'll understand what exactly
34:36the code is doing. It has a very simple
34:38syntax, and this simple syntax can be
34:41implemented to solve simple problems
34:43like addition of two strings, and it can
34:45also be used to solve complex problems
34:48like building machine learning models
34:50and deep learning models. So, ease of
34:52learning is a major factor when it comes
34:54to why Python is chosen for artificial
34:57intelligence, right? Next, we have
34:59platform independent.
35:01So, good thing about Python is that you
35:02can get your project running on
35:04different operating systems, right? And
35:07what happens when you transfer your code
35:09from one operating system to another
35:11operating system is we find a lot of
35:13dependency issues. To solve that, Python
35:16has a couple of packages such as there
35:18is a package known as PyInstaller,
35:20right? This PyInstaller will take care
35:22of all the dependency issues when you're
35:24transferring your code from one platform
35:26to the other platform. So, all of this
35:29support is provided by Python. The last
35:31reason is massive community support.
35:33This is a very important point because
35:36it is important that you have a large
35:38community that will help you out with
35:40any errors or with any sort of problems
35:43in your code, right? So, Python has
35:46several communities and several forums
35:48and groups on Facebook. So, if you have
35:50any doubts regarding any error, you can
35:53just post those errors in these groups,
35:55and you'll have like a bunch of people
35:56helping you out. Right? So, guys, these
35:59are a couple of reasons as to why Python
36:01is chosen for artificial intelligence.
36:03It's actually considered the most
36:04popular and the most used language for
36:07data science, AI, machine learning, and
36:09deep learning. To prove that to you,
36:11here is a stat from Stack Overflow.
36:13Stack Overflow recently stated that
36:16Python is the fastest growing
36:17programming language. If you look at the
36:19graph, you can see that it has taken
36:21over JavaScript and Java and C#, C++,
36:25and PHP, right? So, Python is actually
36:28growing at an exponential rate,
36:30especially when it comes to data science
36:32and artificial intelligence. A lot of
36:34developers are very comfortable with the
36:35Python language because, you know, it's
36:37a general-purpose language, first of
36:38all. So, most of the developers are
36:40already aware of Python. And then, using
36:43the same language in order to solve
36:45complex problems like artificial
36:47intelligence, machine learning, and deep
36:49learning is something every developer
36:51wants, right? They want a simple
36:52language in order to code all the
36:54complex algorithms or the complex
36:56models. Right? So, that's why Python is
36:59the best choice for artificial
37:00intelligence.
37:02For those of you who are not aware of
37:03Python programming and don't know much
37:05about Python, I'm going to leave a
37:07couple of links in the description box.
37:09Right? You can go through those links
37:11and study a little bit more about how
37:12Python works or how the coding part
37:15works. Right? I'm going to be focusing
37:17mainly on artificial intelligence, and
37:19I'll be showing you a lot of demos. So,
37:21those of you are not aware of Python,
37:23make sure you check the description box.
37:25Right?
37:26Next, I'm going to discuss the different
37:28Python packages for artificial
37:30intelligence. Now, these are the
37:31packages that are specifically for
37:33machine learning, deep learning, natural
37:35language processing, and so on. So,
37:37let's take a look at all these packages.
37:39So, first, we have TensorFlow. If you
37:42are currently working on a machine
37:44learning project in Python, then you
37:46must have heard of this popular
37:48open-source library known as TensorFlow.
37:51Right? This library was developed by
37:52Google in collaboration with Brain team.
37:55TensorFlow is used in almost every
37:57Google application for machine learning.
37:59Now, let me just discuss a few features
38:01of TensorFlow. It has a responsive
38:03construct, meaning that with TensorFlow,
38:06we can easily visualize each and every
38:08part of the graph, which is not an
38:10option when you're using other packages
38:12such as NumPy or scikit. Right? Another
38:15feature is that it's very flexible. Now,
38:17one of the most important TensorFlow
38:19features is that it is flexible in
38:22operability. Meaning that it has
38:24modularity and the parts of which you
38:27want to make standalone, it offers you
38:29that option. Right? It's very flexible
38:31in that way. It'll give you exactly what
38:33you want. Now, good feature about
38:35TensorFlow is that you can train it on
38:37both CPU and GPU. Right? So, for
38:39distributed computing, you can have both
38:42these options. Also, it supports
38:44parallel neural network training. So,
38:46TensorFlow offers pipelining in the
38:49sense that you can train multiple neural
38:51networks and multiple GPUs, which makes
38:54the models very efficient on any
38:56large-scale system. Right? So, parallel
38:58neural network training is supported by
39:00TensorFlow. Right? This is one of the
39:03most important features of TensorFlow.
39:05Apart from this, it has a very large
39:07community. And needless to say, if it
39:09has been developed by Google, then
39:11there's already a large team of software
39:13engineers who work on stability,
39:16improvements, and all of that. Right?
39:19The next library I'm going to talk about
39:20is scikit-learn.
39:22Now, scikit-learn is a Python library
39:24that is associated with NumPy and SciPy.
39:26Right? That's why it has the name
39:28scikit-learn. Now, this is considered to
39:30be one of the best uh libraries for
39:32working with complex data. And there are
39:34a lot of changes that are being made in
39:36this library. And one modification is
39:39the cross-validation feature, which
39:41provides the ability to use more than
39:43one metric. Right? Cross-validation is
39:46one of the most important and one of the
39:47most easiest methods for checking the
39:50accuracy of a model. Right? So,
39:51cross-validation is being implemented in
39:53scikit-learn. And apart from that,
39:56again, there are a large spread of
39:57algorithms that you can implement by
39:59using scikit-learn. Right? These include
40:01unsupervised learning algorithms,
40:03starting from clustering, factor
40:05analysis, principal component analysis,
40:08to all the unsupervised neural networks.
40:11Scikit-learn is also very essential uh
40:13for feature extracting in images and
40:16text.
40:17So, mainly scikit-learn is used for
40:19implementing all the standard machine
40:21learning and data mining tasks like
40:24reducing dimensionality, classification,
40:26regression, clustering, and model
40:28selection. Next up, we have NumPy. Now,
40:30NumPy is considered as one of the most
40:33popular machine learning libraries in
40:35Python. Now, let me tell you that
40:36TensorFlow and other libraries, they
40:38make use of NumPy internally for
40:41performing multiple operations on
40:43tensors. The most important feature of
40:47NumPy is the array interface. It
40:49supports multi-dimensional arrays.
40:51Right? That's one of the most important
40:53features of NumPy. Another feature is uh
40:55it makes complex mathematical
40:57implementations very simple. Right? It's
41:00mainly known for computing mathematical
41:03data. So, NumPy is a package that you
41:05should be using for any sort of
41:07statistical analysis or data analysis
41:10that involves a lot of math. Apart from
41:12that, it makes coding very easy and
41:15grasping the concept is extremely easy
41:16with NumPy. Now, NumPy is mainly used
41:20for expressing images, sound waves, and
41:22other mathematical computations.
41:24All right? Moving on to our next
41:26library, we have Theano. Theano is a
41:29computational framework which is used
41:32for computing multi-dimensional arrays.
41:34Right? Theano actually works very
41:36similar to TensorFlow, but the only
41:38drawback is that you can't fit Theano
41:41into production environments. But apart
41:43from that, Theano allows you to define,
41:46optimize, and evaluate mathematical
41:48expressions that involve
41:49multi-dimensional arrays. Right? This is
41:52another library that lets you implement
41:54multi-dimensional arrays. Features of
41:56Theano include tight integration with
41:58NumPy. An advantage of Theano is that
42:01you can easily implement NumPy arrays in
42:03Theano. Right? That's why there's a
42:05connection between Theano and NumPy
42:07because both of them effectively use
42:09multi-dimensional arrays. Transparent
42:11use of GPU. Now, performing data
42:14intensive computations are much faster
42:16when it comes to uh Theano because of
42:18its use of GPU, right? Theano also lets
42:21you detect and diagnose multiple types
42:24of errors and any sort of ambiguity in
42:27the model. So, guys, Theano was actually
42:29designed to handle the types of
42:31computations required for large neural
42:34network algorithms, right? It was mainly
42:36built for deep learning and neural
42:38networks. It was one of the first
42:41libraries of its kind and it is
42:43considered as an industry standard for
42:46deep learning research and development.
42:48Theano is being used in multiple neural
42:50networks projects and the popularity of
42:52Theano is only going to grow with time,
42:54right? A lot of people actually haven't
42:56heard of Theano, but let me tell you
42:58that this is one of the best ways to
42:59implement deep learning and neural
43:01network models.
43:03Moving on, uh we have Keras. Now, Keras
43:05is considered to be the most popular
43:08Python package. It provides some of the
43:10best functionalities for compiling
43:12models, processing your data sets, and
43:15visualizing graphs. It is also popular
43:17in the implementation of neural
43:19networks, right? It is considered to be
43:21the simplest package uh with which you
43:23can implement neural networks. In fact,
43:25in our today's demo for deep learning,
43:27we'll be implementing Keras in order to
43:29understand how neural networks work. Few
43:32of the features of Keras include that it
43:34runs very smoothly on both CPU and GPU.
43:37It supports almost all the models of the
43:40neural network, right? From fully
43:42connected, convolutional, pooling,
43:44recurrent, embedding, all of these
43:46models are supported by Keras.
43:48And not only that, you can combine these
43:50models to build more complex models.
43:53Keras is completely Python-based, which
43:55makes it very easy to debug and explore,
43:58right? Since Python has a huge community
44:00of followers, it's very simple in order
44:03to debug any sort of error that you find
44:05while implementing Keras. So, the
44:07libraries that I discussed so far were
44:09dedicated to machine learning and deep
44:11learning. For natural language
44:13processing, we have the most famous
44:15library known as the Natural Language
44:17Toolkit, which is an open-source Python
44:19library, mainly used for natural
44:21language processing, text analysis, and
44:24text mining. The main features include
44:26that it studies and analyzes natural
44:28language text in order to draw useful
44:31information from all this natural
44:32language text. It performs text analysis
44:36and sentimental analysis by performing
44:38tasks such as stemming, lemmatization,
44:41tokenization, and so on. Now, don't
44:44worry if you don't know what any of
44:45those terms mean. I'll be discussing all
44:47of those terms with you by the end of
44:49today's session. So guys, these were a
44:51couple of Python-based libraries, which
44:53are very essential for implementing
44:55machine learning and deep learning and
44:58artificial intelligence when you're
45:00using Python, right? These libraries are
45:02perfect for implementing AI. So guys, if
45:05any of you have any doubts regarding the
45:07libraries or if you want to learn more
45:09about the libraries, I will leave a
45:11couple of links in the description box.
45:13You can go through those videos as well.
45:15So now, let's move on to the main topic
45:17of discussion, which is artificial
45:18intelligence.
45:20Now, before we get started with the
45:22demand of artificial intelligence, let
45:25me tell you that AI was invented long
45:27ago. AI goes back to the 19th century.
45:29It was not something that was recently
45:31invented, even though AI has recently
45:33gained a lot of popularity. We can say
45:36that in the past decade, AI has gained
45:39the maximum popularity. But, it was
45:41actually invented in the 19th century.
45:44Now, especially in the year 1950, there
45:46was somebody known as Alan Turing. I'm
45:48sure a lot of you have heard about the
45:50Turing test. The Turing test is
45:52basically used to determine whether or
45:54not a machine is artificially
45:57intelligent, meaning that whether a
45:59machine can think intelligently like a
46:01human being.
46:02Right? This was the first proposition
46:04and this was one of the most important
46:07breakthroughs in artificial
46:08intelligence. Right? Somebody known as
46:10Alan Turing, he published a landmark
46:13paper in which he speculated about the
46:15possibility of creating machines that
46:17think. Right? So, the Turing test was
46:20the first serious proposal in the
46:22philosophy of artificial intelligence.
46:24This was done in 1950. Right? After
46:27this, we had eras of AI. We had the game
46:30AI which was in 1951. Now, since the
46:33emergence of AI in 1950s, we have seen
46:37an exponential growth in its potential.
46:39Right? AI covers domains like machine
46:41learning, deep learning, neural
46:43networks, natural language processing,
46:45knowledge base, and so on. It's also
46:47made its way into computer vision and
46:49image processing. But, the question is
46:51if AI has been here for over half a
46:54century,
46:55why has it suddenly gained so much
46:57importance? Right? Why are we talking
47:00about artificial intelligence now? The
47:02main reasons for the vast popularity of
47:04AI are the following. Right? The first
47:06reason is more computational power. Now,
47:09AI requires a lot of computing power.
47:12Recently, many advances have been made
47:14and complex deep learning models can be
47:16deployed. And one of the greatest
47:18technology that made this possible are
47:20GPUs. Since the invention of GPUs, we
47:23can compute much more with our
47:25computers. Initially, we could barely
47:27process 1 GB of data. Right? We only had
47:30hard disk to store additional memory and
47:32all of that. Now, our computers can
47:34process tons and tons of data. So, now
47:37we have more computational power, which
47:39is one of the main reasons behind why AI
47:41became so popular. So, by having more
47:44computational power, it becomes much
47:46easier to implement artificial
47:47intelligence. Next reason is more data.
47:50Now, big data is one of the most
47:52important reasons behind the development
47:55of artificial intelligence. Now, AI and
47:58data science and machine learning, deep
48:00learning, all of these processes are
48:03here only because we have a lot of data
48:05at present. Now, the main idea behind
48:08all these technologies is to draw useful
48:10insights from data. Now, since we start
48:12generating a lot of data, we need to
48:14find a method that can process this much
48:17data and draw useful insights from data
48:20such that it benefits an organization or
48:23it grows a business. That's why
48:25artificial intelligence and machine
48:27learning comes into the picture. Right?
48:28So, more data led to the demand of
48:31artificial intelligence. Apart from
48:33this, we also have better algorithms
48:35now, right? We have state-of-the-art
48:37algorithms. Most of them are based on
48:40the idea of neural networks and these
48:42are constantly getting better. Neural
48:43networks are actually one of the most
48:45significant discoveries in artificial
48:48intelligence because with neural
48:50networks, you can take in thousand
48:52layers of input data. Right? You can
48:54take in a lot of input data to perform
48:56computations. So, through neural
48:58networks, we are actually able to solve
49:00a lot of problems including healthcare
49:02problems, fraud detection problems, and
49:04so on. Another reason is broad
49:06investment. So, our universities and
49:09governments and startups and any tech
49:12giants like Google, Amazon, and
49:14Facebook, they are all investing heavily
49:16in artificial intelligence, which also
49:18led to the demand of AI. So, AI is
49:21rapidly growing both as a field of study
49:23and also as an economy. Right? It's
49:26adding a lot to the economy and I think
49:29this is the perfect time for you to get
49:30into the field of artificial
49:31intelligence because right now AI is in
49:34a really high demand. AI, machine
49:36learning, data science, all of this are
49:38of really high demand at present. All
49:41right. So, this is the perfect time for
49:42you to get started with artificial
49:43intelligence. Now, let me tell you that
49:45the term artificial intelligence was
49:47first coined in the year 1956
49:50by a scientist known as John McCarthy.
49:53Now, John McCarthy defined artificial
49:56intelligence as the science and
49:57engineering of making intelligent
50:00machines. So, now let's move on and talk
50:02about how artificial intelligence is
50:05different from machine learning and deep
50:06learning. A lot of people uh tend to
50:08assume that artificial intelligence,
50:10machine learning, and deep learning are
50:12the same because they have common
50:14applications, right? For example, Siri
50:17is an application of AI, machine
50:19learning, and deep learning. So, how are
50:22these technologies uh related, right? Or
50:24how are they different from each other?
50:26Now, artificial intelligence is the
50:28science of getting machines to mimic the
50:30behavior of human beings. Machine
50:33learning is the subset of artificial
50:36intelligence that focuses on getting
50:38machines to make decisions by feeding
50:40them data. Deep learning, on the other
50:43hand, is a subset of machine learning
50:45that uses the concept of neural networks
50:48to solve complex problems.
50:50So, to sum it up to you, artificial
50:52intelligence, machine learning, and deep
50:54learning are heavily interconnected
50:56fields, right? Machine learning and deep
50:58learning aids artificial intelligence by
51:00providing a set of algorithms and neural
51:03networks to solve data-driven problems.
51:05However, AI is not restricted to only
51:08machine learning and deep learning,
51:09right? It covers a vast domain of fields
51:12which include natural language
51:13processing, object detection, computer
51:15vision, robotics, expert systems, and so
51:18on, right? So, AI is a very vast field.
51:20Guys, I hope I cleared the difference
51:22between AI, machine learning, and deep
51:24learning. Also, a lot of you might be
51:26confused about data science. Data
51:28science is now an umbrella term, right?
51:31Data science basically means to derive
51:33useful insights from data. So, data
51:35science actually uh uses AI, machine
51:38learning, and deep learning, right? So,
51:40it implements all of these three
51:42technologies in order to derive useful
51:44insights from data, right? Now, let's
51:47move on to the most interesting topic in
51:49artificial intelligence, which is
51:51machine learning. Now guys, the term
51:53machine learning was first coined by a
51:55scientist known as Arthur Samuel in the
51:58year 1959.
51:59Looking back, that year was probably the
52:02most significant in terms of
52:03technological advancements.
52:05In order to define machine learning, if
52:08you browse the internet for what is
52:10machine learning, you'll get at least
52:11100 different definitions.
52:13In simple terms, machine learning is a
52:16subset of artificial intelligence, which
52:18provides machines the ability to learn
52:21automatically and improve from
52:23experience without being explicitly
52:26programmed to do so. In a sense, it is
52:29the practice of getting machines to
52:31solve problems by gaining the ability to
52:33think. Now the question here is, can a
52:36machine think or can a machine make
52:38decisions? Well, if you feed a machine a
52:41good amount of data, it will learn how
52:43to interpret, process, and analyze this
52:46data by using something known as machine
52:48learning algorithms. To give you a basic
52:51idea of how the machine learning process
52:53works, look at the figure on this slide.
52:56A machine learning process always begins
52:58by feeding the machine lots and lots of
53:00data. Now by using this data, the
53:03machine is trained to detect any hidden
53:05insights and trends in the data. These
53:08insights are then used to build a
53:10machine learning model by using a
53:12machine learning algorithm in order to
53:14solve a problem. The basic aim of
53:17machine learning is to solve a problem
53:19or find a solution by using data. Now
53:22moving ahead, I'll be discussing the
53:23machine learning process in depth,
53:25right? So don't worry if you haven't got
53:27the exact idea of what machine learning
53:29is.
53:30Now the machine learning process
53:32involves building a predictive model
53:34that can be used to find a solution for
53:37a particular problem. A well-defined
53:39machine learning process will have
53:41around seven steps. It always begins
53:44with defining the objective followed by
53:46data gathering or data collection. Then
53:49we have something known as preparing
53:51data, which is also called data
53:53pre-processing. Then we have data
53:55exploration or exploratory data
53:57analysis. This is followed by building a
54:00machine learning model.
54:02Then we have model evaluation and
54:04finally predictions. This is how the
54:06process of machine learning works. To
54:08understand the machine learning process,
54:10let's assume that you've been given a
54:12problem that needs to be solved by using
54:14machine learning. Let's say that the
54:16problem is to predict the occurrence of
54:18rain in your local area by using machine
54:21learning. Now, the first step is to
54:23define the objective of the problem.
54:25Right? At this step we must understand
54:27what exactly needs to be predicted. In
54:29our case, the objective is to predict
54:31the possibility of rain by studying the
54:34weather conditions. So, at this stage it
54:36is essential to take mental notes on
54:39what kind of data can be used to solve
54:41this problem or the type of approach
54:43that you must follow to get to the
54:45solution.
54:46The questions you should be asking
54:48yourself is what are we trying to
54:50predict? Right? Here we're trying to
54:51predict whether it'll rain or not.
54:54Right? You need to understand what are
54:56the target features. Target features are
54:59basically the variable that you need to
55:01predict. Here we need to predict a
55:03variable that'll show us whether it's
55:04going to rain tomorrow or not. Then you
55:07must also understand what kind of data
55:09you'll need to solve this problem. Apart
55:11from that, you need to know what kind of
55:13problem you're facing. Is it a binary
55:15classification problem or is it a
55:17clustering problem? Now, if you don't
55:19know what classification and clustering
55:21is, don't worry. I'll be talking about
55:23all of these things in the upcoming
55:24slides. So, your first step is to define
55:27the objective of your problem. You need
55:30to understand what exactly needs to be
55:32done here. Right? How can you solve this
55:34problem?
55:35Moving on, your next step is to gather
55:37the data that you need. At this stage,
55:39you must be asking questions such as
55:42what kind of data is needed to solve
55:44this problem. Is the data available to
55:46me? And if it's not available, how can I
55:48get the data? Right? Once you know the
55:51type of data that is required, you must
55:53understand how you can derive this data.
55:56Data collection can be either done
55:58manually or it can be done by web
56:00scraping. But don't worry if you're a
56:02beginner and you're just looking to
56:04learn machine learning, you don't have
56:05to worry about getting the data.
56:08There are thousands of data resources on
56:10the web. You can just download the data
56:12set and you can get going.
56:14Coming back to the problem at hand, the
56:16data needed for weather forecasting
56:18includes measures such as humidity
56:20level, your temperature, the pressure,
56:23the locality, whether or not you live in
56:26a hill station, and so on.
56:28Such data must be collected and it has
56:30to be stored for analysis. This is where
56:33you collect all the data. Now, moving on
56:35to step number three is data
56:37preparation. The data that you collected
56:40is almost never in the right format. All
56:43right, even if you collect it from a
56:45internet resource, if you download it
56:47from some website, even then your data
56:50is not going to be clean. Right? It's
56:52not going to be in the correct format.
56:53There's always going to be some sort of
56:55inconsistencies in your data.
56:58Inconsistencies include any missing
57:00values or any redundant variables,
57:03duplicate values. All of these are
57:05inconsistencies.
57:07Removing all of this is very essential
57:09because they might lead to any wrongful
57:11computation. Therefore, at this stage,
57:13you can scan the entire data set for any
57:16missing values and you have to fix them
57:18here itself.
57:20Now, actually, this is one of the most
57:21time-consuming steps in a machine
57:23learning process. If you ask a data
57:26scientist which step he hates the most
57:28or which step is, you know, the most
57:30time-consuming, they're probably going
57:32to tell you data processing and data
57:33cleaning. Right? It's one of the most
57:36tiresome task because you need to look
57:38at all the values that are there. You
57:39need to find any missing values, any
57:41data that is not relevant to you. Right?
57:43All of this has to be removed so that
57:45you can analyze the data in a better
57:47way.
57:48Now, step number four is exploratory
57:50data analysis.
57:52So, guys, this stage is all about
57:54getting deep into your data and finding
57:57all the hidden data mysteries.
58:00EDA or exploratory data analysis is like
58:03the brainstorming stage of machine
58:05learning.
58:06Data exploration involves understanding
58:08the patterns and the trends in your
58:09data. So, at this stage all the useful
58:12insights are drawn and any correlations
58:15between the variables are understood.
58:17For example, in the case of predicting
58:19rainfall, we know that there is a strong
58:22possibility of rain if the temperature
58:24has fallen low. Such correlations have
58:27to be understood and mapped at this
58:29stage.
58:30EDA is actually the most important step
58:32in a machine learning process because
58:34here is where you understand your data.
58:37You understand how your data is going to
58:39help you predict the outcome.
58:41Moving on to step number five, we have
58:44building a machine learning model. So,
58:46all the insights and all the patterns
58:49that you got from your data exploration
58:51stage, those insights are used to build
58:53the machine learning model. So, this
58:55stage always begins by splitting the
58:57data set into two parts, that is
59:00training and testing data. Now, remember
59:02that the training data will be used to
59:05build and analyze the model.
59:07The model is basically the machine
59:09learning algorithm that predicts the
59:11output by using the data that you feed
59:13to it. An example of machine learning
59:15algorithm is logistic regression and
59:18linear regression. All of these are
59:20machine learning algorithms.
59:22Now, don't worry about choosing the
59:23right algorithm. Right? First, we'll
59:25focus on what the machine learning
59:26process is.
59:28But anyway, choosing the right algorithm
59:30will depend on several factors, right?
59:32It depends on the type of problem you're
59:34trying to solve, the data set, and the
59:36level of complexity of the problem.
59:39In the upcoming sections, we'll discuss
59:40all the different types of problems that
59:42can be solved by using machine learning.
59:44Moving on to step number six, we have
59:46model evaluation and optimization.
59:49Now, after you build a model by using
59:51the training data set, it is finally
59:53time to put the model to a test. The
59:56testing data set is used to check the
59:58efficiency of the model and how
1:00:00accurately it can predict the outcome.
1:00:03Now, once the accuracy is calculated and
1:00:06any further improvements in the model,
1:00:08they have to be implemented at this
1:00:10stage. Methods like parameter tuning and
1:00:13cross-validation can be used to improve
1:00:15the performance of the model.
1:00:17Before I move any further, I don't know
1:00:19if all of you know what training and
1:00:21testing data set means. In machine
1:00:23learning, the input data is always
1:00:25divided into two sets. We have something
1:00:28known as the training data set, and we
1:00:29have something known as the testing data
1:00:31set.
1:00:32So, in machine learning, you always
1:00:34split the data into two parts, right?
1:00:36This process is known as a data
1:00:38splicing. Now, the training data set
1:00:40will be used to build the machine
1:00:42learning model, and the testing data set
1:00:44will be used to test the efficiency of
1:00:47the model that you built. This is what
1:00:49training and testing data set is.
1:00:51They're not any different data that you
1:00:53derive. They're the same as the input
1:00:55data set. The only thing is you are
1:00:57splitting the data set so that you can
1:00:59train the model on one data and test the
1:01:01model on another data.
1:01:03Now, remember that the training data set
1:01:05is always larger in size when compared
1:01:08to the testing data set. Because
1:01:10obviously, you are training and building
1:01:12the model by using the training data
1:01:13set. The testing data set is just for
1:01:16evaluating the performance of your
1:01:18model.
1:01:19Now, let's move on and understand step
1:01:21number seven, which is predictions. Now,
1:01:24once a model is evaluated and you've
1:01:26improved the model, it is finally used
1:01:29to make predictions. The final output
1:01:31can be a categorical variable or it can
1:01:34be a continuous quantity. Right? All of
1:01:36this depends on the type of problem
1:01:38you're trying to solve. Don't worry,
1:01:39I'll be discussing the type of problems
1:01:41that can be solved using machine
1:01:43learning in the upcoming slides. In our
1:01:45case for predicting the occurrence of
1:01:47rainfall, the output will be a
1:01:49categorical variable.
1:01:51Categorical variable is anything that
1:01:53has some categorical value. For example,
1:01:56gender is a categorical variable. Gender
1:01:59has either male, female, or other.
1:02:02It has a defined set of values. That is
1:02:04a categorical variable. So guys, that
1:02:06was the entire machine learning process.
1:02:09Now, as we continue with this tutorial,
1:02:11in the upcoming sections, I will be
1:02:14running a demo in Python, in which we
1:02:16will be performing weather forecasting.
1:02:18So, make sure you remember all these
1:02:20steps that I spoke about because I'll be
1:02:21going through all these steps by using
1:02:24Python. We'll be coding all of this that
1:02:26we just spoke about.
1:02:28Now, the next topic we're going to
1:02:29discuss is the types of machine
1:02:31learning.
1:02:32A machine can learn to solve a problem
1:02:34by following any one of the three
1:02:37approaches.
1:02:38You can say that there are three ways in
1:02:40which a machine learns. The three ways
1:02:43are supervised learning, unsupervised
1:02:45learning, and reinforcement learning.
1:02:48These are the three methods in which you
1:02:50can train a machine to learn.
1:02:52So first, let's discuss supervised
1:02:54learning.
1:02:55So, what is supervised learning?
1:02:57Supervised learning is a technique in
1:02:59which we teach or train the machine by
1:03:01using data which is labeled. To
1:03:04understand this better, let's consider
1:03:06an analogy.
1:03:07As kids, we all needed guidance to solve
1:03:10math problems. At least I had a really
1:03:12tough time solving math problems. Yeah,
1:03:15so our teachers always helped us
1:03:17understand what addition is and how it
1:03:19is done. Similarly, you can think of
1:03:21supervised learning as a type of machine
1:03:24learning that involves a guide. The
1:03:26label data set is a teacher that will
1:03:28train the machine to understand the
1:03:30patterns in the data. The label data set
1:03:33is nothing but the training data set.
1:03:35So, to better understand this, consider
1:03:37the figure. Right here, we're feeding
1:03:40the machine images of Tom and Jerry, and
1:03:42the goal is for the machine to identify
1:03:45and classify the images into two
1:03:47separate groups. Basically, one group
1:03:49will contain Tom images, and the other
1:03:51group will contain images of Jerry. Now,
1:03:54pay attention to the training data set.
1:03:56The training data set that is fed to a
1:03:58model is labeled. As in, we're telling
1:04:01the machine, "Listen, this is how Tom
1:04:03looks, and this is how Jerry looks."
1:04:05But, basically, labeling each data point
1:04:08that we're feeding to the machine.
1:04:10Right? If the image is of Tom's, we've
1:04:12labeled it as Tom, and if the image is a
1:04:15Jerry image, then we're going to label
1:04:17it as Jerry. By doing this, you're
1:04:19training the machine by using labeled
1:04:22data. So, to sum it up, in supervised
1:04:24learning, there is a well-defined
1:04:26training phase done with the help of
1:04:28labeled data. Right? The rest of the
1:04:30process is the same. After you feed the
1:04:33machine labeled data, you're going to
1:04:34perform data cleaning, then exploratory
1:04:37data analysis, followed by building the
1:04:39machine learning model, and then model
1:04:41evaluation, and finally, your
1:04:43predictions. Also, one more point to
1:04:45remember is that the output that you're
1:04:47going to get in a supervised learning
1:04:49algorithm is a labeled output. This
1:04:52Jerry will be labeled as Jerry, and this
1:04:54Tom will be labeled as Tom. Basically,
1:04:56you'll get a labeled output. Now, let's
1:04:59understand what is unsupervised
1:05:01learning.
1:05:02Unsupervised learning involves training
1:05:05by using unlabeled data and allowing the
1:05:07model to act on that information without
1:05:10any guidance. So, think of unsupervised
1:05:13learning as a smart kid that learns
1:05:15without any guidance. In this type of
1:05:17machine learning, the model is not fed
1:05:20with any label data. As in, the model
1:05:22has no clue that this image is Tom and
1:05:25this image is Jerry. It figures out
1:05:28patterns and the differences between Tom
1:05:30and Jerry on its own by taking in tons
1:05:32of data. For example, it identifies
1:05:35prominent features of Tom such as pointy
1:05:38ears, bigger in size, and so on to
1:05:40understand that this image is of type
1:05:42one.
1:05:43Similarly, it finds such features in
1:05:45Jerry and knows that this is another
1:05:47type of image, [clears throat] maybe
1:05:48type two. Right? Therefore, it
1:05:50classifies the images into two different
1:05:52clusters without knowing who is Tom and
1:05:55who is Jerry. Now, the main idea behind
1:05:57unsupervised learning is to understand
1:05:59the patterns in your data set and form
1:06:01clusters based on feature similarity.
1:06:04Basically, it'll feature similar images
1:06:06or similar data points into one cluster,
1:06:09and it'll form another cluster which is
1:06:11totally different from the first
1:06:13cluster. So, look at the output over
1:06:15here. The unlabeled output is basically
1:06:17clusters or groups of two different
1:06:20data.
1:06:21Next, we have something known as the
1:06:22reinforcement learning. Now,
1:06:24reinforcement learning is comparatively
1:06:26different, right? It's pretty different
1:06:28from supervised and unsupervised.
1:06:31It is basically a part of machine
1:06:32learning where you put an agent in an
1:06:35environment, and this agent learns to
1:06:37behave in the environment by performing
1:06:40certain actions and observing the
1:06:42rewards which it gets from these
1:06:44actions. To understand reinforcement
1:06:46learning, imagine that you were dropped
1:06:49off at an isolated island. What would
1:06:51you do? Initially, we'd all panic,
1:06:54right? But, as time passes by, you will
1:06:56learn how to live on the island. You
1:06:59will explore the environment. You will
1:07:01understand the climate conditions.
1:07:03You'll understand the type of food that
1:07:05grows there. You'll know what is
1:07:07dangerous to you and what is not. You'll
1:07:09understand which food is good for you
1:07:11and which is not. This is exactly how
1:07:13reinforcement learning works. It
1:07:15involves an agent, which is basically
1:07:17you stuck on the island, that is put in
1:07:20an unknown environment, which is the
1:07:22island, where the agent must learn by
1:07:24observing and performing actions that
1:07:27result in rewards.
1:07:28Reinforcement learning is mainly used in
1:07:30advanced machine learning areas such as
1:07:33self-driving cars, AlphaGo, and so on.
1:07:35So, guys, that sums up the types of
1:07:37machine learning. Before we go any
1:07:39further, I'd like to discuss the
1:07:40difference between supervised,
1:07:42unsupervised, and reinforcement
1:07:43learning. Now, first of all, we have the
1:07:45definition. Supervised learning is all
1:07:48about teaching a machine by using
1:07:50labeled data. Unsupervised learning,
1:07:53like the name suggests, there is no
1:07:55supervision over here. The machine is
1:07:57trained on unlabeled data without any
1:08:00guidance.
1:08:01Reinforcement learning is totally
1:08:02different. Here, you have an agent who
1:08:04interact with the environment by
1:08:06producing actions and discover some
1:08:09errors and rewards.
1:08:10Now, the type of problem that is solved
1:08:12using supervised learning is regression
1:08:14and classification problems. We'll
1:08:16discuss what regression, classification,
1:08:18and clustering is in the upcoming slide,
1:08:21right? So, don't worry if you don't know
1:08:22what it is. Unsupervised learning is
1:08:24mainly to solve association and
1:08:26clustering problems.
1:08:28Reinforcement learning is for
1:08:29reward-based problems.
1:08:31Now, what is the type of data in
1:08:33supervised learning? It is labeled data.
1:08:35That is the main difference between
1:08:37supervised and any other type of machine
1:08:39learning.
1:08:40In supervised, you have labeled data. In
1:08:42unsupervised, we have unlabeled data.
1:08:44Whereas, in reinforcement learning, we
1:08:46have no predefined data at all. The
1:08:49machine has to perform everything from
1:08:51scratch. It has to collect data,
1:08:53analyze, do everything on its own. Now,
1:08:55the training in supervised learning is
1:08:57external supervision, meaning that we
1:09:00have external supervision in the form of
1:09:02the labeled training data set. In
1:09:05unsupervised, there is obviously no
1:09:07supervision. There is an unlabeled data
1:09:09set, therefore there's no supervision.
1:09:11In reinforcement learning, there is no
1:09:13supervision at all. Now, the approach to
1:09:15solving supervised learning problem is
1:09:17basically you're going to map your
1:09:19labeled input to your known output. In
1:09:22unsupervised learning, the machine is
1:09:23going to understand the patterns and
1:09:25discover the output on its own.
1:09:27Reinforcement learning, here the agent
1:09:30will follow something known as a trial
1:09:32and error method. Right? It's totally
1:09:34based on the concept of trial and error.
1:09:37Popular algorithms under supervised
1:09:39learning are linear regression, logistic
1:09:42regression, support vector machines, and
1:09:44so on. Under unsupervised learning, we
1:09:46have the famous K-means clustering
1:09:48algorithm. Under reinforcement learning,
1:09:50we have the Q-learning algorithm, which
1:09:52is one of the most important algorithms.
1:09:55It is basically the logic behind the
1:09:57famous AlphaGo game. I'm sure all of you
1:09:59have heard of that. So, guys, these were
1:10:01the differences between supervised,
1:10:03unsupervised, and reinforcement
1:10:04learning. Now, let's move on and discuss
1:10:07the type of problems that you can solve
1:10:09by using machine learning.
1:10:11Now, there are three types of problems
1:10:13in machine learning. Now, any problem
1:10:15that needs to be solved in machine
1:10:17learning can fall into one of these
1:10:19three categories. Now, what is a
1:10:22regression? In this type of problem, the
1:10:24output is a continuous quantity. For
1:10:27example, if you want to predict the
1:10:29speed of a car given the distance. That
1:10:32means it is a regression problem.
1:10:34First of all, what is a continuous
1:10:36quantity? A continuous quantity is any
1:10:38variable that can hold a continuous
1:10:40value. A continuous variable is any
1:10:43variable that can have infinite number
1:10:45of values.
1:10:47For example, the height of a person or
1:10:49the weight of a person is a continuous
1:10:51quantity. Right? I can have a weight of
1:10:5350.1 kgs or 50.12 or 50.112 kg.
1:10:59This is a continuous quantity.
1:11:01Regression problems can be solved by
1:11:03using supervised learning algorithms.
1:11:06Another type of problem is a
1:11:07classification problem. Here, the output
1:11:10is always a categorical value.
1:11:13Classifying emails into two classes, for
1:11:15example, classifying your email as spam
1:11:17and non-spam is a classification
1:11:19problem.
1:11:20Here again, you'll be using supervised
1:11:22learning classification algorithms such
1:11:24as support vector machines, naive bias,
1:11:27logistic regression, and so on.
1:11:29Then we have clustering problem. And
1:11:31this type of problem involves assigning
1:11:34the input into two or more clusters
1:11:36based on feature similarity. For
1:11:39example, clustering the viewers into
1:11:41similar groups based on their interest
1:11:44or based on their age or geography can
1:11:46be done by using unsupervised learning
1:11:48algorithms like K-means clustering.
1:11:51One thing you need to understand is
1:11:53under supervised learning, you can solve
1:11:55regression and classification problems.
1:11:58Under unsupervised learning, you can
1:11:59solve clustering problems. Reinforcement
1:12:02learning is something else altogether,
1:12:05right? You can solve reward-based
1:12:06problems and more complex and deep
1:12:08problems. So, now let's move on and
1:12:12understand the different machine
1:12:13learning algorithms. Now, I will not be
1:12:16going into depth for machine learning
1:12:18algorithms because there are a lot of
1:12:19algorithms to cover, but we have content
1:12:22around almost every machine learning
1:12:24algorithm out there. So, I'll be leaving
1:12:26a couple of links in the description
1:12:28box, right? You can check out all these
1:12:30links and understand how each of these
1:12:32machine learning algorithms work in
1:12:34depth. So, I'm just going to show you a
1:12:36hierarchical diagram of how the
1:12:38algorithms are structured. So, under
1:12:40machine learning, we have three types of
1:12:42learning. We have supervised,
1:12:43unsupervised, and reinforcement. Under
1:12:45supervised learning, we have regression
1:12:47and classification problems. And under
1:12:50unsupervised learning, we have
1:12:51clustering problems. Reinforcement
1:12:54learning is completely different. I'll
1:12:55be leaving a link in the description
1:12:57specifically for reinforcement learning.
1:12:59You can check out the entire content of
1:13:01reinforcement learning there.
1:13:03Now, regression problems can be solved
1:13:05by using linear regression algorithm
1:13:07such as linear regression, decision
1:13:10trees, and random forest can also be
1:13:11used in regression problems. But,
1:13:14usually decision trees and random
1:13:16forest, all of these are used to solve
1:13:18classification problems. Famous
1:13:20classification algorithms include
1:13:22K-nearest neighbor, which is basically
1:13:24KNN, decision trees and random forest,
1:13:27logistic regression, naive bias, support
1:13:30vector machines. All of these are
1:13:32classification algorithms. Coming to
1:13:34unsupervised learning, we have
1:13:36clustering and association analysis. And
1:13:39clustering problems can be solved by
1:13:40using K-means. And association analysis
1:13:43can be solved by using a priori
1:13:45algorithm. A priori algorithm is mainly
1:13:48used in market basket analysis. Right?
1:13:51For this algorithm as well, I'll be
1:13:52leaving a link in the description. We've
1:13:55performed a very excellent demo where in
1:13:57we've shown how market basket analysis
1:13:59can be done by using a priori algorithm.
1:14:02Markov model is also explained in one of
1:14:05the videos. I'll be leaving that link in
1:14:07the description box.
1:14:08Now, to sum up machine learning to you,
1:14:10I'll be running a small demonstration in
1:14:13Python. Right? Like I promised earlier,
1:14:16I'll be using Python to understand the
1:14:18whole machine learning process. All
1:14:20right. So, let's get started with that
1:14:21demo.
1:14:22So guys, for those of you who don't know
1:14:24Python, I will leave a couple of links
1:14:26in the description box so that you
1:14:27understand Python. But, apart from that,
1:14:30Python is pretty understandable. If you
1:14:31just look at the code, you'll know what
1:14:33exactly I'm talking about. Right? So,
1:14:35don't worry. And also, I'll be
1:14:36explaining everything in the code.
1:14:39So, I'm using PyCharm in order to run
1:14:42the demo.
1:14:43Right? So guys, like I said, if you
1:14:44don't know Python, I'll leave a couple
1:14:46of links in the description box. You can
1:14:48go through those videos as well. The
1:14:51main aim of our demo is to build a
1:14:53machine learning model that will predict
1:14:55whether or not it will rain tomorrow by
1:14:58studying the past data set. Now, this
1:15:00data set contains around 145,000
1:15:03observations on the daily weather
1:15:05conditions as observed in Australia.
1:15:08Right, the data set has around 24
1:15:10features and we will be using 23
1:15:13features out of that to predict the
1:15:15target variable which is rain tomorrow.
1:15:18So, this data set I collected from
1:15:19Kaggle. Right, for those of you don't
1:15:21know Kaggle is a online platform where
1:15:23you can find hundreds of data sets and
1:15:26you know, there are a lot of
1:15:26competitions held by machine learning
1:15:28engineers and all of that. It's an
1:15:31interesting website.
1:15:33Now, the problem statement itself is to
1:15:35build a machine learning model that will
1:15:37predict whether or not it will rain
1:15:39tomorrow.
1:15:40This is clearly a classification
1:15:42problem. The machine learning model has
1:15:44to classify the output into two classes,
1:15:47that is either yes or no. Yes will stand
1:15:50for it will rain tomorrow and no will
1:15:52basically denote that it will not rain
1:15:54tomorrow. Right, this is a
1:15:55classification problem.
1:15:57So, I hope the objective is clear.
1:15:59Right, so we'll begin the demonstration
1:16:01by importing the required libraries. So,
1:16:04first of all, for mathematical
1:16:06computations, we'll be importing the
1:16:07NumPy library. We'll also be importing
1:16:10the Pandas library for data processing.
1:16:13Next, we will load the CSV file.
1:16:15Basically, my data is stored in a CSV
1:16:18format in this file. weatherAUS.csv is
1:16:21my data set.
1:16:22So, basically I've saved this file in
1:16:24this path. Right, so that's what I'm
1:16:26doing here. I'm loading my data set and
1:16:28I'm storing it in a variable known as
1:16:31DF.
1:16:32Next, what we'll do is we'll see the
1:16:33size of our data frame.
1:16:36Let's print the size of the data frame.
1:16:38We'll also display the first five
1:16:40observations in our data frame.
1:16:42Let's look at the output.
1:16:45Basically, around 145,000
1:16:47observations and 24 features. Now, 24
1:16:51features are basically the variables
1:16:53that are there in my data set. You know,
1:16:55for example, date is a variable,
1:16:57location is a variable, minimum
1:16:59temperature till rain tomorrow. All of
1:17:01these are variables. So, I have around
1:17:0224 features in my data set, right? Now,
1:17:05the variable that I have to predict is
1:17:07rain tomorrow. Okay? If the value of
1:17:10rain tomorrow is no, it denotes that it
1:17:12will not rain tomorrow. But, if the
1:17:14value is yes, then it will denote that
1:17:16it will rain tomorrow.
1:17:18So, rain tomorrow is basically my target
1:17:20variable.
1:17:21Right? I'll be finding out whether it's
1:17:23going to rain tomorrow or not. So, this
1:17:25is my target variable, also known as
1:17:27your output variable.
1:17:28My input variables will be the other 23
1:17:31variables. Date, location, minimum
1:17:33temperature, rain today, risk, all of
1:17:36this will be my input variables. Now,
1:17:38these variables are also known as
1:17:40predictor variables. Basically, they're
1:17:42used to predict your outcome. So, these
1:17:44are also known as predictor variables.
1:17:47Now, the next step is checking for null
1:17:49values.
1:17:50This is basically data pre-processing.
1:17:53Let me just comment it for you.
1:17:55This is data preparation or data
1:18:01Right. So, this stage is data
1:18:03pre-processing. Right here, we start
1:18:05checking for any null values or any
1:18:07missing values. This is exactly what I'm
1:18:09doing over here. I am checking for any
1:18:11missing or null values in my data set.
1:18:14If you notice the output, it shows that
1:18:16the first four columns have more than
1:18:1840% null values. Right? So, it's always
1:18:21best for us to remove features or such
1:18:23variables because they will not help us
1:18:25in our prediction.
1:18:26Now, during data pre-processing, it is
1:18:28always necessary to remove the variables
1:18:31that are not significant. Unnecessary
1:18:33data will just increase our
1:18:35computations. That's why it's always
1:18:37best if you remove the unwanted or
1:18:39unnecessary variables.
1:18:41Now, apart from removing these four
1:18:42variables, we'll also remove the
1:18:44location variable, and we will remove
1:18:47the date variable. Right, I'll come to
1:18:49this variable in a minute. We'll also be
1:18:51removing location and date variable
1:18:54because both of these variables are not
1:18:56needed in order to predict whether it'll
1:18:57rain tomorrow. Right, we do not need to
1:19:00know the location and the date.
1:19:02Now, we'll also be removing this a risk
1:19:04mm variable. Risk mm variable basically
1:19:07tells us the amount of rain that might
1:19:09occur the next day.
1:19:11Right, now this is a very informative
1:19:13variable, and it might actually leak
1:19:15some information to our model.
1:19:17By using this variable, we'll easily be
1:19:19able to predict, and there's no point of
1:19:21doing that. So, this variable will give
1:19:23us too much information. And so, that's
1:19:25why we're going to remove this variable
1:19:26as well. It'll leak a lot of
1:19:28information. So, after that, if you
1:19:30print the shape of your data frame,
1:19:33we have only 17 variables and so many
1:19:37observations.
1:19:38Now, after this, we'll just uh look at
1:19:40any null values, and we'll remove them.
1:19:43This drop.any function will just remove
1:19:45all the null values. Right, then if you
1:19:47print the shape of your data frame,
1:19:49we'll have around 112,000 rows with 17
1:19:53variables. This is the shape of the data
1:19:55set after removing all the null values
1:19:57and all the redundant or unnecessary
1:20:00variables.
1:20:01Now, it's time to remove the outliers in
1:20:03the data. So, after you remove any null
1:20:05values, we should also check our data
1:20:07set for any outliers. An outlier is a
1:20:11data point that is very different from
1:20:13your other observations.
1:20:15Outliers usually occur because of
1:20:17miscalculations while collecting the
1:20:19data.
1:20:20These are some sort of errors in your
1:20:22data set.
1:20:23So, in this whole code snippet, we're
1:20:25just getting rid of outliers.
1:20:29This is the output that we get. All our
1:20:31outliers.
1:20:32Next, what we'll be doing is we will be
1:20:35assigning zeros and ones in the place of
1:20:38yes and no.
1:20:39The only thing is we're going to change
1:20:40the categorical variables from yes and
1:20:42no to zero and one. Right, that's
1:20:44exactly what we're doing over here.
1:20:46Now, if there are any unique values such
1:20:48as any character values which are not
1:20:50supposed to be there, we'll be changing
1:20:52them into integer values.
1:20:54That's all we're doing over here.
1:20:56After this, we'll be normalizing our
1:20:58data set.
1:20:59This is a very important step because in
1:21:02order to avoid any biasness in your
1:21:04output, you have to normalize your input
1:21:06variables.
1:21:08Right, to do this, we can make use of
1:21:09the min-max scalar function which Python
1:21:11provides in a package known as
1:21:13scikit-learn. You can use that package
1:21:16in order to normalize your data set.
1:21:18So, after normalizing our data set, this
1:21:20is what our data set looks like.
1:21:23This is before normalization. You can
1:21:25see that these are in two digits,
1:21:27whereas these values are in single
1:21:29digits. Right? This causes a lot of
1:21:30biasness. But once we normalize the
1:21:33values, we know that all of the values
1:21:35are in a similar range. We have
1:21:37everything in decimals. Right, so
1:21:39normalization is something that has to
1:21:41be performed because if you have a data
1:21:43set like this, your output is not going
1:21:44to be correct. And that's why we perform
1:21:46normalization.
1:21:48So, now that we are done with
1:21:50pre-processing, what we're going to do
1:21:51is it's time for exploratory data
1:21:54analysis.
1:21:55Let me just comment it for you. This is
1:21:58exploratory data analysis.
1:22:01So, basically here, what we're going to
1:22:02do is we're going to analyze and
1:22:04identify the significant variables that
1:22:07will help us predict the outcome. To do
1:22:09this, we'll be using the select key best
1:22:12function which is present in the
1:22:13scikit-learn library.
1:22:15There's a predefined function in Python
1:22:17called select key best, which will
1:22:19basically select the most significant
1:22:21predictor variables in our data set.
1:22:24When we run that line of code,
1:22:26we get these three variables to be the
1:22:29most significant variables in our data
1:22:31set. Right, the main aim of this demo is
1:22:33to make you understand how machine
1:22:35learning works. That's why to simplify
1:22:37the competition, we'll assign only one
1:22:39of these significant variables as the
1:22:42input. Instead of taking all three
1:22:44variables as input, we'll select one
1:22:46variable and we'll take that as the
1:22:48input and the output is the rain
1:22:50tomorrow variable.
1:22:52So, basically, we are creating a data
1:22:54frame of all the significant variables.
1:22:56Basically, we're choosing this variable
1:22:58in order to predict our outcome.
1:23:00Obviously, our outcome is rain tomorrow
1:23:02variable.
1:23:03So, our input is humidity level and our
1:23:06output is to detect whether it'll rain
1:23:08tomorrow.
1:23:09The next step is data modeling. All of
1:23:11you are aware of what data modeling is.
1:23:13To solve this, we'll be using
1:23:15classification algorithms over here.
1:23:17We'll use logistic regression. We will
1:23:20use random forest classifier, which is
1:23:23another machine learning algorithm.
1:23:25We'll also use the decision tree
1:23:26classifier and support vector machine.
1:23:30Right, we'll be using all of these
1:23:31algorithms in order to predict the
1:23:33outcome. We'll also check which
1:23:35algorithm gives us the best accuracy.
1:23:38So guys, we're just using multiple
1:23:40algorithms or multiple classification
1:23:42algorithms on the same data set. We're
1:23:44not doing anything very complex over
1:23:46here.
1:23:47So, we start by importing all the
1:23:48necessary libraries for the logistic
1:23:50regression algorithm.
1:23:52We're also going to import time because
1:23:54we'll be calculating the accuracy and
1:23:56the time taken by the algorithm to get
1:23:58the output.
1:23:59So, the first step is data splicing.
1:24:01I've already mentioned data splicing is
1:24:03splitting your data set into your
1:24:05testing data set and into your training
1:24:07data set. That's exactly what we're
1:24:09doing over here. So, 25% of your data is
1:24:12assigned for the testing data and the
1:24:14remaining 75% is your training data.
1:24:17Here, you're creating the instance of
1:24:19the logistic regression algorithm. This
1:24:21is an instance that you created. Then
1:24:23you'll fit the model by using your
1:24:25training data set. So basically, to
1:24:27build your machine learning algorithm,
1:24:29you'll be fitting your training data
1:24:30set. So X_train and Y_train variables
1:24:33have your training data set.
1:24:35After that, you will be evaluating the
1:24:37model by using your testing data set.
1:24:40Then you'll calculate the accuracy
1:24:42score. Right? I'll also be printing the
1:24:44accuracy using logistic regression and
1:24:47the time taken using logistic
1:24:48regression. Let's look at the accuracy.
1:24:51Don't worry about these warnings. They
1:24:53are not important. So accuracy using
1:24:55logistic regression is around 0.83%,
1:24:59which is 83% accuracy, approximately
1:25:0184%. And this is the time taken. So the
1:25:05accuracy is actually pretty good, right?
1:25:0684% is a good number. Then we have
1:25:09random forest classifier. Here again,
1:25:11we'll import the libraries that are
1:25:13needed to run random forest classifier.
1:25:16Then we're again calculating the
1:25:17accuracy and the time taken by the
1:25:19classifier.
1:25:20Data splicing, like I mentioned,
1:25:22splitting the data into testing and
1:25:24training data set. Then you're just
1:25:26building the model by using the training
1:25:28data set. After that, you'll evaluate
1:25:31the model by using the testing data set
1:25:33and you'll finally calculate the
1:25:35accuracy. The accuracy using random
1:25:37forest is again approximately 84%, which
1:25:40is a really good number. Then we have
1:25:42decision tree classifier. Here again,
1:25:45we'll be importing the libraries needed
1:25:47for this classifier. We'll be
1:25:49calculating the accuracy and the time
1:25:51taken by this classifier. Data splicing
1:25:54followed by building the model by using
1:25:56the training data set, evaluating the
1:25:58model by using the testing data set, and
1:26:00finally calculating the accuracy and
1:26:02printing the accuracy.
1:26:04So let's see the accuracy using decision
1:26:06tree classifier. Again, we have an
1:26:09accuracy of around 83 to 84%.
1:26:12This is a pretty good number. And last,
1:26:14we're going to do this by using another
1:26:16classification algorithm known as
1:26:18support vector machine.
1:26:20Here again, we're importing the needed
1:26:22libraries. Then we're calculating the
1:26:24accuracy and the time, performing data
1:26:26splicing.
1:26:27Then we're building the model by using
1:26:29the training data set, testing the model
1:26:31using the testing data set, and finally
1:26:33printing the accuracy.
1:26:35So guys, all the classification models
1:26:38gave us an accuracy score of
1:26:39approximately 84% to 83%.
1:26:43So this is exactly how a machine
1:26:44learning process works. Right? You begin
1:26:47by importing all your data, then you
1:26:49perform data pre-processing or data
1:26:51cleaning. After that, you perform
1:26:53exploratory data analysis, where you
1:26:55understand the important patterns or the
1:26:58important variables in your data set.
1:27:00After that, you build a model, then you
1:27:03will evaluate the model by using the
1:27:05testing data set, and finally calculate
1:27:07the accuracy. I showed you all the steps
1:27:09in the machine learning process by using
1:27:11a practical demonstration in Python.
1:27:14So guys, give yourself a pat on the back
1:27:16because we just understood the whole
1:27:17machine learning process with a small
1:27:19implementation in Python.
1:27:22Now let's move on to our next topic,
1:27:24which is limitations of machine
1:27:26learning.
1:27:27Before we understand what deep learning
1:27:29is, it's important to know the
1:27:30limitations of machine learning, and why
1:27:33these limitations gave rise to the
1:27:35concept of deep learning. One major
1:27:37problem in machine learning is machine
1:27:40learning algorithms and models are not
1:27:42capable of handling high-dimensional
1:27:45data.
1:27:46Right? We can take in data with 20 to 30
1:27:49feature variables, but when it comes to
1:27:51data sets which have thousands of
1:27:53variables, machine learning does not
1:27:55work. Machine learning is not capable
1:27:58enough to process that much data.
1:28:01So high-dimensional data cannot be
1:28:03analyzed, processed, and modeled by
1:28:05using machine learning.
1:28:07Another limitation is that it cannot be
1:28:09used in image recognition and object
1:28:12detection because these applications
1:28:14require the implementation of
1:28:16high-dimensional data. Another major
1:28:19challenge in machine learning is to tell
1:28:21the machine what are the important
1:28:23features it should look for in order to
1:28:25precisely predict the outcome. So,
1:28:28basically you're selecting the important
1:28:29features for the machine learning model
1:28:31and you're telling them like these are
1:28:32the important features and this is what
1:28:35you should use in order to build the
1:28:36model.
1:28:37This process is known as feature
1:28:39extraction.
1:28:40Now, in machine learning this is a
1:28:41manual process. You're going to manually
1:28:43input as a programmer, you're going to
1:28:45tell that these are the important
1:28:47predictor variables. But, what happens
1:28:49when your data set has hundreds of
1:28:51variables?
1:28:52How are you going to sit and choose
1:28:54every variable and perform analysis on
1:28:56each variable to understand which is a
1:28:58really significant variable? That's
1:29:01going to become a very tedious task,
1:29:03right? It's not possible for you to
1:29:04manually sit down with 100 variables,
1:29:06check the correlation with each variable
1:29:08and understand which variable is
1:29:09significant in predicting the output.
1:29:12So, performing feature extraction
1:29:14manually is very tedious and that is one
1:29:16of the major limitations of machine
1:29:17learning. Now, deep learning comes to
1:29:20the rescue to all of these problems.
1:29:22So, let's understand what deep learning
1:29:25is and why we have deep learning in the
1:29:27first place.
1:29:28So, deep learning is actually one of the
1:29:30only methods by which we can overcome
1:29:33the challenge of feature extraction.
1:29:35This is because deep learning models are
1:29:37capable of learning to focus on the
1:29:39right features by themselves requiring
1:29:42minimal human intervention. Meaning that
1:29:45feature extraction will be performed by
1:29:47the deep learning model itself. You
1:29:49don't have to manually tell that this
1:29:50feature is important, that feature is
1:29:52important, choose this feature for
1:29:54predicting the output. All of this is
1:29:56not needed in deep learning. The model
1:29:58itself will learn which features are
1:30:00most significant in predicting the
1:30:02output.
1:30:03Also, deep learning is mainly used to
1:30:05deal with high-dimensional data, right?
1:30:08It is based on the concept of neural
1:30:10networks and is often used in object
1:30:13detection and image processing. This is
1:30:15exactly why we need deep learning. It
1:30:17solves the problem of processing
1:30:19high-dimensional data and manual feature
1:30:22extraction.
1:30:23Now, how exactly does deep learning
1:30:25work? Now, deep learning mimics the
1:30:27basic component of the human brain
1:30:29called the brain cell. The brain cell is
1:30:32also known as a neuron.
1:30:34So, inspired from a neuron, an
1:30:37artificial neuron was developed. Deep
1:30:39learning is based on the functionality
1:30:41of a biological neuron. So, let's
1:30:44understand how we mimic this
1:30:46functionality in an artificial neuron.
1:30:49Now guys, an artificial neuron is also
1:30:51known as a perceptron.
1:30:53Let's understand what this biological
1:30:55neuron does and how deep learning is
1:30:57based on this concept.
1:30:59In a biological neuron, you can see
1:31:01these dendrites, right? In this image,
1:31:03you see something known as dendrites.
1:31:06These dendrites are used to receive any
1:31:08input. These inputs are summed in the
1:31:11cell body and through the axon, it is
1:31:13passed on to the next neuron.
1:31:16So, similar to the biological neuron, a
1:31:18perceptron or a artificial neuron
1:31:20receives multiple inputs, applies
1:31:23various transformations and functions,
1:31:25and provides an output.
1:31:27Right? So, that's how artificial neural
1:31:29networks or that's how deep learning
1:31:31works.
1:31:32Now guys, the human brain consists of
1:31:34multiple connected neurons called a
1:31:36neural network. Similarly, by combining
1:31:39multiple perceptrons, we've developed
1:31:41what is known as deep neural networks.
1:31:44The main idea behind deep learning is
1:31:46neural networks and that's what we're
1:31:48going to learn about. So now, let's
1:31:50understand what exactly deep learning
1:31:52is.
1:31:53Deep learning is a collection of
1:31:55statistical machine learning techniques
1:31:57used to learn feature hierarchies based
1:32:00on the concept of artificial neural
1:32:02networks. So, the main idea behind deep
1:32:05learning is to use the concept of neural
1:32:08networks.
1:32:09A deep neural network will have three
1:32:11layers. Okay, there's something known as
1:32:13the input layer followed by the hidden
1:32:15layers and then we have the output
1:32:17layer. The input layer is basically the
1:32:19first layer and it receives all the
1:32:21inputs. So, all the inputs are fed into
1:32:24this input layer.
1:32:25The last layer is obviously the output
1:32:27layer. This layer will provide your
1:32:30desired output. Now, all the layers
1:32:32between the input and your output layer
1:32:34are known as the hidden layers.
1:32:36Now, the number of hidden layers in a
1:32:39deep learning network will depend on the
1:32:41type of problem you're trying to solve
1:32:42and the data that you have.
1:32:44We'll get into depth of what exactly a
1:32:46hidden layer does, but for now this is
1:32:48how a neural network is structured in
1:32:51deep learning. So, guys uh deep learning
1:32:53is used in highly computational use
1:32:56cases such as face verification,
1:32:58self-driving cars, and so on. Right? So,
1:33:00let's understand the importance of deep
1:33:02learning by looking at a real-world use
1:33:05case.
1:33:06So, I'm sure all of you have heard of
1:33:07the company PayPal. Now, PayPal makes
1:33:10use of deep learning to identify any
1:33:13possible fraudulent activities.
1:33:16So, the company makes use of deep
1:33:18learning for fraud detection. Now,
1:33:20PayPal recently processed over 235
1:33:24billion dollars in payments from 4
1:33:27billion transactions by its more than
1:33:30170 million customers. So, basically it
1:33:33processed this much data by using deep
1:33:35learning. PayPal uses machine learning
1:33:38and deep learning algorithms to mine
1:33:40data from the customers purchasing
1:33:42history in addition to reviewing
1:33:44patterns of any sort of fraud stored in
1:33:47the database and it will do this to
1:33:49predict whether a particular transaction
1:33:51is fraudulent or not. Now, the company
1:33:53has been relying on deep learning and
1:33:55machine learning technology for around
1:33:5710 years.
1:33:59Initially, the fraud monitoring team
1:34:01used simple linear models, right? They
1:34:03used machine learning, but over the
1:34:05years the company switched to more
1:34:07advanced machine learning technology
1:34:09called deep learning. This shows how
1:34:11deep learning is used in more advanced
1:34:14and more complicated use cases.
1:34:17The fraud risk manager and the data
1:34:19scientist at PayPal, he quoted that what
1:34:23we enjoy from more modern advanced
1:34:25machine learning is its ability to
1:34:27consume a lot more data, handle layers
1:34:30and layers of abstraction, and be able
1:34:32to see things that a simpler technology
1:34:35would not be able to see. Even human
1:34:37beings might not able to see. This is
1:34:40exactly what he quoted. He said that a
1:34:42simple linear model is capable of
1:34:44consuming around 20 variables, but with
1:34:48deep learning technology, you can run
1:34:50thousands of data points.
1:34:52He also quoted that there is a magnitude
1:34:55of difference. You'll be able to analyze
1:34:57a lot more information and identify
1:35:00patterns that are a lot more
1:35:02sophisticated.
1:35:03So, by implementing deep learning
1:35:05technology, PayPal can finally analyze
1:35:07millions of transactions to identify any
1:35:10fraudulent activity.
1:35:12This is how PayPal makes use of deep
1:35:14learning.
1:35:15Not only PayPal, we also have Facebook,
1:35:17right? Facebook makes use of deep
1:35:19learning technology for face
1:35:21verification.
1:35:22You've all seen the tagging feature at
1:35:24Facebook where we tag our friends in
1:35:26photos. All of that is based on deep
1:35:29learning and machine learning.
1:35:31So guys, that was a real-world use case
1:35:33to make you understand how important
1:35:35deep learning is.
1:35:36Now, let's move on and look at what
1:35:39exactly a perceptron is, right? We'll be
1:35:41going in depth about deep learning.
1:35:44A perceptron is basically a single-layer
1:35:47neural network that is used to classify
1:35:49linear data.
1:35:51It is the most basic component of a
1:35:53neural network. Now, a perceptron has
1:35:55four important components. It has
1:35:58something known as inputs, weights, and
1:36:00bias, summation functions, activation
1:36:03and transformation functions.
1:36:05These are four important parts of a
1:36:07perceptron.
1:36:09Now, before I discuss this diagram with
1:36:11you, let me tell you the basic logic
1:36:13behind a perceptron.
1:36:15There is something known as inputs,
1:36:17right? The input X here, you can see X1,
1:36:19X2 till Xn.
1:36:21So, let me explain the structure of a
1:36:22perceptron. What you're going to do is
1:36:24you're going to input variables into the
1:36:26perceptron, right? This X1, X2 till Xn
1:36:29basically stands for input. W1, W2 till
1:36:33Wn stands for the weight assigned to
1:36:36each of these inputs.
1:36:38Right? There is a specific weight
1:36:40that'll be randomly initialized in the
1:36:42beginning for each of your input.
1:36:45Next, you have something known as the
1:36:46summation element. Here, what you do is
1:36:48you multiply the respective input with
1:36:51the respective weight, and you add all
1:36:54these products. Right? That is basically
1:36:57your summation function. After this is
1:36:59what is your transfer function, also
1:37:01known as activation function.
1:37:03Right? The activation function map your
1:37:05input to your desired output. So, your
1:37:08input will go through these processes.
1:37:11It'll go through summation and
1:37:12activation function in order to get to
1:37:14the output. So, guys, remember that the
1:37:17neural networks work the same way as a
1:37:19perceptron. So, if you want to
1:37:21understand how deep neural networks
1:37:23work, you need to understand what a
1:37:24perceptron does.
1:37:26A deep neural network is nothing but
1:37:27multiple perceptrons.
1:37:29So, let me tell you how the entire works
1:37:31once again.
1:37:32So, basically, all your inputs are
1:37:34multiplied with their respective
1:37:36weights. Now, you add all the multiplied
1:37:39values, and you call them as a weight
1:37:41sum. You use the summation function to
1:37:43add all of this. After that, you apply
1:37:46the weighted sum to the correct
1:37:48activation transfer function. Activation
1:37:50function is very similar to a function
1:37:53in our brain. The neurons become active
1:37:56in our brain after a certain potential
1:37:59is reached. That threshold is known as
1:38:01the activation potential.
1:38:03So, mathematically, there are a few
1:38:05functions which represent the activation
1:38:07function. Basically, the signum, the
1:38:09sigmoid, the tan h, all of these are
1:38:11activation functions. You can think of
1:38:13activation function as a function that
1:38:15maps the input to the respective output.
1:38:19Then, I spoke about something known as
1:38:20weights and biases. All right, now you
1:38:23must be wondering, why do we have to
1:38:25assign weights to each of our input?
1:38:28Weights basically show the strength of a
1:38:30particular input or how important a
1:38:33particular input is for predicting the
1:38:35output. In simple words, the weightage
1:38:38denotes the importance of an input. Bias
1:38:41is basically a value which allows you to
1:38:44shift the activation function curve in
1:38:47order to get a precise output.
1:38:50All right, so that's exactly what
1:38:51weights are.
1:38:52I hope all of you are clear with inputs,
1:38:54weights, summation, and activation
1:38:57function. Also, one important thing I
1:38:59forgot to mention in a perceptron is a
1:39:01single layer perceptron will have no
1:39:04hidden layers. All right, there'll only
1:39:05be an input layer, an output layer, and
1:39:07a couple of transformation function in
1:39:09between. That's all will be there in a
1:39:11perceptron. Now, perceptron, like I
1:39:13mentioned, is used to solve only linear
1:39:15problems. If you look at this data
1:39:17distribution, how do you think we can
1:39:19solve this? This data is not linearly
1:39:22separable. So, you cannot use a single
1:39:24layer perceptron to separate this data.
1:39:27All right, that's why we need something
1:39:28known as a multi-layer perceptron with
1:39:31backpropagation.
1:39:32I'll be explaining this in the next
1:39:34slide. So, complex problems that involve
1:39:38a lot of parameters and high-dimensional
1:39:41data can be solved by using multiple
1:39:43layer perceptron. Now, a multi-layer
1:39:46perceptron is the same as a single layer
1:39:48perceptron. The only difference is that
1:39:50a multi-layer perceptron will have
1:39:52hidden layers.
1:39:53So, the number of hidden layers in a
1:39:56model depends upon various factors. I
1:39:58told you it depends on the complexity of
1:40:00the problem you're trying to solve. It
1:40:01depends on the number of inputs in your
1:40:03data and so on. So, it works in the same
1:40:05way. All your inputs are multiplied with
1:40:08your weights and then you do the
1:40:09summation and then there is a
1:40:11transformation function or a activation
1:40:14function.
1:40:15While designing a neural network, in the
1:40:17beginning itself I told you we
1:40:19initialize weights with some random
1:40:21values. We do not have some specific
1:40:23[clears throat] value for each
1:40:24weightage. Initially, we've selected
1:40:27random values.
1:40:28It is always important that whatever
1:40:31weight values we have selected will be
1:40:33correct.
1:40:34Now, whatever weight values we've
1:40:36assigned to each input, it denotes the
1:40:38importance of that input variable. So,
1:40:41we need to assign the weights in such a
1:40:43way or we need to update the weights in
1:40:45such a way that it denotes the
1:40:47significance of that particular input.
1:40:50So, initially we're selecting some
1:40:52random value for weight. And let's say
1:40:54that we use this weight value to get our
1:40:57output.
1:40:58Now, what happens is the output is
1:41:00actually very different or it is not
1:41:02precise when compared to our actual
1:41:04output. Basically, the error value is
1:41:07very huge.
1:41:08So, how will you reduce the error?
1:41:10The main thing in a neural network is
1:41:13the weightage that you give to a input
1:41:15variable, right? Depending on the
1:41:16weightage that you give to a input
1:41:18variable, you're telling the neural
1:41:19network how important that variable is.
1:41:22Now, what if you randomly give some
1:41:24weightage and your output is wrong? The
1:41:28first thing that comes into your mind is
1:41:29that you need to change the weight
1:41:31because the weight signifies the
1:41:33importance of a variable.
1:41:35So, basically what we need to do is we
1:41:37need to somehow explain to the model to
1:41:39change the weight in such a way that the
1:41:41error becomes minimum. Let's put it in
1:41:44another way. So, basically, we need to
1:41:46train a model. One way to train a model
1:41:48is called as backpropagation.
1:41:51So, in backpropagation, what happens is
1:41:54once you've initialized a weight to each
1:41:56of the input, you calculate the output.
1:41:59Right? You get an output, and let's say
1:42:01you have a very high error value in that
1:42:03output. What you do is you'll
1:42:05backpropagate as in you'll go back to
1:42:08the weight, and you'll keep updating the
1:42:10weight in such a way that your error
1:42:12becomes minimum.
1:42:14This is exactly what backpropagation is.
1:42:16You'll be going back to the first layer.
1:42:18You'll be updating each of the weights
1:42:20in such a way that your output is more
1:42:23precise. So, guys, basically, the weight
1:42:26and the error in a neural network is
1:42:29highly related. By updating the weight
1:42:32in a particular way, your error will
1:42:33decrease. So, you need to figure out how
1:42:36you need to update the weight. Do you
1:42:38have to increase the weight or decrease
1:42:40the weight? Once you figure out whether
1:42:41you have to increase or decrease the
1:42:43weight, you have to just follow that
1:42:45direction in such a way that your error
1:42:47is minimized. And that's exactly what
1:42:50backpropagation is. So, the final output
1:42:53of backpropagation is you're going to
1:42:55select the weight that minimizes the
1:42:57error function. And then you're going to
1:42:59use that weight to solve the whole
1:43:01problem.
1:43:02Right? This is what backpropagation is
1:43:04about.
1:43:05Now, in order to make you understand
1:43:07deep neural networks, let's look at a
1:43:09practical implementation.
1:43:12So, again, guys, I'll be using Python to
1:43:14run the demo. If you don't have a good
1:43:16idea about Python, check the
1:43:17description. I'll leave a couple of
1:43:19links about Python programming. Now, in
1:43:21this demo, I'll be walking you through
1:43:23one of the most important applications
1:43:25of deep learning. I will demonstrate how
1:43:28you can construct a high-performance
1:43:30model to detect credit card fraud.
1:43:32Right? We'll be using deep learning
1:43:34models to do this. Now, before that, let
1:43:36me just tell you something about our
1:43:37data set. Right? The data set contains
1:43:40transactions made by credit cards in the
1:43:43year September 2013 by European card
1:43:46holders. This data set presents
1:43:48transactions that occurred in 2 days,
1:43:51where we have 492 frauds out of 285,000
1:43:56transactions. Approximately 285,000
1:44:00transactions. Out of these transactions,
1:44:02492 were frauds, and the data set is
1:44:05quite unbalanced. Right? The positive
1:44:07class accounts for 0.172%.
1:44:11So, the positive class, basically the
1:44:13fraudulent class. So, again, we're going
1:44:15to start by importing the required
1:44:17packages. We're going to import Keras,
1:44:19Matplotlib library, Seaborn library, and
1:44:22scikit-learn for preprocessing. Right?
1:44:24Again, min-max scalar, which is for
1:44:26normalization. We're going to import our
1:44:29data set and store it in this variable.
1:44:31Right? This is the path to my data set.
1:44:34My data set is in the CSV format, or
1:44:36also known as comma-separated version.
1:44:39Now, we're going to print out the first
1:44:41five rows of our data set. Right? Let's
1:44:44take a look at the output.
1:44:47So, here is the time of the transaction.
1:44:49V1, V2, V3, etc. These are all the
1:44:52features of our data set. I'm not going
1:44:55to go into depth of what these features
1:44:57stand for, because this demo is all
1:44:59about understanding deep learning. Now,
1:45:01these V1, V2, V3, these are all
1:45:03predictor variables, which will help us
1:45:05predict our class.
1:45:07So, guys, don't worry about what these
1:45:09features are. These features are just
1:45:11information and details about your
1:45:13transaction, such as the amount you
1:45:15spend, or the time of transaction, and
1:45:18so on.
1:45:19So, here we have the amount variable,
1:45:20which denotes the amount spent. After
1:45:23that, we have the class variable. Now,
1:45:25this class variable is your output
1:45:27variable or your target variable. So,
1:45:30your class is basically your output
1:45:32variable. Value zero denotes that there
1:45:34has been no fraudulent activity, but if
1:45:36you get a class of one, it means that
1:45:39this transaction is a fraudulent
1:45:41transaction. For example, this
1:45:43transaction is not fraudulent, and
1:45:45that's why we have a value of zero over
1:45:47here. All right, so this is our data
1:45:49set.
1:45:51Next what we're doing is we're counting
1:45:53the number of samples for each class.
1:45:55Right, we have class zero and class one,
1:45:57where in class zero denotes the normal
1:45:59transaction, which is non-fraudulent
1:46:02transaction, and class one will denote
1:46:04the fraudulent transactions. Right, so
1:46:06we have around 492 fraudulent
1:46:08transactions and around 284,315
1:46:13non-fraudulent transactions.
1:46:15So, when you see this, you know that our
1:46:17data set is highly unbalanced. Highly
1:46:19unbalanced means that one class has a
1:46:22really small number when compared to the
1:46:24other class. Right, there's no balance
1:46:25between the two classes.
1:46:27So, here what we're doing is we are
1:46:29starting the data set by class for
1:46:31stratified sampling. Stratified sampling
1:46:34is a statistical technique for sampling
1:46:37your data set. Now, this type of
1:46:39sampling is always good if you have an
1:46:41unbalanced data set.
1:46:43Next what we're going to do is we're
1:46:44going to perform data pre-processing.
1:46:47Data pre-processing in deep learning
1:46:49mainly has a method known as dropout
1:46:52method.
1:46:53Next what we're going to do is we're
1:46:54going to uh drop out the entire time
1:46:57column. We do not need the time of the
1:47:00transaction in order to understand if
1:47:02the transaction was fraudulent or not.
1:47:05Right, so that's why we're getting rid
1:47:06of unnecessary variables. Right, so
1:47:08we're dropping out that variable.
1:47:12So, after dropping out the time
1:47:14variable, we're going to assign the
1:47:16first 3,000 samples to our new data
1:47:18frame. Right, this DF sample will have
1:47:20our first 3,000 samples, and we're going
1:47:24to use those 3,000 samples.
1:47:26So, here we're just counting the number
1:47:27of class for each of these samples.
1:47:30After that, we're just counting the
1:47:31number of samples for each of the class.
1:47:33Like we're doing the same thing again
1:47:34and here we get class zero has 2,508
1:47:38samples and class one has 492 samples.
1:47:41Now, this makes the data set quite
1:47:43balanced, right? It's very balanced when
1:47:45compared to our old data set.
1:47:48Next, we'll just randomly shuffle our
1:47:49data set, right? In order to remove any
1:47:52sort of biasness in the data.
1:47:54After that, we'll split our data set
1:47:56into two parts. One is for training and
1:47:59your other data set is for testing,
1:48:01right? This is also known as data
1:48:03splicing.
1:48:05Then, we'll be splitting each data frame
1:48:06into feature and label, meaning that
1:48:09your input and your output.
1:48:12We'll be doing this for your training
1:48:13data and for your testing data, right?
1:48:15All you're doing is you're separating
1:48:17your input from your output.
1:48:19Next, we're looking at our training data
1:48:21set, right? We're printing the shape of
1:48:23our training data set. The training data
1:48:25set has around 2,400
1:48:27observations and 29 variables or 29
1:48:31features.
1:48:32Similarly, we'll be printing out the
1:48:34size of our test data frame, right?
1:48:36That's exactly what we're doing over
1:48:37here.
1:48:38After that, we'll perform normalization,
1:48:41right? For this, we'll be using the
1:48:42min-max scaler.
1:48:44So, in normalization, we'll basically be
1:48:46scaling all our predictor variables
1:48:48around the same range so that there is
1:48:50no biasness in our prediction.
1:48:53After this, we'll be plotting a function
1:48:55for each of the learning curves. For
1:48:57your training phase and for your testing
1:48:59phase, you'll be plotting a learning
1:49:01curve. Now, I'll show you the output of
1:49:03this in a couple of minutes. For now,
1:49:05let's move on to the main part, which is
1:49:07model creation, right? In this demo,
1:49:10we'll use three fully connected layers.
1:49:13We'll also use dropout technique. Now,
1:49:15dropout is a type of regularization
1:49:18technique that is used to avoid any sort
1:49:20of overfitting in a neural network. It
1:49:23is a technique where you select neurons
1:49:25and you drop them during the training
1:49:26phase.
1:49:27We'll be using the ReLU as the
1:49:29activation function, which is a type of
1:49:31activation function just like sigmoid
1:49:33and tanh.
1:49:35So, the type of model that we'll be
1:49:36using is the sequential model. Right,
1:49:39sequential is the easiest way to build a
1:49:41model in Keras. Right, we're using the
1:49:43Keras library over here. If you
1:49:45remember, I imported that in the
1:49:47beginning. Right, it allows you to build
1:49:49a model layer by layer. So, each layer
1:49:52has weights that correspond to the layer
1:49:54that follows it. After this, you'll use
1:49:57the add function to add the dense
1:50:00layers. Basically, your hidden layers
1:50:01you're going to add over here. So, in
1:50:03our model, we'll be adding two dense
1:50:05layers or hidden layers, you can say.
1:50:08So, here what we're doing is we're
1:50:10adding the first dense layer. Now guys,
1:50:12a dense layer is standard layer type
1:50:15that works for most cases. Right, in a
1:50:18dense layer, all the nodes in the
1:50:19previous layer connect to the nodes in
1:50:21the current layer.
1:50:23So guys, don't get too involved into
1:50:25what exactly is happening here. All I'm
1:50:27doing is I'm creating a sequential
1:50:29model, and what is happening is I'm just
1:50:32assigning the number of inputs for each
1:50:34of the dense layer or for each of the
1:50:36hidden layer. I'm also assigning dropout
1:50:39value. Dropout is basically to prevent
1:50:41overfitting. Overfitting might occur
1:50:43when your model memorizes the training
1:50:45data set. Overfitting basically reduces
1:50:48the accuracy of a model. That's why
1:50:50we're using the dropout method to
1:50:52prevent overfitting. So, in the first
1:50:54hidden layer, we have around 200 units.
1:50:56Right, we have the activation function
1:50:58ReLU. Then, we're adding the second
1:51:01dense layer with again 200 neurons and
1:51:03the ReLU activation function. Kernel
1:51:06initializer is uniform, meaning that
1:51:08it's just sequential and normal. Then,
1:51:10we're again adding a dropout layer of
1:51:120.5. The dropout value of a network has
1:51:16to be chosen very wisely. Okay, a value
1:51:19that is too low will result in a minimal
1:51:21effect and a value that is too high will
1:51:23result in under learning by the network.
1:51:26So, 0.5 is a standard dropout value.
1:51:30Now, this last layer is our output
1:51:32layer. In the output layer, we'll
1:51:33obviously have only one neuron. We'll
1:51:35have one neuron that will show us the
1:51:37output class, either zero or one. Zero
1:51:41will show us non-fraudulent transactions
1:51:43and one will denote fraudulent
1:51:45transaction. Right, that's why we have
1:51:46only one neuron over here. And the
1:51:49activation function here is sigmoid.
1:51:51Right, since the number of neurons is
1:51:52only one.
1:51:54After that, we're printing the model
1:51:55summary. Now, I'll show you the summary
1:51:58and everything. Before that, let us
1:51:59understand what exactly optimization
1:52:02functions are. We'll understand what
1:52:04this optimization function does. Now, an
1:52:06optimizer takes care of the necessary
1:52:09computations that are used to change the
1:52:12network's weights and bias. So,
1:52:14basically, your optimizers will take
1:52:16care of all your computations such as
1:52:19changing the weight or updating the
1:52:21weight. If you all remember, I spoke
1:52:22about backpropagation, right? Where
1:52:24you'll update the weight and all of
1:52:25that. That is done by using optimizers.
1:52:28Here, we're selecting an optimizer known
1:52:30as the Adam optimizer.
1:52:33So, Adam optimizer is one of the current
1:52:35default optimizers in deep learning.
1:52:38Right, it stands for adaptive moment
1:52:40estimation. We don't have to get into
1:52:42the depth of all of this. Right, all of
1:52:44these are predefined optimizers in our
1:52:46Keras package itself. After this, we're
1:52:48going to fit our model by using the
1:52:50training features. We're also setting
1:52:53200 epochs and also there's something
1:52:56known as epochs and batch size. Right,
1:52:58we're setting epochs as 200 and batch
1:53:00size as 500. I'll tell you what exactly
1:53:03this means. Now, batch sizes are
1:53:05basically used so that we don't overfit
1:53:08our model, right? We're going to
1:53:09basically split our data set into 500
1:53:12batches. So, our input will be going in
1:53:14the form of batches, right? And our
1:53:17batch size is 500 inputs per batch, and
1:53:20we'll be going through 200 epochs.
1:53:23Meaning that our training will iterate
1:53:25200 times. This is basically the number
1:53:27of times that training our model. All
1:53:29right, that's what epoch and batch size
1:53:31is. After that, we're just showing our
1:53:34training history. I mean, just printing
1:53:35the accuracy curve for our training
1:53:37phase. We're also going to print our
1:53:39loss curves for our training phase,
1:53:41basically the error curves. And then
1:53:43finally, we have the evaluation. Here,
1:53:45we'll be testing our model by using our
1:53:47testing data set.
1:53:49Then we're finally printing the accuracy
1:53:51on our testing data set. After that,
1:53:54we're just going to plot a heat map,
1:53:56which I'll be showing y'all. Let me just
1:53:58show you the output.
1:54:00So guys, in this entire line of code,
1:54:01all we're doing is we're printing an
1:54:03accuracy plot. All right, basically
1:54:05we're printing a heat map. I'll show you
1:54:07what the heat map looks like.
1:54:10This is just to check the accuracy.
1:54:12We're comparing all the correctly
1:54:14predicted values to our incorrectly
1:54:16predicted values. So, this is our
1:54:19training history. Here, blue stands for
1:54:21our training phase, and this is our
1:54:22validation or our prediction stage.
1:54:26That was our training curve, and this is
1:54:28our loss curve. Now, when you compare it
1:54:30to the actual validation stage, it's
1:54:32quite similar, right? Meaning that our
1:54:34model is doing pretty well. So guys,
1:54:37this is the heat map that I was talking
1:54:38about. This is basically going to give
1:54:40us the class for each of our
1:54:41predictions. All right, it basically
1:54:43plots the classes that we correctly
1:54:46predicted, right? Basically, for each
1:54:48data point, it's just going to tell us
1:54:49whether we predicted it correctly or
1:54:51not. It's sort of a confusion matrix in
1:54:53the form of a heat map. So guys, these
1:54:56are all our epochs, basically the 200
1:54:59iterations that we went through, right?
1:55:01This is the 50th iteration is showing us
1:55:03our loss, it's showing us our accuracy
1:55:05as well. And here we have 88%, 90%, 92%.
1:55:10Now, if you carefully look at the epoch
1:55:12accuracy values, you see that as we
1:55:15train our model even more, our accuracy
1:55:17keeps increasing. Initially, our
1:55:19accuracy was around 83, right? At epoch
1:55:22number 15, our accuracy was around 83%.
1:55:26But as we kept training our model a
1:55:29little bit more, our accuracy kept
1:55:31increasing. We have 90, we have 91, 94,
1:55:3595, 96, and so on. All right, so
1:55:38basically, the more you train your
1:55:39model, the better it's going to be. So
1:55:42guys, this was our entire demo. Now, in
1:55:44the end, I'm printing out the false
1:55:46positive rate and the false negative
1:55:48rate. All of this basically denotes how
1:55:50many of the data points was I correctly
1:55:53able to predict as fraudulent and how
1:55:55many did I predict wrongly. That's all
1:55:58the false negative and the false
1:56:00positive rate denotes. So guys, this was
1:56:02the entire demo on deep learning. Now,
1:56:06if you have any doubts regarding the
1:56:07deep learning demo, please mention them
1:56:09in the comment section and I will solve
1:56:11your queries. All right, now let's look
1:56:13at our last topic for the day, which is
1:56:15natural language processing. Now, before
1:56:18we understand what is natural language
1:56:20processing, let's understand the need
1:56:21for natural language processing and a
1:56:24process known as text mining. Text
1:56:26mining and natural language processing
1:56:28are heavily correlated. All right, I'll
1:56:29talk about both of these in the upcoming
1:56:32slides. For now, let me tell you why we
1:56:34need natural language processing or text
1:56:36mining. So guys, the amount of data that
1:56:38we're generating these days is
1:56:40unbelievable. It is a known fact that
1:56:42we're creating 2.5 quintillion bytes of
1:56:46data every day, and this number is only
1:56:48going to grow. With the evolution of
1:56:50communication through social media, we
1:56:52generate tons and tons of data. All
1:56:55right, the numbers are on your screen.
1:56:57So, basically, we post around 1.7
1:56:59million pictures on Instagram per
1:57:01minute. Right, I'm talking about post
1:57:04per minute. All of these numbers are per
1:57:06minute values. These are the amount of
1:57:08tweets, 347,000
1:57:10tweets per minute. Right, this is a lot
1:57:13of data. We're generating data while
1:57:15we're watching YouTube videos, when
1:57:17we're sending emails, when we are
1:57:19chatting, and all of that. Right, even
1:57:21the IoT devices at our house, right, we
1:57:24have Alexa all of This is generating a
1:57:25lot of data. A single click on your
1:57:27phone is generating a lot of data. Now,
1:57:30not only that, out of all the data that
1:57:32we generate, only 21% of the data is
1:57:35structured and well formatted. Right,
1:57:38the remaining of the data is
1:57:39unstructured. And the major sources of
1:57:42unstructured data include text messages
1:57:44from WhatsApp, Facebook likes, comments
1:57:46on Instagram, the bulk emails, and all
1:57:49of this. Right, all of this accounts for
1:57:51the unstructured data that we have
1:57:53today. Now, the data we generate is used
1:57:56to grow a business. So, by analyzing and
1:57:58mining the data, we can add more value
1:58:01to a business. This is exactly what
1:58:04natural language processing and text
1:58:06mining is all about. Text mining and NLP
1:58:09is a subset of artificial intelligence,
1:58:11wherein we try and understand the
1:58:14natural language text that we get from
1:58:16text messages and so on, in order to
1:58:19derive useful insights and grow
1:58:21businesses by using these insights. So,
1:58:23what exactly is text mining? Text mining
1:58:26is a process of deriving meaningful
1:58:28insights or information from natural
1:58:31language text. So, all the data that we
1:58:34generate through text messages, emails,
1:58:36and documents are written in natural
1:58:38language text. Right, and we're going to
1:58:40use text mining and natural language
1:58:42processing to draw useful insights or
1:58:44patterns from such data in order to grow
1:58:47a business.
1:58:48Now, let's understand where exactly do
1:58:50we make use of natural language
1:58:51processing and text mining? Now, have
1:58:53you ever noticed that if you start
1:58:55typing a word on Google, you immediately
1:58:58get suggestions, right? This feature is
1:59:00known as auto complete. It will
1:59:02basically suggest the rest of the word
1:59:04to you. We also have something known as
1:59:05spam detection, right? Here's an example
1:59:08of how Google recognizes this
1:59:10misspelling Netflix and shows results
1:59:13for the keyword that matches your
1:59:14misspelling. Let me show you a couple of
1:59:17more examples. We also have predictive
1:59:20typing and spell checkers and features
1:59:22like auto correct, email classification.
1:59:25So, predictive typing and spell
1:59:27checkers, all of these are applications
1:59:30of natural language processing. All of
1:59:32this basically involves processing the
1:59:34natural language that we use and
1:59:36deriving some useful information from
1:59:38it, right? Or running businesses from
1:59:40it. Netflix uses natural language
1:59:42processing in a really good fashioned
1:59:45way, right? It basically studies the
1:59:47reviews that customer gives for a
1:59:49particular movie and it tries to figure
1:59:51out if that movie is good or bad
1:59:53depending on the review. So, Netflix
1:59:55actually uses NLP in a very interesting
1:59:58manner. It tries to understand the type
2:00:00of movies that a person likes by the way
2:00:03a person has rated the movie or by the
2:00:05way the person has reviewed a movie. So,
2:00:08by understanding what type of review a
2:00:10person is giving to a movie, Netflix
2:00:12will recommend more movies that you
2:00:15like. That's how important NLP has
2:00:17become. Now, let's look at what exactly
2:00:19NLP is. NLP, which also stands for
2:00:22natural language processing, is a part
2:00:24of computer science and artificial
2:00:26intelligence, which deals with human
2:00:28language. Right? It's basically the
2:00:30process of processing natural language
2:00:32in order to derive some useful
2:00:34information from it. For those of you
2:00:36who have studied natural language
2:00:38processing or have heard of natural
2:00:40language processing, there is a huge
2:00:42confusion between text mining and
2:00:44natural language processing. So, text
2:00:46mining is the process of deriving high
2:00:48quality information from text. But, the
2:00:51overall goal is to turn the text into
2:00:53data for analysis by using natural
2:00:56language processing. So, basically text
2:00:58mining is implemented by using natural
2:01:01language processing techniques. Right?
2:01:03There are various techniques in natural
2:01:04language processing that can help us
2:01:06perform text mining. That's how text
2:01:08mining and natural language processing
2:01:10are related. Natural language processing
2:01:12is the techniques that are used to solve
2:01:15the problem of text mining, text
2:01:17analysis, and all of that. Let's look at
2:01:19a couple more applications. Sentimental
2:01:22analysis is one of the major
2:01:23applications of natural language
2:01:25processing. You see Twitter performs
2:01:27sentimental analysis, Facebook, Google,
2:01:30all of these perform sentimental
2:01:31analysis. Sentimental analysis mainly
2:01:34used to analyze social media content
2:01:36that can help us determine the public
2:01:38opinion on a certain topic. Then we have
2:01:41chatbots. Now, chatbots use natural
2:01:43language processing to convert human
2:01:45language into desirable actions. We also
2:01:48have machine translation. NLP is used in
2:01:51machine translation by studying the
2:01:52morphological analysis of each word and
2:01:55translating it to another language.
2:01:58Advertisement matching is also done
2:01:59using NLP in order to recommend ads
2:02:02based on your history. Right? These are
2:02:04few of the applications of NLP. Now, let
2:02:07me tell you the basic terminologies
2:02:08under natural language processing. So,
2:02:10tokenization is the most basic step in
2:02:13natural language processing.
2:02:14Tokenization means breaking down the
2:02:17data into smaller chunks or tokens so
2:02:20that they can be easily analyzed. So,
2:02:22the first step is you'll break a complex
2:02:24sentence into words, then you'll
2:02:26understand the importance of each of the
2:02:28word with respect to that sentence in
2:02:31order to produce a structural
2:02:33description on an input sentence. So,
2:02:35for example, take this sentence. How
2:02:37would I perform tokenizations on the
2:02:39sentence?
2:02:40Let's say that tokens are simple is a
2:02:43sentence and I want to perform
2:02:44tokenization on the sentence. This is
2:02:47what I'm going to do. I'm going to split
2:02:48the sentence into different words. I'm
2:02:50going to understand each word with
2:02:52respect to that sentence. Right? This is
2:02:55done to simplify operations in natural
2:02:57language processing. Right? It's always
2:02:59simpler to analyze a single token
2:03:02instead of analyzing an entire sentence.
2:03:04Then we have something known as
2:03:05stemming. Now look at this example.
2:03:08Right here we have words such as
2:03:10detection, detecting, detected, and
2:03:12detections. We all know that the root
2:03:15word for all of these words is detect.
2:03:17So stemming algorithm basically does
2:03:20that. It works by cutting off the end or
2:03:23the beginning of the word and taking
2:03:25into account a list of common prefixes
2:03:28and suffixes that can be found in an
2:03:30inflicted word. Stemming basically helps
2:03:33us in analyzing a lot of words. We know
2:03:36that detections, detected, and detection
2:03:38basically mean the same thing. So all
2:03:40we're doing is we're going to ease our
2:03:42analysis by removing prefixes and
2:03:44suffixes which not make sense. Right? We
2:03:47just need to understand the
2:03:48morphological analysis of the word.
2:03:50Right? So that's why we're randomly
2:03:51cutting the prefixes and suffixes in
2:03:53such a way that we only get the
2:03:55important part of the word. This is
2:03:57called stemming. Now this cutting of
2:03:59words can be successful in some
2:04:02occasions, but not always. That is why
2:04:04we say that stemming approach has a few
2:04:08limitations. In order to get over these
2:04:11limitations, we have a process known as
2:04:13lemmatization. Right? Lemmatization on
2:04:16the other hand takes into consideration
2:04:18the morphological analysis of the words.
2:04:21It does not randomly cut the word in the
2:04:23beginning and the ending. It understands
2:04:25what the word means and only then it
2:04:27cuts the word. For example, let's
2:04:29consider the word recap. If we perform
2:04:32stemming on the word recap, we'll get
2:04:35cap. Right? The output will be cap. But,
2:04:38cap and recap do not have the same
2:04:40meaning, do they? They have absolutely
2:04:41different meanings. That's why stemming
2:04:43is sometimes not considered to be the
2:04:45right thing to do. But, when it comes to
2:04:47lemmatization, it's going to understand
2:04:49the meaning of recap. Only then will it
2:04:52perform any sort of change in the word,
2:04:54or it'll cut down the word. So,
2:04:56basically, it groups together different
2:04:58inflected forms of a word called lemma.
2:05:01Lemmatization is similar to stemming
2:05:03because it maps several words into one
2:05:06common root. But, the output of a
2:05:09lemmatization process is always a proper
2:05:11word. An example of lemmatization is to
2:05:15map gone, going, and went into go. Gone,
2:05:18going, went, all of them mean go. So,
2:05:21basically, by lemmatization, you can
2:05:22just output the words as go. That is
2:05:25what lemmatization is. Next, we have
2:05:27something known as stop words, right?
2:05:29Stop words are basically a set of
2:05:31commonly used words in any language,
2:05:34right? Not just English, any language.
2:05:36The reason why stop words are critical
2:05:38to many applications is that if we
2:05:41remove the words that are very commonly
2:05:44used in a given language, we can finally
2:05:46focus on the important words. For
2:05:48example, in the context of Let's say you
2:05:51open up Google and you look for
2:05:53strawberry milkshake recipe. Instead of
2:05:55typing strawberry milkshake recipe,
2:05:57let's say you type how to make
2:05:59strawberry milkshake. Now, here, what
2:06:02Google will do is it'll find results for
2:06:04how, to, and make. Instead, if you just
2:06:07type strawberry milkshake recipe, you'll
2:06:10get the most desired output. That's why
2:06:13it's always considered a good practice
2:06:15in natural language processing to get
2:06:17rid of stop words, right? Stop words
2:06:19will just increase our computation, and
2:06:21it'll just add additional work to us.
2:06:23They are not very helpful when we're
2:06:25analyzing important documents, right? We
2:06:27need to focus on the important keywords
2:06:29in the documents instead of all of these
2:06:31commonly used words. Example of stop
2:06:34words include the, how, when, why, not,
2:06:38yes, no. All of these are stop words,
2:06:41right? So, in order to better analyze
2:06:43our data, we need to get rid of stop
2:06:45words. Now, the last terminology I'm
2:06:47going to discuss is document term
2:06:49matrix. It is important to create
2:06:51something known as the document term
2:06:53matrix in natural language processing. A
2:06:56DTM or a document term matrix is
2:06:59basically a matrix that shows the
2:07:01frequency of words in a particular
2:07:03document. Let's say that we're trying to
2:07:05understand if the sentence this is fun
2:07:09is available in one of my documents.
2:07:12So, if it is there in my document one,
2:07:14I'm going to put a one corresponding to
2:07:16each of the words that is available in
2:07:18my document. For example, in document
2:07:20two, I have this is, but I do not have
2:07:23the word fun.
2:07:24Similarly, in document four, I have the
2:07:27word this, but I do not have the word is
2:07:29and fun. So, basically, a document term
2:07:31matrix is like the frequency matrix of a
2:07:34document. So, during text analysis, you
2:07:37always begin by building a document term
2:07:39matrix, right? Here, you try to
2:07:40understand which words frequently occur
2:07:43and which words are important and not
2:07:45important in the document. So, guys,
2:07:47these were a couple of terminologies in
2:07:49natural language processing.
2:07:57Artificial intelligence and machine
2:07:59learning are not just trending
2:08:00technologies anymore.
2:08:02They're becoming the backbone of every
2:08:04industry. And right now, the demand for
2:08:06AI and ML engineers is exploding
2:08:09worldwide. In India, AI engineers earn
2:08:12anywhere from 8 lakhs to 36 lakhs per
2:08:15year, depending on skills and
2:08:17experience. In the United States, the
2:08:19same roles can start from $120,000 and
2:08:23can go all the way up to $250,000 for
2:08:26senior and specialized positions.
2:08:28So, if you have been thinking about
2:08:30getting into AI and ML, switching
2:08:32careers, or upskilling for higher-paying
2:08:35opportunities, there has never been a
2:08:37better time. And in this video, I am
2:08:39giving you a complete,
2:08:41beginner-friendly, and deeply practical
2:08:43AI and ML engineer roadmap that shows
2:08:46you exactly what to learn and how to
2:08:48grow in this booming field.
2:08:50And now, the first step of your AI
2:08:53journey starts with a strong
2:08:55foundations.
2:08:56And no, you don't need to be a
2:08:57mathematician. You simply need the
2:09:00essentials. So, start with Python,
2:09:02because Python is the language that
2:09:04powers almost every modern AI system.
2:09:07So, focus on the basics like variables,
2:09:10loops, functions, lists, dictionaries,
2:09:14file handling, and how to work with
2:09:16APIs.
2:09:17So, these skills are enough to write
2:09:18simple programs and understand AI code.
2:09:21Next, learn data handling, because AI is
2:09:25built on data. Use Pandas to clean and
2:09:27organize data, NumPy to perform
2:09:30calculation, and Matplotlib or Seaborn
2:09:33to visualize patterns. Even simple tasks
2:09:36like removing missing values or
2:09:38analyzing sales trends will prepare you
2:09:40for the real AI projects.
2:09:42Then, learn the essential math behind
2:09:44AI. Not heavy equations, just the basic
2:09:47understanding. Understand mean, median,
2:09:50variance, probability basic,
2:09:52correlations, and what vectors and
2:09:54matrices are. Learn what gradient
2:09:57descent means conceptually, so you
2:09:59understand how models learn without
2:10:01getting buried in complex math. So, once
2:10:03your foundations are ready, move to
2:10:05machine learning. ML is simply teaching
2:10:08computers to learn from examples.
2:10:10Instead of writing instructions, show
2:10:12the model real-world data and let it
2:10:14find patterns.
2:10:16So, learn key ML concepts like training
2:10:18and testing, accuracy and precision,
2:10:21underfitting and overfitting, cross
2:10:23validation, and feature engineering. So,
2:10:26these concepts helps you understand how
2:10:28to build, tune, and improve models.
2:10:31Then, learn the core email algorithms
2:10:33that companies use every single day. So,
2:10:36you need to start with linear
2:10:38regression, logistic regression,
2:10:40decision trees, random forest, SVM,
2:10:44naive Bayes, K-means clustering, and
2:10:46PCA.
2:10:47These algorithms cover most practical
2:10:49business problems like predicting sales,
2:10:52detecting fraud, segmenting customers,
2:10:55and identifying patterns in large data
2:10:57sets.
2:10:58Then, build small email projects such as
2:11:00house price predictor, spam email
2:11:03classifier, credit score predictor, or
2:11:06customer segmentation model. So, these
2:11:08projects give you confidence and make
2:11:10your portfolio job ready. After ML, move
2:11:13into deep learning, the technology
2:11:15behind ChatGPT, self-driving cars, and
2:11:18medical AI.
2:11:19Start by understanding how neural
2:11:21networks work. Learn what neurons,
2:11:24layers, activation functions, and loss
2:11:26functions are.
2:11:28You don't need to memorize formulas,
2:11:30just understand how the network adjusts
2:11:32itself to improve predictions.
2:11:34Pick either TensorFlow or PyTorch as
2:11:37your deep learning framework because
2:11:39both are used by companies in
2:11:40production, so you only need to choose
2:11:42one. And then, build deep learning
2:11:44projects like digit recognition, image
2:11:46classification, sentiment analysis, or
What is Machine Learning?
2:11:49go for fake news classification. So,
2:11:51these projects teach you how to use
2:11:53neural networks in real scenarios.
2:11:56So, AI is no longer about learning
2:11:59everything. It's about choosing your
2:12:01specialization.
2:12:02So, here we have options. So, option A
2:12:05is NLP and LLMs, which has the highest
2:12:08demand. So, if you want to work with
2:12:10chatbots, smart assistants, or language
2:12:13models like ChatGPT, choose NLP and
2:12:16LLMs. And all you need to learn is
2:12:19tokenization, embeddings, transformers,
2:12:22BERT, GPT models, prompt engineering,
2:12:25fine-tuning, RAG, and agentic
2:12:28architectures.
2:12:29And you can build projects like AI
2:12:31chatbots, document search tools,
2:12:34question and answer systems, or customer
2:12:36support bots.
2:12:37Option B is computer vision. If you like
2:12:40working with images and videos, choose
2:12:42computer vision. Learn CNNs, YOLO,
2:12:45object detection, and segmentation. And
2:12:48you can build real-world projects like
2:12:50face detection, medical image analysis,
2:12:53CCTV monitoring systems, or vehicle
2:12:55counting tools. Option C is generative
2:12:58AI. If you enjoy creativity, choose
2:13:01generative AI. Learn GANs, VAEs, and
2:13:05diffusion models. Build applications
2:13:07like AI art generators, product design
2:13:10tools, image-to-image systems, or video
2:13:12generation models. Option D is MLOps. If
2:13:16you prefer infrastructure and
2:13:18deployment, then choose MLOps. Learn
2:13:21Docker, Kubernetes, MLflow, CI/CD
2:13:23pipelines, cloud deployment, and model
2:13:26monitoring. And you can build projects
2:13:28that focus on deploying ML and LLM
2:13:31models into real environments. All
2:13:33right. So, agentic AI is the biggest
2:13:36trend of 2026.
2:13:38These are not just models, these are the
2:13:41intelligent agents that can reason,
2:13:43plan, use tools, and take actions. So,
2:13:46learn how agents work with frameworks
2:13:48like LangChain, LangGraph, and
2:13:51crew-based agent architectures.
2:13:53Understand tool calling, memory systems,
2:13:56planning, and multi-agent collaboration.
2:13:59Also, build agentic projects like an AI
2:14:02research assistant, an autonomous email
2:14:04automation agent, a financial analysis
2:14:07agent, or a customer service automation
2:14:10agent. So, these projects stand out in
2:14:13the interviews because companies want
2:14:15people who can build intelligent
2:14:17workflows and not just models.
2:14:19So, now that you have skills, you need
2:14:21projects that prove it. So, your
2:14:23portfolio should have two machine
2:14:25learning projects, two deep learning
2:14:27projects, two specialization projects,
2:14:29and one real end-to-end AI system. And
2:14:32this final project could be a chatbot
2:14:34with rack, a vision-based attendance
2:14:36system, an AI assistant with memory, or
2:14:40a complete ML pipeline deployed on
2:14:43cloud. And finally, upload your work on
2:14:46GitHub, write clear documentation, and
2:14:48add deployment links so employers can
2:14:51test your work instantly. So, with this
2:14:53roadmap, you can apply for the most
2:14:55in-demand roles in 2026,
2:14:58such as AI engineer, machine learning
2:15:00engineer, LLM engineer, NLP engineer,
2:15:04generative AI engineer, computer vision
2:15:06engineer, MLOps engineer, or AI
2:15:09automation specialist.
2:15:11AI is not just the future, it's the
2:15:14career shift of today. If you follow
2:15:16this roadmap step by step, you will
2:15:18build the skills, the projects, and the
2:15:20confidence to enter the world of AI and
2:15:23machine learning.
2:15:27>> [music]
2:15:30>> So guys, let's see what we are going to
2:15:32explore today. So today, we are going to
2:15:34explore real-world examples of machine
2:15:36learning, starting your journey, how you
2:15:38can start your journey to machine
2:15:40learning, and key concepts and impacts
2:15:42of machine learning in your day-to-day
2:15:44life. So guys, to help you navigate
2:15:47through this video, here is a quick
2:15:49rundown for you guys. Table of contents.
2:15:51What is machine learning? How does
2:15:53machine learning works? Five features of
2:15:55machine learning, types of machine
2:15:57learning, what skills one should have to
2:15:59learn machine learning, machine learning
2:16:01applications, and last but not the least
2:16:04guys, that is future of machine
2:16:06learning.
2:16:07So guys, I'm going to amaze you with
2:16:10this best example of Google Translator
2:16:12and AI which converts from one language
2:16:15to the another. I have chosen one
2:16:17language for you guys since many of you
2:16:19watch animes and all. So, I'm going to
2:16:21convert from Japanese language to the
2:16:24English language. Buckle up. Let's get
2:16:26started. So guys, I'm going to write a
2:16:29word in Japanese that is "Ohayo
2:16:32gozaimasu".
2:16:34If you know what "Ohayo gozaimasu"
2:16:36means, then please comment down below.
2:16:38So, I'm going to convert "Ohayo
2:16:40gozaimasu" from Japanese to English. So,
2:16:43let's copy this and paste this in here.
2:16:46So, "Ohayo gozaimasu" in Japanese, but
2:16:50in English it means good morning. A
2:16:52Google Translator works on a neural
2:16:54network. And a neural network is a
2:16:57machine learning algorithm which learns
2:16:59from Japanese language, keeps on
2:17:01improving itself, and gives you the
2:17:04optimal output that is good morning in
2:17:06English. This is how a language
2:17:08translator works. Let's see one more
2:17:11example. If I write here one more thing
2:17:14that is "Hajimemashite".
2:17:18"Hajimemashite" in Japanese, if you
2:17:20know, then please comment down below.
2:17:22Let me see what does it mean.
2:17:25So, "Hajimemashite" in Japanese, but in
2:17:27English it means nice to meet you. Here
2:17:30also the same. The neural network is
2:17:32learning from Japanese language and
2:17:34showing you what does it mean in English
2:17:37language. So, let me properly explain
2:17:40you what it actually does. So guys, a
2:17:43translator is nothing, but it's an AI
2:17:45machine that converts from one language
2:17:47to the another using a machine learning
2:17:50algorithm called neural network. Now,
2:17:52what neural network does is it learns
2:17:54from one language, makes some mistake or
2:17:56errors, then keeps on improving itself
2:17:58so that it can convert to to language
2:18:00such as from Japanese to English or from
2:18:03English to Japanese.
2:18:04So guys, the translator which you are
2:18:06using works on neural networks and this
2:18:09is how a translation works.
2:18:13So guys, you must be thinking then what
2:18:15is machine learning? So moving on to
2:18:17what is machine learning we have a
2:18:19machine learning is nothing but a
2:18:21training from historical data or
2:18:23experiences or in layman terms you can
2:18:25say learning from data to predict future
2:18:28or required output is called machine
2:18:30learning.
2:18:31Guys, a simple machine learning
2:18:33algorithm works in a way that you have a
2:18:35data of any kind and you give it to a
2:18:37machine. Now what does machine does is
2:18:40it learns from it in different ways
2:18:42using some different algorithms and
2:18:44gives you the required amount of future
2:18:47or predicts the required amount of
2:18:48output you wanted.
2:18:51Now guys, you must have understood what
2:18:53is machine learning. Let's deep dive a
2:18:55bit. Let's see what are the features of
2:18:57machine learning. We have five features
2:18:59of machine learning that is predictive
2:19:01modeling, automation, scalability,
2:19:05generalization, and adaptiveness. These
2:19:07are the five main features of any
2:19:09machine learning.
2:19:11So guys, tighten your seat belts. Let's
2:19:13move to these one by one. Predictive
2:19:15model. In the predictive model, what it
2:19:17does it it uses some mathematical
2:19:19functions and statistical techniques on
2:19:21the historical data and gives you the
2:19:24future predictions. The best example I
2:19:26can give you guys is the stock market
2:19:28prediction app where it uses some
2:19:29graphs, straight line graphs, and charts
2:19:32to show you the prediction based on the
2:19:34historical data using some statistical
2:19:36and mathematical functions. This is how
2:19:39a predictive model works.
2:19:42Moving on to our next topic that is
2:19:44automation. Automation is one of the
2:19:46best feature to save money. You know
2:19:48why? Because the companies which are
2:19:51having the less domains and cannot hire
2:19:53the employees, they can automate
2:19:55different machines to do the same work
2:19:58as the employee does. Such as if you
2:20:00want a developer, but you don't have the
2:20:02cost to pay, then you can automate a
2:20:04machine that can develop for you. This
2:20:06is the best example I can give you for
2:20:08the automation feature.
2:20:11So guys, moving on to our next feature,
2:20:13that is scalability. Scalability has its
2:20:15own importance because a machine
2:20:17learning algorithm, if it is not
2:20:19scalable, then it cannot handle larger
2:20:21amount of data sets or bigger data sets.
2:20:24I can give you the best example, that is
2:20:26Amazon.
2:20:28So guys, here I am at the Amazon
2:20:30website, and you can see a lot of
2:20:32product, and not only you can see, but
2:20:34whole world can see who are using Amazon
2:20:36app. Now, this system is a scalable
2:20:38system. No matter how many customers are
2:20:41here, and they are buying, the system
2:20:43will never crash. It can handle that
2:20:45amount of larger data sets. So, this is
2:20:48the best example of a scalable system.
2:20:51So guys, moving on to our next feature,
2:20:53that is generalization. In
2:20:55generalization, what it does is it's the
2:20:58ability of the model to generalize
2:21:01things, to forecast new data. Suppose
2:21:03your model is trained on a data set, and
2:21:06you're going to test it on some another
2:21:08data set, which is not there in the
2:21:09training, but still your model is giving
2:21:1295% of accuracy. Means your model is
2:21:15generalizing, your model is summarizing,
2:21:17and giving you the best and optimal
2:21:19result.
2:21:20Now guys, moving on to our last feature,
2:21:22that is adaptiveness. You can take it as
2:21:25a survival of the fittest thing, because
2:21:26if your model is not surviving the
2:21:28real-time environments or the new
2:21:30problems, then your model is not
2:21:32adaptive or good. It will going to
2:21:33extinct. Suppose there is a model which
2:21:36is built on traditional model, and still
2:21:38giving you best and advanced solutions
2:21:40on the real-time problems, then your
2:21:42model is adaptive and is the optimal
2:21:44model you can have.
2:21:46So guys, clear your mind because we are
2:21:48going to go in the types of machine
2:21:50learning. There are four different types
2:21:52of machine learning. First, we have is
2:21:54supervised or guided machine learning.
2:21:57In the supervised or guided machine
2:21:58learning, what it does is if there is a
2:22:00data set and having some values and you
2:22:03are labeling it as a specifying the
2:22:05value to the machine, then it will
2:22:07recognize those data through the labels
2:22:10and giving you the optimal
2:22:11classification. For example, if you have
2:22:14the pictures of cats and dogs and you
2:22:16have to classify it, then you will label
2:22:18it as cats and dogs. Then your machine
2:22:20will recognize those labels and classify
2:22:23and give you the optimal result.
2:22:26So, in supervised learning, what we have
2:22:28is a supervised algorithm. We have some
2:22:30labels and we are putting those labels
2:22:32to a data set. And after this, the
2:22:35machine easily recognizes and giving you
2:22:37the optimal results.
2:22:39So guys, moving on to our next type,
2:22:41that is unsupervised or unguided
2:22:44learning. In unsupervised or unguided
2:22:46learning, what machine does is you are
2:22:48giving the data which is not labeled.
2:22:50And after few trainings and making some
2:22:52errors, the machine easily recognizes
2:22:54this. So, what we have is a unsupervised
2:22:57algorithm. We have some unlabeled data
2:23:00and in those unlabeled data, your
2:23:02machine is trying to find patterns.
2:23:03After few errors, it will give you the
2:23:06optimal results.
2:23:07Moving on to our next type, that is
2:23:09semi-supervised learning. In
2:23:11semi-supervised learning, the algorithm
2:23:13uses both unsupervised and supervised in
2:23:16combined form giving you the optimal
2:23:18result. For example, we have a
2:23:20semi-supervised algorithm and we are
2:23:22trying to find out patterns from the
2:23:24data sets which are both labeled and
2:23:26unlabeled. So, this is how a
2:23:28semi-supervised learning works.
Types of Machine Learning Models
2:23:30Moving on to our last type, that is
2:23:33reinforcement learning. The
2:23:34reinforcement learning you can
2:23:36understand in a way like when you are
2:23:37playing a game, you make some mistake,
2:23:39then learn from them, then again make
2:23:41some mistake in a level, learn from them
2:23:43and reach your goal. This This how
2:23:45reinforcement learning works in a
2:23:47software. The software is being trained
2:23:49multiple times making some errors and
2:23:51learning from them and giving you the
2:23:53optimal results. This is the best
2:23:55example I can give you for the
2:23:57reinforcement learning that is the
2:23:58gaming system. So guys, what we have is
2:24:01a reinforcement algorithm and an
2:24:03environment. We are taking some actions
2:24:05in that environment, making some errors
2:24:07and getting some results. Then again
2:24:09making some errors and getting some
2:24:11results. This is how a reinforcement
2:24:13model works in a machine learning.
2:24:16So guys, moving on to the examples of
2:24:18machine learning, we have a voice
2:24:19recognition system and a image
2:24:22recognition system. These both you are
2:24:24using in your phone, in your laptop
2:24:26every day and machine learning is being
2:24:28used. Now guys, you have learned so
2:24:30much. Now you must be thinking that how
2:24:33should I start my journey? What skills I
2:24:35should have? So these are the skills
2:24:38required to learn machine learning. That
2:24:40is SQL, structured query language,
2:24:42JavaScript, C++, R programming for those
2:24:46who are moving with machine learning to
2:24:47the data scientist and Python, one of
2:24:50the best programming language for the
2:24:51machine learning. And last but not the
2:24:53least, if you are moving a bit deep down
2:24:56in the machine learning towards the deep
2:24:57learning, then you need NLP or natural
2:25:01language processing. These skills are
2:25:03required for the one who want to learn
2:25:05machine learning.
2:25:07So guys, moving on to the real-life
2:25:09examples of machine learning I have,
2:25:11that is Google searches. Google searches
2:25:14uses our history, track it down, then
2:25:16giving you the prediction based on your
2:25:18histories. For example, let me show you.
2:25:21So guys, I came here at Google. Now I'm
2:25:24going to type Amazon and it's giving me
2:25:28that results which are already there in
2:25:30my history. I must have searched before
2:25:32a month ago, uh 2 months ago, then 4
2:25:34months ago and all that history combined
2:25:37form is being predicted in here. Now
2:25:39Amazon Prime is there, Amazon videos are
2:25:42there. So, this is how a predictive
2:25:44model is working behind this Google
2:25:45searches.
2:25:47Moving on to our next real-life example,
2:25:49guys, that is Instagram. You swipe
2:25:52Instagram every day, every night. So,
2:25:54all that is based on your past data
2:25:57only. If you are seeing some videos of
2:25:58cats, then further swipes will be the
2:26:00cats only. So, this is how your history
2:26:03is being tracked down and watched by the
2:26:05Instagram machine learning algorithms
2:26:07and giving you the results of the same.
2:26:10So, this is how a real feeder in
2:26:13Instagram works. Moving on to our next
2:26:15example, that is movie recommendation
2:26:17system on any movie watching website
2:26:19such as Netflix or anime websites such
2:26:22as Watch Anime, Anycon, etc.
2:26:25So, guys, if I take you to Any watch and
2:26:28show you the anime recommendation
2:26:30trending ones are based on the people's
2:26:32choices they are watching more based on
2:26:34your histories only. If I watch any of
2:26:36these animes and I watch them regularly,
2:26:39then the recommendations will show me
2:26:41the same. So, this is how a
2:26:43recommendation system works in Any watch
2:26:46or you can say in Netflix based on your
2:26:48choices in past data.
2:26:51So, guys, moving on to the future of
2:26:53machine learning, it will be using
2:26:55everywhere. The advanced techniques in
2:26:57medical field, architecture field,
2:26:59electrical field, and yes, in the
2:27:01computer science field. Machine learning
2:27:03will be everywhere. It will be
2:27:04revolutionizing the world.
2:27:12Now, let's explore different types of ML
2:27:14models.
2:27:16So, not all data is structured the same
2:27:18way. And different problems require
2:27:19different approaches.
2:27:21So, for example, predicting stock prices
2:27:23requires the models that learn from
2:27:25historical trends.
2:27:26And then identifying objects in images
2:27:29needs models that recognize patterns in
2:27:31visual data.
2:27:32Next, the chatbots and voice assistants
2:27:35rely on the models trained to understand
2:27:37and generate human language.
2:27:39So, to tackle these challenges, as I
2:27:41discussed previously, that ML is divided
2:27:43into different learning models, such as
2:27:45supervised, unsupervised, and
2:27:48reinforcement learning. And each has its
2:27:50own strengths, and it is used depending
2:27:52on the problem at hand.
2:27:54Since we know why different ML models
2:27:56are needed, let's see how they play a
2:27:58crucial role in generative AI.
2:28:00Well, generative AI is one of the most
2:28:03exciting applications of machine
2:28:04learning. And unlike traditional ML
2:28:07models that make predictions or
2:28:08classifications, generative models
2:28:10create entirely new content. And here's
2:28:13how ML enables AI to generate.
2:28:15So, first here we have text. A language
2:28:18models, like GPT, generate human-like
2:28:20text for chatbots, content writing, and
2:28:22coding.
2:28:23Next is the image.
2:28:25So, AI-powered tools, like DALL-E, can
2:28:27create realistic images from textual
2:28:30descriptions.
2:28:31Next is videos. So, advanced ML models
2:28:34synthesize lifelike video content,
2:28:37transforming media, marketing, and even
2:28:39filmmaking.
2:28:40So, these advancements in generative AI
2:28:42are reshaping creativity and automation,
2:28:45proving that machine learning is not
2:28:46just about making decision, it's about
2:28:48creating new possibilities.
2:28:50So, now that we have seen how ML models
2:28:52enable AI to create new content. So, now
2:28:55let us briefly understand the different
2:28:56types of machine learning models.
2:28:58So, here, the first type of machine
2:29:00learning model is supervised learning.
2:29:03Supervised learning trains a model using
2:29:05labeled data, where each input has a
2:29:07corresponding correct output. And this
2:29:09makes it ideal for tasks where
2:29:10historical data can be used to predict
2:29:12future outcomes.
2:29:14For example, let's say spam detection.
2:29:17Email services, like Gmail, use a
2:29:19supervised learning to classify emails
2:29:21as spam or not spam by learning from
2:29:24past labeled examples.
2:29:26The next example is the price
2:29:27predictions.
2:29:29So, real estate platforms use regression
2:29:31models to predict house prices based on
2:29:33the features like location, size, and
2:29:36amenities.
2:29:37Now, let us see some of the popular
2:29:39algorithms.
2:29:40So, first let's discuss on decision
2:29:42trees. These models break down the data
2:29:45into a tree-like structure, where each
2:29:47node represent a decision based on a
2:29:49feature.
2:29:50So, they are easy to interpret and work
2:29:52well for both classification. For
2:29:54example, deciding if an email is a spam
2:29:57or not. And regression example,
2:29:59predicting house price.
2:30:01However, they can become overly complex.
2:30:04Next is the support vector machines.
2:30:07So, SVMs are powerful for classification
2:30:09task, as they find the optimal boundary,
2:30:11also called a hyperplane. And that best
2:30:14separates different classes in the data.
2:30:17They work well for high-dimensional
2:30:19spaces and cases where the distinction
2:30:21between categories is clear, such as
2:30:23handwriting, facial recognition, or
2:30:26medical diagnosis.
2:30:27So, now that we have seen how labeled
2:30:29data is used. So, now let's explore how
2:30:31unsupervised learning finds patterns
2:30:33without labels.
Machine Learning Algorithm
2:30:35Well, unsupervised learning works with
2:30:37unlabeled data, identifying hidden
2:30:39patterns and relationships without
2:30:41predefined categories.
2:30:42So, here we have some of the popular
2:30:44algorithms. So, first is the K-means
2:30:47clustering. This algorithm partitions
2:30:49data into a predefined number of
2:30:51clusters.
2:30:52By grouping similar data points based on
2:30:54their attributes. It works well for
2:30:57tasks like customer segmentation. Where
2:30:59businesses can group customer based on
2:31:02purchasing behavior. However, it assumes
2:31:04clusters are spherical and may struggle
2:31:07with irregular shaped data. Next we have
2:31:10autoencoders.
2:31:12So, these are specialized neural
2:31:13networks designed to learn efficient
2:31:15data representations by encoding and
2:31:18reconstructing input data. Let us see
2:31:20some of the examples.
2:31:22So, first example here we have is
2:31:24customer segmentation.
2:31:26Where e-commerce platforms group
2:31:28customer based on their shopping
2:31:29behavior to offer personalized
2:31:31recommendations.
2:31:32The next example is market analysis.
2:31:35Businesses analyze purchasing trends to
2:31:37find associations such as which products
2:31:40are frequently brought together.
2:31:42Now we have covered both labeled and
2:31:43unlabeled learning. So let's see how
2:31:45semi-supervised learning combines the
2:31:47best of both worlds.
2:31:49So semi-supervised learning bridges the
2:31:52gap between the supervised and
2:31:53unsupervised learning by using a small
2:31:55amount of data along with large amount
2:31:58of unlabeled data.
2:31:59So for example, let's say AI assistant
2:32:01medical diagnosis.
2:32:03Labeled medical images such as x-rays
2:32:06with diagnosis are scarce, but large
2:32:08amounts of unlabeled images exist.
2:32:11Semi-supervised learning help AI learn
2:32:14patterns from both labeled and unlabeled
2:32:16data improving accuracy in disease
2:32:18detection. All right. Now let's explore
2:32:21the reinforcement learning where AI
2:32:23learns through trial and error.
2:32:25Well, reinforcement learning is inspired
2:32:28by the concept of learning through trial
2:32:30and error. So models interact with an
2:32:32environment, receive rewards or
2:32:34penalties for actions and refine their
2:32:37strength over time. For example, let's
2:32:39say gaming.
2:32:40Mario AI developed using reinforcement
2:32:42learning learns to navigate levels by
2:32:45optimizing actions through trial and
2:32:47error. The next example is robotics.
2:32:50Where robots learn to walk, balance, or
2:32:52perform tasks through reinforcement
2:32:54learning by maximizing positive
2:32:56outcomes.
2:32:57Also, reinforcement learning uses
2:32:59agents, actions, and rewards to improve
2:33:02decision-making.
2:33:03Making it ideal for tasks requiring
2:33:05continuous learning and adaptation.
2:33:08So now that we have covered all the
2:33:09types of machine learning models. So
2:33:11let's go over some of the key tips to
2:33:13help you choose the right one for your
2:33:14needs.
2:33:16So here are the tips.
2:33:17When it comes to supervised learning
2:33:19classifying emails as spam or not and
2:33:22diagnosing diseases from patient data.
2:33:25Next is the unsupervised learning. So,
2:33:27unsupervised learning is best when
2:33:29you're grouping shoppers by behavior and
2:33:31detecting fraud in banking. Next, we
2:33:33have semi-supervised learning.
2:33:35And this is best when you're improving
2:33:37speech recognition with limited label
2:33:39data and identifying fake news.
2:33:42And finally, the reinforcement learning.
2:33:45This will be best when you're training
2:33:46self-driving cars to navigate,
2:33:48optimizing AI in video games like Mario.
2:33:51So, whether it's supervised,
2:33:53unsupervised, semi-supervised, or
2:33:55reinforcement learning, each model plays
2:33:57a crucial role in shaping AI's future.
2:34:00So, as generative AI continues to
2:34:02evolve, these models are driving
2:34:03innovation in text, images, and video
2:34:06generation.
2:34:07So, which machine learning model do you
2:34:09find the most fascinating? Let me know
2:34:11in the comments below.
2:34:15>> [music]
2:34:18>> Let me connect you to the real life and
2:34:20tell you what all are the things which
2:34:22you can easily do using the concepts of
2:34:23machine learning.
2:34:25So, you can easily get answer to the
2:34:26questions like which types of house lies
2:34:28in this segment or what is the market
2:34:30value of this house? Or is this a mail a
2:34:33spam or not a spam? Is there any fraud?
2:34:36Well, these are some of the question you
2:34:37could ask to the machine. But for
2:34:38getting an answer to these, you need
2:34:40some algorithm. The machine need to
2:34:42train on the basis of some algorithm.
2:34:44Okay, but how will you decide which
2:34:46algorithm to choose and when?
2:34:48Okay, so the best option for us is to
2:34:50explore them one by one.
2:34:53So, the first is classification
2:34:54algorithm where the category is
2:34:56predicted using the data. If you have
2:34:58some question like is this person a male
2:35:01or a female? Or is this a mail a spam or
2:35:04not a spam? Then these category of
2:35:06question would fall under the
2:35:07classification algorithm.
2:35:09Classification is a supervised learning
2:35:10approach in which the computer program
2:35:13learns from the input given to it and
2:35:14then uses this learning to classify new
2:35:17observation. Some examples of
2:35:19classification problems are speech
2:35:21organization, handwriting recognition,
2:35:23biometric identification, document
2:35:25classification, etc.
2:35:28Shall we move ahead?
2:35:30Okay.
2:35:32So, next is the anomaly detection
2:35:34algorithm where you identify the unusual
2:35:37data point. So, what is anomaly
2:35:38detection? Well, it's a technique that
2:35:40is used to identify unusual pattern that
2:35:43do not conform to expected behavior. Or
2:35:45you can say the outliers.
2:35:47It has many application in business like
2:35:49intrusion detection, like identifying
2:35:51strange patterns in the network traffic
2:35:53that could signal a hack, or system
2:35:55health monitoring, that is spotting a
2:35:56deadly tumor in the MRI scan.
2:35:59Or you can even use it for fraud
2:36:01detection in credit card transaction, or
2:36:03to deal with fault detection in
2:36:04operating environment.
2:36:06So, next comes the clustering algorithm.
2:36:08You can use this clustering algorithm to
2:36:10group the data based on some similar
2:36:12condition. Now, you can get answer to
2:36:14which type of houses lies in this
2:36:16segment, or what type of customer buys
2:36:18this product. The clustering is a task
2:36:20of dividing the population or data
2:36:22points into a number of groups such that
2:36:24the data point in the same groups are
2:36:26more similar to other data points in the
2:36:28same group than those in the other
2:36:30groups. In simple words, the aim is to
2:36:33segregate groups with similar trait and
2:36:35assign them into cluster.
2:36:37Now, this clustering is a task of
2:36:38dividing the population or data points
2:36:40into a number of groups such that the
2:36:42data points in the X group is more
2:36:44similar to the other data points in the
2:36:46same group rather than those in the
2:36:48other group. In other words, the aim is
2:36:50to segregate the groups with similar
2:36:52traits and assign them into different
2:36:54clusters. Let's understand this with an
2:36:56example. Suppose you're the head of a
2:36:58rental store and you wish to understand
2:37:00the preference of your customer to scale
2:37:02up your business. So, is it possible for
2:37:04you to look at the detail of each
2:37:05customer and design a unique business
2:37:08strategy for each of them?
2:37:09Definitely not. Right?
2:37:12But what you can do is to cluster all
2:37:14your customer saying to 10 different
2:37:16groups based on their purchasing habit
2:37:18and you can use a separate strategy for
2:37:20customers in each of these 10 different
2:37:22groups. And this is what we call
2:37:24clustering.
2:37:26Next we have regression algorithm where
2:37:28the data itself is predicted. Question
2:37:30you may ask to this type of model is
2:37:32like what is the market value of this
2:37:34house or is it going to rain tomorrow or
2:37:36not?
2:37:37So regression is one of the most
2:37:39important and broadly used machine
2:37:40learning and statistics tool.
2:37:43It allows you to make prediction from
2:37:44data by learning the relationship
2:37:46between the features of your data and
2:37:48some observed continuous valued
2:37:49response. Regression is used in a
2:37:52massive number of application.
2:37:54You know what? Stock prices prediction
2:37:55can be done using regression.
2:37:57Now you know about different machine
2:37:59learning algorithm. How will you decide
2:38:01which algorithm to choose and when?
2:38:03So let's cover this part using a demo.
2:38:05So in this demo part, what we'll do,
2:38:07we'll create six different machine
2:38:08learning model and pick the best model
2:38:11and build the confidence such that it
2:38:12has the most reliable accuracy.
2:38:16So for our demo part, we'll be using the
2:38:17Iris data set. This data set is quite
2:38:20very famous and is considered one of the
2:38:22best small project to start with.
2:38:24You can consider this as a hello world
2:38:25data set for machine learning. So this
2:38:27data set consists of 150 observation of
2:38:30Iris flower.
2:38:31There are four columns of measurement of
2:38:33flowers in centimeters. The fifth column
2:38:35being the species of the flower
2:38:36observed. All the observed flowers
2:38:38belong to one of the three species of
2:38:40Iris setosa, Iris virginica and Iris
2:38:43versicolor.
2:38:44Well, this is a good project because it
2:38:46is so well to understand. The attributes
2:38:48are numeric so you have to figure out
2:38:49how to load and handle the data. It is a
2:38:51classification problem thereby allowing
2:38:53you to practice with perhaps an easier
2:38:55type of supervised learning algorithm.
2:38:57It has only four attributes and 150 rows
2:38:59meaning it is very small and can easily
2:39:01fit into the memory.
2:39:02And even all of the numeric attributes
2:39:04are in same unit and the same scale. It
2:39:07means you do not require any special
2:39:08scaling or transformation to get
2:39:10started.
2:39:12So, let's start coding and as I told
2:39:14earlier for the demo part, I'll be using
2:39:16Anaconda with Python 3.0 installed on
2:39:18it. So, when you install Anaconda, how
2:39:20your navigator would look like. So,
2:39:22there's my home page of my Anaconda
2:39:23navigator. On this I'll be using the
2:39:26Jupiter notebook, which is a web-based
2:39:27interactive computing notebook
2:39:29environment, which will help me to write
2:39:30and execute my Python codes on it. So,
2:39:32let's hit the launch button and execute
2:39:34our Jupiter notebook.
2:39:36So, as you can see that my Jupiter
2:39:37notebook is starting on localhost 8890.
2:39:41Okay? So, this is my Jupiter notebook.
2:39:42What I'll do here, I'll select new
2:39:44notebook Python 3.
2:39:48There's my environment where I can write
2:39:50and execute all my Python codes on it.
2:39:52So, let's start by checking the version
2:39:54of the libraries. In order to make this
2:39:56video short and more interactive and
2:39:57more informative, I've already done the
2:39:59set of code. So, let me just copy and
2:40:01paste it down. I'll explain you then one
2:40:03by one.
2:40:04So, let's start by checking the version
2:40:05of the Python libraries.
2:40:07Okay? So, there's the code. Let's just
2:40:10copy it.
2:40:11Copied and let's paste it. Okay. First,
2:40:14let me summarize things for you. What we
2:40:16are doing here, we are just checking the
2:40:17version of the different libraries.
2:40:19Starting with Python, we'll first check
2:40:20what version of Python we are working
2:40:22on, then we'll check what are the
2:40:23version of SciPy we are using, then
2:40:25NumPy, Matplotlib, then Pandas, then
2:40:27scikit-learn. Okay? So, let's execute
2:40:29the run button and see what are the
2:40:30various version of libraries which we
2:40:32are using. Hit the run. So, we are
2:40:33working on Python 3.6.4, SciPy 1.0,
2:40:37NumPy 1.14, Matplotlib 2.12, Pandas
2:40:400.22, and scikit-learn of version 0.19.
2:40:44Okay?
2:40:45So, these are the version which I'm
2:40:46using. Ideally, your version should be
2:40:48more recent or it should match. But,
2:40:50don't worry if you lag few versions
2:40:52behind as the APIs do not change so
2:40:54quickly. Everything in this tutorial
2:40:56will very likely still work for you.
2:40:58Okay? But, in case you're getting an
2:41:00error, stop and try to fix that error.
2:41:03In case you're unable to find the
2:41:04solution for the error, feel free to
2:41:06reach out Edureka even after this class.
2:41:08Let me tell you this, if you're not able
2:41:09to run the script properly, you will not
2:41:11be able to complete this tutorial, okay?
2:41:13So, whenever you get a doubt, reach out
2:41:15to Edureka and just resolve it.
2:41:17Now, if everything is working smoothly,
2:41:19then now it's the time to load the data
2:41:21set. So, as I said, I'll be using the
2:41:23Iris flower data set for this tutorial.
2:41:25But, before loading the data set, let's
2:41:27import all the modules, function, and
2:41:29the object which we are going to use in
2:41:31this tutorial. Same, I've already
2:41:32written the set of code, so let's just
2:41:34copy and paste them. Let's load all the
2:41:36libraries.
2:41:38So, these are the various libraries
2:41:39which we'll be using in our tutorial.
2:41:42So, everything should work fine without
2:41:43an error. If you get an error, just
2:41:45stop. You need to work on your SciPy
2:41:46environment before you continue any
2:41:48further. So, I guess everything should
2:41:50work fine. Let's hit the run button and
2:41:51see.
2:41:53Okay, it worked. So, let's now move
2:41:56ahead and load the data. We can load the
2:41:58data direct from the UCI machine
2:41:59learning repository. First of all, let
2:42:01me tell you, we are using Panda to load
2:42:03the data.
2:42:04Okay?
2:42:05So, let's say my URL is this. So, this
2:42:08is my URL for the UCI machine learning
2:42:09repository from where I'll be
2:42:10downloading the data set, okay?
2:42:13Now, what I'll do, I'll specify the name
2:42:14of each column when loading the data.
2:42:16This will help me later to explore the
2:42:18data, okay?
2:42:19So, I'll just copy and paste it down.
2:42:22Okay?
2:42:23So, I'm defining a variable names which
2:42:25consists of various parameters including
2:42:27sepal length, sepal width, petal length,
2:42:29petal width, and class. So, these are
2:42:31just the name of column from the data
2:42:32set, okay? Now, let's define the data
2:42:35set. So, data set equals panda.read_csv.
2:42:39Inside that, we are defining URL and the
2:42:41names, that is equal to name.
2:42:44As I already said, we'll be using Panda
2:42:46to load the data, all right?
2:42:49So, we are using panda.read_csv, so we
2:42:51are reading the CSV file, and inside
2:42:53that, from where that CSV is coming?
2:42:54From the the Which URL? So, this is my
2:42:56URL. Okay?
2:42:58And names equal names. It's just
2:43:00specifying the names of the various
2:43:01columns in that particular CSV file.
2:43:03Okay?
2:43:04So, let's move forward and execute it.
2:43:06So, even our data set is loaded.
2:43:09In case you have some network issues,
2:43:11just go ahead and download the Iris data
2:43:13file into your working directory and
2:43:14load it using the same method. But yeah,
2:43:16make sure that you change the URL to the
2:43:18local name, or else you might get an
2:43:19error. Okay. Yeah, our data set is
2:43:22loaded. So, let's move ahead and check
2:43:23our data set. Let's see how many columns
2:43:25or rows we have in our data set. Okay.
2:43:28So, let's print the number of rows and
2:43:30columns in our data set. So, our data
2:43:32set is data set.shape.
2:43:35What this will do, it will just give you
2:43:37the numbers of total number of rows and
2:43:39total number of column, or you can say
2:43:40the total number of instances or
2:43:42attributes in your data set. Fine?
2:43:44So, print data set.shape. What are you
2:43:46getting? 150 and 5. So, 150 is the total
2:43:49number of rows in your data set, and 5
2:43:50is the total number of columns. Fine?
2:43:53So, moving on ahead, what if I want to
2:43:55see the sample data set? Okay. So, let
2:43:58me just print the first 30 instances of
2:43:59the data set. Okay? So, print
2:44:03data set.head.
2:44:07What I want is the first 30 instances.
2:44:09Fine? This will give me the first 30
2:44:11result of my data set. Okay? So, when I
2:44:13hit the run button, what I'm getting is
2:44:15the first 30 result. Okay?
2:44:180
2:44:20to 29. So, this is how my sample data
2:44:22set looks like.
2:44:23Sepal length, sepal width, petal length,
2:44:25petal width, and the class. Okay?
2:44:28So, this is how our data set looks like.
2:44:31Now, let's move on and look at the
2:44:32summary of each attribute. What if I
2:44:34want to find out the count, mean, the
2:44:37minimum and the maximum values, and some
2:44:39other percentiles as well. So, what
2:44:40should I do then?
2:44:41For that, print
2:44:43data set.describe.
2:44:47What it will give,
2:44:48let's see.
2:44:50So, you can see that all the numbers are
2:44:52the same scales of similar range between
2:44:540 to 8 cm, right? The mean value, the
2:44:57standard deviation, the minimum value,
2:44:59the 25th percentile, 50th percentile,
2:45:0175th percentile, the maximum value, all
2:45:03these values lies in the range between 0
2:45:05to 8 cm.
2:45:07Okay.
2:45:08So what we just did is we just took a
2:45:10summary of each attribute. Now let's
2:45:13look at the number of instances that
2:45:14belong to each class. So for that, what
2:45:17we'll do print data set first of all.
2:45:21So let's print data set and I want to
2:45:24group it
2:45:25group by using
2:45:28class
2:45:30and I want the size of it, size of each
2:45:32class. Fine?
2:45:34And let's hit the run.
2:45:47Okay. So what I want to do, I want to
2:45:49print print what? Data set. How I want
2:45:52to get it? I want it by class. So group
2:45:54by class.
2:45:56Okay. Now I want the size of each class.
2:45:59Find the size of each class. So group by
2:46:01class.size. Execute the run.
2:46:04So you can see that I have 15 instances
2:46:06of Iris setosa, 15 instances of Iris
2:46:08versicolor, and 15 instances of Iris
2:46:10virginica. Okay? All are of data type
2:46:13integer of base 64. Fine? So now we have
2:46:16a basic idea of our data. Now let's move
2:46:18ahead and create some visualization for
2:46:20it. So for this we are going to create
2:46:22two different types of plot. First would
2:46:23be the univariate plot and the next
2:46:25would be the multivariate plot. So we'll
2:46:26be creating univariate plots to better
2:46:28understand about each attribute. And the
2:46:30next we'll be creating the multivariate
2:46:32plot to better understand the
2:46:33relationship between different
2:46:34attributes. Okay? So we start with some
2:46:36univariate plot. That is plot of each
2:46:38individual variable. So given that the
2:46:40input variables are numeric, we can
2:46:41create box and whiskers plot for it.
2:46:43Okay? So let's move ahead and create a
2:46:44box and whiskers plot. So data set.plot.
2:46:47What kind I want? It's a box.
2:46:50Okay. And do I need a subplot? Yeah, I
2:46:53need subplots for that. So, subplots
2:46:55equal true. What type of layout do I
2:46:57want? So, my layout structure is 2 cross
2:47:012.
2:47:02Next, do I want to share my coordinates,
2:47:04X and Y coordinates? No, I don't want to
2:47:06share it. So, share X equal false.
2:47:09And even share Y, that too equals false.
2:47:13Okay. So, we have our dataset.plot kind
2:47:16equal box. My subplots is true, layout 2
2:47:19cross 2. And then what I want to do it,
2:47:21I want to see it. So, plot.show.
2:47:23Whatever I created, show it. Okay.
2:47:26Execute it.
2:47:29Now, this gives us a much clearer idea
2:47:30about the distribution of the input
2:47:32attribute. Now, what if I had given the
2:47:33layout to 2 cross 2 instead of that, I'd
2:47:36have given it 4 cross 4. So, what it
2:47:39will result? Just see. Fine. Everything
2:47:41would be printed in just one single row.
2:47:43Hold on, guys. Arya has a doubt. He's
2:47:44asking that why we are using the share X
2:47:46and share Y values. What are these? Why
2:47:48we have assigned false values to it?
2:47:50Okay, Arya. So, in order to resolve this
2:47:52query, I need to show you what will
2:47:54happen if I give true values to them.
2:47:55Okay. So, be with me. So, share X equal
2:47:58true and share Y, that equals true. So,
2:48:01let's see what result we'll get.
2:48:04You're getting it. The X and Y
2:48:05coordinates are just shared among all
2:48:07the four visualization, right? So, Arya,
2:48:09you can see that the sepal length and
2:48:10sepal width has Y values ranging from
2:48:130.0 to 7.5 which are being shared among
2:48:15both the visualization. So, is with the
2:48:17petal length, it has shared value
2:48:19between 0.0 to 7.5. Okay. So, that is
2:48:22why I don't want to share the value of X
2:48:24and Y. It's just giving us a cluttered
2:48:26visualization. So, Arya, why I'm doing
2:48:28this? I'm just doing it cuz I don't want
2:48:31my X and Y coordinates to be shared
2:48:33among any visualization. Okay. That is
2:48:35why my share X and share Y value are
2:48:37false. Okay. Let's execute it.
2:48:40So, this is a
2:48:41pretty much clear visualization which
2:48:43gives a clear idea about the
2:48:44distribution of the input attributes.
2:48:46Now, if you want, you can also create a
2:48:48histogram of each input variable to get
2:48:50a clear idea of the distribution. So,
2:48:52let's create a histogram for it. So,
2:48:53dataset.hist, okay? I would need to
2:48:56proceed. So, plot.show. Let's see. So,
2:48:58this is my histogram and it seems that
2:49:00we have two input variables that have a
2:49:02Gaussian distribution. So, this is
2:49:04useful to note as we can use the
2:49:06algorithms that can exploit this
2:49:07assumption, okay? So, next comes the
2:49:09multivariate plot. Now that we have
2:49:10created the univariate plot to
2:49:12understand about each attribute, let's
2:49:14move on and look at the multivariate
2:49:16plot and see the interaction between the
2:49:18different variables. So, first let's
2:49:20look at the scatter plot of all the
2:49:21attribute. This can be helpful to spot
2:49:23structured relationship between input
2:49:24variables, okay? So, let's create a
2:49:26scatter matrix. So, for creating a
2:49:27scatter plot, we need scatter matrix and
2:49:31we need to pass our dataset into it,
2:49:33okay? And then, what I want, I want to
2:49:35see it. So, plot.show. So, this is how
2:49:37my scatter matrix looks like. It's like
2:49:39that the diagonal grouping of some pair,
2:49:41right? So, this suggests a high
2:49:42correlation and a predictable
2:49:44relationship, all right? This was our
2:49:45multivariate plot. Now, let's move on
2:49:47and evaluate some algorithm. Now, it's
2:49:49time to create some model of the data
2:49:51and estimate the accuracy on the base of
2:49:53unseen data, okay? So, now we know all
2:49:56about our dataset, right? We know how
2:49:58many instances and attributes are there
2:49:59in our dataset. We know the summary of
2:50:01each attribute. Now, I guess we have
2:50:03seen much about our dataset. Now, let's
2:50:05move on and create some algorithm and
2:50:07estimate their accuracy based on the
2:50:09unseen data. Okay. Now, what we'll do,
2:50:11we'll create some model of the data and
2:50:13estimate the accuracy based on the some
2:50:15unseen data, okay? So, for that, first
2:50:17of all, let's create a validation
2:50:18dataset. What is a validation dataset?
2:50:20Validation dataset is your training
2:50:22dataset that will be using it to train
2:50:24our model, fine? All right. So, how
2:50:26we'll create a validation dataset? For
2:50:28creating a validation dataset, what we
2:50:30are going to do is we are going to split
2:50:31our dataset into two part, okay? So, the
2:50:33very first thing we'll do is to create a
2:50:35validation dataset. So, why do we even
2:50:37need a validation data set? So, we need
2:50:39a validation data set to know that the
2:50:41model we created is any good. Later,
2:50:43what we'll do, we'll use the statistical
2:50:45method to estimate the accuracy of the
2:50:47model that we create on the unseen data.
2:50:49We also want a more concrete estimate of
2:50:51the accuracy of the best model on unseen
2:50:53data by evaluating it on the actual
2:50:55unseen data. Okay? Confused? Let me
2:50:57simplify this for you. What we'll do,
2:50:59we'll split the loaded data into two
2:51:00parts. The first 80% of the data we'll
2:51:03use it to train our model. And the rest
2:51:0520% we'll hold back as the validation
2:51:07data set that we'll use it to verify our
2:51:09trained model. Okay? Fine. So, let's
2:51:11define an array. This is my array. What
2:51:14it will consist of? It will consist of
2:51:16all the values from the data set. So,
2:51:17data set.values.
2:51:19Okay? Next, I'll define a variable X
2:51:22which will consist of all the column
2:51:25from the array from zero to four.
2:51:28Starting from zero to four. And the next
2:51:30variable Y which would consist of the
2:51:33array starting from this. So, first of
2:51:37all, we'll define a variable X that will
2:51:39consist of the values in the array
2:51:41starting from the beginning zero till
2:51:43four. Okay? So, these are the column
2:51:45which we'll include in the X variable.
2:51:47And for a Y variable, I'll define it as
2:51:49a class or the output. So, what I need,
2:51:51I just need the fourth column that is my
2:51:53class column. So, I'll start it from the
2:51:55beginning and I just want the fourth
2:51:57column. Okay? Now, I'll define the my
2:51:59validation size.
2:52:01validation_size.
2:52:04I'll define it as 0.20 and I'll use a
2:52:07seed.
2:52:08I'll define seed equals six.
2:52:11So, this method seed sets the integer
2:52:13starting value used in generating random
2:52:15number. Okay? I'll define the value of
2:52:17seed equals six. I'll tell you what is
2:52:19the importance of it later on. Okay? So,
2:52:21let me define first few variables such
2:52:23as X_train, test, Y_train,
2:52:27and Y_test.
Linear Regression Algorithm
2:52:30Okay? So, what we want to do is select
2:52:32some model. Okay. So, model underscore
2:52:34selection. But, before doing that, what
2:52:36we have to do is split our training data
2:52:37set into two halves. Okay. So, dot train
2:52:39underscore test underscore split. What
2:52:42we want to split is the value of X and
2:52:45Y. Okay. And my test size is
2:52:49equals to validation size.
2:52:52Which is a 0.20. Correct? And my random
2:52:55state
2:52:57is equal to seed. So, what the seed is
2:52:59doing here, it's helping me to keep the
2:53:01same randomness in the training and
2:53:03testing data set. Fine. So, let's
2:53:05execute it and see what is our result.
2:53:08Let's execute it. Next, we'll create a
2:53:10test harness. For this, we'll use
2:53:1210-fold cross-validation to estimate the
2:53:14accuracy.
2:53:16So, what it will do, it will split our
2:53:17data set into 10 parts. Train on the
2:53:20nine part and test on the one part. And
2:53:22this will repeat for all combination of
2:53:24train and test splits. Okay. So, for
2:53:26that, let's define again
2:53:29my seed that was six, already defined,
2:53:32and scoring
2:53:34equals accuracy.
2:53:37Fine.
2:53:37So, we are using the metric of accuracy
2:53:39to evaluate the model. So, what is this?
2:53:42This is a ratio of number of correctly
2:53:44predicted instances divided by the total
2:53:46number of instances in the data set
2:53:48multiplied by 100, giving a percentage.
2:53:50Example, it's 98% accurate or 99%
2:53:54accurate, things like that. Okay. So,
2:53:55we'll be using the scoring variable when
2:53:57we run the build and evaluate each model
2:54:00in the next step. So, next part is
2:54:02building model.
2:54:04Till now, we don't know which algorithm
2:54:05would be good for this problem or what
2:54:07configuration to use. So, let's begin
2:54:09with six different algorithm. I'll be
2:54:11using logistic regression, linear
2:54:13discriminant analysis, K-nearest
2:54:15neighbor, classification and regression
2:54:17trees, Naive Bayes, and support vector
2:54:19machine. Well, these algorithms which
2:54:21I'm using is a good mixture of simple
2:54:23linear or non-linear algorithms. In
2:54:25simple linear which included the
2:54:26logistic regression and the linear
2:54:28discriminant analysis or the non-linear
2:54:30part which included the KNN algorithm,
2:54:32the CART algorithm, the Naive Bayes, and
2:54:34the support vector machines. Okay. So,
2:54:36we reset the random number seed before
2:54:38each run to ensure that evaluation of
2:54:40each algorithm is performed using
2:54:42exactly the same data splits. It ensures
2:54:44the result are directly comparable.
2:54:46Okay. So, let me just copy and paste it.
2:54:49Okay.
2:54:53So, what we're doing here, we're
2:54:54building five different types of model.
2:54:56We're building a logistic regression,
2:54:58linear discriminant analysis, K-nearest
2:55:00neighbor, decision tree, Gaussian Naive
2:55:02Bayes, and the support vector machine.
2:55:04Okay. Next, what we'll do, we'll
2:55:05evaluate model in each turn. Okay.
2:55:08So, what is this? So, we have six
2:55:10different model and accuracy estimation
2:55:12for each one of them. Now, we need to
2:55:14compare the model to each other and
2:55:15select the most accurate of them all.
2:55:17So, running this script, we saw the
2:55:19following result. So, we can see some of
2:55:21the result on the screen. What is this?
2:55:23It is just the accuracy score using
2:55:24different set of algorithms. Okay. When
2:55:27we are using logistic regression, what
2:55:28is the accuracy rate? When we are using
2:55:30linear discriminant algorithm, what is
2:55:32the accuracy? And so on and so. Okay.
2:55:34So, from the output, it seems that LD
2:55:36algorithm was the most accurate model
2:55:38that we tested. Now, we want to get an
2:55:40idea of the accuracy of the model on our
2:55:42validation set or the testing data set.
2:55:44So, this will give us a independent
2:55:46final check on the accuracy of the best
2:55:47model. It is always valuable to keep a
2:55:50testing data set for just in case you
2:55:52made a over-fitting to the testing data
2:55:54set or you made a data leak. Both will
2:55:56result in a overly optimistic result.
2:55:58Okay.
2:55:59You can run the LD model directly on the
2:56:01validation set and summarize the result
2:56:03as a final score, a confusion matrix,
2:56:06and a classification report.
2:56:09>> [music]
2:56:13>> Let us understand what regression in
2:56:15machine learning is.
2:56:16So, what exactly is regression?
2:56:18The main goal of regression is the
2:56:20construction of an efficient model to
2:56:22predict the dependent attributes from a
2:56:24bunch of attribute variables.
2:56:26A regression problem is where the output
2:56:28variable is either real or a continuous
2:56:30value like salary, weight, area, etc.
2:56:33We can also define regression as a
2:56:35statistical means that is used in
2:56:36applications like housing, investing,
2:56:38etc. to predict the relationship between
2:56:40a dependent variable and a bunch of
2:56:42independent variables.
2:56:44For example, let's say in the finance
2:56:46application or investing, we can
2:56:48actually predict the values of certain
2:56:50stock prices or you know those values
2:56:52depending on the independent variables
2:56:55like how many years it takes for a stock
2:56:57to you know actually mature or how many
2:56:59days will it take to grow or those
2:57:01variables that you have in investing and
2:57:04depending upon that we can make a
2:57:05possible outcome or a possible
2:57:06prediction of how a stock is going to be
2:57:09invested in a profit state or a loss
2:57:11state or all those things or we can take
2:57:13another example like housing. We can
2:57:15take different parameters like number of
2:57:17years it's been there, how many people
2:57:19have used it or what is the area of the
2:57:22house depending on all these factors or
2:57:24how many rooms does the house have, we
2:57:26can predict the price of a house.
2:57:28So this is basically what regression
2:57:30really is.
2:57:31So let us take a look at the various
2:57:32types of regression techniques that we
2:57:34have.
2:57:35We have simple linear regression, then
2:57:36we have polynomial regression, support
2:57:38vector regression, decision tree
2:57:40regression, we have random forest
2:57:42regression and we have logistic
2:57:43regression as well. That is also a type
2:57:45of regression that we have. But for now
2:57:47we'll be focusing on simple linear
2:57:49regression.
2:57:50So let's talk about how or what exactly
2:57:52is simple linear regression first. So
2:57:54one of the most interesting and common
2:57:56regression technique is a simple linear
2:57:57regression. In this we predict the
2:57:59outcome of a dependent variable Y based
2:58:02on the independent variables X. So the
2:58:04relationship between the variables is
2:58:06linear, hence the word linear
2:58:08regression.
2:58:09Then comes the polynomial regression.
2:58:11So in this regression technique, we
2:58:13transform the original features into a
2:58:15polynomial feature of a given degree and
2:58:17then perform regression on it. So, this
2:58:19is basically polynomial regression.
2:58:22After this, we have support vector
2:58:23machine regression or we can also call
2:58:25it SVR. We identify a hyperplane with
2:58:28maximum margin such that the maximum
2:58:31number of data points are within those
2:58:33margins.
2:58:34It is also quite similar to the support
2:58:35vector machine classification algorithm.
2:58:38Then we have decision tree regression.
2:58:41A decision tree can be used for both
2:58:42regression and classification. But, in
2:58:44this case of regression, we use the ID3
2:58:47algorithm, which is iterative
2:58:48dichotomizer 3, to identify the
2:58:51splitting node by reducing the standard
2:58:53deviation.
2:58:54After this, we have a random forest
2:58:55regression, which is basically an
2:58:57ensemble of predictions of several
2:58:59decision tree regressions.
2:59:01So, this is all about the types of
2:59:02regressions for now. We're going to
2:59:04focus on simple linear regression.
2:59:06So, let's take a look at what exactly is
2:59:08a simple linear regression.
2:59:10Simple linear regression is a regression
2:59:12technique in which the independent
2:59:14variable has a linear relationship with
2:59:16the dependent variable.
2:59:18The straight line in the diagram is the
2:59:19best fit line, and the main goal of the
2:59:21simple linear regression is to consider
2:59:24the given data points and plot the best
2:59:25fit line to fit the model in the best
2:59:27way possible.
2:59:28So, if you talk about a real-life
2:59:30analogy to explain linear regression, we
2:59:32can take an example of a car resale
2:59:34value. So, we have different parameters,
2:59:36you know, when we are talking about
2:59:37resale value of a car. Like how many
2:59:40years the car has been there in the
2:59:41market, and how many kilometers it has
2:59:44been driven,
2:59:45the kind of mileage the car gives, and
2:59:47then we have different parameters we can
2:59:49focus upon. And all these independent
2:59:51variables somehow are linearly connected
2:59:53or interconnected to the price of the
Logistic Regression Algorithm
2:59:55car.
2:59:56So, that is one example to understand
2:59:58linear regression. We'll be doing that
2:59:59in the use case. I'll be telling you
3:00:01about how you can predict the price of
3:00:02car.
3:00:03Now, talking about linear regression
3:00:05terminologies, there are a few
3:00:07terminologies that you have to be
3:00:08thorough with to begin with linear
3:00:10regression.
3:00:11So, first of all, we have to talk about
3:00:13cost function.
3:00:14So, the best fit line can be based on
3:00:16the linear equation that is given here.
3:00:18So, in this, the dependent variable that
3:00:20is to be predicted is denoted by Y.
3:00:22A line that touches the Y axis is
3:00:24denoted by the intercept B0. The B1 is
3:00:27the slope of the line, and X represents
3:00:29the independent variables that determine
3:00:31the prediction of Y.
3:00:33The error in the resultant prediction is
3:00:35denoted by E.
3:00:36Now, talking about cost function, the
3:00:38cost function provides the best possible
3:00:40values for B0 and B1 to make the best
3:00:43fit line for the data points.
3:00:45We do this by converting this problem
3:00:46into a minimization problem to get the
3:00:49best values for B0 and B1.
3:00:51So, with this, the error is minimized in
3:00:53this problem between the actual value
3:00:55and the predicted value, and we choose
3:00:57the function above to minimize.
3:00:59Now, we square the error difference and
3:01:01sum the error over all the data points.
3:01:04The division between the total number of
3:01:05data points and the produced value
3:01:07provides the average square error for
3:01:09all the data points.
3:01:11It is also known as mean squared error,
3:01:13and we can change the values of B0 and
3:01:15B1 so that the MSE or the mean squared
3:01:17error value is settled at the minimum.
3:01:20So, this is one terminology that is cost
3:01:22function that we use in linear
3:01:23regression.
3:01:24Then, we have the gradient descent.
3:01:27So, the next important terminology to
3:01:28understand linear regression is gradient
3:01:30descent, of course, and it is a method
3:01:32of updating B0 and B1 value to reduce
3:01:35the MSE, which is the mean squared
3:01:36error.
3:01:37The idea behind this is to keep
3:01:39iterating the B0 and B1 values until we
3:01:41reduce the MSE to the minimum.
3:01:44Now, to update B0 and B1, we take the
3:01:45gradients from the cost function, and to
3:01:48find these gradients, we take partial
3:01:50derivatives with respect to B0 and B1.
3:01:53And these partial derivatives are the
3:01:55gradients and are used to update the
3:01:57values of B0 and B1.
3:01:59I'm sure guys, this is might be a little
3:02:00confusing for you guys if you are new to
3:02:03this, like gradient descent and cost
3:02:04function, but you don't have to worry
3:02:06about this because in Python when we're
3:02:08using linear regression, we're going to
3:02:09be using the scikit-learn or the
3:02:11scikit-learn library, so you don't have
3:02:12to worry about this. You just have to
3:02:14integrate your model with the linear
3:02:15regression model that we have already
3:02:17over there, and you'll be done with it.
3:02:19And when I'm implementing the linear
3:02:21regression model, you'll see how easy it
3:02:23is to actually implement linear
3:02:25regression in Python.
3:02:26So, after this, let's talk about a few
3:02:28advantages and disadvantages of linear
3:02:30regression.
3:02:32So, talking about the advantages first,
3:02:34linear regression performs exceptionally
3:02:36well for linearly separable data. And it
3:02:38is actually very easy to implement,
3:02:40interpret, and very efficient to train
3:02:43as well.
3:02:44And even though the linear regression is
3:02:45prone to overfitting, it handles it
3:02:48pretty well using dimension reduction
3:02:49techniques, regularization, and
3:02:51cross-validation. And one more advantage
3:02:54is that the extrapolation beyond a
3:02:56specific data set.
3:02:58So, these are all the advantages that we
3:02:59have with linear regression. Let's talk
3:03:01about a few disadvantages as well.
3:03:03So, one of the most common disadvantage
3:03:05with linear regression is that it takes
3:03:07the assumption of linearity between
3:03:09dependent and independent variables. The
3:03:11next disadvantage is it is often very
3:03:14prone to noise and overfitting as well,
3:03:16which is not a very good sign for any
3:03:18model if you're doing regression or
3:03:19classification in machine learning.
3:03:22The next disadvantage is it is very
3:03:24quite sensitive to outliers as well.
3:03:27And the last one is that it is very
3:03:28prone to multicollinearity.
3:03:30So, these are all the advantages and
3:03:32disadvantages of linear regression.
3:03:35>> [music]
3:03:39>> So, let's understand the what and why of
3:03:41logistic regression. Now, this algorithm
3:03:44is most widely used when the dependent
3:03:46variable, or you can say the output, is
3:03:47in the binary format. So, here you need
3:03:50to predict the outcome of a categorical
3:03:52dependent variable. So, the outcome
3:03:54should be always discrete or categorical
3:03:56in nature. Now, by discrete, I mean the
3:03:58value should be binary, or you can say
3:04:00you just have two values. It can either
3:04:02be zero or one. It can either be yes or
3:04:05a no. Either be true or false. Or high
3:04:07or low. So only these can be the
3:04:09outcomes. So the value which you need to
3:04:12predict should be discrete or you can
3:04:13say categorical in nature. Whereas in
3:04:16linear regression we have the value of Y
3:04:18or you can say the value you need to
3:04:19predict is in a range. So that is how
3:04:21there's a difference between linear
3:04:22regression and logistic regression. Now
3:04:24you must be having a question, why not
3:04:26linear regression? Now guys, in linear
3:04:28regression the value of Y or the value
3:04:30which you need to predict is in a range.
3:04:32But in our case, as in the logistic
3:04:34regression, we just have two values. It
3:04:36can be either zero or it can be one. It
3:04:39should not entertain the values which is
3:04:40below zero or above one. But in linear
3:04:43regression we have the value of Y in the
3:04:45range. So here, in order to implement
3:04:47logistic regression, we need to clip
3:04:48this part. So we don't need the value
3:04:51that is below zero or we don't need the
3:04:52value which is above one. So since the
3:04:54value of Y will be between only zero and
3:04:57one, that is the main rule of logistic
3:04:58regression, the linear line has to be
3:05:00clipped at zero and one. Now once we
3:05:02clip this graph, it would look somewhat
3:05:04like this. So here you're getting a
3:05:06curve which is nothing but three
3:05:07different straight lines. So here we
3:05:09need to make a new way to solve this
3:05:11problem. So this has to be formulated
3:05:13into equation and hence we come up with
3:05:15logistic regression. So here the outcome
3:05:17is either zero or one, which is the main
3:05:20rule of logistic regression. So with
3:05:21this our resulting curve cannot be
3:05:23formulated. So hence our main aim to
3:05:25bring the values to zero and one is
3:05:26fulfilled. So that is how we came up
3:05:28with logistic regression. Now here, once
3:05:31it gets formulated into an equation, it
3:05:33looks somewhat like this.
3:05:35So guys, this is nothing but a S curve
3:05:36or you can say the sigmoid curve or
3:05:38sigmoid function curve. So this sigmoid
3:05:41function basically converts any value
3:05:43from minus infinity to infinity to your
3:05:45discrete values which a logistic
3:05:47regression wants or you can say the
3:05:48values which are in binary format,
3:05:50either zero or one. So if you see here
3:05:53the values are either zero or one. And
3:05:55this is nothing but just a transition of
3:05:57it. But guys, there's a catch over here.
3:05:59So, let's say I have a data point that
3:06:01is 0.8. Now, how can you decide whether
3:06:04your value is zero or one? Now, here you
3:06:07have the concept of threshold, which
3:06:09basically divides your line. So, here
3:06:11threshold value basically indicates the
3:06:13probability of either winning or losing.
3:06:16So, here by winning I mean the values
3:06:18equals to one, and by losing I mean the
3:06:20values equals to zero. But how does it
3:06:22do that? Let's say I have data point
3:06:24which is over here. Let's say my cursor
3:06:26is at 0.8. So, here I'll check whether
3:06:28this value is less than my threshold
3:06:30value or not. Let's say if it is more
3:06:33than my threshold value, it should give
3:06:34me the result as one. If it is less than
3:06:36that, then it should give me the result
3:06:38as zero. So, here my threshold value is
3:06:400.5. Now, I need to define that if my
3:06:43value, let's say 0.8, it is more than
3:06:450.5, then the value shall be rounded off
3:06:48to one. And let's say if it is less than
3:06:500.5, let's say I have a value 0.2, then
3:06:52it should reduce it to zero. So, here
3:06:55you can use the concept of threshold
3:06:56value to find the output. So, here it
3:06:59should be discrete, it should be either
3:07:00zero or it should be one.
3:07:02So, I hope you caught this curve of
3:07:03logistic regression. So, the guys, this
3:07:05is the sigmoid S curve.
3:07:08So, to make this curve, we need to make
3:07:10an equation. So, let me address that
3:07:11part as well.
3:07:13So, let's see how an equation is formed
3:07:14to imitate this functionality. So, over
3:07:17here we have an equation of a straight
3:07:18line, which is Y is equals to MX + C.
3:07:21So, in this case, I just have only one
3:07:23independent variable. But let's say if
3:07:25we have many independent variable, then
3:07:27the equation becomes M1 X1 + M2 X2 + M3
3:07:30X3 and so on till MN XN. Now, let us put
3:07:34in B and X. So, here the equation
3:07:36becomes Y is equals to B1 X1 + B2 X2 +
3:07:39B3 X3 and so on till BN XN + C.
3:07:44So, guys, the equation of the straight
3:07:45line has a range from minus infinity to
3:07:47infinity. But in our case, or you can
3:07:50say in logistic equation, the value
3:07:52which we need to predict or you can say
3:07:53the Y value, it can have the range only
3:07:55from zero to one. So, in that case, we
3:07:57need to transform this equation. So, to
3:08:00do that, what we had done, we had just
3:08:02divide the equation by 1 - Y. So, now
3:08:04when Y is equals to zero, so zero over 1
3:08:07- 0 which is equals to 1. So, zero over
3:08:091 is again zero. And if we take Y is
3:08:12equals to 1, then 1 over 1 - 1 which is
3:08:15zero. So, 1 over zero is infinity. So,
3:08:17here my range is now between zero to
3:08:19infinity. But, again we want the range
3:08:21from minus infinity to infinity. So, for
3:08:24that, what we'll do, we'll have the log
3:08:25of this equation. So, let's go ahead and
3:08:27have the logarithmic of this equation.
3:08:29So, here we have just transform it
3:08:31further to get the range between minus
3:08:33infinity to infinity. So, over here we
3:08:35have log of Y over 1 - 1 and this is
3:08:38your final logistic regression equation.
3:08:40So, guys, don't worry, you don't have to
3:08:42write this formula or memorize this
3:08:44formula. In Python, you just need to
3:08:46call this function which is logistic
3:08:47regression and everything will be
3:08:49automatically for you. So, I don't want
3:08:51to scare you with the maths and the
3:08:52formulas behind it, but it's always good
3:08:54to know how the formula was generated.
3:08:57Moving ahead, let us see the various use
3:08:58cases wherein logistic regression is
3:09:00implemented in real life.
3:09:03So, the very first is weather
3:09:04prediction.
3:09:05Now, logistic regression helps you to
3:09:06predict your weather. For example, it is
3:09:09used to predict whether it is raining or
3:09:10not, whether it is sunny, is it cloudy
3:09:13or not. So, all these things can be
3:09:15predicted using logistic regression.
3:09:17Whereas, you need to keep in mind that
3:09:19both linear regression and logistic
3:09:20regression can be used in predicting
3:09:22weather. So, in that case, linear
3:09:24regression helps you to predict what
3:09:25will be the temperature tomorrow.
3:09:27Whereas, logistic regression will only
3:09:29tell you whether it's going to rain or
3:09:30not or whether it's cloudy or not,
3:09:32whether it's going to snow or not. So,
3:09:34these values are discrete. Whereas, if
3:09:36you apply linear regression, you're
3:09:37predicting things like what is the
3:09:39temperature tomorrow or what is the
3:09:41temperature day after tomorrow and all
3:09:43those things. So, these are the slight
3:09:44differences between linear regression
3:09:46and logistic regression. Now moving
3:09:47ahead, we have classification problem.
3:09:50So Python performs multi-class
3:09:51classification. So here it can help you
3:09:53tell whether it's a bird or it's not a
3:09:55bird. Then you classify different kind
3:09:57of mammals. Let's say whether it's a dog
3:09:59or it's not a dog. Similarly, you can
3:10:01check it for reptile whether it's a
3:10:03reptile or not a reptile. So in logistic
3:10:05regression, it can perform multi-class
3:10:07classification. So this point I've
3:10:09already discussed that it is used in
3:10:10classification problems. Next, it also
3:10:13helps you to determine the illness as
3:10:14well. So let me take an example. Let's
3:10:17say a patient goes for a routine checkup
3:10:19in hospital. So what doctor will do it
3:10:21it will perform various tests on the
3:10:22patient and will check whether the
3:10:24patient is actually ill or not. So what
3:10:26will be the features? So doctor can
3:10:29check the sugar level, the blood
3:10:30pressure, then what is the age of the
3:10:32patient? Is it very small or is it a old
3:10:34person? Then what is the previous
3:10:36medical history of the patient? And all
3:10:38of these features will be recorded by
3:10:40the doctor. And finally, doctor checks
3:10:42the patient data and determines the
3:10:44outcome of the illness and the severity
3:10:46of illness. So using all the data, a
3:10:48doctor can identify whether a patient is
3:10:51ill or not. So these are the various use
3:10:53cases in which you can use logistic
3:10:54regression. Now I guess enough of theory
3:10:57part, so let's move ahead and see some
3:10:59of the practical implementation of
3:11:00logistic regression.
3:11:02So over here I'll be implementing two
3:11:04projects wherein I have the data set of
3:11:06a Titanic. So over here we'll predict
3:11:08what factors made people more likely to
3:11:10survive the sinking of the Titanic ship.
3:11:12And in my second project, we'll see the
3:11:14data analysis on the SUV cars. So over
3:11:16here we have the data of the SUV cars,
3:11:18who can purchase it, and what factors
3:11:21made people more interested in buying
3:11:23SUV.
3:11:24So these will be the major questions as
3:11:25to why you should implement logistic
3:11:27regression and what output will you get
3:11:29by it. So let's start by the very first
3:11:31project that is Titanic data analysis.
3:11:33So some of you might know that there was
3:11:35a ship called as Titanic which basically
3:11:37hit an iceberg and it sank to the bottom
3:11:39of the ocean. And it was a big disaster
3:11:42at that time because it was the first
3:11:44voyage of the ship and it was supposed
3:11:45to be really, really strongly built and
3:11:47one of the best ships of that time. So,
3:11:49it was a big disaster of that time and
3:11:51of course there's a movie about this as
3:11:53well. So, many of you might have watched
3:11:55it. So, what we have we have data of the
3:11:57passengers, those who survived and those
3:11:59who did not survive in this particular
3:12:00tragedy. So, what you have to do you
3:12:02have to look at this data and analyze
3:12:04which factors would have been
3:12:05contributed the most to the chances of a
3:12:08person's survival on the ship or not.
3:12:10So, using the logistic regression we can
3:12:12predict whether the person survived or
3:12:14the person died. Now, apart from this we
3:12:16also have a look with the various
3:12:17features along with that. So, first let
3:12:19us explore the data set. So, over here
3:12:21we have the index value. Then the first
3:12:24column is passenger ID. Then my next
3:12:26column is survived. So, over here we
3:12:28have two values, a zero and a one. So,
3:12:31zero stands for did not survive and one
3:12:33stands for survived. So, this column is
3:12:35categorical where the values are
3:12:37discrete. Next we have passenger class.
3:12:39So, over here we have three values, one,
3:12:41two, and three. So, this basically tells
3:12:43you that whether a passenger is
3:12:45traveling in the first class, second
3:12:47class, or third class. Then we have the
3:12:49name of the passenger, we have the sex
3:12:51or you can say the gender of the
3:12:52passenger, whether passenger is a male
3:12:54or female. Then we have the age, we have
3:12:56the sib SP. So, this basically means the
3:12:59number of siblings or the spouses aboard
3:13:01the Titanic. So, over here we have
3:13:03values such as 1, 0, and so on. Then we
3:13:06have parch. So, parch is basically the
3:13:09number of parents or children aboard the
3:13:11Titanic. So, over here we also have some
3:13:13values.
3:13:14Then we have the ticket number, we have
3:13:16the fare, we have the cabin number, and
3:13:18we have the embarked column. So, in my
3:13:20embarked column we have three values, we
3:13:22have S, C, and Q. So, S basically stands
3:13:25for Southampton, C stands for Cherbourg,
3:13:27and Q stands for Queenstown.
3:13:30So, these are the features that we'll be
3:13:31applying our model on. So, here we'll
3:13:33perform various steps and then we'll be
3:13:35implementing logistic regression. So,
3:13:37now these are the various steps which
3:13:39are required to implement any algorithm.
3:13:41So now in our case we are implementing
3:13:43logistic regression. So very first step
3:13:45is to collect your data or to import the
3:13:47libraries that are used for collecting
3:13:49your data and then taking it forward.
3:13:51Then my second step is to analyze your
3:13:53data. So over here I can go through the
3:13:55various fields and then I can analyze
3:13:57the data. I can check did the females or
3:13:59children survive better than the males
3:14:01or did the rich passengers survive more
3:14:03than the poor passenger or did the money
3:14:05matter as in who paid more to get into
3:14:08the ship were they evacuated first and
3:14:10what about the workers? Does the worker
3:14:12survive or what is the survival rate if
3:14:15you were the worker in the ship and not
3:14:16just a traveling passenger? So all of
3:14:18these are very very interesting
3:14:20questions and you would be going through
3:14:21all of them one by one. So in this stage
3:14:24you need to analyze your data and
3:14:25explore your data as much as you can.
3:14:28Then my third step is to wrangle your
3:14:29data. Now data wrangling basically means
3:14:32cleaning your data. So over here you can
3:14:34simply remove the unnecessary items or
3:14:36if you have a null values in the data
3:14:38set you can just clear that data and
3:14:40then you can take it forward. So in this
3:14:42step you can build your model using the
3:14:44train data set and then you can test it
3:14:46using the test. So over here you will be
3:14:48performing a split which basically split
3:14:50your data set into training and testing
3:14:52data set and finally you will check the
3:14:54accuracy so as to ensure how much
3:14:56accurate your values are. So I hope you
3:14:58guys got these five steps that we're
3:15:00going to implement in logistic
3:15:01regression. So now let's go into all
3:15:03these steps in detail. So number one we
3:15:05have to collect your data or you can say
3:15:07import the libraries. So let me show you
3:15:09the implementation part as well. So I'll
3:15:11just open my Jupiter notebook and I'll
3:15:13just implement all of these steps side
3:15:15by side.
3:15:17So guys this is my Jupiter notebook. So
3:15:19first let me just rename Jupiter
3:15:21notebook to let's say Titanic data
3:15:23analysis.
3:15:27Now our first step was to import all the
3:15:29libraries and collect the data. So let
3:15:31me just import all the libraries first.
3:15:33So, first of all, I'll import pandas.
3:15:35So, pandas is used for data analysis.
3:15:38So, I'll say import pandas as pd. Then,
3:15:40I'll be importing NumPy. So, I'll say
3:15:42import NumPy as np. So, NumPy is a
3:15:45library in Python which basically stands
3:15:47for numerical Python. And it is widely
3:15:49used to perform any scientific
3:15:51computation. Next, we'll be importing
3:15:53seaborn. So, seaborn is a library for
3:15:55statistical plotting. So, I'll say
3:15:57import seaborn as sns. I'll also import
3:16:00matplotlib.
3:16:01So, matplotlib library is again for
3:16:03plotting. So, I'll say import
3:16:05matplotlib.pyplot
3:16:07as pld.
3:16:09Now, to run this library in Jupyter
3:16:10Notebook, all I have to write in is
3:16:12percentage matplotlib inline.
3:16:15Next, I'll be importing one module as
3:16:17well. So, as to calculate the basic
3:16:20mathematical functions. So, I'll say
3:16:22import maths. So, these are the
3:16:23libraries that I'll be needing in this
3:16:25Titanic data analysis. So, now let me
3:16:27just import my dataset. So, I'll take a
3:16:29variable, let's say Titanic data. And
3:16:32using the pandas, I will just read my
3:16:34CSV. Or you can say the dataset.
3:16:37I'll write the name of my dataset, that
3:16:38is titanic.csv.
3:16:40Now, I have already showed you the
3:16:42dataset. So, over here, let me just
3:16:43print the top 10 rows. So, for that,
3:16:45I'll just say I'll take the variable
3:16:47Titanic data. head and I'll say the top
3:16:5010 rows. Now, I'll just run this. So, to
3:16:52run this, I just have to press shift
3:16:54plus enter. Or else, you can just
3:16:56directly click on the cell.
3:16:58So, over here, I have the index. We have
3:17:00the passenger ID, which is nothing but
3:17:02again the index which is starting from
3:17:03one. Then, we have the survived column
3:17:05which has the categorical values or you
3:17:07can say the discrete values, which is in
3:17:09the form of zero or one. Then, we have
3:17:11the passenger class. We have the name of
3:17:13the passenger, sex, age, and so on. So,
3:17:15this is the dataset that I'll be going
3:17:17forward with. Next, let us print the
3:17:18number of passengers which are there in
3:17:20this original dataset. So, for that,
3:17:22I'll just simply type in print. I'll say
3:17:25number of passengers.
3:17:31And using the length function, I can
3:17:32calculate the total length. So, I'll say
3:17:34length and inside this I'll be passing
3:17:36this variable which is Titanic data. So,
3:17:38I'll just copy it from here. I'll just
3:17:40paste it {dot} index.
3:17:42And next, let me just print this one.
3:17:45So, here the number of passengers which
3:17:46are there in the original data set we
3:17:48have is 891. So, around this number were
3:17:51traveling in the Titanic ship. So, over
3:17:54here my first step is done. We have just
3:17:56collected data, imported all the
3:17:57libraries, and find out the total number
3:17:59of passengers which are traveling in
3:18:01Titanic. So, let me just go back to
3:18:03presentation and let's see what is my
3:18:04next step.
3:18:05So, we're done with the collecting data.
3:18:07Next step is to analyze your data. So,
3:18:09over here we'll be creating different
3:18:11plots to check the relationship between
3:18:13variables as in how one variable is
3:18:15affecting the other. So, you can simply
3:18:17explore your data set by making use of
3:18:19various columns and then you can plot a
3:18:21graph between them. So, you can either
3:18:23plot a correlation graph, you can plot a
3:18:25distribution graph. It's up to you guys.
3:18:27So, let me just go back to my Jupiter
3:18:29notebook and let me analyze some of the
3:18:30data. Over here my second part is to
3:18:32analyze data. So, I'll just put this in
3:18:34header two.
3:18:36Now, to put this in header two, I just
3:18:37have to go on code, click on markdown,
3:18:39and I'll just run this.
3:18:41So, first let us plot a count plot where
3:18:43you can compare between the passengers
3:18:44who survived and who did not survive.
3:18:46So, for that I'll be using the seaborn
3:18:47library. So, over here I have imported
3:18:50seaborn as sns. So, I don't have to
3:18:52write the whole name. I'll simply say
3:18:53sns.count plot.
3:18:58I'll say x is equal to survive and the
3:19:00data that I'll be using is the Titanic
3:19:01data. Or you can say the name of
3:19:03variable in which you have stored your
3:19:04data set. So, now let me just run this.
3:19:07So, over here as you can see I have
3:19:09survived column on my x-axis and on the
3:19:11y-axis I have the count. So, zero
3:19:13basically stands for did not survive and
3:19:15one stands for the passengers who did
3:19:17survive. So, over here you can see that
3:19:19around 550 of the passengers who did not
3:19:22survive and there were around 350
3:19:24passengers who only survived. So here
3:19:26you can basically conclude that there
3:19:28are very less survivors than
3:19:29non-survivors. So this was the very
3:19:32first plot. Now let us plot another plot
3:19:34to compare the sex as to whether out of
3:19:36all the passengers who survived and who
3:19:38did not survive, how many were men and
3:19:40how many were female. So to do that I'll
3:19:42simply say sns.countplot.
3:19:47I'll add the hue as sex.
3:19:49So I want to know how many females and
3:19:51how many males survived.
3:19:53Then I'll be specifying the data. So I'm
3:19:54using Titanic data set.
3:19:57And let me just run this.
3:19:59Okay, I've done a mistake over here.
3:20:01So over here you can see I have survived
3:20:02column on the x-axis and I have the
3:20:04count on the y.
3:20:06Now so here your blue color stands for
3:20:07your male passengers and orange stands
3:20:09for your female.
3:20:11So as you can see here the passengers
3:20:13who did not survive, that has a value
3:20:14zero.
3:20:15So we can see that majority of males did
3:20:18not survive. And if we see the people
3:20:20who survived, here we can see the
3:20:22majority of females survived. So this
3:20:24basically concludes the gender of the
3:20:25survival rate. So it appears on average
3:20:28women were more than three times more
3:20:30likely to survive than men. Next let us
3:20:32plot another plot where we have the hue
3:20:34as the passenger class. So over here we
3:20:36can see which class that the passenger
3:20:38was traveling in, whether it was
3:20:39traveling in class one, two or three.
3:20:42So for that I'll just write the same
3:20:44command. I'll say
3:20:45sns.countplot.
3:20:49I'll keep my x-axis as only. I'll change
3:20:52my hue to passenger class.
3:20:54So my variable named as pclass.
3:20:57And the data set that I'll be using is
3:20:58Titanic data. So this is my result. So
3:21:01over here you can see I have blue for
3:21:03first class, orange for second class and
3:21:05green for the third class.
3:21:07So here the passengers who did not
3:21:09survive were majorly of the third class
3:21:11or you can see the lowest class or the
3:21:12cheapest class to get into the Titanic.
3:21:15And the people who did survive majorly
3:21:16belong to the higher classes. So here
3:21:18one and two has more rise than the
3:21:20passenger who were traveling in the
3:21:21third class. So here we have concluded
3:21:24that the passengers who did not survive
3:21:26are majorly of third class or you can
3:21:27say the lowest class. And the passengers
3:21:30who were traveling in first and second
3:21:31class would tend to survive more. Next
3:21:34let us plot a graph for the age
3:21:35distribution. Over here I can simply use
3:21:37my data. So we'll be using pandas
3:21:39library for this. I'll declare a array
3:21:42and I'll pass in the column that is age.
3:21:45So I plot and I want a histogram so I'll
3:21:47say plot.hist.
3:21:51So you can notice over here that we have
3:21:53more of young passengers or you can see
3:21:55the children between the ages zero to
3:21:5710. And then we have the average age
3:21:59people. And if you go ahead lesser would
3:22:01be the population. So this is the
3:22:03analysis on the age column. So we saw
3:22:06that we have more young passengers and
3:22:08more mediocre age passengers who are
3:22:10traveling in the Titanic.
3:22:11So next let me plot a graph of fare as
3:22:13well. So I'll say Titanic data.
3:22:17I'll say fare.
3:22:18And again I'll plot a histogram so I'll
3:22:20say hist.
3:22:23So here you can see the fare size is
3:22:25between zero to 100. Now let me add the
3:22:27bin size so as to make it more clear.
3:22:30So over here I'll say bin is equals to
3:22:32let's say 20 and I'll increase the
3:22:34figure size as well. So I'll say fix
3:22:36size. Let's say I'll give the dimensions
3:22:39as 10 by 5.
3:22:41So it is bins. So this is more clear
3:22:43now. Next let us analyze the other
3:22:45columns as well.
3:22:47So I'll just type in Titanic data. And I
3:22:50want the information as to what all
3:22:51columns are left.
3:22:54So here we have passenger ID which I
3:22:56guess it's of no use. Then we have to
3:22:58see how many passengers survived and how
3:23:00many did not. We also see the analysis
3:23:02on the gender basis. We saw whether
3:23:03female tend to survive more or the men
3:23:05tend to survive more. Then we saw the
3:23:07passenger class where the passenger is
3:23:09traveling in in first class, second
3:23:10class or third class. Then we have the
3:23:12name, so in name we cannot do any
3:23:14analysis. We saw the sex, we saw the age
3:23:17as well. Then we have sibsp. So this
3:23:20stands for the number of siblings or the
3:23:22spouses which are aboard the Titanic. So
3:23:24let us do this as well. So I'll say
3:23:26sns.countplot.
3:23:30I'll mention x as sibsp.
3:23:33And I'll be using the Titanic data.
3:23:36So you can see the plot over here. So
3:23:38over here you can conclude that it has
3:23:40the maximum value on zero. So you can
3:23:42conclude that neither a children nor a
3:23:44spouse was on board the Titanic. The
3:23:46second most highest value is one. And
3:23:49then we have very less values for two,
3:23:51three, four, and so on.
3:23:53Next, if I go above, we saw this column
3:23:55as well. Similarly, you can do for
3:23:56parch.
3:23:57So next we have parch, or you can say
3:23:59the number of parents or children which
3:24:00were aboard the Titanic. So you
3:24:02similarly can do this as well. Then we
3:24:04have the ticket number. So I don't think
3:24:06so any analysis is required for ticket.
3:24:08Then we have fare. So fare we have
3:24:10already discussed as in the people who
3:24:12tend to travel in the first class
3:24:13usually pay the highest fare. Then we
3:24:15have the cabin number, and we have
3:24:17embarked. So these are the columns that
3:24:19we'll be doing data wrangling on.
3:24:21So we have analyzed the data, and we
3:24:22have seen quite a few graphs in which we
3:24:24can conclude which variable is better
3:24:27than the another or or what are the
3:24:28relationship they hold.
3:24:29So third step is my data wrangling. So
3:24:31data wrangling basically means cleaning
3:24:33your data. So if you have a large data
3:24:36set, you might be having some null
3:24:37values, or you can say NaN values. So
3:24:39it's very important that you remove all
3:24:41the unnecessary items that are present
3:24:43in your data set. So removing this
3:24:45directly affects your accuracy. So I'll
3:24:47just go ahead and clean my data by
3:24:49removing all the NaN values and
3:24:51unnecessary columns which has a null
3:24:52value in the data set. So next I'll be
3:24:54performing data wrangling.
3:25:02So first of all, I'll check whether my
3:25:03data set is null or not. So I'll say
3:25:06Titanic data, which is the name of my
3:25:07data set, and I'll say is null. So, this
3:25:10will basically tell me what all values
3:25:12are null, and it will return me a
3:25:13Boolean result. So, this basically
3:25:15checks the missing data, and your result
3:25:17will be in Boolean format, as in the
3:25:19result will be in true or false. So,
3:25:20false mean if it is not null, and true
3:25:23means if it is null. So, let me just run
3:25:25this.
3:25:27Over here, you can see the values as
3:25:29false or true. So, false is where the
3:25:31value is not null, and true is where the
3:25:34value is null. So, over here, you can
3:25:35see in the cabin column, we have the
3:25:37very first value, which is null. So, we
3:25:39have to do something on this.
3:25:41So, you can see that we have a large
3:25:43data set.
3:25:44The counting does not stop, and we can
3:25:47actually see the sum of it. We can
3:25:48actually print the number of passengers
3:25:50who have the NaN value in each column.
3:25:52So, I'll say Titanic_data
3:25:55is null, and I want the sum of it. So,
3:25:57I'll say dot sum. So, this will
3:25:59basically print the number of passengers
3:26:01who have the NaN values in each column.
3:26:03So, we can see that we have missing
3:26:05values in age column, that is 177. Then,
3:26:07we have the maximum value in the cabin
3:26:09column, and we have very less in the
3:26:11embarked column, that is two.
3:26:13So, here, if you don't want to see these
3:26:15numbers, you can also plot a heat map,
3:26:17and then you can visually analyze it.
3:26:19So, let me just do that as well. So,
3:26:20I'll say sns.heatmap
3:26:26and say yticklabels
3:26:31to false. So, I'll just run this. So, as
3:26:33we have already seen that there were
3:26:35three columns in which missing data
3:26:36value was present. So, this might be
3:26:38age. So, over here, almost 20% of age
3:26:41column has a missing value. Then, we
3:26:43have the cabin columns. So, this is
3:26:44quite a large value, and then we have
3:26:46two values for embarked column as well.
3:26:49Add a cmap for color coding. So, I'll
3:26:51say cmap
3:26:55So, if I do this, so the graph becomes
3:26:57more attractive. So, over here, your
3:26:59yellow stands for true, or you can say
3:27:01the values are null.
3:27:03So, here we have concluded that we have
3:27:05the missing value of age. We have a lot
3:27:07of missing values in the cabin column
3:27:09and we have very less value, which is
3:27:11not even visible in the embark column as
3:27:13well.
3:27:14So, to remove these missing values, you
3:27:16can either replace the values and you
3:27:18can put in some dummy values to it or
3:27:20you can simply drop the column.
3:27:22So, here let us first pick the age
3:27:24column. So, first let me just plot a box
3:27:26plot and then we analyze with having a
3:27:28column as age. So, I'll say sns.
3:27:31boxplot
3:27:33I'll say x is equals to passenger class.
3:27:36So, it's pclass. I'll say y is equals to
3:27:38age.
3:27:39And the data set that I'll be using is
3:27:41Titanic set. So, I'll say data is equals
3:27:43to Titanic data.
3:27:45You can see the age in first class and
3:27:47second class tends to be more older
3:27:49rather than we have it in the third
3:27:50class. Well, that depends on the
3:27:52experience, how much you earn or might
3:27:54be there any number of reasons.
3:27:56So, here we concluded that passengers
3:27:57who were traveling in class one and
3:27:59class two attend to be older than what
3:28:01we have in the class three.
3:28:03So, we have found that we have some
3:28:05missing values in M.
3:28:06Now, one way is to either just drop the
3:28:08column or you can just simply fill in
3:28:10some values to that. So, this method is
3:28:12called as imputation.
3:28:14Now, to perform data wrangling or
3:28:16cleaning, let us first print the head of
3:28:17the data set. So, I'll say Titanic.head.
3:28:20Sorry, it's Titanic_data.
3:28:23Let's say I just want the five rows.
3:28:25So, here we have survived, which is
3:28:27again categorical. So, in this
3:28:28particular column, I can apply logistic
3:28:30regression. So, this can be my y value
3:28:33or the value that you need to predict.
3:28:35Then, we have the passenger class. We
3:28:36have the name. Then, we have ticket
3:28:38number, fare, cabin. So, over here we
3:28:41have seen that in cabin we have a lot of
3:28:43null values or you can say the NaN
3:28:44values, which is quite visible as well.
3:28:46So, first of all, we'll just drop this
3:28:48column. So, for dropping it, I'll just
3:28:50say Titanic_data.
3:28:52And I'll simply type in drop and the
3:28:53column which I need to drop. So, I have
3:28:56to drop the cabin column.
3:28:58I mention the axis equals to one and
3:29:00I'll say in place also to true.
3:29:04So, now again I'll just print the head
3:29:05and let us see whether this column has
3:29:07been removed from the data set or not.
3:29:09So, I'll say Titanic .head.
3:29:12So, as you can see here we don't have
3:29:14cabin column anymore.
3:29:16Now, you can also drop the NA values.
3:29:18So, I'll say Titanic data
3:29:20.drop all the NA values or you can say
3:29:22NaN which is not a number and I'll say
3:29:25in place is equals to true.
3:29:27Let's Titanic
3:29:29So, over here let me again plot the heat
3:29:31map and let's see all the values which
3:29:33were before showing a lot of null values
3:29:35has it been removed or not. So, I'll say
3:29:37sns.heatmap I'll pass in the data set.
3:29:41I'll check if this null.
3:29:43I'll say white labels is equals to
3:29:45false.
3:29:47And I don't want color coding, so again
3:29:49I'll say false.
3:29:51So, this will basically help me to check
3:29:53whether my values has been removed from
3:29:55the data set or not. So, as you can see
3:29:56here I don't have any null values. So,
3:29:59it's entirely black.
3:30:01Now, you can actually know the sum as
3:30:02well, so I'll just go above.
3:30:05So, I'll just copy this part and I'll
3:30:07just use the sum function to calculate
3:30:09the sum.
3:30:10So, here that tells me the data set is
3:30:12clean as in the data set does not
3:30:14contain any null value or any NaN value.
3:30:18So, now we have wrangled our data or you
3:30:20can say clean our data.
3:30:21So, here we have done just one step in
3:30:23data wrangling that is just removing one
3:30:25column out of it. Now, you can do a lot
3:30:27of things. You can actually fill in the
3:30:28values with some other values or you can
3:30:31just calculate the mean and then you can
3:30:32just fit in the null values.
3:30:34But, now if I see my data set So, I'll
3:30:37say Titanic data.head.
3:30:39But, now if I see over here I have a lot
3:30:41of string values. So, this has to be
3:30:43converted to a categorical variables in
3:30:45order to implement logistic regression.
3:30:47So, what we will do we will convert this
3:30:49to categorical variable into some dummy
3:30:51variables and this can be done using
3:30:53pandas because logistic regression just
3:30:55take two values.
3:30:57So whenever you apply machine learning,
3:30:58you need to make sure that there are no
3:31:00string values present because it won't
3:31:02be taking these as your input variables.
3:31:05So using string, you don't have to
3:31:06predict anything. But in my case, I have
3:31:08the survived columns, so I need to
3:31:09predict how many people tend to survive
3:31:11and how many did not. So zero stands for
3:31:13did not survive and one stands for
3:31:15survive.
3:31:16So now let me just convert these
3:31:17variables into dummy variables.
3:31:20So let's use pandas and I'll say
3:31:22pd.get_dummies.
3:31:24You can simply press tab to auto
3:31:26complete. I'll say Titanic data.
3:31:29And I'll pass the sex.
3:31:30So you can just simply click on shift
3:31:32plus tab to get more information on
3:31:34this.
3:31:35So here we have the type data frame and
3:31:37we have the passenger ID survived and
3:31:39passenger class.
3:31:40So if you run this, you'll see that zero
3:31:42basically stands for not a female and
3:31:44one stand for it is a female. Similarly
3:31:46for male, zero stands for it's not male
3:31:48and one stand for it's male. Now we
3:31:50don't require both these columns because
3:31:52one column itself is enough to tell us
3:31:55whether it's male or you can say female
3:31:56or not. So let's say if I want to keep
3:31:58only male, I'll say if the value of male
3:32:01is one, so it is definitely a male and
3:32:03it is not a female. So that is how it
3:32:05you don't need both of these values. So
3:32:07for that, I'll just remove the first
3:32:09column, let's say a female.
3:32:10So I'll say drop first
3:32:13and true.
3:32:15So over here it has given me just one
3:32:17column which is male and has the value
3:32:19zero and one. Let me just set this as a
3:32:22variable, let's say sex. So over here I
3:32:24can say sex.head.
3:32:26I just want to see the first five rows.
3:32:29Sorry, it's dot.
3:32:31So this is how my data looks like.
3:32:34Now here we have done it for sex, then
3:32:35we have the numerical values in age, we
3:32:37have the numerical values in spouses,
3:32:39then we have the ticket number, we have
3:32:41the fare and we have embarked as well.
3:32:43So in embarked, the values are in S, C
3:32:45and Q. So here we can apply this get
3:32:48dummy function.
3:32:49So, let's say I'll take a variable,
3:32:51let's say embarked.
3:32:53I'll use the pandas library.
3:32:57I'll enter the column name, that is
3:32:59embarked.
3:33:03So, let me just print the head of it.
3:33:04So, I'll say embarked.head.
3:33:07So, over here we have C, Q, and S. Now,
3:33:09here also we can drop the first column
3:33:11because these two values are enough
3:33:13whether the passenger is either
3:33:15traveling from Q, that is Queenstown, S
3:33:17for Southampton. And if both the values
3:33:18are zero, then definitely the passenger
3:33:20is from Cherbourg, that is the third
3:33:22value. So, you can again drop the first
3:33:24value. So, I'll say drop
3:33:27and true.
3:33:28Let me just run this. So, this is how my
3:33:29output looks like. Now, similarly you
3:33:31can do it for passenger class as well.
3:33:33So, here also we have three classes,
3:33:35one, two, and three.
3:33:36So, I'll just copy the whole statement.
3:33:42So, let's say I want the variable name,
3:33:44let's say PCL.
3:33:47I'll pass in the column name, that is P
3:33:48class, and I'll just drop the first
3:33:51column. So, here also the values would
3:33:53be one, two, or three, and I'll just
3:33:55remove the first column. So, here we
3:33:57just left with two and three. So, if
3:33:59both the values are zero, then
3:34:00definitely the passenger is traveling in
3:34:02the first class. Now, we have made the
3:34:04values as categorical. Now, my next step
3:34:06would be to concatenate all these new
3:34:09rows into a data set.
3:34:11I can say Titanic data. Using the
3:34:13pandas, we'll just concatenate all these
3:34:15columns. So, I'll say pd.concat.
3:34:18And I'll say we have to concatenate sex,
3:34:21we have to concatenate embarked and PCL.
3:34:24And then I'll mention the axis to one.
3:34:26I'll just run this.
3:34:28Okay, I need to print the head.
3:34:30So, over here you can see that these
3:34:32columns have been added over here. So,
3:34:34we have the male column which basically
3:34:36tells whether a person is male or it's a
3:34:38female. Then we have the embarked which
3:34:40is basically Q and S. So, if it's
3:34:43traveling from Queenstown, value would
3:34:44be one, else it would be zero. And if
3:34:46both of these values are zero, it is
3:34:48definitely traveling from Cherbourg.
3:34:50Then we have the passenger class as two
3:34:52and three. So, if the value of both
3:34:54these is zero, then the passenger is
3:34:56traveling in class one.
3:34:58So, I hope you got this till now.
3:35:00Now, these are the irrelevant columns
3:35:02that we have it over here. So, we can
3:35:04just drop these columns. We're dropping
3:35:05P class,
3:35:07the embarked column,
3:35:09and the sex column.
3:35:10So, I'll just type in Titanic data
3:35:13{dot} drop. I'll mention the columns
3:35:15that I want to drop. So, I'll say
3:35:22I'll even delete the passenger ID
3:35:24because it's nothing but just the index
3:35:25value, which is starting from one.
3:35:27So, I'll drop this as well.
3:35:29Then I don't want name as well, so I'll
3:35:31delete name as well.
3:35:32Then what else we can drop? We can drop
3:35:34the ticket as well.
3:35:37And then I'll just mention the axis.
3:35:40I'll say in place is equals to true.
3:35:43Okay, so the my column name starts from
3:35:45upper case.
3:35:47So, these have been dropped. Now, let me
3:35:48just print my data set again.
3:35:51So, this is my final data set, guys. We
3:35:52have the survived column, which has the
3:35:54value zero and one. Then we have the
3:35:56passenger class. Oh, we forgot to drop
3:35:58this as well. So, no worries. I'll drop
3:36:00this again.
3:36:05So, now let me just run this.
3:36:08So, over here we have the survived, we
3:36:10have the age, we have the sibsp, we have
3:36:12the parch, we have fare, male, and these
3:36:15we have just converted.
3:36:17So, here we have just performed data
3:36:18wrangling, or you can say clean the
3:36:20data. And then we have just converted
3:36:22the values of gender to male, then
3:36:25embarked to Q and S, and the passenger
3:36:27class to two and three. So, this was all
3:36:29about my data wrangling, or just
3:36:30cleaning the data.
3:36:32Then my next step is training and
3:36:33testing your data. So, here we will
3:36:35split the data set into train subset and
3:36:37test subset. And then what we'll do
3:36:39we'll build a model on the train data
3:36:41and then predict the output on your test
3:36:43data set. So let me just go back to
3:36:45Jupiter and let us implement this as
3:36:46well.
3:36:47Over here I need to train my data set.
3:36:49So I'll just put this in date heading
3:36:51three.
3:36:54So over here you need to define your
3:36:55dependent variable and independent
3:36:57variable.
3:36:58So here my Y is the output or you can
3:37:00say the value that I need to predict.
3:37:03So over here I'll write Titanic data.
3:37:06I'll take the column which is survive.
3:37:09So basically I have to predict this
3:37:11column whether the passenger survived or
3:37:12not. And as you can see we have the
3:37:14discrete outcome which is in the form of
3:37:16zero and one. And rest all the things we
3:37:19can take it as a features or you can say
3:37:21independent variable. So I'll say
3:37:22Titanic data
3:37:24.drop.
3:37:26So we'll just simply drop the survive
3:37:28and all the other columns will be my
3:37:29independent variable.
3:37:31So everything else are the features
3:37:32which leads to the survival rate. So
3:37:34once we have defined the independent
3:37:36variable and the dependent variable,
3:37:38next step is to split your data into
3:37:39training and testing subset. So for that
3:37:42we'll be using SK learn. I'll just type
3:37:44in from SK learn.cross_validation
3:37:48import train_test_split.
3:37:54Now here if you just click on shift and
3:37:56tab, you can go to the documentation and
3:37:59you can just see the examples over here.
3:38:02I'll click on plus to open it. And then
3:38:04I'll just go to examples and see how you
3:38:06can split your data. So over here you
3:38:08have X_train, X_test, Y_train, Y_test.
3:38:12And then using this train_test_split you
3:38:13can just pass in your independent
3:38:15variable and dependent variable and just
3:38:17define a size and a random state to it.
3:38:19So let me just copy this.
3:38:21And I'll just paste it over here.
3:38:24Over here we'll train_test. Then we have
3:38:27the dependent variable train and test.
3:38:29And using the split function we'll pass
3:38:30in the independent and dependent
3:38:32variable and then we'll set a split
3:38:34size. So, let's say I'll put it at 0.3.
3:38:37So, this basically means that your data
3:38:38set is divided in 0.3. That is in 70/30
3:38:41ratio. And then I can add any random
3:38:43state to it. So, let's say I'm applying
3:38:45one. This is not necessary. If you want
3:38:48the same result as that of mine, you can
3:38:49add the random state. So, this will
3:38:51basically take exactly the same sample
3:38:53every time.
3:38:55Next, I have to train and predict by
3:38:57creating a model. So, here logistic
3:38:59regression will grab from the linear
3:39:01regression. So, next I'll just type in
3:39:03from sklearn.linear_model
3:39:07import LogisticRegression.
3:39:09Next, I'll just create the instance of
3:39:11this logistic regression model. So, I'll
3:39:13say log model
3:39:15is equals to LogisticRegression.
3:39:17Now, I just need to fit my model. So,
3:39:19I'll say log model.fit and I'll just
3:39:22pass in my X_train
3:39:25and Y_train.
3:39:28All right. So, here it gives me all the
3:39:30details of logistic regression.
3:39:32So, here it gives me the class weight,
3:39:34dual, fit intercept and all those
3:39:35things.
3:39:36Then, what I need to do, I need to make
3:39:38prediction. So, I'll take a variable,
3:39:40let's say predictions, and I'll pass on
3:39:42the model to it. So, I'll say log
3:39:44model.predict
3:39:46and I'll pass in the value that is
3:39:48X_test. So, here we have just created a
3:39:50model, fit that model, and then we have
3:39:52made predictions. So, now to evaluate
3:39:54how my model has been performing, so you
3:39:56can simply calculate the accuracy or you
3:39:58can also calculate the classification
3:40:00report. So, don't worry, guys. I'll be
3:40:02showing both of these methods. So, I'll
3:40:04say from sklearn.metrics
3:40:08import classification_report.
3:40:11So, over here I'll use
3:40:12classification_report and inside this
3:40:14I'll be passing in Y_test and the
3:40:16predictions.
3:40:21So, guys, this is my classification
3:40:22report. So, over here I have the
3:40:24precision, I have the recall, we have
3:40:26the F1 score, and then we have support.
3:40:29So, here we have the value of precision
3:40:31as 75, 72, and 73, which is not that
3:40:34bad. Now, in order to calculate the
3:40:36accuracy as well, you can also use the
3:40:38concept of confusion matrix. So, if you
3:40:40want to print the confusion matrix, I'll
3:40:42simply say from SK learn {dot} metrics
3:40:46import confusion matrix first of all,
3:40:48and then we'll just print this.
3:40:51So, here my function has been imported
3:40:52successfully, so I'll say confusion
3:40:54matrix.
3:40:56And I'll again pass in the same
3:40:57variables, which is Y test and
3:40:59predictions.
3:41:01So, I hope you guys already know the
3:41:03concept of confusion matrix. So, can you
3:41:05guys give me a quick confirmation as to
3:41:07whether you guys remember this confusion
3:41:09matrix concept or not? So, if not, I can
3:41:11just quickly summarize this as well.
3:41:14Okay, Jagriti says a yes.
3:41:16Okay, Swati is not clear with this. So,
3:41:18I'll just tell you in a brief what
3:41:19confusion matrix is all about.
3:41:22So, confusion matrix is nothing but a 2
3:41:24by 2 matrix, which has a four outcomes.
3:41:27This basically tells us that how
3:41:28accurate your values are. So, here we
3:41:30have the column as predicted no,
3:41:32predicted yes,
3:41:34and we have actual no and an actual yes.
3:41:38So, this is the concept of confusion
3:41:40matrix. So, here let me just feed in
3:41:42these values which we have just
3:41:43calculated. So, here we have 105,
3:41:48105, 21, 25, and 63.
3:41:53So, as you can see here, we have got
3:41:55four outcomes. Now, 105 is the value
3:41:58where our model has predicted no, and in
3:42:00reality it was also a no. So, here we
3:42:03have predicted no and an actual no.
3:42:05Similarly, we have 63 as a predicted
3:42:07yes, so here the model predicted yes,
3:42:09and actually also it was a yes. So, in
3:42:12order to calculate the accuracy, you
3:42:13just need to add the sum of these two
3:42:15values and just divide the whole by the
3:42:18sum. So, here these two values tells me
3:42:20where the model has actually predicted
3:42:22the correct output. So, this value is
3:42:24also called as true negative. This is
3:42:26called as false positive. This is called
3:42:28as true positive and this is called as
3:42:30false negative. Now, in order to
3:42:31calculate the accuracy, you don't have
3:42:33to do it manually. So, in Python, you
3:42:35can just import accuracy score function
3:42:38and you can get the results from that.
3:42:39So, I'll just do that as well. So, I'll
3:42:41say from sklearn.metrics
3:42:44import accuracy score
3:42:46and I'll simply print the accuracy.
3:42:49I'll pass in the same variables, that is
3:42:51Y test and predictions.
3:42:53So, over here it tells me the accuracy
3:42:54as 78, which is quite good. So, over
3:42:57here if you want to do it manually, you
3:42:59have to plus these two numbers, which is
3:43:01105 + 63. So, this comes out to almost
3:43:04168.
3:43:06And then you have to divide it by the
3:43:07sum of all the four numbers. So, 105 +
3:43:1063 + 21 + 25. So, this gives me a result
3:43:13of 214. So, now if you divide these two
3:43:16number, you'll get the same accuracy
3:43:18that is 78% or you can say 0.78.
3:43:21So, that is how you can calculate the
3:43:23accuracy. So, now let me just go back to
3:43:25my presentation and let's see what all
3:43:27we have covered till now.
3:43:28So, here we have first split our data
3:43:30into train and test subset. Then we have
3:43:32built our model on the train data and
3:43:34then predicted the output on the test
3:43:36data set. And then my fifth step is to
3:43:38check the accuracy. So, here we have
3:43:40calculated accuracy to almost 78%, which
3:43:43is quite good. You cannot say that
3:43:45accuracy is bad.
3:43:46So, here it tells me how accurate your
3:43:48results are. So, here my accuracy score
3:43:50defines that and hence we got a good
3:43:52accuracy.
3:43:54So, now moving ahead, let us see the
3:43:55second project, that is SUV data
3:43:57analysis.
3:43:58So, in this a car company has released
3:44:00new SUV in the market. And using the
3:44:03previous data about the sales of their
3:44:05SUV, they want to predict the category
3:44:07of people who might be interested in
3:44:08buying this. So, using the logistic
3:44:10regression, you need to find what
3:44:12factors make people more interested in
3:44:14buying this SUV.
3:44:15So, for this let us see our data set
3:44:17where I have user ID, I have gender as
3:44:19male and female.
3:44:21Then we have the age, we have the
3:44:22estimated salary.
3:44:24And then we have the purchased column.
3:44:26So this is my discrete column, or you
3:44:28can say the categorical column. So here
3:44:30we just have the value that is zero and
3:44:31one. And this column we need to predict
3:44:33whether a person can actually purchase a
3:44:35SUV or not. So based on these factors,
3:44:38we will be deciding whether a person can
3:44:40actually purchase a SUV or not.
3:44:42So we know the salary of a person, we
3:44:43know the age. And using these, we can
3:44:46predict whether a person can actually
3:44:47purchase a SUV or not. So let me just go
3:44:49to my Jupiter notebook and let's
3:44:51implement logistic regression. So guys,
3:44:53I will not be going through all the
3:44:54details of data cleaning and analyzing
3:44:56the part. So that part, I'll just leave
3:44:58it on you. So just go ahead and practice
3:45:00as much as you can.
3:45:02All right. So my second project is SUV
3:45:04predictions.
3:45:07All right. So first of all, I have to
3:45:08import all the libraries. So I say
3:45:10import NumPy as NP.
3:45:13And similarly, I'll do the rest of it.
3:45:21All right.
3:45:22So now let me just print the head of
3:45:24this data set.
3:45:26So this we have already seen that we
3:45:27have columns as user ID, we have gender,
3:45:30we have the age, we have the salary, and
3:45:32then we have to calculate whether a
3:45:33person can actually purchase a SUV or
3:45:35not.
3:45:36So now let us just simply go on to the
3:45:38algorithm part. So we'll directly start
3:45:40off with the logistic regression and how
3:45:42you can train a model. So for doing all
3:45:44those things, we first need to define
3:45:46your independent variable and dependent
3:45:47variable. So in this case, I want my X,
3:45:50that is our independent variable. I say
3:45:52dataset.iloc.
3:45:54So here I'll be specifying all the rows.
3:45:56So colon basically stands for that. And
3:45:58in the columns, I want only two and
3:46:01three. dot values. So here it should
3:46:04fetch me all the rows and only the
3:46:05second and third column, which is age
3:46:07and estimated salary. So these are the
3:46:09factors which will be used to predict
3:46:11the dependent variable that is purchase.
3:46:13So, here my dependent variable is
3:46:15purchase.
3:46:16And the dependent variable is of age and
3:46:17salary. So, I'll say
3:46:20dataset.iloc
3:46:21I'll have all the rows and I just want
3:46:24fourth column that is my purchase
3:46:25column. dot values. All right, so I just
3:46:28forgot one
3:46:29one square bracket over here. All right.
3:46:31So, over here I have defined my
3:46:33independent variable and dependent
3:46:34variable. So, here my independent
3:46:37variable is age and salary and dependent
3:46:39variable is the column purchase.
3:46:41Now, you must be wondering what is this
3:46:42iloc function. So, iloc function is
3:46:44basically an indexer for pandas data
3:46:46frame and it is used for integer-based
3:46:48indexing or you can also say selection
3:46:50by index.
3:46:52Now, let me just print these independent
3:46:53variables and dependent variables. So,
3:46:56if I print the independent variable, I
3:46:57have the age as well as the salary.
3:47:00Next, let me print the dependent
3:47:01variable as well. So, over here you can
3:47:03see I just have the values in zero and
Linear Regression Vs Logistic Regression
3:47:05one. So, zero stands for did not
3:47:07purchase. Next, let me just divide my
3:47:10dataset into training and test subset.
3:47:12So, I'll simply write in from
3:47:13sklearn.cross_split
3:47:15dot cross_validation
3:47:18import train_test.
3:47:20Next, I'll just press shift and tab. And
3:47:23over here I'll go to the examples and
3:47:25just copy the same line.
3:47:27So, I'll just copy this.
3:47:30I'll remove the points.
3:47:31Now, I want the test size to be let's
3:47:33say 25. So, I have divided the train and
3:47:35test split in 75 25 ratio.
3:47:38Now, let's say I'll take the random
3:47:40state as zero. So, random state
3:47:42basically ensures the same result or you
3:47:44can say the same samples taken whenever
3:47:46you run the code. So, let me just run
3:47:48this. Now, you can also scale your input
3:47:50values for better performing and this
3:47:52can be done using standard scalar. So,
3:47:54let me do that as well. So, I'll say
3:47:56from sklearn.preprocessing
3:48:00import standard scalar.
3:48:02Now, why do we scale it? Now, if you see
3:48:04our dataset we are dealing with large
3:48:07numbers.
3:48:08Well, although we are using a very small
3:48:10data set, so whenever you're working in
3:48:12a broad environment, you'll be working
3:48:13with large data set where you'll be
3:48:15using thousands and hundred thousands of
3:48:17tuples. So, there scaling down will
3:48:19definitely affect the performance by a
3:48:20large extent. So, here let me just show
3:48:22you how you can scale down these input
3:48:24values. And then the pre-processing
3:48:26contains all your methods and
3:48:27functionality which is required to
3:48:29transform your data.
3:48:30So, now let us scale down for test as
3:48:32well as our training data set. So, I'll
3:48:34first make an instance of it. So, I'll
3:48:36say
3:48:37standard scalar.
3:48:39Then I'll have X_train. I'll say sc.fit
3:48:43fit_transform.
3:48:46I'll pass in my X_train variable.
3:48:52And similarly, I can do it for test
3:48:53wherein I'll pass the X_test.
3:48:57All right.
3:48:58Now, my next step is to import logistic
3:49:00regression. So, I'll simply apply
3:49:01logistic regression by first importing
3:49:03it. So, I'll say from sklearn
3:49:06from sklearn.linear_model
3:49:09import logistic regression.
3:49:12Now, over here I'll be using classifier.
3:49:14So, I'll say classifier. is equals to
3:49:16logistic regression.
3:49:18So, over here I'll just make an instance
3:49:19of it. So, I'll say logistic regression.
3:49:22And over here I'll just pass in the
3:49:23random state which is zero.
3:49:26And now I'll simply fit the model.
3:49:31And I'll simply pass in X_train and
3:49:32Y_train.
3:49:34So, here it tells me all the details of
3:49:36logistic regression.
3:49:39Then I have to predict the values. So,
3:49:41I'll say Y_pred
3:49:42is equals to classifier
3:49:45then predict function
3:49:47and then I'll just pass in X_test.
3:49:50So, now we have created the model, we
3:49:52have scaled down our input values, then
3:49:53we have applied logistic regression, we
3:49:55have predicted the values, and now we
3:49:57want to know the accuracy. So, to know
3:49:59the accuracy, first we need to import
3:50:01accuracy score. So, I'll say from
3:50:03sklearn.metrics
3:50:06import accuracy score.
3:50:08And using this function we can calculate
3:50:09the accuracy. Or you can manually do
3:50:12that by creating a confusion matrix.
3:50:14So, I'll just pass in my Y_test and my
3:50:16Y_predicted.
3:50:19All right. So, over here I get the
3:50:20accuracy as 89%. So, if we want to know
3:50:23the accuracy in percentage, so I just
3:50:24have to multiply it by 100. And if I run
3:50:26this,
3:50:27so it gives me 89%. So, I hope you guys
3:50:30are clear with whatever I have taught
3:50:31you today. So, here I have taken my
3:50:33independent variables as age and salary.
Decision Tree Algorithm
3:50:35And then we have calculated that how
3:50:37many people can purchase the SUV. And
3:50:40then we have calculated our model by
3:50:42checking the accuracy. So, over here we
3:50:44get the accuracy as 89, which is great.
3:50:51Let's compare the two models.
3:50:54So, first of all, let's look at the
3:50:55definition of linear regression and
3:50:57logistic regression. So, the main aim of
3:51:00linear regression is to predict a
3:51:01continuous dependent variable based on
3:51:04the values of the independent variables.
3:51:07But when it comes to logistic
3:51:08regression, the aim is to predict a
3:51:10categorical dependent variable based on
3:51:13the values of independent variables.
3:51:15These are the main aim of each of these
3:51:17models. Now, let's look at the variable
3:51:20type. Now, in linear regression, the
3:51:22dependent variable is always continuous.
3:51:25All right. This is very important to
3:51:26remember, because this is the main
3:51:28objective of linear regression. Okay, it
3:51:30makes use of continuous dependent
3:51:32variables to predict continuous values.
3:51:34Similarly, when it comes to logistic
3:51:36regression, you're going to use
3:51:38categorical dependent variable to
3:51:40predict a categorical value. All right.
3:51:43Now, let's look at the estimation
3:51:44method. So guys, linear regression is
3:51:47based on the least square estimation,
3:51:49which basically says that the regression
3:51:51coefficients should be chosen in such a
3:51:54way that it minimizes the sum of the
3:51:56squared distance of each observed
3:51:58response. Okay? Now, this is in-depth
3:52:01about linear regression, so that's why
3:52:02I'm going to leave a link in the
3:52:03description. Now, logistic regression on
3:52:06the other hand is based on maximum
3:52:08likelihood estimation. Okay? This
3:52:10basically says that the coefficients
3:52:12should be chosen in such a way that it
3:52:15maximizes the probability of Y given
3:52:18some value of X. Next is the equation.
3:52:21Earlier, we discussed this equation
3:52:22where for linear regression, we have Y
3:52:24is equal to B not plus B1 into X plus E.
3:52:28Okay? Similarly, this is the equation
3:52:30for logistic regression.
3:52:32So, the next difference is a best fit
3:52:33line. So, guys, linear regression aims
3:52:36at finding the best fitting straight
3:52:38line, which is also called the
3:52:40regression line. All right? But, when it
3:52:42comes to logistic uh regression, if you
3:52:44try and uh map the relationship between
3:52:46the dependent and independent variable,
3:52:48you're going to get a curve, which is
3:52:50also known as the sigmoid curve. All
3:52:52right? So, in linear regression, the
3:52:53relationship between the dependent and
3:52:55independent variable is represented
3:52:58using a straight line, but when it comes
3:52:59to logistic regression, the relationship
3:53:01between the dependent and independent
3:53:03variable is represented using a sigmoid
3:53:06curve. Now, let's look at the
3:53:07relationship between dependent and
3:53:09independent variable. Now, when it comes
3:53:11to linear regression, there has to be a
3:53:13linear relationship between the two. So,
3:53:16when I say linear, I mean that the
3:53:17variables have to vary linearly. Okay?
3:53:20That's how the straight line is formed
3:53:22in the first place. All right? When it
3:53:24comes to logistic regression, it's not
3:53:26necessary to have a linear relationship.
3:53:29Now, the output of linear regression is
3:53:31always going to be a predicted integer
3:53:33value or basically a continuous value.
3:53:36So, when it comes to logistic
3:53:37regression, the output has to be a
3:53:39binary value. Okay? So, it should either
3:53:41be class A or class B or it should be
3:53:43zero or one, something like that.
3:53:45Finally, we have applications. Now,
3:53:47linear regression is mainly used to
3:53:49predict outcomes like the expected
3:53:52number of sales. And you know, it's
3:53:54always used to predict some continuous
3:53:55value. All right, but when it comes to
3:53:57logistic regression, it's mainly used in
3:53:59classification. So, when you want to
3:54:01classify a data set into two different
3:54:03classes, then you use logistic
3:54:05regression. You can find a lot of
3:54:07applications of linear regression in the
3:54:09business domain. And logistic regression
3:54:12is mainly used in the cybersecurity,
3:54:14image processing, and classification
3:54:16domain.
3:54:17>> [music]
3:54:22>> What is classification?
3:54:25And if I have to tell you about
3:54:26classification,
3:54:29like for example, what happens is like
3:54:32we have two type of when we talk about
3:54:34machine learning, machine learning is
3:54:36nothing but you know, like a series of
3:54:38instructions you give it to the computer
3:54:41so that it can learn the patterns from
3:54:42your data set, right? To give you an
3:54:44example, imagine that there is a
3:54:46trending topic, for example, you found
3:54:48it you want to find it out whether
3:54:51Prime Minister Modi will be the second
3:54:53Prime Minister once again the Prime
3:54:55Minister for the country or not, okay?
3:54:57So, now what you will do is you will
3:54:59collect the data set from multiple
3:55:00different sources. And you will you will
3:55:04actually build it a like a algorithm
3:55:06where you you will get a label as yes or
3:55:09no. Yes, he will continue as a next
3:55:11Prime Minister or no, he will not
3:55:12continue as a next Prime Minister. So,
3:55:14you will collect the data set and you
3:55:16will feed this data set to the computer
3:55:17and this process is called as a machine
3:55:20learning, right? So, now in this case
3:55:22what happens is um
3:55:24So, this this is about the
3:55:26classification, right? So, now machine
3:55:29learning is basically of two type. One
3:55:31is called as a supervised machine
3:55:33learning, another is called as a
3:55:34unsupervised machine learning, and the
3:55:36third one is a reinforcement machine
3:55:38learning. So, when we speak about
3:55:40supervised machine learning, as the name
3:55:42suggests, it provides some supervision,
3:55:44right? For example, the teacher teaching
3:55:46the kid is a supervised machine
3:55:48learning. Right? So, we will give the
3:55:50trained examples. We will give the
3:55:52trained data set with the pure label on
3:55:54top of that. Uh this is called as a
3:55:57supervised machine learning. So, if I
3:55:58draw in front of you, this type of
3:56:00machine learning look like this,
3:56:02supervised machine learning. Where what
3:56:04happens is like you would have the data
3:56:06set, which is a structured data set, and
3:56:08you would have one column, which is
3:56:10called as a label. What you want to
3:56:12predict, okay? And you would have a uh
3:56:15various predictors by which you want to
3:56:17predict. To give you an example, imagine
3:56:19that you want to predict the pricing of
3:56:21a community. Okay? You want to predict
3:56:23what would be the pricing of apartment
3:56:25in a particular community, right? Now,
3:56:27these can be the variable like you can
3:56:29see that uh what would be the number of
3:56:31how many floors it has. You can have a
3:56:33variable like what is a pollution level,
3:56:35how how many educational institution are
3:56:37nearby, right? So, based upon that, the
3:56:39pricing will change. But this type of
3:56:42supervised machine learning, why it is
3:56:43called supervised machine learning
3:56:45because we provide the independent
3:56:47variable or we provide the predictors.
3:56:50Also, we provide the label data set.
3:56:52Okay? Now, the supervised machine
3:56:54learning is basically of two type. This
3:56:57supervised machine learning, one type
3:56:59one first is called as the, you know,
3:57:02regression based supervised machine
3:57:03learning.
3:57:05Okay? And the second one is called as a
3:57:07classification based supervised machine
3:57:09learning.
3:57:10Now, what is a difference between
3:57:11regression based and a classification
3:57:13based supervised machine learning?
3:57:15Regression based supervised machine
3:57:17learning is that machine learning where
3:57:19what you want to predict
3:57:22is continuous in nature. Okay? Imagine
3:57:25that you want to predict the
3:57:27community prices, right? Which is a
3:57:29continuous value. If it is a continuous
3:57:32value, then we will go ahead with
3:57:33regression based supervised machine
3:57:35learning. Okay? Whereas, if you want to
3:57:38predict something which is the discrete
3:57:40outcome, to give you an example, you
3:57:42want to predict that whether I will win
3:57:44the match or not.
3:57:46Okay? I want to predict whether the
3:57:48particular employee will turn out from
3:57:50the company or not. You want to predict,
3:57:53you know, like whether the person will
3:57:55have a cancer or not.
3:57:57You're getting my point, right? So, if
3:57:59you have the output which you want to
3:58:01predict is in the form of yes or no or
3:58:04true or false, right? This is called as
3:58:07the supervised machine learning, but a
3:58:10classification-based supervised machine
3:58:12learning. Okay? So, classification-based
3:58:15supervised machine learning is the
3:58:16process of dividing the data set into
3:58:19different categories or group by adding
3:58:21a label. Okay, so always remember that
3:58:24whenever you guys want to predict the
3:58:27classes in the data set, whenever you
3:58:29want to predict, you know, like whether
3:58:31this will happen or not, whether the
3:58:33whether a person will do a credit card
3:58:36fraud or not. You're getting my point,
3:58:38right? Whether the employee will turn
3:58:39out from the company or not, right?
3:58:42Whether the particular person will have
3:58:43a diabetes as a disease or not. All
3:58:45these questions, wherever you want to
3:58:47find it out yes or no or true or false
3:58:50or you want to predict classes in the
3:58:51data set, this is called as a
3:58:53classification-based supervised machine
3:58:56learning. Okay? So, this is what is
3:58:58called as a classification-based
3:59:00supervised machine learning and
3:59:02today I will teach you, you know, I will
3:59:05tell you about various form of
3:59:06classification-based supervised machine
3:59:08learning. Although we will do a deep
3:59:10dive into decision tree. Right? Now, you
3:59:13will be able to understand that decision
3:59:15tree how decision tree is connected to
3:59:17the network what we are learning today
3:59:19with classification-based supervised
3:59:22machine learning. Okay? So, now
3:59:25we have various algorithms. Algorithm is
3:59:28nothing but a set of mathematical
3:59:30equations for classification-based
3:59:33supervised machine learning. We first of
3:59:36all, we have something called as
3:59:37decision tree. Then we would have
3:59:39something called as random forest, then
3:59:41naive base, and then KNN, which is
3:59:43called as a K nearest neighbors. Okay?
3:59:46So, let me give you few statements about
3:59:48this algorithm, and we have many others,
3:59:50but today we will focus on one of them,
3:59:52which is decision tree.
3:59:54So, now first let's start with decision
3:59:56tree. Now, what is a decision tree? What
3:59:58I was telling you is, believe me or not,
4:00:00decision tree is something you use every
4:00:02day in your um in your daily life. For
4:00:05example, you take decisions, and for ex-
4:00:08today also you took a decision to attend
4:00:10this webinar, right? But, how do you
4:00:13decide a decision based on various
4:00:15further decisions, right? For example,
4:00:18for today joining the webinar, you have
4:00:20seen that, okay,
4:00:21uh when this webinar is about, okay? So,
4:00:24you said it is weekday or weekend. Then
4:00:26you might have said check it out, right?
4:00:28What is a time of the webinar? Then you
4:00:30might have checked it out, what is a
4:00:31topic of the webinar, right? So, and who
4:00:34is conducting this webinar? So, based
4:00:36upon this, you took a decision, shall I
4:00:38go ahead or not go ahead? You're getting
4:00:40my point, right? So, this is what is
4:00:42called as a decision tree. We call it as
4:00:44decision tree because it is a graphical
4:00:46representation of all the possible
4:00:48solution to a decision. It's like a
4:00:49tree. Like, decision tree is like a
4:00:52tree. Why? Because a tree also start
4:00:55with a root, and then it emerge into
4:00:57various branches. Similarly, you have a
4:00:59decision tree, which I'm showing to you
4:01:01here as a simplest algorithm, which is
4:01:03being used for machine learning
4:01:05purposes.
4:01:06Now, it is something like this. Imagine
4:01:08that you want to find it out that, you
4:01:11know, like, do you want to go to a
4:01:13restaurant, or do you want to buy an
4:01:15hamburger? Okay? So, you have two
4:01:17choices. Either you can go for a
4:01:19restaurant, or either you can buy a
4:01:21hamburger. Now, how would you decide
4:01:23which one you would follow? So, you will
4:01:25start with what is called as a root
4:01:27node. You will start with the root node
4:01:28that whether I'm hungry or not, right?
4:01:32If I am hungry, right? Then only I will
4:01:35go for all these activity. If I am not
4:01:38hungry at all, then simply go and sleep.
4:01:41You got my point, right? So, this is how
4:01:43you will start with the first node,
4:01:45which is to find it out whether I'm
4:01:47feeling hungry or not. If I'm feeling
4:01:48hungry, then I will decide that, "Well,
4:01:51I Do I have money, which is around $25
4:01:54worth?" If I have money, then I will go
4:01:56for a restaurant. If I don't have a
4:01:58money, then I will buy a hamburger.
4:02:00You understood this? This is simplest
4:02:03representation of decision tree.
4:02:05Basically, you decide something on the
4:02:07basis of the previous outcomes. And you
4:02:10can imagine any sort of example here.
4:02:12Imagine that you want to find it out
4:02:14whether the person will do a credit card
4:02:16fraud or not. Again, it will depend upon
4:02:18the previous circumstances that, you
4:02:20know, for example, it will depend upon
4:02:23how much is the salary of the person.
4:02:24You It will depend upon what is a job
4:02:26profile of the person. It will depend
4:02:28upon the fact that, you know, like, for
4:02:29example, how many fraud or how many card
4:02:32this person have, right? So, on [snorts]
4:02:34basis of you will decide that whether
4:02:35this will do a credit card fraud or not.
4:02:39So, this is the simplest form of
4:02:41classification-based algorithm.
4:02:43Then we have the next algorithm, which
4:02:45is called as a random forest. Now, what
4:02:48is a random forest? As the name
4:02:50suggests, you built one decision tree in
4:02:52the last example, right? But
4:02:55one decision tree can sometime be
4:02:57over-fitting as a case, right? So,
4:02:59people said that, you know, like, "Why
4:03:01should I only trust one decision tree?"
4:03:03For example, whenever you take a bold
4:03:05decision in your life, you don't trust
4:03:08only a single voice, right? You want to
4:03:10hear [clears throat] it from multiple
4:03:12different people to make your decision
4:03:14more stronger, right? Go to a doctor. If
4:03:16the doctor says to you that you have
4:03:18this type of a disease, you don't
4:03:21believe in that. What you do is you also
4:03:23talk to second doctor, third doctor,
4:03:25fourth doctor to confirm that this is
4:03:27true or not. Okay? So, this is what is
4:03:30called as a random forest. As the name
4:03:32suggests, why it is called random
4:03:33forest? It is called random forest
4:03:36because now you're building the various
4:03:38number of decision trees here. It's like
4:03:40a forest of all the trees, right? It is
4:03:43no longer a single tree, it is a forest
4:03:45of complete trees.
4:03:47So, for example, you can imagine that if
4:03:49it is my training data set, I will I
4:03:51will split my training data set into
4:03:53multiple examples, multiple decision
4:03:55trees will be built, and then based upon
4:03:58the majority, I will decide whether
4:04:00should I do this or not. Okay? So, it is
4:04:03also called as a bagging sort of
4:04:04methodology where we bring the outcome
4:04:07of various models, or we bring the
4:04:09outcome of various trees all together to
4:04:12make the
4:04:13uh to make a powerful decision. Okay?
4:04:16So, this is another example This is
4:04:18another algorithm which is called as a
4:04:20random forest.
4:04:23Then, after random forest, we have
4:04:25something called as naive base. Okay?
4:04:28Now, what is a naive base algorithm?
4:04:31Naive base is a also a simplest
4:04:33algorithm, but naive base is basically
4:04:36based on the base theorem. Okay? So,
4:04:38this algorithm is basically based on the
4:04:40base theorem, and base theorem is based
4:04:42on the conditional probability.
4:04:45Right? So, it is based on the
4:04:47conditional probability, which is your
4:04:49naive base. Now, in naive base, what
4:04:52happens is like we decide that whether
4:04:56something will happen or not on the
4:04:58basis of probability. So, let me
4:05:01illustrate you with this example.
4:05:03Imagine that you want to find it out
4:05:06whether I would have a disease or not.
4:05:09Okay? So, first of all, the probability
4:05:12of having a disease is 0.10. And the
4:05:16probability of not having a disease is
4:05:180.90. Okay? So, if there is a
4:05:21probability of having a disease is 0.10,
4:05:24you will further find it out that what
4:05:26is a probability that my test will be
4:05:30positive given I'm diseased.
4:05:33What is the probability that my test to
4:05:35diagnose the disease is negative given I
4:05:37have a disease? Right? Similarly, if you
4:05:40go into this direction, if there is a
4:05:42probability of not having a disease is
4:05:440.90, you will find it out that what is
4:05:47a probability of [snorts] having a
4:05:49disease or test being positive with
4:05:51having with no disease. And what is a
4:05:54probability of no disease given I don't
4:05:57have the disease, which is 0.90. So,
4:05:59basically, here you check the outcomes.
4:06:03Here you check the outcome of all the
4:06:06all the possible combinations. This is
4:06:08what is called as the naive base
4:06:10algorithm or base theorem or
4:06:13you know, conditional probability base
4:06:15theorem. Okay?
4:06:17Then we also have something which is
4:06:19called as a K nearest neighbor, which is
4:06:22the one of the finest algorithm
4:06:25which also helps you in deciding the
4:06:28classification, right? Now, what is K
4:06:31nearest neighbor? K nearest neighbor, as
4:06:33the name suggests, what happens in this
4:06:35case is we try to build, you know, we
4:06:38try to build
4:06:40basically, it's like a neighbor. So,
4:06:42it's something like this. Like if I give
4:06:44you the data set, say I give you the
4:06:46customer data set. Okay?
4:06:49And I have It is a transaction data set.
4:06:51So, customer number one
4:06:53has bought a product number one, right?
4:06:56From a particular vendor at a particular
4:06:58rate, and this is a profit this customer
4:07:00has given to us. And this is the
4:07:02revenues. Okay? This is my class
4:07:04customer number one. Similarly, you
4:07:07would have customer number two, customer
4:07:09number three, customer number four,
4:07:11right? Now, if I ask you that which
4:07:13customer profiling is same
4:07:15is nearby same, right? So, what you can
4:07:18do is you can group your customers or
4:07:20you will get to know that which customer
4:07:23behave in the similar way. So, you can
4:07:25say that customer one, customer three,
4:07:27and customer four, they behave in a
4:07:29similar way because they are giving us
4:07:30the high profit and high revenue
4:07:32margins.
4:07:34You understood? So, this is what happens
4:07:37in the case of K nearest neighbor. So, K
4:07:39nearest neighbor, what you generally do
4:07:41is you find it out that what uh I would
4:07:45be able to, you know,
4:07:47uh find it out who is my nearest
4:07:49neighbor, like what are the similarities
4:07:51in the patterns we have, right? This is
4:07:53what we use for K nearest neighbors. And
4:07:56this you can see that, for example, the
4:07:58algorithm based on the distance-based
4:08:00mechanism find it out that how many
4:08:02people or how what type of audience is
4:08:05similar to the outcomes, right?
4:08:08So, then moving further here
4:08:11the with the second topic, let's go into
4:08:13detail about what is decision tree all
4:08:15about, right? So, so far I have touched
4:08:18base on uh classification-based
4:08:20algorithm, right? And I was teaching you
4:08:22different type of classification-based
4:08:24algorithm, but since the focus of
4:08:26today's class is decision tree, let's
4:08:28take a deep dive into the decision tree.
4:08:31Okay? So, let's get started with
4:08:33decision tree.
4:08:34A decision tree is a graphical
4:08:36representation of all possible solution
4:08:39to a decision based on a certain
4:08:40condition. What does it mean? It means
4:08:43that it is as simple as that. Imagine
4:08:45that this is a tree, right? So, it is
4:08:47like a tree where
4:08:49you have a problem statement that should
4:08:51I accept a job offer or not. Imagine
4:08:54that you want to find it out that should
4:08:56I accept a job offer or not. This is a
4:08:58problem statement which is there in your
4:09:00mind. Now, how would you solve this
4:09:01problem statement with the help of a
4:09:03decision tree?
4:09:04First of all, you will start with what
4:09:06we call as a root node. We will start
4:09:08with a root node starting with, you
4:09:11know, to find it out what is the salary.
4:09:13Okay, what I'm what I what is the salary
4:09:15I'm getting. So if my salary is equal to
4:09:19or greater than equal to 50,000, I'll go
4:09:22here. If my salary is not greater than
4:09:25equal to 50,000, I will say I will not
4:09:28accept this offer.
4:09:29Okay? Now imagine that you say that your
4:09:32salary is greater than 50,000, then you
4:09:34will check another
4:09:36another variable here. You will check it
4:09:38out whether I have to commute more than
4:09:401 hour. If I have to commute more than 1
4:09:43hour, I will decline the offer.
4:09:46You're getting my point, right? Then you
4:09:48will if you don't have to commute more
4:09:50than 1 hour, then you will still
4:09:51consider this option and then you will
4:09:53check it out. For example, in this case
4:09:54we're checking it out whether you're
4:09:56getting the free offers also like coffee
4:09:59or some snacks or other things like
4:10:00that. If yes, then you will accept
4:10:03finally the offer, otherwise you will
4:10:05decline the offer.
4:10:07You got my point. This is how our
4:10:09decision tree works. Basically, decision
4:10:11tree will keep on splitting, keep on
4:10:13splitting unless and until you are able
4:10:16to find it out your decision.
4:10:18Okay? So here the decision was shall I
4:10:21accept this or not, right? So it will
4:10:23keep on splitting, keep on splitting
4:10:25unless and until you get your decision
4:10:27whether you do this or not. Like this
4:10:30example which I have to I have explained
4:10:32you. Okay?
4:10:34Now with this what happens is like let
4:10:37me if I go further and explain you
4:10:39further on this decision tree, let's
4:10:41understand this more importantly, right?
4:10:43Let's understand this
4:10:45one by one. So imagine that this is my
4:10:48data set. Okay? Now in my data set you
4:10:51can see that
4:10:53I have various [clears throat] colors
4:10:54given to you. It's like a green color,
4:10:55yellow color, red color, red color and a
4:10:58yellow color. I'm saying if this is a if
4:11:01there is a green color fruit with a
4:11:03diameter of three, right? And I can say
4:11:06it is a mango.
4:11:08Whereas I'm saying if it is a yellow
4:11:10color and a diameter three, still will
4:11:12be called as a mango. If it is a red
4:11:14color with one diameter, then it is a
4:11:17grape.
4:11:18And it can be a red color with a
4:11:20diameter of one still can be a grape,
4:11:22but even a yellow with a diameter of
4:11:25three can be lemon. Now, what is
4:11:27happening in this case is if you have
4:11:29this type of a data set, where you want
4:11:31to predict the label of the fruit,
4:11:33right? You want to predict whether a
4:11:35fruit will be mango, whether a fruit
4:11:37will be lemon, right? Or whether a fruit
4:11:40in this case is mango and lemon or
4:11:42grape.
4:11:43I have to build a classifier, which is a
4:11:45decision tree classifier on the basis of
4:11:47this data set. Imagine this is a problem
4:11:50statement given to you. Now, first of
4:11:52all, what you will do here is you will
4:11:54take this data set, right? You will take
4:11:56this data set and you will start with a
4:11:59root node. Root node imagine here is
4:12:01like, is my diameter of the fruit
4:12:03greater than equal to three or not?
4:12:06Okay? Is my diameter of the fruit
4:12:08greater than equal to three or not? If
4:12:11it is greater than equal to three,
4:12:13right? Now, if it is greater than equal
4:12:16to three, you can see that you have
4:12:18three fruits here, three rows of the
4:12:19data set here. One is green with three,
4:12:22mango. Yellow with three, lemon. Yellow
4:12:25with three, mango. And wherever it fails
4:12:28in the condition, you are left with
4:12:30where the diameter is not greater than
4:12:32equal to three, it is less than equal to
4:12:34three, then what which are the data set
4:12:36we have? Red one, grape and red one,
4:12:39grape.
4:12:41Okay? So, I hope you are understanding
4:12:43this, right? How we have starting with
4:12:45the decision tree with the base on one
4:12:47condition here, right? Which is
4:12:49diameter. Now, based upon this, what I
4:12:52will do is I have to split it further.
4:12:54Here in this case, I don't have to split
4:12:56it because I already got the result. So,
4:12:58if my diameter is greater than equal to
4:13:00three in my data set, if the diameter is
4:13:03not equal to greater than equal to
4:13:05three, I know it is a grape. Right?
4:13:08Whereas, if the diameter is greater than
4:13:10equal to three, then it can be a mango
4:13:13or it can be a lemon. I'm not sure on my
4:13:15decision. So, what I will do here is I
4:13:17will split it further. So, I have to
4:13:20split this further.
4:13:22Here, the splitting is not required.
4:13:24But, if I split this further, I may
4:13:26check it out now. Is the color equal to
4:13:29yellow or not? Now, if the color is
4:13:31equal to yellow, then I have two fruit
4:13:34which is
4:13:35your this row, which is, you know, um
4:13:38your mango or lemon. And if the color is
4:13:42not equal to yellow, then you will have
4:13:44the another row of the data set, which
4:13:46is, you know, which you're left with,
4:13:48right? So, this is how you have done
4:13:50this work
4:13:52in the case of your decision tree.
4:13:55Okay? And how you will find it out at
4:13:57which type of the criteria or which type
4:14:00of the algorithm I should or which type
4:14:02of criteria or variable I should choose
4:14:04here to split the tree is on the basis
4:14:06of Gini index and the information gain,
4:14:09which I will illustrate and show you in
4:14:11the few slides from now. Okay? So, I
4:14:14hope with this you understand the how
4:14:17your decision tree works. Basically, on
4:14:18the basis of the condition, and if you
4:14:20get a pure subset, then no need to split
4:14:23it further. If there is no pure subset,
4:14:25keep on splitting, keep on splitting
4:14:27unless and until you get a pure subset.
4:14:30Okay?
4:14:32So, in this case, what has happened is
4:14:34you got 100% mango. Here, you got 50%
4:14:37mango and 50% lemon. Okay?
4:14:41So, now
4:14:43this is what I have told you in the
4:14:45previous slide, like how does it work,
4:14:48right?
4:14:50Now, this all depend upon the Gini index
4:14:53basis, right? Now, in the next
4:14:55subsequent slide, let me tell you and
4:14:58explain you about Gini index and
4:15:00information gain. How does it work?
4:15:02Right? How does that something works on
4:15:05the basis of uh your Gini index?
4:15:08So, imagine that, you know, like imagine
4:15:11that I have a feature here is the color
4:15:14green or not, right? Basis of whether
4:15:16the feature whether you have a color
4:15:18green or not, what would happen here is
4:15:21if the color is green,
4:15:22you will get this row here. If it is
4:15:24not, you will get this these two rows
4:15:26here, right? In the next, it may decide
4:15:29on the other basis, right? It can be is
4:15:31the diameter which is greater than equal
4:15:33to three or not and many other things.
4:15:36Okay?
4:15:37Now, let's go ahead the example and the
4:15:40questions which is coming to your mind
4:15:42which is related to, you know, decision
4:15:44tree terminologies. I'm pretty sure you
4:15:47would have these questions in your mind
4:15:49and you would be thinking that how
4:15:50should I decided which feature I should
4:15:52use and which feature shouldn't I be
4:15:54using? This is the question which you
4:15:55had. And this is a excellent question
4:15:57for understanding the decision tree. But
4:16:00before I go on to that, I have to tell
4:16:02you about some of the terminologies
4:16:05which we commonly use while we build the
4:16:07decision tree. So, first of all,
4:16:09decision tree looks like this type of a
4:16:11tree-based structure, okay? So, every
4:16:13decision tree will have its root node.
4:16:16So, root node is where the decision tree
4:16:18will start. So, it represent the entire
4:16:20population or sample and it is further
4:16:23get divided or into two or more
4:16:25homogeneous sets. So, as you know that
4:16:26this will be the first feature on the
4:16:28basis of your tree will start. Like tree
4:16:31start with a root, your in this case,
4:16:33your decision tree will also start with
4:16:35a root node.
4:16:37Okay? Now, once you have the root, after
4:16:39that in a tree, what happens is you will
4:16:41start getting the branches, right? I
4:16:43hope you understand what is branches. In
4:16:45this case, we will keep on splitting,
4:16:47keep on splitting unless and until we
4:16:49get a decision whether this will happen
4:16:50or not. So, it's like a branches.
4:16:53Okay? Then, we would also have a parent
4:16:55or a child node. So, child node is
4:16:57nothing but when you have a branches,
4:17:00then the branches can also have the
4:17:01outcomes. So, in the previous example,
4:17:04like for example, we have whether the
4:17:06diameter is greater than equal to three
4:17:09or not. You remember? Whether the
4:17:10diameter is greater than equal to three
4:17:12or not. If it is true, what is
4:17:14happening? If it is a false, what is
4:17:15happening? So, this is a child or child
4:17:18node. Basically, this is a intermediate
4:17:20node. This is not the final decision
4:17:22which is being made. So, this is called
4:17:24as a parent or the child node.
4:17:27Then, we also have other terminology
4:17:29here, which is splitting, which you know
4:17:31that we will keep on splitting unless
4:17:33and until you get a desired node. And
4:17:35finally, the tree end with a leaf node.
4:17:37So, always remember one thing, you will
4:17:39start your tree with a root node. Root
4:17:42node is a node with which we will start
4:17:44the decision tree. And you will end your
4:17:46decision as decision tree at a leaf node
4:17:49where you will get a decision that you
4:17:51should do this or not.
4:17:53Okay? And now,
4:17:56uh pruning is a activity where
4:17:58uh you you will cut down the decision
4:18:00tree if it is pruned lot amount of
4:18:03times. I will even explain you this. It
4:18:05is a case of overfitting. You shouldn't
4:18:08build thousand You can shouldn't build
4:18:09thousand uh branches of the tree when it
4:18:12is not even required. So, I'll I'll
4:18:13explain you this point.
4:18:17Now, let's move further here and let's
4:18:20see this with the help of an example,
4:18:23which was your
4:18:24uh Gini index and your information gain.
4:18:28Right?
4:18:29So, this is where you were asking me
4:18:32that which question to ask and when,
4:18:34right? How would you decide that which
4:18:37feature has to be taken first? Right?
4:18:39Let's take up an example here and I
4:18:42would encourage everyone of you to hear
4:18:45me really fine here because this is on
4:18:47the basis how would you decide or how
4:18:49the algorithm decide to break this
4:18:51further. And and I'm explaining you with
4:18:53the help of one
4:18:55simplest example, and we will do some
4:18:57maths over it.
4:18:58So, imagine that
4:19:00there is a data set where I want to find
4:19:03it out whether I will play the match or
4:19:06not.
4:19:08Okay, there is a cricket match or there
4:19:10is a football match or whatever it is. I
4:19:12want to find it out whether I will play
4:19:14the match or not. Okay? Now,
4:19:17how do the decision tree look like?
4:19:19Decision tree look like this. If the
4:19:21outlook if the outlook is humid
4:19:25if the outlook is humid
4:19:27and if the outlook is humid and humidity
4:19:30is very high, then I will not play the
4:19:33match. On the other hand, if the outlook
4:19:36is humid and humidity is normal, I may
4:19:38play the match.
4:19:40Okay?
4:19:41If the outlook is
4:19:42absolutely clear, then I will always
4:19:44play. On the other hand, if outlook is
4:19:47windy and the winds are very strong, I
4:19:49may not play the match. On the other
4:19:51hand, if the outlook is windy and the
4:19:54winds are weak, I may play the match.
4:19:56So, now what is happening again is that
4:19:58you are not sure how you decided with
4:20:00outlook, how you decided with these
4:20:03nodes here, how you decided with these
4:20:05features here. Let me try to show you
4:20:07this entire data set. So, this entire
4:20:09data set look like this.
4:20:11Okay? So, this is basically your 14 days
4:20:14data set. And this exactly happens in
4:20:17the case of your classification based
4:20:18algorithm, you will get a data set like
4:20:20this.
4:20:21So, you can imagine that you want to
4:20:23find it out I would like whether I will
4:20:26play or not. Right? This is what you
4:20:29want to essentially find it want to find
4:20:31it out whether I would play or not. On
4:20:33basis of what? On basis of four
4:20:35different variables you have. Four
4:20:37different variables or features in the
4:20:39data set is outlook
4:20:41temperature, humidity, and wind. Right?
4:20:45So, now
4:20:46it can be like if the outlook is sunny,
4:20:49temperature is hot, humidity is high,
4:20:52wind is not there, I will not play the
4:20:55match. This is how you will read this
4:20:57one data point, right? Similarly, I have
4:20:59multiple data point. Now, I have to
4:21:01decide how can I build a decision tree
4:21:04and out of these four features, which
4:21:07feature should I use first as my root
4:21:09node? Right? So, this is what I have to
4:21:12decide and
4:21:14let's do it
4:21:16accordingly, right? So, now what I'll do
4:21:18from here is
4:21:21uh we I will illustrate you what is Gini
4:21:24index, what is information gain, and how
4:21:27does this happens, right? And this is
4:21:29basically used to build your decision
4:21:31tree.
4:21:32So, before I make you understand about
4:21:35Gini index or information gain, one
4:21:38thing which you should always remember
4:21:39is to understand the concept of
4:21:42impurity. What is impurity? Impurity is
4:21:44nothing but you can see there is a
4:21:46basket where you have apples, right?
4:21:49Now, if you have a basket which is of
4:21:50apple and another
4:21:52tray, it is written the label as apple.
4:21:55Now, in this case, you will never make a
4:21:57mistake. You will never make a mistake
4:22:00because
4:22:01here you have an apple, here you have
4:22:03only one label. So, everything will be
4:22:05perfect. It will be 100%. Basically,
4:22:08there is no impurity, there is no
4:22:10problem in your data set, right? On the
4:22:13other hand, let me flip the story. In
4:22:16this case, imagine that I have different
4:22:18fruits in the basket, which is like
4:22:20apple, you have a banana, you have a
4:22:22grapes, you have a you know, cherries,
4:22:25and many other things, right? And you
4:22:26have many apples, many labels here. In
4:22:29this case, try imagining you have to
4:22:32match each fruit with its label.
4:22:36Right? Now, in this case, the impurity
4:22:38cannot be equal to zero. The impurity
4:22:41will not be equal to zero in this case
4:22:43because what would happen here is that
4:22:45you have a chances of misclassification.
4:22:48Right? This is a very important concept,
4:22:50right? When you have perfect thing, the
4:22:53misclassification will not happen. But,
4:22:55whereas, if you have a multiple labels
4:22:57with multiple fruit, misclassification
4:22:59or impurities will not be equal to zero.
4:23:03Right? So, this is associated with a
4:23:06term called as entropy. Maybe in your
4:23:08childhood days, you have learned about
4:23:10entropy in your chemistry class. In a
4:23:12simple sense, what is entropy? Entropy
4:23:15is a randomness of the space sample
4:23:18space. Whenever you are not sure on your
4:23:21decision, then entropy will be more.
4:23:24Imagine that I'm giving you the data set
4:23:26where it is like 51% of doing this
4:23:29thing, 49% of not doing this thing. 51%
4:23:33chance that the employee may leave the
4:23:35organization, 49% chance that employee
4:23:37may not leave the organization. So,
4:23:39basically, you are not sure on your
4:23:40decision. If you are not sure on your
4:23:42decision, then the entropy will be very
4:23:45high. On the other hand, if you are very
4:23:47sure on your decision, then the entropy
4:23:50will be very low. So, basically, we need
4:23:53that feature which can provide us lowest
4:23:56entropy rather than the highest entropy
4:24:00uh to select as that as a good feature.
4:24:04So, we generally find it out entropy by
4:24:06the help of this formula. What we simply
4:24:09do is don't get scared with this formula
4:24:11because everything happens automatically
4:24:13in R or Python, right? Imagine just look
4:24:17at this formula. What is this formula?
4:24:18This formula says that what is the
4:24:21probability that something will happen
4:24:23multiplied by log base two probability
4:24:27that something will happen subtract this
4:24:29with probability that something will not
4:24:31happen into log base two probability
4:24:34that something will not happen. Okay,
4:24:36let's take up an example. Don't worry
4:24:38about it. Let's take up an example that
4:24:41probability that I will win the match
4:24:43you will apply here and probability that
4:24:46I will not play the match or win the
4:24:47match you will apply here and then you
4:24:49will calculate the entropy. Let's take
4:24:53you know an example how you can do this
4:24:56in our case.
4:24:58So, in our case let me show you.
4:25:05Yeah, how we will build the decision
4:25:07tree in our case. In our case you just
4:25:10see here what is happening is that we
4:25:12have
4:25:1414 instances
4:25:16or 14 [clears throat] rows of the data
4:25:17set where nine times if you see it I
4:25:20will play the match and five times I
4:25:23will not play the match. Okay, so if you
4:25:25see carefully there are nine labels
4:25:27where I'm playing the match and there
4:25:29are five labels where I'm not playing
4:25:31the match. Right? So, first of all I
4:25:33have to find it out the total entropy.
4:25:36How I will find the total entropy? This
4:25:38is being determined by probability that
4:25:41I will play the match sub multiply this
4:25:44with log base power two probability that
4:25:47I will play the match. What are the
4:25:49chances that I will play the match? Nine
4:25:51out of 14.
4:25:53All of you will be with me, right? This
4:25:55is nine out of 14. Multiply with log
4:25:58base two nine out of 14.
4:26:00Subtract this with what is the
4:26:01probability that I will not play the
4:26:03match? Five out of 14. Multiply this
4:26:06with log base two five out of 14.
4:26:09So, once you calculate this you will
4:26:11find it out the entropy of this entire
4:26:14system, entropy of this entire system is
4:26:170.94. Okay, so this is the entropy of
4:26:21your entire sample space. This is the
4:26:23first thing. Now, how this entropy will
4:26:26help you in selecting which features you
4:26:28will take or not? So, let's go further.
4:26:32So, now we will take each feature one by
4:26:34one. Whether I should take outlook,
4:26:36whether I should take temperature,
4:26:38whether I should pick up humidity, or
4:26:40whether should I pick up windy, right?
4:26:42Let's go one by one. Now, first I'm
4:26:44plotting for outlook. Imagine for
4:26:47outlook, how many times is what are the
4:26:50distinct value of outlook? Outlook can
4:26:51be sunny,
4:26:53outlook can be overcast, or outlook can
4:26:55be rainy. Right? These are the three
4:26:57different combinations you can have for
4:26:59outlook. Now, if the outlook is sunny,
4:27:02two times I'm playing, three times I'm
4:27:04not playing the match. If the outlook is
4:27:06overcast, I'm always playing the match.
4:27:08If the outlook is rainy, three times I'm
4:27:11playing, two times I'm not playing the
4:27:12match.
4:27:13Right? This is how I have bifurcated it.
4:27:16What I will do in the next iteration is,
4:27:18let me find it out the entropy of
4:27:21outlook.
4:27:22Okay? So, if we start with outlook,
4:27:26remember this formula which is, you
4:27:28know, probability that I will play
4:27:30multiply with the probability that I
4:27:32will play, right? And subtract with
4:27:35probability which I will not play, and
4:27:37log base two of not playing. So, 2 by 5
4:27:40is a chances that I will not I will play
4:27:42into log base two 2 by 5. This is a
4:27:45subtraction here, right? With I will
4:27:48play and log 3 by 3 by 5 I will play.
4:27:52Right? So, you will calculate this
4:27:54entropy when outlook is sunny.
4:27:57Yeah? So, you got this entropy when the
4:27:59outlook is sunny is 0.971.
4:28:02Accordingly, you will proceed with
4:28:04calculating the entropy when the outlook
4:28:07is overcast. Outlook is overcast, always
4:28:10you are playing.
4:28:11Right? If the outlook is overcast, every
4:28:13time you are playing, so it means that
4:28:16you will get a probability of zero. If
4:28:18you apply in that formula, you will get
4:28:20zero.
4:28:20And third, what what what would happen
4:28:23if the
4:28:24if the outlook is sunny? In that case,
4:28:27you will again multiply and you will put
4:28:29the formula and you will get it out
4:28:310.971, right? So, you will find it out
4:28:34entropy for each and every distinct
4:28:37combination of your feature.
4:28:40You got my point, right? Outlook being
4:28:42sunny, outlook being overcast, outlook
4:28:44being sunny
4:28:46uh you know,
4:28:47overcast, sunny, and rainy. This should
4:28:49be replaced here, right? And then
4:28:52what you will do is you will finally
4:28:54calculate the information gain.
4:28:56Information gain is nothing but what you
4:28:58will do is you will pick it up the
4:29:00chances, you will pick it up the entire
4:29:03chances when you are playing, right?
4:29:05Which is five out of 14. You remember
4:29:07five out of 14 were the total chances
4:29:09that I will play into if it is sunny,
4:29:13plus four out of 14 if it is overcast,
4:29:16five out of 14 if it is rainy, right?
4:29:20And you will calculate the information
4:29:22from this outlook. And once you
4:29:24calculate the information, you will
4:29:25subtract this information from your
4:29:28total entropy which I found it out in
4:29:30the last slide which was 0.94. You will
4:29:33subtract this and you will get the
4:29:35information gain. Or this is also called
4:29:37as a information gain from a particular
4:29:40feature. So, basically these type of a
4:29:43calculation first of all, don't get
4:29:44scared away that you have to do this
4:29:46calculation. But what I'm trying to
4:29:48explain you is this is how your
4:29:50algorithm will work for each and every
4:29:53feature. It will calculate the
4:29:55information gain from your feature.
4:29:59Right? If you have the more information
4:30:01gain, it means that this variable is of
4:30:04very very important in predicting that
4:30:07something will happen or not. Okay? So,
4:30:10for outlook, I got the information gain
4:30:13as 0.247 with all the calculation.
4:30:16Remember then we will proceed with wind
4:30:19if the wind is there or not, right? And
4:30:23then I will proceed with wind and I will
4:30:25calculate the information gain and say
4:30:27information gain I found it out is
4:30:290.048, right? Similarly, we will
4:30:31calculate for all the four of them. Let
4:30:34me put it together for all of you. So,
4:30:37this is what happens here. Now, if I put
4:30:40in front of all of you, these were the
4:30:42four different variable I have, outlook,
4:30:45temperature, humidity and wind, right?
4:30:47I'm calculating the information gain for
4:30:50each one of them and the information
4:30:52gain I got for outlook is 0.247.
4:30:56So, if the information gain is highest
4:30:58in a particular feature, that feature
4:31:01will become your root node.
4:31:04You got my point? So, therefore, we will
4:31:06pick it up outlook as our root node.
4:31:09Similarly, for when the tree get
4:31:11started, later on as a branch node also,
4:31:14your information gain will be calculated
4:31:17and wherever whichever feature is giving
4:31:19you more information gain, that will be
4:31:20picked up later in the
4:31:23your tree also.
4:31:25So, this is how you will build your
4:31:26decision tree and you will finally get a
4:31:29decision tree like this, okay? So, now
4:31:32what I will do is
4:31:34you know, like I will quickly show you
4:31:36how do we do the decision tree
4:31:39in your Python, okay? So, I'll show you
4:31:42how do you do this in Python and then I
4:31:44will summarize for all of you that why
4:31:47decision trees or tree based algorithms
4:31:50are better than your
4:31:52other algorithm, okay? So, how do you
4:31:55choose
4:31:56basically that which algorithm you will
4:31:58select and when, okay? So, I'll I'll
4:32:00describe you this, but before let's jump
4:32:03on to Python and go there.
4:32:06So, what I have done here is like I had
4:32:08built the decision tree in front like
4:32:12you know already. So, quickly I will
4:32:15walk you through the commands. Okay? So,
4:32:17what happens here is that
4:32:19in Python
4:32:21as you would be well versed with this
4:32:22that we generally import packages in
4:32:25Python. So, I'm importing NumPy. I'm
4:32:27importing Matplotlib for plotting the
4:32:29chart. I'm importing various packages
4:32:32from your scikit-learn which is
4:32:35for machine learning purposes, right?
4:32:36So, I'm importing your label encoder,
4:32:39your decision tree classifier,
4:32:40classification report, and I also I'm
4:32:43importing your
4:32:44tree, right? So, we have all this which
4:32:48I'm importing, right? After I import,
4:32:50what I will do is I'm reading my my data
4:32:53set. So, I'm showing you this with the
4:32:55help of a Iris data set which is one of
4:32:58the very popular data set for building
4:33:00the a decision tree, right? For any
4:33:03particular source. So, imagine that this
4:33:04is my data set and I'm just showing you
4:33:06six rows of the data set where you know
4:33:08like
4:33:09I want to find it out whether a
4:33:11particular flower species will be
4:33:13setosa, versicolor, or virginica. I have
4:33:16three different flowers which I want to
4:33:18predict and on the basis of sepal
4:33:20length, petal length, sepal width, and
4:33:22petal width. So, basically I have
4:33:24different dimensions of flowers length
4:33:27and width and based on that I want to
4:33:29find it out whether the particular
4:33:31species will be setosa, versicolor, or
4:33:33virginica. You got my point, right? So,
4:33:35this is a data set. So, what I will do
4:33:37is I as you know with every machine
4:33:39learning data set we play with the data
4:33:41set. So, I'm checking the information
4:33:43here
4:33:44like what type of the data type it is.
4:33:46So, sepal length it is a float, petal
4:33:48length it is a float, and species is a
4:33:51object.
4:33:52Why object? Because this is a
4:33:54categorical column.
4:33:56So, after this I'm also checking whether
4:33:58there is any null value present in the
4:34:00data set or not because if there is any
4:34:02null value, then we have to get rid of
4:34:04that null value or we have to replace
4:34:06that null value with some imputed value,
4:34:09right? This is what I'm checking here.
4:34:11Once I do that, I'm also plotting this
4:34:13because we usually do visualization,
4:34:15right? To understand the data set
4:34:17better. So, what I'm doing here in the
4:34:18with the help of SNS, which is your
4:34:20SNS.pairplot, I'm plotting all the
4:34:23possible plots. So, basically, sepal
4:34:25length with sepal width, what type of
4:34:27the combination look like. So, setosa,
4:34:29versicolor, and virginica, there are
4:34:31three different species you can see in
4:34:32the data set. And this is what I'm
4:34:35getting a trend between your sepal
4:34:36length and sepal width.
4:34:38Similarly, this is basically from sepal
4:34:40width to petal length.
4:34:42Right? So, this is how I'm understanding
4:34:45the patterns or I'm understanding the
4:34:46relationship between the variables in
4:34:48the data set. So, this is what I'm
4:34:50understanding here. I'm also checking
4:34:52whether there's a correlation or not.
4:34:54Higher the shade, it means there will be
4:34:56a strong correlation. So, all of these
4:34:58things is being done as a part of
4:35:00exploratory data analysis before even
4:35:02you start your machine learning model,
4:35:04right? Once you do this, after that what
4:35:06I have to do is after that what I'm
4:35:08trying to do is I'm taking your species
4:35:11column, what I want to predict as my
4:35:13target variable, which is your dependent
4:35:15variable. And what I So, this is my the
4:35:18dependent variable and with the help of
4:35:20what I want to predict will be my
4:35:22independent variable. So, I'm calling X
4:35:25all my independent variable and I'm
4:35:27calling target as my dependent variable.
4:35:30Once I do this, then you will also think
4:35:32about it that your
4:35:35uh the variable which I want to predict
4:35:36is the flower species, right? But I want
4:35:39to convert this into zero and one class,
4:35:42zero, one, two class because your
4:35:44computer cannot understand text, right?
4:35:46Computer can only understand the
4:35:48numbers. So, what I'm doing here is I'm
4:35:51uh changing it to the
4:35:52uh I'm changing it to the class.
4:35:55Right? So, what I'm trying to do here is
4:35:57I'm saying, "Wherever it is setosa, it
4:35:59will will zero. Wherever it is
4:36:01virginica, it will become one. And it is
4:36:04versicolor as a third category, it will
4:36:06become two. So, now imagine my data set,
4:36:08my label to predict becomes zero, one,
Random Forest
4:36:11or two instead of three flower which was
4:36:13setosa, versicolor, and virginica. This
4:36:15is what I'm doing with the encoder here.
4:36:18Once I convert this, and this become my
4:36:20target, what I will do is as I told you
4:36:22that I will split this data set into
4:36:24training and test data set. So,
4:36:26basically 80% of the data is going into
4:36:29the training data set and remaining 20%
4:36:32I'm taking as a test data set.
4:36:34Okay?
4:36:36Now, I will call decision tree
4:36:37classifier.
4:36:38You know, like I want to make a decision
4:36:40tree, and I want to fit this on my
4:36:42training data set and the test data set.
4:36:45And I will start building the decision
4:36:48tree and with the help of, you know,
4:36:50from So, here is when I have created the
4:36:53decision tree, and here is when I'm
4:36:55checking the prediction, how accurate my
4:36:57decision tree is. So, I got my
4:37:00precision, which is good. I got my
4:37:01recall. I got my F1 score, and I also
4:37:04get the support, right? So, all these
4:37:06matrices are being used to calculate
4:37:08your
4:37:09how accurate your predictions are. So,
4:37:12higher the precision, better the results
4:37:14would be. Right? And finally, I'm
4:37:16showing to you that how does the tree
4:37:18look like? So, I had tried to plot this
4:37:21decision tree in front of you. So,
4:37:22basically it start with petal length,
4:37:24and you can see the gini index coming up
4:37:26here or the information gain, which is
4:37:28information gain and gini index are
4:37:30reciprocal to each other. If you want to
4:37:31use information gain, you will get that
4:37:33score. If you don't want to use
4:37:35information gain, you will get gini
4:37:36index. So, they are both reciprocal of
4:37:38each other. It is one in the same thing.
4:37:40You use gini index or you use
4:37:42information gain. They are the two
4:37:44different metrics to build your decision
4:37:46tree.
4:37:47So, you can see that
4:37:49based on a petal length, the tree gets
4:37:51splitted like this. Then based on the
4:37:53petal length
4:37:54of different dimensions, your tree
4:37:56further split, then it further split,
4:37:59then it further split, and finally you
4:38:01get to know that whether the particular
4:38:03species will be versicolor or virginica.
4:38:07Okay? So, this is how your entire
4:38:10decision tree is built in Python. Okay?
4:38:14So, this is how it's so simple. Uh I
4:38:16know it takes time to build this thing,
4:38:19but once you are a good data scientist
4:38:21and you understand all these things, it
4:38:23is very simple to build all these things
4:38:26very easily in Python or R.
4:38:29Right? So, now
4:38:31uh finally going back how would you
4:38:34decide how would you decide that which
4:38:36algorithm, you know, which algorithm
4:38:39will be taken when?
4:38:41So, uh what happens is this is on the
4:38:44basis of scikit-learn. So, it it starts
4:38:48something like this, okay? So, it is on
4:38:50the basis of like this that first of all
4:38:52you see here that uh
4:38:56whether how many samples you have, how
4:38:58many data points you have in the data
4:39:00set, right? If you have more than 50
4:39:02samples, then you will go here. If you
4:39:05have less than 50 data point, then you
4:39:07know, you will go here. So, you will get
4:39:10more data set, right? So, if you have
4:39:12more data If you have
4:39:13greater than 50 data point, then you
4:39:15will go further. You split on your data
4:39:17set. If it is not, then you will you
4:39:19just the kind of advice that get more
4:39:21data set, okay? So, then you will decide
4:39:24that what you want to predict. If you
4:39:25have a labeled data set, then you will
4:39:27go here, right? And do clustering. If
4:39:30you don't have the labeled data set,
4:39:31then you will see whether you want to
4:39:33predict a quantity. If yes, you will go
4:39:35in regression. If you want to predict uh
4:39:37if you just want to do exploratory
4:39:39analysis, you can do dimensional
4:39:40reduction. If you have a labeled data
4:39:42set, uh you know, for classification,
4:39:45you can do all this classification. So,
4:39:47basically, this is a cheat sheet which
4:39:49we generally use to decide that what we
4:39:51have to do with the data set and when.
4:39:59Let's understand what is a random
4:40:01forest.
4:40:02A random forest is constructed by using
4:40:05multiple decision trees and the final
4:40:08decision is obtained by majority votes
4:40:11of these decision trees. So, let me make
4:40:13things very simple for you by taking an
4:40:16example. Now, suppose we have got three
4:40:18independent decision trees. Here we are
4:40:20just taking three decision trees and
4:40:22I've got an unknown fruit and I want
4:40:24that these trees would give me a result
4:40:27of what exactly this fruit is. So, I
4:40:29pass this fruit to the first decision
4:40:31tree, the second decision tree, and the
4:40:33third decision tree. Now, a random
4:40:35forest is nothing but a combination of
4:40:37these decision trees. So, the results
4:40:40are being fed into the random forest
4:40:42algorithm. So, what it sees is that,
4:40:45okay, the first decision tree classifies
4:40:47it as peach, the second decision tree
4:40:49says that it is an apple, and the third
4:40:51one says that it is a peach. So, random
4:40:54forest classifier says that, okay, I've
4:40:56got the result as two peach and one for
4:41:01an apple. So, I would say that the
4:41:03unknown fruit is an peach.
4:41:06All right, so this is based on the
4:41:08majority voting of the decision trees
4:41:10and that is how a random forest
4:41:12classifier comes to a decision of
4:41:14predicting the unknown value.
4:41:16Okay, so this was a classification
4:41:18problem, so it took the majority vote.
4:41:20Now, suppose if it was in regression
4:41:22problem, it would have taken mean of it,
4:41:24okay? So, now let's move on further to
4:41:27understanding what is a decision tree.
4:41:29But before that, we should understand
4:41:30that random forest the building blocks
4:41:33are decision trees and that's why
4:41:35studying decision tree becomes important
4:41:37because if we understand one decision
4:41:40tree, we can apply the same concept to
4:41:42random forest, okay? So, now let's move
4:41:45on forward and understand the important
4:41:47terms in random forest. And this will
4:41:49also help us consolidate whatever we
4:41:51have learned so far. So, we have taken
4:41:53the same small decision tree of the
4:41:55previous example, and let's understand
4:41:57these are also the important terms which
4:41:59will be relevant to random forest also.
4:42:01So, the first is the root node. Now,
4:42:03here what happens is that the entire
4:42:05training data has been fed to the root
4:42:07node. And then we've got here that each
4:42:10node will ask either true or false
4:42:12question with respect to one of the
4:42:14feature. And then in response to that
4:42:16question, it will partition the data set
4:42:18into different subsets. That's what it
4:42:20is it is doing here based on the
4:42:23condition that it if the mass body mass
4:42:25is greater than equal to 2500, it ask a
4:42:27question either yes or no. And based on
4:42:30that, again further partition is done.
4:42:32And if not, then it just classifies the
4:42:34species. And then again, what happens is
4:42:37that the splitting Now, this is very
4:42:39important here. The splitting takes
4:42:41place either with the help of a genie or
4:42:43entropy methods. And these helps to
4:42:45decide the optimal split.
4:42:48And we will be discussing about
4:42:49splitting methods very soon, right?
4:42:51Okay. And then we've got the decision
4:42:53nodes which provide the link to the leaf
4:42:56nodes. And these are really important
4:42:57because then only the leaf nodes will
4:43:00tell us what actually the real
4:43:02predictions are to which class does the
4:43:05species belong. So, now coming to the
4:43:08leaf node and these are the end points
4:43:10where no further division will take
4:43:11place and we will obtain our
4:43:13predictions. Okay? So, now coming up to
4:43:17another important thing here is working
4:43:19of random forest. So, now for working of
4:43:22random forest, we will have to
4:43:23understand a few important concepts like
4:43:26random sampling with replacement,
4:43:28feature selection, and also the ensemble
4:43:30technique which is used in random forest
4:43:33and that is bootstrap aggregation which
4:43:35is also known as bagging. So, we will
4:43:37understand this with the help of an
4:43:39example which will be very simple. And
4:43:42then we will go on understanding how
4:43:44feature selection is done in both the
4:43:46classification and the regression
4:43:48problem. Actually, how random forest
4:43:50select features for the construction of
4:43:53decision trees. Well, in random forest,
4:43:55the best split is chosen based on Gini
4:43:57impurity or information gain methods.
4:44:00So, this also we will understand. Now,
4:44:03let us first understand random sampling
4:44:05with replacement. Now, what happens here
4:44:07is that we have got a small subset of
4:44:09the same penguin data set, wherein we
4:44:11have got some six rows and four
4:44:14features, that means four columns. And
4:44:16the arrows that you can see is that now
4:44:18we will be creating three subsets from
4:44:21this small subset, right? And these
4:44:23three subsets will become our decision
4:44:25trees. And then we'll be constructing
4:44:27decision trees from these subsets. So,
4:44:29let us create our first subset. And you
4:44:32can see here that the subset is randomly
4:44:34being created. And for convenience'
4:44:36sake, let me just also show you the
4:44:39different subsets here. Okay. So, now
4:44:41for better understanding, let us
4:44:42understand this that in the first
4:44:44subset, if we focus, we've got certain
4:44:47random rows here, and we have got
4:44:49certain feature. But we do not know how
4:44:52this feature has been selected. We got
4:44:54island and we got body mass. But in the
4:44:56second subset, we got island and flipper
4:44:59length. And in the third subset, we got
4:45:02body mass and flipper length, right?
4:45:04Now, let's look at the rows. Now, when I
4:45:07am talking about these features, I will
4:45:09say this is feature selection, and
4:45:11remember this term. Now, coming to the
4:45:13second concept, that is random sampling.
4:45:16Now, random sampling is nothing but
4:45:17selecting randomly from your subset. So,
4:45:21I'm selecting randomly certain rows from
4:45:24my subset and creating further subset,
4:45:27okay? So, what is replacement here?
4:45:30Replacement is can be seen here and can
4:45:32be understood with the second subset. We
4:45:35see here that the Gentoo species, this
4:45:37is being repeated again. And this is
4:45:40replacement. That means that when we are
4:45:43working with repeated rows, and this row
4:45:46can be repeated again in the second or
4:45:49the third subset, then this is random
4:45:51sampling with replacement. That means my
4:45:53random forest can use a row multiple
4:45:56times in multiple decision trees, right?
4:45:59So, this is the basic concept of random
4:46:01sampling with replacement and feature
4:46:03selection in random forest. Another
4:46:06important term which I would like to
4:46:08bring into the notice is that
4:46:10when we are working with these type of
4:46:13small subsets, these are also known as a
4:46:16bootstrap data sets. And when we
4:46:18aggregate the results of all these data
4:46:20set, it becomes bootstrap aggregation.
4:46:23So, just filling in the gaps so that
4:46:25later on the concepts become more clear.
4:46:28So, now let's move on to drawing
4:46:30decision trees of these subsets, okay?
4:46:33So, let's draw the decision tree of the
4:46:35first subset. Again, we are taking body
4:46:38mass as the first root node, and then
4:46:40based on a decision like if the mass is
4:46:43greater than equal to 3,500, then take a
4:46:45decision either yes or no. If it is no,
4:46:47then the specie is chinstrap. And if it
4:46:50is yes, then again you partition based
4:46:52on island. And if it is Torgersen, then
4:46:55it is Adélie. And if it is Biscoe, then
4:46:58it is Gentoo species. Okay? So, this is
4:47:01how we will construct two more decision
4:47:03trees of the remaining subsets. So, in
4:47:05the second subset, let us just again
4:47:07create decision tree. And here now we
4:47:10are taking flipper length, and then
4:47:12based on a condition that if the flipper
4:47:15length is greater than equal to 190,
4:47:17then make a split. If it is yes, then
4:47:19the specie become Gentoo. And if it is
4:47:22no, that means again make a decision
4:47:25based on island. And if it is Torgersen,
4:47:27it is Adélie. And if the island is Dream
4:47:31Island, then it is a chinstrap species.
4:47:32So, this is how the decision tree of the
4:47:35second subset has been created and this
4:47:37is how it will take decisions, right?
4:47:40Based on the tree length, depth, and
4:47:42also the features it is selecting, okay?
4:47:46So, now let's create the third decision
4:47:47tree of the third subset and we get a
4:47:49decision tree something like this
4:47:51wherein body mass if it is greater than
4:47:534,000 and if it is yes, then clearly it
4:47:56is a Gentoo species and if it is no,
4:47:59then again make a partition with the
4:48:01with respect to flipper length, another
4:48:03feature here, and then if it is again
4:48:06greater than equal to 190, then the
4:48:08species would be Adélie, else it would
4:48:10be chinstrap. So, this is how decision
4:48:13tree three will make a decision.
4:48:15Now, let's just keep these decision
4:48:17trees with us, okay? And we will make
4:48:20sense of these trees just in a while.
4:48:23Okay?
4:48:24But before that, let us understand how
4:48:26feature selection is done in a random
4:48:28forest. How am I selecting the columns?
4:48:31So, for classification, by default the
4:48:33feature selection is taken as the square
4:48:35root of total number of all the
4:48:36features. Now, suppose I've got here
4:48:39four features, so it is a classification
4:48:41problem, I will take the square root of
4:48:43these four features, which becomes two.
4:48:45So, decision tree would be constructed
4:48:46based on two features each. If suppose I
4:48:49had 16 features, then it would be square
4:48:51root of 16, that would be four. So, four
4:48:53features would be taken in each decision
4:48:55tree. All right? And suppose if this
4:48:58would have been a regression problem,
4:49:00then by default what would happen? The
4:49:01features would be selected by taking the
4:49:04total number of features and dividing
4:49:06them by three, okay? So, this is how by
4:49:08default the feature selection is being
4:49:10done by a random forest. Okay, now let
4:49:13us move on forward to consolidating our
4:49:15learning.
4:49:16So, now we are coming to ensemble
4:49:18techniques, that is also known as
4:49:20bootstrap aggregation.
4:49:22Random forest uses ensemble techniques.
4:49:25And what is ensembling? It just means
4:49:27that you're aggregating the result of
4:49:29the decision trees and taking the
4:49:31majority vote in case of classification
4:49:34and the mean in case of regression
4:49:35problems and giving the output. Okay. So
4:49:38now we have again plotted all our
4:49:41decision trees here. And below we can
4:49:44see that there's an unknown data and I
4:49:47want to predict the species of this
4:49:48data. So what will happen is that again
4:49:51let us just feed this problem to each of
4:49:54the decision trees. And let's see what
4:49:57each decision tree makes the prediction.
4:49:59So I just feed this unknown data to
4:50:01decision tree one and it says that okay
4:50:03the species seems to be chinstrap. Okay.
4:50:06And then decision tree two says that
4:50:08based on the data it has been found that
4:50:11the species Adélie. And then decision
4:50:14tree three says that no I I with my
4:50:16decision tree this species is chinstrap.
4:50:19Okay. Now all these data is being fed to
4:50:23random forest classifier. And it says
4:50:25that okay for chinstrap I've got two
4:50:28votes for Adélie it's got one vote. So
4:50:31the new species would be chinstrap,
4:50:33right? So this is how the bootstrap
4:50:35aggregation is done based on the
4:50:38majority voting and the decisions taken
4:50:41by different decision trees they have
4:50:43been combined together aggregated and we
4:50:46get an ensemble result in the random
4:50:49forest. Okay. So this was very simple
4:50:51concept of ensemble techniques which has
4:50:53been used in random forest.
4:50:56Okay. So now let's move on forward to
4:50:58splitting methods. So what are the
4:51:00splitting methods that we use in random
4:51:02forest? So splitting methods are many
4:51:04like Gini impurity, information gain or
4:51:07chi-square. So let's discuss about Gini
4:51:09impurity. So Gini impurity is nothing
4:51:12but it is used to predict the likelihood
4:51:14that a randomly selected example would
4:51:16be incorrectly classified by a specific
4:51:19node. And it is called impurity metric
4:51:21because it shows how the model differs
4:51:24from a pure division, right? And another
4:51:26interesting fact about Gini impurity is
4:51:28that the impurity ranges from zero to
4:51:31one with zero indicating that all of the
4:51:33elements belong to a single class and
4:51:36one indicates that only one class exist.
4:51:39Now value which is like 0.5, this
4:51:42indicates that the elements they are
4:51:44uniformly distributed across some
4:51:46classes, right? Now moving on forward to
4:51:49information gain. Now this is another
4:51:51method which random forest can use and
4:51:54information gain utilizes entropy. So
4:51:57entropy is nothing but it is a measure
4:51:59of uncertainty. So information gain
4:52:01let's talk about that first. So the
4:52:04features they are selected that provide
4:52:06most of the information about a class,
4:52:08right? And this utilizes the entropy
4:52:10concept. So let's see what is entropy.
4:52:14This is a measure of randomness or
4:52:16uncertainty in the data, right? So we
4:52:19will understand this entropy with the
4:52:20help of a small example. So don't worry
4:52:22about it. So let's understand this
4:52:24entropy. Now suppose there's a fruit
4:52:26fruit tray with four different fruits,
4:52:28right? And what do you feel about the
4:52:31entropy here? That means the randomness
4:52:33of the data. Is it really easy to
4:52:36classify these fruits into the
4:52:38respective class? So this becomes really
4:52:40uncertain and the data looks messy here.
4:52:43But what if we just split here these
4:52:45into two trays where in the first tray
4:52:48would have peaches and oranges and the
4:52:50second tray will have apples and lemons.
4:52:53So now this becomes a little more
4:52:55certain. We get low randomness here and
4:52:58this is called as low entropy. So when
4:53:01we move down from tree, that means from
4:53:03root node to the leaf nodes, the entropy
4:53:06reduces and we can also calculate
4:53:09information gain from this entropy. That
4:53:12is the difference in entropy before and
4:53:14after to split that is known as
4:53:16information gain. Okay? So, once we move
4:53:19down the tree and start reducing the
4:53:21randomness from the data, the entropy
4:53:23becomes lower and that is what we want
4:53:26in our data. If there's low entropy,
4:53:28that means we are likely that the
4:53:30predictions would be more accurate and
4:53:33we can make predictions very easily as
4:53:35compared to very messy data which has
4:53:38high entropy. Okay? So, that was about
4:53:40entropy and now let us just move on to
4:53:44the practical demonstration or a
4:53:45hands-on on random forest.
4:53:48Okay. So, now it's time for a hands-on
4:53:50on random forest. So, let us just import
4:53:53a few basic libraries of Python in our
4:53:55Jupyter notebook. And we will run this.
4:53:58We will import pandas as pd, numpy as
4:54:00np, and seaborn as sns. Now, seaborn is
4:54:04needed here because we want to load a
4:54:05data set, that is a penguins data set
4:54:08with the help of seaborn. And this has
4:54:10already been preloaded in seaborn. This
4:54:12is already loaded data set and seaborn
4:54:15has got multiple data sets, you know,
4:54:16for practice for beginners. So, it is a
4:54:18good way to practice for data sets. Now,
4:54:21we can see this asterisk sign that means
4:54:23it is telling us to wait. So, let us
4:54:25just let it get loaded. So, we get got
4:54:27our data in an object called df and we
4:54:29can see the first five entries here. And
4:54:32this uh data frame is shown in the form
4:54:34of a table, rows and columns. And we see
4:54:37here some species, island, bill length,
4:54:39bill depth, flipper length, body mass,
4:54:41and the sex of the penguin. So, our task
4:54:43is to specify or to classify these
4:54:46species of penguins into their
4:54:48respective correct species, right? So,
4:54:51we see the shape of our data and we see
4:54:53that it is like 344 rows and seven
4:54:55columns.
4:54:56And we will see the info. So, we see
4:54:58df.info and this gives us, along with
4:55:01the non-null count, we also get the data
4:55:04type of the values. So, we have got
4:55:07species, island as the object data type,
4:55:09whereas the bill length, bill depth,
4:55:11flipper length, and body mass are in
4:55:13floating point, or or you can say
4:55:15floating data type. And the sex is in
4:55:18object data type, right? So, now moving
4:55:20on forward to calculating how many null
4:55:22values are there with the help of
4:55:24df.isnull.sum.
4:55:26So, we get certain like some around two
4:55:29null values in all these columns, as you
4:55:32can see the features like bill length,
4:55:34bill depth, flipper length, and body
4:55:36mass. Whereas there are 11 null values
4:55:38in sex feature, right? So, what we do is
4:55:40since they are very small null values,
4:55:42we can just drop it, or you can also
4:55:44ignore them. So, here in this data
4:55:46frame, what I'm doing is I'm just
4:55:47dropping these null values, and let us
4:55:50just check whether they are they are
4:55:51being dropped or not with the help of
4:55:53again the same function {dot} is null
4:55:55{dot} sum. And then we see that yes,
4:55:58they are being dropped from our data
4:55:59frame. Now, let us do some feature
4:56:01engineering with our data. Now, we have
4:56:03seen that we have got some object data
4:56:05type in our data frame. And before
4:56:07feeding it into algorithm that is random
4:56:10forest, we have to transform the
4:56:13categorical data or the object data type
4:56:15into the numeric. So, we are using here
4:56:17one-hot encoding to convert the
4:56:19categorical data into numeric. Now,
4:56:21there are various ways in Python which
4:56:22we can do that, like one-hot encoding or
4:56:25you can also use mapping function in
4:56:27Python, but here we are using one-hot
4:56:29encoding. So, let us just do that.
4:56:31And we find here first of all, let us
4:56:33apply it on the sex column. And here we
4:56:36see that we have got two unique values
4:56:37in sex, that is male and female. And we
4:56:40use pandas here to get dummies, that is
4:56:42how we will apply this one-hot encoding
4:56:45because this is how get dummies work.
4:56:47So, what happens is here is that the new
4:56:50unique values are converted into the
4:56:52respective columns in the data frame.
4:56:54So, we see here we have got two unique
4:56:56values, males and females, and they are
4:56:58being converted into the columns. Okay?
4:57:00So, one thing to note here is that we
4:57:03also get a problem of dummy trap because
4:57:06here we see only two unique values. Now,
4:57:08suppose if I had six or seven unique
4:57:10values and I do this one hot encoding, I
4:57:13would have lots of features in my data
4:57:16frame and that would lead to several
4:57:18complexities. So, what I do is
4:57:21to keep things simple, I can use one hot
4:57:23encoding when my data frame or my unique
4:57:25counts are low, when my unique values
4:57:27are less. So, since I had just two or
4:57:30three, I can use it. So, I'm using here.
4:57:33So, what I do is again, now one row, one
4:57:35column as we can see here that it is
4:57:37redundant, giving me extra information,
4:57:39so I will just drop it. So, I drop this
4:57:42first column and what I get in this data
4:57:44frame is only male. So, let us just
4:57:47infer whether I can also infer females
4:57:49from this or not. So, if the value is
4:57:51one, that means the penguin is a male
4:57:53and if the value is zero, that means the
4:57:55penguin is a female. Okay? So, only one
4:57:58column is needed for this data frame.
4:58:01So, I just kept one and dropped the
4:58:02other one. Okay, now apply again one hot
4:58:05encoding to the island feature. So, in
4:58:08island if we check the unique values,
4:58:10we've got three unique values here.
4:58:12Torgersen, Biscoe and Dream Island and
4:58:14the object is the data type, right? So,
4:58:17again we will use pandas, pd.get_dummies
4:58:21and we will use apply it on the feature
4:58:24island and let's get the head of it. So,
4:58:27we get here again the unique values were
4:58:29converted into columns and we get here
4:58:31expected three columns. And then again
4:58:33we will just drop the first column to
4:58:35get the remaining two columns. So, here
4:58:38also we can infer that if the island is
4:58:40Torgersen, if it is one, then it is not
4:58:43Dream, neither Biscoe, right? So, this
4:58:45is how we can read it from the data
4:58:46frame and understand that. Now, remember
4:58:48this thing that these two island and
4:58:50here sex, these are two independent data
4:58:53frames. These are not yet included in
4:58:55the main data frame. So, what we will do
4:58:57now is we will concatenate the above two
4:58:59data frames into the original data
4:59:01frame. So, what we do, we again create a
4:59:03new data frame that is new data and let
4:59:05us just concat with the help of
4:59:06pd.concat function and we will concat
4:59:09what? df.island and sex. And axis is one
4:59:13that means in the column. Okay. So, when
4:59:15we will run this, let's see the head of
4:59:17it. So, everything gets concatenated in
4:59:20a single data frame which is good for
4:59:22the feeding this data into or splitting
4:59:24the data into test and train data. So,
4:59:26now we have this new data frame and
4:59:29we've got some repeated columns here
4:59:31which needs to be deleted. So, what we
4:59:33do is we will delete sex and island here
4:59:36which are just repeating because we've
4:59:37got here male and we have also got here
4:59:40dream and togerson. So, we do not
4:59:42require this island column neither this
4:59:44sex. So, we just drop it with the help
4:59:46of new data.drop and the column names x
4:59:49is one in place equals to true, right?
4:59:51And let's see the head of this data
4:59:53frame. Head of the data frame gives me
4:59:55five unique values, right? And now it is
4:59:58time to create a separate target
5:00:00variable. And what we'll do is we will
5:00:03store in a variable called y only
5:00:05species. So, what we do is from this new
5:00:08data.species, we will just store the
5:00:10species in this y. And we see this
5:00:12y.head that is the first five species
5:00:15and we got the values here. That means
5:00:17another target variable has been created
5:00:19now.
5:00:20So, and you can also see the y.unique
5:00:23values as Adelie, Chinstrap and Gentoo.
5:00:25So, now we see here three unique values
5:00:28of the penguin that is Chinstrap, Adelie
5:00:30and Gentoo. And the data type is object
5:00:32here. So, again we need to convert this
5:00:34object into the numeric data type. So,
5:00:36now what we are doing is we are using
5:00:38the map function in Python and what we
5:00:40do is we map Adelie to zero, Chinstrap
5:00:42to one and Gentoo to two. So, this is
5:00:45how we see then all the values are being
5:00:48mapped to numeric. This is another way
5:00:50to convert a categorical value into a
5:00:52numeric value in Python. Now, what we do
5:00:54is let us just drop the target value
5:00:57species from our main data frame. So, we
5:00:59just drop it and let's see our new data
5:01:01frame. So, we see that we don't have any
5:01:04target species here, right? Okay. So, in
5:01:07X, let's store this new data and perform
5:01:10the splitting of the data. So, what we
5:01:13do is from sklearn.model_selection,
5:01:15we will import our train_test_split
5:01:18and we will split our training data into
5:01:2070% and 30%. So, test data becomes 30%
5:01:23and training data is some 70%. And this
KNN Algorithm
5:01:26random state is zero, which means that
5:01:28I'm not fixing any random state. And
5:01:31this is also useful for the code
5:01:33reproducibility. Now, suppose if I again
5:01:35run this code, I will get the same
5:01:36result. It will not change. You can set
5:01:39this random state to any of the random
5:01:40number as per your choice and result
5:01:43would differ. Okay. So, now let us print
5:01:45the shape of X_train, Y_train, X_test
5:01:48and Y_test. So, we see here that it has
5:01:50been splitted into 70 and 30% and we get
5:01:53X_train as 233 values here and seven
5:01:56features. And X_test has 100 values and
5:01:59seven features. Similarly, Y_train you
5:02:01can see 233 values and Y_test has 100
5:02:04values. That means the species. Okay.
5:02:07So, that has been perfectly splitted
5:02:09into 70 and 30%. Now, what we do is we
5:02:12will train the random forest classifier
5:02:14on the training set. How do we do it? We
5:02:17will import the random forest classifier
5:02:19from sklearn.ensemble.
5:02:21So, we've already dealt with what is
5:02:23ensemble. And then in classifier, we
5:02:26will store this random forest and this
5:02:28n_estimators is nothing but decision
5:02:30tree. So, we are creating some five
5:02:31decision trees here. And the criteria is
5:02:33entropy. And again, random state is set
5:02:36to zero. So, let's see. And then we will
5:02:38fit this X_train and Y_train. So, this
5:02:40has been fitted and the criteria is
5:02:42entropy here. All right. So, now let's
5:02:44make some predictions and let's create a
5:02:46variable called Y_predict and we will
5:02:49just predict it on X_test. And we've
5:02:51also printed this Y prediction and now
5:02:54let's bring the confusion matrix to
5:02:56check the accuracy of random forest
5:02:58algorithm. And what we do is from
5:03:00matrices as sklearn matrices, we will
5:03:02import classification report and
5:03:04confusion matrix and also the accuracy
5:03:06score. So, we will just import them and
5:03:09then in CM variable, we will print the
5:03:12confusion of Y test and Y predictions.
5:03:15So, we will print it and we see here the
5:03:17accuracy score also, which is 98%. So,
5:03:21our random forest classifier is giving
5:03:23us a very good accuracy of 98% and you
5:03:26can see your confusion matrix that only
5:03:27two cases have been misclassified. Rest
5:03:30all the cases have been correctly
5:03:32classified by random forest classifier.
5:03:34Okay. So, now let's move on to printing
5:03:36the classification report of Y test and
5:03:38Y prediction. Let's see and we get the
5:03:41precision as 96%. That means
5:03:43the two predictions by the algorithm is
5:03:4596%. The recall or the true prediction
5:03:48rate is 100% which is very nice and Evan
5:03:51score is also good which is 98%. So,
5:03:54this is giving us a good result. But
5:03:56what if if we change the criteria from
5:03:58entropy to gini? So, let's just
5:04:00experiment with that too. So, let's try
5:04:03this with the different number of trees
5:04:04and change the criteria to gini
5:04:07coefficient. So, now again from
5:04:09sklearn.ensemble, we will import random
5:04:11forest classifier and fit it, okay? And
5:04:14here what we are doing is just we are
5:04:16using seven trees. Previously we used
5:04:18five and now in the criteria, we will
5:04:20use gini coefficient and random state is
5:04:23zero. So, let's run this and see whether
5:04:25there's a change in accuracy or not and
5:04:27let's predict this and let's check the
5:04:30accuracy score. What is the accuracy
5:04:32score for this random forest classifier
5:04:34with seven trees? So, we get 99%
5:04:36accuracy with changing the criteria and
5:04:39changing the number of trees. So, you
5:04:40can just experiment with different
5:04:42number of trees and different number of
5:04:44decision trees. Let's just experiment
5:04:45with, you know, 12 decision trees and
5:04:48see what happens.
5:04:50So, you can see the accuracy reduced to
5:04:5198%. Okay? With seven, we were getting
5:04:5599. So, let's just keep seven because it
5:04:57is giving us really good accuracy. So,
5:04:59this is about random forest classifier
5:05:02and how it works with several trees and
5:05:05different criteria to give us very good
5:05:07accuracy on our training and test data.
5:05:15>> [music]
5:05:15>> The case that we are discussing is
5:05:19basically the KNN algorithm, which is
5:05:22an algorithm which we used for mostly
5:05:25machine
5:05:26learning, which is where you have
5:05:28labeled data.
5:05:30Okay, so
5:05:33KNN algorithm. KNN algorithm is K
5:05:36nearest neighbors algorithm.
5:05:38And it is an example of supervised
5:05:40learning algorithm where
5:05:42basically, you try to classify a new
5:05:45data point based on the neighbors of
5:05:48that data point, which is basically
5:05:50which data points are closer to it. For
5:05:53example, here, as you can see,
5:05:56you have on one side couple of cats and
5:05:59on the other side you have couple of On
5:06:02one side you have dogs
5:06:04and on the other side you have cats.
5:06:07Right? Now, if a new data's point is
5:06:11given to us, there is a
5:06:13picture of a new animal.
5:06:15And
5:06:17if it is lying somewhere here, right?
5:06:20Then, we know that it is nearer to the
5:06:23cats, right? And therefore, we will
5:06:24classify it as cat.
5:06:26Whereas,
5:06:27if it is
5:06:29sort of
5:06:30uh
5:06:31nearer to the dogs, then we classify it
5:06:33as dog, right? So, that's the like, you
5:06:36know, neighborhood for the dog. And
5:06:38therefore, we sort of uh classify it
5:06:42that new animal or the new picture as as
5:06:45being
5:06:46all that of a dog. And this is quite
5:06:50you know, this is something which is
5:06:52even seen our in our regular day-to-day
5:06:54life, right? You know, we have had
5:06:57examples where our parents keep telling
5:06:59us, "Okay, don't play with those kinds
5:07:02of you know, children or something
5:07:04because they are not good in their
5:07:06studies or probably they are not
5:07:08so good in their behavior because you
5:07:10would become like them, right?" So, it's
5:07:12again an example from real life of
5:07:14classifying a particular person based on
5:07:16the company that they keep, right? Or
5:07:18from the
5:07:20with the kind of people that they are.
5:07:22So, that's that's sort of
5:07:26the example and now let me actually go
5:07:28back.
5:07:30Okay.
5:07:31What are the features of of K nearest
5:07:33neighbors algorithm?
5:07:36Okay, so
5:07:37let's talk about
5:07:39the features of KNN. So, as as I said,
5:07:42KNN is a supervised learning algorithm.
5:07:45Typically, it is you know,
5:07:47used for supervised learning kind of
5:07:49problems.
5:07:50It's very simple as we mentioned.
5:07:52Intuitively, you can you know, it's
5:07:55about you know, what kind of neighbors
5:07:56do you have? So, your class is predicted
5:07:59based on the your nearest neighbors as
5:08:01the name suggests. And then it's a
5:08:04non-parametric technique. So, I would
5:08:06like to spend a couple of minutes here
5:08:09to discuss about what we mean by by
5:08:12non-parametric. So, typically, you know,
5:08:15the supervised machine learning
5:08:16algorithms
5:08:18are of two kinds, right? One is the
5:08:20parametric types and the second one is
5:08:22the non-parametric type. When we say
5:08:24parametric, what we mean is basically
5:08:27that the machine or the algorithm
5:08:29assumes
5:08:31that there is an underlying function or
5:08:34a distribution that is known of the
5:08:36particular data set. Like for example,
5:08:39the linear regression would assume that
5:08:40the relationship between
5:08:43two two things X and Y is is linear in
5:08:46nature, right?
5:08:47Um and and and similarly for a
5:08:50distribution a Gaussian distribution it
5:08:52would assume a normal distribution of
5:08:54the data points and so on. So,
5:08:57a lot of the supervised machine learning
5:08:59algorithms they assume some kind of
5:09:01function association between the
5:09:04predictor which is basically the root
5:09:06causes which help you in predicting and
5:09:09the variable that you're trying to
5:09:10predict.
5:09:12Whereas there are certain algorithms
5:09:14like the KNN
5:09:15or the Parzen window or the linear
5:09:18discriminant analysis which is of the
5:09:20kind which is called non-parametric
5:09:22because it does not assume
5:09:24any particular kind of distribution or
5:09:26any particular kind of functional
5:09:28relationship of the data that you're
5:09:30trying to you know predict or you're
5:09:33trying to
5:09:34learn the the pattern of.
5:09:37So, KNN is one of the kind as I said
5:09:39Parzen window is another one
5:09:42uh which is basically where you uh in in
5:09:46Parzen window essentially, you know, the
5:09:49the volume of the data set uh or the
5:09:53area that the data set covers is known
5:09:55whereas you're trying to find the K
5:09:57there where uh which is the number of
5:09:59data points within that area or the
5:10:02volume, right? Whereas in case of
5:10:03non-parametric technique like KNN
5:10:06it's the opposite where the K is known
5:10:09which is basically you would like to
5:10:12associate the class of of
5:10:14the data point that you're trying to
5:10:15predict based on the K number of
5:10:19neighbors around it. Now, K can be
5:10:21three, four, five, whatever, right? 10,
5:10:2420, and so on.
5:10:26Essentially, the difference between
5:10:27Parzen window and KNN is that in KNN you
5:10:31already know the K and then from the K
5:10:33you try to find out the volume and
5:10:35therefore then you try to find the the
5:10:37probability of the density underlying
5:10:39distribution.
5:10:40And then there is the the discriminant
5:10:43analysis which is off again two kinds
5:10:45linear discriminant and multiple
5:10:46discriminant analysis where basically
5:10:48what you do is you you transform the
5:10:52underlying data the features into a
5:10:54higher dimension and in such a way that
5:10:57in the new feature space after you have
5:10:59transformed the data
5:11:01you now try to apply a parametric
5:11:03approach like for example you will try
5:11:05to project the features onto a line or
5:11:09if it is on a sub sub space which is
5:11:13higher dimension than a line then
5:11:15essentially it becomes multiple
5:11:16discriminant analysis. So
5:11:18basically
5:11:19um those are the three kinds of
5:11:21non-parametric techniques. So even if
5:11:23you were not able to sort of get the
5:11:25full hang of
5:11:27what these three types are
5:11:29what you need to keep in mind is that
5:11:31non-parametric technique does not assume
5:11:34any kind of distribution or any kind of
5:11:36functional relationship of the
5:11:38underlying data
5:11:39and therefore it gives us a lot of
5:11:40flexibility whereas the parametric
5:11:43techniques they basically assume some
5:11:45kind of functional relationship
5:11:48between the data points
5:11:50or they assume some kind of
5:11:51distribution.
5:11:53So where would you use a non-parametric
5:11:55versus a parametric technique right?
5:11:57So basically you would use
5:11:59non-parametric technique you know where
5:12:02you do not know about the functional
5:12:03relationship that is one
5:12:05second is that you know there is
5:12:09maybe let's say large amount of data and
5:12:12so on. And thirdly
5:12:14non-parametric technique like KNN does
5:12:16not work in very high dimensional data.
5:12:20So you would also use parametric
5:12:22techniques in that case.
5:12:24Whereas if you have you know smaller
5:12:26data sets, you would use, uh, you know,
5:12:29typically non-parametric techniques. And
5:12:32then also, if you know understand that
5:12:34the relationship between the data points
5:12:36might be, for example, linear or
5:12:39something like that, you would use a
5:12:40parametric technique. So, you're
5:12:42assuming that there is some kind of
5:12:43relationship, uh, like a linear
5:12:45relationship or something like that in
5:12:47between the data points.
5:12:48So, that's where you will use the
5:12:50parametric technique. So,
5:12:51again, just to summarize, basically
5:12:53non-parametric is an approach where you
5:12:55do not assume any kind of distribution
5:12:57or functional relationship, whereas
5:12:59parametric assumes a functional
5:13:00relationship or basically a distribution
5:13:02between the data points.
5:13:05The other feature of KNN is that it is a
5:13:07lazy algorithm. So, what lazy algorithm?
5:13:10Actually, in most of the supervised, uh,
5:13:13learning algorithms,
5:13:15uh, you basically train your model on
5:13:17the training data set,
5:13:19then you have your model,
5:13:21and then you apply this particular model
5:13:23to the test data set to then classify or
5:13:26predict.
5:13:27Uh, you know, for example, whether a new
5:13:29image is that of a cat or a dog, right?
5:13:33So, this can be, you know, some kind of
5:13:35algorithm like support vector machine
5:13:37or regression or logistic regression or
5:13:39whatever, right? So,
5:13:40you run this logistic regression or you
5:13:42run this regression or support vector
5:13:45machine on the training data set,
5:13:47and it learns the features or it learns
5:13:50the parameters of the model from that
5:13:52data set, and then applies this learned
5:13:55model on the test data set.
5:13:58Whereas, in case of KNN, actually, there
5:14:01is no training step at all. That's why
5:14:04it's called a lazy algorithm.
5:14:06Because what it does is,
5:14:07at the time that you are actually now
5:14:10want to predict,
5:14:12at that point in time, actually, it will
5:14:13go and it will do all the calculations
5:14:15of the distance of the new data point,
5:14:18like, for example, the new image of the
5:14:20cat or dog, from all the other data
5:14:23points that you have.
5:14:25So, it will calculate all the distances
5:14:27and then will check for the those data
5:14:29points which are or the K data points
5:14:31which are nearest to this.
5:14:33So, that's why it's called the lazy
5:14:35algorithm because nothing happens
5:14:38till the point or no calculations happen
5:14:40till the point you are actually trying
5:14:41to predict something.
5:14:43So, there is no training step involved.
5:14:45Okay, and then
5:14:47it's used for both classification and
5:14:49regression as we just mentioned. So, it
5:14:51can be used to predict the values as
5:14:53well as be able to classify something
5:14:56like, you know, okay, whether it is a
5:14:57cat or dog or if you're trying to, let's
5:15:00say, predict some value
5:15:02some forecast or something that for that
5:15:04also you can use it. And then it is
5:15:06based on feature similarity, which is
5:15:07basically what do we mean by feature
5:15:10similarity? So, feature can be things
5:15:11like if, for example, you know, you are
5:15:14looking at classifying cats versus dog,
5:15:16right? So, is the eyes like a dog? That
5:15:20can be one of the features. Is the what
5:15:22how do the ears look? That can be one of
5:15:24the features.
5:15:25What about the tongue, the face, and so
5:15:27on? So, there can be multiple such
5:15:28features.
5:15:30And how similar it these features are
5:15:33between two data points,
5:15:35which is used basically by the KNN
5:15:38algorithm.
5:15:39And then
5:15:40as I said, there is no training step
5:15:41involved. So, these are the features of
5:15:44KNN algorithm.
5:15:45And therefore, now let's look at
5:15:47actually just some simple examples of
5:15:49how it works.
5:15:52So, as you can see here in this slide,
5:15:54we have
5:15:56two classes of data. So, one is the all
5:15:58these blue data points.
5:16:00And there's another one which is the
5:16:02orange data points.
5:16:04Now, if you have a new data point, which
5:16:06is this
5:16:08pink one here, which class should it
5:16:10belong to?
5:16:12Should it belong to class A or should it
5:16:14belong to class B? So, what you would do
5:16:16is you would actually start calculating
5:16:18the distance of this pink data point
5:16:21from every square or blue triangle data
5:16:24point. And then
5:16:26you will decide you will have to assume
5:16:28a particular K. Let's say
5:16:30K is three.
5:16:32Right? Which is I'm looking at the
5:16:34nearest three data points. And in that
5:16:38case, basically, as you can see if we
5:16:41draw the circle, right?
5:16:42Then we see that two of the nearest data
5:16:46points within that circle is of the
5:16:49square orange kind. So, basically, we
5:16:51will predict that this particular new
5:16:53data point belongs to class A.
5:16:55Whereas, K value was seven, right? As in
5:16:59this particular example now.
5:17:01You would see that four out of the seven
5:17:03is actually of the blue triangle kind.
5:17:05And therefore, we will now classify it
5:17:07as belonging to class B.
5:17:09So, essentially, this prediction
5:17:12changes, as you can see here, depending
5:17:15on what is the K value.
5:17:17So,
5:17:18therefore, the question is what should
5:17:20be the value of K, right? And typically,
5:17:22what happens is you run a trial and
5:17:24error, and basically, you will come up
5:17:27with okay, what is the best K value. But
5:17:29essentially, what one needs to
5:17:31understand is that as the value of K
5:17:35increases, basically, the the partition
5:17:38line starts moving towards becoming more
5:17:41and more linear. So, it starts becoming
5:17:43less flexible, and it starts assuming
5:17:45some kind of a linear dividing line or
5:17:48something like that. So,
5:17:50what happens is that in that as you
5:17:53increase the K,
5:17:54your bias
5:17:56basically, increases, but your variation
5:18:00reduces. So, we know that in
5:18:03classification problems or in machine
5:18:04learning problems, bias and variance are
5:18:07two things that we are trying to manage,
5:18:09right? Bias is basically how close you
5:18:12are to the actual class or how or to the
5:18:15actual
5:18:16value. Whereas variation is how much
5:18:19variability is there in your prediction.
5:18:22So,
5:18:22as the K increases, the bias
5:18:25sort of increases, but the variance
5:18:27reduces. And then it is vice versa. So,
5:18:29if your K decreases, let's say
5:18:31at K is equal to 1, where you are only
5:18:34looking at just one nearest neighbor and
5:18:36then
5:18:37predicting based on that, actually the
5:18:40bias is the least, which means it is the
5:18:42most flexible.
5:18:43K is equal to 1 is the will give you the
5:18:45most flexible sort of demarcating line
5:18:49or function.
5:18:50Whereas the variability will be the
5:18:52maximum.
5:18:54So, that that's the sort of the
5:18:55trade-off.
5:18:57And that's how we actually determine K.
5:19:00So, we have to get a K value in such a
5:19:02way based on trial and error that
5:19:04sort of maximizes our sort of or reduces
5:19:07the bias as well as the variance. And
5:19:10and that's the kind of optimization we
5:19:11are trying to do.
5:19:13Okay, so how do we calculate the
5:19:14distance itself, right? And distance
5:19:17typically can be of many kinds. So, you
5:19:20know, the example here is of the
5:19:21Euclidean distance, but you can have
5:19:23other kinds of distances like Manhattan
5:19:25distance or Mahalanobis distance and you
5:19:28can look up references for other kinds
5:19:31of distances.
5:19:32Now, Euclidean distance is calculated
5:19:35for the point P1 and P2 as given here.
5:19:37Essentially, Euclidean distance is
5:19:40nothing but the you know, square root of
5:19:42sum of the X coordinates of these two
5:19:45points P1 and P2 and then Y coordinate
5:19:48square of
5:19:49of of these two points P1 and P2.
5:19:52And
5:19:53this is just an example of one kind of
5:19:55distance and
5:19:56other kinds of distances like Manhattan
5:19:59or Mahalanobis are also there.
5:20:01And then this calculating this distance
5:20:03becomes quite challenging,
5:20:05especially in cases where
5:20:07you know, you are trying to for example
5:20:09calculate, let's say, how close two
5:20:11LinkedIn profiles are, right? Or trying
5:20:13to classify uh the category of
5:20:16electrocardiogram and so on so forth.
5:20:18So, there we have to bring in more
5:20:20creativity to just decide what kind of
5:20:22distance to use.
5:20:24Okay. Now, let's move ahead. We will now
5:20:27talk of some use cases where KNN can be
5:20:29used and this is an example of how KNN
5:20:32can be used for book recommendation. So,
5:20:34if you have purchased books on Amazon or
5:20:36whatever, right?
5:20:38Some of these recommendations are based
5:20:39on on KNN algorithm. And then, you know,
5:20:43as we said, you know, KNN is like based
5:20:45on features, right? So, maybe let's say
5:20:48what will be the nearest neighbors of a
5:20:50particular book. It can be based on who
5:20:51is the author, what is the topic, and so
5:20:54on so forth.
5:20:55And then there are other use cases like
5:20:58I mentioned. So, for classifying
5:20:59satellite images, for classifying
5:21:01handwritten digits,
5:21:03uh on on image analytics or or for
5:21:07classifying electrocardiograms,
5:21:09um etc. Uh you know, typically KNN can
5:21:12be used.
5:21:13Okay. So, now actually we will get into
5:21:15some hands-on.
5:21:18Okay. So, to start the hands-on session,
5:21:21I'll go to this Jupyter notebook that I
5:21:25already have installed on my system.
5:21:28And I have a certain
5:21:31code written, which uh we will take two
5:21:33examples.
5:21:34Both the examples are based on data sets
5:21:37which are available in the open source.
5:21:39So, you can easily get access to that
5:21:42data.
5:21:43So, what we do is we start by importing
5:21:47the necessary libraries.
5:21:49So, we import pandas, seaborn, numpy,
5:21:54and matplotlib. Basically, pandas and
5:21:56numpy are there for doing the data
5:21:59manipulation and also for storing data
5:22:02as matrices or as arrays and and be able
5:22:05to perform some mathematical procedures
5:22:08on them.
5:22:09And then seaborn is basically used for
5:22:11plotting and matplotlib for plotting as
5:22:13well.
5:22:14And this line here get IPython just
5:22:17helps us to run the images that we'll be
5:22:20creating in line with Jupiter notebook
5:22:22instead of opening up a new window.
5:22:25So let's run this.
5:22:27And what it will do is it will import
5:22:28all these packages for us which we are
5:22:30going to use.
5:22:32And then we will first import the breast
5:22:35cancer data that is available in your
5:22:38scikit-learn datasets.
5:22:40So we import that. And then let's just
5:22:44initialize that data into a variable
5:22:47here called cancer. So cancer here
5:22:49represents all the load the breast
5:22:51cancer data.
5:22:53And now we will let's actually look at
5:22:55what this data is.
5:22:58It's a bunch of attributes in this
5:23:01dictionary here. So you have data, the
5:23:03target which is basically nothing but
5:23:05whether it is a cancer or not. So
5:23:08whether it is malignant or benign.
5:23:09Malignant means it's a bad cancer and
5:23:12benign means well it's just a tumor,
5:23:13it's not cancerous.
5:23:15Target name description, feature names
5:23:18which is basically
5:23:19the features that will tell us whether a
5:23:21particular
5:23:23case belongs to cancerous or or
5:23:25non-cancer or malignant or benign.
5:23:28And then
5:23:29actually let's just print the
5:23:30description of this particular data
5:23:32here.
5:23:34So as you can see we can use this
5:23:36command to print the description here.
5:23:39And then we see that there are 569 data
5:23:42points with about 30 attributes.
5:23:45And these attributes are radius,
5:23:46texture, perimeter, etc.
5:23:48And
5:23:50the the max and min values of those are
5:23:53given here.
5:23:55And then now let's look at some of the
5:23:57feature names.
5:23:59So, these are the feature names, radius,
5:24:03texture, and so on.
5:24:06And now, let's actually set up a data
5:24:08frame of this particular data here
5:24:13using pandas, this function here. So,
5:24:15there are 569
5:24:18data points.
5:24:19And all these are basically your
5:24:21features, as we talked about.
5:24:25And let's look at the target variable,
5:24:28which is nothing but whether it is
5:24:29telling us whether a particular data of
5:24:31point belongs to malignant or benign.
5:24:33So, zero is cancerous and one is
5:24:35non-cancerous.
5:24:37And then,
5:24:38we convert the target into a data frame
5:24:41as well.
5:24:42And then, let's look at the couple of
5:24:45examples of how the data points look
5:24:48like. This is, you know, one row of the
5:24:51data points which with all the several
5:24:53feature values that we have.
5:24:56So, basically, we use this um package
5:25:00called standard scalar from scikit-learn
5:25:02for pre-processing and for standardizing
5:25:04the variables.
5:25:06And we initialize this standard scalar
5:25:09into a variable called scalar.
5:25:11So, standardizing is nothing but, you
5:25:13know, basically, bringing all the
5:25:15samples to essentially the same range,
5:25:18right? Because
5:25:20uh what might happen is some of the data
5:25:22point, like, for example, temperature
5:25:23might be from zero to 100 and some price
5:25:26might be from, let's say, 1,000 to
5:25:30100,000 or whatever, right? So, the
5:25:32absolute values can can lead to some
5:25:34issues with respect to the prediction.
5:25:36Therefore, we have to standardize it or
5:25:37bring it between, let's say, minus one
5:25:40and one. So, and then, a mean of zero,
5:25:42right? So, we have to bring everything
5:25:44to the same scale to be able to compare
5:25:46the samples.
5:25:47So, we first of all, we fit the this
5:25:51standardization um or normalization on
5:25:54the data set we have.
5:25:56And that is we calculate the the
5:25:59variance and and the means and then we
5:26:02actually apply it on the data set to
5:26:04transform it to the actual values.
5:26:07And then
5:26:09if we look at the scale values now, so
5:26:11let's look at the scale values.
5:26:14And this will give an example of the top
5:26:17five rows here.
5:26:18So we can see now the values are between
5:26:20minus one and one.
5:26:22Or rather it is standardized.
5:26:24Essentially with a normal distribution.
5:26:27And then we divide this data into
5:26:30test and train. So basically we will
5:26:34train the model and then we will test it
5:26:36on a separate data set. If you use the
5:26:39same random state, you should be able to
5:26:41get the same result. Otherwise you may
5:26:43get a different result here. And
5:26:44essentially we are keeping the testing
5:26:47size to 30 which means that we are
5:26:48dividing the entire data set into two
5:26:50parts. The train part which is having
5:26:5270% of the data and the test part which
5:26:55is having 30% of the data. And again we
5:26:56are using this package called train test
5:26:59split from the scikit-learn package.
5:27:02So we get the X and the Ys which are
5:27:05basically nothing but your train and the
5:27:09X's are your predictors and Y is your
5:27:11predicted variable whether it is
5:27:13cancerous or not.
5:27:15And then now let's import the K nearest
5:27:17neighbors classifier. This is the actual
5:27:19algorithm
5:27:20which we are importing from scikit-learn
5:27:23package.
5:27:24And now we initialize this particular
5:27:28algorithm.
5:27:30And then we fit it on the
5:27:33data.
5:27:37And some of the parameters as you can
5:27:40see
5:27:41is basically what is the leaf size and
5:27:42so on so forth.
5:27:44The nearest N neighbors we are taking.
5:27:46So we are taking K is equal to one here
5:27:48basically as of now. We will see the
5:27:50results based on that and then we will
5:27:52change it and see how the results vary.
5:27:56And we now
5:27:59run it on the We now try to predict it.
5:28:03And then we will now try to evaluate
5:28:06what the results look like.
5:28:08So, we have imported the classification
5:28:10report and confusion matrix,
5:28:12which is basically trying to see whether
5:28:15we were able to correctly classify the
5:28:18cancerous as cancerous and non-cancerous
5:28:20as non-cancerous as or not.
5:28:22So, we can see that this is the actual
5:28:24and this is the predicted. So, basically
5:28:26some data points here five and four are
5:28:29classified wrongly, otherwise all the
5:28:31others are classified well. So, if we
5:28:33look at the accuracy calculated accuracy
5:28:37actually, so
5:28:39we see that the precision, which is true
5:28:42alarm, right? Which is basically from
5:28:44the cancerous
5:28:45samples, how many were you able to
5:28:47actually predict as cancerous?
5:28:49If we see the accuracy is quite high,
5:28:50almost 94 95%.
5:28:54And then the recall, which is from all
5:28:57of the cancerous samples, how many were
5:28:59you able to actually predict accurately
5:29:01is about again 94 95%. And F1 score is
5:29:04nothing but a combination of both
5:29:06precision as well as recall. And that's
5:29:08quite good as well. So, with K is equal
5:29:10to one, you're able to get some already
5:29:12some good results. Now, let's try to see
5:29:15how to choose the K value, right? So,
5:29:17this is basically nothing but a
5:29:20a bunch of code that actually runs the K
5:29:23value from one to 40 and then tries to
5:29:26check the accuracy.
5:29:28And this is just like doing a trial and
5:29:30error to see where we get the best
5:29:32results so that we can then use the best
5:29:34K value. So, if you can see this
5:29:36particular plot here after we plot the
5:29:38result from the running the trial and
5:29:41error from one to 40, we see that the
5:29:43error actually starts decreasing and
5:29:45somewhere around this K is going to 21,
5:29:48we get the minimum value of error. So,
5:29:50for us, the best K value is
5:29:5221. So, now if we compare the results
5:29:55between K is equal to 1 and K is equal
5:29:57to 21, we we should be able to see the
5:29:59prediction results. So, as you can see,
5:30:00this was the result with K is equal to
5:30:021, which is
5:30:03we get about 94 95% accuracy.
5:30:06And then with K is equal to 21,
5:30:09we will see whether the accuracy
5:30:10improves, right? So, we see that yes,
5:30:12the accuracy has now gone up to almost
5:30:1599%, which is we earlier had nine
5:30:19misclassified data points
5:30:21out of all of the points. And then here
5:30:24we have just two data points which are
5:30:26misclassified from the test data set.
5:30:28So, now this was one example of applying
5:30:31KNN on the cancer data set, which is
5:30:33available freely.
5:30:35And now let's look at another example,
5:30:37which is the Iris data set.
5:30:39And
5:30:40again, available freely as well.
5:30:43So, Iris is a type of flower.
5:30:45And we will see what flower is it. So,
5:30:48just give me a moment here. So, we again
5:30:49start by importing
5:30:51the necessary libraries. And then
5:30:54we'll look at what this Iris data set
5:30:57is. So, the Iris data set comprises of
5:31:0050 samples of three species of Iris
5:31:03flower, which is Iris
5:31:04setosa, Iris virginica, and Iris
5:31:06versicolor. These are
5:31:08the three types of Iris flowers. And if
5:31:10you run this, basically we will find
5:31:12that
5:31:16Okay.
5:31:17Sorry, we did not copy the entire the
5:31:20code here. So,
5:31:21it was giving an issue.
5:31:23Let's just run it again.
5:31:28Okay. So, we see that it is this
5:31:29particular flower, which is Iris setosa.
5:31:32So, the Iris setosa, you can see.
5:31:35And
5:31:36now let's look at the other two kinds of
5:31:38flowers here. So, which is
5:31:40Iris versicolor.
5:31:44So, this is Iris versicolor. And then
5:31:47you have the Iris virginica.
5:31:55So, we see that this one is Iris
5:31:57virginicas. So, essentially we now will
5:32:01import the sort of data set. We have
5:32:04already done that. And we will now use
5:32:06the seaborn package to actually plot
5:32:09some of this data and see
5:32:12how it looks like.
5:32:14Which is basically do some kind of
5:32:15exploratory data analysis. So, if we
5:32:18look at the data itself, so this is the
5:32:20top five rows from the data, right? So,
5:32:22essentially the data consists of
5:32:25basically what is the sepal length,
5:32:27sepal width,
5:32:28petal length, and petal width.
5:32:30And this is nothing but basically your
5:32:33petal is your this colored part of the
5:32:35flower and the sepal is basically your
5:32:37green part here, right? So, it is
5:32:39talking about what what is the sepal
5:32:40length, sepal width, and petal length,
5:32:42and petal width of each of the species,
5:32:44whether it is setosa, virginica, or
5:32:47versicolor. And we have we will see
5:32:49whether we can use KNN to actually
5:32:52classify these
5:32:54these flowers into the data points into
5:32:56these categories of flowers.
5:32:58So, let's do some quick exploratory data
5:33:01analysis
5:33:02on this. So, we are running a pair plot
5:33:05on the data set. And uh
5:33:08in the meantime, I'll just copy another
5:33:11part of the code here.
5:33:13Okay, so now the plot has come up. So,
5:33:16as we can see that the green is
5:33:18basically your setosa flower. And
5:33:22we can see that this pair plot actually
5:33:24just plots
5:33:25the sepal length, sepal width,
5:33:28and petal length, petal width of each of
5:33:30the samples of setosa, verse- color and
5:33:33virginica. And we see that
5:33:34the green dots which are the setosa
5:33:36flower is actually quite separable from
5:33:38the others. It's when you plot
5:33:40let let's say for example sepal length
5:33:42and
5:33:43petal length, right? We see that this is
5:33:45quite separate from the other data
Naive Bayes Classifier
5:33:47points. So, let's see whether you know,
5:33:49we can actually
5:33:50be able to classify it using KNN
5:33:53or not. And here we are running a kernel
5:33:56density estimation function on the
5:33:58setosa flower to check
5:34:01what kind of distribution it has. So,
5:34:03this is the kernel density estimation
5:34:06plot using the SNS package. So, only for
5:34:09the setosa flower. So, if we plot the
5:34:12sepal length and sepal width, we get
5:34:13something distribution like this. So,
5:34:15essentially we see that the maximum
5:34:17centered around here and then there is a
5:34:19distribution as you can see here. So,
5:34:21there's some kind of a linear
5:34:22relationship here.
5:34:24Okay, so now we will again do the same
5:34:27standardization of the variables
5:34:30that we had done in the cancer data set
5:34:32case. So, we are importing the standard
5:34:34scalar function
5:34:36from the scikit-learn preprocessing. So,
5:34:39we will initialize that.
5:34:41So, we are again basically doing the
5:34:42standardization or normalization of the
5:34:45data.
5:34:46And we will do the standardization on
5:34:48everything except the species which is a
5:34:50categorical value, right? So, it's
5:34:52categorical whether it is which kind of
5:34:54flower it is. So, we have removed that
5:34:56and then we have
5:34:58done the standardization or
5:35:00normalization on rest of the data.
5:35:02So, we now convert this into a data
5:35:05frame, pandas data frame. And if we look
5:35:08at the top five rows, now it's all
5:35:09converted or transformed. So, the values
5:35:12are now normally distributed basically.
5:35:15Okay, so now we
5:35:18divide the data again into train and
5:35:20test.
5:35:22And we again have training of about 70%
5:35:27and test data set of about 30%. Um so we
5:35:31are dividing that entire data set into
5:35:32these two buckets.
5:35:34And we will now use KNN
5:35:38to see if we can use KNN to classify
5:35:40them.
5:35:42Again, the same and K is equal to 1.
5:35:46And we will check the results and then
5:35:47we will do a trial and error
5:35:50to check what is the best value of K.
5:35:52So here K is equal to 1.
5:35:57And now we are going to predict on the
5:36:00test data set
5:36:02and look at the results.
5:36:05So we are now importing the
5:36:06classification report and the confusion
5:36:08matrix.
5:36:12So if you look at the confusion matrix,
5:36:18we see that
5:36:19kind of already we are getting quite
5:36:21good
5:36:22uh prediction. So just two misclassified
5:36:24points.
5:36:27And uh if we look at the accuracy,
5:36:32we see that the accuracy is quite high,
5:36:34around 96%
5:36:37already.
5:36:38Now we choose we have to see what is the
5:36:41best value of K.
5:36:43So
5:36:45essentially we will
5:36:49we will again run K is equal to 1 to 40
5:36:52and check which is the best value.
5:36:56So let's plot the errors when we vary
5:36:59the K from 1 to 40. And we see that the
5:37:02actually the error decreases and then
5:37:04increases. So basically the error is
5:37:07minimum with K is equal to let's say
5:37:08three or even five or maybe 11. So let's
5:37:12choose one of these values. So let's say
5:37:14K is equal to three.
5:37:16And let's see how the results look like.
5:37:19Does it improve the
5:37:21accuracy or not? So we now see that
5:37:24even the two data points which are
5:37:25misclassified earlier is now classified
5:37:27properly. So, the accuracy improves to
5:37:29100%.
5:37:31So, that's the
5:37:32example of how you can choose K.
5:37:36>> [music]
5:37:40>> What is Naive Bayes?
5:37:42Let us understand Naive Bayes with an
5:37:44example. Here, I just cannot seem to
5:37:47figure out which are the best days to
5:37:49play football with my friend. Can you
5:37:52please help us out?
5:37:54All possible conditions are given to us.
5:37:56There is
5:37:58summer, monsoon, and winter, which is
5:38:00nothing but the outlook.
5:38:03Am I correct in saying that?
5:38:04Summer, monsoon, and winter is nothing
5:38:06but the outlook. Then we have sunny or
5:38:09not sunny. So, that basically is the
5:38:12humidity, right? And then we have windy
5:38:15or no windy. That speaks about
5:38:18the winds. How are the winds?
5:38:20Right?
5:38:21So, if you look at these combinations,
5:38:24okay, we will look at this using Naive
5:38:25Bayes on how do we decide whether we can
5:38:28play or not.
5:38:29So, if I have noted down all the days it
5:38:31was good, bad to play football, and the
5:38:33combination of weather matrices on that
5:38:35day, that will be perfect, right? That
5:38:37is perfect, and we will be able to do
5:38:39Naive Bayes classifiers using that. Now,
5:38:42Naive Bayes classifier comes from the
5:38:44Naive Bayes theorem, and Naive Bayes
5:38:46theorem is purely and purely based on
5:38:50the assumption of independence.
5:38:53So, what does it mean? When I say
5:38:55independence, what it means is that
5:38:59this variable has no relationship, no
5:39:03association with this variable. Now,
5:39:06when I talk about this in a linear
5:39:09context, obviously in summer, you will
5:39:12see that we have more sunny days.
5:39:16Yes? So, if you look at it from the
5:39:18correlation area, from the linear
5:39:21algebra concepts, linearly these two are
5:39:25correlated to each other. Am I correct
5:39:28in saying that?
5:39:29Obviously, in monsoon we have less sunny
5:39:32days. In winter we further have less
5:39:34sunny days.
5:39:36Yeah?
5:39:36So,
5:39:37although there is a relation,
5:39:39Naive Bayes theorem
5:39:42says that all these variables are
5:39:45independent of each other.
5:39:49What does the that mean? If this is
5:39:51causing any kind of an effect, if this
5:39:54is causing any kind of an impact,
5:39:57this should not matter.
5:40:00Okay? There's going to be no
5:40:01relationship. Summer, monsoon, winter,
5:40:04it has its own weightage. Okay? And it
5:40:07has nothing to do with the other
5:40:09conditions. Every condition is equally
5:40:12significant.
5:40:14All right? So, what happens in Naive
5:40:15Bayes is we estimate the posterior
5:40:18probability of every event happening.
5:40:21Here, we calculate the posterior
5:40:24probability of an event happening.
5:40:27Okay? So, here if you see, if you look
5:40:29at the sunny conditions, what we have
5:40:31done, sunny conditions we have this
5:40:33distribution. That there is no play
5:40:35happening in summer, there is play
5:40:36happening in monsoon, there is play
5:40:38happening in winter.
5:40:39Right? So, base of the season, we are
5:40:42figuring that out. Similarly for windy
5:40:44conditions, we are doing that.
5:40:46Right? Again, then we do it for a
5:40:49combination.
5:40:50Right? Whether when windy conditions are
5:40:53yes and no, what happens to play?
5:40:56Okay? So, here
5:40:57what at the end of the day, what gets
5:41:00selected is the one which has a
5:41:02posterior probability of greater than
5:41:04five. Now, when I talk about posterior
5:41:07probability,
5:41:08what do I mean by posterior probability?
5:41:10Let me have a blank slate. Here you go.
5:41:13Actually, it's given. So, I need not
5:41:15show you that. What is the simplistic
5:41:17probabilistic classifier here? What is
5:41:19the probability of an event A happening
5:41:22given
5:41:23B.
5:41:24So, if you look at our problem context,
5:41:27what is it that we are trying to figure
5:41:29out?
5:41:29We are trying to figure out what is the
5:41:32probability of
5:41:36play happening
5:41:38given
5:41:45the outlook is sunny,
5:41:50{comma}
5:41:53the
5:41:54uh
5:41:57winds
5:42:03are normal,
5:42:07and there is no rain.
5:42:13On any given day,
5:42:15on any given day when there is no rain,
5:42:19there is no wind,
5:42:21and the outlook is sunny,
5:42:23whether play will happen or not. So,
5:42:26what we end up doing is we calculate the
5:42:28posterior probability of play happening
5:42:30given these conditions. We also
5:42:32calculate the posterior probability of
5:42:35play not happening
5:42:38given these conditions, and then we
5:42:40normalize these probabilities.
5:42:43Mathematically, we do all of these
5:42:45calculations to figure out naive Bayes.
5:42:47Okay? Now, here
5:42:49see,
5:42:50here we are talking about one event.
5:42:53Here, we have three events. We have
5:42:55outlook,
5:42:57we have winds,
5:42:59and we have rains.
5:43:00So, what this becomes is
5:43:04probability of three independent events.
5:43:07So, what we will do, we figure out what
5:43:10is the probability of play happening
5:43:16given
5:43:18outlook is sunny.
5:43:22We also figure out
5:43:26what is the probability of
5:43:29play happening.
5:43:34Multiply this with the probability of
5:43:37wind as no.
5:43:41Given no, then probability of play
5:43:44happening
5:43:49given
5:43:51rain is no.
5:43:53We figure out all the in three
5:43:56independent probabilities multiplied.
5:43:59All right? This is what we do
5:44:01mathematically.
5:44:03This is what is done mathematically in
5:44:06Naive Bayes theorem.
5:44:08Ultimately, the posterior probability
5:44:12the posterior probability probability of
5:44:14an event A happening given
5:44:17B conditions is calculated by first
5:44:21calculating the class probability.
5:44:24What is the class probability here?
5:44:25Probability of it raining given play was
5:44:29happening in that day.
5:44:31This is multiplied by the total
5:44:33probability of play happening and
5:44:35divided by the total probability of
5:44:37sunny conditions.
5:44:41Here, as you see, this is the posterior
5:44:43probability.
5:44:44This is the class probability.
5:44:47Here is the predictors probability. This
5:44:49is the outlook. X is the outlook. C is
5:44:53what we are trying to predict.
5:44:55Right?
5:44:56And here is the likelihood.
5:44:59So, how is this formulated into our
5:45:01table? Now, if you correlate this to our
5:45:03graph,
5:45:04what will be the likelihood?
5:45:07What will be the likelihood of sunny
5:45:10conditions given play happens?
5:45:13Sunny conditions play happens.
5:45:152 / 3
5:45:18out of
5:45:19Sorry, how many sunny conditions do we
5:45:21have? Six
5:45:22conditions.
5:45:24In six con-
5:45:25ditions, how many days does play happen?
5:45:27Two days.
5:45:292 / 6 1 / 3. What is the class
5:45:32probability? So, of all the events that
5:45:34are given to us, how many days does play
5:45:36happen?
5:45:37Okay? This is how this is calculated.
5:45:40So, here you see what we have done is
5:45:43Let me go back to the previous slide.
5:45:45Here you go.
5:45:46Okay? Here, what we have done is we have
5:45:48calculated this table. Now, it speaks
5:45:51about a data set. Where can you get this
5:45:54data set?
5:45:55You can look for golf play days data set
5:45:59online. You can look for golf play days
5:46:02data set.
5:46:03Okay?
5:46:04In this data set, you will find all the
5:46:07data
5:46:08which is required for this particular
5:46:11example to be done. So, what I will be
5:46:13doing is I will be doing this example,
5:46:15this Naive Bayes classification, with
5:46:18you in Python. All right? So, whatever
5:46:20calculations are being done here,
5:46:23okay? I will do the same activity in
5:46:26Python with pen and paper. And instead
5:46:28of doing this
5:46:32in a numeric way where I'm doing lot of
5:46:34probabilistic calculations,
5:46:37I will achieve this simply in
5:46:40very limited lines of code.
5:46:43Very limited lines of code with Python.
5:46:47Once I'm I have done that, I will come
5:46:49and explain this prob- probability table
5:46:52to all of us.
5:46:53I'm going to use some basic libraries.
5:46:56All right. So, here, what I will be
5:46:58doing is using some very basic libraries
5:47:01for this activity. All right?
5:47:57Done. Now, let me quickly go and uh
5:48:01read the data set. So, for that what I
5:48:04will do is quickly
5:48:08change my working directory.
5:48:43And now, let me quickly go and read my
5:48:44data. So, my data frame is pd.
5:49:09>> Here you go. This is my data set.
5:49:12Right? So, if you look at this data set,
5:49:14in this data set, you have 13 14 days.
5:49:18In these 14 days, you have the outlook,
5:49:21overcast,
5:49:22rainy, and sunny. You have temperature,
5:49:26temperature is hot, cool, and mild. You
5:49:29have humidity, you have wind, and you
5:49:32have play. Right? So, I will not be
5:49:35using uh
5:49:36Okay, let us use all four. In this
5:49:39example, they're using only three
5:49:40variables, but in our
5:49:43hands-on, okay? In this hands-on, what I
5:49:46will be doing is I will be uh
5:49:49using all four variables. Let us do
5:49:51that.
5:49:52Okay?
5:49:53So, before I do that, let me convert
5:49:56everything into a category.
5:49:58If you look at your data frame right
5:49:59now, it's not everything is not into a
5:50:02categorical variable.
5:50:04Here you go, see.
5:50:05Okay? So, let me quickly go and convert
5:50:07everything into a category.
5:50:26And once I have done this, uh
5:50:29let me create a new data frame in which
5:50:31I have everything as a category code.
5:50:35So, that I have numbers. I'll show you
5:50:37what What do I mean by this?
5:50:52>> So, let me execute this. Here you go.
5:50:55See, now I have two data frames. In the
5:50:57first data frame I have all these
5:50:59values. These are now categorical
5:51:01variables, but in the second data frame
5:51:03I have all ones and zeros. So, wherever
5:51:06you see there is sunny conditions, now I
5:51:08have a code two.
5:51:09Rainy conditions, code one.
5:51:11Similarly, when play happens I have a
5:51:15one. When play does not happen I have a
5:51:17zero.
5:51:19This is what I have done.
5:51:21This data frame is available online.
5:51:24Okay? You can get this data frame
5:51:27online.
5:51:29All right? Now, my data frame is ready.
5:51:32So, now what I'm going to do is I'm
5:51:34going to divide my data frame
5:51:37into training and testing. I have 14
5:51:39records.
5:51:41So, let's take 10 records for training.
5:51:44I will give 10 records as an input.
5:51:47And I will give four records, last four
5:51:49records
5:51:57as my test data frame. All right?
5:52:00Uh
5:52:00so, now I will need to create my X and
5:52:03my Y.
5:52:04So, how I will do that is
5:52:06I'll say Y {underscore} train
5:52:09is equal to from train
5:52:12I don't want the play variable.
5:52:14That play variable should be my Y. As
5:52:16simple as that.
5:52:18And I will say X {underscore} train is
5:52:21equal to train.
5:52:23Okay? And I will do the same thing for
5:52:25my test data frame also.
5:52:29Now, those who are new to Python will
5:52:32find this a little bit strange. Please
5:52:36bear with me.
5:52:37But these are the only calculations,
5:52:39only steps which need to be performed
5:52:41every time.
5:52:42You are trying to achieve maybe bias
5:52:45algorithm or any kind of an algorithm.
5:52:49All right. So, here now you see this is
5:52:52my training data frame in which play
5:52:54variable is not there.
5:52:56Play variable is not there. This is my Y
5:53:00in which only play variable is there.
5:53:02This is my training data set. So, both
5:53:04of them have 10 records with the
5:53:06matching index.
5:53:08Similarly, test data frame four records
5:53:12four records with the matching index.
5:53:15Right? So, that we know which data frame
5:53:18is where.
5:53:20Now, multinomial naive bias. Very
5:53:23simple, three lines of code and my model
5:53:26will be done.
5:53:28Okay? First, I initialize my model.
5:53:32Here you go. I have initialized my
5:53:33model.
5:53:34In this model, I fit my data.
5:53:39In this model, I will fit my data. So,
5:53:42to do that, what I say is fit
5:53:45X underscore
5:53:50train comma
5:53:53comma Y underscore train.
5:53:56Done.
5:53:57Your model object is now ready.
5:54:00And now you can simply get the
5:54:03classification outcomes. So, we have in
5:54:06our
5:54:07test data frame, if you look at our test
5:54:09data frame, this is our test data frame.
5:54:11We have three four conditions. All four
5:54:14are sunny,
5:54:16high temperature, low humidity, and
5:54:19windy. Right? And if you look at their
5:54:22outcomes, these are their outcomes.
5:54:24On the first two days, play is not
5:54:26happening. On the next two days, play is
5:54:28happening. Let us look at what is the
5:54:30prediction of our model for this. So, to
5:54:33do do that, what I simply do is
5:54:37X out is equal to
5:54:40model.predict
5:54:46To this I give my X {underscore} test.
5:54:50Here you go. Okay? Now you have your Y Y
5:54:54out variable, so this is the prediction
5:54:56for all the
5:54:59four inputs that you give, and this is
5:55:00the prediction.
5:55:02First day
5:55:03first day we say play does not happen.
5:55:06Let us match it.
5:55:08Let us try to match it with our
5:55:10here.
5:55:11See?
5:55:12Out of four records, three records we
5:55:14are predicting correctly.
5:55:16Three records we are predicting
5:55:18correctly. If you want to check the
5:55:20accuracy, what is the accuracy of your
5:55:23model? What you can simply do is print
5:55:25Let us print the accuracy on both
5:55:27training and testing.
Support Vector Machine
5:55:34Training accuracy. How do I get the
5:55:36training accuracy? Very simple. model.
5:55:39score
5:55:43And here I give my X {underscore} train
5:55:47{comma} Y {underscore} train.
5:55:50And then we do the
5:55:52testing accuracy also.
5:56:06Here you go. So here you can see for our
5:56:10model training we have 80% accuracy, and
5:56:13for testing we have 75% accuracy. Okay?
5:56:18So this is the advantage of doing this
5:56:20activity in Python. But what is
5:56:22happening in the back end?
5:56:23Now let us go and also understand that
5:56:25in terms of naive Bayes classifier.
5:56:28We have successfully
5:56:30we have successfully implemented the
5:56:33Naive Bayes classifier in Python
5:56:35programming language. But, here let us
5:56:38try to understand Bayes theorem, what is
5:56:41happening. So, from this data set all
5:56:43the tabulated data frequency tables are
5:56:45calculated.
5:56:47Once the frequency tables are
5:56:48calculated, they are substituted in our
5:56:51formula to calculate the probabilistic
5:56:53scores. So, what is the probability of
5:56:56summer given it is playing conditions?
5:57:00Total how many playing conditions are
5:57:02there? Total there are nine playing
5:57:03conditions. Nine days play happened.
5:57:06That becomes our denominator. Out of
5:57:08those days, how many days was summer is
5:57:10our numerator. That is how for this we
5:57:13get a probability of 0.33.
5:57:16Then we calculate the class probability
5:57:18where we look at how many days was it
5:57:20summer? Out of total 14 days, five days
5:57:23was summer, so that's the probability
5:57:25and the class probability is 0.64. Put
5:57:28everything into our equation.
5:57:30Put everything into our equation and
5:57:32this is what we get.
5:57:34Okay? So, we do this for each and every
5:57:37condition.
5:57:38So, here we calculate it for winter.
5:57:42All right? Once we have done it for all
5:57:44three days,
5:57:47winter, sunny and windy days,
5:57:49we substitute those here
5:57:51and that gives us the probability which
5:57:53is
5:57:54more than 0.5. Thus, now we can say that
5:57:58if
5:57:59it is winter, sunny and conditions are
5:58:02sunny and there are winds. Conditions
5:58:04are not sunny and there are winds. Play
5:58:07can happen.
5:58:09Look at another example. If a single
5:58:11card is drawn from a standard deck of
5:58:14playing cards, the probability that card
5:58:16is a king is 4/52
5:58:18since there are four kings in a standard
5:58:21deck.
5:58:22King is the event. This card is a king.
5:58:25This is the event. The prior probability
5:58:28of this is 1 by 13. If evidence is
5:58:31provided, for instance, someone looks at
5:58:33the card that the single card is a face
5:58:35card, then the posterior probability can
5:58:38be calculated using Bayes' theorem.
5:58:47Okay? Since every king is also a face
5:58:49card, the probability of face happening
5:58:52given you getting a face card given it's
5:58:54a king is one. Since there are three
5:58:57face cards in each suit,
5:58:59all right? It's actually four. Ace is
5:59:01also a face card. So, it's jack, king,
5:59:04queen, and uh ace. The probability of
5:59:06the face card is 4 by 13.
5:59:08Okay? So, if you combine these three
5:59:10likelihoods, what you get is 13 by 4.
5:59:13So, using Bayes' theorem, this is the
5:59:15probability that you get.
5:59:18>> [music]
5:59:23>> What is support vector machine?
5:59:25Support vector machine comes under
5:59:27supervised machine learning.
5:59:30And we use it specifically for
5:59:32performing the task of classification.
5:59:35So, support vector machine is a
5:59:36discriminative classifier
5:59:38that is formally designed by a separate
5:59:41hyperplane.
5:59:43Okay? It is a representation of examples
5:59:45as points in a space that are mapped so
5:59:48that the points of different categories
5:59:50are separated by a gap as wide as
5:59:53possible.
5:59:55So, in this case of the support vector
5:59:56machine,
5:59:57let's say I have some data points. So,
6:00:00there are some data points of X, and
6:00:02there are data points of circle.
6:00:05Now, this support vector machine
6:00:07is a type of machine learning algorithm
6:00:10where if I have the collection of
6:00:12points, so here in this data points, I
6:00:14have two classes. One is X, and the
6:00:17another one is circle.
6:00:19Now, given this kind of data points,
6:00:22okay? Given this kind of binary
6:00:23classification problem,
6:00:26so the expectation is
6:00:29in case of support vector machine, I'm
6:00:31going to draw a hyperplane
6:00:33which separates as much as possible.
6:00:37Okay? So, I'm going to draw a hyperplane
6:00:40which separates these two classes as
6:00:43much as possible.
6:00:46So, it says that the I'm going to draw
6:00:49draw hyperplane
6:00:51and it it will be separated by a gap as
6:00:54wide as possible.
6:00:56So, that is the intuition behind support
6:00:58vector machine.
6:01:00Okay. Now that you have an intuition
6:01:02behind what is support vector machine,
6:01:05let's understand as how does this SVM,
6:01:08that is support vector machine, would
6:01:09work.
6:01:10So, in case of support vector machine,
6:01:13so here there is one more example. I
6:01:15have the set of
6:01:17points which is green green color and I
6:01:19have another set of points which are in
6:01:21red color. So, these two
6:01:24points are belonging to the different
6:01:26different classes.
6:01:28Now, what I'm going to do is I'm going
6:01:30to draw a hyperplane which separates
6:01:34these two classes data points as much as
6:01:36possible. And when I'm drawing the
6:01:38hyperplane, I'll make sure that this
6:01:40hyperplane is
6:01:42as this hyperplane is equidistant from
6:01:46my support vectors.
6:01:48Now, the support vectors is nothing but
6:01:51the point which is closer to my
6:01:52hyperplane.
6:01:54Now, here in this example that you're
6:01:55seeing,
6:01:57the this data point and this data point
6:02:01are called as support vectors because
6:02:03these are the data points which are
6:02:05nearest from my hyperplane that I've
6:02:06just drawn.
6:02:09In if I'm trying to make use of this SVM
6:02:12model, it is going to draw this kind of
6:02:14hyperplane
6:02:16to make sure that it is separating two
6:02:18classes. The two classes that we have
6:02:20over here in this example is red and
6:02:22green. It's going to separate these two
6:02:24classes as much as possible and it will
6:02:28be equidistant from my support vectors
6:02:31and the support vectors are nothing but
6:02:33the nearest point to my hyperplane.
6:02:36And that is how I'm going to separate
6:02:39between two classes when it comes to
6:02:40support vector machines.
6:02:44Now, here in this example that you're
6:02:46currently seeing,
6:02:47the hyperplane that I've just drawn, so
6:02:49this is a simple linear hyperplane.
6:02:53Just like a straight line that I'm
6:02:54trying to draw if I want to separate two
6:02:57classes of data points.
6:02:59Now, apart from drawing this straight
6:03:01line, we also have other kind of lines
6:03:05as well which we can draw.
6:03:07So, let's see how we can do that.
6:03:11So, the types of line that we can draw
6:03:13or the hyperplane that we can draw is
6:03:16called as SVM kernels, that is support
6:03:18vector machine kernels. The example that
6:03:20we have seen, it's an example for linear
6:03:23SVM kernels.
6:03:25So, let's see what are the other types
6:03:27of kernels that we have. So, when it
6:03:29comes to SVM kernels, we have linear
6:03:31kernels,
6:03:33radial basis function kernel and along
6:03:35with that, we also have polynomial
6:03:38kernel.
6:03:40Now, in case of linear kernel, I'm going
6:03:42to draw a hyperplane which is like a
6:03:44straight line.
6:03:45In case of polynomial kernel, I can draw
6:03:48my hyperplane on the basis of polynomial
6:03:50function that I have created on the
6:03:52basis of number of variables that I have
6:03:54and the degree that I have over there in
6:03:57case of polynomial. And in case of
6:03:59radial basis function, so I'll make use
6:04:01of radial basis to separate my data
6:04:04points.
6:04:07Okay, so these three are the important
6:04:10kernels that we have in SVM and this is
6:04:12one of the commonly asked interview
6:04:13question when it comes to the topic of
6:04:15support vector machines.
6:04:18Now, let's look at some of the use cases
6:04:21or the way we can do where we can use
6:04:24this SVM to uh
6:04:27work or let's look at some of the use
6:04:29cases where we can use this SVM.
6:04:32Okay.
6:04:34So, we can use this SVM
6:04:38on many of the use cases. So, to name a
6:04:41few, we can use it in face detection.
6:04:44We can use it in text and hypertext
6:04:46categorization. We can use the SVM if
6:04:49I'm trying to classify any images. I can
6:04:52make use in bioinformatics.
6:04:55And if I'm trying to detect something,
6:04:57so I can
6:04:59in the in an example here, remote
6:05:01homology detection, handwriting
6:05:03detection. Or in general, we can make
6:05:06use of this generalized predictive
6:05:08control. So, wherever we are dealing
6:05:10with the task of classification, we can
6:05:13use this SVM model. Okay. Now that we
6:05:17have a theoretical understanding as what
6:05:19is SVM and how it is actually going to
6:05:22look like and how it will be,
6:05:25let's have a quick walk through as how
6:05:28we can implement this SVM. Now, to
6:05:31implement this SVM,
6:05:33these are the common steps that we are
6:05:34going to follow.
6:05:36We are going to load the data.
6:05:39We'll explore the data.
6:05:41And once we have explored the data, we
6:05:43are going to split the data into two
6:05:45parts. The reason is simple. One, I have
6:05:47training, so I'll be making use of my
6:05:50training data.
6:05:51And once my training is complete, I'll
6:05:54check how my model has been trained with
6:05:56the help of my test data. So,
6:05:58I'm going to split the data.
6:06:00Now, once that is complete, we are going
6:06:02to train this SVM model. And finally, we
6:06:05can evaluate the model and observe as
6:06:08how model is working.
6:06:11So, this is the overview of the
6:06:13implementation of support vector
6:06:15machines.
6:06:17So, let's do one thing. Let's
6:06:21work it out and let's create the
6:06:23notebook in Google Colab and let's see
6:06:25it in action as how we can implement
6:06:27this SVM.
6:06:28I'll come back to my Google Colab.
6:06:31So, this is the notebook that I have
6:06:32already prepared and I'll give you a
6:06:35walk-through as we proceed along.
6:06:37Now, here in my first cell, I'm
6:06:39importing my NumPy library, Pandas
6:06:41library, and along with that, for
6:06:43creation of plots, I'm importing my
6:06:45Matplotlib library. Now, if you're
6:06:48comfortable with Seaborn, you can use
6:06:50the Seaborn library as well. So, in my
6:06:52example, I'm just making use of
6:06:54Matplotlib because we are not interested
6:06:57in creation of visualization, but we
6:06:59want to understand as how model is being
6:07:02working.
6:07:05Okay.
6:07:06And I'm going to execute this cell.
6:07:09So, this is going to take care of
6:07:10necessary imports. I'm importing my
6:07:12necessary libraries.
6:07:14And once that is done,
6:07:16here I'm importing this SVM. So, this
6:07:20SVM model is available inside my
6:07:22scikit-learn library. So, I've mentioned
6:07:24as
6:07:25scikit-learn .svm
6:07:29and from scikit-learn.svm, I'm importing
6:07:32my SVC. Okay? So, I'm importing my SVC.
6:07:35So, I'll show you what is this SVC.
6:07:38Um SVM SVC
6:07:45So, it's C means support vector
6:07:47classification.
6:07:48Okay? Now, here when I'm instantiating
6:07:52this SVC, I can mention what is the
6:07:54kernel that I want to use. And if I'm
6:07:57working with any polynomial kernel, then
6:07:59I can also mention what is the degree of
6:08:01polynomial that I want to use while
6:08:03performing the fit for my data set.
6:08:06So, I'm importing my SVC. And along with
6:08:09that, I'm also importing the data sets.
6:08:12So, in the scikit-learn library itself,
6:08:14we have a data set. So, it the
6:08:17scikit-learn host already like it it
6:08:19actually scikit-learn has many toy data
6:08:22set which will actually help us in our
6:08:24learning journey. So, we are going to
6:08:25use one of the data set, the famous Iris
6:08:28data set. We use that for multi-class
6:08:31classification.
6:08:32So, I'm going to load that Iris data
6:08:34set.
6:08:35And I'm just going to extract only two
6:08:37features. So, the two features that I'm
6:08:39extracting is petal length and petal
6:08:41width because I don't want to complicate
6:08:43it. I just want to visualize the data.
6:08:45So, in order to help in visualization, I
6:08:48I'm just getting only two features of my
6:08:50given data.
6:08:52And I'm separating my Y as
6:08:55Iris target. So, whatever the target
6:08:56variable that I had, I'm assigning to my
6:08:59variable of Y.
6:09:01Then,
6:09:02I'm going to uh
6:09:04do this check whether it is setosa or
6:09:07versicolor.
6:09:08That means this default data set, which
6:09:11is in multi-class classification, I'm
6:09:13just going to convert it into a binary
6:09:15classification task.
6:09:17You'll get a better understanding once I
6:09:19execute this next cell. So, this going
6:09:21to prepare my data set and once the data
6:09:24set is prepared, if I create a scatter
6:09:26plot, so I'm just creating the scatter
6:09:28plot to show us
6:09:30what and how my data set looks like. So,
6:09:33this is how my data set looks like.
6:09:36On my X axis, I think I'm having petal
6:09:38length. On my Y axis, I'm having petal
6:09:40width.
6:09:41And here,
6:09:42the blue points refers to the class zero
6:09:45and the orange points refers to the
6:09:48class of one.
6:09:50Okay? So, this is how my data set looks
6:09:54like.
6:09:55You can clearly see that I have one set
6:09:57of points in one region and I have
6:09:59another set of points in another region.
6:10:01Now, this is a classic example to
6:10:03understand about the SVM. How does it uh
6:10:06draw a hyperplane?
6:10:09So, we have the data set ready.
6:10:11And as I mentioned already, in order to
6:10:15fit this model, so when I say support
6:10:17vector machine, I'm going to draw a
6:10:18line.
6:10:20This line that I have drawn, it will be
6:10:23equidistant from my support vectors.
6:10:25Now, here in this example, the support
6:10:27vector is this because this is the only
6:10:30point which is nearest to my line. And
6:10:32here, I think this is the data point
6:10:34which is nearest to my SVM SVM line,
6:10:37that is this twisted line hyperplane
6:10:38line.
6:10:39So, I'll be placing this hyperplane such
6:10:42that it is equidistant from the support
6:10:45vectors.
6:10:47That is how I'll be drawing this support
6:10:49vector line.
6:10:51So, we now have an intuition. Let's see
6:10:53whether we get the same outcome as we
6:10:55are expecting.
6:10:57So, here I'm initializing my model. So,
6:11:00for initialization, I'm saying it as
6:11:02SVC. Use the kernel as linear because
6:11:06I'm able to draw a line effectively. We
6:11:08were We just seen. And I'm using the C
6:11:11as infinity, that means it should be a
6:11:13hard classifier. So, hard classifier
6:11:15means I make I want the 100% result. I
6:11:18mean, I don't want any loosens. I want
6:11:20to draw a line which passes which
6:11:23clearly separates two classes. So, I'm
6:11:25saying it as C as infinity to mention
6:11:27this as a hard classifier.
6:11:30And once I initialize any model,
6:11:32here in this scenario, SVM model, I'm
6:11:35performing the fit on my data set. Now,
6:11:36this is the common flow that we follow
6:11:39whenever we are performing the fit. So,
6:11:41we'll initialize the model and then we
6:11:43perform the fit on a data set.
6:11:46Now, since this SVM being a supervised
6:11:49machine learning model, I have to
6:11:51specify both my input X as well as my
6:11:55output Y.
6:11:56Hence,
6:11:57SVM classifier.fit
6:12:00X, Y.
6:12:01So, this is going to perform the fit for
6:12:03my data set.
6:12:05I'll just execute this. So, this has
6:12:08performed the fit and here it is giving
6:12:10me the confirmation as this is the
6:12:13parameter that are being used to perform
6:12:15the fit.
6:12:17Okay. Now, once I have drawn and once I
6:12:20have found this fit,
6:12:22next,
6:12:23if I want to display the weight terms,
6:12:26so, I can say it as SVM
6:12:28classifier.coefficients.
6:12:30So, these are the weight terms. And if I
6:12:32want to display my bias term or the
6:12:34intercept, it is minus 3.78.
6:12:38Now, this means the line that I've just
6:12:40drawn, so that line has the
6:12:44uh
6:12:45that line has the C term as or the W not
6:12:48term as minus 3.78 and W1, W2 are 1.29
6:12:53and 0.82, respectively.
6:12:56So, that's how the data is distributed
6:12:59for us. That's how the values has been
6:13:02formed for our scenario.
6:13:05Next, in order to get the better
6:13:07visualization,
6:13:08here I have created a function that is
6:13:10called as plot SVC decision boundary and
6:13:14this takes my SVM model,
6:13:17the X min and the X max.
6:13:21Now, W and B I'm extracting from the
6:13:24coefficient and the intercept parameter
6:13:26that we have over here.
6:13:28So, we are extracting from this
6:13:30uh
6:13:31at from this attributes that we have
6:13:33from this model.
6:13:35And now, if I want to draw a decision
6:13:37boundary,
6:13:38so,
6:13:39I need the set of points. So, in order
6:13:42to get the points, I'm saying it as X
6:13:43not is equal to np.linspace X max, X
6:13:46min, X max, 200. And I'm specifying as
6:13:50how does my decision boundary should
6:13:52look like.
6:13:53My decision boundary is given by w
6:13:56naught into x naught plus w one into x
6:13:58one plus b is equal to zero. So, this is
6:14:00what my
6:14:01decision boundary would look like. So, I
6:14:03know what is x naught. I know w naught.
6:14:06I also have w one and I also have b. So,
6:14:09the only term that I do not have is my
6:14:12x one.
6:14:13Okay? So, the only term that I do not
6:14:15have over here in this example is x one.
6:14:18And the x one if I want it, so I just
6:14:20have to substitute it. So, x one is
6:14:22equal to minus w zero divided by w one
6:14:25into x naught minus b divided by w one.
6:14:30Now, I'm specifying the same equation
6:14:33over here for my x two.
6:14:35So, my x naught and the decision
6:14:37boundary will give me the pair of input
6:14:40and output. Okay?
6:14:42Now, along with this
6:14:44there is a property in SVM. Okay? So,
6:14:47the property is given by whenever I have
6:14:49a margin, so that margin is given by one
6:14:53over w one.
6:14:55Okay? So, the margin is nothing but the
6:14:58distance between my hyper plane and the
6:15:01support vector. So, that is given by ma
6:15:04one by w one.
6:15:06Hence, I have mentioned as gutter up and
6:15:08down. Gutter up means one line or the
6:15:11one line where the support vector lies.
6:15:13So, that is given by decision boundary
6:15:16plus margin.
6:15:17And one line below my
6:15:19one line below my hyper plane.
6:15:22That is where another support vector
6:15:24would lie. So, I mentioned as decision
6:15:26boundary minus margin.
6:15:29Okay?
6:15:30Then
6:15:31I'm defining where exactly my support
6:15:34vectors are present.
6:15:37My support vectors uh I can access the
6:15:40support vectors coordinates by saying it
6:15:42as
6:15:43by accessing the attribute of my train
6:15:45model support underscore vectors
6:15:47underscore.
6:15:49Now, I'm specifying where exactly those
6:15:51support vectors are present with the
6:15:53help of a simple scatter plot by
6:15:55highlighting my support vectors.
6:15:57And I'm specifying where exactly my
6:15:59decision boundary is present. And I'm
6:16:02also mentioning where is my line that is
6:16:04gutter up and gutter down. So, let's do
6:16:06one thing. I'll just execute this. This
6:16:08is going to create me a function.
6:16:11I'm going to call my function
6:16:13support vector machine classification.
6:16:16And I will specify my range of X and Y
6:16:18as
6:16:19here, yeah, X min and X max as 0 {comma}
6:16:225.5.
6:16:24I'll just execute this.
6:16:28So,
6:16:29what we have done just now is we have
6:16:32created this hyperplane.
6:16:36So, the middle one, the solid line that
6:16:38you're seeing over here, so this solid
6:16:40line is called as your hyperplane.
6:16:43And these points that you're seeing over
6:16:45here, so these two points which are
6:16:47highlighted, these two points are
6:16:49actually called as support vectors.
6:16:54Okay? So, this dotted line that you're
6:16:57seeing, so this dotted line refers to my
6:17:00gutter up and gutter down which I've
6:17:02found right here.
6:17:05Let's do one thing. I'll add some label
6:17:07so that you'll get some more
6:17:09visualization in the plot itself. I'll
6:17:11say label and I'll mention it as
6:17:14hyperplane.
6:17:19Okay.
6:17:21And
6:17:23there is one more, yeah.
6:17:36These are support vectors.
6:17:38And I'll say
6:17:42plt.legend.
6:18:00So, this clearly says which are all my
6:18:04hyperplane and which are all my support
6:18:06vectors.
6:18:09So, this is the intuition behind support
6:18:11vector machines.
6:18:13So, we'll be drawing a hyperplane which
6:18:16separates the points that we have.
6:18:18Okay? And whichever the point which is
6:18:20nearest to my hyperplane, we call that
6:18:22point as a support vector. Now, to
6:18:25access that support vector, we make use
6:18:27of the attribute. So, let's do one
6:18:29thing. Let's explore the same the
6:18:31attributes.
6:18:32svm.
6:18:33support_vectors_.
6:18:36So, this is going to tell me where
6:18:37exactly my support vectors are present.
6:18:40So, one point is given by 1.9.0.4.
6:18:43I think this is the point that I'm
6:18:45talking about.
6:18:46And the another support vector that we
6:18:48have is at the location 3, 1.1. So, 3
6:18:51and 1.1. This is where we have another
6:18:54support vector.
6:18:56So, using all these attributes, we have
6:18:58been able to create this visualization.
6:19:03Okay.
6:19:05Now,
6:19:06whenever we are working with the support
6:19:07vector machines,
6:19:09it's very important that we scale the
6:19:11data first. If I do not scale the data,
6:19:14I'll not be able to get a better fit of
6:19:17my SVM model.
6:19:19So, here I've given one more example
6:19:22where I have my X
6:19:24uh is given as 1, 55, 23, 80. As you can
6:19:28clearly see, it's it's not scaled. Okay?
6:19:32So, I'm going to execute this cell. So,
6:19:35this is going to tell me and give me a
6:19:37visualization as how the
6:19:39fit will be in case of scaled and
6:19:42unscaled.
6:19:44See, if it is unscaled
6:19:47I'll If it is not scaled, okay? That
6:19:49means if it is unscaled, we can clearly
6:19:51see that the hyperplane that I'm drawing
6:19:55and the distance from my hyperplane,
6:19:57it's very close to each other.
6:20:01And whenever I'm working, it's It's It
6:20:03will be difficult for me to separate
6:20:04those two data points.
6:20:08But, if I scale them correctly
6:20:11Now, here for scaling, I have made use
6:20:12of a scalar standard scalar. Now, if I
6:20:15scale it correctly, then in that
6:20:17scenario, it will be easier for me and
6:20:20it would actually work better when I
6:20:22have scaled data.
6:20:26Okay? So
6:20:28this is about using the linear SVM model
6:20:32to perform the fit on my given data set.
6:20:36Now, if I go below, we have some more
6:20:38examples about non-linear classifiers as
6:20:41well.
6:20:42Now, in order to test out the same
6:20:45here
6:20:46I'm creating an example data set and
6:20:48that data set that I'm generating is
6:20:50called as make moons data set and this
6:20:53has been generated with the help of a
6:20:55scalar data set generator.
6:20:57Now, as you can clearly see, I cannot
K- Means Clustering Algorithm
6:21:00make use of linear classifier. So,
6:21:02linear classifier is nothing but a
6:21:05classifier, okay? Which is an SVM model
6:21:08where I'm drawing or where I'm using a
6:21:10straight line to split my data points. I
6:21:13can clearly see that I Wherever I I join
6:21:16or wherever I try to draw a line over
6:21:19here, I cannot split the data in an
6:21:21effective manner.
6:21:23Now, this brings us the challenge. Now,
6:21:24if I have a data set in this way where I
6:21:27cannot linearly separate it, how can we
6:21:30go about and fit uh perform the fit on
6:21:33our SVM model?
6:21:34So, in order to save us, we have a model
6:21:37that is called as uh SVM model, and from
6:21:40that SVM model, we can actually create a
6:21:43polynomial uh
6:21:45polynomial kernel. So, we can make use
6:21:46of polynomial kernel, and using that
6:21:48polynomial kernel, I can actually uh
6:21:52create it like this. I mean, using
6:21:53polynomial kernel, I can perform
6:21:55polynomial regression.
6:21:57Or I can draw a line like this. Now, to
6:21:59show you how it works,
6:22:01I'm getting some data like this. So,
6:22:03this is some uh random data.
6:22:06And I'm making use of pipeline.
6:22:08So, this pipeline is going to take care
6:22:10of my stan- standard scaler as well as
6:22:13kernel.
6:22:14I'll do one thing, I'll just come below.
6:22:16So, this is what we are currently
6:22:17interested in.
6:22:23So, here,
6:22:26I'm importing the polynomial features,
6:22:29and I'm generating the polynomial
6:22:31features for my data.
6:22:33I'm performing the fit and transform my
6:22:35polynomial data. That means, I'm just
6:22:37modifying my existing data, and I am
6:22:39sending it
6:22:41for my uh
6:22:43X, okay? So, this is how my pair of
6:22:45input X and Y looks like.
6:22:47Now, I'll use my X. I'm going to
6:22:49transform it with the help of my
6:22:51polynomial features,
6:22:53and then, I'm going to scale it with the
6:22:56help of my standard scaler,
6:22:58and I'm going to send it inside my
6:23:01classifier, that is SVM classifier.
6:23:08Okay? So, I'm going to
6:23:11combine it together like this.
6:23:14Now, observe what would happen.
6:23:17Now, once that is complete,
6:23:19see?
6:23:20With the help of my polynomial uh
6:23:24polynomial features that have applied on
6:23:26my given linear data.
6:23:28So, I have increased the degrees by
6:23:32which my model can learn.
6:23:36Now, instead of straight line, my model
6:23:37is also having the ability to learn this
6:23:40complex representation as well.
6:23:43Because I have increased the model
6:23:45complexity by adding my polynomial
6:23:47features.
6:23:50And while doing it, to make sure that we
6:23:51follow a
6:23:53clear path, so I have defined this is
6:23:56scalar's pipeline. So, if you're new to
6:23:58data science machine learning, I highly
6:24:00recommend you to learn this concept of a
6:24:02scalar pipeline. Now, this is scalar
6:24:04pipeline helps us to combine multiple
6:24:07operations in a single call.
6:24:10So, here we have created a pipeline.
6:24:12This pipeline is going to add some
6:24:15polynomial features for my input data.
6:24:18And on top of it, this is going to
6:24:19perform scaling. And then I'm going to
6:24:21perform this binomial classification
6:24:24using this SVM.
6:24:27And finally,
6:24:29I'm performing the fit on my data set.
6:24:31See, when I perform the fit, it takes my
6:24:33input X and it's going to do all these
6:24:36activities. It is going to chain all
6:24:38these activities together, and then it
6:24:41is going to perform the fit for my data
6:24:43Y.
6:24:45Once the fit has been complete, so we
6:24:47can validate how my model is performing.
6:24:50>> [music]
6:24:55[music]
6:24:56>> So, what is a clustering technique?
6:24:57Clustering technique is something that
6:24:59we will use it for grouping purpose.
6:25:03So, especially there's a very easy way
6:25:06to understand what is clustering
6:25:07technique. You would have seen such a
6:25:09while we are going through an
6:25:11COVID-19 situation, the governments has
6:25:13came up with creating some containment
6:25:15zones.
6:25:16As all of you must be knowing.
6:25:19So, how on what criteria government has
6:25:21taken that okay, which area supposed to
6:25:23be a containment zone or which area
6:25:25supposed to be applied with some some
6:25:26restrictions and which areas can be can
6:25:29be considered as normal? On what
6:25:30criteria that they have created? So,
6:25:32that's what using clustering technique.
6:25:35Which means if the governments or when I
6:25:37say government means that the people who
6:25:38will be taking the final decision in
6:25:40such criteria, either prime minister
6:25:41either either the chief ministers of
6:25:43that particular state will be be taking
6:25:45decisions whether to go for lockdown
6:25:47whether to not to go for lockdown or
6:25:49which areas has to be considered as
6:25:51containment zones or non-containment
6:25:52zones.
6:25:53So, those high-level decisions are
6:25:55something which will be taken based on
6:25:57the clustering technique output which is
6:25:59generated by the these algorithms.
6:26:02Based on a number of inputs, okay, what
6:26:04is the population in a particular area?
6:26:06How many number of people are affected?
6:26:08How many number of hospitals which are
6:26:10present?
6:26:11How many number of
6:26:13people who are been recovered? So,
6:26:16likewise based on this these multiple
6:26:18criteria, people will do some clustering
6:26:20technique on top of the data and
6:26:22according to that people will be
6:26:24segregated or the areas will be
6:26:26segregated.
6:26:27So, that saying that okay, these are the
6:26:28observations which belong to one
6:26:29cluster, these are the observations
6:26:31which belong to one cluster like that so
6:26:32that people can cluster them which can
6:26:34make organizations to take decisions on
6:26:38a very high level.
6:26:40Okay? That's what is all clustering
6:26:41technique.
6:26:42Which clustering technique output will
6:26:45contain the different different groups.
6:26:46It itself will group the different
6:26:47different components
6:26:49based on whatever the number of clusters
6:26:51that you want to generate. That may not
6:26:53give the direct output. On top of the
6:26:55generated output, people will be taking
6:26:56business related decisions. That's what
6:26:58is all about clustering techniques.
6:27:01Okay. So, now what we will do?
6:27:03Let's take an example of within
6:27:06clustering techniques, what are the
6:27:07different types of clustering techniques
6:27:08we have?
6:27:09So, what What different types of
6:27:10clustering that we have?
6:27:13So, there are multiple types of there
6:27:14are multiple types of ways based on the
6:27:17type of output that we want to produce.
6:27:18There are multiple different types of
6:27:20clustering techniques we have, but out
6:27:21of which the let's try to understand
6:27:23about what are the very famous and most
6:27:25widely used clustering technique
6:27:26algorithm. Out of which we have
6:27:28something called K-means clustering
6:27:30algorithm is one of the very famous and
6:27:32most widely used. More than 90% of the
6:27:35people will end up with using K-means
6:27:36clustering algorithm, which is very very
6:27:38famous in clustering techniques.
6:27:40Right? Which is very very famous in
6:27:42clustering techniques.
6:27:43So, what are these clustering
6:27:44techniques? As I said, how this
6:27:46clustering technique will work.
6:27:48So, K-means clustering is nothing but
6:27:49always remember one thing. If any one of
6:27:51you were going to work in machine
6:27:52learning or anywhere in anywhere,
6:27:55wherever you see a notation called K,
6:27:57by default K is nothing but you are you
6:28:01are supposed to as a user, you are
6:28:03supposed to provide what is the input of
6:28:06K. Which means wherever you see there
6:28:08are multiple techniques that we have in
6:28:09machine learning like K-means clustering
6:28:11technique, K nearest neighbor is one of
6:28:13the algorithm, K-fold cross validation,
6:28:15likewise. Wherever you see a notation
6:28:17called K, what is what does a K means?
6:28:19It's an input that you are supposed to
6:28:21provide. Always remember this.
6:28:24It's an input that you are supposed to
6:28:25provide
6:28:27to your algorithm. Your algorithm cannot
6:28:29identify that K value. Of course,
6:28:30everything else will be taken care by
6:28:31your algorithm, but whenever you see K,
6:28:33which means in K-means clustering, what
6:28:35is the meaning of K-means clustering?
6:28:37How many number of clusters that you
6:28:38want to provide? That is something that
6:28:40you have to input it to your algorithm.
6:28:43That is something that you have to input
6:28:44to your algorithm.
6:28:45Right?
6:28:47Here, the meaning of K is how many
6:28:50clusters that you want to generate.
6:28:52So, how many clusters that you want to
6:28:54generate? Okay, when you have 1,000
6:28:55observations which are present, when you
6:28:57have 1,000 in input column input records
6:28:59which are present in your historical
6:29:01data, how many number of clusters that
6:29:03you want to provide? Do you want to go
6:29:04for one cluster? Obviously, one cluster
6:29:06means that the entire data set will be
6:29:07considered as is.
6:29:09Do you want to create two clusters out
6:29:11of the data?
6:29:12Do you want to create three clusters out
6:29:14of the data? Four clusters, five
6:29:15clusters, or 10 clusters?
6:29:17So, how this can be done?
6:29:18There are multiple steps that are
6:29:20involved in generating K-means
6:29:22clustering algorithm. So, you can see,
6:29:23choose the number of clusters. This is
6:29:25what is nothing but your first step. It
6:29:26means you need to decide what is your
6:29:29K is nothing but number of clusters that
6:29:31you want to produce.
6:29:32And then, there is an initialization of
6:29:35centroids will happen as a one-time
6:29:37activity.
6:29:38Right? So, there is an initialization of
6:29:40centroids which will be which will be
6:29:42declared that will that will be used as
6:29:44your initial step for your machine. And
6:29:45then, assign the clusters, move the
6:29:47centroids, and optimization, and then
6:29:49converge the the
6:29:51all the clusters into one component.
6:29:53Yes, I know it will be very difficult to
6:29:54understand by looking at this thing. So,
6:29:56let me show you a very simple example
6:29:58how exactly it will be done. Maybe let
6:29:59me take a simple diagram for you to show
6:30:01how exactly that's going to work.
6:30:05Okay? So, let's say for example, I'm
6:30:07going to take some historical data just
6:30:09to explain you on how exactly the
6:30:11K-means clustering algorithm will work.
6:30:13So, what is that it is written in the
6:30:14first step?
6:30:16What is the data is written in the first
6:30:17step?
6:30:18So, the choose the number of clusters.
6:30:20Choose number of clusters. Now, it means
6:30:22say for example, we need to take some
6:30:24historical data. I'm considering some
6:30:25historical data here. Let's assume this
6:30:27is the historical data.
6:30:29Let's assume this is the historical data
6:30:31that we have.
6:30:32So, as you can see, there are multiple
6:30:35historical data points. As you can see,
6:30:37the first step, choose number of
6:30:38clusters, which means let's assume to
6:30:40make this thing simple, I'm going to
6:30:42choose that we want to have two clusters
6:30:43generated. Okay? So, this is which means
6:30:46we want to have two clusters created.
6:30:47So, because my number of clusters that I
6:30:49want to generate is two, I'm going to
6:30:51consider that there are two centroids
6:30:52which are present. This is my first
6:30:54step. How this algorithm will do?
6:30:56How algorithm will come to come with the
6:30:58number of clusters? So, this is how it
6:30:59will happen. So, the first step is to
6:31:01choose the number of centroid and then
6:31:03initialize your centroid. That's the
6:31:04second step. So, now what is the third
6:31:06step?
6:31:07Let's assume this is observation number
6:31:09one. This is our data point one.
6:31:12So, now what what is the next step?
6:31:14We will take the distance from every
6:31:16observation to every centroid and every
6:31:18observation to every centroid, which
6:31:19means now tell me which observation
6:31:23For this observation number one, which
6:31:24centroid is more closer?
6:31:26Is the green color centroid is more
6:31:27closer or the red color centroid is more
6:31:29closer for this?
6:31:31Centroid means that data points, the one
6:31:33which I have highlighted, that's what we
6:31:34call a centroid, data points.
6:31:37We have chosen two centroids only
6:31:39because that is the as a user you're
6:31:41supposed to input what's supposed to be
6:31:42your K.
6:31:43What's supposed to be your K value,
6:31:45that's what I said. You need to know how
6:31:47what is your K is nothing but how many
6:31:48number of clusters that you want to
6:31:49provide. Usually your business users are
6:31:51going to provide that. In case if you're
6:31:53going to work in this kind of
6:31:54algorithms, they will provide that. Or
6:31:56else there are some other methods
6:31:58available like
6:31:59elbow method available other things.
6:32:01Yeah, we'll talk about that.
6:32:03Now, this observation is more closer to
6:32:05red color. So, now what happens? What
6:32:07what is the next step? The algorithm
6:32:08will assign this particular observation
6:32:10one as a red color for now.
6:32:12As considering that this is belong to
6:32:13red color. Likewise for the second
6:32:15observation, when the second observation
6:32:16appears, what is the distance from the
6:32:18second observation to both the
6:32:20centroids? Now, which one is more
6:32:21closer? I see green color is more
6:32:23closer. So, now I'll mark this as green
6:32:25color.
6:32:26If both if what if the distance is same
6:32:28equal? So, then the algorithm will force
6:32:31any of the observation to get into any
6:32:33of the centroid.
6:32:35So, number of clusters number of
6:32:37centroids that you will choose based on
6:32:38the K value as I said.
6:32:41Now, you're going to Likewise, you will
6:32:43repeat the same process and whatever the
6:32:45observation which is more closer to
6:32:47whatever the centroid it is, you will
6:32:48mark them as with their so-called mark
6:32:51like this. Now, you're going to
6:32:52initially mark them as observations into
6:32:54either into green color either into red
6:32:56color.
6:32:56So, now, these observations are now
6:32:59considered as green color observations,
6:33:00and these observations are now
6:33:02considered as red color observations.
6:33:05This is step number one.
6:33:07What is the step number two?
6:33:09Step number two is segregate all these
6:33:11red color observations and take the
6:33:13average value of X and Y coordinates,
6:33:15and repeat the same process for
6:33:17calculating average coordinates of X and
6:33:18Y coordinates for the green color
6:33:20observations.
6:33:21Take the average of all these
6:33:22observations, calculate average, and
6:33:24take the all these observations, take
6:33:25the average. You end up with getting a
6:33:27new centroid positions called XY. Which
6:33:30means, now you ended up with getting a
6:33:32new centroids in the initial step that
6:33:35we have taken random centroid. Now, you
6:33:37got the centroids that you can use based
6:33:40on the previous iteration. Now, you
6:33:41ended up with getting a new centroids.
6:33:44You repeat the same process again.
6:33:46Again, you repeat the same process. Take
6:33:47the distance from every observation to
6:33:49every centroid and assign the
6:33:50observation based on the nearest nearest
6:33:52to distance, and continue to mark every
6:33:55observation either into red color or
6:33:56either green color or whatever it is,
6:33:58and you repeat the process until you
6:34:00will be able to see there will be no
6:34:02change applicable for your clusters.
6:34:05You repeat the process. Which means, in
6:34:07every step, you might end up with
6:34:08changing your centroid. Every step will
6:34:10continue to change your centroid. Now,
6:34:12the centroid might become like this.
6:34:13Then, later, your centroid will become
6:34:15like this.
6:34:17Likewise, your centroids will be keep on
6:34:19moving. Initially, you have taken it
6:34:20like this, but it it might continue to
6:34:22move like this. Somewhere, it will be
6:34:23fixed. And after that, there will be no
6:34:25change that you will notice if you are
6:34:26repeating the same process. At this
6:34:28particular stage, whatever the
6:34:30observations which are marked into which
6:34:32are
6:34:33grouped into green color, you say like
6:34:35these are the green color observations.
6:34:37Whatever the observations which are
6:34:39marked into red color, you will you will
6:34:40mark them as okay, these are the
6:34:41observations which belong to red color.
6:34:44These are the observations which belong
6:34:45to red color. Likewise, you will
6:34:47segregate all these observations either
6:34:49into red color or green color.
6:34:51Right? So, so, we can generate these
6:34:54cluster techniques. That's how the
6:34:55K-means clustering algorithm will
6:34:56generate these algorithms
6:34:58output of using this algorithm.
6:35:00Okay? So, that's how K-means clustering
6:35:02algorithm will work.
6:35:04Okay?
6:35:05So, now what likewise there are multiple
6:35:06algorithms that we have. So, like
6:35:09when we talk about machine learning, so
6:35:10there are multiple types of algorithms
6:35:12that we have. So, like how we have the
6:35:14how does the K value K value will occur
6:35:16as you can see on the PPTs which are
6:35:17also mentioned. So, you're going to
6:35:19choose some randomly generated K value
6:35:21and you will be choosing the number of
6:35:22case over here and you can see that it
6:35:24will assign them based on the number of
6:35:26these easy
6:35:27most nearest distances and according to
6:35:29that you will change your centroids.
6:35:31Once you change your centroids, you will
6:35:33repeat the same process until you're
6:35:34able to change that your centroids don't
6:35:37move further and then once it has been
6:35:39finalized, you will say like this is the
6:35:40final centroid.
6:35:42Final cluster output that we can
6:35:44generate out of this.
6:35:46Right?
6:35:47So, likewise we also have different
6:35:48types of cluster technique and the
6:35:49second type of cluster technique that we
6:35:51have is a fuzzy or C-means clustering.
6:35:53So, what is fuzzy or C-means clustering?
6:35:54The output will remain same. So, C-means
6:35:56clustering means that there are places
6:35:58that one or two observations can belong
6:36:00to one or two different clusters.
6:36:03Like one or two different clusters,
6:36:05which means usually the primary
6:36:06difference between
6:36:09the primary difference between your
6:36:10C-means and K-means clustering technique
6:36:12is in case if there are any observations
6:36:15which are having equal amount of
6:36:16distance, usually in K-means clustering
6:36:18what we will do, we will force this
6:36:20observation to be part of any of the
6:36:22cluster. But in C-means clustering,
6:36:25based on the distance that we see, there
6:36:27are chances that an observation can go
6:36:29to or can belong to one or two clusters.
6:36:33So, it purely depends on the business
6:36:34use case who is going to decide to
6:36:36either to go for K-means clustering or
6:36:38C-means clustering based on the the
6:36:39business use case.
6:36:41Based on our business outcome, so if
6:36:43people are going to decide whether to go
6:36:44for C-means clustering or K-means
6:36:45clustering. So, there are n-number of
6:36:48observations which might It's not
6:36:50mandatory. Which might can belong to one
6:36:52or more clusters. That can happen.
6:36:55So, the third is agglomerative
6:36:57clustering, so which which is the third
6:36:58type of clustering technique that we
6:37:00have. Okay? So, which is also one of one
6:37:03of the clustering technique that we
6:37:04have.
6:37:06Okay?
6:37:07So, now which can also be used here.
6:37:09Okay? So, now what is this agglomerative
6:37:12clustering? So, this is what we also
6:37:14call it as hierarchical clustering. So,
6:37:16in times you'll also call them as
6:37:18hierarchical clustering. These
6:37:19clustering techniques are built using
6:37:20H-clustering. In short, we call it as
6:37:22H-clustering.
6:37:23And also people will also call it as
6:37:24hierarchical clustering techniques. So,
6:37:25what are this? So, based on the type of
6:37:28algorithm
6:37:29they will try to There is a There is a
6:37:31mathematical expression which are
6:37:32involved in it. But, then considering
6:37:33that the limited time that we have, I'm
6:37:35not going to take you through all the in
6:37:36in detail depth of it. So, considering
6:37:38the way how the data points are being
6:37:39segregated, we'll it will build a kind
6:37:41of dendrogram. So, on top of this
6:37:43dendrogram, your observations can be
6:37:45classified here like this, as you can
6:37:46see on the screen. Is K-means clustering
6:37:48sensitive to outlier?
6:37:50You need to understand one thing.
6:37:52When we are talking about unsupervised
6:37:54learning algorithms, as I said, you may
6:37:56or may not have clarity on the data.
6:37:59Which means
6:38:01your assumption is that at the least
6:38:02level
6:38:04you don't have clarity on the data.
6:38:06Then if you don't When you don't have
6:38:08clarity on the data, how can you say
6:38:09that This is an outlier or this is not
6:38:11an outlier?
6:38:12When you have clarity on the data,
6:38:13that's especially when you're working on
6:38:14supervised learning algorithms, you can.
6:38:16But, you don't have output column also.
6:38:18How you will be able to evaluate how the
6:38:20so-and-so-called output column can be
6:38:22can be evaluated because this is not
6:38:24being classified properly, this is not
6:38:25being clustered properly based on the
6:38:27historical data because there is no
6:38:28output column.
6:38:29So, those type of concepts are something
6:38:31which you don't need to worry about when
6:38:32you are working on supervised learning.
6:38:34They will be primarily they'll be
6:38:35constrained when you are working on
6:38:37supervised learning algorithms.
6:38:38And of course, in case if you see that
6:38:40there are outliers which are present,
6:38:41obviously it has to be It will be
6:38:43considered It will be considered as one
6:38:44of the cluster in of any any of any of
6:38:46the cluster it belongs to the data. But
6:38:48anyhow, that's the characteristic of the
6:38:49data. You don't need to
6:38:51You don't need to take care because you
6:38:52might be killing the actual original
6:38:54values which are present in the data.
6:38:55But you cannot expect that outlier can
6:38:57be recognized
6:38:58for all the cases that you have in
6:38:59unsupervised learning.
6:39:01So, the next type of clustering
6:39:02technique that we have is division
6:39:04clustering. So, what is this division
6:39:05clustering? As a division clustering is
6:39:07also creates the data in a form of
6:39:09dendrogram. But the difference is you
6:39:11can see that the starts with all data
6:39:12points in one cluster, splits the root
6:39:14into child recursively based on the
6:39:16dendrogram, and stops when there is a
6:39:18single term clusters which are created.
6:39:19Which means that for every cluster there
6:39:21will be one observation which belong to.
6:39:23So, likewise the clustering technique
6:39:24will work.
6:39:25Now, there is one more technique that we
6:39:27have in terms of building a clustering
6:39:28technique, that's what we call it the
6:39:29mean shift clustering. So, what we will
6:39:31do we'll take an average of every
6:39:33cluster that we have. What we will do
6:39:35we'll end up with reducing their means
6:39:36into the into a single density of items,
6:39:39and then we'll continue you to repeat
6:39:40the process to see to that which
6:39:42observations mean will belong to the
6:39:43same cluster, and according to that
6:39:45we'll continue to identify which
6:39:47observation will belong to a cluster a
6:39:48particular cluster.
6:39:50So, likewise we can apply different
6:39:52types of clustering technique that can
6:39:53help us to cluster the data which is
6:39:55part of unsupervised learning.
6:39:57Right? So, now let's take a let's jump
6:40:00into something called a small hands-on.
6:40:01Okay, let's let's take a small
6:40:03Python example, and we will see how do
6:40:05we build that particular a clustering
6:40:07technique on top of the data using one
6:40:09of the simple data that we have.
6:40:11Jupiter notebook, let me open uh how
6:40:15we can use scikit-learn using one of the
6:40:17example. So, meanwhile let me show you
6:40:19some of the data set also.
6:40:22Let me show you a data set that I'm
6:40:23going to use as well.
6:40:25So, I'm What I'm going to do is I'm
6:40:27going to take this movie metadata
6:40:28information. Let me open this data set.
6:40:32I'm going to take this example of movie
6:40:34metadata information where this this
6:40:37data set has got a number of
6:40:38observations which are present.
6:40:41Okay, let me open this. Okay, you can
6:40:43see that there are a number of movies
6:40:44related information as you can see the
6:40:45movie names. So, Avatar, Pirates of the
6:40:48Caribbean, Spectre, The Dark Knight,
6:40:50Star Wars, etc. etc. John Carter,
6:40:52Spider-Man 3, Tangled, or etc. etc. We
6:40:55have a lot of movies.
6:40:56And about every movie we got a lot of
6:40:58information which is present like who is
6:40:59the director, who is the actor, what are
6:41:01the director Facebook likes, what are
6:41:03the actor Facebook likes, likewise we
6:41:04got we got a lot of information which is
6:41:06present as part of these particular
6:41:08every observation.
6:41:10So, now what we will do, we will try to
6:41:12cluster these data points. You can see a
6:41:13lot of observations which are given.
6:41:15What is the gross of the movie? What is
6:41:16the number of reviewers? What is the
6:41:17IMDb rating? What is the so-and-so
6:41:19called IMDb score? What is the movie
6:41:21span? What is the gross? What is the
6:41:22budget? And everything.
6:41:24So, now what I'll be doing, I'll be
6:41:26reading this data set using one of the
6:41:27Pandas library that we have. Okay, so
6:41:30I'll be reading this data set. Let me
6:41:31open the Jupiter notebook.
6:41:33I'll be using something called Pandas.
6:41:35So, as you can see, read.pandas.csv.
6:41:37I'll be reading this data set where you
6:41:38can see that this is the data set that
6:41:39I'm able to read. I got a number of
6:41:41observations that are present. I can see
6:41:43that director Facebook likes and actor
6:41:45three Facebook likes.
6:41:47So, now what I what is it I want to do
6:41:48is instead of building this custom
6:41:50technique on top of every observation,
6:41:52so what I will do is I will take this
6:41:54columns called
6:41:55called number of Facebook likes on
6:41:57director and number of actor Facebook
6:41:58likes versus director Facebook likes.
6:42:00So, where I can see that if I want to
6:42:02select a director Facebook likes alone,
6:42:03I'll be able to choose this director
6:42:05Facebook likes alone. You can see for
6:42:07every movie you got the number of
6:42:08director Facebook likes that are
6:42:09present. So, if there are more number of
6:42:11Facebook likes, what does it mean?
6:42:13The director is famous person.
6:42:15Correct?
6:42:17Or rather if there are more number of
6:42:18famous Facebook likes that the actor has
6:42:20got, which means that the person is or
6:42:22the actor is very famous person. That's
6:42:24what you can understand. So, now you can
6:42:25see that I'll try to extract these
6:42:27independent component that we talking
6:42:29about. I will extract all the records in
6:42:31all the director Facebook likes versus
6:42:33actor Facebook likes where I'm just
6:42:35going to form an object called new data
6:42:37by applying some I location as a filter.
6:42:39I location I LOC stands for index
6:42:41location where I can filter out what are
6:42:43the records that I wanted to what are
6:42:45the column that I wanted to using this I
6:42:47location function.
6:42:49Now I got all the so and so called
6:42:51number of director Facebook likes versus
6:42:52actor Facebook likes which are present
6:42:54as part of this. Okay, so this has been
6:42:56loaded into an object called new data.
6:42:59So now what is that I'll be doing? After
6:43:00that I'm importing something called SK
6:43:02learn K means cluster. So this algorithm
6:43:04is available as part of this
6:43:05scikit-learn algorithm. What are the
6:43:07number of steps that we have discussed
6:43:09you are not supposed to execute all
6:43:10these steps manually and of course if
6:43:12you want you can also write such a
6:43:13program as well, but what is that we'll
6:43:15be doing in scikit-learn
6:43:17library there is a Python library called
6:43:19scikit-learn which has got most of the
6:43:20algorithm present and we'll be importing
6:43:22this K means clustering algorithm. And
6:43:25for this K means clustering I'm
6:43:26providing the C is equal to number of
6:43:28clusters is equal to five.
6:43:29So here if you are if you are aware of
6:43:31uh
6:43:32object-oriented programming using Python
6:43:34you'll be able to correlate. I'm
6:43:35importing this K means clustering which
6:43:37is implemented as a class here where I'm
6:43:39creating an object called K means
6:43:40clustering by providing an input called
6:43:42N number of scores clusters is equal to
6:43:44five which means that what is the
6:43:45meaning of five? I want to build a five
6:43:47clusters out of this. So where once you
6:43:49create an object using K means I'm
6:43:51calling this method called fit the
6:43:53method by providing new data as my
6:43:55independent variables.
6:43:57I'm for calling this fit method which
6:43:59means fit is a method that we're going
6:44:00to invoke what supposed to be the
6:44:01process that needs to be executed that's
6:44:04going to build my clustering technique
6:44:05algorithm by taking number of clusters
6:44:07is equal to five.
6:44:08So now I'm able to generate my algorithm
6:44:10where by looking at my model I'll be
6:44:12able to extract what are the centroids
6:44:13that I got final centroids because
6:44:15initially you'll be taking some random
6:44:16centroids, but at the end you'll end up
6:44:18with getting a final centroid position
Hierarchical Clustering
6:44:20somewhere fixed to it that is nothing
6:44:22but a center point for every cluster
6:44:24that we got. These are the final
6:44:25centroids that we got.
6:44:27Even if I want to print the what is the
6:44:29labels which are generated labels is
6:44:30nothing but it will extract the outcome.
6:44:32You can see that these are the label
6:44:33numbers which are added here. Out of
6:44:355,000 movies we got the labels which are
6:44:37added as an array. But we
6:44:39[clears throat] won't be able to see it
6:44:39like that. What is it I'm trying to do?
6:44:41I'm trying to get all the unique values
6:44:43present in this labels with respect to
6:44:45two counts. Now you can see that my data
6:44:47is now clustered into five different All
6:44:49the movies are clustered into five
6:44:50different clusters as you can see.
6:44:52Cluster number zero has got 4,700 Most
6:44:55of the observations are moved into
6:44:56cluster number zero.
6:44:57104 movies went into observation number
6:45:00one. 11 movies are moved into cluster
6:45:02cluster number two. And 87 movies are
6:45:04moved into cluster number three and 67
6:45:06into cluster number four.
6:45:08That is how the data properties are
6:45:09being distributed and that's how the
6:45:11clustering technique has divided the
6:45:12data into five clusters.
6:45:14Now you can see what is it I'm trying to
6:45:15do? I'm trying to put this into a new
6:45:17data of cluster which means what are the
6:45:19labels that are generated here.
6:45:21This is my output column. I'm going to
6:45:23create this into as a new column in my
6:45:25new data. Where I'm using this Allen
6:45:27plot that will print whatever the
6:45:29columns that I have called director
6:45:31Facebook likes versus actor Facebook
6:45:32likes.
6:45:33And I'm choosing this data is equal to
6:45:35new data that will help me to that will
6:45:37help me to identify what column can be
6:45:39considered as few so that I'll be
6:45:42printing it in a cool warm type chart.
6:45:44You can see that palette type is equal
6:45:46to cool warm type which will help me to
6:45:49identify based on the cluster that you
6:45:52have created.
6:45:53Now if I print this you can see that
6:45:54pretty much the every observation is now
6:45:56categorized into individual cluster you
6:45:58can see. These are the movies which are
6:45:59now graphical representation. This is
6:46:01the graphical representation that we are
6:46:03using to see how these movies are being
6:46:05segregated. You can see these movies are
6:46:07nothing but cluster number zero.
6:46:09You can see there are a few more movies
6:46:11which are extracted over here. This is
6:46:12the nothing but based on the color
6:46:14indication this is cluster number three.
6:46:16And these movies are created as a
6:46:17cluster number one.
6:46:19Cluster number two. These movies are
6:46:21something which are created as cluster
6:46:22number one.
6:46:24And these movies are created as cluster
6:46:25number four.
6:46:27Zero.
6:46:31One.
6:46:32This is nothing but two. Cluster number
6:46:34three. Cluster number
6:46:35three and then cluster number four like
6:46:37this.
6:46:38Now you can clearly see that how is the
6:46:39cluster came clustering has came up. The
6:46:41movies which are made by new people with
6:46:44the new directors, new actors are making
6:46:46films with new directors.
6:46:48Or new directors are making films with
6:46:50new actors.
6:46:51These are all You will see more number
6:46:53of movies will fall into this category
6:46:54because you'll end up with getting new
6:46:56people into the into the into the film
6:46:57industry most of the cases.
6:47:00Lot of movies are being made with new
6:47:02directors with new actors.
6:47:06And you can see that these movies are
6:47:07the clusters which are being segregated
6:47:09very clearly. The
6:47:11famous actors are making films with new
6:47:13directors. Very famous actors are making
6:47:15films with new directors.
6:47:17And these movies are something where
6:47:18very famous directors are less I mean
6:47:21average paying directors are making
6:47:23films with
6:47:24some new actors. You can see these
6:47:26movies are nothing but very famous
6:47:27actors are making directors are making
6:47:29films with very
6:47:31new actors.
6:47:33You can see very famous directors are
6:47:35making films with very famous actors.
6:47:37Like Liber be making a film with a Tom
6:47:40Cruise or something like that.
6:47:42So likewise you can clearly see that
6:47:44where instead of if you do this activity
6:47:46manually it might take a little longer
6:47:48time for you to segregate each and every
6:47:49component. But within five minutes we
6:47:51are able to cluster this activity.
6:47:53Right? So that's what the beauty of
6:47:54algorithms. You don't need to manually
6:47:56do this activity where it can provide
6:47:58your data automatically based on the
6:47:59properties of the data your algorithm
6:48:01itself will will this clustering output
6:48:02out of the data.
6:48:04All right. So that's how we will be able
6:48:05to build a clustering technique on top
6:48:07of the given data. So that's what we
6:48:09have as part of one of the hands on
6:48:10example.
6:48:15>> [music]
6:48:17>> What is hierarchical clustering?
6:48:20So, hierarchical clustering, it is also
6:48:21known as HCA or hierarchical cluster
6:48:24analysis, and this is a method of
6:48:26cluster analysis, as we have seen. So,
6:48:28what happens here is that this
6:48:30clustering allows us to build the tree
6:48:33structure from data similarities, like
6:48:35we have built X and Y, and we have
6:48:38created trees, and these trees are
6:48:40actually called known as dendrogram. So,
6:48:42the way you represent a hierarchical
6:48:44cluster or a hierarchical clustering is
6:48:47through dendrogram. So, we actually drew
6:48:49a dendrogram, okay? So, this is how the
6:48:52clustering is being formed, and this is
6:48:54how the clusters are being made, and
6:48:56this is how the relationship among the
6:48:58clusters is being shown by a dendrogram.
6:49:01So, based on these things, now we will
6:49:03go on further to understanding what is
6:49:05agglomerative clustering. So, when this
6:49:08hierarchical clustering follows a
6:49:09bottom-up approach, this is called
6:49:12agglomerative clustering. And when this
6:49:14is following up a top-down approach,
6:49:16this is used in divisive clustering. So,
6:49:19now let us understand what is
6:49:20agglomerative clustering. So, the types
6:49:22of hierarchical clustering are two, that
6:49:24is agglomerative and divisive. So, now
6:49:26moving on further, what is agglomerative
6:49:28clustering? So, agglomerative
6:49:30hierarchical clustering, this is also
6:49:32known as AGNES, which means
6:49:33agglomerative nesting hierarchical
6:49:35clustering, and it follows a bottom-up
6:49:37approach, which means that clustering or
6:49:40clusters, they are formed from the
6:49:41bottom and are again clustered till a
6:49:44complete single cluster is formed. And
6:49:47what happens then is that the clustering
6:49:50continues until we obtain a single
6:49:52cluster, and we will see how we obtain a
6:49:54single cluster, and we also represent
6:49:56it. So, individual data points, they are
6:49:58clustered based on similarity, and we go
6:50:01on clustering until there is only one
6:50:03single cluster left. So, let us just
6:50:05plot this agglomerative clustering and
6:50:07make things really simple for us.
6:50:10Now, suppose I have got these data
6:50:12points scattered here A to G and these
6:50:15data points have to be clustered. So
6:50:17another important thing is that now we
6:50:19will form clusters. So how clusters are
6:50:21formed? So we can see that based on some
6:50:23similarity like because of the distance
6:50:26nearby distance A and B can be grouped
6:50:28together in a single cluster. So I'm
6:50:31just doing that.
6:50:32C and D I form another cluster because
6:50:34they are near. So I just club them and
6:50:37again I would just club E and F based on
6:50:39their distance and G is separate so I
Apriori Algorithm Explained
6:50:42will just form a separate cluster. Now
6:50:44what happens in agglomerative clustering
6:50:47is that I have to plot all these data
6:50:49points like A B C right? D
6:50:54E F and G. So these are separate
6:50:57clusters. The clustering starts from the
6:50:59bottom and each data point is treated as
6:51:02a single cluster which we will also
6:51:04understand with the help of an example
6:51:06further but now for simplicity let's
6:51:08take A B C D. So then what happens is
6:51:11that since A and B are grouped as one
6:51:14cluster so this is how I just group them
6:51:17all right? And C and D is grouped as one
6:51:19cluster this is how I group them. E and
6:51:21F are grouped in one single cluster.
6:51:24This is how now
6:51:26which is not been grouped into any of
6:51:28the cluster. So now clustering I said
6:51:30that it is it continues until a single
6:51:33cluster is left. So now I would have to
6:51:36have another level of clustering. That
6:51:38means that E F G since G is very close
6:51:42to E F I will cluster it in one single
6:51:44cluster okay? And since I can see that
6:51:47both these A B and C D pairs these
6:51:49clusters are again together. So I will
6:51:52just cross this line and I will make one
6:51:55single cluster of these four points
6:51:58right? So what I do is since these two
6:52:00are connected I connect them with the
6:52:02help of this line figure that is tree
6:52:04structured dendrogram and this G E and
6:52:08F, they are connected somehow. I connect
6:52:09them. All right? Now, what happens is
6:52:12that I have got two big clusters, and
6:52:14clustering continues until a single
6:52:16cluster is obtained. So, in the end, I
6:52:19will have to cluster everything into a
6:52:22single cluster, and this is how I do
6:52:25that. And to join it, I will again join
6:52:27this entire graph. So, this is when it
6:52:31follows a bottom-up to up approach. This
6:52:34is called as agglomerative hierarchical
6:52:37clustering.
6:52:38Okay? So, now we will go and see an
6:52:41example of this hierarchical
6:52:43agglomerative clustering, right? Okay.
6:52:46So, let us understand what is
6:52:48agglomerative clustering with this
6:52:50example. Now, we see that here the
6:52:53clustering takes place from bottom to
6:52:55up, and we have taken an example of
6:52:57population, wherein we go on clustering
6:52:59until we get population. So, here from
6:53:02the bottom, the individual professions
6:53:04are being plotted, and we see that let's
6:53:07let's take for a convenience the
6:53:09left-hand side. And on the left-hand
6:53:11side, the red dots, as you see, this is
6:53:14individual profession in public sector.
6:53:16And on the another side, which we see in
6:53:19the brown circles, is the private sector
6:53:21employment. So, somehow there's
6:53:23similarity between private sector, so uh
6:53:25they are being clustered as one single
6:53:27cluster, and the private sector as a
6:53:29another cluster.
6:53:31These again are being clustered into one
6:53:33single cluster, and that is employment
6:53:35cluster. Whereas on the another side, we
6:53:37can see another cluster, which is
6:53:39different, and that is unemployed
6:53:41section of cluster of people. Now, they
6:53:44again share one similarity, and that is
6:53:46that they all belong to a single gender,
6:53:49that is male. So, everything is being
6:53:51clustered into one single cluster, that
6:53:53is male. And then again, male and female
6:53:56clusters are being clustered together to
6:53:58form one single cluster, that is
6:54:00population. Similarly, we also divide on
6:54:02the right-hand side the individual
6:54:05professions of women clubbed into
6:54:07private and public sector. And then
6:54:09again, we have separate clusters of
6:54:11employed and unemployed women. And they
6:54:13have been grouped into one single
6:54:15cluster and that is women. Again, we
6:54:17merge the two big clusters into one
6:54:20single cluster that is population. So,
6:54:23this is how the clustering is taking
6:54:25place. The levels are being increasing
6:54:27from bottom to up. So, when we are using
6:54:30this bottom-up approach, this is
6:54:31agglomerative clustering.
6:54:35>> [music]
6:54:39>> Now, many of us have visited retail
6:54:41shops such as Walmart or Target for our
6:54:43household needs. Or let's say that we
6:54:45are planning to buy the new iPhone from
6:54:47Target.
6:54:48What we would typically do is search for
6:54:50the model by visiting the mobile section
6:54:52of the store and then select the product
6:54:54and head towards the billing counter.
6:54:57But in today's world, the goal of the
6:54:59organization is to increase the revenue.
6:55:01Can this be done by just pitching one
6:55:03product at a time to the customer? Now,
6:55:05the answer to this is clearly no.
6:55:07Hence, organization began mining data
6:55:10relating to frequently bought items.
6:55:13So, market basket analysis is one of the
6:55:15key techniques used by large retailers
6:55:18to uncover associations between items.
6:55:21Now, examples could be the customers who
6:55:23purchase bread have a 60% likelihood to
6:55:26also purchase jam.
6:55:28Customers who purchase laptops are more
6:55:30likely to purchase laptop bags as well.
6:55:33They try to find out associations
6:55:35between different items and products
6:55:37that can be sold together, which gives
6:55:40assisting in the right product
6:55:41placement.
6:55:43Typically, it figures out what products
6:55:45are being bought together and
6:55:47organizations can place products in a
6:55:49similar manner.
6:55:50For example, people who buy bread also
6:55:53tend to buy butter, right?
6:55:55And the marketing team at retail stores
6:55:57should target customers who buy bread
6:56:00and butter and provide an offer to them
6:56:02so that they buy a third item, suppose
6:56:05eggs.
6:56:06So, if a customer buys bread and butter
6:56:08and sees a discount offer on eggs, he
6:56:10will be encouraged to spend more and buy
6:56:12the eggs. And this is what market basket
6:56:15analysis is all about. This is what we
6:56:17are going to talk about in this session,
6:56:19which is association rule mining and the
6:56:21a priori algorithm.
6:56:23Now, association rule can be thought of
6:56:25as an if-then relationship. Just to
6:56:28elaborate on that, we have come up with
6:56:31a rule, suppose if an item A is being
6:56:33bought by the customer, then the chances
6:56:35of item B being picked by the customer
6:56:38too under the same transaction ID is
6:56:40found out. You need to understand here
6:56:43that it's not a causality, rather it's a
6:56:46co-occurrence pattern that comes to the
6:56:48force.
6:56:49Now, there are two elements to this
6:56:50rule. First is the if and second is the
6:56:53then.
6:56:54Now, if is also known as antecedent.
6:56:57This is an item or a group of items that
6:57:00are typically found in the item set. And
6:57:02the later one is called the consequent.
6:57:06This comes along as an item with an
6:57:08antecedent group or the group of
6:57:10antecedents are purchased.
6:57:13Now, if you look at the image here A
6:57:14arrow B, it means that if a person buys
6:57:17an item A, then he will also buy an item
6:57:19B. Or he will most probably buy an item
6:57:21B.
6:57:22Now, the simple example that I gave you
6:57:24about the bread and butter and the eggs
6:57:27is just a small example. But what if you
6:57:29have thousands and thousands of items?
6:57:32If you go to any professional data
6:57:34scientist with that data, you can just
6:57:36imagine how much of profit you can make
6:57:39if the data scientist provides you with
6:57:41the right examples and the right
6:57:42placement of the items which you can do.
6:57:45And you can get a lot of insights. That
6:57:47is why association rule mining is a very
6:57:50good algorithm which helps the business
6:57:52make profit. So, let's see how this
6:57:55algorithm works.
6:57:56So, association rule mining is all about
6:57:58building the rules. And we have just
6:58:01seen one rule that if you buy A, then
6:58:05there's a slight possibility or there's
6:58:07a chance that you might buy B also. This
6:58:10type of relationship in which we can
6:58:12find the relationship between these two
6:58:14items is known as single cardinality.
6:58:17But what if the customer who bought A
6:58:20and B also wants to buy C?
6:58:23Or if a customer who bought A, B, and C
6:58:25also wants to buy D? Then in these
6:58:28cases, the cardinality usually increases
6:58:30and we can have a lot of combination
6:58:33around these data.
6:58:35And if you have around 10,000 or more
6:58:37than 10,000 data or items, just imagine
6:58:40how many rules you're going to create
6:58:42for each product. That is why
6:58:44association rule mining has such
6:58:47measures so that we do not end up
6:58:49creating tens of thousands of rules.
6:58:52Now, that is where the a priori
6:58:54algorithm comes in. But before we get
6:58:56into the a priori algorithm, let's
6:58:58understand what's the maths behind it.
6:59:01Now, there are three types of matrices
6:59:03which help to measure the association.
6:59:06We have support, confidence, and lift.
6:59:09So, support is the frequency of item A
6:59:11or the combination of item A or B.
6:59:14It's basically the frequency of the
6:59:16items which we have bought and what are
6:59:18the combination of the frequency of the
6:59:20item we have bought. So, with this, what
6:59:22we can do is filter out the items which
6:59:25have been bought less frequently.
6:59:28This is one of the measures which is
6:59:29support.
6:59:30Now, what confidence tells us?
6:59:32So, confidence gives us how often the
6:59:34items A and B occur together given the
6:59:37number of times A occur.
6:59:39Now, this also helps us solve a lot of
6:59:41other problems because if somebody is
6:59:43buying A and B together and not buying
6:59:45C, we can just rule out C at that point
6:59:47of time.
6:59:48So, this solves another problem is that
6:59:51we obviously do not need to analyze the
6:59:53products which people just buy barely.
6:59:56So, what we can do is according to the
6:59:58sales, we can define our minimum support
7:00:01and confidence. And when you have set
7:00:03these values, we can put these values in
7:00:05the algorithm and we can filter out the
7:00:08data and we can create different rules.
7:00:11And suppose even after filtering, you
7:00:13have like 5,000 rules. And for every
7:00:16item, we create these 5,000 rules. So,
7:00:19that's practically impossible. So, for
7:00:22that, we need the third calculation
7:00:24which is the lift. So, lift is basically
7:00:26the strength of any rule.
7:00:28Now, let's have a look at the
7:00:29denominator of the formula given here.
7:00:32And if you see here, we have the
7:00:34independent support values of A and B.
7:00:37So, this gives us the independent
7:00:39occurrence probability of A and B. And
7:00:42obviously, there's a lot of difference
7:00:44between this random occurrence and
7:00:46association. And if the denominator of
7:00:49the lift is more, what it means is that
7:00:53the occurrence of randomness is more
7:00:55rather than the occurrence because of
7:00:58any association.
7:00:59So, lift is the final verdict where we
7:01:01know whether we have to spend time on
7:01:03this particular rule what we have got
7:01:06here or not. Now, let's have a look at a
7:01:08simple example of association rule
7:01:10mining.
7:01:11So, suppose we have a set of items A, B,
7:01:14C, D, and E and a set of transactions
7:01:17T1, T2, T3, T4, and T5.
7:01:20And as you can see here, we have the
7:01:21transactions T1 in which we have ABC, T2
7:01:24ACD, T3 BCD, T4 ADE, and T5 BCE.
7:01:30Now, what we generally do is create some
7:01:33rules or association rules such as A
7:01:36gives D or C gives A, A gives C, B and C
7:01:40gives A. What this basically means is
7:01:43that if a person buys A, then he's most
7:01:45likely to buy D. And if a person buys C,
7:01:48then he's most likely to buy A. And if
7:01:50you have a look at the last one, if a
7:01:51person buys B and C, he's most likely to
7:01:54buy the item A as well.
7:01:56Now, if we calculate the support,
7:01:58confidence, and lift using these rules,
7:02:00as you can see here in the table, we
7:02:02have the rule and the support,
7:02:04confidence, and the lift values.
7:02:06Now, let's discuss about a priori.
7:02:09So, a priori algorithm uses the frequent
7:02:12item sets to generate the association
7:02:14rule.
7:02:15And it is based on the concept that a
7:02:17subset of a frequent item set must also
7:02:20be a frequent item set itself.
7:02:23Now, this raises the question, what
7:02:24exactly is a frequent item set?
7:02:26So, a frequent item set is an item set
7:02:29whose support value is greater than the
7:02:30threshold value. Now, just now we
7:02:32discussed that the marketing team,
7:02:34according to the sales, have a minimum
7:02:36threshold value for the confidence as
7:02:39well as the support.
7:02:41So, frequent item set is that item set
7:02:43whose support value is greater than the
7:02:44threshold value already specified.
7:02:47Now, example, if A and B is a frequent
7:02:49item set, then A and B should also be
7:02:52frequent item sets individually.
7:02:55Now, let's consider the following
7:02:56transaction to make the things a little
7:02:59easier. Suppose we have transactions 1 2
7:03:023 4 5 and these items are there.
7:03:04So, T1 has 1 3 and 4, T2 has 2 3 and 5,
7:03:08T3 has 1 2 3 5, T4 2 5, and T5 1 3 and
7:03:125. Now, the first step is to build a
7:03:15list of item sets of size one by using
7:03:18this transactional data. And one thing
7:03:20to note here is that the minimum support
7:03:23count, which is given here, is two.
7:03:25Let's suppose it's two.
7:03:27So, the first step is to create item
7:03:29sets of size one and calculate their
7:03:31support values.
7:03:32So, as you can see here, we have the
7:03:33table C1 in which we have the item sets
7:03:361 2 3 4 5, and the support values.
7:03:39If you remember the formula of support,
7:03:41it was frequency divided by the total
7:03:43number of occurrence.
7:03:45So, as you can see here, for the item
7:03:46set 1, the support is three.
7:03:49As you can see here, the item set 1
7:03:51appears in T1, T3, and T5.
7:03:54So, as you can see, its frequency is 1,
7:03:562, and 3.
7:03:57Now, as you can see here, the item set 4
7:04:00has a support of 1 as it occurs only
7:04:02once in transaction 1.
7:04:04But, the minimum support value is two.
7:04:06That's why it's going to be eliminated.
7:04:09So, we have the final table, which is
7:04:10the table F1, in which we have the item
7:04:13sets 1, 2, 3, and 5, and we have the
7:04:15support values 3, 3, 4, and 4.
7:04:19Now, the next step is to create item
7:04:20sets of size two and calculate the
7:04:22support values.
7:04:24Now, all the combination of the item
7:04:25sets in the F1, which is the final table
7:04:29in which you discarded the four, are
7:04:31going to be used for this iteration. So,
7:04:33we get the table C2. So, as you can see
7:04:35here, we have 1, 2, 1, 3, 1, 5, 2, 3, 2,
7:04:395, and 3, 5.
7:04:40Now, if you calculate the support here
7:04:42again, we can see that the item set 1, 2
7:04:45has a support of 1, which is again less
7:04:48than the specified threshold. So, we're
7:04:51going to discard that.
7:04:52So, if we have a look at the table F2,
7:04:55we have 1, 3, 1, 5, 2, 3, 2, 5, and 3,
7:04:595.
7:05:00Again, we're going to move forward and
7:05:02create the item set of size three and
7:05:05calculate the support values.
7:05:07Now, all the combinations are going to
7:05:08be used from the item set F2 for this
7:05:11particular iterations.
7:05:13Now, before calculating support values,
7:05:15let's perform pruning on the data set.
7:05:18Now, what is pruning? Now, after the
7:05:20combinations are being made, we divide
7:05:21C3 item sets to check if there is
7:05:23another subset whose support is less
7:05:26than the minimum support value.
7:05:28That is what frequent item set means.
7:05:31So, if you have a look here, the item
7:05:33sets we have is 1 2 3, 1 2, 1 3, 2 3 for
7:05:38the first one. Because as you can see
7:05:40here, if we have a look at the subsets
7:05:42of 1 2 3, we have 1 {comma} 2 as well.
7:05:46So, we are going to discard this whole
7:05:48item set.
7:05:49Same goes for the second one. We have 1
7:05:512 5. We have 1 2 in that, which was
7:05:53discarded in the previous set or the
7:05:55previous step. That's why we're going to
7:05:57discard that also.
7:05:59Which leaves us with only two factors,
7:06:01which is 1 3 5 item set and the 2 3 5.
7:06:05And the support for this is two and two
7:06:07as well.
7:06:08Now, if we create the table C4 using
7:06:12four elements, we're going to have only
7:06:14one item set, which is 1 2 3 and 5. And
7:06:18if we have a look at the table here, the
7:06:20transaction table, 1 2 3 and 5 appears
7:06:23only once. So, the support is one.
7:06:25And since C4, the support of the whole
7:06:28table C4 is less than two, so we're
7:06:30going to stop here and return to the
7:06:32previous item set that is three.
7:06:34So, the frequent item sets are 1 3 5 and
7:06:372 3 5.
7:06:38Now, let's assume our minimum confidence
7:06:40value is 60%.
7:06:42For that, we're going to generate all
7:06:43the non-empty subsets for each frequent
7:06:46item sets.
7:06:47Now, for I equals 1 {comma} 3 {comma} 5,
7:06:50which is the item set, we get the subset
7:06:521 3, 1 5, 3 5, 1, 3, and 5.
7:06:58Similarly, for 2 3 5, we get 2 3, 2 5, 3
7:07:015, 2, 3, and 5.
7:07:04Now, this rule states that for every
7:07:06subset S of I, the output of the rule
7:07:09gives something like S gives I to S.
7:07:13That implies S recommends I of S.
7:07:16And this is only possible if the support
7:07:18of I divided by the support of S is
7:07:20greater than equal to the minimum
7:07:22confidence value.
7:07:24Now, applying these rules to the item
7:07:26set of F3, we get rule one, which is 1,3
7:07:30gives 1,3,5
7:07:32and 1,3.
7:07:33It means one and three gives five.
7:07:36So, the confidence is equal to the
7:07:39support of 1,3,5 / support of 1,3, that
7:07:44equals 2 / 3, which is 66% and which is
7:07:47greater than the 60%.
7:07:49So, the rule one is selected.
7:07:51Now, if we come to rule two, which is
7:07:531,5, it gives 1,3,5
7:07:56and 1,5. It means if we have one and
7:07:59five, it implies we also going to have
7:08:02three. Now, if we calculate the
7:08:03confidence of this one, we're going to
7:08:05have support 1,3,5 / support 1,5, which
7:08:08gives us 100%, which means rule two is
What is Deep Learning?
7:08:11selected as well. But again, if you have
7:08:13a look at rule five and rule six over
7:08:15here, similarly, if it select three
7:08:18gives 1,3,5 and three, it means if you
7:08:21have three, we also get one and five.
7:08:23So, the confidence for this comes at
7:08:2550%, which is less than the given 60%
7:08:29target, so we're going to reject this
7:08:31rule. And same goes for the rule number
7:08:33six.
7:08:34Now, one thing to keep in mind here is
7:08:37that although the rule one and rule five
7:08:39look a lot similar, they are not.
7:08:42So, it really depends what's on the
7:08:44left-hand side of the arrow and what's
7:08:45on the right-hand side of the arrow.
7:08:47It's the if-then possibility.
7:08:49I'm sure you guys can understand what
7:08:52exactly these rules are and how to
7:08:54proceed with the rules.
7:08:55So, let's see how we can implement the
7:08:57same in Python, right?
7:09:00So, for that, what I'm going to do is
7:09:01create a new Python file and
7:09:06I'm going to use the Jupyter Notebook.
7:09:08You're free to use any sort of IDE.
7:09:11I'm going to name it as Apriori.
7:09:14So, the first thing what we're going to
7:09:16do is we'll be using the online
7:09:19transactional data of a retail store for
7:09:21generating association rules. So,
7:09:23firstly, what we need to do is get the
7:09:25pandas and mlxtend libraries imported
7:09:27and read the file.
7:09:30So, as you can see here, we are using
7:09:32the online retail.xlsx
7:09:34format file.
7:09:36And from mlxtend, we're going to import
7:09:38a priori and association rules. It all
7:09:40comes under mlxtend.
7:09:43So, as you can see here, we have the
7:09:45invoice, the stock code, the
7:09:47description, the quantity, the invoice
7:09:49data, unit price, customer ID, and the
7:09:52country.
7:09:53Now, next in this step, what we're going
7:09:54to do is do data cleanup, which includes
7:09:57removing the spaces from some of the
7:09:59descriptions, and drop the rows that do
7:10:01not have invoice numbers, and remove the
7:10:03credit card transactions, because that
7:10:06is of no use to us.
7:10:13So, as you can see here, the output in
7:10:15which we have like 532,000
7:10:19rows with eight columns.
7:10:21So, after the cleanup, we need to
7:10:23consolidate the items into one
7:10:25transaction per row with each product.
7:10:27For the sake of keeping the data set
7:10:29small, we are only looking at the sales
7:10:31for France.
7:10:36So, as you can see here, we have
7:10:38excluded all the other sales. We are
7:10:40just looking at the sales for France.
7:10:42Now, there are a lot of zeros in the
7:10:44data, but we also need to make sure any
7:10:46positive values are converted to one,
7:10:48and anything less than zero is set to
7:10:49zero.
7:10:52So, as you can see here, we have still
7:10:53392 rows.
7:10:56We're going to encode it and see.
7:10:59Check again.
7:11:00Now that you have structured the data
7:11:02properly, in this step, what we're going
7:11:03to do is generate frequent item sets
7:11:05that have support at least 7%.
7:11:08Now, this number is chosen so that you
7:11:10can get close enough and generate the
7:11:12rules with the corresponding support,
7:11:13confidence, and lift.
7:11:19So, guys, as you can see here, the
7:11:20minimum support is 0.7. And what if we
7:11:23add another constraint on the rules,
7:11:25such as the lift is greater than six and
7:11:28the confidence is greater than 0.8?
7:11:31So as you can see here, we have the
7:11:33left-hand side and the right-hand side
7:11:34of the association rule, which is the
7:11:36antecedent and the consequence.
7:11:39We have the support, we have the
7:11:40confidence, the lift, the leverage, and
7:11:42the conviction.
7:11:43So guys, that's it for this session.
7:11:45That is how you create association rules
7:11:48using the Apriori algorithm, which helps
7:11:50a lot in the marketing business.
7:11:53It runs on the principle of market
7:11:55basket analysis, which is exactly what
7:11:57big companies like Walmart, you have
7:11:59Reliance, and Target too. Even IKEA does
7:12:03it.
7:12:04>> [music]
7:12:10>> But the first thing that we need to
7:12:11focus on is why artificial intelligence.
7:12:13Why do we need artificial intelligence?
7:12:16This with an example. So nowadays, if
7:12:18you have noticed, if your car exceeds
7:12:20the speed limit, so you'll get a letter
7:12:22or basically a challan at your home. How
7:12:24do you think that happens? Do you think
7:12:26that there's a person who is sitting in
7:12:27a chair and actually noting down all the
7:12:29number plates that crosses the speed
7:12:31limit? Well, that is not possible
7:12:33because there might be millions of cars
7:12:35that pass through that road. And at
7:12:36once, there might be many cars that will
7:12:38be passing through that road. So for a
7:12:40human being to actually do this task is
7:12:42next to impossible. Now, let us see
7:12:44another approach to this particular
7:12:46problem. So what we can do, we can
7:12:48actually make use of cameras that will
7:12:50click the picture of the car that
7:12:52exceeds the speed limit. And then we
7:12:54could convert that picture into a text.
7:12:56For example, we have a UK plate.
7:12:59So in this way, the human error, the
7:13:01risk of human error has been reduced.
7:13:03And at the same time, machines, they
7:13:05never get tired. So because of that, you
7:13:07can capture all the images of cars that
7:13:09actually crosses the speed limit.
7:13:11Similarly, you can think of uh many
7:13:13other examples as well. It is used in
7:13:15order to recognize a sign that is in
7:13:16banks if you want to authenticate
7:13:18whether that person is the bank customer
7:13:20or not. Yeah, apart from that, it is
7:13:22even used for self-driving cars as well.
7:13:24So, in US around 30,000 people die every
7:13:27year because of road accidents. So, that
7:13:29can be completely removed if we use the
7:13:31self-driving cars, which is based on the
7:13:33concept of artificial intelligence. And
7:13:35let me tell you guys, you might find it
7:13:36very fascinating that people in MIT are
7:13:39using artificial intelligence in order
7:13:41to predict the future. So, you can
7:13:43imagine why we need artificial
7:13:44intelligence. If you have any questions,
7:13:46any doubts, you can ask me. It is even
7:13:48used in places where humans can't reach.
7:13:50For example, uh deep oceans or
7:13:52navigation in Mars. So, in those places
7:13:55we need machines which are smart enough
7:13:57to carry out tasks.
7:13:58So, this is why we need artificial
7:14:00intelligence. Let us move forward and
7:14:01understand what exactly is artificial
7:14:03intelligence.
7:14:05Now, artificial intelligence, I know the
7:14:06word sounds pretty complex and there are
7:14:08a lot of Hollywood movies that are based
7:14:10on artificial intelligence. If you have
7:14:12seen Terminator or Matrix, all these
7:14:14movies are based on artificial
7:14:15intelligence. But, you don't need to
7:14:17worry about it because till now we
7:14:18haven't reached that level as they have
7:14:20shown in movies like Terminator.
7:14:22But, yeah, the concept is pretty
7:14:23similar. So, basically, we want systems
7:14:26and softwares in such a way that they
7:14:28can imitate the human behavior.
7:14:30Now, what happens in artificial
7:14:31intelligence? Artificial intelligence is
7:14:34accomplished by studying how human brain
7:14:36thinks and how human brain learns,
7:14:38decide, and work while trying to solve a
7:14:41problem. And then we use outcome of this
7:14:43study as the basis of development of
7:14:45intelligent software and systems.
7:14:48So, our major goal is to have systems or
7:14:50softwares that can imitate the human
7:14:53behavior. The way they think, the way
7:14:55they decide, the way they solve a
7:14:57problem. So, in that similar fashion, we
7:14:59want our machines to do that. So, this
7:15:02is basically artificial intelligence in
7:15:03a nutshell.
7:15:05Let us move forward and look at various
7:15:06applications of artificial intelligence.
7:15:09So, this slide basically talks about the
7:15:11application of artificial intelligence.
7:15:13Now, I've listed only three of them, but
7:15:15there are millions of applications. For
7:15:17example, it is used in speech
7:15:18recognition. So, whenever you search
7:15:20something on Google, so you can just
7:15:21tell Google and it'll search it for you.
7:15:23Similarly, it is used for understanding
7:15:24natural language as well as for image
7:15:26recognition as well. And there are many,
7:15:28many other applications in which
7:15:30artificial intelligence finds its use.
7:15:32For example, it can be used in
7:15:34self-driving cars. It can be used in
7:15:36Siri for recommending some products. And
7:15:38even when you go to websites like
7:15:39YouTube or Pandora, YouTube knows which
7:15:41video you want to next. Pandora knows
7:15:43which song you want to listen to. How do
7:15:45you think this happens? It happens all
7:15:47because of artificial intelligence. So,
7:15:49all of these are a few examples of
7:15:51artificial intelligence, but nowadays,
7:15:53it is used almost everywhere, guys.
7:15:55Trust me on that. Now, let us move
7:15:57forward and understand how to achieve
7:15:59artificial intelligence.
7:16:01Now, in order to achieve artificial
7:16:02intelligence, there were a few
7:16:04technologies that First came a machine
7:16:06learning. Now, there were certain
7:16:08limitations of machine learning. In
7:16:10order to overcome those limitations,
7:16:12came a deep learning. Now, let me tell
7:16:14you guys, the concept of artificial
7:16:16intelligence is not new. It was first
7:16:18coined in 1956,
7:16:20but it was just a theoretical concept.
7:16:22Then in '80s and '90s, we were talking
7:16:24about neural networks. But since we
7:16:26didn't have enough computational power,
7:16:29so we couldn't utilize it properly. But
7:16:31in late '90s and 2000s, we started using
7:16:33neural networks for machine learning.
7:16:35Then in 2006, the term deep learning was
7:16:38coined for the first time that overcame
7:16:40the limitations of machine learning. And
7:16:42from 2010, deep learning was used
7:16:44commercially as well. So, this was just
7:16:47a small history about artificial
7:16:48intelligence, machine learning, and deep
7:16:50learning. Now, in order to understand
7:16:52this deep learning, we need to first
7:16:54look at machine learning and what were
7:16:55the biggest limitations of machine
7:16:57learning that led to the evolution of
7:16:58deep learning. So, we'll move forward
7:17:00and understand what exactly is a machine
7:17:03learning.
7:17:04Now, what is machine learning? So,
7:17:05machine learning is nothing but a type
7:17:07of artificial intelligence or you can
7:17:09say a subset of artificial intelligence.
7:17:11And it provides computers with the
7:17:13ability to learn without being
7:17:15explicitly programmed. So, you don't
7:17:16need to hardcode your machine for that.
7:17:19Let us understand this with an example.
7:17:21So, we have a problem statement in which
7:17:23whenever you give a certain input, we
7:17:25need to determine the species of the
7:17:26flower. And what is that input? That
7:17:28input will be sepal length, sepal width,
7:17:31petal length, and petal width. So,
7:17:33whenever we get these four parameters or
7:17:35these four variables, our machine should
7:17:37be able to predict what sort of a flower
7:17:39it is. Now, how do you think that will
7:17:41happen? First, what we need to do, we
7:17:44need to train our machine on the basis
7:17:46of the data that we have. So, in this
7:17:49data, we have sepal length, sepal width,
7:17:51petal length, and petal width, and we
7:17:52have species. So, our machine will learn
7:17:55from this data. It'll determine what
7:17:57should be the length and width of the
7:17:59sepal and petal in order to classify it
7:18:01as setosa or versicolor or other species
7:18:04of flowers as well. Now, what happens
7:18:06next? So, you have trained your data.
7:18:08So, you have trained your machine from
7:18:09the data set. Then, what happens?
7:18:11Whenever you give a new input to this
7:18:13particular machine, it'll predict the
7:18:15species of the flower.
7:18:17So, this is how machine learning works.
7:18:18It is nothing but machine learning in a
7:18:19nutshell.
7:18:21So, basically, I'll just revise it once
7:18:23more.
7:18:24So, you have a data set. So, you split
7:18:26that data into training and testing
7:18:28data. So, what happens with the help of
7:18:30training data? You train your particular
7:18:32machine. And after that, you test it in
7:18:35order to determine the accuracy. And
7:18:37once it is done, whenever you give the
7:18:39new input, it'll predict the outcome or
7:18:41the desired outcome. So, this is how
7:18:43machine learning works, guys. Let us
7:18:45move forward and understand various
7:18:47types of machine learning.
7:18:49So, the first type is called supervised
7:18:51learning. Now, in supervised learning,
7:18:54what happens? You have input variables X
7:18:56and an output variable Y.
7:18:58And you can use an algorithm to learn
7:19:00mapping function from the input to the
7:19:02output. Now, let me simplify it for you.
7:19:04So, what happens in supervised learning,
7:19:06the data that you have already contain
7:19:09the classification. Now, let me talk
7:19:11about the previous example itself. So,
7:19:13from our data set, we knew that if we
7:19:15have this width, this length of our
7:19:17sepal and petal,
7:19:19so that will be the species of flower.
7:19:21So, the classifications are already
7:19:23defined. So, that will be under
7:19:25supervised learning. Now, let me tell
7:19:27you how it actually works.
7:19:28So, you have data. You divide that data
7:19:30into training data as well as test data.
7:19:33So, on the basis of this training data,
7:19:35you train your machine. And after that,
7:19:37you create a model. So, as you can see
7:19:39that this phase is called training
7:19:40phase. And after that, you create a
7:19:42model. Now, in order to check this model
7:19:45to get the accuracy, you have test data.
7:19:47So, you'll pass that test data and
7:19:49you'll see the accuracy. That is nothing
7:19:51but the actual output minus the output
7:19:54that is present in the test data. So,
7:19:56with that, you can get the accuracy. So,
7:19:58this is nothing but uh supervised
7:20:00learning. And if you have any questions,
7:20:02you can ask me right now.
7:20:04So, we'll move forward and understand
7:20:06unsupervised learning.
7:20:07Now, in unsupervised learning, unlike
7:20:09supervised learning, you don't have any
7:20:11predefined classes. So, what happens,
7:20:12you have data. So, on the basis of that
7:20:15data, you try to create your own class.
7:20:18You try to make sure that whatever class
7:20:20you create has high intra-class
7:20:21similarities and have a low inter-class
7:20:24similarities. That means, if I've
7:20:26created two class like this, class one
7:20:28and class two, so the elements of this
7:20:31particular class should have high
7:20:33similarity, but at the same time, it
7:20:35should have low similarity with the
7:20:37elements of class two.
7:20:39So, you can think of examples as well of
7:20:40unsupervised learning. For example, if I
7:20:43have a data about my customers. So, if I
7:20:46have a website and there are millions of
7:20:47visitors on my website, and I want to
7:20:49make sure that I group people on various
7:20:52criteria. For example, I can group
7:20:54people on the basis of willingness to
7:20:56purchase a product that is there on my
7:20:57website or where they are coming from,
7:21:00what is the source, all those things.
7:21:01So, I want to group my customers and I
7:21:04want to make sure that I have certain
7:21:05high priority customers and I have low
7:21:07priority customers and I have medium
7:21:09priority customers. So, with the help of
7:21:10unsupervised learning, I can actually do
7:21:13that. I can make a certain classes of
7:21:14people on whom I should focus more on as
7:21:17compared to the other class. So, this
7:21:18was just an example, guys. You can use
7:21:20it in various other fields as well. So,
7:21:22in marketing, this is how you can use
7:21:24unsupervised learning.
7:21:25So, this brings us to our next uh type
7:21:27of machine learning, which is called a
7:21:29reinforcement learning. Now, this is
7:21:31reinforcement learning, guys. Now, what
7:21:33happens in reinforcement learning, the
7:21:34machine learns by interacting with space
7:21:36or an environment. So, it learns with
7:21:38its experience, with its past
7:21:40experience, and also by new choice
7:21:42exploration. Now, I'll take the analogy
7:21:44of dogs. So, if you have any dog or a
7:21:46pet at your home, so if you have trained
7:21:48your dog in order to get the newspaper,
7:21:50and if it gets it, then you reward it
7:21:51with some chocolate or things that the
7:21:54dog likes, right? So, the dog will know
7:21:56whatever he has done, he's actually
7:21:57rewarded for that. So, it'll continue
7:21:59doing that. But, apart from that, if he
7:22:01does something else, if instead of the
7:22:02newspaper, he brings something else. So,
7:22:04what do you do? You might even punish
7:22:05it. So, because of that, the dog will
7:22:07come to know that it has to get a
7:22:09newspaper every morning.
7:22:10Now, the same example is there in front
7:22:12of your screen. So, you have this
7:22:14machine. So, it has two choices, either
7:22:16to touch the fire or touch the water.
7:22:19Now, first what it does, it goes on and
7:22:21touch the fire. So, because of that, it
7:22:23gets some burning sensation. Now, it has
7:22:24only other option, that is to touch the
7:22:27water. So, when it touch the water, it
7:22:29gets some reward. So, because of that,
7:22:31it'll understand that it does not have
7:22:32to touch fire ever again.
7:22:35Now, there's a diagram that is there in
7:22:36front of your screen. So, what happens
7:22:38you have an agent, all right? That agent
7:22:40performs some action. And on the basis
7:22:42of that action, it'll be exposed to some
7:22:44sort of an environment. Now, if that
7:22:46action is correct, then it'll be
7:22:48rewarded with that. But if it is not,
7:22:50then it will change its choice, and it
7:22:52will again perform some action. So, this
7:22:54process will keep on repeating. So, this
7:22:56is how reinforcement learning works.
7:22:59So, let us move forward and understand
7:23:01when we have machine learning, why do we
7:23:03need deep learning? That is, we'll look
7:23:05at various uh limitations of machine
7:23:07learning.
7:23:08Now, the first limitation is high
7:23:09dimensionality of the data. Now, the
7:23:11data that is now generated is huge in
7:23:14size. So, we have a very large number of
7:23:16inputs and outputs. So, due to that,
7:23:18machine learning algorithms fail. So,
7:23:20they cannot deal with high
7:23:21dimensionality of data, or you can say
7:23:22data with large number of inputs and
7:23:25outputs.
7:23:26Now, there's another problem as well, in
7:23:27which it is unable to solve the crucial
7:23:29AI problems, which can be natural
7:23:31language processing, image recognition,
7:23:32and uh things like that.
7:23:34Now, one of the biggest challenges with
7:23:35machine learning models is feature
7:23:37extraction. Now, let me tell you what
7:23:39are features. So, in statistics, we
7:23:40consider features as variables, but when
7:23:42we talk about artificial intelligence,
7:23:44these variables are nothing but the
7:23:45features.
7:23:46Now, what happens because of that? The
7:23:48complex problems such as object
7:23:50recognition or handwriting recognition
7:23:52becomes a huge challenge for machine
7:23:53learning algorithms to solve. Now, let
7:23:55me give you an example of this uh
7:23:56feature extraction. Suppose, if you want
7:23:58to predict that whether there'll be a
7:24:00match today or not. So, it depends on a
7:24:02various features. It depends on the
7:24:03whether the weather is sunny, whether it
7:24:05is windy, all those things. So, we have
7:24:07provided all those features in our data
7:24:09set. But, we have forgot one particular
7:24:11feature that is humidity. And now, our
7:24:13machine learning models are not that
7:24:14efficient that they will automatically
7:24:16generate that particular feature. So,
7:24:18this is one huge problem, or you can say
7:24:20limitation, with machine learning. Now,
7:24:22obviously, we have limitation, and it
7:24:23won't be fair that if I don't give you
7:24:25the solution to this particular problem.
7:24:27So, we'll move forward and understand
7:24:28how deep learning solves these kind of
7:24:29problems.
7:24:31Now, as you can see that the first line
7:24:32on your slide, which says that deep
7:24:34learning models are capable to focus on
7:24:36the right features by themselves,
7:24:37requiring little guidance from the
7:24:38programmer. So, with the help of little
7:24:40guidance, what these deep learning
7:24:42algorithms can do. They can generate the
7:24:44features on which the outcome will
7:24:46depend on. And at the same time, it also
7:24:48solves the dimensionality problem as
7:24:50well. If you have very large number of
7:24:51inputs and outputs, you can make use of
7:24:53a deep learning algorithm. Now, what
7:24:55exactly is deep learning? Again, since
7:24:57we know that it has been evolved by
7:24:59machine learning and machine learning is
7:25:01nothing but a subset of artificial
7:25:02intelligence. And the idea behind
7:25:03artificial intelligence is to imitate
7:25:05the human behavior. The same idea is for
7:25:07the deep learning as well is to build
7:25:09learning algorithms that can mimic
7:25:11brain.
7:25:12Now, let us move forward and understand
7:25:14deep learning what exactly it is.
7:25:16Now, the deep learning is implemented
7:25:18with the help of neural networks. And
7:25:19the idea or the motivation behind neural
7:25:21networks are nothing but neurons. What
7:25:23are neurons? These are nothing but your
7:25:24brain cells. Now, here's a diagram of
7:25:26neuron. So, we have dendrites here,
7:25:28which are used to provide input to a
7:25:30neuron. As you can see, we have multiple
7:25:32dendrites here. So, these many inputs
7:25:34will be provided to a neuron. Now, this
7:25:35is called cell body and inside the cell
7:25:37body, we have a nucleus, which performs
7:25:39some function. After that, that output
7:25:41will travel through axon and it will go
7:25:44towards the axon terminals. And then,
7:25:46this neuron will fire this output
7:25:48towards the next neuron. Now, the
7:25:50studies tell us that the next neuron now
7:25:52or you can say the two neurons are never
7:25:54connected to each other. There's a gap
7:25:55between them. So, that is called a
7:25:57synapse. So, this is how basically a
7:26:00neuron works like. And on the right hand
7:26:02side of your slide, you can see an
7:26:03artificial neuron. Now, let me explain
7:26:05you that. So, over here, similar to
7:26:07neurons, we have multiple inputs. Now,
7:26:09these inputs will be provided to a
7:26:11processing element like a cell body. And
7:26:14over here in the processing element,
7:26:15what will happen? Summation of your
7:26:17inputs and weights. Now, when it moves
7:26:20on, then what will happen? This input
7:26:22will be multiplied with our weights. So,
7:26:24in the beginning, what happens? These
7:26:25weights are randomly assigned. So, what
7:26:27will happen if I take the example of X1?
7:26:29So, X1 multiplied by W1 will go towards
7:26:32the processing element. Similarly, X2
7:26:34and W2 will go towards the processing
7:26:36element. And, similarly, the other
7:26:38inputs as well. And, then summation will
7:26:40happen which will generate a function of
7:26:41S, that is f of S.
7:26:43After that comes the concept of
7:26:45activation function. Now, what is
7:26:47activation function? It is nothing but
7:26:49in order to provide a threshold. So, if
7:26:51your output is above the threshold, then
7:26:52only this neuron will fire, otherwise it
7:26:54won't fire. So, you can use a step
7:26:56function as an activation function, or
7:26:57you can even use a sigmoid function as
7:26:59your activation function. So, this is
7:27:01how an artificial neuron it looks like.
7:27:03So, a network will be multiple neurons
7:27:05which are connected to each other will
7:27:06form an artificial neural network. And,
7:27:08this activation function can be a
7:27:10sigmoid function or a step function,
7:27:12that totally depends on your
7:27:13requirement.
7:27:14Now, once it exceeds the threshold, it
7:27:16will fire. After that, what will happen?
7:27:18It will check the output. Now, if this
7:27:20output is not equal to the desired
7:27:22output, so these are the actual outputs,
7:27:24and we know the real output. So, we'll
7:27:26compare both of that, and we'll find the
7:27:28difference between the actual output and
7:27:30the desired output. On the basis of that
7:27:32difference, we are again going to update
7:27:34our weights. And, this process will keep
7:27:36on repeating until we get the desired
7:27:39output as our actual output. Now, this
7:27:41process of updating weight is nothing
7:27:43but your back propagation method.
7:27:45So, this is neural networks in a
7:27:47nutshell. So, we'll move forward and
7:27:49understand what are deep networks. So,
7:27:51basically, deep learning is implemented
7:27:53by the help of deep networks, and deep
7:27:54networks are nothing but neural networks
7:27:57with multiple hidden layers. Now, what
7:27:59are hidden layers? Let me explain you
7:28:00that. So, you have inputs that comes
7:28:03here. So, this will be your input layer.
7:28:05After that, some process happens, and
7:28:07it'll go to the next node, or you can
7:28:09say to the hidden layer nodes. So, this
7:28:11is nothing but your hidden layer one.
7:28:13So, every node is interconnected if you
7:28:16can notice. After that, you have one
7:28:18more hidden layer where some function
7:28:19will happen. And, as you can see that
7:28:22again these nodes are interconnected to
7:28:23each other. After this hidden layer two
7:28:26comes the output layer, and this output
7:28:28layer again we are going to check the
7:28:30output whether it is equal to the
7:28:31desired output or not. If it is not, we
7:28:33are again going to update the weights.
7:28:35So, this is how a deep network looks
7:28:37like. Now, there can be multiple hidden
7:28:39layers. There can be hundreds of hidden
7:28:41layers as well. But, when we talk about
7:28:43machine learning, that was not the case.
7:28:45We were not able to process multiple
7:28:47hidden layers when we talk about machine
7:28:49learning. So, because of deep learning,
7:28:51we have multiple hidden layers at once.
7:28:54Now, let us understand this with an
7:28:55example. So, we'll take an image which
7:28:57has four pixels. So, if you can notice,
7:28:59we have four pixels here, among which
7:29:01the top two pixels are bright, that is
7:29:03they are black in color, whereas bottom
7:29:04two pixels are white. Now, what happens?
7:29:07We'll divide these pixels and we'll send
7:29:08these pixels to each and every node. So,
7:29:11for that, we need four nodes. So, this
7:29:13particular pixel will go to this node,
7:29:14it will go to this node, this pixel will
7:29:16go to this node, and finally this pixel
7:29:18will go to this particular node that I'm
7:29:19highlighting with my cursor. Now, what
7:29:21happens? We provide them random weights.
7:29:24So, these white lines actually represent
7:29:25the positive weights, and these black
7:29:27lines represents the negative weights.
7:29:29Now, this particular brightness, when we
7:29:31display high brightness, we'll consider
7:29:32it as negative. Now, what happens? When
7:29:35you see the next output or the next
7:29:36hidden layer, it'll be provided with the
7:29:38input with this particular layer. So,
7:29:40this will provide an input with positive
7:29:42weight to this particular node, and the
7:29:44second input will come from this
7:29:45particular node. Since both of them are
7:29:47positive, so we'll get this kind of a
7:29:49node. Similarly, this node as well. Now,
7:29:51when I talk about these two nodes, the
7:29:52first node over here, so this is getting
7:29:54input from this node as well as from
7:29:56this node. Now, over here we have a
7:29:58negative weight. So, because of that,
7:30:00the value will be negative, and we have
7:30:02represented that with black color.
7:30:04Similarly, over here as well, we're
7:30:06getting one input from here which has a
7:30:07negative weight, and the another input
7:30:09from here which has again has a negative
7:30:10weight. So, accordingly, we get again a
7:30:13negative value here. So, these two
Artificial Neural Network
7:30:14becomes black in color. Now, if you
7:30:17notice what'll happen next, we'll
7:30:19provide one input here, which will be
7:30:21negative and a positive weight, which
7:30:23will be again negative, and this will be
7:30:25also negative and a positive weight. So,
7:30:27that will again come out to be negative.
7:30:29So, that is why we have got this kind of
7:30:31a structure. If you notice this this is
7:30:33nothing but the inverse of this
7:30:34particular image. When I talk about this
7:30:36node over here, we are getting the
7:30:38negative value with a positive weight,
7:30:40which is negative, and a negative value
7:30:41with a negative weight, which is
7:30:42positive. So, we are getting something
7:30:44which is positive here.
7:30:46Now, obviously, I want this particular
7:30:47image to get inverse. I want these black
7:30:50strips to come up. So, what I'll do,
7:30:52I'll actually calculate the inverse by
7:30:54providing a negative weight like this.
7:30:55Over here, I've provided a negative
7:30:57weight, it'll come up. So, when I
7:30:58provide a positive weight, so it'll stay
7:31:00wherever it is. After that, it'll
7:31:02detect, and the output you can see will
7:31:04be a horizontal image, not a solid, not
7:31:06a vertical, not a diagonal, but a
7:31:08horizontal. And after that, we are going
7:31:10to calculate the difference between the
7:31:11actual output and the desired output,
7:31:13and we are going to update the weights
7:31:14accordingly. Now, this is just an
7:31:16example, guys. So, guys, this is one
7:31:18example of deep learning, where what
7:31:19happens, we have images here. We provide
7:31:22these raw data to the first layer to the
7:31:24input layer.
7:31:25Then, what happens, these input layers
7:31:27will determine the patterns of local
7:31:28contrast, or it'll fixate those patterns
7:31:30of local contrast, which means that
7:31:32it'll differentiate on the basis of
7:31:34colors and luminosity and all those
7:31:36things. So, it'll differentiate those
7:31:37things. And after that, in the following
7:31:39layer, what will happen, it'll determine
7:31:41the face features, it'll fixate those
7:31:43face features. So, it'll form nose,
7:31:45eyes, ears, all those things. Then, what
7:31:47will happen, it'll accumulate those
7:31:49correct features for the correct face,
7:31:51or you can say that and fixate those
7:31:52features on the correct face template.
7:31:55So, it'll actually determine the faces
7:31:56here, as you can see it over here. And
7:31:58then, it'll be sent to the output layer.
7:32:00Now, basically, you can add more hidden
7:32:02layers to solve more complex problem.
7:32:04For example, if I want to find out a
7:32:06particular kind of face, for example, a
7:32:08face which has large eyes, or which has
7:32:10light complexion. So, I can do that by
7:32:12adding more hidden layers. And I can
7:32:14increase the complexity also at the same
7:32:16time, if I want to find which image
7:32:18contains a dog. So, for for also, I can
7:32:20have one more hidden layer. So, as and
7:32:22when hidden layer increases, we are able
7:32:23to solve more and more complex problem.
7:32:25So, this is just a general overview of
7:32:27how a deep network looks like. So, we
7:32:29have first patterns of local contrast in
7:32:31the first layer. Then what happens, we
7:32:33fixate these patterns of local contrast
7:32:35in order to form the face features such
7:32:37as eyes, nose, ears, etc. And then we
7:32:39accumulate these features for the
7:32:41correct face and then we determine the
7:32:43image. So, this is how uh deep learning
7:32:46network or you can say deep network
7:32:47looks like.
7:32:48So, we'll move forward and I'll give you
7:32:50some applications of deep learning. So,
7:32:52here are few applications of deep
7:32:53learning. It can be used in self-driving
7:32:55cars. So, you must have heard about
7:32:57self-driving cars. So, what happens,
7:32:59it'll capture the images around it.
7:33:00It'll process that huge amount of data
7:33:02and then it'll decide what action should
7:33:04it take. Should it take left, right?
7:33:05Should it stop? So, accordingly it'll
7:33:07decide what action should it take and
7:33:09that will reduce the amount of accidents
7:33:10that happens every year. Then when we
7:33:12talk about voice control assistants, I'm
7:33:14pretty sure you must have heard about
7:33:15Siri. All the iPhone users know about
7:33:17Siri, right? So, you can tell Siri
7:33:19whatever you want to do. It'll search it
7:33:20for you and display for you.
7:33:22Then when we talk about automatic image
7:33:24caption generation. So, what happens in
7:33:26this, whatever image that you upload,
7:33:27the algorithm is in such a way that
7:33:29it'll generate the caption accordingly.
7:33:31So, for example, if you have say blue
7:33:33colored eyes, so it'll display a blue
7:33:35colored eye caption uh at the bottom of
7:33:37the image.
7:33:38Now, when I talk about automatic machine
7:33:39translation, so we can convert English
7:33:42language into Spanish. Similarly,
7:33:44Spanish to French. So, basically
7:33:46automatic machine translation, you can
7:33:47convert one language to another language
7:33:49with the help of deep learning. And
7:33:51these are just few examples, guys. There
7:33:52are many, many other examples of deep
7:33:54learning. It can be used in game
7:33:56playing. It can be used in many other
7:33:58things. And let me tell you one very
7:33:59fascinating thing that I've told you in
7:34:00the beginning as well. With the help of
7:34:02deep learning, MIT is trying to predict
7:34:04future. So, yeah, I know it is growing
7:34:06exponentially right now, guys.
7:34:14So, this is the problem statement, guys.
7:34:15We need to figure out if the bank notes
7:34:17are real or fake. And for that, we'll be
7:34:19using artificial neural network. And
7:34:21obviously, we need some sort of data in
7:34:23order to train our network. So, let us
7:34:25see how the data set looks like. So,
7:34:27over here I've taken a screenshot of the
7:34:29data set with few of the rows. In it,
7:34:31data were extracted from images that
7:34:33were taken from genuine and forged bank
7:34:35note-like specimens.
7:34:37After that, wavelet transform tools were
7:34:39used to extract features from those
7:34:41images. And these are few features that
7:34:43I'm highlighting with my cursor. And the
7:34:45final column or the last column actually
7:34:47represents the label.
7:34:48So, basically, label tells us to which
7:34:50class that pattern represents, whether
7:34:52that pattern represents a fake note or
7:34:54it represents a real note. Let us
7:34:56discuss these features and labels one by
7:34:58one.
7:34:59So, the first feature or the first
7:35:00column is nothing but variance of a
7:35:02wavelet transformed image. The second
7:35:04column is about skewness. The third is
7:35:06kurtosis of wavelet transformed image.
7:35:08And finally, fourth one is entropy of
7:35:10the image. After that, when I talk about
7:35:12label, which is nothing but my last
7:35:13column, over here if the value is one,
7:35:15that means the pattern represents a real
7:35:17note. Whereas, when value is zero, that
7:35:19means it represents a fake note. So
7:35:21guys, let's move forward and we'll see
7:35:22what are the various steps involved in
7:35:24order to implement this use case.
7:35:26So, over here we'll first begin by
7:35:28reading the data set that we have. We'll
7:35:29define features and labels.
7:35:32After that, we are going to encode the
7:35:33dependent variable. And what is a
7:35:35dependent variable? It is nothing but
7:35:36your label.
7:35:37Then, we are going to divide the data
7:35:39set into two parts, one for training,
7:35:41another for testing.
7:35:42After that, we'll use TensorFlow data
7:35:44structures for holding features, labels,
7:35:46etc. And TensorFlow is nothing but a
7:35:48Python library that is used in order to
7:35:50implement deep learning models or you
7:35:52can say neural networks.
7:35:53Then, we'll write the code in order to
7:35:55implement the model. And once this is
7:35:57done, we will train our model on the
7:35:58training data. We'll calculate the
7:36:00error. The error is nothing but your
7:36:02difference between the model output and
7:36:04the actual output.
7:36:05And we'll try to reduce this error. And
7:36:07once this error becomes minimum, we'll
7:36:09make prediction on the test data and
7:36:11we'll calculate the final accuracy.
7:36:13So guys, let me quickly open my PyCharm
7:36:15and I'll show you how the output looks
7:36:16like.
7:36:18So this is my PyCharm, guys. Over here,
7:36:19I've already written the code in order
7:36:21to execute the use case. I'll go ahead
7:36:23and run this and I'll show you the
7:36:24output.
7:36:29So over here, as you can see, with every
7:36:30iteration, the accuracy is increasing.
7:36:33So let me just stop it right here.
7:36:35All right. Till now, any questions, any
7:36:37doubts with respect to what is our use
7:36:38case, what is the data set about? Any
7:36:41questions, guys? You can go ahead and
7:36:42ask me.
7:36:44Okay, there's a question from Arpan.
7:36:46He's asking, "Can you explain the code?"
7:36:47Definitely, Arpan. I'll be doing that at
7:36:49the end of this class when you are done
7:36:51with all the fundamentals of neural
7:36:52networks. I'll explain you the entire
7:36:53code, how I've written that, and how
7:36:55I've used TensorFlow in order to
7:36:56implement a neural network.
7:36:58I hope I you are satisfied with the
7:37:00answer. Okay, he's fine with it. Any
7:37:02other questions, any other doubts, guys?
7:37:03Just go ahead and ask me. Over here, you
7:37:05don't need to worry about code right
7:37:06now, guys, because I'll explain this
7:37:08later in the session. So what I'll do,
7:37:10I'll open my slides once more and we'll
7:37:12discuss the fundamentals of neural
7:37:13networks that are required in order to
7:37:15implement this use case.
7:37:16So in order to understand why we need
7:37:18neural networks, we are going to compare
7:37:20the approach before and after neural
7:37:21networks. And we'll see what were the
7:37:23various problems that were there before
7:37:25neural networks. So earlier,
7:37:26conventional computers use an
7:37:28algorithmic approach. That is, the
7:37:30computer follows a set of instructions
7:37:33in order to solve a problem. And unless
7:37:35the specific steps that the computer
7:37:37needs to follow are known, the computer
7:37:39cannot solve the problem. So obviously,
7:37:42we need a person who actually knows how
7:37:44to solve that problem and then he or she
7:37:45can provide the instructions to the
7:37:47computer as to how to solve that
7:37:48particular problem, right? So we first
7:37:50should know the answer to that problem,
7:37:52or we should know how to overcome that
7:37:54challenge or problem which is there in
7:37:55front of us. Then only we can provide
7:37:57instructions to the computer.
7:37:59So this restricts the problem-solving
7:38:00capability of conventional computers to
7:38:03problems that we already understand and
7:38:05know how to solve. But what about those
7:38:07problems whose answer we have no clue
7:38:09of? So, that's where our traditional
7:38:11approach was a failure. So, that's why
7:38:14neural networks were introduced. Now,
7:38:15let us see what was the scenario after
7:38:17neural networks.
7:38:18So, neural networks basically process
7:38:20information in a similar way the human
7:38:22brain does.
7:38:23And these networks, they actually learn
7:38:25from examples. You cannot program them
7:38:27to perform a specific task. They will
7:38:29learn from their examples, from their
7:38:31experience. So, you don't need to
7:38:33provide all the instructions to perform
7:38:34a specific task, and your network will
7:38:36learn on its own with its own
7:38:38experience.
7:38:39All right. So, this is what basically
7:38:40neural network does.
7:38:42So, even if you don't know how to solve
7:38:43a problem, you can train your network in
7:38:45such a way that with experience, it can
7:38:47actually learn how to solve the problem.
7:38:50So, that was a major reason why neural
7:38:52networks came into existence.
7:38:54We'll move forward and we'll understand
7:38:56what is the motivation behind neural
7:38:58networks.
7:38:59So, these neural networks are basically
7:39:01inspired by neurons, which are nothing
7:39:02but your brain cells.
7:39:04And the exact working of the human brain
7:39:06is still a mystery, though.
7:39:08So, as I've told you earlier as well
7:39:09that neural networks work like human
7:39:10brain and so the name.
7:39:13And similar to a newborn human baby, as
7:39:15he or she learns from his or her
7:39:17experience, we want our network to do
7:39:19that as well. But, we want it to do very
7:39:21quickly.
7:39:22So, here's a diagram of a neuron.
7:39:24Basically, a biological neuron receives
7:39:26input from other sources, combines them
7:39:29in some way, perform a generally
7:39:31non-linear operation on the result, and
7:39:33then outputs the final result. So, here
7:39:36if you notice these dendrites, these
7:39:37dendrites will receive signals from the
7:39:39other neurons. Then, what will happen?
7:39:41It will transfer it to the cell body.
7:39:43The cell body will perform some
7:39:44function. It can be summation, it can be
7:39:46multiplication. So, after performing
7:39:48that summation on the set of inputs, via
7:39:50axon it is transferred to the next
7:39:52neuron.
7:39:53Now, let's understand what exactly are
7:39:55artificial neural networks.
7:39:58It is basically a computing system that
7:40:00is designed to simulate the way the
7:40:02human brain analyzes and process the
7:40:04information. Artificial neural networks
7:40:06has self-learning capabilities that
7:40:08enable it to produce better results as
7:40:11more data becomes available. So, if you
7:40:13train your network on more data, it will
7:40:14be more accurate.
7:40:16So, these neural networks, they actually
7:40:17learn by example.
7:40:19And you can configure your neural
7:40:20network for specific applications. It
7:40:22can be pattern recognition or it can be
7:40:24data classification, anything like that,
7:40:26all right?
7:40:27So, because of neural networks, we see a
7:40:29lot of new technology has evolved.
7:40:31From translating web pages to other
7:40:33languages to having a virtual assistant
7:40:35to order groceries online to conversing
7:40:37with chatbots. All of these things are
7:40:40possible because of neural networks.
7:40:43So, in a nutshell, if I need to tell
7:40:44you, artificial neural network is
7:40:46nothing but a network of various
7:40:48artificial neurons.
7:40:50All right? So, let me show you the
7:40:51importance of neural network with two
7:40:53scenarios, before and after neural
7:40:55network.
7:40:56So, over here we have a machine and we
7:40:58have trained this machine on the four
7:41:00types of dogs, as you can see where I'm
7:41:02highlighting with my cursor.
7:41:03And once the training is done, we
7:41:05provide a random image to this
7:41:06particular machine which has a dog. But
7:41:09this dog is not like the other dogs on
7:41:11which we have trained our system on.
7:41:13So, without neural networks, our machine
7:41:15cannot identify that dog in the picture,
7:41:17as you can see it over here. Basically,
7:41:19our machine will be confused. It cannot
7:41:21figure out where the dog is. Now, when I
7:41:23talk about neural networks, even if you
7:41:25have not trained our machine on this
7:41:26specific dog, but still it can identify
7:41:29certain features of the dogs that we
7:41:31have trained on and it can match those
7:41:33features with the dog that is there in
7:41:34this particular image and it can
7:41:36identify that dog. So, this happens all
7:41:39because of neural networks. So, this is
7:41:41just an example to show you how
7:41:42important are neural networks. Now, I
7:41:44know you all must be thinking how neural
7:41:47networks work.
7:41:48So, for that, we'll move forward and
7:41:50understand how it actually works.
7:41:52So, over here I'll begin by first
7:41:54explaining a single artificial neuron
7:41:56that is called as perceptron.
7:41:58So, this is an example of a perceptron.
7:42:00Over here we have multiple inputs X1,
7:42:02X2, {dash} {dash} {dash} till Xn. And we
7:42:05have corresponding weights as well. W1
7:42:07for X1, W2 for X2, similarly Wn for Xn.
7:42:10Then what happens, we calculated the
7:42:12weighted sum of these inputs. And after
7:42:15doing that, we pass it through an
7:42:16activation function. This activation
7:42:19function is nothing but it provides a
7:42:20threshold value. So, above that value my
7:42:23neuron will fire, else it won't fire.
7:42:26So, this is basically an artificial
7:42:27neuron. So, when I talk about a neural
7:42:29network, it involves a lot of these
7:42:31artificial neurons with their own
7:42:33activation function and their processing
7:42:35element.
7:42:36Now, we'll move forward and we'll
7:42:38actually understand various modes of
7:42:40this perceptron or single artificial
7:42:42neuron. So, there are two modes in a
7:42:44perceptron. One is training, another is
7:42:45using mode. In training mode, the neuron
7:42:48can be trained to fire for particular
7:42:50input patterns, which means that we'll
7:42:52actually train our neuron to fire on
7:42:54certain set of inputs and to not fire on
7:42:56the other set of inputs. That's what
7:42:58basically training mode is. When I talk
7:43:00about using mode, it means that when a
7:43:01taught input pattern is detected at the
7:43:03input, its associated output becomes the
7:43:05current output, which means that once
7:43:07the training is done and we provide an
7:43:09input on which the neuron has been
7:43:11trained on, so it will detect the input
7:43:14and will provide the associated output.
7:43:16So, that's what basically using mode is.
7:43:18So, first you need to train it, then
7:43:19only you can use your perceptron or your
7:43:21network.
7:43:23So, these were the two modes, guys. And
7:43:24next up we'll understand what are the
7:43:25various activation functions available.
7:43:28So, these are the three activation
7:43:29functions, although there are many more,
7:43:30but I've listed down three. Step
7:43:32function. So, over here the moment your
7:43:34input is greater than this particular
7:43:35value, your neuron will fire, else it
7:43:37won't. Similarly for sigmoid and sign
7:43:39function as well. So, these are three
7:43:41activation functions. There are many
7:43:42more that I've told you earlier as well.
7:43:44So, yeah, these are the three majorly
7:43:45used activation functions. Next up what
7:43:48we are going to do, we are going to
7:43:49understand how a neuron learns from its
7:43:51experience. So, I'll give you a very
7:43:53good analogy in order to understand
7:43:55that. And later on when we talk about
7:43:57the neural networks or you can say
7:43:58multiple neurons in a network, I'll
7:44:00explain you the math behind it. I'll
7:44:02explain you the math behind learning how
7:44:04it actually happens. So, right now I'll
7:44:05explain you with an analogy. And guys,
7:44:07trust me that analogy is pretty
7:44:09interesting.
7:44:10So, I know all of you must have guessed
7:44:12it. So, these are two beer mugs and all
7:44:14of you who love beer can actually relate
7:44:15to this analogy a lot.
7:44:17And I know most of you actually love
7:44:19beer, so that's why I've chosen this
7:44:20particular analogy so that all of you
7:44:22can relate to it.
7:44:24All right, jokes apart. So, fine guys,
7:44:26so there's a beer festival happening
7:44:27near your house.
7:44:29And you want to badly go there. But your
7:44:31decision actually depends on three
7:44:32factors. First is how is the weather,
7:44:34whether it is good or bad. Second is
7:44:37your wife or husband is going with you
7:44:38or not. And the third one is any public
7:44:40transport is available. So, on these
7:44:43three factors your decision will depend
7:44:44whether you will go or not. So, we'll
7:44:46consider these three factors as inputs
7:44:49to our perceptron. And we'll consider
7:44:51our decision of going or not going to
7:44:53the beer festival as our output. So, let
7:44:55us move forward with that. So, the first
7:44:57input is how is the weather, we'll
7:44:58consider it as X1. So, when weather is
7:45:00good, it'll be one and when it is bad,
7:45:02it'll be zero.
7:45:03Similarly, your wife is going with you
7:45:05or not, so that'd be your X2. If she is
7:45:08going then it's one, if she's not going
7:45:10then it's zero. Similarly for public
7:45:11transport, if it is available then it is
7:45:13one, else it is zero.
7:45:14So, these are the three inputs that I'm
7:45:15talking about. Let's see the output. So,
7:45:18output will be one when you're going to
7:45:19the beer festival and output will be
7:45:21zero when you want to relax at home. You
7:45:23want to have beer at home only, you
7:45:25don't want to go outside. So, these are
7:45:27the two outputs, whether you're going or
7:45:28you're not going.
7:45:30Now, what a human brain does. Over here,
7:45:32okay, fine. I need to go to the beer
7:45:34festival, but there are three things
7:45:36that I need to consider. But will I give
7:45:38importance to all these factors equally?
7:45:41Definitely not. There'll be certain
7:45:43factors which will be of high priority
7:45:45for me. I'll focus on those factors
7:45:47more. Whereas few factors won't affect
7:45:50that much to me. All right. So, let's
7:45:52prioritize our inputs or factors. So,
7:45:54here our most important factor is
7:45:56weather. So, if weather is good, I love
7:45:58beer so much that I don't care even if
7:45:59my wife is going with me or not or if
7:46:01there is a public transport available.
7:46:03So, I love beer that much that if
7:46:05weather is good, then definitely I'm
7:46:06going there. That means when X1 is high,
7:46:08output will be definitely high.
7:46:11So, how we do that? How we actually
7:46:12prioritize our factors or how we
7:46:14actually give importance more to a
7:46:16particular input and less to another
7:46:18input in a perceptron or in a neuron?
7:46:21So, we do that by using weights. So, we
7:46:23assign high weights to the more
7:46:24important factors or more important
7:46:27inputs and we assign low weights to
7:46:28those particular inputs which are not
7:46:30that important for us.
7:46:31So, let's assign weights, guys. So,
7:46:33weight W1 is associated with input X1,
7:46:36W2 with X2, and similarly W3 with X3.
7:46:39Now, as I've told you earlier as well
7:46:41that weather is a very important factor,
7:46:42so I'll assign a pretty high weight to
7:46:43weather and I'll keep it at six.
7:46:45Similarly, W2 and W3 are not that
7:46:47important, so I'll keep it as two two.
7:46:49After that, I've defined a threshold
7:46:51value as five, which means that when the
7:46:53weighted sum of my input is greater than
7:46:55five, then only my neuron will fire or
7:46:58you can say then only I'll be going to
7:46:59the beer festival.
7:47:00All right. So, I'll use my pen and we'll
7:47:03see what happens when weather is good.
7:47:06So, when weather is good, our X1 is one.
7:47:09Our weight is six, we'll multiply it
7:47:10with six.
7:47:11Then,
7:47:14if my wife decides that she is going to
7:47:16stay at home and she will probably be
7:47:18busy with cooking and she doesn't want
7:47:20to drink beer with me. So, she's not
7:47:22coming. So, that input becomes zero.
7:47:24Zero into two will actually make no
7:47:26difference because it'll be zero.
7:47:29Then again, there's no public transport
7:47:30available also. Then also this will be
7:47:32zero into two.
7:47:35So, what output I get here?
7:47:37I get here as six.
7:47:40I notice the threshold value that is
7:47:42five. So, definitely six is greater than
7:47:44five.
7:47:46That means my output
7:47:49will be one or you can say my neuron
7:47:51will fire or I'll actually go to the
7:47:53beer festival.
7:47:54So, even if these two inputs are zero
7:47:57for me, that means my wife is not
7:47:58willing to go with me and there is no
7:48:00public transport available, but weather
7:48:02is good, which has very high weight
7:48:03value and it actually matters a lot to
7:48:05me.
7:48:06So, if that is high, it doesn't really
7:48:08matter whether the two inputs are high
7:48:09or not. I'll go to the beer festival.
7:48:11All right? Now, I'll explain you a
7:48:13different scenario. So, over here our
7:48:15threshold was five, but what if I change
7:48:17this threshold to three? So, in that
7:48:20scenario, even if my weather is not
7:48:22good, uh I'll give it a zero. So, zero
7:48:25into six.
7:48:26But, my wife and public transport both
7:48:30are available.
7:48:31All right? So, one into two
7:48:34plus
7:48:35one into two.
7:48:38Which is equal to four.
7:48:41And it is definitely greater than three.
7:48:45Then also my output will be one. That
7:48:48means I will definitely go to the beer
7:48:49festival even if weather is bad.
7:48:52And my neuron will fire. So, these are
7:48:54the two scenarios that I've discussed
7:48:56with you. All right? So, there can be
7:48:57many other ways in which you can
7:48:59actually assign weight to your problem
7:49:02or to your learning algorithm.
7:49:04So, these are the two ways in which you
7:49:05can assign weight and prioritize your
7:49:07inputs or factors on which your output
7:49:09will depend.
7:49:10So, obviously on real life all the
7:49:12inputs or all the factors are not as
7:49:14important for you. So, you actually
7:49:16prioritize them. And how you do that in
7:49:17perceptron, you provide high weight to
7:49:19it. This is just an analogy so that you
7:49:22can relate to a perceptron to a real
7:49:24life. We'll actually discuss the math
7:49:26behind it later in the session as to how
7:49:28a network or a neuron learns. All right?
7:49:31So, how the weights are actually updated
7:49:33and how the output is changing, that all
7:49:36those things we'll be discussing later
7:49:37in this session. But my aim is to make
7:49:40you understand that you can actually
7:49:42relate to a real life problem with that
7:49:44of a perceptron. All right? And in real
7:49:47life problems are not that easy. They
7:49:49are very very complex problems that uh
7:49:51we actually face. So in order to solve
7:49:53those problems, a single neuron is
7:49:55definitely not enough. So we need
7:49:57networks of neuron. And that's where
7:49:59artificial neural network, or you can
7:50:01say multi-layer perceptron, comes into
7:50:03the picture. Now let us discuss that.
7:50:06Multi-layer perceptron or artificial
7:50:07neural network.
7:50:09So this is how an artificial neural
7:50:10network actually looks like. So over
7:50:12here we have multiple neurons in present
7:50:14in different layers. The first layer is
7:50:16always your input layer. This is where
7:50:18you're actually feed in all of your
7:50:19inputs. Then we have the first hidden
7:50:21layer. Then we have second hidden layer,
7:50:24and then we have the output layer.
7:50:25Although the number of hidden layers
7:50:26depend on your application, on what are
7:50:28you working, what is your problem. So
7:50:30that actually determines how many hidden
7:50:32layers you'll have.
7:50:33So let me explain you what is actually
7:50:34happening here. So you provide in some
7:50:36input to the first layer, which is
7:50:38nothing but your input layer. You
7:50:39provide inputs to these neurons. All
7:50:41right? And after some function, the
7:50:43output of these neurons will become the
7:50:45input to the next layer, which is
7:50:46nothing but your hidden layer one. Then
7:50:48these hidden layers also have various
7:50:50neurons. These neurons will have
7:50:51different activation functions. So
7:50:53they'll perform their own function on
7:50:54the inputs that it receives from the
7:50:56previous layer, and then the output of
7:50:58this layer will be the input to the next
7:51:00hidden layer, which is hidden layer two.
7:51:02Similarly, the output of this hidden
7:51:04layer will be the input to the output
7:51:06layer. And finally, we get the output.
7:51:09So this is how basically an artificial
7:51:10neural network looks like. Now let me
7:51:12explain you this with an example.
7:51:14So over here I'll take an example of
7:51:16image recognition using neural networks.
7:51:19So over here what happens, we feed in a
7:51:21lot of images to our input layer.
7:51:23Now this input layer will actually
7:51:25detect the patterns of local contrast.
7:51:28And then we'll feed that to the next
7:51:29layer, which is hidden layer one. So in
7:51:31this hidden layer one, the face features
7:51:34will be recognized. They'll recognize
7:51:36eyes, nose, ears, things like that. And
7:51:38then, that will be again fed as input to
7:51:41the next hidden layer.
7:51:42And in this hidden layer, we'll assemble
7:51:44those features and we'll try to make a
7:51:45face. And then, we'll get the output
7:51:48that is the face will be recognized
7:51:50properly. So, if you notice here, with
7:51:52every layer, we're trying to get a more
7:51:54abstract version or the generalized
7:51:55version of the input. So, this is how
7:51:58basically an artificial neural network
7:52:00work, how it works. All right.
7:52:02And there's a lot of training and
7:52:03learning which is involved that I'll
7:52:04show you now.
7:52:06Training a neural network. So, how we
7:52:07actually train our neural network? So,
7:52:09basically, the most common algorithm for
7:52:11training a network is called back
7:52:12propagation.
7:52:14So, what happens in back propagation?
7:52:16After the weighted sum of inputs and
7:52:17passing through an activation function
7:52:19and getting the output, we compare that
7:52:21output to the actual output that we
7:52:22already know. We figure out how much is
7:52:24the difference. We calculate the error.
7:52:27And based on that error, what we do, we
7:52:28propagate backwards. And we'll see what
7:52:31happens when we change the weight. Will
7:52:33the error decrease or will it increase?
7:52:35And if it increases, when it increases
7:52:37by increasing the value of the variables
7:52:39or by decreasing the value of variables.
7:52:41So, we kind of calculate all those
7:52:43things and we update our variables in
7:52:45such a way that our error becomes
7:52:47minimum. And it takes a lot of
7:52:49iterations. Trust me, guys. It takes a
7:52:51lot of iterations. We get output a lot
7:52:53of times and then we compare it with the
7:52:54model with the actual output. Then
7:52:56again, we propagate backwards. We change
7:52:58the variables. Then again, we calculate
7:52:59the output. We compare it again with the
7:53:01desired output or the actual output.
7:53:03Then again, we propagate backwards. So,
7:53:05this process keeps on repeating until we
7:53:06get the minimum value.
7:53:08All right. So, there's an example that
7:53:10is there in front of your screen. Don't
7:53:11be scared of the terms that I used. I'll
7:53:13actually explain you with an example.
7:53:15So, this is the example over here. We
7:53:16have zero, one, and two as inputs. And
7:53:18our desired output or the output that we
7:53:20already know is zero, one, and four. All
7:53:22right. So, over here, we can actually
7:53:23figure out that desired output is
7:53:25nothing but twice of your input. But I'm
7:53:27training a computer to do that, right?
7:53:29The computer is not a human.
7:53:31So, what happens? I actually initialize
7:53:33my weight. I keep the value as three.
7:53:35So, the model output will be 3 * 0 is 0,
7:53:383 * 1 is 3, 3 * 2 is 6. Now, obviously
7:53:42it is not equal to your desired output.
7:53:43So, we check the error. Now, the error
7:53:46that we have got here is 0, 1, and 2,
7:53:48which is nothing but your difference.
7:53:49So, 0 - 0 is 0, 3 - 2 is 1, 6 - 4 is 2.
7:53:53Now, this is called an absolute error.
7:53:55After squaring this error, we get square
7:53:57error, which is nothing but 0, 1, and 4.
7:54:00All right? So, now what we need to do,
7:54:01we need to update the variables. We have
7:54:03seen that the output that we got is
7:54:05actually different from the desired
7:54:07output. So, we need to update the value
7:54:08of the weight. So, instead of three, our
7:54:11computer makes it as four. After making
7:54:13the value as four, we get the model
7:54:15output as 0, 4, and 8.
7:54:17And then we saw that the error has
7:54:19actually increased. Instead of
7:54:20decreasing, the error has increased. So,
7:54:22after updating the variable, the error
7:54:24has increased. So, you can see that
7:54:26square error is now 0, 4, and 16, and
7:54:28earlier it was 0, 1, and 4. That means
7:54:30we cannot increase the weight value
7:54:32right now. But if we decrease that, make
7:54:34it as two, we get the output, which is
7:54:37actually equal to desired output. But is
7:54:39it always the case that we need to only
7:54:41decrease the weight? Definitely not.
7:54:44So, in this particular scenario,
7:54:45whenever I'm increasing the weight,
7:54:46error is increasing, and when I'm
7:54:47decreasing the weight, error is
7:54:49decreasing. But as I've told you earlier
7:54:51as well, this is not the case every
7:54:52time. Sometimes you need to increase the
7:54:54weight as well. So, how we determine
7:54:56that? All right. Fine, guys. This is how
7:54:58basically a computer decide whether it
7:54:59has to increase the weight or decrease
7:55:01the weight. So, what happens here? This
7:55:02is a graph of square error versus
7:55:04weight.
7:55:05So, over here, what happens? Suppose
7:55:07your square error is somewhere here.
7:55:09And your computer, it starts increasing
7:55:11the weight in order to reduce the square
7:55:13error. And it notices that whenever it
7:55:14increases the weight, square error is
7:55:16actually decreasing.
7:55:17So, it'll keep on increasing until the
7:55:19square error reaches a minimum value.
7:55:22And after that, when it tries to still
7:55:24increase the weight, the square error
7:55:26will increase. So, at that time, our
7:55:28network will recognize that whenever it
7:55:30is increasing the weight after this
7:55:31point, error is increasing. So,
7:55:33therefore, it will stop right there, and
7:55:34that will be our weight value.
7:55:36Similarly, there can be one more
7:55:38scenario. Suppose if we increase the
7:55:40weight, but then also the square error
7:55:41is increasing. So, at that time, we
7:55:44cannot increase the weight. At that
7:55:45time, computer will realize, "Okay,
7:55:46fine. Whenever I'm increasing the
7:55:47weight, the square error is increasing.
7:55:49So, it'll go in the opposite direction."
7:55:51So, it'll start decreasing the weight,
7:55:52and it'll keep on doing that until the
7:55:54square error becomes minimum. And the
7:55:56moment it decreases more, the square
7:55:58error is again increases. So, our
7:56:00network will know that
7:56:01whenever it decreases the weight value,
7:56:04the square error is increasing. So, that
7:56:05point will be our final weight value.
7:56:08So, guys, this is what basically
7:56:09backpropagation in a nutshell is. Fine.
7:56:12So, we'll move forward, and now is the
7:56:14correct time to understand how to
7:56:15implement the use case that I was
7:56:17talking about in the beginning. That is,
7:56:19how to determine whether a node is fake
7:56:20or real. So, for that, I'll open my
7:56:22PyCharm.
7:56:24This is my PyCharm again, guys. Let me
7:56:26just close this. All right.
7:56:28So, this is the code that I've written
7:56:30in order to implement the use case. So,
7:56:32over here, what we do, we import the
7:56:33first important libraries which are
7:56:35required. Matplotlib is used for
7:56:36visualization. TensorFlow, we know, in
7:56:38order to implement the neural network.
7:56:40NumPy for arrays, Pandas for reading the
7:56:42data set. Similarly, scikit-learn for
7:56:44label encoding as well as for shopping,
7:56:46and also to split the data set into
7:56:48training and testing parts.
7:56:49All right. Fine, guys. So, we'll begin
7:56:51by first reading the data set, as I've
7:56:52told you earlier as well when I was
7:56:54explaining the steps. So, what I'll do,
7:56:56I'll use Pandas in order to read the CSV
7:56:58file, which has the data set.
7:57:00After that, I'll define features and
7:57:02labels. So, X will be my feature, and Y
7:57:04will contain my label. So, basically, X
7:57:06includes all the columns apart from the
7:57:08last column, which is the fifth one. And
7:57:10because the indexing starts from zero,
7:57:12that's why we have written zero till
7:57:14fourth. So, it won't include the fourth
7:57:16column. All right? And so, our last
7:57:18column will actually be our label.
7:57:21Then, what we need to do, we need to
7:57:22encode the dependent variable.
7:57:24So, dependent variable, as I've told,
7:57:26nothing but your label. So, I've
7:57:28discussed encoding in TensorFlow
7:57:29tutorial, you can go through it, and you
7:57:31can actually get to know why and how we
7:57:32do that. Then, what we have done, we
7:57:35have uh read the data set. Then, what we
7:57:37need to do is to split our data set into
7:57:38training and testing. And uh these are
7:57:41all optional steps. You can print the
7:57:42shape of your training and test data. If
7:57:44you don't want to do it, it's still
7:57:45fine.
7:57:46Then, we have defined learning rate. So,
7:57:47learning rate is actually the steps in
7:57:50which the weights will be updated, all
7:57:52right? So, that is what basically
7:57:53learning rate is. Then, when we talk
7:57:55about epochs means iterations.
7:57:58Then, we have defined cost history, that
7:57:59will be an empty NumPy array, and its
7:58:02shape will be one, and it will include
7:58:03the float type object. Then, we have
7:58:05defined N dim, which is nothing but your
7:58:07X shape of axis one, which means your
7:58:09column. Then, we'll print that. After
7:58:12that, we have defined the number of
7:58:13classes. So, there can be only two
7:58:14class, whether the note can be fake or
7:58:16it can be real. And this model path I've
7:58:19given in order to save my model. So,
7:58:21I've just given a path where I need to
7:58:23save it. So, I'll just save it here
7:58:24only, in the current working directory.
7:58:26Now is the time to actually define our
7:58:29neural network. So, we'll first make
7:58:31sure that we have defined the important
7:58:33parameters like hidden layers, number of
7:58:35neurons in hidden layers. So, I'll take
7:58:3610 neurons in every hidden layer, and
7:58:37I'm taking four layers like that. Then,
7:58:40X will be my placeholder, and the shape
7:58:42of this particular placeholder is none,
7:58:43{comma} N {underscore} dim. N
7:58:45{underscore} dim value I'll get it from
7:58:47here, and none can be at any value. I'll
7:58:49define one variable W, and I'll
7:58:51initialize it with zeros, and this will
7:58:53be the shape of my weight. Similarly,
7:58:56for bias as well, this will be the
7:58:57particular shape. And there will be one
7:58:59more placeholder Y dash, which will
7:59:01actually be used in order to provide us
7:59:03with the actual output of the model.
7:59:05There'll be one model output, and
7:59:07there'll be one actual output, which we
7:59:08use in order to calculate the
7:59:09difference, right? So, we'll feed in the
7:59:11actual values of the labels in this
7:59:13particular placeholder Y dash.
7:59:16And now we'll define the model. So, over
7:59:18here we have named the function as
7:59:20multilayer perceptron, and in it we'll
7:59:22first define the first layer. So, the
7:59:24first hidden layer, and we are going to
7:59:26name it as layer underscore one, which
7:59:28will be nothing but the a matrix
7:59:30multiplication of X and weights of H1,
7:59:33that is the hidden layer one.
7:59:35And that'll be added to your biases B1.
7:59:37After that, we'll pass it through a
7:59:39sigmoid activation function. Similarly,
7:59:40in layer two as well, matrix
7:59:42multiplication of layer one and weights
7:59:45of H2. So, if you can notice, layer one
7:59:47was the network layer just before the
7:59:50layer two, right? So, the output of this
7:59:52layer one will become input to the layer
7:59:53two. And that's why we have written
7:59:55layer underscore one. It'll be
7:59:56multiplied by weights H2, and then we'll
7:59:58add it with the bias.
8:00:00Similarly, for this particular hidden
8:00:01layer as well, and this particular layer
8:00:03as well. But, over here we are going to
8:00:05use a ReLU activation function instead
8:00:07of sigmoid. Then, we are going to define
8:00:09the weights and biases. So, this is how
8:00:12we basically define weights. This is how
8:00:13we basically define weights. So, weights
8:00:15H1 will be a variable which will be a
8:00:18truncated normal with the shape of N
8:00:20underscore dim and N underscore hidden
8:00:22underscore one. So, these are nothing
8:00:24but your shapes. All right.
8:00:26And after that, what we have done, we
8:00:27have defined biases as well. Then, we
8:00:29need to initialize all the variables.
8:00:31So, all the guys, in brief, let's talk
8:00:34about TensorFlow.
8:00:36Since in TensorFlow, we need to
8:00:37initialize a variable before we use it.
8:00:42That's how we
8:00:43do it. We first initialize it.
8:00:46And then we need to run it. That's when
8:00:48your variables will be initialized.
8:00:51After that, we are going to
8:00:52create a stateful object, and then
8:00:55finally, I'm going to call my model.
8:00:58And then comes it
8:01:00part where the training happens. Cost
8:01:02function. Cost function
8:01:04is nothing but you can say an error that
8:01:07will be calculated between the actual
8:01:09output
8:01:10and the model output.
8:01:13All right, so Y is nothing but our model
8:01:15output and
8:01:17that is nothing but actual output or the
8:01:19output that we already know.
8:01:21All right, and then we are going to use
8:01:23a gradient
8:01:24descent optimizer to reduce the error.
8:01:27Then
8:01:28we are going to create a session object
8:01:30and uh finally we are going to run the
8:01:32session.
8:01:33So,
8:01:34this is how we basically for every
8:01:35calculated change as
8:01:38well as the accuracy that comes after
8:01:40every
8:01:41the epoch on the training data.
8:01:44After we have calculated the accuracy on
8:01:46the training data, we are going to plot
8:01:47it for every
8:01:49accuracy is.
8:01:51And after
8:01:52after plotting that we have accuracy
8:01:53with our tell using the same prediction
8:01:55on the test and after the print the and
8:01:57the mean score. So, let's do this, guys.
8:02:00All right, so training and what
8:02:01See, accuracy epochs see has 99%. So,
8:02:05with every epoch it is actually
8:02:06increasing apart from a couple of
8:02:08instances, it is actually keep on
8:02:09increasing. So, the more data you train
8:02:11your model on, it will be more accurate.
8:02:14Let me just close it. So, now the model
8:02:16has also been saved where I wanted it to
8:02:18be. This is my final test accuracy and
8:02:21this is the mean squared error. All
8:02:23right, so these are the files that will
8:02:24appear once you save your model.
8:02:26These are the four files that I've
8:02:27highlighted. Now, what we need to do is
8:02:29restore this particular model and I've
8:02:32explained this in detail how how to re-
Convolutional Neural Network
8:02:35restore a model that you have already
8:02:37saved. So, over here what I'll take I've
8:02:39taken before to 768.
8:02:41So, all the values in the row of 754 and
8:02:45768 will be fed to our model and our
8:02:48model will make prediction on that. So,
8:02:50let us go ahead and run this.
8:02:53So, when I'm restoring my model, it
8:02:55seems that my model is 100% I'll use a
8:02:57value
8:02:58fed in. So, whatever values that I have
8:03:00actually given as input to my model, it
8:03:02has correctly identified its class,
8:03:04whether it's a
8:03:06fake note or a real note, because fake
8:03:09note and one stands for real note, okay?
8:03:11So, original class is nothing but a set,
8:03:13so it is zero already. And what
8:03:15prediction my model has made is zero,
8:03:17that means it is fake. percent.
8:03:19Similarly, for other values as well.
8:03:23Fine, guys. So, this is how we basically
8:03:24implement the use case that we saw in
8:03:26the beginning.
8:03:27So, in this slide you can notice that
8:03:29I've listed down only two applications,
8:03:30although there are many more.
8:03:32So, neural networks in medicine.
8:03:34Artificial neural networks are currently
8:03:36a very hot research area in medicine,
8:03:38and it is believed that they will
8:03:39receive extensive application to
8:03:42biomedical systems in the next few
8:03:43years. And currently, the research is
8:03:46mostly on modeling parts of human body
8:03:48and uh recognizing diseases from various
8:03:50scans. For example, it can be
8:03:51cardiograms, CAT scans, ultrasonic
8:03:53scans, etc.
8:03:55And uh currently, the research is going
8:03:57uh mostly on uh two major areas. First
8:03:59is modeling and diagnosing the
8:04:00cardiovascular system.
8:04:02So, neural networks are used
8:04:03experimentally to model the human
8:04:05cardiovascular system.
8:04:07Diagnosis can be achieved by building a
8:04:08model of the cardiovascular system of an
8:04:10individual and comparing it with the
8:04:12real-time physiological measurements
8:04:14taken from the patient. And trust me,
8:04:16guys, if this routine is carried out
8:04:18regularly, potential harmful medical
8:04:21conditions can be detected at an early
8:04:23stage and thus, make the process of
8:04:25combating disease much easier.
8:04:27Apart from that, it is currently being
8:04:29used in electronic noses as well.
8:04:31Electronic noses have several potential
8:04:33applications in telemedicine. Now, let
8:04:36me just give you an introduction to
8:04:37telemedicine. Telemedicine is a practice
8:04:39of medicine over long distance via a
8:04:41communication link. So, what the
8:04:43electronic noses will do, they would
8:04:45identify odors in the remote surgical
8:04:47environment. These identified odors
8:04:50would then be electronically transmitted
8:04:52to another site, wherein odor generation
8:04:54system would recreate them.
8:04:57Because the sense of the smell can be an
8:04:58important sense to the surgeon,
8:05:00tele-smell would enhance tele-present
8:05:02surgery.
8:05:04So, these are the two ways in which you
8:05:05can use it in medicine. You can use it
8:05:08in business as well, guys. So, business
8:05:10is basically a diverted field with
8:05:12several general areas of specialization
8:05:14such as accounting or financial
8:05:16analysis. Almost any neural network
8:05:18application would fit into one business
8:05:20area or financial analysis.
8:05:22Now, there is some potential for using
8:05:24neural networks for business purposes
8:05:25including resource allocation and
8:05:27scheduling. I've listed down two major
8:05:29areas where it can be used. One is
8:05:31marketing.
8:05:32So, there is a marketing application
8:05:33which has been integrated with a neural
8:05:35network system.
8:05:37The airline marketing tactician is a
8:05:39computer system made of various
8:05:41intelligent technologies including
8:05:43expert systems. A feedforward neural
8:05:45network is integrated with the AMT,
8:05:47which is nothing but airline marketing
8:05:49tactician, and was trained using
8:05:51backpropagation to assist the marketing
8:05:53control of airline seat allocation.
8:05:56So, it has wide applications in
8:05:58marketing as well.
8:06:00Now, the second area is credit
8:06:01evaluation. Now, I'll give you an
8:06:02example here. The HNC company has
8:06:05developed several neural network
8:06:06applications, and one of them is a
8:06:08credit scoring system which increases
8:06:10the profitability of existing model up
8:06:12to
8:06:13So, these are few applications that I'm
8:06:15telling you guys. Neural network is
8:06:17actually the future.
8:06:19People are talking about neural networks
8:06:21everywhere, and especially after the
8:06:23intro-
8:06:24duction of GPUs and the amount of data
8:06:26that we have now, neural network is
8:06:28actually spreading like plague right
8:06:29now.
8:06:31>> [music]
8:06:37[music]
8:06:40>> So, this is an
8:06:41image of New York's this picture. So,
8:06:43when a human will see this image, he'll
8:06:45a lot of buildings in different colors
8:06:47and stuff like that. But, how are this
8:06:49image? So, So there'll be three
8:06:51channels. Red,
8:06:52another will be green and finally we
8:06:53have blue channel which is popularly
8:06:55known as RGB. So all each of these
8:06:58channels will they have their own
8:06:59respective pixel values as you can see
8:07:01it over here. So when I say size is B
8:07:04cross A cross 3, it means that there are
8:07:07B
8:07:08rows, A columns and three channels. All
8:07:10right? So So if somebody tells you that
8:07:13the size of an image is 28 cross 28
8:07:15cross three pixels, it means that it has
8:07:1728 rows, 28 columns and three channels.
8:07:20So this is how
8:07:21this is for colored images for we have
8:07:23only two channels. So let's move forward
8:07:25and we'll see why can't we use for image
8:07:27classification.
8:07:28So consider an image which has 28 three
8:07:30pixels.
8:07:32So when I feed in this image to a fully
8:07:33con-
8:07:34nected network like this, then the total
8:07:36number of weights required in the fully
8:07:38connected 2,352.
8:07:40You can just go ahead and multiply it
8:07:41your-
8:07:42self. All right?
8:07:44But in real life the images are not that
8:07:46small. All right? So whatever images
8:07:47that we have, they are definitely above
8:07:49200 cross 200 cross three pixels.
8:07:51So if I take an image which has 200
8:07:53cross 200 cross three pixels and I feed
8:07:55it to a fully connected network, that
8:07:57time the number of weights required
8:07:58itself will be 120,000 guys. So we need
8:08:00to deal with such huge amount of
8:08:02parameters and obviously we require more
8:08:04number of neurons. So that can
8:08:05eventually lead to overfitting. So
8:08:07that's why we can't use network for
8:08:09image classification. Let's see why we
8:08:10need convolutional neural networks.
8:08:12Basically in convolutional neural
8:08:14network in the layer will only be
8:08:15connected to a small region of the layer
8:08:17before it. So if you consider this
8:08:18particular neuron which I'm highlighting
8:08:20right now is only connected to three
8:08:21other neurons. Unlike the fully
8:08:23connected network where this particular
8:08:24neuron will be connected to all these
8:08:25five neurons. Because of this we need to
8:08:27handle less amount of weights and in
8:08:29turn we need less number of neurons as
8:08:31well. So let us understand what exactly
8:08:33is convolutional neural network. So
8:08:35convolutional neural networks are
8:08:36special type
8:08:38of feedforward artificial neural
8:08:40networks which is inspired from visual
8:08:43cortex. So visual cortex
8:08:45This but a small region in our brain
8:08:47brain which is present somewhere here
8:08:49where you can see the bulb and basically
8:08:52what happened
8:08:53was an experiment conducted and people
8:08:55got to know that visual cortex is small
8:08:57regions of cells that are sensitive to
8:08:58specific regions of visual field.
8:09:01So what I'm
8:09:02example some neurons in the visual
8:09:03cortex exposed to vertical edges. Some
8:09:05will fire when exposed to horizontal
8:09:07edges. Some will fire when exposed to
8:09:08diagonal edges and that is nothing but
8:09:10the motivation behind convolutional
8:09:12neural network. So now let us
8:09:14convolutional neural network work force.
8:09:15So generally a collect work has three
8:09:17layers convolution labeling layer and
8:09:18fully connected layer. We'll understand
8:09:20each of these layers one by one. We'll
8:09:22take an example of a classifier that can
8:09:24classify an image of an X as well as an
8:09:26O. So with this example we'll be
8:09:28understanding all these four layers. So
8:09:30let's begin guys. Now there are certain
8:09:32trickier cases. So what I mean by that
8:09:33is X can be represented in these four
8:09:36forms as well, right? So these are
8:09:38nothing but the deformed images of X.
8:09:39Similarly for O as well. So these are
8:09:42deformed images. So even I want to
8:09:43classify these images either X or O. All
8:09:46right, because even this is X, this is
8:09:47X, this is X, this is X. But all these
8:09:49are deformed images. But they are in
8:09:52turn X, right? So I want my classifier
8:09:53to classify them as X. So basically
8:09:56that's what I want. So if you can notice
8:09:58here this is a proper image of an X and
8:10:00which is actually equal to this
8:10:02particular X which is a deformed image.
8:10:03Same goes for this O as well. So now
8:10:05what we are going to do is we know that
8:10:07a computer understands an image using
8:10:08numbers at each pixels. So what we'll
8:10:10do, whatever the white pixels that we
8:10:12have we are going to assign a value
8:10:13minus one to it and whatever the black
8:10:15pixels we have we are going to assign a
8:10:16value one to it. When we use normal
8:10:18techniques to compare these two images,
8:10:20one is a proper image of X and another
8:10:21is a deformed image of X, we got to know
8:10:23that a computer is not able to classify
8:10:25the deformed image of X correctly. Why?
8:10:27Because it is comparing it with the
8:10:29proper image of X, right? So when you go
8:10:31ahead and add the pixel values of both
8:10:33of these images you get something like
8:10:35this. So basically our computer is not
8:10:37able to recognize whether it is an X or
8:10:39not. Now what we do with the help of CNN
8:10:41we take small patches of our image. So,
8:10:43these patches or these pieces are known
8:10:46as nothing but features or filters. So,
8:10:48what we do, by finding rough feature
8:10:50matches in roughly the same positions in
8:10:52two images, CNN gets a lot better at
8:10:54seeing the similarity between the whole
8:10:56image matching schemes. What I mean by
8:10:57that is, we have these filters, right?
8:10:59We have these filters that you can see.
8:11:01So, consider this first filter. This is
8:11:03exactly equal to the feature or the part
8:11:05of the image in the deformed image as
8:11:07well. So, this is our proper image and
8:11:08this is our deformed image, all right?
8:11:10Right? So, this particular feature or
8:11:11this particular part of the image is
8:11:13actually equal to this particular part
8:11:14of the image. Same goes for this
8:11:16particular feature or filter as well.
8:11:18And similarly, we have this filter as
8:11:19well, which is actually equal to this
8:11:21particular part of the deformed image,
8:11:24all right? So, let's move forward and
8:11:25we'll see we're taking in our example.
8:11:27So, we'll be considering these three
8:11:29features or filters. This is a diagonal
8:11:31filter, this is again a diagonal filter
8:11:32and this is nothing but a small x. So,
8:11:34we'll take these three filters and we'll
8:11:36move forward. So, what we are going to
8:11:37do is we are going to compare these
8:11:39features, the small pieces of the bigger
8:11:41image, we are going to put it on the
8:11:43input image and if it matches, then the
8:11:45image will be classified correctly. Now,
8:11:47we'll begin, guys. The first layer is
8:11:48convolution layer. So, these are the
8:11:50beginning two steps of this particular
8:11:51layer. First, we need to line up the
8:11:53feature in the image and then multiply
8:11:54image by the corresponding feature
8:11:56pixel. Now, let me explain you with an
8:11:57example. So, this is our first diagonal
8:11:59feature that we'll take. We are going to
8:12:01put this particular feature on our image
8:12:04of x, all right? And we're going to
8:12:05multiply the corresponding pixel value.
8:12:07So, one will be multiplied with one,
8:12:09we'll get one and we'll put it in
8:12:10another matrix. Similarly, we are going
8:12:13to move forward and we're going to
8:12:14multiply minus one with minus one. We're
8:12:16going to multiply minus one with minus
8:12:18one, as you can see. Similarly, we
8:12:19multiply this result, minus one into
8:12:21minus one, then again minus one into
8:12:22minus one. So, we are going to complete
8:12:24this whole process and we're going to
8:12:25finish up this matrix, all right? And
8:12:27once we are done finishing up the
8:12:28multiplication of all the corresponding
8:12:30pixels in the feature as well as in the
8:12:32image, we need to follow two more steps.
8:12:34We need to add them up and divide by the
8:12:36total number of the pixels in the
8:12:37feature. So, what I mean by that is
8:12:39after the multiplication of the
8:12:41corresponding pixel values, what we do,
8:12:43we add all these values, we divide by
8:12:45the total number of pixels, and we get
8:12:47some value, right? And then now our next
8:12:49step is to create a map and put the
8:12:51value of the filter at that particular
8:12:53place. We saw that after multiplying the
8:12:54pixel value of a feature with the
8:12:56corresponding pixel value of with that
8:12:58of our image, we get the output which is
8:13:00one. So, we place one here. Similarly,
8:13:03we are going to move this filter
8:13:05throughout the image. Next up, we are
8:13:06going to move this filter here. After
8:13:08that, we're going to move it here, here,
8:13:09here, everywhere on the image we are
8:13:11going to move it and we're going to
8:13:12follow the same process. All right, so
8:13:14yeah, this is one more example where
8:13:15I've moved my filter in between and
8:13:17after doing that, I've got the output
8:13:19something like this, 1 1 -1 and all. So,
8:13:22over here if you notice, I've got couple
8:13:24of times -1 as well, due to which my
8:13:26output that comes is 0.55, right? So,
8:13:29I'm going to place 0.55 here. Similarly,
8:13:31after moving the pixel after moving the
8:13:33filter throughout the image, I got this
8:13:35particular matrix. All right? And this
8:13:37is for one particular feature. After
8:13:39performing the same process for the
8:13:41other two filters as well, I've got
8:13:43these two values. So, we have these
8:13:45three values after passing through the
8:13:46convolution layer. Let me give you a
8:13:48quick recap of what happens in
8:13:49convolution layer. So, basically we have
8:13:51taken three features, all right? And one
8:13:53by one we'll take one feature, move it
8:13:55through the entire image, and when we
8:13:56are moving it, at that time we are
8:13:57multiplying the pixel value of the image
8:13:59with that of the corresponding pixel
8:14:00value of the filter, adding them up,
8:14:02dividing by the total number of pixels
8:14:04to get the output. So, when we do that
8:14:07for all the filters, we get we got these
8:14:09three outputs, all right? So, let's move
8:14:10forward and we'll see what happens in
8:14:12ReLU layer. So, this is ReLU layer,
8:14:14guys, and people who have gone through
8:14:15the previous tutorial actually know what
8:14:17it is. So, let me just give you a quick
8:14:18introduction of ReLU layer. So, ReLU is
8:14:20nothing but a activation function. All
8:14:22right? So, what I mean by that is it
8:14:23will only activate a node if the input
8:14:26is above a certain quantity. While the
8:14:28input is below zero, the output is also
8:14:30zero, all right? And when the input
8:14:32rises above the certain threshold, it
8:14:34has a linear relationship with the
8:14:36dependent variable. Now, I'll explain
8:14:38you with an example. We have a graph of
8:14:40ReLU function here. So, my function says
8:14:42that when f of x is equal to zero if x
8:14:45is less than zero, and it is equal to x
8:14:47when x is greater than zero. All right?
8:14:49So, whatever values that I have which
8:14:51are below zero will actually in turn
8:14:53become zero, and whatever values that
8:14:54are above zero, our function value will
8:14:57also be equal to that particular value.
8:14:59So, f of x will be equal to x if it is
8:15:01greater than or equal to zero, and it
8:15:03will be zero if it is less than zero.
8:15:05So, if I have x value as minus three, so
8:15:07definitely it is less than zero, so f of
8:15:09x becomes zero. Similarly, if I have
8:15:11minus five x value, then that again it
8:15:12is less than zero, so my f of x value
8:15:14becomes zero. But, when I consider three
8:15:16as my x value, then my f of x becomes
8:15:19equal to x, which is nothing but three.
8:15:20So, over here I'll have three. Again, if
8:15:22I take my x value as five, then
8:15:24obviously it is greater than or equal to
8:15:26zero, then my f of x becomes equal to x,
8:15:30so my f of x value becomes five. So,
8:15:32this is how our ReLU function works. So,
8:15:34why are we using ReLU function here is
8:15:35we want to remove all the negative
8:15:37values from our output that we got
8:15:39through the convolution layer. So, we'll
8:15:41only take the first output that we got
8:15:43by moving one feature throughout the
8:15:45image. So, this is the output that we
8:15:47have got for only one filter. All right?
8:15:49So, over here I'm going to remove all
8:15:50negative values. So, over here you can
8:15:52see that it it was minus point one one
8:15:54before, and I've converted that to zero.
8:15:56Similarly, I'm going to repeat the whole
8:15:57process for the entire matrix. And once
8:16:00I'm done with that, I get this
8:16:01particular value. Now, remember this is
8:16:03only for the output that we got through
8:16:05one feature. All right? So, when we were
8:16:07doing convolution at that time, we were
8:16:08you
8:16:09So, this is the output only for one
8:16:10filter. After doing it for the output of
8:16:12the other two filters as well, we have
8:16:14got these two values more. So, totally
8:16:16we have these three values after passing
8:16:18through ReLU activation function. Next
8:16:20up, we'll see what exactly is pooling
8:16:21layer. So, in pooling layer what we do,
8:16:23we take a window size of two, and we
8:16:25move it across the entire matrix that we
8:16:27have got after passing through ReLU
8:16:28layer. And we take only the maximum
8:16:31value from there so that we can shrink
8:16:33the image. So what we are actually doing
8:16:35is we are reducing the size of our
8:16:37image. So let me explain you with an
8:16:38example. So this is basically one output
8:16:41that we have got after passing through
8:16:42ReLU layer. And over here we have taken
8:16:44a window size of two cross two. So when
8:16:46we keep this window at this particular
8:16:48position, we see that one is the highest
8:16:50value. So we're going to keep one here.
8:16:52And we are going to repeat the same
8:16:54process for this particular window as
8:16:56well. So over here the maximum value is
8:16:570.33 so 0.33 will come. So if you notice
8:17:00here, earlier we had
8:17:02seven cross seven matrix and now we have
8:17:04reduced that to four cross four matrix.
8:17:06So after doing that for the entire
8:17:08image, we have got this as our output.
8:17:11This output we have got after moving our
8:17:13window throughout the image that we have
8:17:16got after passing through ReLU layer,
8:17:17right? And when we repeat this process
8:17:19for all the three outputs that we have
8:17:21got after the ReLU layer, then we get
8:17:23this particular output after pooling
8:17:25layer. Right? So basically we have
8:17:27shrinked our image to a four cross four
8:17:29matrix. Now comes the tricky part. So
8:17:31what we are going to do now is stack up
8:17:33all these layers. So we have discussed
8:17:34convolution layer, ReLU layer and
8:17:36pooling layer. So I'll just give you a
8:17:38brief recap of what all things we have
8:17:39discussed. In convolution layer what we
8:17:41did, we took three features and then
8:17:43after that one by one we moved each
8:17:45filter throughout the image. And when we
8:17:47were moving it, we were continuously
8:17:49multiplying the image pixel value with
8:17:51that of the corresponding filter pixel
8:17:53value and then we were dividing it by
8:17:55the total number of pixels. All right?
8:17:57With that we got three output after
8:17:58passing through the convolution layer.
8:18:00Then those three output we passed
8:18:02through a ReLU layer where we have
8:18:03removed the negative value. All right?
8:18:05And after removing negative value again
8:18:07we have got the three outputs.
8:18:09Then those three outputs we passed
8:18:10through pooling layer. So basically
8:18:11we're trying to shrink our image. And
8:18:13what we did, we took a window size of
8:18:15two cross two, moved it through all the
8:18:17three outputs that we have got through
8:18:19ReLU layer. And after doing that, we
8:18:21were only taking the maximum value pixel
8:18:23value in that particular window and then
8:18:26we were putting it in a different matrix
8:18:27so that we get a shrinked image. And
8:18:29after passing it through pooling layer,
8:18:30we have got a four cross four matrix.
8:18:32Since we took three features in the
8:18:34beginning, so therefore we have got the
8:18:36three outputs after passing through
8:18:37pooling layer. All right. Next up, we
8:18:39are going to stack up all the layers,
8:18:41all right? So, let's do that. So, after
8:18:43passing through convolution, relu and
8:18:44pooling, we have got this four cross
8:18:46four matrix. This was our input image.
8:18:48Now, when we add one more layer of
8:18:50convolution, relu and pooling, we have
8:18:52shrinked our image from four cross four
8:18:53to two cross two as you can notice here.
8:18:55Now, we are going to use fully connected
8:18:57layer. Now, what happens in fully
8:18:58connected layer? The actual
8:18:59classification happens here, guys, okay?
8:19:01So, what we are doing here is we are
8:19:03going to take the shrinked images and
8:19:05put it into a single list. So, basically
8:19:08this is what we have got after passing
8:19:10through two layers of convolution, relu
8:19:12and pooling and this is what we have
8:19:13got. So, basically we're converting into
8:19:15a single list or a vector. How we do
8:19:17that? We take the first value one, then
8:19:18we take 0.55, then we take 0.55, then we
8:19:21take one again. Then we take one, then
8:19:22we take 0.55, 0.55, 0.55. Then we again
8:19:25take 0.55, one, one and 0.55. So, this
8:19:29is nothing but a vector or you can say a
8:19:31list. If you notice here that there are
8:19:33certain values in my list which are high
8:19:35for X and similarly if I repeat the
8:19:37entire process that we have discussed
8:19:39for O, there'll be certain will be high.
8:19:41So, for X we have first, fourth, fifth,
8:19:4410th and 11th element vector values are
8:19:47high. For O we have second, third, ninth
8:19:51and 12th element vector which are high.
8:19:53So, basically we know now if if we have
8:19:55an input image
8:19:57which has a first, fourth, 10th and 11th
8:20:00element vector values high, we know that
8:20:03we can classify it as X. Similarly, if
8:20:05our input image has a list which has the
8:20:07second, third,
8:20:09ninth and 12th element vector values
8:20:12high, then we can classify it as zero.
8:20:13Now, let me explain you with an example.
8:20:15So, after the training is done, after
8:20:17the after doing the entire process for
8:20:19both X and O, you know that our model is
8:20:21trained now, okay? So, we have given one
8:20:23a new input image and that input image
8:20:25passes through all the layers. And once
8:20:26it has passed through all the layers, we
8:20:28have got this 12-element vector. Now, it
8:20:30has 0.9, 0.65, all these values, right?
8:20:33Now, how do we classify it whether it is
8:20:35an X or O? So, what we do, we'll compare
8:20:37this with a list of X and O, right? So,
8:20:39we have got the list in the previous uh
8:20:41slide, if you notice. We have got two
8:20:43different lists for X and O. We are
8:20:45going to compare this new input image
8:20:47list that we have got with that of X and
8:20:49O, right? So, first let us compare that
8:20:51with X. Now, as I've told you earlier as
8:20:54well, for X there are certain values
8:20:55which will be higher, which is nothing
8:20:57but first, fourth, fifth, 10th, and 11th
8:20:59value, right? So, I'm going to sum
8:21:01first, fourth, fifth, 10th, and 11th
8:21:03value and I've got five. 1 + 1 + 1 + 1
8:21:06and + 1. So, five times one, I've got
8:21:08five. And now, I'm going to sum the
8:21:10corresponding values of my input image
8:21:11vector as well. So, the first value is
8:21:130.9. Then, the fourth value is 0.87,
8:21:16fifth value is 0.96, 10th value is 0.89,
8:21:19and the 11th value is 0.94. So, after
8:21:21this doing the sum of these values, I've
8:21:23got 4.56. When I divide this by five, I
8:21:25got 0.91, right? Now, this is for X.
8:21:28Now, when I do the same process for O,
8:21:30so in O, if you notice, I have second,
8:21:32third, ninth, and 12th element vector
8:21:35values is high. So, when I sum these
8:21:37values, I get four.
8:21:38And when I do the sum of the
8:21:40corresponding values in my input image,
8:21:41I've got 2.07. When I divide that by
8:21:44four, I got 4 for 0.51.
8:21:46So, now we notice that 0.91 is a higher
8:21:49value compared to 0.51. So, we have when
8:21:51we have compared our input image with
8:21:52the values of X, we got a higher value
8:21:55than the value that we have got after
8:21:56comparing the input image with the
8:21:58values of O. So, the input image is
8:22:00classified as X. All right, so now let
8:22:02us move towards our use case. So, this
8:22:04is our use case, guys. So, over here,
8:22:06what we are going to do is we are going
8:22:08to train our model on different types of
8:22:11dogs and cats images and then we are
8:22:14going to provide an input and it will
8:22:16classify whether the input is of a dog
8:22:19or a cat. Now, let me tell you the steps
8:22:21involved in it. So, what we are going to
8:22:23do in the beginning is obviously first
8:22:24we need to download the data set. After
8:22:26that we are going to write a function to
8:22:28encode the labels. Labels are nothing
8:22:30but the dependent variable that we are
8:22:32trying to predict. So, in our training
8:22:34data and testing data, obviously we know
8:22:36the labels, right? So, on that basis
8:22:38only we can train our model. So, we are
8:22:39going to encode those labels. After that
8:22:41we'll resize the image to 50 cross 50
8:22:43pixel and we are going to read it as a
8:22:45grayscale image. Then we are going to
8:22:47split the data, 24,000 images for
8:22:50training and 50 for testing. Once this
8:22:52is done, we are going to reshape the
8:22:54data appropriately for TensorFlow. Now,
8:22:56TensorFlow, I think everyone knows about
8:22:58TensorFlow. It's nothing but a Python
8:23:00library for implementing deep learning
8:23:01models. Then we are going to build the
8:23:03model, calculate the loss. It is nothing
8:23:05but categorical cross entropy. Then we
Recurrent Neural Networks
8:23:07are going to reduce the loss by using
8:23:10Adam optimizer with a learning rate set
8:23:11up set to point double zero one. Then we
8:23:14are going to train the train the deep
8:23:15neural network for 10 epochs and finally
8:23:18we are going to make predictions. All
8:23:19right, so I'll just quickly open my
8:23:21PyCharm and I'll show you the code how
8:23:22it looks like.
8:23:24So, this is the code that I've written
8:23:25in order to implement the use case. In
8:23:26the beginning I need to import the
8:23:28libraries that I require.
8:23:31And once it is done, what I mean my
8:23:32training data and the testing data. So,
8:23:36train and test one contains as well as
8:23:39testing data respectively. Then I've
8:23:41taken my image size as 50 and learning
8:23:44rate I've defined here and I've given a
8:23:46name to my model. You can give whatever
8:23:47name you want. All right, the first
8:23:49thing that we saw we need to encode the
8:23:51dependent variable. That's what we are
8:23:52doing here. We are encoding our
8:23:54dependent variable. So, whenever the
8:23:57label is cat, then it will be converted
8:23:59to an array of one comma zero and when
8:24:01it is dog it will be converted to an
8:24:02array of zero comma one. So, why why we
8:24:04are actually encoding the label? Because
8:24:06our code cannot understand the
8:24:07categorical variable. So, we need to
8:24:09encode it. Right? Next, what I'm doing
8:24:11is I'm resizing my image to 50 cross 50
8:24:14and I am converting it to a grayscale
8:24:15image. Right? And once this is done, I'm
8:24:17going to split my data set into training
8:24:20and testing parts.
8:24:21So, yeah, we are basically splitting the
8:24:22data set into two parts for training and
8:24:25testing.
8:24:26And here
8:24:27we are defining a model. So, you can
8:24:29just I can just go ahead and throw in a
8:24:32comment here.
8:24:35Building the model. Yeah. So, so this is
8:24:38where we are building the model. So,
8:24:40basically what we have done here is we
8:24:41have resized our image to 50 cross 50
8:24:44cross one matrix and that is the size of
8:24:46the input that we are using, right?
8:24:48Talking about. Then, what we have done
8:24:50here we have defined two filters and a
8:24:53stride of five with an activation
8:24:55function. After that
8:24:57that we have added a pooling layer, max
8:24:59pool layer. Okay? What we have done, we
8:25:01have repeated the same process, but over
8:25:04here we are taking 64 filters and five
8:25:07passing it through a real activation
8:25:08function. And after that we have
8:25:10have a
8:25:11repeated the 128 filters. After that we
8:25:12have repeated for 64 filters, then for
8:25:1432 filters. Then after that we are using
8:25:17a fully connected layer with 1024
8:25:19neurons. And finally we are using the
8:25:21dropout layer with key probability of
8:25:240.8 to finish our models. This is where
8:25:26our model is actually finished. And then
8:25:28what we are doing is we are using the
8:25:30Adam optimizer to optimize our model.
8:25:33So, basically whatever the loss that we
8:25:35have, we are trying to reduce it. And
8:25:37this is basically for your TensorBoard.
8:25:39So, we are creating some log files and
8:25:41then with that log file TensorBoard will
8:25:43create a pretty fancy graphs for us that
8:25:46helps us to visualize the entire model.
8:25:48And then what we are doing is we are
8:25:50trying to fit the model. And we have
8:25:51defined epochs as 10, that is the number
8:25:53of iterations that will happen will be
8:25:5610. And yeah, so this is pretty much it.
8:25:58Model name we have given. Then input is
8:26:00X score to check the accuracy. Similarly
8:26:04uh the target will be Y test labels
8:26:07associated with that test data will be a
8:26:09Y test and which we have encoded
8:26:11basically. So this is how we are going
8:26:13to actually calculate the accuracy and
8:26:16we'll try to reduce the loss as much as
8:26:17possible in 10 epochs. So till now our
8:26:20model is complete. We are done with it.
8:26:22Next what I'm doing is I'm feeding in
8:26:24some random input from the test data and
8:26:26I'm validating whether my model is
8:26:28predicting it correct or not. All right.
8:26:30So I've already trained the model
8:26:32because it takes a lot of time and yeah
8:26:34I cannot do it here. So I've already
8:26:36trained the model and you can see that
8:26:38the loss that came after the 10th epoch
8:26:41is 0.2973
8:26:43and the accuracy is somewhere around 88%
8:26:45which is pretty good guys and yeah and
8:26:48I've done the prediction on the test
8:26:49data as well. So let me just show it to
8:26:51you that. So this is the prediction that
8:26:53it has done on few of the images in the
8:26:55test data. So yeah it is a cat predicted
8:26:57as cat cat predicted as a cat cat cat
8:26:59cat and dogs as well. There are certain
8:27:01dogs as well.
8:27:04>> [music]
8:27:08>> Why can't we use feed forward networks?
8:27:11Now let us take an example of a feed
8:27:13forward network that is used for image
8:27:14classification. So we have trained this
8:27:17particular network for classifying
8:27:18various images of animals. Now if you
8:27:20feed in an image of a dog it'll identify
8:27:23that image and will provide a relevant
8:27:25label to that particular image.
8:27:26Similarly if you feed in an image of an
8:27:28elephant it'll provide relevant label to
8:27:31that particular image as well. Now if
8:27:32you notice the new output that we have
8:27:34got that is classifying an elephant has
8:27:37no relation with the previous output
8:27:39that is of a dog. Or you can say that
8:27:41the output at time T is independent of
8:27:44output at time T minus one. As we can
8:27:46see that there is no relation between
8:27:48the new output and the previous output.
8:27:50So we can say that in feed forward
8:27:52networks outputs are independent to each
8:27:54other. Now But are few scenarios where
8:27:56we actually need the previous output to
8:27:57get the new output. Let us discuss one
8:28:00such scenario.
8:28:01Now what happens when you read a book?
8:28:03You'll understand that book only on the
8:28:05understanding of your previous words.
8:28:07All right, so if I use a feedforward
8:28:08network and try to predict the next word
8:28:10in a sentence, I can't do that. Why
8:28:12can't I do that? Because my output will
8:28:15actually depend on the previous outputs.
8:28:17But in the feedforward network, my new
8:28:20output is independent of the previous
8:28:21outputs. That is, output at t plus one
8:28:24has no relation with output at t minus
8:28:26two, t minus one and at t. So basically,
8:28:28we cannot use feedforward networks for
8:28:30predicting the next word in a sentence.
8:28:32Similarly, you can think of many other
8:28:34examples where we need the previous
8:28:36output, some information from the
8:28:37previous output, so as to infer the new
8:28:39output. This is just one small example.
8:28:41There are many other examples that you
8:28:43can think of. So we'll move forward and
8:28:45understand how we can solve this
8:28:46particular problem. So over here, what
8:28:48we have done, we have input at t minus
8:28:50one. We'll feed it to our network, then
8:28:53we'll get the output at t minus one.
8:28:55Then at the next time stamp, that is at
8:28:57time t, we have input at time t. That
8:28:59will be given to our network along with
8:29:01the information from the previous time
8:29:03stamp, that is t minus one, and that
8:29:06will help us to get the output at t.
8:29:08Similarly, at output for t plus one, we
8:29:11have two inputs. One is new input that
8:29:13we give. Another is the information
8:29:15coming from the previous time stamp,
8:29:16that is t, in order to get the output at
8:29:19time t plus one. Similarly, it can go
8:29:21on. So over here, I have just written a
8:29:23generalized way to represent it. There's
8:29:25a loop where the information from the
8:29:27previous time stamp is flowing. This is
8:29:29how we can solve this particular
8:29:30challenge. Now let us understand what
8:29:32exactly are recurrent neural networks.
8:29:35So for understanding recurrent neural
8:29:36network, I'll take an analogy. Suppose
8:29:38your gym trainer has made a schedule for
8:29:40you. The exercises are repeated after
8:29:42every third day.
8:29:44Now this is the order of your exercises.
8:29:46First day you'll be doing shoulder,
8:29:47second day you'll be doing biceps, third
8:29:49day you'll be doing cardio. And all
8:29:50these exercises are repeated in a proper
8:29:52order. Now, what happens when we use a
8:29:54feedforward network for predicting the
8:29:56exercise today? So, we'll provide in the
8:29:58input such as day of the week, month of
8:30:00the year, and health status. All right,
8:30:02and we need to train our model or our
8:30:04network on the exercises that we have
8:30:06done in the past. After that, there'll
8:30:07be a complex voting procedure involved
8:30:09that will predict the exercise for us.
8:30:11And that procedure won't be that
8:30:13accurate. So, whatever output we'll get
8:30:15won't be as accurate as we want it to
8:30:17be. Now, what if I change my inputs and
8:30:20I make my inputs as what exercise I've
8:30:22done yesterday? So, if I've done
8:30:23shoulder, then definitely today I'll be
8:30:24doing biceps. Similarly, if I've done
8:30:26biceps yesterday, today I'll be doing
8:30:28cardio. Similarly, if I've done cardio
8:30:30yesterday, today I'll be doing shoulder.
8:30:32Now, there can be one scenario where you
8:30:34are unable to go to gym for 1 day. Due
8:30:36to some personal reasons, you could not
8:30:38go to the gym. Now, what will happen at
8:30:40that time?
8:30:41We'll go one time step back and we'll
8:30:43feed in what exercise that happened day
8:30:45before yesterday. So, if the exercise
8:30:47that happened day before yesterday was
8:30:49shoulder, then yesterday there were
8:30:50biceps exercises. All right, similarly,
8:30:52biceps happened day before yesterday,
8:30:54then yesterday would have been cardio
8:30:56exercises. Similarly, if cardio would
8:30:57have happened day before yesterday,
8:30:59yesterday would have been shoulder
8:31:00exercises. All right, and this
8:31:02prediction, the prediction for the
8:31:04exercise that happened yesterday, will
8:31:06be fed back to our network and these
8:31:08predictions will be used as inputs in
8:31:10order to predict what exercise will
8:31:12happen today. Similarly, if you have
8:31:14missed your gym, say for 2 days, 3 days,
8:31:16or 1 week, so you need to roll back. You
8:31:19need to go to the last day when you went
8:31:21to the gym. You need to figure out what
8:31:23exercise you did on that day, feed that
8:31:25as an input, and then only you'll be
8:31:26getting the relevant output as to what
8:31:28exercise will happen today.
8:31:30Now, what I'll do, I'll convert these
8:31:31things into a vector. Now, what is a
8:31:33vector? Vector is nothing but a list of
8:31:35numbers. All right, so this is the new
8:31:37information, guys, along with the
8:31:39information from the prediction at the
8:31:40previous time step. So, we need both of
8:31:43these in order to get the prediction at
8:31:44time t. Imagine if I've done shoulder
8:31:47exercises yesterday, so this will be
8:31:49one, this will be zero, this will be
8:31:50zero. Now, the prediction that will
8:31:52happen will be biceps exercise because
8:31:53if I have done shoulder yesterday, it's
8:31:55related to biceps. So, my output will be
8:31:57zero, one, and zero. And this is how
8:31:59vectors work, guys. So, I hope you have
8:32:01understood this, guys. Now, this is how
8:32:03a neural network looks like, guys. We
8:32:05have new information along with the
8:32:08information from the previous time step.
8:32:10The output that we have got in the
8:32:11previous time step will certain
8:32:13information from that. We'll feed into
8:32:15our network as inputs, and then that
8:32:17will help us to get the new output.
8:32:19Similarly, this new output that we have
8:32:21got will take some information from
8:32:23that, feed in as an input to our network
8:32:25along with the new information to get
8:32:26the new prediction, and this process
8:32:28keeps on repeating.
8:32:29Now, let me show you the math behind the
8:32:31recurrent neural networks.
8:32:33So, this is the structure of a recurrent
8:32:34neural network, guys. Let me explain you
8:32:36what happens here. Now, consider at time
8:32:38t equals to zero, we have input x
8:32:40naught, and we need to figure out what
8:32:41is x naught. So, according to this
8:32:43equation, h of zero is equal to w i,
8:32:47weight matrix, multiplied by our input x
8:32:49of zero plus w r into h of zero minus
8:32:54one, which is h of minus one, and time
8:32:56can never be negative, so we this
8:32:58particular equation cannot be applied
8:33:00here, plus a bias. So, w i into x of
8:33:03zero plus b h passes through a function
8:33:05g of h to get h of zero over here. After
8:33:08that, I want to calculate y naught. So,
8:33:10for y naught, I'll multiply h of zero
8:33:12with the weight matrix w i, and I'll add
8:33:14a bias to it and pass it through a
8:33:16function g of i to get y naught. Now, in
8:33:18the next time step, that is at time t
8:33:20equals to one, things become a bit
8:33:22tricky. Now, let me explain you what
8:33:24happens here. So, at time t equals to
8:33:26one, I have input x one, I need to
8:33:27figure out what is x one. So, for that,
8:33:29I'll use this equation. So, I'll
8:33:31multiply w i, that is the weight matrix,
8:33:34by the input x one plus w r into h of
8:33:38one minus one, which is zero. H of zero,
8:33:40we know what we got from here. So, WR
8:33:42into H of zero plus the bias, pass it
8:33:45through a function G of H to get the
8:33:47output as H1. Now, this H1 will use to
8:33:50get Y1. We'll multiply H1 with WY plus a
8:33:53bias and we'll pass it through a
8:33:55function G of Y to get Y1.
8:33:57Similarly, the next time stamp, that is
8:33:59at time T equals to two, we have input
8:34:01X2. We need to figure out what will be
8:34:03H2. So, we'll multiply the weight matrix
8:34:05WI with X of two plus WR into H of one
8:34:08that we have got here plus B of H and
8:34:11pass it through a function G of H to get
8:34:13H of two. From H of two, we'll calculate
8:34:15Y of two. WY into H of two plus BY, that
8:34:18is the bias, pass it through a function
8:34:20G of Y to get Y2. And this is how
8:34:22recurrent neural network works, guys.
8:34:24Now, you must be thinking how to train a
8:34:26recurrent neural network.
8:34:28So, a recurrent neural network uses back
8:34:29propagation algorithm for training. But
8:34:31back propagation happens for every time
8:34:34stamp. That is why it is commonly called
8:34:36as back propagation through time.
8:34:38Over here, I won't be discussing back
8:34:39propagation in detail. I'll just give
8:34:41you a brief introduction of what it is.
8:34:44Now, with back propagation, there are
8:34:45certain issues, namely vanishing and
8:34:47exploding gradients. Let us see those
8:34:49one by one.
8:34:50So, in vanishing gradient, what happens?
8:34:52When you use back propagation, you tend
8:34:54to calculate the error, which is nothing
8:34:56but the actual output that you already
8:34:58know minus the model output, output that
8:35:01you got through your model, and the
8:35:02square of that.
8:35:03So, you figure out the error. With that
8:35:05error, what do you do? You tend to find
8:35:08out the change in error with respect to
8:35:10change in weight or any variable. So,
8:35:13we'll call it weight here. So, change of
8:35:15error with respect to weight multiplied
8:35:17by learning rate will give you the
8:35:18change in weight. Then you need to add
8:35:20that change in weight to the old weight
8:35:22to get the new weight. All right? So,
8:35:25obviously, what we are trying to do, we
8:35:26are trying to reduce the error. So, for
8:35:28that, we need to figure out what will be
8:35:30the change in error if my variables are
8:35:32changed, right? So, that way we can get
8:35:34the change in in variable and add it to
8:35:36our old variable to get the new
8:35:37variable. Now, over here, what can
8:35:39happen if the value dE by dW, that is a
8:35:42gradient, or you can say the rate of
8:35:44change of error with respect to our
8:35:45variable weight, becomes very small than
8:35:47one, like it is 0.00 something. So, if
8:35:50you multiply that with the a learning
8:35:52rate, which is definitely smaller than
8:35:54one, then you get the change of weight,
8:35:56which is negligible. All right? So,
8:35:59there might be certain examples where,
8:36:00you know, you are trying to predict,
8:36:02say, a next word in a sentence, and that
8:36:03sentence is pretty long. For example, if
8:36:05I say, "I went to France {dash} {dash}
8:36:08{dash} I went to France." Then there are
8:36:10certain words. Then I say, "Few of them
8:36:13speak {dash}." Now, I need to predict
8:36:15speak, what will come after speak. So,
8:36:18for that, I need to go back in time and
8:36:20check what was the context, which will
8:36:22be very complex. And due to that,
8:36:24there'll be a lot of iterations. And
8:36:26because of that, this error, this change
8:36:28in weight, will become very small, very
8:36:31small. So, the new weight that we'll get
8:36:32will be actually almost equal to your
8:36:35old weight. So, there won't be any
8:36:37updation of weight that will be
8:36:38happening. And that is nothing but your
8:36:40vanishing gradient. All right, I'll
8:36:42repeat it once more. So, what happens in
8:36:44back propagation, you first calculate
8:36:46the error. This error is nothing but the
8:36:48difference between the actual output and
8:36:49the model output and the square of that.
8:36:52With that error, we figure out what will
8:36:53be the change in error when we change a
8:36:55particular variable, say, weight. So, dE
8:36:58by dW, multiply it with learning rate to
8:37:00get the change in the variable or change
8:37:02in the weight. Now, we'll add that
8:37:03change in the weight to our old weight
8:37:05to get the new weight. This is back
8:37:07propagation, is guys, all right? I'm
8:37:08just giving you a small introduction to
8:37:10back propagation. Now, consider a
8:37:12scenario where you need to predict the
8:37:14next word in a sentence. And your
8:37:15sentence is something like this. "I have
8:37:18been to France." Then there are a lot of
8:37:20words. After that, few people speak. And
8:37:24then you need to predict what comes
8:37:25after speak. Now, if I need to do that,
8:37:27I need to go back and understand the
8:37:29context, what is it talking about?
8:37:32And that is nothing but your long-term
8:37:34dependencies. So, what happens during
8:37:35long-term dependencies if this DE by DW
8:37:38becomes very small? Then, when you
8:37:40multiply it with N, which is again
8:37:41smaller than one, you get delta W, which
8:37:44will be very, very small. That will be
8:37:46negligible. So, the new weight that
8:37:48you'll get here will be almost equal to
8:37:50your old weight. So, I hope you're
8:37:52getting my point. So, this new weight
8:37:54So, there will be no updation of
8:37:56weights, guys. This new weight will
8:37:58definitely be will always be almost
8:38:00equal to our old weight. There won't be
8:38:02any learning here. So, that is nothing
8:38:04but your vanishing gradient problem.
8:38:06Similarly, when I talk about exploding
8:38:08gradient, it is just the opposite of
8:38:09vanishing gradient. So, what happens
8:38:11when your gradient or DE by DW becomes
8:38:13very uh large, becomes greater than
8:38:15greater than one? All right? And you
8:38:17have some long-term dependencies. So, at
8:38:19that time, your DE by DW will keep on
8:38:22increasing. Delta W will become large.
8:38:24And because of that, your weights, the
8:38:26new weight with that will come will be
8:38:28very different from your old weight. So,
8:38:30these two are the problems with back
8:38:31propagation. Now, let us see how to
8:38:33solve these problems.
8:38:35Now, exploding gradients can be solved
8:38:37with the help of truncated BPTT, back
8:38:38propagation through time. So, instead of
8:38:40starting back propagation at the last
8:38:42time stamp, we can choose a smaller time
8:38:44stamp like 10. Or we can clip the
8:38:47gradients at a threshold. So, there can
8:38:48be a threshold value where we can, you
8:38:50know, clip the gradients. And we can
8:38:52adjust the learning rate as well. Now,
8:38:53for vanishing gradient, we can use a
8:38:55ReLU activation function. We have
8:38:56discussed ReLU activation function in
8:38:58artificial neural network tutorial,
8:38:59guys. Similarly, we can also use LSTM
8:39:02and GRUs. In this tutorial, we'll be
8:39:04discussing LSTMs that are long
8:39:06short-term memory units. Now, let us
8:39:09understand what exactly are LSTMs.
8:39:12So, guys, we saw what are the two
8:39:13limitations with the recurrent neural
8:39:15networks. Now, we'll understand how we
8:39:17can solve that with the help of LSTMs.
8:39:19Now, what are LSTMs? Long short-term
8:39:21memory networks, usually called as
8:39:23LSTMs, are nothing but a special kind of
8:39:25recurrent neural network. And these
8:39:27recurrent neural networks are capable of
8:39:29learning long-term dependencies. Now,
8:39:31what are long-term dependencies? I've
8:39:33discussed on the previous slide, but
8:39:35I'll just explain it to you here as
8:39:36well. Now, what happens sometimes we
8:39:38only need to look at the recent
8:39:39information to perform the present task.
8:39:42Now, let me give you an example.
8:39:43Consider a language model trying to
8:39:45predict the next word based on the
8:39:47previous ones. If we are trying to
8:39:49predict the last word in the sentence,
8:39:51say, "The clouds are in the sky." So, we
8:39:54don't need any further context. It's
8:39:55pretty obvious that the next word is
8:39:57going to be sky. Now, in such cases
8:39:59where the gap between the relevant
8:40:01information and the place that it's
8:40:03needed is small, RNNs can learn to use
8:40:06the past information. And at that time,
8:40:08there won't be such problems like
8:40:09vanishing and exploding gradient. But,
8:40:11there are few cases where we need more
8:40:14context. Consider trying to predict the
8:40:16last word in the text, "I grew up in
8:40:19France." Then, there are some words.
8:40:20After that comes, "I speak fluent
8:40:23French." Now, recent information
8:40:25suggests that word is probably the name
8:40:27of a language. But, if we want to narrow
8:40:29down which language, we need the context
8:40:32of France from further back. And it's
8:40:35entirely possible for the gap between
8:40:37the relevant information and the point
8:40:39where it is needed to become very large.
8:40:41And this is nothing but long-term
8:40:42dependencies. And the LSTMs are capable
8:40:45of handling such long-term dependencies.
8:40:47Now, LSTMs also have a chain-like
8:40:50structure like recurrent neural
8:40:51networks. Now, all the recurrent neural
8:40:53networks have the form of a chain of
8:40:54repeating modules of neural networks.
8:40:56Now, in standard RNNs, the repeating
8:40:58module will have a very simple structure
8:41:00such as a single tan h layer that you
8:41:01can see. Now, this tan h layer is
8:41:03nothing but a squashing function. Now,
8:41:05what I mean by squashing function is to
8:41:07convert my values between minus one and
8:41:10one. All right, that's why we use tan h.
8:41:12And this is an example of an RNN. Now,
8:41:15we'll understand what exactly are LSTMs.
8:41:17Now, this is a structure of an LSTM. All
8:41:20If you notice, LSTM also have a chain
8:41:22like structure. But the ripple has
8:41:24different structures. Instead of having
8:41:26single neural network here, there are
8:41:28four interacting in a very special way.
8:41:30Now, the key to LSTM is the cell state.
8:41:32Now, this particular line that I'm
8:41:34highlighting, this is what what is
8:41:36called the cell state. The horizontal
8:41:38line running through the top of the
8:41:39diagram. So, this is nothing but your
8:41:40cell state. Now, you can consider the
8:41:42cell state as a kind of a conveyor belt.
8:41:45It runs straight down the entire chain
8:41:47with only some minor linear
8:41:48interactions. Now, what I'll do, I'll
8:41:50give you a walk through of LSTM step by
8:41:52step, all right? So, we'll start with
8:41:54the first step.
8:41:55All right, guys. So, the first step in
8:41:57our LSTM is to decide what information
8:41:59we are going to throw away from the cell
8:42:01state. And you know what is the cell
8:42:03state, right? I've discussed in the
8:42:04previous slide. Now, this decision is
8:42:06made by the sigmoid layer. So, the layer
8:42:09that I'm highlighting with my cursor, it
8:42:10is the sigmoid layer. Called the forget
8:42:12gate layer. It looks at HT minus one,
8:42:15that is the information from the
8:42:16previous time step, and XT, which is the
8:42:19new input, and outputs a number between
8:42:21zeros and ones for each number in the
8:42:23cell state, CT minus one, which is
8:42:25coming from the previous time step. A
8:42:27one represents completely keep this,
8:42:29while a zero represents completely get
8:42:31rid of this. Now, if we go back to our
8:42:33example of a language model trying to
8:42:35predict the next word based on all the
8:42:37previous ones, in such a problem, the
8:42:39cell state might include the gender of
8:42:41the present subject so that the correct
8:42:43pronouns can be used. When we see a new
8:42:45subject, we want to forget the gender of
8:42:47the old subject, right? We want to use
8:42:50the gender of the new subject. So, we'll
8:42:52forget the gender of the previous
8:42:53subject here. This is just an example to
8:42:56explain you what is happening here.
8:42:58Uh now, let me explain you the equations
8:42:59which I've written here. So, FT will be
8:43:02uh combining with the cell state later
8:43:04on, that I'll tell you. So, currently,
8:43:06FT will be nothing but the weight matrix
8:43:09multiplied by HT minus one and XT, and
8:43:13uh plus the bias, and this equation is
8:43:15passed through a sigmoid layer. All
8:43:17right? And we get an output that is zero
8:43:19and one. Zero means completely get rid
8:43:21of this and one means completely keep
8:43:23this. All right, so this is what
8:43:24basically is happening in the first
8:43:26step. Now, let us see what happens in
8:43:28the next step. So, the next step is to
8:43:30decide what information we are going to
8:43:32store. In the previous step, we decided
8:43:34what information we are going to keep,
8:43:35but here we are going to decide what
8:43:37information we are going to store here.
8:43:39All right, what new information we are
8:43:41going to store in the cell state. Now,
8:43:42this has two parts. First, a sigmoid
8:43:45layer, this is called a sigmoid layer
8:43:46and which is also known as an input gate
8:43:48layer, decide which values will update.
8:43:51All right, so what values we need to
8:43:52update. Then there's also a tan h layer
8:43:54that creates a vector of the candidate
8:43:56values c bar of t minus one that will be
8:44:00added to the state later on. All right,
8:44:02so let me explain it to you in a simpler
8:44:03terms. So, whatever input that we are
8:44:05getting from the previous time stamp and
8:44:07the new input, it will be passed through
8:44:09a sigmoid function, which will give us i
8:44:11of t. All right, and this i of t will be
8:44:14multiplied by c t but coming from the
8:44:17previous time stamp and the new input
8:44:19with that is passed through a tan h that
8:44:21will result in c t. And this will be
8:44:23later added on to our cell state. In the
8:44:25next step, we'll combine these two to
8:44:27update the states. Now, let me explain
8:44:28the equations. So, i of t will be what?
8:44:31Weight matrix and then we have h t minus
8:44:33one comma x t multiplied by the weight
8:44:35matrix plus the bias pass it through a
8:44:37sigmoid function, we get i of t. c bar
8:44:39of t will get by passing a weight matrix
8:44:41h t minus one x t plus bias through a
8:44:44tan h square function and we'll get c
8:44:46bar of t. All right, so as I've told you
8:44:48earlier as well in the next step, we'll
8:44:49combine these two to update the state.
8:44:51Let us see how we do that. So, now is
8:44:54the time to update the old cell state c
8:44:56t minus one with the new cell state c t.
8:44:59All right, in the previous steps, we
8:45:00have already decided what to do. We just
8:45:02need to actually do it. So, what we'll
8:45:04do, we'll multiply the old cell state c
8:45:06t minus one with f t that we got in the
8:45:08first step for getting the things that
8:45:10we decided to forget earlier in the
8:45:12first step if you can recall. Then what
8:45:14we do, we add it to IT and CT. Then we
8:45:18add it by the term that will come after
8:45:20multiplication of IT and C bar T. And
8:45:22this new candidate value scaled by how
8:45:24much we decided to update each state
8:45:26value. All right? So, in the case of the
8:45:29language model that we are discussing,
8:45:30this is where we would actually drop the
8:45:32information about the old subject gender
8:45:35and add the new information as we
8:45:36decided in the previous steps. So, I
8:45:38hope you are able to follow me guys. All
8:45:40right? So, let us move forward and we'll
8:45:42see what is the next step. Now, our last
8:45:44step is to decide what we are going to
8:45:46output. And this output will depend on
8:45:48our cell state, but [snorts] it will be
8:45:50a filtered version. Now, finally what we
8:45:52need to do is we need to decide what we
8:45:53are going to output. And this output
8:45:55will be based on our cell state. First,
8:45:57we need to pass HT minus one and XT
8:45:59through a sigmoid activation function so
8:46:02that we get output that is OT. All
8:46:04right? And this OT will be in turn
8:46:06multiplied by the cell state after
8:46:08passing it through an NH squashing
8:46:10function or an activation function. And
8:46:12why we do that? Just to push the values
8:46:14between minus one and one. So, after
8:46:17multiplying OT, that is this value, and
8:46:20a tan at CT, we'll get the output H2,
8:46:23which will be our new output. And that
8:46:25will only output the part that we
8:46:27decided to. Whatever we have decided in
8:46:29the previous steps, it will only output
8:46:30that value. All right? Now, I'll take
8:46:32the example of that language model
8:46:34again. Since it just saw a subject, it
8:46:36might want to output information
8:46:38relevant to a verb and in case that's
8:46:40what is coming next.
8:46:42For example, it might output whether the
8:46:44subject is singular or plural. So, that
8:46:46we know what form of a verb should be
8:46:48conjugated into. All right? And uh you
8:46:51can see from the uh you can see the
8:46:52equations as well. Again, we have a
8:46:54sigmoid function. Then that uh whatever
8:46:57output we get from there, we multiply it
8:46:59with tan at CT to get the new output.
8:47:01All right, guys? So, this is basically
8:47:03uh LSTMs in a nutshell. So, in the first
8:47:06step, we decided what we need to forget.
8:47:08In the next step, we decided what are we
8:47:10going
8:47:11to our cell state, what new information
8:47:13going to add to cell state, and we were
8:47:15taking example of the gender throughout
8:47:17this whole process. All right? And in
8:47:18the third step, what we do, we actually
8:47:20combined it to get the new cell state.
8:47:22Now, in the fourth step, what we did, we
8:47:24finally got the output that we want. And
8:47:27how we did that? Just by passing HT - 1
8:47:29and HT through a sigmoid function,
8:47:31multiplying it with the tan H CT, the
8:47:33tan H new cell state, and we get the new
8:47:36output. Fine, guys? So, this is what
8:47:38basically LSTM is, guys. Now, we'll look
8:47:40at a use case where we'll be using LSTM
8:47:43to predict the next word in a sentence.
8:47:45All right? Let me show you how we are
8:47:46going to do that.
8:47:48So, this is what we are trying to do in
8:47:49our use case, guys. We'll feed LSTM with
8:47:52correct sequences from the text of three
8:47:54symbols. For example, had a general and
8:47:57a label that is counsel in this
8:47:59particular example. Eventually, our
8:48:01network will learn to predict the next
8:48:03symbol correctly. So, obviously, we need
8:48:05to train it on something. Let us see
8:48:06what we are going to train it on.
8:48:08So, we'll be training LSTM to predict
8:48:10the next word using a sample short story
8:48:12that you can see over here.
8:48:14All right? So, it has basically 112
8:48:16unique symbols. So, even comma and full
8:48:18stop are considered as symbols. All
8:48:20right? So, this is what we are going to
8:48:22train it on.
8:48:23So, technically, we know that LSTMs can
8:48:25only understand real numbers. All right?
8:48:27So, what we need to do is we need to
8:48:29convert these unique symbols into a
8:48:31unique integer value based on the
8:48:33frequency of occurrence. And like that,
8:48:35we'll create a dictionary. For example,
8:48:37we have had here that will have value
8:48:3920. A will have value six. General will
8:48:42have value 33. All right? And then, what
8:48:45happens, our LSTM will create a 112
8:48:48element vector that will contain the
8:48:50probability of each of these words or
8:48:53each of these unique integer values. All
8:48:55right? So, since 0.6 has the highest
8:48:57probability in this particular vector,
8:48:59it'll pick the index value of 0.6. Then,
8:49:02it will see it what symbol is attached
8:49:04to that particular integer value. So, 37
8:49:06is attached to counsel. So, this will be
8:49:08our prediction, which is absolutely
8:49:09correct as the label is also counsel
8:49:11according to our training data. All
8:49:13right. So, this is what we are going to
8:49:15do in our use case. So, guys, this is
8:49:17what we'll be doing in our today's use
8:49:18case. Now, I'll quickly open my PyCharm
8:49:21and I'll show you how you can implement
8:49:22it using Python. We'll be using
8:49:24TensorFlow, which is a popular Python
8:49:26library for implementing deep neural
8:49:28networks or neural networks in general.
8:49:30All right. So, I'll quickly open my
8:49:31PyCharm now. So, guys, this is my
8:49:33PyCharm and I over here I've already
8:49:35written the code in order to execute the
8:49:36use case that we have. So, first we need
8:49:38to do is import the libraries, NumPy for
8:49:41arrays, TensorFlow we know,
8:49:42tensorflow.contrib from that we need to
8:49:44import RNN and random collections and
8:49:46time. All right. So, this particular
8:49:48block of code is used to evaluate the
8:49:51time taken for the training. After that,
8:49:53we have log_path and this log_path is
8:49:56basically telling us the path where the
8:49:58graph will be stored. All right. So,
8:49:59there will be a graph that will be
8:50:00created and then that graph will be
8:50:02launched. Then only our RNN model will
8:50:04be executed. Then that's how TensorFlow
8:50:06works, guys.
8:50:07So, that graph will be created in this
8:50:09particular path. All right. And we are
8:50:11using summary writer. So, that will
8:50:13actually create the log file that will
8:50:14be used in order to display the graph
8:50:17using TensorBoard. All right. So, then
8:50:19we have defined training_file, which
8:50:21will have our story on which we'll train
8:50:23our model on. Then what we need to do is
8:50:25read this file. So, how are we going to
8:50:27do that? First is read line by line
8:50:29whatever content that we have in our
8:50:31file. Then we are going to strip it.
8:50:33That means we are going to remove the
8:50:35first and the last white space. Then
8:50:37again, we are splitting it just to
8:50:40remove all the white spaces that are
8:50:41there. After that, we're creating an
8:50:43array and then we're reshaping it. Now,
8:50:45in during the reshape, if you notice
8:50:46this minus one value tells us the
8:50:48compatibility. All right. So, when
8:50:50you're reshaping it, you need to make
8:50:51sure that
8:50:53you know, we are providing in the
8:50:54correct parameters to reshape it. So,
8:50:56you can convert a three cross two matrix
8:50:58to a two cross three matrix, like right?
8:51:01So, just to make sure that that it is
8:51:02compatible enough, we add this minus one
8:51:04and it'll be done automatically. All
8:51:06right? Then, return content. After that,
8:51:09what we are doing, we are feeding in the
8:51:11training data that we have, training
8:51:12{underscore} file. We are feeding in our
8:51:14story and calling the function read
8:51:16{underscore} data. Then, what we are
8:51:17doing, we are creating a dictionary.
8:51:19What is a dictionary? We all know, key
8:51:20value pairs based on the frequency of
8:51:22occurrences of each symbol. All right?
8:51:24So, from here, collections.counter
8:51:26words.most_common. So, most common words
8:51:29with their frequency of occurrence,
8:51:30there'll be a dictionary created. And
8:51:32after that, uh we'll call this dict
8:51:34function and this dict function will
8:51:36feed in word and which is equal to
8:51:38length of dictionary. That means
8:51:40whatever the length of that particular
8:51:42dictionary, how many time it is
8:51:43repeated. So, we'll have the frequency
8:51:45as well as a symbol. That'll be our key
8:51:47value pair and we're reversing it as
8:51:49well.
8:51:50Then, what we are doing, we are calling
8:51:51it build {underscore} data set and we're
8:51:54feeding in our training data there. This
8:51:56is our vocabulary size, which is nothing
8:51:57but the length of your dictionary. Then,
8:51:59we have defined various parameters such
8:52:01as learning rate, uh iterations or
LSTM Explained
8:52:03epochs. Then, we have display step and
8:52:05{underscore} input. Now, learning rate,
8:52:07we all know what it is, uh the steps in
8:52:09which our variables are updated.
8:52:11Training {underscore} iterations is
8:52:12nothing but your epochs, the total
8:52:14number of iterations. So, we have given
8:52:1550,000 iterations here. Then, we have
8:52:17display {underscore} step, that is
8:52:191,000, which is basically your batch
8:52:20size. So, batch size is what? After
8:52:23every 1,000 epochs, you'll see the
8:52:24output. All right? So, it'll be
8:52:25processing it in batches of 1,000
8:52:27iterations. Then, we have n {underscore}
8:52:29input as three. Now, the number of units
8:52:31in the RNN cell, we'll keep it as 512.
8:52:34Then, we need to define X and Y. So, X
8:52:37will be our placeholder that will have
8:52:38the input values and Y will have all the
8:52:41labels. All right? vocab size.
8:52:44So, X is a placeholder where we'll be
8:52:46feeding in our input dictionary.
8:52:47Similarly, Y is also one more
8:52:49placeholder and it'll have a shape of
8:52:51none {comma} vocab size. Vocab size we
8:52:53have defined earlier.
8:52:55As you can see, which is nothing but the
8:52:56length of your dictionary. Then we're
8:52:57defining weights as well as biases.
8:53:00After that, we have defined our model.
8:53:02All right. So, this is how we are going
8:53:03to define it. We'll
8:53:05create a function RNN when we'll have X
8:53:08weights and biases. And after that, we
8:53:10are calling in RNN.multi_rnn_cell
8:53:12function. And this is basically to
8:53:14create a two-layer LSTM. And each layer
8:53:17has n_hidden_units.
8:53:19After that, what we are doing, we are
8:53:20generating the predictions. But once we
8:53:22have generated the prediction, there are
8:53:24n_input_outputs,
8:53:25but we only want the last output. For
8:53:28that, we have written this particular
8:53:29line. And then finally, we are making a
8:53:30prediction. We are calling this RNN and
8:53:32function feeding in X weights and
8:53:34biases. After that, we are calculating
8:53:36the loss as and then we are optimizing
8:53:38it. For calculating the loss, we are
8:53:41using reduce_mean softmax_cross_entropy.
8:53:44And this will give us basically the
8:53:46probability of each symbol. And then we
8:53:48are optimizing it using RMS uh prop
8:53:50optimizer. All right. And this gives
8:53:52actually a better accuracy than Adam
8:53:54optimizer. And that's the reason why we
8:53:56are using it. Then we are going to
8:53:57calculate the accuracy. And after that,
8:54:00we are going to initialize the variables
8:54:01that we have used. As we have seen in
8:54:03TensorFlow, that we need to initialize
8:54:04all the variables, unlike constants and
8:54:06placeholders in TensorFlow. All right.
8:54:08And once we are done with that, we are
8:54:10feeding in our values, then calculating
8:54:12the accuracy, how accurate it is. And
8:54:14then when optimization is done, we are
8:54:16calculating the elapsed time as well.
8:54:18So, that will give us how much time it
8:54:20took in order to train our model. Then
8:54:22this is just to run the TensorBoard on
8:54:24our local host 6006. And yeah, and this
8:54:28particular block of code is is used in
8:54:30order to handle the exceptions. So,
8:54:32exceptions can be like whatever word
8:54:34that we are putting in might not be
8:54:36there in our dictionary or might not be
8:54:37there in our training data. So, those
8:54:39exceptions will be handled here. And if
8:54:41it is not there in our dictionary, then
8:54:42it will print word not in our
8:54:44dictionary. All right. So, fine guys.
8:54:46Let's
8:54:47input some values and we'll have some
8:54:49fun with this model. All right? So, the
8:54:51first thing that I'm going to feed in is
8:54:53had general. So, whenever I feed in
8:54:56these three values, had a general,
8:54:58there'll be a story that will be
8:54:59generated by feeding back the predicted
8:55:01output as the next symbol in the inputs.
8:55:04All right? So, when I feed in had a
8:55:05general, so it'll predict the correct
8:55:07output as counsel. And this counsel will
8:55:10be fed back as a part of the new input
8:55:12and our new input will be a general
8:55:14counsel. So, it'll be a general counsel.
8:55:16All right? So, these three words will
8:55:18become our new input to predict the new
8:55:19output, which is two. All right? And so
8:55:21on. So, surprisingly, LSTM actually
8:55:24creates a story that, you know, somehow
8:55:26makes sense. So, let's just read it. Had
8:55:28a general counsel to consider what
8:55:30measures they could take to outwit their
8:55:32common enemy, the cat. By this means, we
8:55:35should always know when she was about
8:55:37and could easily. All right? So, somehow
8:55:39it actually makes sense when you feed in
8:55:40that. So, what'll happen when you feed
8:55:42in these three inputs, it'll predict the
8:55:44next word, that is counsel. After that,
8:55:46it'll take counsel and it'll feed back
8:55:48as an input along with a general. So, a
8:55:50general counsel will be your next input
8:55:53to predict two. Similarly, in the next
8:55:55iteration, it'll take general counsel
8:55:57two and predict counsel for us. And this
8:55:59will keep on repeating.
8:56:02>> [music]
8:56:09>> processing and why do we even need it?
8:56:12You see, natural language
8:56:13analysis in both audible data as well as
8:56:16the text document. NLP system can
8:56:19capture meaning from an input such as
8:56:21sentences, paragraphs, pages, and give
8:56:23out a desired output based on our
8:56:25application. So, why do we need NLP? You
8:56:28see, natural language processing helps
8:56:29computer communicate with humans in
8:56:31their own language and scale other
8:56:33language-related task. For example, NLP
8:56:36makes it possible for computers to read
8:56:38text, hear speeches, interpret it,
8:56:41measure the sentiment, and then
8:56:42determine which part of it are
8:56:44important. Today's machines can analyze
8:56:46more language-based data than humans.
8:56:48That too with consistency, accuracy, and
8:56:51in an unbiased manner. More or less, we
8:56:53all know that there is a staggering
8:56:55amount of unstructured data that is
8:56:57generated every day. Be it from a
8:56:59medical record or to a social media.
8:57:01Automating NLP task will critically be
8:57:04helpful in future for analyzing text and
8:57:06speech data efficiently. So, moving
8:57:09ahead, let us now see the ways we can
8:57:10process our textual data. We can process
8:57:13our textual data in one of two ways. One
8:57:15is a machine learning way, and other one
8:57:17is a deep learning method. In machine
8:57:19learning, we can make use of algorithms
8:57:21such as bag of words, TF-IDF to classify
8:57:24and predict the desired output. But, the
8:57:26drawback of this is that these machine
8:57:28learning algorithms do not consider the
8:57:30context of the word or a sequence. Here,
8:57:32the way it works is on the base on
8:57:34number of times a word is repeating, a
8:57:36probability is derived out of it, and
8:57:38then performs a classification task.
8:57:41This is the reason why we have deep
8:57:42learning model for NLP tasks. Speaking
8:57:45about deep learning model, we have
8:57:46something like recurrent neural network,
8:57:48LSTM, transformer network, Google's BERT
8:57:51algorithm, and many more. This deep
8:57:53learning model learns the pattern of a
8:57:55word or a sequence, and then tries to
8:57:57predict the desired outcome of the task.
8:57:59We can perform NLP using deep learning
8:58:01in one of two ways. One by pre-trained
8:58:03models such as Google's Word2vec or
8:58:06global vector models. And the other way
8:58:08is to train our own model. If you're
8:58:10trying to train our own model, it would
8:58:12require a very huge amount of data and
8:58:14also a compute power to support it. In
8:58:16most of the cases, we'll be using
8:58:18pre-trained models. Moving ahead, let us
8:58:20now discuss recurrent neural networks.
8:58:23As I mentioned earlier, the bag of word
8:58:25or TF-IDF model for processing our text
8:58:28is very inefficient. As I mentioned
8:58:30earlier, the bag of word or TF-IDF model
8:58:32that was used in machine learning to
8:58:34process our textual data is very
8:58:36inefficient as it takes one word at a
8:58:38time and also the context of the word in
8:58:41which it is being spoken about is
8:58:42totally ignored. Although this would
8:58:44give us some prediction, but we can
8:58:45expect lot of loss. This is why we use
8:58:48RNN model or recurrent neural network.
8:58:50Here it requires a sequential data.
8:58:53Now you might be wondering what does
8:58:54this sequential data mean, right? In
8:58:56simple words, sequential data is
8:58:58dependent on the past value. What I'm
8:59:00trying to say here is that we can read
8:59:02our document, right? We can read our
8:59:04document or English documents only from
8:59:06left to right side. Or take example of
8:59:08stock price prediction. We cannot
8:59:10randomly place the dates, right? If you
8:59:12have to make a prediction for next 2-3
8:59:14months, we'll obviously refer to the
8:59:16data that was previously recorded. So
8:59:18this is what a sequential data means. So
8:59:21what makes RNN capable of handling
8:59:23sequential data? Well, you see RNN model
8:59:26makes use of something called a state.
8:59:28This is nothing but a temporary memory
8:59:30that stores the previous data. As you
8:59:32can see here in an image, this is the
8:59:34general architecture of a recurrent
8:59:36neural network. So let me now move to my
8:59:38canvas and show you how recurrent neural
8:59:40network works and what are its internal
8:59:42workings.
8:59:44All right. So as we have seen in an
8:59:45image, right? Back in our slide, you saw
8:59:47that, you know, recurrent neural network
8:59:49have something called as states and then
8:59:51we also had some boxes, right? So what
8:59:54does this boxes represents? So let me
8:59:56quickly draw over here and show you what
8:59:57does this box represents. So if you
9:00:00remember, right? It goes something like
9:00:02we have a box here. Okay? And then we
9:00:04also had another box.
9:00:07And then let's consider like two more
9:00:09boxes that would be sufficient. Okay?
9:00:11And then the way we provide input for
9:00:13our recurrent neural network is over
9:00:15here.
9:00:16Okay? Now based on our application, we
9:00:18can demand our recurrent neural network
9:00:20to provide output in either one of three
9:00:22ways. It can either be like we can have
9:00:24multiple inputs, as you can see here, or
9:00:26we can have single input and multiple
9:00:28outputs, or the other way around is we
9:00:30can have single input and single output.
9:00:32So here what we'll consider is we are
9:00:34using multiple inputs and also we have
9:00:36multiple outputs.
9:00:38Okay. So now let us see as this is a
9:00:40supervised learning model, right?
9:00:42Obviously it will have some kind of
9:00:43input and then it will also have labels
9:00:45to clarify the data. So let's take
9:00:48something like if this is our X data,
9:00:49right? So if this is the X data, so
9:00:52let's take this as a list and then over
9:00:53here we'll have data which would
9:00:55represent something like X1, then we'll
9:00:57have X2,
9:00:59then we'll have X3,
9:01:00and then we'll have something like Xn.
9:01:04Okay? So what does this X over here
9:01:07represents? See, here let's take an
9:01:09example that Uri, which was a movie, is
9:01:12a good movie. We can have couple more
9:01:13X's and we'll just put it here as good
9:01:16movie.
9:01:27Okay. So this is nothing but the inputs.
9:01:30Each of these words. Now what our model
9:01:32over here does, let's take an example
9:01:34that we are trying to find name entity
9:01:36prediction, which stands for NER, right?
9:01:38So what does name entity prediction does
9:01:40is, you know, trains the model in such a
9:01:42manner that it finds a pattern. So as
9:01:44you can see here, we this is Uri, right?
9:01:47Uri, it's it's something like a name,
9:01:48right? So this this goes as one.
9:01:50And is is not a name. So this would be
9:01:52zero. Good is this is also not a name
9:01:55and movie is also not a name. So the
9:01:57output over here would be something like
9:01:581000. All right? So now let's see what
9:02:01are the inputs and how they would look
9:02:03like. So first off we have inputs. So if
9:02:06this is our X data, so we'll have input
9:02:08something like Xi. This should be in
9:02:10lower case.
9:02:11Okay? So this would be X1,
9:02:13then we'll have X2,
9:02:15we'll have X3, and then we'll have X4.
9:02:19Similarly, the Y the output over here
9:02:21would be something like we'll give it Y
9:02:23hat of 1, [snorts]
9:02:25Y hat of 2,
9:02:27Y hat off three, at the same time we'll
9:02:29also have why hat off four. Now if
9:02:31you're wondering what does this why hat
9:02:33represents, why hat over here is nothing
9:02:34but
9:02:35predicted values.
9:02:39And why is nothing but you know the
9:02:41labeled value or trained values.
9:02:45And now as we all know that the main
9:02:48thing or the main feature behind
9:02:50recurrent neural network is nothing but
9:02:52the states. And the way the states are
9:02:54represented is by using the state vector
9:02:56or it can also be called as context
9:02:57vector. So the state vector over here
9:02:59starts with A. Let me give it as a
9:03:01different color here. Let me give it as
9:03:02blue. So here we'll have A and this
9:03:05should be zero. Okay? And now this A
9:03:08would be passed down to this. Okay, so
9:03:10this would be A of one and over here
9:03:12we'll have A of two,
9:03:14A of three and then finally A of four.
9:03:18I'm pretty sure you might be wondering
9:03:19what does this A contains, right? So
9:03:21over here as I have mentioned, let me
9:03:23take this very example. So here Uri is
9:03:26good movie, okay? And so what does A1
9:03:29will contain? A1 will contain Uri. See,
9:03:33if we take the normal algorithm, what's
9:03:34going to happen is it will just consider
9:03:36only this one particular block. Okay? If
9:03:39this was just a normal algorithm or
9:03:41something which is used in olden days,
9:03:42it will just consider this particular
9:03:44block and it won't be considering the
9:03:45previous values. So this previous value
9:03:48is being stored over here in the form of
9:03:50a memory. So now A1 will contain Uri and
9:03:53what A2 will contain over here? So A2
9:03:57will be having something like Uri is.
9:04:00And similarly as it goes through here,
9:04:02it will collect each and every word and
9:04:04finally A4 will be complete context
9:04:06vector. So it will have the entire
9:04:07sentence Uri is a good movie.
9:04:12Okay, I hope you understood what does
9:04:14this A signifies over here. Let me
9:04:16quickly erase all of these. Okay, so the
9:04:19another important part which goes over
9:04:20here in any machine learning or deep
9:04:22learning model is nothing but the
9:04:24weights.
9:04:25So, let's consider that we have weights
9:04:27which is nothing but W or let's take it
9:04:29as U. Let me give another color here, U,
9:04:32V, and W. In this RNN model, right? What
9:04:35we're going to do is we're going to have
9:04:37a single or we'll have just only one
9:04:39weights, right? So, the U weight is same
9:04:42for all of these inputs here. V weight
9:04:45over here is same for all the outputs.
9:04:47The weight W is same for all our state
9:04:50metrics. Okay? So, this is how the
9:04:52weights are determined.
9:04:54So, now to get a better understanding of
9:04:56what's happening over here and to derive
9:04:58a mathematical equation for feedforward
9:05:00network, let's take a single block over
9:05:02here. And let's see how this internal
9:05:04working of this is working, okay? So,
9:05:06over here we'll have a box.
9:05:09Okay? And this will be our input. So,
9:05:11let's give our input as X of T because
9:05:14this is a generalized model, right? And
9:05:16now what you're going to do over here is
9:05:17this would be A
9:05:19of T minus 1. That's because A of T is
9:05:22will be present over here. And now if
9:05:24you have an output which is present over
9:05:26here and this would be nothing but Y hat
9:05:29of T.
9:05:30All right? So, as I've mentioned earlier
9:05:31that we will be having weights. So, the
9:05:33weights over here is nothing but U, V,
9:05:35and W. So, what's happening over here?
9:05:38Okay. So, first off, there will be a
9:05:41matrix multiplication between these two
9:05:42values.
9:05:44Okay? And then there will be a matrix
9:05:45multiplication between these two values.
9:05:47And once they are done, we'll add these
9:05:49two values and then give a activation
9:05:51function to it. Okay? And the activation
9:05:54function that will be used over here is
9:05:55tan H. So, if I have to put it in an
9:05:58equation form over here, so the product
9:06:00of these two, X of T
9:06:02times or it would be a dot product
9:06:05of U, okay? And the sum of these two, so
9:06:09it will be T minus 1 times W. And now
9:06:12what we're going to do is we're going to
9:06:13pass an activation function which is
9:06:14nothing but tan H over here.
9:06:16So, let's just give a small dotted
9:06:18notation over here.
9:06:20So, whatever the output which comes
9:06:22it'll be performing an addition of these
9:06:25two and also give an activation function
9:06:27which will represent here by f.
9:06:29All right? So, this is how we get an
9:06:31activation function over here. So, what
9:06:33does this value signifies this? This is
9:06:35nothing but this output over here, a of
9:06:37t. So, I hope you understand what is a
9:06:39of t. A of t is this value. Okay, let me
9:06:42quickly highlight that for you. So, a of
9:06:44t is this value. Okay? So, a of t is
9:06:47represented over here. And the way we
9:06:49get this is by this particular equation.
9:06:52Now, what about the value for y hat of t
9:06:54or the prediction of y of t? So, for y
9:06:57of t, it's going to be something like
9:06:58this. So, y hat of t, this would be
9:07:01nothing but we have to multiply this
9:07:03particular a of t with this matrix over
9:07:06here. Okay? Or the weights I can say.
9:07:08So, it's going to be nothing but a of t
9:07:11times the weight matrix that is v. And
9:07:14then we have to pass activation
9:07:16function. Usually, the activation
9:07:17function that is going to be used over
9:07:19here is softmax or sigmoid. It totally
9:07:21depends upon the what kind of output
9:07:24you're expecting. All right? And this is
9:07:26how it works. And now, an important
9:07:28thing that I would like to mention over
9:07:30here is that we also have to add
9:07:32something called as bias. So, it would
9:07:34be b over here. I'll just represented
9:07:36this by a different color.
9:07:38So, this is an equation for our
9:07:40recurrent neural network in a
9:07:41mathematical form.
9:07:43So, now what's going to happen is in
9:07:44order to find the loss, right? We have
9:07:46to subtract whatever value we have
9:07:47predicted with the given value. So, the
9:07:50predicted value over here, let's say
9:07:51this is as capital y. Okay? And we also
9:07:54have been given the train value. So,
9:07:55this would be looking something like
9:07:57this.
9:07:58So, we'll obviously have the train
9:07:59values.
9:08:01Let's just give a random output, but
9:08:03there'll be four. So, it'll be 0 but the
9:08:05train values. And now we'll have the
9:08:07other, this is nothing but the given
9:08:08label data. So, this would be correct
9:08:10values. So, 1 0 0 1. As be correct
9:08:12values. So, 1 0 0 1. As this is a
9:08:15supervised learning, so we obviously
9:08:17will be given this label data over here.
9:08:19So, now what this will do is this will
9:08:20subtract each of these values and
9:08:23calculate a loss. So, how do you
9:08:25calculate a loss, right?
9:08:27Okay. So, as you can see here, if I give
9:08:30this L, like let me change the color
9:08:32here. So, if I say L is loss, so this
9:08:35would be nothing but Y hat of first
9:08:37value minus the actual value. Okay, so
9:08:40this is nothing but the predicted value
9:08:42and this is nothing but the given value.
9:08:44So, if I try to find out the loss of
9:08:46this, then what this would look like is
9:08:48this is just for the one value, right?
9:08:49So, if I have to do it for all the
9:08:50values, then it would be the summation.
9:08:52So, it would be nothing but loss is
9:08:54equal to summation I which ranges from 1
9:08:57to n, right? And then we'll have loss
9:09:01and then theta I. Okay, so this LI over
9:09:04here represent these values. Okay, so
9:09:06this is just for one. So, if you want me
9:09:07to explain you this in detail, so over
9:09:09here we have Y1, Y2, Y3 and Y4. How this
9:09:12would look like is something So, this
9:09:14would be for one plus Y hat of two minus
9:09:18actual value of Y of two, then summation
9:09:21predicted value of Y3 minus the actual
9:09:24value of Y3 and then it'll be predicted
9:09:28value of Y4 minus the actual value of
9:09:30Y4. And when I perform addition over
9:09:33here or when I perform summation, I can
9:09:35generalize this equation into this form.
9:09:37But we're not yet done over here. As you
9:09:39can see, we have something called as
9:09:41theta values.
9:09:42So, what does this theta values
9:09:43represents?
9:09:44You see, the main agenda behind finding
9:09:47a loss is to increase our accuracy,
9:09:48right? So, if I say this is my gradient
9:09:51descent and this is the lowest global
9:09:54minima, right? So, my agenda over here
9:09:56is to reach this global minima. So, now
9:09:59in order for me to do this, in order to
9:10:01increase my accuracy, I have to
9:10:02obviously change my values. So, how do I
9:10:05do that? How do I increase an accuracy?
9:10:07It's obviously by these weights. Okay?
9:10:09It's by this W, U, and V. So, W, U, and
9:10:13V are nothing but the weights. So, what
9:10:15this theta over here represents, let me
9:10:17quickly erase this.
9:10:19Okay, so what this theta over here
9:10:20represents is nothing but the values of
9:10:22W, U, and V. So, now what we're going to
9:10:25do is we're going to have a partial
9:10:27derivative of the loss with respect to
9:10:31U, and then we'll have a partial
9:10:33derivative of loss with respect to V,
9:10:35and then we'll also have partial
9:10:37derivative of loss with respect to W.
9:10:40These are nothing but weights, and we
9:10:41are trying to train the weights.
9:10:44This is done in order to increase our
9:10:45accuracy.
9:10:46Okay? And the way these models get
9:10:48trained is with the help of back
9:10:50propagation. And the way the back
9:10:51propagation works over here is by
9:10:53partial derivative. So, what I'm trying
9:10:55to say here is as if I get my weights
9:10:57over here, so let me just take another
9:10:59color. So, this V we know that it is
9:11:01totally dependent upon this value over
9:11:03here. This won't be V, this would be W.
9:11:05Okay? So, this W value will be totally
9:11:07dependent on the previous one. And this
9:11:09value will be dependent upon this one,
9:11:10and this value will be dependent upon
9:11:12this one, and this one would be finally
9:11:13dependent on this. So, this is in order
9:11:16to move from here to here to here and
9:11:18then to here, we'll use something called
9:11:20as partial derivatives, right? So, this
9:11:23how we update our old weights.
9:11:25All right. Another important concept
9:11:27that make neural network or recurrent
9:11:29neural network very important is nothing
9:11:31but embedding layer. So, let me quickly
9:11:33draw a boundary over here.
9:11:35Okay. So, embedding layer.
9:11:40All right. So, what does this embedding
9:11:42layer signifies? Okay, so if I take a
9:11:44convention way, right? So, like let's
9:11:46say that, you know, our X or, you know,
9:11:49our input over here, this is nothing but
9:11:51X over here, right? So, the what is the
9:11:52dimension or the shape of this X? The
9:11:54shape of this X is nothing but 1 {comma}
9:11:574, right? 1 {comma} 4. Similarly, it's 1
9:11:59{comma} 4 here, 1 {comma} 4, and 1
9:12:02{comma} 4. But, this is not usually the
9:12:04case when you're working with the real
9:12:06world examples. When you're working with
9:12:08real world examples, you won't be having
9:12:09four different words, right? We'll
9:12:11obviously have lots of values. So, for
9:12:13example, let's say we have X and the
9:12:16shape of this X is something like we
9:12:18have 10,000 values. Okay? So, now if I
9:12:21try to feed this 10,000 values into my
9:12:25into my network over here, obviously I
9:12:26would be using batch propagation. So, it
9:12:29would take a lot of time, right? Because
9:12:31it's 10,000 values after all. So, in
9:12:33order to overcome this, what we're going
9:12:34to do is we're going to reduce the size
9:12:36of this. Okay? So, basically what word
9:12:38embedding layer does is think that we
9:12:40have a matrix. This is our input layer,
9:12:43right? Our input word. So, that we call
9:12:45this as sparse matrix.
9:12:47This is nothing but, you know, an
9:12:49individual value. That's XI or X1, X2,
9:12:51X3. Let's for generalization we'll give
9:12:53you here as XI. And the shape of this is
9:12:55nothing but 1,V. V here represents the
9:12:58end number of dimensions. And one is
9:13:00because it's just going to be one word,
9:13:02right? So, it's going to be one. So, if
9:13:03I consider with respect to this, this is
9:13:05one and this is V. Okay?
9:13:07So, now what will happen is if V is is
9:13:1010,000 or pretty great,
9:13:12it obviously we won't have much
9:13:14we're going to lose a lot of time on
9:13:15computation and also take huge compute
9:13:18power. So, what we'll do is I'll
9:13:19multiply this with an embedding layer.
9:13:24And what this embedding matrix does is
9:13:26it is basically a set of features. So,
9:13:28this is something you know, this is a
9:13:30black box model. And we'll all this
9:13:32would do is this would attract couple of
9:13:34features. If you want me to give you a
9:13:36better analogy of what this is, and if
9:13:38you want me to compare this with respect
9:13:40to a CNN, which is nothing but another
9:13:41great algorithm for image processing
9:13:43using deep learning, right? Over there
9:13:45we are going to use something called as
9:13:46filters or kernels. So, each of those
9:13:48filters or kernels is responsible for
9:13:50extracting one specific feature, right?
9:13:52So, this is what embedding layer does.
9:13:54And what this would do, for example, now
9:13:57let's say that the size of embedding
9:13:58layer is V {comma} K. K is something
9:14:00that we provide an input over here. So,
9:14:02now what this would happen is this would
9:14:05give us a new matrix or embedded matrix
9:14:08whose size would be 1 {comma} K. I'm
9:14:10pretty sure you didn't understand this
9:14:12because over here I'm using, you know,
9:14:13these these letters. So, in order to
9:14:15make you better understand this, what
9:14:16I'm going to do is let me take a matrix
9:14:18over here. Okay? Let me take something
9:14:20like, you know, because this would be in
9:14:22an embedded form, right? So, this would
9:14:23be a sparse matrix. So, it would be like
9:14:250 0 0 0 1 then we'll have 0 0 and so on.
9:14:30Okay? So, now what this would do is
9:14:33we'll also create a matrix over here.
9:14:35The size of this would be something
9:14:36similar to that of
9:14:38V.
9:14:39This is nothing but 1 {comma} V and over
9:14:42here it would be K. So, the matrix shape
9:14:44over here would be V {comma} K, right?
9:14:47To give you a better analogy, let me
9:14:48also draw a couple of boxes here.
9:14:51So, this is the matrix and this is the
9:14:52matrix over here again. So, now what
9:14:54will happen over here is when I try to
9:14:56perform this matrix multiplication,
9:14:58right? What this would do, you know, as
9:15:00everything is zero and only one value is
9:15:02true, so let's say this is the one
9:15:04value, right? And this would be going
9:15:06across, you know, from left to right and
9:15:08this would be from top to bottom. Only
9:15:10one part over here would be marked and
9:15:12rest everything would be zero.
9:15:13Therefore, reducing the dimension. Let
9:15:16me give some random values like 0.5,
9:15:181.8, 0.5, just some random values. So,
9:15:22now this would obviously reduce the size
9:15:24of 1 {comma} K. So, what I'm trying to
9:15:26say here is now, for example, say that I
9:15:29have a size over here as 1 {comma}
9:15:3110,000. Okay, which is a very huge
9:15:33matrix and the shape of this, let's say
9:15:35that it's 10,000 {comma} 200.
9:15:39When I perform this embedding, right? Or
9:15:40embedding, the shape of the new matrix
9:15:43would be nothing but 1 {comma} 200. If I
9:15:45compare this part over here to the
9:15:47embedding whatever we have received over
9:15:49here. So, let me just give a quick
9:15:51brief. So, this is our embedded layer.
9:15:53So, this is really 1,200. So, you will
9:15:56see that we have decreased the
9:15:57dimensions by a drastic amount. Okay, so
9:16:00this is 10,000 and this is only 200. And
9:16:02this would be very efficient when we are
9:16:04trying to feed this to our recurrent
9:16:06neural network. And let me quickly show
9:16:08you how this would go. So, first let me
9:16:10draw our architecture. Let's take this
9:16:14blocks like this.
9:16:17And then we'll have an output over here.
9:16:19Okay. So, this would be our inputs,
9:16:21right? So, let me give something like
9:16:23this.
9:16:24So, initially, we used to provide X
9:16:26values over here, right? Now, we won't
9:16:28be doing that. We won't be providing any
9:16:30X values directly. Instead of that, what
9:16:32I'm going to do is I'll have an
9:16:33embedding layer over here.
9:16:38And this will have the X values. So, let
9:16:41me give here as X of 1, so X of 2, X of
9:16:453, and then we'll have X of 4, and then
9:16:49similarly, let's take this model to be
9:16:51multiple input and single output. So,
9:16:53here we'll be have Y hat of T. And then
9:16:56we'll have weights, obviously. So, this
9:16:58would be U, V, and W. And this is
9:17:02nothing but our matrix over here. So,
9:17:04this would be A of 0, A of 1, A of 2, A
9:17:09of 3,
9:17:10and finally A of 4.
9:17:12Okay, so this is our context matrix. And
9:17:14obviously, we'll be performing an
9:17:15activation function here. So, I'll just
9:17:17give it a F. You can put F, you can put
9:17:19G, it's totally up to you. So, this is
9:17:21how our recurrent neural network would
9:17:22actually work. All right?
9:17:25So, now that we know how RNN works, let
9:17:27us now understand what is LSTM. Or we
9:17:30can also say it as long short-term
9:17:31memory. You see, traditional RNNs are
9:17:34not good at capturing long-range
9:17:36dependencies. What I mean to say here is
9:17:38that when we tend to work with a very
9:17:39huge data set and multiple RNN layer, we
9:17:42are at the risk of vanishing gradient
9:17:44problem. Now, you might be wondering
9:17:46what is this vanishing gradient, right?
9:17:48Well, you see when training a very deep
9:17:50neural network, gradient or the
9:17:52derivatives decrease exponentially as it
9:17:54propagates down the layer. This is known
9:17:56as vanishing gradient problem. These
9:17:58gradients are actually used to update
9:18:00the weights of a neural network. But
9:18:02when the gradients vanish, these weights
9:18:04will not get updated. In the worst case
9:18:06scenario, it will completely stop the
9:18:08neural network from training. This
9:18:10vanishing gradient problem is a common
9:18:12issue in very deep neural networks. So
9:18:15to overcome this vanishing gradient
9:18:16problem in RNNs, long short-term memory
9:18:19was introduced. You see LSTM or long
9:18:22short memory is a modification to RNNs
9:18:24hidden layer. LSTM is capable of
9:18:26remembering RNNs weights and their
9:18:28inputs over a very long period of time.
9:18:31In LSTM, in addition to the hidden
9:18:32state, cell state is passed down to the
9:18:34next block. The way LSTM works is that
9:18:37it can capture long-range dependencies,
9:18:40that is old weights. It can have memory
9:18:42of previous inputs for a very extended
9:18:44time duration. The way LSTM cell does
9:18:46this is by using three main gates. First
9:18:49one is a forget gate. Forget gate
9:18:51removes the information that is no
9:18:52longer useful in the cell state. Then we
9:18:55have input gate. Additional information
9:18:57to the cell state is added by input
9:18:59gate. And finally, we have something
9:19:01called as output gate. Additional useful
9:19:03information to the cell state is also
9:19:05added by an output gate. This gating
9:19:07mechanism of LSTM has allowed network to
9:19:10learn the conditions for when to forget,
9:19:12ignore, or keep information in the
9:19:14memory cell.
9:19:15So let me now quickly move to my Jupiter
9:19:17notebook and show you how I can
9:19:19implement LSTM on name entity
9:19:21prediction. All right, so let me quickly
9:19:23move there. All right, so over here
9:19:25first off, I'll be opening my Google
9:19:28Colab.
9:19:32Okay, so let us give a name for our
9:19:34Google Colab over here.
9:19:36Let's give a short term, right? Name
9:19:38entity prediction. And let's connect our
9:19:40Google Colab to our server.
9:19:42Okay, meanwhile that's connecting. So,
9:19:44now you might be wondering from where am
9:19:46I going to use my data set? So, for me
9:19:48to use my data set, I'll just go for
9:19:49Kaggle, k a g g l e
9:19:52baby names. So, let me just quickly show
9:19:55you how this data set would look like.
9:19:57So, this is a CSV file over here. All
9:19:59right, so as you can see here, we have
9:20:01over 93,889
9:20:03unique values. Okay, so this is a very
9:20:06huge data set. And let's try downloading
9:20:09this. To download this is pretty simple.
9:20:11All you need to do is click this and it
9:20:13will get downloaded. As I've already
9:20:15downloaded this file, let me quickly
9:20:17upload this on my Jupyter notebook. So,
9:20:19let me go here and upload it from here.
9:20:23Okay, so let me go to this upload file.
9:20:26And yeah, so I have my CSV file here and
9:20:29let me open this. As this is a pretty
9:20:31huge data set, it will take some time.
9:20:32Meanwhile that's loading, let's see what
9:20:34we can do.
9:20:36So, first off let's import couple of
9:20:37libraries. So, we'll have import pandas
9:20:42as pd.
9:20:44And then we're going to import
9:20:46NumPy as np. And then we also need to
9:20:50have matplotlib. So, from sklearn
9:20:53All right, and we also need something
9:20:55like label encoder, but I'll show you a
9:20:57shortcut way to you know bypass label
9:20:59encoding. Okay, so let's try to load our
9:21:02cell here. And in order for us to read
9:21:04this data, so it's pretty simple. All
9:21:06we're going to do is let's give this as
9:21:08a data. This would be nothing but
9:21:10pandas.read_csv
9:21:12and then we're going to pass our file
9:21:15name. Let me change this to our root
9:21:16directory
9:21:18by putting a dot over here. Okay, so I
9:21:20won't be executing this as of now
9:21:21because it's trying to load our file.
9:21:26All right, so now that we have
9:21:27successfully loaded our data so, let's
9:21:30try running this cell over here. Okay,
9:21:32so let me close this and let me zoom in
9:21:35over here.
9:21:36So, now what we're going to do is let's
9:21:38see the shape of our data. So, let's see
9:21:40what's the data shape. data.shape
9:21:43and now let's see what it would be like.
9:21:45Okay, so as you can see here, we have
9:21:47five columns. But, the number of rows
9:21:50that we have is 1.8 million. That is
9:21:52approximately 18 lakhs, right? So, this
9:21:55is a pretty huge value. So, now what
9:21:57we're going to do is we'll just see how
9:21:59our data is looking like. So, we'll see
9:22:01data.head.
9:22:03And let's see what we need. So, as you
9:22:05can see here, we have ID, which is of no
9:22:08use for us. Then we have name. Okay,
9:22:10then this year, I don't think it's of
9:22:12any use for us. Then we have gender and
9:22:14count. Count here represents, you know,
9:22:17how many people have the name Mary, how
9:22:19many people have the name Anna, how many
9:22:21people have the name Emma, Elizabeth,
9:22:23and Minnie. This is over here, out of
9:22:25this if you see, right? There are a
9:22:27couple of things that we don't need. We
9:22:28can drop them out. You know, all we need
9:22:30is a name. Okay, and then we also need
9:22:32the gender. Because this is going to be
9:22:35our prediction. We're going to predict a
9:22:36we'll give our own custom name and then
9:22:38we'll see whether the name that is
9:22:40you're giving is male or a female. Okay?
9:22:44So, now what we're going to do is let's
9:22:45see how many unique values we have. So,
9:22:47let me quickly erase this first. Okay,
9:22:50so what I'm going to do is data.names.
9:22:53So, this should give us here name. And
9:22:56then we'll type here as unique.
9:22:58Okay, so this should give me unique
9:23:00values. Okay, so over here I have 93,889
9:23:04unique names. Okay, so now what we're
9:23:07going to do is we want to label encode
9:23:09this, right? So, we want our female, uh
9:23:11which is nothing but F, we want female
9:23:13to be zero and then male to be one or
9:23:15vice versa. So, in order to do that,
9:23:17either we can use label encoder or
9:23:20there's a shortcut method to this. Let
9:23:21me quickly show you how that works. So,
9:23:23first of all, we'll take our data frame,
9:23:25so it's data. And which column do you
9:23:27want to do this for? We want to do this
9:23:29for our gender column, right? So, let me
9:23:32pass this and give gender. And now what
9:23:35we're going to do is
9:23:37Okay?
9:23:38We'll take this as as type.
9:23:41Okay, this would be obviously in the
9:23:42form of category.
9:23:44And now what we'll do is this is cat
9:23:47dot codes.
9:23:50Okay, so this is nothing but panda
9:23:51shortcut, you know, to label encoding.
9:23:53Let's try to execute this and see what
9:23:55it would look like. So, as you can see
9:23:57here, we have couple of zeros and, you
9:23:59know, ones. This is nothing but it's
9:24:00representing females with one and males
9:24:03with zeros. Okay? So, now what we'll do
9:24:05is we have to update this column.
9:24:09So, we'll paste this. And this should be
9:24:11something like this over here. And let
9:24:13me execute this. Okay, so if you want to
9:24:16see how our data would look like now,
9:24:18let me just quickly run this once again.
9:24:20So, you'll see here now the values has
9:24:22been label encoded. Okay? So, now what
9:24:25we're going to do is we obviously need
9:24:27to take the unique names, right? And
9:24:30then we'll obviously group it by, right?
9:24:31So, what we'll do for this is we'll take
9:24:33something like data. We'll group this by
9:24:37the names. So, group by
9:24:39names.
9:24:40All right? And now what we'll do is
9:24:42we'll calculate the mean
9:24:44of the genders.
9:24:46We'll reset the index. The reason why we
9:24:47want to reset the index is because, you
9:24:49know, if you don't give the index then
9:24:50our name over here will become the
9:24:52index, right? So, we'll give reset
9:24:54{underscore} index.
9:24:56All right? So, let's give this to a new
9:24:59data frame and we'll call this as DF.
9:25:02Okay, let me execute this now.
9:25:04And let's see how this DF would look
9:25:05like. Okay, let me execute this right
9:25:08after this.
9:25:09So, as you can see here, it has grouped
9:25:11by by names, all everything in an
9:25:13ascending order. So, if this is all in
9:25:15an alphabetical manner. And yeah.
9:25:19And now only thing that I want to work
9:25:21on is this gender.
9:25:22Okay? The reason is because over here
9:25:24I'm getting a floating point value. I
9:25:26don't want this floating point value. I
9:25:28want to change this to integer value,
9:25:30right? So, what I'll do is
9:25:32DF gender
9:25:34This would be nothing but
9:25:36DF gender. Then I'll all I'm going to do
9:25:39is as type.
9:25:40I'll just put here as int. So, let's now
9:25:42see what this value would look like.
9:25:45Fantastic. We over here have now, you
9:25:47know, ones and zeros, which is nothing
9:25:48but an integer value. Okay? So, if you
9:25:51want to see this, so I either I can
9:25:53write DF or I can also put as head.
9:25:56Okay, so these are the first five
9:25:57values.
9:25:58Okay. So, now the way our neural network
9:26:01is going to work or the recurrent neural
9:26:02network is going to work is that, you
9:26:04know, I hope you remember these boxes,
9:26:06right? So, when I was talking or when I
9:26:09was explaining this RNN, I was saying
9:26:11that I would be passing around the
9:26:13words. But here in this project or in
9:26:16this program, we won't be passing words
9:26:18over here. You know, we won't be passing
9:26:20like Abba or Abida or Adam. We won't be
9:26:23passing these words. Instead of that,
9:26:25we'll be passing letters.
9:26:27So, over here it's going to be like
9:26:29alphabets. So, A, B. It can be any
9:26:32alphabet. It can be Z here. So,
9:26:34basically it depends upon whatever the
9:26:36value is coming here. So, in order to do
9:26:38that, we have to find number of unique
9:26:40alphabets. So, we know how many unique
9:26:41alphabets we have, right? So, we it's
9:26:4326. So, in order to get these alphabets,
9:26:46what we'll do is let me first quickly
9:26:47erase this.
9:26:49Erase all drawing.
9:26:50So, now we have 26 alphabets. We have to
9:26:53create our own vocabulary. So, what I'm
9:26:54going to do is I'm going to import
9:26:56string. So, now I need letters, right?
9:26:59So, l e t t e r s. This would be nothing
9:27:01but list of string
9:27:05.ascii.
9:27:06Okay? And if you want to see what this
9:27:07would give me, this would be nothing but
9:27:10the list of alphabets, which are in
9:27:11lower cases.
9:27:13Okay?
9:27:14And now what we'll do is we'll try to
9:27:15create a label encoding or we have to
9:27:18create a vocabulary, right? So, we'll
9:27:19have something like vocab. This would be
9:27:21nothing but I'll be using dictionary.
9:27:24And now what I want is zip. The way I
9:27:26want over here is, you know, for every
9:27:28individual values of this A B C D, I
9:27:31want to label encode this to 0 1 and
9:27:35whatever the value it is, right? So, it
9:27:36would be from 1 to 27.
9:27:38A unique numbers, right? So, this would
9:27:40be nothing but letters. And then uh
9:27:42we'll be need something like uh range
9:27:451 {comma} 27. So, this would give me the
9:27:48matrix from 1 to 26, right? And let's
9:27:50now see what this would look like. So,
9:27:52we have vocab.
9:27:54And let me execute this. This should
9:27:56give me a dictionary, okay? So, here
9:27:58we'll convert A to 1.
9:28:00Okay? And then B would be 2, C would be
9:28:033, and so on, Z would be 26.
9:28:06And now what we're going to do is uh
9:28:08we'll just try to create the reverse
9:28:10vocabulary. And the reason is we
9:28:12obviously won't be needing this, but uh
9:28:14you know, just in case you want to use
9:28:16it would be something very similar to
9:28:17this. Let me just copy the exact same
9:28:19thing.
9:28:20And paste it over here.
9:28:22So, we'll just do it as reverse, right?
9:28:24So, it will be R {underscore}
9:28:26R {underscore} And here, instead of
9:28:28numbers being second,
9:28:30we'll just cut this letters and we'll
9:28:32pass letters over here.
9:28:34And now you'll see if you're trying to
9:28:36decode whatever we have predicted, you
9:28:37know, we can just pass it down like
9:28:39this.
9:28:40Okay. So, now what what will happen is
9:28:42we need to do something like, you know,
9:28:44all our data, whatever is there, we have
9:28:45to convert them into a lowercase.
9:28:48So, and then once we convert them into a
9:28:50lowercase, we have to encode them into a
9:28:52numbers.
9:28:54So, whatever I'm saying is this A A B A
9:28:57N, right? A ban. So, we this A A
9:28:59obviously first of we have to convert
9:29:00all of these into a lowercase,
9:29:02you know, this value. And then whatever
9:29:04the equivalent value of A, the numerical
9:29:07value of A, so it's obviously going to
9:29:09be one. We'll substitute that with this.
9:29:11And it's going to be a list, right? So,
9:29:12how do I do that? So, for that I'll
9:29:14write a function.
9:29:16So, we'll have DEF word to number,
9:29:18right? Word to
9:29:21So, now what I'm going to do is I'm
9:29:22going to have for loop for I in range.
9:29:26So, this would be nothing but
9:29:29we have to go through the entire shape,
9:29:31right? So, d f dot shape.
9:29:33This should give me a list and I just
9:29:35need the first index.
9:29:37Okay? So, now what I'm going to do is
9:29:39I'll create one new list sequence.
9:29:42This would be nothing but for letters in
9:29:46d f. Obviously, we want the names part.
9:29:50And in this we're going to pass the
9:29:51index value. It's going to be I.
9:29:53Let me just give some space here just so
9:29:55that you better understand this.
9:29:57And now what I'm going to do is, you
9:29:59know, I'll have this vocabulary.
9:30:01vocab See, every time I pass a letter
9:30:03it'll convert it into, you know, this
9:30:05individual letter it'll convert it into
9:30:07a list all the equivalent, you know,
9:30:09numerical representation. It'll be
9:30:11letters and obviously it has to be in
9:30:13lower so it'll be lower.
9:30:15And then we'll just close this bracket
9:30:16here.
9:30:17So, now what we're going to do is before
9:30:19we execute this function, we'll have to
9:30:22append this so it'll be d f.
9:30:24And this is going to be names
9:30:27dot I.
9:30:28We'll replace the name in that index
9:30:30with this particular sequence.
9:30:32Okay? So, now all we need to do is run
9:30:34this function over here.
9:30:36And yeah.
9:30:38This will take some time. The reason is
9:30:39because we have almost around 18 lakh
9:30:42values. So, yeah, this should take some
9:30:44time. Meanwhile, let me just comment
9:30:46this.
9:30:53Okay? So, in the next stage what we're
9:30:54going to do is let's see how our this
9:30:57value over here would look like. So, let
9:31:00us now first execute this.
9:31:03All right. So, let us now see how our
9:31:04data frame will look like. So, let me
9:31:06execute this block now.
9:31:09So, as you can see here, our names have
9:31:11been completely changed or converted
9:31:13into list of numbers. But now, only
9:31:15issue that we are trying to have is the
9:31:18imbalance in the size of the list.
9:31:20Because when we are trying to have the
9:31:22number of boxes, right? We won't be
9:31:24having variable number of boxes. Okay?
9:31:27So, what we're going to do is either we
9:31:28set a value like something like take an
9:31:30average number like 10, 20, or you can
9:31:33take something like, you know, something
9:31:35like you take you either depend on
9:31:36maximum number or the minimum number of
9:31:38list. But the thing is, if you take the
9:31:41maximum number, then we have to pad a
9:31:42lot of zeros, and this would lead to a
9:31:44loss.
9:31:45So, if I reduce the size, this would
9:31:46also decrease the accuracy, right? So,
9:31:48what we're going to do is we'll plot
9:31:49this name and gender in the form of a
9:31:52histogram. So, let's take here X. This
9:31:55would be DF names.
9:31:58And we'll give here as dot values.
9:32:00And then same thing we'll do it for Y,
9:32:02DF gender.
9:32:04And this should be dot values.
9:32:06So, what we'll do is we also need a
9:32:09list, okay? So, now as we are going to
9:32:10plot this on a histogram, and what we're
9:32:12going to see in the histogram is just to
9:32:15analyze, you know, this this is a graph.
9:32:17We want to analyze, you know, where does
9:32:19the highest number of sequence, or if
9:32:22suppose this is a size, if this is size
9:32:23eight, and this is like 8,000 words or
9:32:268,000 names have the size eight, then
9:32:28you know, we can keep our average
9:32:30somewhere near, and then we can also
9:32:32decide, you know, if if the number after
9:32:3410, if not many names have a longer
9:32:36number or the longer length of that
9:32:38name, you know, so we can keep our
9:32:40average somewhere around nine or 10,
9:32:42okay? So, let's now quickly see how we
9:32:44can do that.
9:32:46To get the length of our names, so
9:32:47length X or name length.
9:32:51This would be like list comprehension
9:32:53for I in range 0, DF.shape
9:32:58of 0.
9:33:00And now what we're going to do is we
9:33:01need to find the length. So, this would
9:33:03be length X of I.
9:33:06Or here, you can either give BF or you
9:33:08can also give this X, right? So, it
9:33:10would be length of X.
9:33:11Okay, so let me quickly execute this.
9:33:14Okay, so here we're getting an error.
9:33:15Oh, yeah. It's not O, it's going to be
9:33:17zero, right? So, let me execute this
9:33:19now.
9:33:19Okay, so let me show you how this would
9:33:21look like. Name length, and let me print
9:33:24this off.
9:33:25So, as you can see, this is giving me
9:33:26list of names.
9:33:27So, there are huge amount of names. So,
9:33:29as you can see, first we had five, five,
9:33:31and then nine. So, let's now plot this
9:33:34and see how it would look like. Import
9:33:37Matplotlib as plt.
9:33:39All right, this is perfect. So, now what
9:33:40I'm going to do is I have to plot this,
9:33:42right? So, all I'm going to do is
9:33:44plt.hist.
9:33:45All right, and now I'm just going to
9:33:47give name length.
9:33:49And number of bins, this would be like
9:33:51let's give 20, okay? And then plt.show.
9:33:54Okay, so what do we find from this graph
9:33:57over here? You see, this is nothing but
9:33:58the length of the names. So, two, four,
9:34:01six, eight, 12, all these are length of
9:34:03the names. And this is nothing but zero,
9:34:055,000, 10,000, this is nothing but
9:34:07number of names that have a length four
9:34:09or number of names that have the length
9:34:11six. So, as you can see, right? The
9:34:13there are around almost 25,000 names
9:34:15whose length is six. And then as I cross
9:34:18like 10 or as I cross 12, not many names
9:34:21are there whose length is greater than,
9:34:23you know, 12. So, what I'm going to do
9:34:25here now is now we have to pad, right?
9:34:27Now, we have to pad number of zeros. So,
9:34:29in order to pad zeros, we have something
9:34:31called as built-in function from Keras.
9:34:33So, from Keras or you can also set as
9:34:35Keras.preprocessing
9:34:38.sequence import pad_sequence.
9:34:42So, let me quickly execute this now.
9:34:44And now what I'm going to do is I'm
9:34:45going to create a new list.
9:34:47So, let this be X. This is in lowercase.
9:34:50So, pad_sequence. And the things that it
9:34:52this is going to take is obviously the
9:34:54sequence. We have to give a list of
9:34:55sequence.
9:34:57Okay, and then we'll give something like
9:34:59df.names.
9:35:01And then this is going to be values.
9:35:04And now we want to define the max
9:35:06length. So, this is going to be 10. We
9:35:08also have an option of providing where
9:35:10do you want to do the padding? So, we
9:35:11can also do it as pre or post. We'll
9:35:14obviously be doing pre. So, let's see
9:35:17how do we do that. This would be nothing
9:35:19but, you know, if you can see this
9:35:20sequence over here. So, we have padding
9:35:22is equal to pre. It's so it's by
9:35:24default, right? So, pre. So, let me
9:35:26execute this now. And let's see how this
9:35:28X would look like.
9:35:30Okay. So, as you can see here, X is a
9:35:31matrix whose length is 10. Okay? So,
9:35:34each of these like this is this is
9:35:36nothing but, you know, 19 million cross
9:35:3810. So, there are 10 columns throughout
9:35:40all.
9:35:41So, now what we're going to do is we're
9:35:42going to create our own model. So, for
9:35:46that we'll do from keras.layers
9:35:49import
9:35:51input layer.
9:35:52And then we have to have embedding layer
9:35:54cuz if you don't have embedding layer,
9:35:56then you know, it it would be like each
9:35:57input would be something like 1 comma or
9:36:001.9 million. That is 18 lakhs. So, it's
9:36:02a pretty huge value to compute. So, we
9:36:04don't want that. So, that's why we'll
9:36:05use embedding layer. Then we have dense
9:36:06layer and then we have LSTM.
9:36:09We also, you know, rather than taking
9:36:11this as a sequential model, we'll take
9:36:12it as, you know, feed forward. So, what
9:36:15we'll do is from keras.models
9:36:19import model.
9:36:21So, now what we're going to do is we'll
9:36:23have to create our input layer. So, this
9:36:25would be input is equal to input.
9:36:28And now the shape that we're going to
9:36:29pass over here for for this shape
9:36:33So, how many columns do we have? We
9:36:34obviously have 10 columns, right? So,
9:36:36it's going to be 10.
9:36:37So, now what we're going to do is next
9:36:38we're going to have embedding layer. So,
9:36:40let's say this is EMB and this would be
9:36:43embedding.
9:36:44So, input over here
9:36:47or the input dimension over here is
9:36:48nothing but vocab size.
9:36:50We haven't defined this vocab size, so
9:36:52let's quickly do that. So, vocab
9:36:55size this would be nothing but length
9:36:59of vocabularies
9:37:01plus one. The reason why I'm doing plus
9:37:03one is because we also have zeros over
9:37:05here, right?
9:37:06And this would be like vocab size if you
9:37:08want to see.
9:37:09And let me execute this.
9:37:11So, we have 27, right? So, 26 are the
9:37:13number of alphabets and one is because
9:37:15we have number of zeros. So, we'll pass
9:37:17this as vocab size.
9:37:19Okay? And now we're also going to pass
9:37:21output dimension. So, output dimension,
9:37:23this is nothing but, you know, how many
9:37:24dimensions we want. So, now this is
9:37:26going to be five. All right? And the
9:37:28input for this embedded layer is going
9:37:29to be from INT.
9:37:31Okay? And now we are going to have our
9:37:33first LSTM layer. So, it's going to be
9:37:34LSTM
9:37:36one. So, this would be LSTM layer.
9:37:40Number of units we have to define here.
9:37:42So, units, this is going to be like 32.
9:37:45The units over here does not represent
9:37:46the number of boxes.
9:37:48The units over here represent the A
9:37:49values, right? So, now we have to do
9:37:52return sequence and this is going to be
9:37:54true.
9:37:55And the input for this is going to be
9:37:56from embedded layer.
9:37:57Then we have LSTM second layer.
9:38:00And this is be LSTM units we're going to
9:38:03pass. So, number of units that we're
9:38:05going to pass now is 64.
9:38:07And now the input for this is going to
9:38:09be LSTM one. Finally, we have an output
9:38:12layer.
9:38:13So, at the end, right? We're going to
9:38:14have a dense layer, right? So, dense.
9:38:17So, the number of units or number of
9:38:18neurons at the end we are going to have
9:38:19one.
9:38:20And kind of activation function that I'm
9:38:22going to have here is going to be
9:38:23sigmoid because we have to predict
9:38:25either it's a male or a female. Okay?
9:38:27So, sigmoid.
9:38:29And the input for this is going to be
9:38:31LSTM two.
9:38:33So, finally, we have to add this to our
9:38:35model. So, I'll do it as my model.
9:38:38This would be model.
9:38:40So, now I have to define the inputs. So,
9:38:42I N P U T S, this is going to be inputs
9:38:44I N P
9:38:46and outputs.
9:38:48This is going to be out.
9:38:49Okay? So, let me quickly execute this
9:38:51now.
9:38:52Okay, so we have this error.
9:38:54Please provide either a shape. Okay.
9:38:57Let's see what's Oh, yeah. Oh, here I've
9:38:59given it as pass, right? So, it's not
9:39:00going to be this pass. It's going to be
9:39:01shape.
9:39:02So, let me quickly execute this once
9:39:04again.
9:39:05Okay, as you can see, we have
9:39:06successfully executed this. And now
9:39:08let's see the model. summary.
9:39:10So, as you can see here, first off, we
9:39:12have input layer.
9:39:13Okay. So, we can have a number of
9:39:15values, but there will be only 10
9:39:17features. Okay, that's 10 columns.
9:39:19And then we're going to have once you go
9:39:21through this embedding layer, then you
9:39:23know, instead of having 10
9:39:25you know, instead of having that 1
9:39:27million or whatever it is, 1.8 million,
9:39:29we'll have just 135 parameters. Okay.
9:39:32Similarly, over here and finally at the
9:39:33dense, we have 65 parameters and then we
9:39:35have one.
9:39:36The reason why we have 65 here, just for
9:39:39if you don't know, is because 64 + 1
9:39:41bias. And this will give us 65.
9:39:43Okay. So, now finally, we're going to
9:39:45train our model. But before that, we
9:39:47have to compile it. It's going to be my
9:39:48model.
9:39:49model.compile
9:39:51Okay. So, we have to find an optimizer.
9:39:53So, optimizer, best one that I feel is
9:39:56Adam.
9:39:57Then we have to find the loss.
9:39:59So, as we're going to use just two
9:40:01predictions, right? It's either it's
9:40:02male or female, we'll use binary
9:40:04cross-entropy.
9:40:06And then finally, the matrix that we
9:40:07want to use here is
9:40:10This will be accuracy.
9:40:12So, let me execute this now.
9:40:14And finally, we are going to compile
9:40:15this. So, we have history.
9:40:17So, this will be model.fit.
9:40:20Okay. And now we're going to pass our
9:40:22values. We're going to give X. We're
9:40:23going to give Y.
9:40:25As you know, X is nothing but a matrix
9:40:26which has a padding. And Y is nothing
9:40:28but you know, the the classes. They're
9:40:29telling us either it's ones or zeros.
9:40:31And number of epochs
9:40:34is going to be 10.
9:40:35Batch size, as this is a pretty used
9:40:37data set, we are going to keep a pretty
9:40:38high batch size.
9:40:40So, I'll give a batch size here as 256.
9:40:43And then finally, we need validation
9:40:45split.
9:40:46Okay, it's not going to be my model,
9:40:47it's going to be my model, right? So, my
9:40:49underscore model. So, finally we're
9:40:52going to have validation split here.
9:40:54So, let's give it as 20%. So, it's going
9:40:56to be 0.2.
9:40:58All right. So, finally it's the moment
9:41:00of truth. Let us now execute our code.
9:41:03This will take some time to execute.
9:41:06So, let's see how this would look like.
9:41:11Okay, so if you can analyze this data
9:41:13over here,
9:41:15so as you can see, right? This
9:41:16validation accuracy has to increase.
9:41:19And we cannot see see every time, you
9:41:21know, if our model is over fitting, the
9:41:23accuracy over here will keep on
9:41:24increasing. All right. So, the more
9:41:27reliable source over here to see is
9:41:29nothing but validation accuracy. So, if
9:41:31validation accuracy is increasing, that
9:41:33means our model is neither over fitting
9:41:34or under fitting. And you can also see
9:41:36that we have our validation loss, which
9:41:38is kind of decreasing. And over here as
9:41:40well, we can see the validation loss
9:41:41over here. We can see the loss of our
9:41:43model is decreasing from 60 then 40 then
9:41:4639 39 and 80 38. So, let's now wait for
9:41:49a few more epochs. So, we have four more
9:41:52to go.
9:41:55Okay, so as you can see here, you know,
9:41:57our validation accuracy has been
9:41:59increasing. So, this is a very healthy
9:42:01growth.
9:42:02And even over here, our accuracy of our
9:42:04model is also increasing. Fine? And the
9:42:06loss is decreasing. It's decreased from
9:42:0860% to 37% and over here, our validation
9:42:11loss decreased from 42% to 36%.
9:42:14So, let us now map this like whatever
9:42:16values we have received. So, in order to
9:42:18map this, we have H.
9:42:20You know, this model over here retrieves
9:42:21us the history function, right? So,
9:42:23we'll give hist
9:42:24dot history.
9:42:25And let me execute this.
9:42:27Okay. So, now what we're going to do is
9:42:29this is nothing but key value pairs. So,
9:42:31if I put H,
9:42:32you know, if I give something like
9:42:33accuracy, okay, we let's plot this.
9:42:36So, model dot plot, right? So, plt dot
9:42:40plot.
9:42:41Okay.
9:42:42We'll have accuracy. We'll compare this
9:42:44with respect to accuracy and then
9:42:46plt.plot
9:42:49then we'll have validation accuracy. And
9:42:51we want to show this, right? So,
9:42:52plt.show.
9:42:53So, as you can see here, okay, so just
9:42:55to give a better analogy, let's let's
9:42:57execute one of these first.
9:42:59Okay, so the blue line over here
9:43:00represents the accuracy of our model and
9:43:02then this is nothing but the accuracy of
9:43:04our training data, right? Or testing
9:43:06data. So, as you can see, our model
9:43:07accuracy isn't decreasing. So, it's a
9:43:09very good model and it has trained very
9:43:11well. So, now coming down to the moment
9:43:13of truth, so let's now, you know, take a
9:43:15random name and see whether it can
9:43:17predict whether the name is true or
9:43:18false. Okay, so we'll give here as
9:43:20test_name.
9:43:21So, let's give something like, you know,
9:43:24we'll this will be like name, right? So,
9:43:25we'll give name.lower.
9:43:28And now we're going to pass the name
9:43:29over here.
9:43:31So, this would be, let's say, Tom.
9:43:34Okay, so we have to convert this into
9:43:35letters, right? So, it'll be vocab of I
9:43:39for
9:43:40I in test name.
9:43:42And now we'll give the name here as
9:43:44X_test.
9:43:46So, this would be nothing but we have to
9:43:47pad the sequence, pad sequence. We'll
9:43:50pass this in a form of a tuple, then
9:43:51this would be seq.
9:43:54And then we know that we have to pad 10,
9:43:55right? And anyways, we don't have to say
9:43:57whether it's pre or post because by
9:43:59default it's going to be pre.
9:44:01Fine. So, let us now see how our text
9:44:03data would look like. So, X_test.
9:44:06String attribute has Okay.
9:44:08Oh, yeah. So, I have done a typo over
9:44:10here. It's going to be l o w e r.
9:44:13Let me execute this once again.
9:44:15So, as you can see here, we have a
9:44:17matrix which has a size of 10. And let's
9:44:19now see what this would predict. So,
9:44:22it's the moment of truth. So, y.predict.
9:44:25This would be model.predict.
9:44:28And I'm going to pass here as X_test.
9:44:30And let's see what does this predict.
9:44:32So, we'll give here as y_pred. And yeah,
9:44:35let us execute this.
9:44:37Okay, so we are getting this in a form
9:44:38of a array. So, what this tells us, you
9:44:41know, this tells us that, you know, this
9:44:43is like 70% chances that this name is
9:44:46Tom. In order to make this, you know,
9:44:48layman's stuff, so what we're going to
9:44:49do is we'll have if y pred Let me
9:44:52execute this first.
9:44:54So, let me go down to another block.
9:44:56So, this would be something like if
9:44:59y pred is less than 0.5, then we'll say
9:45:03the name is female, okay?
9:45:08Else print name is masculine or name is
9:45:13Always let let be male.
9:45:15Okay. So, same thing over here, we'll
9:45:16just give it as
9:45:18male.
9:45:19So, let's now see what this thing
9:45:20predicts. Okay, so this thing predicts
9:45:22male.
9:45:23So, let's take another common name.
9:45:26Let's take something like Let's go to
9:45:28Google and see what name can we take.
9:45:31Yeah, we can take up something like Brad
9:45:32Pitt.
9:45:33Okay, so let's execute this and let's
9:45:36execute this again.
9:45:37And then this.
9:45:39So, as you can see here, this is giving
9:45:40me a male name. And let's give a female
9:45:42name over here. Let's give as Mary.
9:45:45And let's see whether this would predict
9:45:47it as male or female. So, as you can
9:45:49see, it's female.
9:45:50Now, as if you have seen the data set,
9:45:52right? It says this is the name from the
9:45:54US kids, right? What about What will
9:45:56happen if I give a Indian-based name?
Transformers Neural Networks Explained
9:45:59So, let me give a Indian-based name like
9:46:02Priyanka.
9:46:03And let me execute this.
9:46:05It's giving me a female name, right? So,
9:46:07this is something pretty astonishing,
9:46:09right? So, why do you think it gave me a
9:46:10female name? Well, the reason is because
9:46:12when we are training this model like RNN
9:46:15using RNN, right? It's not looking at
9:46:17the name. It doesn't know whether the
9:46:18name is female or not. But, as a matter
9:46:21of fact, it is looking at the pattern.
9:46:23Okay? So, it might be looking at, you
9:46:25know, if the name ends with so and so,
9:46:26it is a female. If the name starts with
9:46:29this or if a name have something like
9:46:31this, it means, you know, it's a female
9:46:33or a male. So, to give you a better
9:46:35analogy, let's give something like
9:46:37Julia.
9:46:39So, let me execute this.
9:46:41This will give me a female, right? So,
9:46:43what if I give something like Juneid?
9:46:46So, it give it as male. So, all it's
9:46:48trying to do is it's trying to see, you
9:46:50know, uh recognize a pattern. That's why
9:46:52it's taking individual words at the same
9:46:54time. So, this is what makes NLP using
9:46:57LSTM very effective.
9:47:00All right. So, moving ahead, let's see
9:47:02some of the LSTMs use cases.
9:47:04You see, LSTM is a very popular deep
9:47:06learning algorithm for sequential
9:47:08models.
9:47:09Apple Siri and Google's voice search are
9:47:11some of the real-world examples that
9:47:13have used LSTM. And you won't believe
9:47:15it, LSTM is a success story for those
9:47:17algorithm.
9:47:18So, let us now have a look and see how
9:47:20LSTM changed that technology.
9:47:22Okay. So, starting off with Apple, in
9:47:242003, Apple was the first major tech
9:47:26company to integrate a smart assistant,
9:47:28that is Siri, into their operating
9:47:30system. And the Siri was actually a
9:47:32byproduct of some other company. So,
9:47:34Siri was a company's adoption of a
9:47:36standalone app that has been purchased
9:47:38along with the creators who made it. It
9:47:39was somewhere in 2010. The initial
9:47:42reviews about Siri was that it was
9:47:43intense. But, over the next few months
9:47:45or and years, the users became more
9:47:47impatient with the shortcomings. And all
9:47:49too often, it wrongly interpreted
9:47:51commands. And then, you know, no matter
9:47:54what you do, there was no fix for it.
9:47:56So, this is when Apple moved Siri's
9:47:57voice recognition to a neural-based
9:47:59system.
9:48:00Some of the previous technique remained
9:48:02operational, something like, you know,
9:48:03applying hidden Markov models. But, most
9:48:05of the time, you know, DNN or deep
9:48:07neural network using LSTM was used.
9:48:10Although people did not find any changes
9:48:12on the outside, but from within, it was
9:48:14a supercharged deep learning model.
9:48:16Speaking about Google's implementation,
9:48:19Google implemented Google voice search
9:48:20somewhere around 2009.
9:48:22Google voice transcription had initially
9:48:24used something called as Gaussian
9:48:25mixture model.
9:48:27This was This was nothing but an
9:48:28acoustic model and this was something
9:48:29considered to be a state-of-the-art
9:48:31speech recognition for almost 30 plus
9:48:33years.
9:48:34But it was in 2012, there was a boom in
9:48:36deep neural network.
9:48:37And when Google implemented deep neural
9:48:39network, that too using multiple layer
9:48:41networks, there was a huge performance
9:48:43gap.
9:48:44But things really improved when the
9:48:46recurrent neural network, especially
9:48:47with LSTM RNN, first launched on an
9:48:49Android speech recognition in May 2012.
9:48:52Compared to deep neural network, LSTM
9:48:54RNNs have additional recurrent
9:48:56connections and memory cell that allows
9:48:58them to remember the previous data.
9:49:00All right. So, moving ahead to our last
9:49:01topic of our session, let us now see
9:49:03some of the real-world applications of
9:49:04LSTM RNN networks. First off, we can
9:49:07perform named entity recognition. This
9:49:09is something which we did in our
9:49:10previous demo, right? So, what is this
9:49:12named entity recognition? You see, named
9:49:14entity recognition is a subtask of
9:49:16information extraction that seeks to
9:49:18locate and classify named entity
9:49:20mentioned in an unstructured data. Okay?
9:49:23Next, we have something called as
9:49:24sentiment analysis.
9:49:26Sentiment analysis is a predictive
9:49:28modeling task where model is trained to
9:49:30predict the polarity of a textual data
9:49:32or sentiments like positive, neutral, or
9:49:34negative.
9:49:35Sentiment analysis is performed by
9:49:37various businesses to understand their
9:49:39consumers' behavior towards the product.
9:49:42Then we have machine translation. The
9:49:44task of machine translation consists of
9:49:45reading text in one language and
9:49:47generating text in another language.
9:49:49When neural networks are used for this
9:49:51task, we talk about neural machine
9:49:53translation.
9:49:55Within neural machine translation, an
9:49:56encoder-decoder structure is quite a
9:49:58popular LSTM RNN architecture.
9:50:02>> [music]
9:50:06>> Why should we choose deep learning for,
9:50:08you know, various tasks?
9:50:10So, the big advantage of using deep
9:50:12learning is that we can extract more
9:50:14number of features. And when we have
9:50:16more number of features and when we can
9:50:17work at the same time with huge amount
9:50:19of data, we can perceive an object like
9:50:21a human being does. What I'm trying to
9:50:23say over here is like if you want to
9:50:25perform a classification task between
9:50:27pen and a pencil, you'll obviously know
9:50:29as a human being you'll know the
9:50:30difference because you have look at a
9:50:32pen and a pencil continuous number of
9:50:34times. And now when you're trying to
9:50:36actually classify it, you can do it with
9:50:38ease.
9:50:39Okay? And the reason for this is because
9:50:40you know the features of a pen and you
9:50:42know the features of a pencil. Okay?
9:50:44Similarly, this is how deep learning
9:50:46works. More the data you feed, more the
9:50:48dimensions it can analyze. More the
9:50:50dimensions it can learn. All right? So,
9:50:52as I've already mentioned, one of the
9:50:54most popular application of deep
9:50:56learning is image classification. And
9:50:58when it comes to image classification,
9:51:00it can be something as simple as
9:51:01classifying between two different
9:51:02animals to something as complicated as,
9:51:05you know, hiding data or trying to run
9:51:08automated cars using classification
9:51:10task. Okay? All right. So, next type of
9:51:13application using deep learning is using
9:51:15on sequential data. Sequential data
9:51:18basically refers to something like time
9:51:20series data or having to understand
9:51:22natural language.
9:51:23So, the reason why we call it sequential
9:51:25data is because here the previous word
9:51:28or the previous feature is dependent
9:51:30upon the next feature. Okay? So, as you
9:51:32can see over here, we have what time is
9:51:34it, right? So, if I just say it is like
9:51:37over here what time is and it are
9:51:39basically features, right? And in order
9:51:40for you to make an analogy or to
9:51:42understand, obviously have to know what
9:51:44has happened in the past. So, in order
9:51:45to do this, we use something called as
9:51:47RNNs. Okay? And there are various
9:51:49versions of RNN that go around in order
9:51:51to overcome the disadvantages which
9:51:53we'll look in in sometime. All right.
9:51:56So, moving on to the next application
9:51:57that is GANs. GANs, which stands for
9:52:00generative adversarial network, is an
9:52:02unsupervised part of a deep learning
9:52:04application. Some of the common
9:52:06application which you can see in recent
9:52:07days is nothing but deep fakes and many
9:52:09more. Finally, coming down to performing
9:52:11classification and regression task using
9:52:13multi-layer perceptron. If you remember
9:52:15or if you're well versed with machine
9:52:17learning, in order to perform
9:52:18classification in machine learning, we
9:52:20had algorithms like decision tree,
9:52:21random forest, or something very simple
9:52:24as linear regression or logistic
9:52:26regression. But, let me tell you what.
9:52:28When we try to perform classification
9:52:30using MLP, or multi-layer perceptron, we
9:52:32get a very high accuracy even compared
9:52:34to SVM and decision trees. All right.
9:52:37So, now that we know what exactly is
9:52:38deep learning and why we use it, let's
9:52:41now stream down to understand how can we
9:52:43process natural language data using
9:52:45RNNs.
9:52:46So, what are RNNs, right? Well, RNN
9:52:48basically stands for recurrent neural
9:52:50network. And we usually use this in
9:52:52order to deal with a sequential data.
9:52:54Sequential data can be something like a
9:52:56time series data, or a textual data of
9:52:58any format. So, why should one use RNN,
9:53:00right? Well, this is because there's a
9:53:02concept of internal memory here. RNN can
9:53:04remember important things about the
9:53:06input it has received. Which allows them
9:53:09to be very precise in predicting what
9:53:11can be the next outcome. So, this is the
9:53:13reason why they are performed or
9:53:14preferred on a sequential data
9:53:16algorithm, okay? And some of the
9:53:17examples of sequential data can be
9:53:19something like time series, speech,
9:53:21text, financial data, audio, video,
9:53:23weather, and many more. Although RNN
9:53:26were the state-of-the-art algorithm for
9:53:27dealing with sequential data, they come
9:53:29up with their own drawbacks. And some of
9:53:31the popular drawbacks over here can be
9:53:33like, due to the complication or the
9:53:35complexity of the algorithm, the neural
9:53:37network is pretty slow to train. And as
9:53:39there are a huge amount of dimensions
9:53:40here, the training is very long and
9:53:43difficult to do, okay? Apart from that,
9:53:45the most decisive feature for RNN, or
9:53:47for the improvement in RNN, is that of a
9:53:50vanishing gradient. What this vanishing
9:53:52gradient is is that, you know, when we
9:53:54go deeper and deeper into our neural
9:53:56network, the previous data is lost. This
9:53:59is because of a concept called as
9:54:01vanishing gradient. And due to this, we
9:54:03cannot work on a large or a longer
9:54:05sequence of data. Okay? To overcome
9:54:08this, we came up with some new or
9:54:10upgrades to the current recurrent neural
9:54:12networks or RNNs.
9:54:14Starting off with bidirectional
9:54:15recurrent neural network. You see,
9:54:17bidirectional recurrent neural network
9:54:19connect two hidden layers of opposite
9:54:21direction into the same output. With
9:54:23this form of generative deep learning,
9:54:25the output layer can get information
9:54:27from past future states simultaneously.
9:54:30So, as you can see here, we have two
9:54:32layers over here, and as they are
9:54:34bidirectional, what happens is when the
9:54:36algorithm feels that it is kind of
9:54:37losing its gradients or the previous
9:54:39data, it can go back and get the data
9:54:41from the past. So, why do we need
9:54:44bidirectional recurrent neural network?
9:54:45Well, bidirectional recurrent neural
9:54:47network duplicates RNN processing chain
9:54:50so that the input process both forward
9:54:52and reverse time order,
9:54:53thus allowing bidirectional recurrent
9:54:55neural network to look into future
9:54:57context as well. The next one is long
9:54:59short-term memory. Long short-term
9:55:01memory or also sometime referred to as
9:55:03LSTM is a artificial recurrent neural
9:55:05network architecture used in the field
9:55:07of deep learning. Unlike standard
9:55:09feedforward neural network, LSTM has a
9:55:11feedback connections. It can not only
9:55:13process single data point, but also the
9:55:15entire sequence of data. So, as you can
9:55:17see here, from what I'm trying to say is
9:55:19with LSTM or long short-term memory, it
9:55:22has something like, you know, we can
9:55:23feed a longer sequence compared to what
9:55:25it was with bidirectional RNN or RNNs.
9:55:29So, why is LSTM better than RNN? We can
9:55:31say that when we move from RNN to LSTM,
9:55:34we are introducing more and more control
9:55:36over the sequence of the data that we
9:55:38can provide. The LSTM gives us more
9:55:40control ability and does better results.
9:55:43All right. So, the next type of
9:55:44recurrent neural network is the gated
9:55:46recurrent neural network or also
9:55:48referred to as GRUs. You see, GRU is a
9:55:50type of recurrent neural network that
9:55:52is, in certain cases, is advantageous
9:55:55over long short-term memory. GRU makes
9:55:57use of less memory and also is faster
9:55:59than LSTM. But thing is, LSTMs are more
What are GANs?
9:56:02accurate while using longer data sets.
9:56:05I'm sure by now you might have got a
9:56:07hint about the trend that has led to the
9:56:09improvement, right? So, the trend over
9:56:11here is, you know, the model should be
9:56:13capable of remembering and taking in on
9:56:16a longer input sequence.
9:56:18The game-changer part for the sequential
9:56:20data was developed when we came up with
9:56:22something called as transformers. And
9:56:24this paper was something which is based
9:56:26on a concept called as attention is
9:56:29everything.
9:56:30All right. So, let's take a look at
9:56:31this.
9:56:32The paper attention is all you need
9:56:35introduces a novel architecture called
9:56:37as transformers. Like LSTM, transformers
9:56:40is an architecture for transforming one
9:56:42sequence into another while helping
9:56:44adapt to parts, that is encoders and
9:56:46decoders. But it differs from previously
9:56:48described sequence to sequence model
9:56:50because it does not work like GRUs,
9:56:52okay? So, it does not implements uh
9:56:55recurrent neural networks.
9:56:57Recurrent neural network until now were
9:56:59one of the best ways to capture the
9:57:00timely dependence on a sequence.
9:57:03However, the team presenting this paper,
9:57:05that is attention is all you need,
9:57:06proved that an architecture with only
9:57:08attention mechanism does not use RNN can
9:57:11improve its result in translation task
9:57:14and other NLP task. One of the best
9:57:16examples for transformers is Google's
9:57:18BERT. So, what exactly is this
9:57:20transformer, right? You see, here we
9:57:22have encoder on the top and decoder on
9:57:24the bottom. Both encoder and decoder are
9:57:26comprised of modules that can stick onto
9:57:29the top of each other multiple times.
9:57:31So, what happens here is the inputs and
9:57:33outputs are first embedded into
9:57:35N-dimension space since we cannot use
9:57:37this directly. So, we obviously have to
9:57:39encode our inputs, whatever we are
9:57:41providing here. One slight but important
9:57:43part of this model is the positional
9:57:45encoding of different words. Since we
9:57:47have no recurrent neural network that
9:57:49can remember how sequence are fed into
9:57:51the model, we need to somehow give every
9:57:53word or part of our sequence a relative
9:57:55position since the sequence depends on
9:57:58the order of the elements, okay? These
9:58:00positions are added to the embedded
9:58:02representation of each words. All right.
9:58:04So, this was the brief about
9:58:06transformers. So, let us now move ahead
9:58:08and see some of the popular language
9:58:10models that are available in the market.
9:58:12All right. So, let us now start off by
9:58:14understanding OpenAI's GPT-3. The
9:58:16successor to GPT and GPT-2 is the GPT-3
9:58:20and is one of the most controversial
9:58:22pre-trained models by OpenAI. The
9:58:24large-scale transformer-based language
9:58:26model has been trained on 175 billion
9:58:29parameters, which is 10 times more than
9:58:31any previous non-sparse language model.
9:58:34The model has been trained to achieve
9:58:36strong performance on many NLP data set,
9:58:38including tasks like translation,
9:58:41answering questions, as well as several
9:58:42other tasks. Then we have Google's BERT.
9:58:45BERT stands for bidirectional encoder
9:58:47representations from transformers. It is
9:58:50a pre-trained NLP model, which is
9:58:52developed by Google in 2018. With this,
9:58:54anyone in the world can train either
9:58:56their own question answering module with
9:58:58up to 30 minutes on a single cloud TPU
9:59:01or few hours using single GPU. The
9:59:04company then released this showcasing
9:59:06the performance of 11 NLP tasks,
9:59:08including very competitive Stanford
9:59:10dataset questions.
9:59:12Unlike other language model, BERT has
9:59:14only been pre-trained on 250 million
9:59:16words of Wikipedia and 800 million words
9:59:18of book corpus and has been successfully
9:59:21used as a pre-trained model in deep
9:59:23neural network. According to
9:59:24researchers, BERT has achieved 93%
9:59:27accuracy, which has surpassed any
9:59:28previous language models.
9:59:30Next, we have ELMo. ELMo, also known as
9:59:33embedding for language model, is a deep
9:59:35contextualized word representation that
9:59:38models syntax and semantic words, as
9:59:40well as their logistic context. The
9:59:42model developed by Allen NLP has been
9:59:44pre-trained on a huge text corpus and
9:59:47learned functions from bidirectional
9:59:49models, that is BiLM. ELMo can easily be
9:59:52added to their existing models, which
9:59:54drastically improves the features of
9:59:56functions across vast NLP problem,
9:59:59including answering questions, textual
10:00:01entailment, and sentiment analysis.
10:00:06>> [music]
10:00:09>> What are GANs?
10:00:11So, we're going to start with generative
10:00:13models.
10:00:14So, generative models are nothing but
10:00:16those models that use an unsupervised
10:00:19learning approach.
10:00:20In a generative model, there are samples
10:00:23in the data that is input variables X,
10:00:26but it lacks a output variable Y. And we
10:00:29use the only input variables to train
10:00:31the generative model, and it recognizes
10:00:34patterns from the input variables to
10:00:36generate an output that is unknown and
10:00:39based on the training data only.
10:00:41In supervised learning, we are more
10:00:43aligned towards creating predictive
10:00:45models from the input variables.
10:00:48And this type of modeling is also known
10:00:50as discriminative modeling.
10:00:52And in a classification problem, the
10:00:54model has to discriminate as to which
10:00:57class the example belongs to. And on the
10:00:59other hand, unsupervised models are used
10:01:01to create or generate new examples in
10:01:04the input distribution.
10:01:06To define a generative model in layman
10:01:09terms, we can say generative models are
10:01:12able to generate new examples from the
10:01:15sample that are not only similar to the
10:01:17examples, but are indistinguishable as
10:01:20well.
10:01:21And the most common example of a
10:01:22generative model is a naive Bayes
10:01:24classifier, which is more often used as
10:01:27a discriminative model.
10:01:29Other examples of generative models
10:01:31include Gaussian mixture model and a
10:01:33rather modern example, that is
10:01:35generative adversarial networks.
10:01:38So, let us try to understand what
10:01:39exactly are GANs, or generative
10:01:42adversarial networks.
10:01:44Generative adversarial networks, or
10:01:46GANs, are a deep learning-based
10:01:48generative model that is used for
10:01:50unsupervised learning.
10:01:52It is basically a system where two
10:01:54competing neural networks compete with
10:01:56each other to create or generate
10:01:58variations in the data.
10:02:00It was first described in a paper in
10:02:022014 by Ian Goodfellow and a
10:02:05standardized and much stable model
10:02:07theory was proposed by Alec Radford in
10:02:102016, which is also known as DCGAN.
10:02:15Also known as DCGAN or we can call it as
10:02:18deep convolutional generative
10:02:20adversarial networks.
10:02:22And most of the GANs today use deep
10:02:24convolutional generative adversarial
10:02:25networks.
10:02:27The GANs architecture consists of two
10:02:29sub models known as the generator model
10:02:32and the discriminator model.
10:02:34So, a generator network takes a sample
10:02:36and generates sample of data.
10:02:38A discriminator network decides whether
10:02:40the data is generated or taken from the
10:02:42real sample using a binary
10:02:44classification problem with the help of
10:02:46a sigmoid function that gives the output
10:02:49in the form or the range zero and one.
10:02:52So, let us go ahead and take a look at
10:02:54how GANs actually work.
10:02:57To understand how GANs work, let's break
10:02:59it down.
10:03:00So, generative means that the model
10:03:02follows the unsupervised learning
10:03:04approach and is a generative model.
10:03:07When we talk about adversarial, the
10:03:09model is trained in an adversarial
10:03:11setting.
10:03:12And network simply means for the
10:03:14training of the model, we use the neural
10:03:16networks as artificial intelligence
10:03:18algorithms.
10:03:20In GANs, there is a generator network
10:03:22that takes a sample and generates a
10:03:24sample of data.
10:03:26And after this, the discriminator
10:03:27network decides whether the data is
10:03:29generated or taken from the real sample
10:03:31using a binary classification problem
10:03:33with the help of a sigmoid function that
10:03:36gives the output in the range zero to
10:03:37one.
10:03:38The generative model analyzes the
10:03:40distribution of the data in such a way
10:03:42that after the training phase, the
10:03:44probability of the discriminator making
10:03:46a mistake maximizes. and the
10:03:48discriminator on the other hand is based
10:03:50on a model that will estimate the
10:03:52probability that the sample is coming
10:03:54from the real data or not the generator.
10:03:57The whole process can be formalized in a
10:03:59mathematical formula.
10:04:01So G over here is generator, D is equal
10:04:04to discriminator, P data X is the
10:04:07distribution of real data, P data Z is
10:04:10the distributor of generator, X is the
10:04:13sample from the real data, and Z is the
10:04:16sample from generator. Where DX is the
10:04:18discriminator network and GZ is a
10:04:21generator network.
10:04:23So let's take a look at the flowchart
10:04:24once again, guys.
10:04:25So we have the training data, which is
10:04:27going to give the real sample. And the
10:04:29generator network is going to generate
10:04:31the sample from the random noise or the
10:04:33examples.
10:04:34And then it will go to the discriminator
10:04:36network, where it's going to check if
10:04:38the sample that is coming is real or
10:04:41fake.
10:04:42So that is how a GAN actually work.
10:04:44Now let's take a look at the training
10:04:46phase, like how a generative adversarial
10:04:48network is actually trained.
10:04:50So it happens in two phases, guys.
10:04:53So the first phase is where we train the
10:04:55discriminator and we actually freeze the
10:04:57generator, which means that the training
10:05:00set for the generator is done false and
10:05:02the network will only do the forward
10:05:04pass and there will not be any back
10:05:06propagation.
10:05:08Basically, the discriminator is trained
10:05:10with real data and checks if it can
10:05:12predict them correctly.
10:05:14And the same with the fake data to
10:05:15identify them as fake.
10:05:18After this, there's the second part
10:05:20where we train the generator and freeze
10:05:22the discriminator.
10:05:24So we get the result from the first
10:05:25phase and we use them to make better
10:05:28from the previous state to try and fool
10:05:30the discriminator better.
10:05:32So to understand this in the layman's
10:05:33term, I'm going to tell you a few steps
10:05:35for training, like how you should start.
10:05:38So the first step is you have to define
10:05:40in problem.
10:05:41You've got to define the problem and
10:05:43collect the data.
10:05:44After this, the second step is you have
10:05:47to choose the architecture of GAN.
10:05:49So, in this step, depending on your
10:05:50problem, you have to choose how your GAN
10:05:52should look like.
10:05:54The third step is training the
10:05:55discriminator on real data. So, we train
10:05:58the discriminator with real data to
10:06:00predict them as real for n number of
10:06:02times, so we call it a epochs as well.
10:06:04And then we generate the fake inputs
10:06:06from the generator.
10:06:08So, in this step, we are going to
10:06:09generate the fake samples from the
10:06:11generator.
10:06:12And the next step is we train the
10:06:14discriminator on fake data.
10:06:17So, whatever samples are generated from
10:06:18the generator network, you're going to
10:06:20train the discriminator to predict the
10:06:21generated data as fake.
10:06:24So, that's how we know that
10:06:25discriminator is actually predicting the
10:06:27values as correctly.
10:06:28And the last step is we train the
10:06:30generator with the output of
10:06:31discriminator. So, after getting the
10:06:33discriminator predictions, we train the
10:06:36generator to fool the discriminator.
10:06:39So, that's how we train the GAN to
10:06:41actually get our solution from the
10:06:43problem. Which is like defining the
10:06:45problem.
10:06:46So, you'll understand this when I'm
10:06:47talking about the applications, guys. No
10:06:49worry.
10:06:50Now, let's go ahead and take a look at a
10:06:51few challenges of generative adversarial
10:06:54networks.
10:06:55So, the concept of GANs is rather
10:06:57fascinating, but there are a lot of
10:07:00setbacks that can cause a lot of
10:07:01hindrance in its path.
10:07:03Some of the major challenges faced by
10:07:05GANs are
10:07:06The first one is the stability.
10:07:08So, there has to be a stability that is
10:07:09required between discriminator and the
10:07:11generator network, otherwise the whole
10:07:13network would just fall.
10:07:15For example, in case, let's say if the
10:07:17discriminator is too powerful, the
10:07:20generator will fail to train altogether.
10:07:22Won't be able to push fake samples to
10:07:24that discriminator, and it will always
10:07:27identify them as fake.
10:07:29And let's say if the network is too
10:07:30lenient, the discriminator network is
10:07:32too lenient,
10:07:34so any image that would be generated by
10:07:36the generator network would make the
10:07:38network useless.
10:07:39The next challenge that is faced by GANs
10:07:42is GANs fail miserably in determining
10:07:44the positioning of the objects in terms
10:07:46of how many times the objects should
10:07:48occur at that location. Suppose we have
10:07:50a image in which we have, let's say,
10:07:52three dogs with two eyes and sometimes a
10:07:56GAN will fail to, you know, determine
10:07:58the positioning of the objects in terms
10:07:59of it will generate an image with like
10:08:01one dog and six eyes. So, that's kind of
10:08:04a problem that we face while working on
10:08:06GANs.
10:08:07And the next challenge is 3D perspective
10:08:10troubles GANs as it is not able to
10:08:12understand the perspective also.
10:08:15So, it will often give a flat image for
10:08:16a 3D object. So, that's one challenge
10:08:19that we face with GANs as well.
10:08:21And GANs have a problem of understanding
Future of AI/ML
10:08:24the global objects and it cannot
10:08:26differentiate or understand a holistic
10:08:28structure.
10:08:29Like, if you're talking about trees or
10:08:31if you're talking about flowers, that's
10:08:33a problem that GANs will follow.
10:08:35And last but not least, newer types of
10:08:37GANs are more advanced that are brought
10:08:39about that is deep convolutional
10:08:41generative adversarial networks and are
10:08:44expected to overcome these shortcomings
10:08:46altogether. So, that we don't have to
10:08:47worry about these. These are the
10:08:49shortcomings that we face with normal
10:08:51GANs, uh initial generative adversarial
10:08:53networks. Now that they have become more
10:08:55advanced, they actually overcome these
10:08:58shortcomings, so you don't have to
10:08:59worry, guys.
10:09:00So, last but not the least, I want to
10:09:02talk about a few applications of
10:09:04generative adversarial networks.
10:09:06So, the first one is prediction of next
10:09:08frame in a video.
10:09:10So, let's say the prediction of future
10:09:11events in a video frame is made possible
10:09:13with the help of GANs and DVD GAN or we
10:09:17can call it as dual video discriminator
10:09:19GAN can generate a 256 by 256 videos of
10:09:23notable fidelity up to 48 frames in
10:09:26length.
10:09:27And this can be used for various
10:09:28purposes including surveillance in which
10:09:31we can determine the activities in a
10:09:32frame that gets distorted due to other
10:09:35factors like rain, dust, smoke, etc.
10:09:38So, the possibilities are immense with
10:09:40this if you're able to predict the next
10:09:42frame in a video. That actually helps in
10:09:44a lot of things like surveillance,
10:09:45security, and we can predict outcomes
10:09:48based on these frames that we generate
10:09:50from a video.
10:09:51After this comes the text to image
10:09:53generation.
10:09:54So, basically object-driven attentive
10:09:56GAN, which is also known as object GAN,
10:09:58performs the text to image synthesis in
10:10:01two steps. So, the first step is
10:10:03generating the semantic layout and then
10:10:06generating the image by synthesizing the
10:10:07image by using a deconvolutional image
10:10:10generator is the final step.
10:10:13So, this could be used intensively to
10:10:14generate images by understanding the
10:10:16captions, the layouts, and refine
10:10:18details by synthesizing the words.
10:10:21And there is another study about the
10:10:23story GANs that can synthesize the whole
10:10:25storyboards from mere paragraphs.
10:10:28So, that's actually very good idea if
10:10:30you're talking about GANs. So, you can
10:10:31just give a few layouts and captions.
10:10:35Based on that, it will generate image
10:10:36for us.
10:10:38Talking about the next application, we
10:10:39have image to image translation.
10:10:42So, Pix2Pix is a model which is designed
10:10:44for general purpose image to image
10:10:46translation.
10:10:48So, let's say we have three images.
10:10:50We have a real image.
10:10:51Then we'll be having a generated image,
10:10:54which is basically a fake, and then it
10:10:56will be reconstructed to the previous
10:10:58image which was real.
10:11:00So, this is how image to image
10:11:01translation work, guys.
10:11:03And after this, we have enhancing the
10:11:04resolution of an image.
10:11:06So, super-resolution generative
10:11:08adversarial network, or also known as
10:11:10SRGAN, is a GAN which can generate the
10:11:14super-resolution images from
10:11:16low-resolution images with finer details
10:11:18and better quality.
10:11:20So, this is actually a very good
10:11:22application of GANs, guys. The
10:11:24applications can be immense.
10:11:26So, you imagine a higher quality image
10:11:29with finer details generated from a low
10:11:31resolution image. The amount of help it
10:11:34would produce to identify details in
10:11:36lower resolution images can be used for
10:11:38wider purposes including surveillance.
10:11:41We can use it for documentation
10:11:43security. We can use it for detecting
10:11:45patterns, etc.
10:11:47And last but not least, we have
10:11:48interactive image generation.
10:11:51So, GANs can be used to generate
10:11:53interactive images as well.
10:11:55And computer science and artificial
10:11:56intelligence laboratory also known as
10:11:58CSAIL
10:11:59has developed a GAN that can generate 3D
10:12:02models with realistic lighting and
10:12:04reflections enabled by the shape and
10:12:06texture editing.
10:12:08And more recently, researchers have come
10:12:10up with a model that can synthesize a
10:12:13re-enacted face animated by a person's
10:12:16movement while preserving the appearance
10:12:19of the face at the same time.
10:12:21There are a lot more applications we can
10:12:23work on.
10:12:25>> [music]
10:12:30>> Evolution of AI. So, AI as we know today
10:12:33is entirely different from where it
10:12:34started. Back in the 18th century, it
10:12:37was entirely based on myths,
10:12:39speculations, and fiction. But, it did
10:12:41start taking shape and the real
10:12:43initiation in its truest essence took
10:12:45place in 1956.
10:12:47The AI search began with six major
10:12:49design goals. The first one was teach
10:12:51the machines to reason in accordance to
10:12:53perform sophisticated mental tasks like
10:12:55playing chess, providing mathematical
10:12:57theorems, and others. The second one is
10:13:00knowledge representation for machines to
10:13:02interact with the real world as humans
10:13:04do. Like machines needed to be able to
10:13:06identify objects, people, and languages.
10:13:09Programming language Lisp was developed
10:13:11for this very purpose.
10:13:13The third one is teach the machines to
10:13:15plan and navigate around the world we
10:13:17live in. With this, machines could
10:13:18autonomously move around by navigating
10:13:21themselves.
10:13:22The fourth one is enable the machines to
10:13:24process natural language so that they
10:13:26can understand the language,
10:13:28conversations, and the context of
10:13:29speech.
10:13:31The fifth one is train the machines to
10:13:33perceive the way humans do, like touch,
10:13:36feel, sight, hearing, and taste.
10:13:38And general intelligence that included
10:13:40emotional intelligence, intuition, and
10:13:42creativity was the sixth point. Talking
10:13:44about machine learning,
10:13:46machine learning as we know can be
10:13:47remotely explained with the evolution of
10:13:49robots in the past years. Although,
10:13:51machine learning isn't just a machine
10:13:53that is going to learn stuff. It has a
10:13:55lot more to it. Basically, we have data
10:13:57at our bay. We train and test the model,
10:14:00which in this case can be a robot, and
10:14:02then make it to do task relevant to the
10:14:04learning. And then again, learning can
10:14:05be of different types, which is
10:14:07supervised, unsupervised, reinforcement,
10:14:09etc. To know more about machine learning
10:14:11in detail, refer to our machine learning
10:14:13full course tutorial to get on speed.
10:14:15Now, let us go ahead and take a look at
10:14:17what exactly is AI and machine learning.
10:14:20So, what exactly is AI? According to the
10:14:22Merriam-Webster dictionary, artificial
10:14:24intelligence is a branch of computer
10:14:25science dealing with the simulation of
10:14:27intelligent behavior in computers.
10:14:30AI is a technique that enables machines
10:14:32to mimic human behavior. Artificial
10:14:35intelligence is the theory and
10:14:36development of computer systems able to
10:14:38perform task normally requiring human
10:14:40intelligence, such as visual perception,
10:14:43speech recognition, decision-making, and
10:14:45translation between languages.
10:14:47If you ask me, AI is the simulation of
10:14:49human intelligence done by machines
10:14:51programmed by us. The machines need to
10:14:54learn how to reason and do some
10:14:56self-correction as needed along the way.
10:14:58And artificial intelligence is
10:14:59accomplished by studying how human brain
10:15:02thinks, learns, and decide to work while
10:15:04trying to solve a problem. And then
10:15:06using the outcomes of this study as a
10:15:07basis of developing intelligent software
10:15:09and systems.
10:15:11Now, let us go ahead and take a look at
10:15:12what exactly is machine learning. So,
10:15:14machine learning is a concept which
10:15:16allows the machines to learn from
10:15:18examples and experiences.
10:15:20And that too without being explicitly
10:15:22programmed. So, instead of you writing
10:15:24the code, what you do is you feed the
10:15:26data to the generic algorithm and the
10:15:29algorithm or the machine builds the
10:15:30logic based on the given data.
10:15:32Machine learning algorithms are an
10:15:34evolution of normal algorithms and they
10:15:36make your program smarter by allowing
10:15:38them to automatically learn from the
10:15:40data that you provide. Now that we know
10:15:42what AI and ML actually is, let us go
10:15:44ahead and take a look at a few
10:15:45applications of AI and ML today.
10:15:47So, AI and ML can be widely used in so
10:15:49many applications and I have listed down
10:15:52a few applications to give you a wider
10:15:53perspective. So, first of all, I'm going
10:15:55to talk about the usage of AI and ML in
10:15:57healthcare.
10:15:58AI and ML in healthcare is an angel in
10:16:01disguise. To understand this, imagine
10:16:03you have the past data of millions of
10:16:05patients with diseases in the past. Now,
10:16:07all this data can be put to an effective
10:16:10use in a sense that we would be able to
10:16:12detect a disease in early stages with
10:16:15the help of machine learning and
10:16:16artificial intelligence algorithms.
10:16:18And since the numbers never lie, we can
10:16:21be pretty sure about the accuracy of the
10:16:22results. Even so, if we have any doubts,
10:16:25we can always check the accuracy in
10:16:27almost all the cases.
10:16:29Now, let's talk about the usage of AI
10:16:31and ML in finance. So, artificial
10:16:33intelligence in finance is transforming
10:16:35the way we interact with money.
10:16:37AI is helping the financial industry to
10:16:39streamline and optimize processes
10:16:42ranging from credit decisions to
10:16:44quantitative trading and financial risk
10:16:46management. And if we have that figured
10:16:48out, it saves us from a lot of bad and
10:16:51risky decisions.
10:16:53Now, let us go ahead and take a look at
10:16:54object detection in which we use AI and
10:16:56ML. So, object detection today is
10:16:58playing an important part in the IT
10:17:00industry. For example, surveillance has
10:17:02never looked more tech-savvy. Google
10:17:04Lens is one example that uses image
10:17:06recognition to identify the images on
10:17:08the camera in real time.
10:17:09Now, let us go ahead and take a look at
10:17:11AI and ML in risk detection and
10:17:13predictive analysis. So, predictive
10:17:15analysis has proven its metal in the
10:17:16industry already with almost every
10:17:18organization using it to derive
10:17:19conclusions based on previous data. So,
10:17:22one example is how sports franchises,
10:17:25such as a cricket team, would take
10:17:26account of the performance of a player,
10:17:28let's say a batsman, facing the
10:17:30deliveries of bouncer and short ball
10:17:32deliveries. So, they will be able to
10:17:34figure out the best possible outcomes
10:17:36based on the short selection and
10:17:38strategies using the artificial
10:17:40intelligence and machine learning
10:17:41algorithms.
10:17:42So, this is one way we can use
10:17:43predictive analysis or engine it's just
10:17:45one example. We can use it for many
10:17:47purposes. Like, we can use it to predict
10:17:49stock prices that we can do in finance
10:17:51and we can use it to predict the weather
10:17:53based on the hundreds and hundreds of
10:17:55years of data that we already have.
10:17:57Now, let's talk about AI and ML in
10:17:58marketing and advertising. So, marketing
10:18:01and advertising industry is the most
10:18:03benefited with the evolution of AI and
10:18:05ML. They are able to recognize the
10:18:07browsing patterns of users through data
10:18:09and target users with specific content
10:18:11on the internet. For example, to reach
10:18:13out in a subtle way, you only see those
10:18:15ads which interest you. And you must
10:18:17have felt sometimes like you keep
10:18:19getting ads or you know the content on
10:18:21the internet that you talk about or you
10:18:22were talking about or you were thinking
10:18:24about. It's basically nothing but AI and
10:18:26machine learning that is learning a
10:18:28pattern through the data and the your
10:18:30browsing history or your browsing
10:18:31pattern and reaches out to you in the
10:18:33form of targeted marketing.
10:18:35And now that we have talked about the
10:18:36current applications, let us talk about
10:18:38how AI and ML will shape up in the
10:18:40future and how it would look like 10
10:18:42years, maybe 20 years from now.
10:18:44So, we cannot be sure about if AI and ML
10:18:45would shape in the future like we have
10:18:47seen in the movies, but for now, we will
10:18:49stick to the realistic possibilities.
10:18:51Although when I say realistic
10:18:52possibilities, we are not really sure
10:18:54how it would look like, but we can take
10:18:56a guess.
10:18:57So, when we talk about future of health
10:18:59care, health care would seem pretty
10:19:01reachable and advanced. Even now,
10:19:03researchers are working on detecting
10:19:05diseases in the early stages based on
10:19:07the lifestyle and other relevant data.
10:19:09And in the future, we can expect more
10:19:10advancements in the psychological part
10:19:12as well, where we will be able to
10:19:14identify traits and warnings in the
10:19:16early stages and work on it before it
10:19:18gets better off us.
10:19:20Now imagine being able to cure a disease
10:19:22before even getting the hint of it. That
10:19:24is what researchers are aiming for and
10:19:26it looks pretty promising, guys. Let me
10:19:28tell you. Now let's talk about the
10:19:29future of self-driving cars. So
10:19:31self-driving cars looks like a dream
10:19:33come true today. But in the coming
10:19:35decades, we're going to see a lot of
10:19:37developments in the self-driving cars.
10:19:39We would have overcome all the
10:19:40challenges that we face today and I'm
10:19:42not saying we will have levitating cars
10:19:44driving you to your destinations, but it
10:19:47won't be less than a fascinating
10:19:48experience that may look like a dream
10:19:49today or you might have seen in the
10:19:50movies.
10:19:52Now let's talk about the AI ML in
10:19:53manufacturing that would take the
10:19:55future.
10:19:56So robots in manufacturing is one thing
10:19:58that is going to change the future for
10:20:00us.
10:20:01Manufacturing industries will have the
10:20:02best ever workforce and I'm not talking
10:20:05about the human aspect of it. The
10:20:06manufacturing would be so much easier
10:20:08with the robots and since they don't get
10:20:10tired and they won't even ask for leaves
10:20:13or they might. We We never know. And AI
10:20:15and robotics is a risky slope, although,
10:20:18but that is not entirely true.
10:20:19Researchers and experts are working day
10:20:22and night to make it as safe as
10:20:24possible.
10:20:25And then again, let me talk about future
10:20:26of finance with AI and ML. So managing
10:20:29finance and risk detection would become
10:20:31a piece of cake and to understand this
10:20:33in layman terms, you will be able to do
10:20:35your taxes without even lifting a pen.
10:20:37And the fraud detection and trading
10:20:38would become a lot easier and
10:20:40accessible.
10:20:41Financial advisory would take a much
10:20:43advanced shape as we are already seeing
10:20:45it with a lot of trading applications in
10:20:47the market.
10:20:48Then again, we have computer vision
10:20:49which is going to change the future for
10:20:50us. And I'm sure most of you are aware
10:20:52of the concept of God's eye that we have
10:20:54already seen in the movies. Although it
10:20:56is fictional, but not sure for the wrong
10:20:57reasons, but for the greater good, this
10:20:59might be the possibility in the coming
10:21:01years as computer vision has started to
10:21:03overcome a lot of challenges in the
10:21:04real-time image recognition.
10:21:06And then we have the future of NLP
10:21:08natural language processing and it's
10:21:10going to you know it would open a lot of
10:21:12linguistic barriers in the
10:21:13conversational AI. The conversational AI
10:21:15that we see today is limited to certain
10:21:17tasks, but in coming years it could be
10:21:19like a personal assistant or even a life
10:21:21guide as well.
10:21:23And with the recent advancements we are
10:21:24aiming for a very flexible interface
10:21:27that is going to work for everyone with
10:21:29any linguistic experience or any
10:21:31linguistic expectations.
10:21:34And then we have a rather fictional
10:21:36concept that I'm going to talk about
10:21:37which is immortality through AI. So
10:21:39there are scientists and researchers who
10:21:41are trying to figure out a way to map
10:21:43the brain simulation on a computer. So
10:21:45immortality isn't just living until the
10:21:47very eternity, it is in my opinion
10:21:49leaving a legacy and but in hindsight
10:21:52this task is pretty impossible, but we
10:21:53never know in the future this might be a
10:21:55possibility and we will be able to live
10:21:57through a computer where people would
10:21:59have figured out a way to shift all of
10:22:01our brain simulations onto a computer
10:22:04and we'll be able to think on its own
10:22:05like our own very image. So that is one
10:22:08possibility with the future in AI and ML
10:22:11that many researchers are actually
10:22:13aiming for and we might as well get
10:22:15through with it. So hang in there guys
10:22:18and now let me just talk about a key
10:22:20skills of AI and ML specialist. There
10:22:22was a lot of applications that I just
10:22:24told you about. Now let's take a look at
10:22:25what are the skill sets that are
10:22:27required to become a AI ML specialist in
10:22:29today's world.
10:22:30So first of all you have to be familiar
10:22:31with programming in Python
10:22:33and you must be very well aware of maths
10:22:35and statistics as well because it needs
10:22:36a lot of logic to build algorithms and
10:22:38understand them and there's a lot of
10:22:40applied mathematics behind it as well.
10:22:42So you have to be familiar with maths
10:22:43and statistics and programming language
10:22:45which is Python and the versioning tools
10:22:47like TensorFlow, Keras, PyTorch, etc.
10:22:50And then you must have an expertise in
10:22:51one of the following machine learning
10:22:53domains which is image processing,
10:22:55computer vision, language processing,
10:22:56speech signal processing, etc. And then
Artificial Intelligence Interview Questions & Answers
10:22:59you must have an advanced knowledge in
10:23:00deep learning algorithms as well because
10:23:02it is a very important aspect of AIML.
10:23:04And there has to be an effective
10:23:06communication skills and a you have to
10:23:07be a problem solver because in
10:23:09hindsight, if you get a problem, the
10:23:11only thing that an employer seeks from
10:23:13you is the solution of the problem. So,
10:23:15anyways, you have to be a problem solver
10:23:17in that skill set. And you must be
10:23:20experienced in data visualization tools
10:23:21and methods like Tableau, Matplotlib,
10:23:23Power BI, etc. And there has to be a
10:23:25knowledge in ensemble and online
10:23:27learning. And you must have an
10:23:29experience with SQL and other database
10:23:31related languages. And you have to be
10:23:33familiar with parallel computing using
10:23:35GPUs and unique big systems. So, these
10:23:37are the key skills that are required to
10:23:39become an AIML specialist, guys.
10:23:41Now, let me just walk you through the
10:23:42market trends that looks pretty
10:23:44promising for now. So, first of all, I'm
10:23:46going to talk about the market trends
10:23:47and it looks pretty solid, guys, with
10:23:49the amount of data flowing in each year
10:23:51that it is pretty obvious that it will
10:23:53be opening a lot of doors for skilled
10:23:54professionals. And the correct approach
10:23:56is to get skilled since it is still in
10:23:58the evolution phase and we have to scale
10:24:01a lot more possibilities in these
10:24:02domains. So, if you have the skill set
10:24:04for it, it's going to be a very bright
10:24:06future for you guys. And to be specific,
10:24:08there are a lot of opportunities in
10:24:09health care, finance, conversational AI
10:24:12or we can call it chatbots, and object
10:24:14detection, etc. Almost every industry
10:24:16would move to automating their
10:24:17processes. And what else than machine
10:24:19learning and AI to do your job? So, it
10:24:21is the best time to learn AI and ML if
10:24:23you are looking for a bright career
10:24:25right now. And let us go ahead and take
10:24:27a look at the salary trends as well so
10:24:28you get the perspective of how much
10:24:29you're going to get paid. So, I have
10:24:31categorized the salary trends in a few
10:24:33job profiles in AI and ML. For a machine
10:24:35learning engineer, the takeaway fruits
10:24:37of your labor would look around $114,000
10:24:40a year. And for a machine learning
10:24:41scientist, the average salary looks
10:24:43around $120,000 a year and can go as
10:24:46high as $150,000 a year.
10:24:48And for an AI engineer, the average
10:24:50salary is around a $90,000, but it can
10:24:53go as high as $140,000 a year as well.
10:24:56And for an AI researcher, it goes from a
10:24:58$125,000 average to as high as $150,000
10:25:03a year. And now let us go ahead and take
10:25:04a look at a few companies that are
10:25:06hiring right now for AI and ML
10:25:07specialists. And although there are a
10:25:09lot more companies that are hiring for
10:25:11AI and ML specialists right now, I've
10:25:12just listed down a few over here. So, we
10:25:14have Ford Motors, we have Capgemini,
10:25:16Accenture, Dell, Deloitte, Google,
10:25:18Amazon. And there are a lot of startups
10:25:20as well which are actually artificial
10:25:22intelligence and machine learning based.
10:25:24So, there's a lot of opportunity for you
10:25:26guys. And let me just tell you the best
10:25:28approach to actually get a job in AI and
10:25:30ML industry. So, the best approach to
10:25:33find a job as an AI and ML specialist,
10:25:35even if you are a beginner or an
10:25:37experienced professional, this works for
10:25:39everyone. So, first of all, you have to
10:25:41start with a programming language,
10:25:42preferably Python, because it works best
10:25:44with AI and ML algorithms. And after you
10:25:47are done mastering these basics, start
10:25:49with machine learning and artificial
10:25:50intelligence algorithms. And before
10:25:52that, make sure you are sophisticated
10:25:53enough to work with data. I mean, you
10:25:55can analyze the data, clean it, prepare
10:25:57it for model building, and etc. And try
10:25:59to learn all of them with the
10:26:00implementations on unique data instead
10:26:03of the generic data that you find on the
10:26:04internet. The next step would be to
10:26:06learn the advanced concepts in AI and ML
10:26:08like TensorFlow for object detection,
10:26:10speech recognition, image processing,
10:26:11etc. And after you have mastered the
10:26:14versioning tools, you must make sure
10:26:16that you have a credibility in order to
10:26:17get a job. Because as an employer,
10:26:20anyone would look for credible person
10:26:22proficient enough to do the job. And
10:26:24where will you get that? If you have a
10:26:26relevant master's degree, it is well and
10:26:28good. But if you don't have a degree,
10:26:30you can always go for a certification,
10:26:32which will give you the credibility, and
10:26:34it is going to be the best option to
10:26:35prove your mettle. And then there's one
10:26:38way to look at it. I mean, you can take
10:26:39up the Edureka's post graduate program
10:26:41in artificial intelligence and machine
10:26:43learning, which is going to be a a good
10:26:45deal for you because it is an
10:26:46affiliation with a top college and then
10:26:49you are good to go. You can apply for
10:26:51jobs and you'll get the job easily if
10:26:52you have all the skill sets and you have
10:26:54the experience in making relevant, you
10:26:56know, you are working on real-time
10:26:58projects as well. So, these are going to
10:26:59be very useful for you.
10:27:02>> [music]
10:27:07>> Let's get started with our basic level
10:27:09questions.
10:27:10So, first we have what is the difference
10:27:12between AI, machine learning, and deep
10:27:14learning? I'm sure all of you have this
10:27:16question at the top of your mind because
10:27:18there's a huge confusion between AI,
10:27:19machine learning, and deep learning. So,
10:27:21let's try to understand how they are
10:27:23different. Now, first of all, AI came
10:27:25into existence at around 1950s. All
10:27:28right, this was followed by machine
10:27:29learning and then deep learning was
10:27:31introduced. Now, AI basically represents
10:27:34simulated intelligence in machines,
10:27:36which means that it represents any robot
10:27:39or any machine that can mimic the
10:27:41behavior of a human being. Machine
10:27:43learning on the other hand is a practice
10:27:46of getting machines to make decisions
10:27:48without being explicitly programmed to
10:27:50do so. Now, if you don't program a
10:27:52machine, how are you going to let it
10:27:54make decisions? Now, the way machines
10:27:57learn is through data. So, the most
10:27:59important thing in machine learning is
10:28:01the data. All right, you're going to
10:28:02train machines using data so that they
10:28:04can make their own decisions. Next, we
10:28:06have deep learning. Now, deep learning
10:28:08is basically the process of using
10:28:10artificial neural networks to solve
10:28:12complex problems. So, basically you can
10:28:15think of deep learning as a field that
10:28:17tries to mimic our brain. Okay, so how
10:28:20we have neural networks in our brain,
10:28:22that's exactly how deep learning uses
10:28:24the concepts of artificial neural
10:28:26networks in order to solve problems.
10:28:28Now, AI is a subset of data science. So
10:28:31guys, first of all, data science is the
10:28:32process of deriving useful insights from
10:28:35data. All right, it's a process of
10:28:37extracting information from data that
10:28:39will help you solve problems. So, AI is
10:28:42a subset of data science. Now, on the
10:28:44other hand, machine learning is a subset
10:28:46of AI and data science because machine
10:28:48learning comes after AI. So, basically
10:28:51in AI, you're going to make use of
10:28:53techniques and concepts of machine
10:28:55learning in order to solve problems.
10:28:57Then, we have deep learning. So, it's
10:28:59sort of a hierarchy. First, we have data
10:29:01science, then we have AI, then we have
10:29:03machine learning, and then we have deep
10:29:04learning. Deep learning is a subset of
10:29:06machine learning, AI, and data science.
10:29:09Okay, I hope this is clear. Now, the
10:29:11main aim of artificial intelligence is
10:29:13to build machines in such a way that
10:29:15they're capable of thinking like human
10:29:17beings. All right, so basically, they
10:29:19must be able to mimic the behavior of a
10:29:22human being. Now, the aim of machine
10:29:24learning on the other hand is to make
10:29:26machines learn by providing them a lot
10:29:28of data. Okay, once you make a machine
10:29:30learn through data, it's going to be
10:29:32able to solve complex problems and find
10:29:34solutions. Now, the aim of deep learning
10:29:37is to build neural networks that are
10:29:40able to solve more advanced and complex
10:29:42problems. Okay, now like I mentioned,
10:29:44deep learning is like an artificial
10:29:47brain. All right, you're basically
10:29:48building an artificial brain that is
10:29:50able to think exactly like how we do.
10:29:53Okay, that's what deep learning is. It's
10:29:55a little more advanced than machine
10:29:57learning. Now guys, in short, AI,
10:29:59machine learning, and deep learning are
10:30:01used to solve problems through data. So,
10:30:03basically, AI makes use of techniques
10:30:06and methods of machine learning and deep
10:30:08learning to solve problems or to draw
10:30:10useful insights from data. So, this is
10:30:13the difference between AI, machine
10:30:14learning, and deep learning. I hope all
10:30:16of you are clear with this. Now, let's
10:30:18look at our question number two. The
10:30:20question is, "What is artificial
10:30:22intelligence? Give an example of where
10:30:25AI is used on a daily basis." So, there
10:30:27are a lot of definitions of AI on the
10:30:29internet. A few of them are, "Artificial
10:30:32intelligence is an area of computer
10:30:34science that emphasizes on the creation
10:30:37of intelligent machines that work and
10:30:39react like humans. So, like I said,
10:30:41basically, a machine that is able to
10:30:43mimic the behavior of a human being is
10:30:46known as artificial intelligence.
10:30:48Another such definition is the
10:30:49capability of a machine to imitate the
10:30:51intelligent human behavior. All right?
10:30:54So, artificial intelligence, in short,
10:30:56is basically a machine that we created
10:30:58who can act and think like a human
10:31:00being. Now, where do you think AI is
10:31:02used on a daily basis? There are tons of
10:31:05applications that make use of AI, but
10:31:07one of the most popular applications of
10:31:09AI is a Google search engine. Now, if
10:31:12you just open up Google search and you
10:31:13start typing anything, immediately you
10:31:15get recommendations. These
10:31:17recommendations you derive by using
10:31:19machine learning algorithms, by using
10:31:21deep neural networks, and so on. So, on
10:31:24the top of my head, the most general
10:31:25example of AI is a Google search engine.
10:31:28All of us use Google search engine, and
10:31:30we know how quick it is with its results
10:31:32and how relevant searches it gives us.
10:31:35All this is because of AI. All right?
10:31:38Now, let's look at our next question,
10:31:39which states, "What are the different
10:31:41types of AI?" Now, a lot of people might
10:31:43not be aware of this because there are a
10:31:46couple of types of AI or a couple of
10:31:48types of machines which are
10:31:50hypothetical. Okay, we haven't actually
10:31:52implemented these machines in the real
10:31:54world. We just have a theoretical
10:31:56definition of these. Okay, let's look at
10:31:58what I'm talking about. So, first of
10:32:00all, we have reactive machines AI. Now,
10:32:03these machines are all based on the
10:32:05present actions. Okay, they have no
10:32:07memory or they have no concept of
10:32:09storing memory so that they can learn
10:32:11from that experience. They just react at
10:32:14the moment. Okay, so they're based on
10:32:15present actions, and they cannot use
10:32:18previous experiences to form current
10:32:20decisions and update their memory. Then,
10:32:22we have limited memory AI. Now, this
10:32:25type of AI has some temporary storage of
10:32:27memory in it. Now, if we have some
10:32:29memory stored in a machine, we know that
10:32:31it can look back into the memory and it
10:32:34can try to make decisions based on
10:32:36previous or past experiences.
10:32:38So, limited memory AI makes use of that
10:32:40concept. We have temporary memory here.
10:32:42We do not have permanent memory, but one
10:32:45of the top applications of limited
10:32:47memory AI is the self-driving cars. I'm
10:32:50sure all of you have heard of
10:32:51self-driving cars. They make use of
10:32:53limited memory AI in order to run. Then
10:32:56we have theory of mind AI. Now, like I
10:32:59mentioned earlier, there are a couple of
10:33:01types of artificial intelligent machines
10:33:03which are not actually implemented in
10:33:05the real world. An example of that is
10:33:07theory of mind AI. Okay, this is
10:33:09basically an advanced machine which will
10:33:12have the ability to understand emotions,
10:33:15people, and other things in the real
10:33:16world. We might have come close to this
10:33:19type of AI, but we haven't actually
10:33:20developed something that can understand
10:33:22emotions. Next, we have self-aware AI.
10:33:25Now, this is another such example of a
10:33:27machine that is not built in the real
10:33:29world. This basically includes any
10:33:31machine that has consciousness or that
10:33:34can react just like a human being. Okay,
10:33:36so basically a machine that can take own
10:33:38decisions, that can form own
10:33:40conclusions, and these are machines that
10:33:42have the capability of making their own
10:33:44decisions without any human
10:33:45intervention. Now, this kind of AI is
10:33:47not developed, like I mentioned, because
10:33:49it's going to take up a lot of resources
10:33:51and we still haven't reached that peak
10:33:54of evolution yet. Then we have
10:33:56artificial narrow intelligence. Now,
10:33:58these are the general purpose AI that we
10:34:00see on a daily basis. I'm sure all of
10:34:03you have used Google Assistant, you've
10:34:05used Siri. All of that comes under
10:34:07artificial narrow intelligence. After
10:34:09that, we have artificial general
10:34:11intelligence. Now, these are a little
10:34:13more advanced than the artificial narrow
10:34:15intelligence.
10:34:16Then we have artificial superhuman
10:34:18intelligence. Now, these are one of the
10:34:21most advanced type of AIs that are
10:34:23there. Now, like I mentioned earlier,
10:34:25there are a couple of types of
10:34:27artificial intelligent machines which
10:34:29are not actually implemented in the real
10:34:31world. An example of that is artificial
10:34:34super human intelligence. So guys, these
10:34:36were the different types of AI. Now,
10:34:39let's look at the next question which
10:34:41says explain the different domains of
10:34:43artificial intelligence. Now, AI covers
10:34:45a lot of different domains starting with
10:34:48machine learning. Okay, so machine
10:34:50learning like I mentioned earlier is the
10:34:52science of getting computers to act by
10:34:54feeding them data and by letting them
10:34:56learn a few tricks on their own without
10:34:59being programmed to do so. Okay, so
10:35:01you're not explicitly programming the
10:35:03machine, instead you're feeding it a lot
10:35:04of data so that it understands the data
10:35:07and it makes its own decisions. Then we
10:35:09have neural networks. Now, neural
10:35:11networks are basically a set of
10:35:13algorithms or you can say a set of
10:35:14techniques which are modeled in
10:35:17accordance with a human brain. Okay,
10:35:19like I mentioned earlier, deep learning
10:35:21or neural networks is almost the same
10:35:23thing. Deep learning makes use of neural
10:35:25networks in order to solve complex
10:35:27problems. Now, we have robotics. Now,
10:35:29robotics is a subset of AI which
10:35:32includes different branches and
10:35:33applications of robots. These robots are
10:35:36basically artificial agents which act in
10:35:39a real world environment. Okay, so an AI
10:35:41robot works by manipulating the objects
10:35:43in its surrounding by perceiving,
10:35:45moving, and taking relevant actions.
10:35:48Then we have expert systems. Now, an
10:35:50expert system is basically a computer
10:35:52system that mimics the decision-making
10:35:54ability of a human being. Now, I know
10:35:56all of these domains sound very similar,
10:35:59but they have a very different approach
10:36:01with which they solve the problem. All
10:36:03right, that's the main difference
10:36:04between these domains. Next, we have
10:36:06fuzzy logic systems. Now, traditional
10:36:09systems usually give out output in the
10:36:11form of binary. So, usually if you feed
10:36:14something to a machine, it's always in
10:36:15the binary form. The output is also
10:36:17usually in the form of yes, no, true,
10:36:19false, and so on. But when it comes to
10:36:21fuzzy logic, it tries to give an output
10:36:23in the form of degrees of truth. Okay,
10:36:26so it's very different when compared to
10:36:28the traditional computer systems or the
10:36:29traditional programs. Next, we have
10:36:32natural language processing. Now, this
10:36:34is a field of AI that analyzes natural
10:36:37human language to derive useful insights
10:36:40so that it can solve problems. Now, NLP
10:36:42is used majorly in social media
10:36:44platforms. So, Twitter sentimental
10:36:46analysis is done via NLP. Even Facebook
10:36:49uses NLP in a lot of things. All right,
10:36:52so NLP, fuzzy logic, expert systems,
10:36:54machine learning, neural networks, and
10:36:56robotics are the different domains of
10:36:58AI. I hope all of you are clear with the
10:37:00domains. Now, let's look at our next
10:37:03question. Okay, so how is machine
10:37:05learning related to artificial
10:37:07intelligence? There is a huge confusion
10:37:10between machine learning and AI. A lot
10:37:12of people tend to believe that AI and
10:37:13machine learning is one in the same
10:37:15thing. All right, I would say that you
10:37:17cannot compare AI and machine learning
10:37:19because machine learning is a subset of
10:37:21AI. So, basically, AI makes use of
10:37:24machine learning algorithms and machine
10:37:26learning concepts to solve problems.
10:37:28That's the basic difference or that is
10:37:30where the confusion ends. Machine
10:37:32learning is a technique which is
10:37:34implemented in artificial intelligence
10:37:36in order to solve problems. I hope this
10:37:39is clear. Now, let's look at what are
10:37:41the different types of machine learning.
10:37:43So, there are three types of machine
10:37:45learning. We have supervised,
10:37:46unsupervised, and reinforcement
10:37:48learning. Now, supervised learning is
10:37:51the type of learning in which the
10:37:52machine learns by using labeled data.
10:37:55Now, to make you understand, let's look
10:37:56at an example. Okay, let's say that
10:37:58you've input images of apples and
10:38:01oranges to your machine and you've
10:38:03labeled them. You've told the machine
10:38:05like, "Listen, this is the apple, this
10:38:07is an orange, and the output should also
10:38:09look like this." Okay, so you're
10:38:11labeling the input as apple and an
10:38:13orange, and then you're asking the
10:38:15machine to output an apple and an
10:38:17orange. But, when it comes to
10:38:18unsupervised learning, you're not going
10:38:20to label them. You're just going to give
10:38:22them images of apple and oranges, and it
10:38:24has to figure out on its own. It has to
10:38:26try and understand the difference
10:38:28between apple and oranges, try and
10:38:30understand how they look different, or
10:38:32how they have a different color. So,
10:38:34basically in unsupervised learning, you
10:38:35don't have a labeled data set. Okay,
10:38:37you're going to give it an unlabeled
10:38:38data set, and you're going to ask it to
10:38:40find out and classify which is an apple
10:38:43and which is an orange. Okay, that's the
10:38:44difference between supervised and
10:38:46unsupervised. Now, reinforcement
10:38:47learning is comparatively different.
10:38:50Let's imagine that you were put off in
10:38:52an island. Okay, let's say that you were
10:38:53left in an isolated island. What would
10:38:56you do? Now, initially, we'll all panic,
10:38:58and we won't know what to do. But, after
10:39:00a point, you'll start exploring the
10:39:02island. You'll start adapting to the
10:39:04change in the climate conditions, you'll
10:39:06start looking for food, and then you'll
10:39:08try and understand which food is right
10:39:09for you and which food is wrong for you.
10:39:11You know, you'll learn from your
10:39:12experience. So, in reinforcement
10:39:15learning, basically, an agent interacts
10:39:17with its environment by producing
10:39:19actions and discovers errors or rewards.
10:39:22Now, the type of problems that
10:39:23supervised learning is used to solve is
10:39:25regression and classification. When it
10:39:27comes to unsupervised, it is association
10:39:29and clustering. And in reinforcement
10:39:31learning, it's all the reward-based
10:39:33problems. The type of data for
10:39:35supervised learning is labeled data. For
10:39:37unsupervised, it is unlabeled. And for
10:39:39reinforcement, it is no predefined data.
10:39:42Now, when I say no predefined data, I
10:39:43mean that the reinforcement learning
10:39:45agent has to start collecting the data.
10:39:48So, basically, in reinforcement
10:39:50learning, from data collection to model
10:39:52evaluation, it does everything. In terms
10:39:54of training, supervised learning
10:39:56provides external supervision in the
10:39:58form of labeled data set. In
10:40:00unsupervised learning, there's no
10:40:02supervision. That's why it's called
10:40:03unsupervised learning. Again, in
10:40:05reinforcement learning, there's no
10:40:06supervision at all. The agent has to
10:40:08figure everything out. Now, how
10:40:10supervised learning works is uh you map
10:40:13the labeled input to the known output.
10:40:15So, basically you teach the machine like
10:40:17you tell it that this is the input and
10:40:19this has to be the output. When it comes
10:40:21to unsupervised learning, you just
10:40:22provide data to the machine and it has
10:40:24to understand patterns and it has to
10:40:26discover the output. Now, in
10:40:28reinforcement learning, it has to follow
10:40:30the trial and error method. Okay,
10:40:32there's no particular way in which the
10:40:35agent learns. It just has to explore the
10:40:37environment, try out a few things, and
10:40:39learn from that experience. Popular
10:40:41supervised learning algorithms include
10:40:43linear regression, logistic regression.
10:40:46For unsupervised, we have K-means. And
10:40:47for reinforcement learning, we have
10:40:49Q-learning. So, guys, these were the
10:40:51different types of machine learning and
10:40:53I also discussed the difference between
10:40:55the three. Now, let's move on and look
10:40:57at our next question, which is what is
10:40:59Q-learning?
10:41:00In the previous slide itself, I told you
10:41:02that a type of reinforcement learning
10:41:03algorithm is Q-learning. So, basically,
10:41:06here what happens is an agent tries to
10:41:08learn the optimal policy from its past
10:41:11experience with the environment. The
10:41:13past experience of an agent are a
10:41:15sequence of action, state, and rewards.
10:41:18So, what happens is, first of all, you
10:41:20take an agent and you put it in state
10:41:22zero. Okay, let's say there's some state
10:41:24known as state zero. Now, this agent is
10:41:27going to perform some action A0. On
10:41:30performing this action, it is going to
10:41:32get a reward R1. And if it gets a reward
10:41:35R1, then it's going to move to state S1.
10:41:38But in case the action is wrong, then
10:41:40it's going to get a negative reward, as
10:41:43in some points are going to be reduced.
10:41:45So, guys, think of Q-learning as a game.
10:41:47You're in state zero, and then you do
10:41:49some action, and either you get a reward
10:41:51and go to the next state, or else you
10:41:53lose and you go back to the same state.
10:41:56So, until you learn, you're going to be
10:41:57in the same state. But if you keep
10:41:59learning and if you keep receiving
10:42:01positive rewards, then you're going to
10:42:02move on to state one, and similarly, you
10:42:04move on to state two, three, and so on.
10:42:06This is what Q learning is about.
10:42:08Now, the next question is what is deep
10:42:10learning? Now, deep learning, like I
10:42:12mentioned earlier, basically mimics the
10:42:15way our brain works. Okay, it learns
10:42:17from experience. Now, the main concept
10:42:19behind deep learning is neural networks.
10:42:22In our brain also, we have neural
10:42:23networks. [clears throat] So, what deep
10:42:24learning tries to do is it tries to use
10:42:26the concept of neural networks in order
10:42:29to solve complex problems. So,
10:42:30basically, we're trying to mimic our
10:42:32brain. Any deep neural network will have
10:42:35three types of layers. The first is the
10:42:37input layer. Now, this layer will
10:42:39basically receive all the input, and it
10:42:41will forward them to the hidden layer.
10:42:43Now, in the hidden layer, all the
10:42:45analysis and the computation takes
10:42:47place. All right, once the computation
10:42:49is done, the result is transferred to
10:42:51the output layer. Now, there can be n
10:42:53number of hidden layers depending on the
10:42:55type of problem you're trying to solve.
10:42:57Then, we have the output layer. So,
10:42:59basically, this layer is responsible for
10:43:01transferring the information from the
10:43:03neural network to the outside world. So,
10:43:05it's as simple as that. It's pretty
10:43:07obvious. Input layer will take in the
10:43:09input, hidden layer will perform the
10:43:11computations, and the output layer will
10:43:13give out the output. This is a small
10:43:15explanation of what deep learning is.
10:43:17Now, of course, this is much more
10:43:18complex than this, but in short, this is
10:43:21exactly what deep learning is. Now,
10:43:23let's look at our next question, which
10:43:25is explain how deep learning works. So,
10:43:28basically, deep learning is a concept
10:43:30based on something known as neuron.
10:43:33Okay, neuron is a basic unit of the
10:43:35brain. Inspired from this neuron, they
10:43:37came up with something known as
10:43:39perceptrons or artificial neurons. Now,
10:43:42in this image on the left-hand side, you
10:43:44can see that there is something known as
10:43:46dendrite. These are modules which
10:43:48receive the input. It basically receives
10:43:50all the signals that we send to our
10:43:52brain. Okay, similar to the dendrites
10:43:55are the input layer in our artificial
10:43:57neural networks. Now, in the previous
10:43:59slide, we discussed that the input layer
10:44:00takes in all the input from the outside.
10:44:03That's exactly what a dendrite does. So,
10:44:05basically a perceptron receives multiple
10:44:08inputs. It applies various
10:44:10transformations and functions, and then
10:44:12it provides an output.
10:44:14So, basically guys, just like how our
10:44:16brain contains multiple connected
10:44:18neurons called neural networks, we also
10:44:21have a network of artificial neurons
10:44:23called perceptrons to form a deep neural
10:44:25network. So, basically an artificial
10:44:27neuron or a perceptron, it models a
10:44:30neuron which has a set of inputs, each
10:44:33of which is assigned some specific
10:44:35weight. Okay, all of these inputs will
10:44:37have a specific weight, and the neuron
10:44:39will compute some function on these
10:44:41weighted inputs and give you the output.
10:44:44So, the neuron will basically perform
10:44:45analysis and all of that on these
10:44:47weighted inputs to give you some output.
10:44:50This is a basic concept of deep
10:44:52learning. So, there are inputs which
10:44:54have some weight on it, and these inputs
10:44:56are then formulated and analyzed in
10:44:59order to give you an output. Now, let's
10:45:01look at our next question, which is
10:45:03explain the commonly used artificial
10:45:05neural networks. Now, this is a very
10:45:07theoretical question because in order to
10:45:09make you understand how each of them
10:45:11work will take a lot of time. Okay, so
10:45:13I'm just going to briefly tell you what
10:45:15each of these networks are and what they
10:45:17do. Now, feedforward neural network is
10:45:19the most basic kind of artificial neural
10:45:21network. So, basically the feedforward
10:45:23neural network is unidirectional. The
10:45:26data passes through the input nodes and
10:45:28leaves through the output nodes. In
10:45:30feedforward neural network, usually the
10:45:32number of hidden layers depends on the
10:45:34complexity of the problem. Coming to
10:45:36convolutional neural networks, here
10:45:39basically the input features are taken
10:45:42in small sets. Okay, or they're taken in
10:45:44batches. This will help the network
10:45:47remember better because you're feeding
10:45:49batches of images or you're feeding
10:45:51batches of input to the neural network.
10:45:54Now, this type of neural network is
10:45:56mainly used for signal and image
10:45:57processing. Next, we have recurrent
10:46:00neural networks. These are also known as
10:46:03long short-term memory networks. So,
10:46:05this basically works on the principle of
10:46:07feeding the output of a layer back into
10:46:10the input layer in order to predict the
10:46:12outcomes. Okay, this way it's more
10:46:14precise and it is a little more complex
10:46:16when compared to convolutional networks.
10:46:18Now, one main important point of
10:46:20recurrent neural networks is that they
10:46:22have something known as memory. So,
10:46:24basically each neuron will have some
10:46:26information or some memory stored in
10:46:28them so that if they can use this memory
10:46:31in order to take actions in the future.
10:46:33So, they have some experiences stored in
10:46:36the form of memory so that they can make
10:46:37their decisions based on previous
10:46:39actions. Now, finally, we have
10:46:41autoencoders. Now, autoencoders are
10:46:44mainly used in dimensionality reduction
10:46:46for learning generative models. Okay,
10:46:49and one more important thing about
10:46:50autoencoders is that the number of units
10:46:53in the output layer and the input layer
10:46:55is the same. This is because the output
10:46:57layer has to reconstruct its own inputs.
10:46:59So, these were the different types of
10:47:02artificial neural networks. Now, let's
10:47:04look at our next question, which is what
10:47:06are Bayesian networks? Okay, so a
10:47:08Bayesian network is a statistical model
10:47:11that represents a set of variables and
10:47:13the conditional dependencies in the form
10:47:16of a directed acyclic graph. Now,
10:47:18basically, on the occurrence of any
10:47:20event, a Bayesian network can be used to
10:47:22predict the likelihood that any one of
10:47:25several possible known causes was a
10:47:27contributing factor. An example of this
10:47:30is a Bayesian network could be used to
10:47:32study the relationship between diseases
10:47:34and symptoms. So, given a set of
10:47:36symptoms, the Bayesian network can be
10:47:38used to find out the probability of the
10:47:41presence of any diseases. All right, so
10:47:43the next question is explain the
10:47:45assessment that is used to test the
10:47:47intelligence of a machine. Now, guys,
10:47:49this is a a common question and it is
10:47:52sort of a general knowledge-based
10:47:53question. All right, I'm hoping that
10:47:55most of you know the answer to this. So,
10:47:57let's look at what the answer is. I'm
10:48:00not sure how many of you have heard of
10:48:01Alan Turing. So, Alan Turing was the one
10:48:04who came up with the Turing test. Now,
10:48:06this test is basically to determine
10:48:08whether or not a computer is capable of
10:48:11thinking like a human being.
10:48:13So, if a machine or if a computer passes
10:48:15this exam, it means that that machine is
10:48:18capable of thinking like a human being.
10:48:20It means that it is successfully an
10:48:22artificial intelligent machine, meaning
10:48:24that it can make its own decisions and
10:48:26interpret data and form their own
10:48:28formulations or form their own
10:48:30conclusions about the data. Now, sadly,
10:48:32I don't think there are a lot of
10:48:33machines that have passed the Turing
10:48:35test. In fact, I'm not sure if there is
10:48:37any machine that's passed the Turing
10:48:39test as of now, but in the near future,
10:48:41I'm sure that we'll see machines who are
10:48:44more smarter than human beings and who
10:48:46have passed this test. Now, for a
10:48:48machine, it might be very easy to do
10:48:50computations, but it might be very hard
10:48:52for a machine to just get up and walk
10:48:54around. All right, the simple things
10:48:56that us humans can do is very
10:48:58complicated for a machine. They can do
10:49:00computations which we can do in probably
10:49:02a year, they can do those computations
10:49:04in maybe a week or less than a week.
10:49:07But, doing simple things such as walking
10:49:09up to the fridge or walking up to the
10:49:11kitchen is very hard for the machines.
10:49:14So, to achieve that level of
10:49:15intelligence, we're going to take a
10:49:16while, but in the near future, I'm sure
10:49:18we'll see machines which are way more
10:49:20capable than human beings. Now, let's
10:49:22move on to our next level. Now, here
10:49:25I'll basically be discussing
10:49:27intermediate level artificial
10:49:28intelligence questions. So, let's look
10:49:30at the first question. All right, the
10:49:32first question is how does reinforcement
10:49:35learning work? Explain with an example.
10:49:38Okay, so first of all, reinforcement
10:49:40learning is a type of machine learning.
10:49:42We discussed about reinforcement
10:49:44learning earlier. Reinforcement learning
10:49:46is a type of machine learning wherein
10:49:48there's an agent and you put this agent
10:49:50in an unknown environment. All right,
10:49:52now the agent has to figure out actions,
10:49:54what sort of actions it must take, and
10:49:56how it's going to get rewards so that it
10:49:58can move from state zero to state one.
10:50:01It's sort of like a video game. If
10:50:02you're in a video game, let's say if
10:50:04you're playing Counter-Strike, you're in
10:50:06level zero or state zero. Now, if you
10:50:09perform some action and if you get some
10:50:11rewards, you're going to move to state
10:50:12one. That's exactly how reinforcement
10:50:14learning works. If you perform the
10:50:16relevant actions and the correct
10:50:17actions, you're going to get a reward
10:50:19and you'll move on to the next state.
10:50:21But in case you perform a wrong action,
10:50:23you'll get negative rewards and you'll
10:50:25stay in the same state unless and until
10:50:27you don't learn. All right, so if you
10:50:29learn and achieve, then you'll move to
10:50:31the next state. So, basically a
10:50:33reinforcement learning system will have
10:50:35two main components. It'll have an agent
10:50:38and an environment. Now, the agent I've
10:50:40been repetitively saying an agent An
10:50:42agent is basically the reinforcement
10:50:44learning algorithm. It is the model. The
10:50:47model has to learn everything on its
10:50:48own. It has to collect data on its own.
10:50:51It has to draw useful insights on its
10:50:53own. Okay, you're not going to feed any
10:50:55predefined data to this reinforcement
10:50:57learning agent. All right, he has to
10:50:59figure out everything on its own. So,
10:51:01let's look at an example of
10:51:02Counter-Strike. Okay, I'm not sure how
10:51:04many of you play the game, but yeah.
10:51:06What happens here is the reinforcement
10:51:08learning agent or the player one
10:51:10collects a state S0 from the
10:51:12environment. Okay, so let's suppose that
10:51:14you're playing Counter-Strike and you're
10:51:16in state zero. Now, you'll perform some
10:51:19action A0. All right, initially it's
10:51:21going to be a random action. So,
10:51:23obviously if you're put in an unknown
10:51:24environment, your first action is going
10:51:26to be random, correct? Because you don't
10:51:28know what's right, you don't know what's
10:51:29wrong. So, in your state zero, you'll
10:51:31take an action A0. This will result in a
10:51:34new state S1. And on achieving state S1,
10:51:38the agent will get a reward R1, okay,
10:51:40from the environment. Now, in the case
10:51:43of Counter-Strike games, if you've
10:51:45observed, whenever you win a state or
10:51:47you pass a level, you're going to get
10:51:49some rewards. Maybe you'll get more
10:51:51weapons or you'll get more points. Okay,
10:51:53just like that, in reinforcement
10:51:55learning problem, you'll get some reward
10:51:57R1. Okay, it's basically a plus point.
10:52:00You might get a negative reward or a
10:52:02positive reward based on the action that
10:52:04you take. Now, this loop will go on
10:52:07until the agent is dead or it reaches
10:52:09the destination. So, in Counter-Strike,
10:52:11until you have failed the level, you
10:52:14will keep playing the game, right?
10:52:15You'll keep moving from state one, state
10:52:17two, state three, and so on. Or, if
10:52:19you've reached the destination, then
10:52:20it's the end game. That's exactly how it
10:52:23works in reinforcement learning. If the
10:52:25agent has explored the entire
10:52:26environment and reached the end state,
10:52:29that's when the loop will end. All
10:52:31right, that's exactly how reinforcement
10:52:33learning works. It is very similar to
10:52:35the games that we play. All right, it's
10:52:37very understandable. Now, let's move on
10:52:39and discuss the next question. So, the
10:52:41next question is explain Markov decision
10:52:44process with an example. Now, the
10:52:46solution for a reinforcement learning
10:52:48problem is achieved through the Markov
10:52:50decision process. It's basically a
10:52:53mathematical approach that maps the
10:52:55solution in reinforcement learning.
10:52:57Okay, so now to understand this, there
10:52:59are a couple of parameters in a Markov
10:53:02decision process. They're going to be a
10:53:04set of actions called A. Okay, you can
10:53:06name them A. A set of states, there's
10:53:09going to be reward, there's going to be
10:53:11policy, and there's going to be value.
10:53:13To sum it up, what exactly happens in a
10:53:15Markov decision process is that the
10:53:18agent takes an action A to transition
10:53:21from the start state to the end state.
10:53:23Now, while doing so, the agent receives
10:53:26some reward R for each action that he
10:53:28takes. The series of actions taken by
10:53:30the agent will define a policy or an
10:53:33approach. And the rewards collected will
10:53:36define the value. So, the main goal in a
10:53:39Markov decision process is to maximize
10:53:41the rewards by choosing the most optimum
10:53:44policy. Meaning that you're going to
10:53:46choose the best path or the best
10:53:47solution in order to get the most number
10:53:50of rewards. Now, in order to make you
10:53:52all understand this better, let's solve
10:53:54the shortest path problem by using
10:53:56Markov decision process. I'm sure all of
10:53:59you have heard of shortest path problem.
10:54:01This was I think taught to us when we
10:54:03were in 11th or 12th, I'm not sure. Look
10:54:06at the diagram that is over here. This
10:54:08is basically a representation of our
10:54:11problem. Given this representation, our
10:54:13goal here is to find the shortest path
10:54:16between the node A and node D. All
10:54:18right, you can see nodes A, B, C, and D.
10:54:21We have to find the shortest path
10:54:23between node A and node D. Now, the link
10:54:25between these two nodes has a number on
10:54:28it. Okay, for example, between A and C
10:54:30you can see there's a number 15. Okay,
10:54:32this basically denotes the cost to
10:54:34traverse that edge. So, if you want to
10:54:36go from A to C, you'll spend around 15
10:54:39points. So, our end goal here is to
10:54:41travel between node A and node D with
10:54:44minimal possible cost. We should travel
10:54:47between A to D in such a way that our
10:54:49cost is minimal. Now, in this problem if
10:54:51you notice that we have a set of states.
10:54:54Okay, these are denoted by the nodes A,
10:54:56B, C, D. Now, like I mentioned earlier,
10:54:59a Markov decision process has a set of
10:55:01states. Similarly, in this problem the
10:55:03set of states are A, B, C, D. The action
10:55:06is to traverse from one node to the
10:55:08other. So, going from A to B is
10:55:10basically an action. Going from A to C
10:55:12is another action. Going from A to D is
10:55:15another action, and so on. Now, reward
10:55:17is represented by the cost on each of
10:55:20these links. And the policy is the path
10:55:22which is taken to reach the destination.
10:55:25So, our aim here is to choose a policy
10:55:28that gets us to node D in the minimum
10:55:30cost possible. So, how do you think you
10:55:33can solve this problem? All right, you
10:55:34can start off at node A and you can take
10:55:37baby steps to your destination. Now,
10:55:39initially only the next possible node is
10:55:41visible to you. Like I mentioned
10:55:43earlier, the initial action taken in a
10:55:46reinforcement learning problem is always
10:55:48random. So, at random you'll choose any
10:55:50node. Let's say you take A to B. Now, if
10:55:53you go from A to B, you can go B to D
10:55:55and you'll reach the destination. So,
10:55:57policy is the path which is taken to
10:56:00reach the destination. All right, so it
10:56:02can go from A to B to D or you can go
10:56:04from A to C to D or you can go A C B D.
10:56:08All right, now it's up to you to figure
10:56:09out which is the shortest path. All
10:56:11right, you have to choose a path in such
10:56:13a way that the cost between A to D is
10:56:16minimized. So guys, this was a simple
10:56:18problem of how Markov decision process
10:56:21is used to solve the shortest path
10:56:22problem. Now, let's move on and look at
10:56:25our next question. All right, now the
10:56:27next question is explain reward maximiza
10:56:30tion in reinforcement learning. So,
10:56:32basically a reinforcement learning agent
10:56:35works based on the theory of reward
10:56:37maximization. Okay, in the previous
10:56:39question itself I told you that the main
10:56:41aim of reinforcement learning is to
10:56:43maximize the reward. So, that's why a
10:56:46reinforcement learning agent must be
10:56:48trained in such a way that he takes the
10:56:50best action so that the reward is
10:56:52maximum. Okay, this is exactly what
10:56:54reward maximization means. He has to
10:56:57choose the best policy in such a way
10:56:59that the reward is maximum. Now, let me
10:57:02explain this with a small game. So, in
10:57:04the figure you can see a fox, you can
10:57:06see some meat and you can see a tiger.
10:57:08Now, our reinforcement learning agent is
10:57:10the fox. His end goal is to eat the
10:57:13maximum amount of meat before being
10:57:15eaten by the tiger. Okay, so he has to
10:57:18explore around, eat the maximum number
10:57:20of meat that he can eat before the tiger
10:57:22kills him. Since the fox is a clever
10:57:25fellow, he eats the meat that is closer
10:57:27to him. Okay, so rather than eating the
10:57:30meat which is close to the tiger, he
10:57:32eats the meat which is only close to
10:57:33him. This is because the closer he gets
10:57:35to the tiger, the higher are his chances
10:57:38of getting killed. So, as a result of
10:57:40this, the rewards near the tiger, even
10:57:43if they are bigger meat chunks, will be
10:57:45discounted. So, because the fox is not
10:57:48going closer to the tiger and eating the
10:57:50meat chunks closer to the tiger, this
10:57:52reward will get discounted. Now, I know
10:57:55you're all wondering what discounted is.
10:57:57Now, this is done because of the
10:57:59uncertainty factor that the tiger might
10:58:01kill the fox. So, what is discounting of
10:58:04reward? Okay, how does it work? To
10:58:06understand this, we define a discount
10:58:09rate called gamma. Okay, this is a
10:58:10parameter and the value of gamma always
10:58:14ranges between zero and one. So, the
10:58:16smaller the gamma, the larger the
10:58:18discount and so on. So, guys, this was
10:58:20reward maximization. So, here basically
10:58:23the fox will try to get as much as meat
10:58:25chunks as he can and he'll also try to
10:58:28avoid getting killed because that will
10:58:30end the reinforcement learning loop. We
10:58:32also discussed the discounted factor.
10:58:34All right, now that is not needed to
10:58:36understand reward maximization, but I
10:58:38just thought I'll add on some extra
10:58:39info.
10:58:40Now, let's look at the next question
10:58:42which is what is exploitation and
10:58:44exploration trade-off?
10:58:46So, basically exploration is, like the
10:58:49name suggests, it is about exploring and
10:58:52capturing more information about an
10:58:53environment. Now, on the other hand,
10:58:56exploitation is about using the already
10:58:58known exploited information to heighten
10:59:01the rewards. So, consider the same
10:59:03example that we discussed in the
10:59:04previous question. Here, the fox only
10:59:07eats the meat chunks which are close to
10:59:09him. Okay, he does not eat the bigger
10:59:11chunks because even though the bigger
10:59:13chunks would give him more rewards, it
10:59:15would get him killed. Okay, he does not
10:59:17go towards the tiger itself. Now, if the
10:59:19fox only focuses on the closest reward,
10:59:23he will never reach the big chunks of
10:59:24meat. Okay, this is what exploitation
10:59:27is. He's sticking only to the
10:59:29information that he knows and he's
10:59:31trying to get the most number of rewards
10:59:32from it. But if the fox decides to
10:59:35explore a bit, it can find the bigger
10:59:37rewards. Okay, the bigger rewards are
10:59:39basically the big chunks of meat which
10:59:41are near the tiger. And this is exactly
10:59:43what exploration is. Okay, exploitation
10:59:46is about using the already known
10:59:48information to heighten your rewards.
10:59:51Exploration on the other hand is about
10:59:53exploring and capturing more information
10:59:55about an environment.
10:59:57All right, so that was about
10:59:58exploitation and exploration. Now, let's
11:00:01move on to our next question. So, this
11:00:04is a difference question which ask the
11:00:07difference between parametric and
11:00:09non-parametric models. So, a parametric
11:00:12model basically uses a fixed number of
11:00:14parameters to build the model. Now,
11:00:16first of all guys, what are parameters?
11:00:18Now, parameters are basically predictor
11:00:21variables that are used to build a
11:00:23machine learning model or build any
11:00:25predictive analytics model. Now that we
11:00:27know what parameters are, let's try to
11:00:29understand the difference between a
11:00:31parametric and a non-parametric model.
11:00:34Now, parametric model basically uses a
11:00:35fixed number of parameters to build the
11:00:37model. A non-parametric model uses
11:00:40flexible number of parameters to build
11:00:42the model. When it comes to a parametric
11:00:44model, the assumptions about the data
11:00:46are very strong. In a non-parametric
11:00:49model, there are fewer assumptions about
11:00:51the data. A parametric model has the
11:00:53fixed number of parameters. Everything
11:00:55is defined over here. So, the
11:00:57computation is very fast. Okay, you know
11:00:59what sort of variables you'll need to
11:01:01predict the outcome. Okay, you have a
11:01:03defined set of variables or a defined
11:01:05set of predictor variables that will
11:01:07compute your outcome. So, that's why the
11:01:09computation is a bit faster. When you
11:01:12compare to non-parametric models, there
11:01:14are a lot of parameters taken into
11:01:16account. Now, when it comes to a
11:01:18non-parametric model, you do not have a
11:01:20fixed number of parameters. All right,
11:01:22you do not have a fixed number of
11:01:24predictor variables that will help you
11:01:26get to the outcome. So, the computation
11:01:28is a bit slower. Now, parametric models
11:01:31require lesser data and non-parametric
11:01:34require more data. Example of parametric
11:01:37models include logistic regression and
11:01:39naive bias. And for non-parametric
11:01:41models, we have KNN and decision tree
11:01:43models. Now, logistic and naive bias
11:01:46models are very strong models because
11:01:49they have a fixed number of parameters
11:01:51or a fixed number of predictor variables
11:01:53and they will give you an immediate
11:01:55output. Okay, when it comes to
11:01:57non-parametric models like decision tree
11:01:58models and KNN, you might even observe a
11:02:01little bit of overfitting. Okay, this
11:02:03happens because you have a fewer number
11:02:06of assumptions about the data and also
11:02:08because your parameters are not fixed.
11:02:10Now, that's not the reason for
11:02:12overfitting, but it's seen that in some
11:02:14of the non-parametric models,
11:02:16overfitting occurs more often. Now,
11:02:18let's discuss the next question, which
11:02:21is what is the difference between
11:02:22hyperparameters and model parameters?
11:02:25Now, model parameters are the predictor
11:02:27variables that I was speaking about
11:02:29earlier. Hyperparameters, let's discuss
11:02:32what they are. Okay, model parameters
11:02:34are the features of training data that
11:02:36will learn on its own during training.
11:02:38Whereas model hyperparameters are the
11:02:40parameters that determine the training
11:02:42process.
11:02:43Now, let's see that you want to
11:02:45determine the height of an individual
11:02:47depending on his weight. The height and
11:02:50weight will become your model
11:02:52parameters. But, your hyperparameter is
11:02:55basically the learning rate. It's the
11:02:57rate at which your model is going to
11:02:59learn this correlation between the
11:03:01height and the weight. So, this is the
11:03:03difference between model parameters and
11:03:05hyperparameters. Model parameters are
11:03:07the ones that you find in your data.
11:03:10These are all the variables that you use
11:03:12to predict your outcomes.
11:03:13Hyperparameters will define your
11:03:15training process. There is a huge
11:03:17difference between model and
11:03:18hyperparameters. Another difference is
11:03:21that they are internal to the model and
11:03:23their value can be estimated from the
11:03:24data. Hyperparameters are external to
11:03:27the model and their value cannot be
11:03:29estimated from data. Now, like I said,
11:03:31model parameters are derived from your
11:03:33data itself. Okay, these are the
11:03:35parameters that are there in your data.
11:03:37Hyperparameters are the ones that you
11:03:39define in order to train your entire
11:03:42data. So, that is the difference between
11:03:44hyperparameters and model parameters.
11:03:47So, next question is what are
11:03:48hyperparameters in deep neural networks?
11:03:51So, guys, like I mentioned in the
11:03:53previous example, hyperparameters are
11:03:56variables such as the learning rate.
11:03:58This will define how your entire data
11:04:00training process goes. For those of you
11:04:02don't know, in order to build a model,
11:04:04you first need to train the model and
11:04:06then you need to test it. Okay, now
11:04:08while training the model, you're going
11:04:09to make the model learn a lot of things.
11:04:12You're going to give it a lot of data.
11:04:14It has to figure out relations between
11:04:16various variables and how these
11:04:18variables are affecting the output. All
11:04:21of this training will depend on a few
11:04:23variables such as the learning rate.
11:04:25Okay, these are basically called
11:04:26hyperparameters in deep neural networks.
11:04:29So, these parameters will define the
11:04:31number of hidden layers that are present
11:04:33in the network. Okay, and more the
11:04:35number of hidden layers, the more
11:04:36accurate your network is going to be.
11:04:39Whereas, if you have less a number of
11:04:40units, you may cause underfitting in
11:04:42your data. Underfitting will also result
11:04:45in inaccurate predictions. So, that's
11:04:47why you need to make sure that the
11:04:49number of hidden units in your hidden
11:04:50layers are perfect, are ideal. Okay, and
11:04:53this is determined by your learning rate
11:04:56or by your hyperparameters. Not by your
11:04:58learning rate specifically, but by the
11:05:00number of hyperparameters you have.
11:05:02These number of hidden layers are
11:05:04determined by the hyperparameters. Okay,
11:05:07that's why hyper parameters are very
11:05:08important in deep neural networks. I
11:05:11hope you all are clear with this. Now,
11:05:13let's look at our next question, which
11:05:14is explain the different algorithms used
11:05:17for hyperparameter optimization. We'll
11:05:20discuss the three methods, which are
11:05:21grid search, random search, and Bayesian
11:05:24optimization. Okay, now grid search
11:05:26basically will train the network only on
11:05:29the two sets of hyper parameters, which
11:05:31are learning rate and the number of
11:05:32layers. Okay, so it's going to use every
11:05:34combination of these two sets in order
11:05:37to train the network. Okay, after that
11:05:39it'll evaluate the efficiency of the
11:05:41model by using the cross-validation
11:05:43techniques. Cross-validation is the best
11:05:46improvement method. Okay, it's the best
11:05:48way to check if your model is optimal or
11:05:51not. Then we have random search. Now,
11:05:54this will randomly select samples, and
11:05:56it will evaluate sets for a particular
11:05:58probability distribution. Now, in random
11:06:01search there is no fixed number of hyper
11:06:03parameters that it's going to evaluate.
11:06:05So, it'll randomly select a set of hyper
11:06:08parameters. Okay, for example, now
11:06:10instead of checking your entire sample
11:06:13or your entire Let's say that you have
11:06:1510,000 samples. Instead of checking all
11:06:17of these samples, it'll randomly select
11:06:19100 parameters that can be checked.
11:06:21Okay, and then it'll use this to build
11:06:23the model.
11:06:24After that, we have Bayesian
11:06:25optimization.
11:06:27Okay, now Bayesian optimization
11:06:29basically uses something known as a
11:06:31Gaussian process. Basically, the
11:06:33Gaussian process will help in model
11:06:35tuning. Okay, model tuning or you can
11:06:37also say parameter tuning. So, parameter
11:06:39tuning will help you tweak the
11:06:41parameters a little bit in order to
11:06:43improve the efficiency of the model.
11:06:45Okay, so Bayesian optimization basically
11:06:47makes use of the Gaussian process, which
11:06:50will provide model tuning to your
11:06:51algorithm and thus improve the
11:06:53efficiency. Now guys, the one of the
11:06:55most important ways to improve the
11:06:57efficiency of a model is by
11:06:59hyperparameter optimization. Okay, if
11:07:01you're tuning your hyper parameters and
11:07:03if you're trying to check in which way
11:07:05these hyper parameters will give you the
11:07:07most accurate outcome, that's when your
11:07:10result will be very good. Okay, so
11:07:12that's the best way to improve the
11:07:13efficiency of the model. All right, now
11:07:15let's look at our next question. The
11:07:18next question is how does data
11:07:20overfitting occur and how can it be
11:07:22fixed? Now guys, this is a very common
11:07:24question in a machine learning or in an
11:07:27artificial intelligence interview. Okay,
11:07:29people expect you to understand what
11:07:31data overfitting is and how you can fix
11:07:34these problems. Okay, because data
11:07:36overfitting occurs pretty often,
11:07:37especially if you're using decision
11:07:40trees or if you're using random forest.
11:07:43Okay, random forest actually reduces
11:07:44overfitting, but sometimes with these
11:07:46complex models you can get data
11:07:48overfitting. Now, to answer this
11:07:49question, first of all, let's understand
11:07:51what overfitting really is. So,
11:07:54overfitting occurs when a machine
11:07:56learning algorithm captures the noise of
11:07:58the data. Okay, this causes an algorithm
11:08:01to show low bias, but high variance in
11:08:03the outcome. Now, what overfitting
11:08:05really means is you have trained your
11:08:08model way too many times on the training
11:08:10data. Okay, so basically the model has
11:08:13memorized the training data. It has
11:08:15memorized the noise in the training
11:08:17data. Okay, so if you feed new data to
11:08:20the model during the testing stage, it
11:08:22will not be able to recognize the noise
11:08:24or it will not be able to recognize any
11:08:26sort of correlation in that data. Okay,
11:08:28that's why it won't be able to get a
11:08:30proper outcome. Okay, that's when
11:08:32overfitting happens. You have trained
11:08:34the model way too much with the training
11:08:36data and this has resulted in inaccurate
11:08:39outcome during the testing phase. Okay,
11:08:42that's what overfitting is about. Now,
11:08:44how do you avoid overfitting? First of
11:08:46all, is cross-validation. Now, before
11:08:49this also I mentioned that
11:08:50cross-validation is the best way to
11:08:52obtain a more optimal solution. Now, the
11:08:55general idea behind cross-validation is
11:08:58to split the training data in order to
11:09:00generate multiple mini train test
11:09:03splits. Okay, these splits can be used
11:09:05to tune your model. Okay, so you're
11:09:07basically splitting the training data in
11:09:09such a way that, you know, the model
11:09:11does not just use the entire training
11:09:13data and memorize it. Instead, it's
11:09:15going to check the different sets in the
11:09:17training data and the different sets in
11:09:18the testing data and learn from it.
11:09:20Okay, so cross validation is one of the
11:09:22best ways to prevent overfitting.
11:09:24Another method to prevent overfitting is
11:09:27by training the model with more data.
11:09:30So, feeding more data to the machine
11:09:31learning model will help in better
11:09:33analysis and classification. However,
11:09:35this method is not always going to work,
11:09:38but yeah, this is also one of the ways
11:09:40to prevent overfitting. Okay, next we
11:09:42have removing features. Now, many times
11:09:45the data set contains irrelevant
11:09:47features or predictor variables, which
11:09:49are not needed for analysis. Such
11:09:51features will only increase the
11:09:53complexity of the model. Therefore, it
11:09:55lead to possibilities of data
11:09:57overfitting. Okay, so if you have
11:09:59irrelevant data, like for example, if
11:10:02you're trying to understand the weight
11:10:04of a person depending on its height, and
11:10:06you have another variable, let's say,
11:10:08you have a variable like the name of the
11:10:10person. Okay, now the name of the person
11:10:12is not relevant in understanding the
11:10:14height of an individual. So, if you have
11:10:17irrelevant predictor variables, then it
11:10:20will just increase the complexity of the
11:10:21model because you have an extra
11:10:23irrelevant variable. All right, this
11:10:25will only increase the complexity of the
11:10:27model. It will not help the model in any
11:10:29way. So, make sure you remove irrelevant
11:10:31features or you remove redundant
11:10:33features. Okay, the next method is early
11:10:36stopping. Now, a machine learning model
11:10:38is trained iteratively. This will allow
11:10:41us to check how well each iteration of
11:10:43the model performs. But, after a certain
11:10:46number of iterations, the model's
11:10:48performance starts to saturate. Further
11:10:51training will only result in
11:10:52overfitting. Okay, so like I mentioned,
11:10:55if you train the model with the same
11:10:57data and you make the model memorize the
11:11:00data, then it'll just saturate. It won't
11:11:02be able to predict any outcomes after a
11:11:04point. What you have to do is you have
11:11:06to understand where you need to stop
11:11:09training the model. So, this can be
11:11:11achieved by using a mechanism known as
11:11:13early stopping. So, at this point you
11:11:15know that you have to stop training the
11:11:16model because this might result in
11:11:18overfitting. Now, regularization is one
11:11:21of the most common ways to prevent
11:11:22overfitting. Regularization can be done
11:11:25in n number of ways. Okay, the method
11:11:27will always depend on the type of
11:11:29learner you're implementing. For
11:11:30example, pruning is performed on
11:11:32decision trees. Now, pruning is a type
11:11:35of regularization. Similarly, the
11:11:37dropout technique can be used on neural
11:11:39networks. And also, there are other
11:11:40methods like parameter tuning which can
11:11:42help to solve overfitting.
11:11:44The next way to prevent overfitting is
11:11:46by using ensemble models. Now, ensemble
11:11:49learning is a technique that is used to
11:11:51create multiple machine learning models
11:11:53which are then combined to produce more
11:11:56accurate results. So, basically if you
11:11:58have one problem statement in machine
11:12:00learning, you're going to use like five
11:12:03to 10 different models and then you're
11:12:05going to calculate the accuracy
11:12:07depending on the average of the result
11:12:09from each of these models. By this way,
11:12:11you will reduce overfitting. Now,
11:12:13ensemble models is one of the best ways
11:12:16to prevent overfitting. An example is
11:12:18the random forest. Random forest uses
11:12:21ensemble of decision trees to make more
11:12:23accurate predictions and to avoid
11:12:25overfitting. So, basically random forest
11:12:28is a set of decision trees. So, here
11:12:30you're going to train the model by using
11:12:32a set of decision trees and this way
11:12:34you'll have different data sets and on
11:12:36each of these data sets you'll have a
11:12:38different decision tree model. Okay,
11:12:40this will reduce overfitting to a very
11:12:42large extent. That's why in most of the
11:12:44cases when you see a decision tree
11:12:46having overfitting issues, you'll be
11:12:48asked to use random forest. So guys,
11:12:51those were the different ways to prevent
11:12:52overfitting. Now the next question is
11:12:55mention a technique that helps to avoid
11:12:57overfitting in a neural network. Now the
11:13:00most famous method to prevent
11:13:02overfitting in neural networks is
11:13:04dropout technique. Okay, now dropout is
11:13:06a type of regularization technique which
11:13:08is used to avoid overfitting in a neural
11:13:10network. So here what you do is you
11:13:12randomly select neurons and you drop
11:13:15them during the training phase. Right?
11:13:17So the dropout value also has to be
11:13:19chosen very carefully because a higher
11:13:21dropout value will result in under
11:13:23learning by the network. So if you're
11:13:25dropping out too many predictor
11:13:26variables or if you're dropping out too
11:13:28many neurons in a neural network, then
11:13:30the model will not learn enough. Okay,
11:13:33because there's not enough predictor
11:13:34variables or not enough neurons. But if
11:13:37you have too much of a low rate for a
11:13:39dropout value, then this might have a
11:13:41very minimal effect. So make sure your
11:13:43dropout value is very optimal depending
11:13:45on the problem you're trying to solve.
11:13:47Okay, so dropout is the technique which
11:13:49is used to avoid overfitting in a neural
11:13:51network. Next question is what is the
11:13:54purpose of deep learning framework such
11:13:56as Keras, TensorFlow, and PyTorch? So
11:13:59Keras is basically an open-source neural
11:14:01network library which is written in
11:14:03Python. So basically it is designed to
11:14:06enable fast experimentation with deep
11:14:08neural networks. Now TensorFlow is
11:14:10another open-source software library for
11:14:12data flow programming. TensorFlow is
11:14:14mainly used in machine learning
11:14:16applications. Similarly, PyTorch is
11:14:18again an open-source machine learning
11:14:20library for Python. Its applications are
11:14:23mainly in the field of natural language
11:14:24processing. Now I'd say that these three
11:14:27deep learning frameworks are the most
11:14:29important when it comes to machine
11:14:30learning and deep learning because they
11:14:32have a varied set of functions in them
11:14:35which help in building a better machine
11:14:36learning model or a better deep learning
11:14:39network. Now let's look at question
11:14:41number 24, which is differentiate
11:14:43between NLP and text mining. So guys,
11:14:46NLP stands for natural language
11:14:47processing for those of you who don't
11:14:49know. Now first of all, let me clear out
11:14:51a confusion between text mining and
11:14:53natural language processing. A lot of
11:14:55people tend to think that text mining
11:14:57and NLP are the same thing, but text
11:14:59mining is the broader field and NLP is
11:15:02basically an application of text mining
11:15:04or it's basically a technique used in
11:15:06text mining. So the aim of text mining
11:15:09is to extract useful insights from
11:15:10structured and unstructured text.
11:15:12Whereas the aim of NLP is to understand
11:15:15what is conveyed in these texts. Now
11:15:17text mining can be done using text
11:15:19processing languages like Perl and NLP
11:15:22can be achieved using advanced machine
11:15:24learning models such as deep neural
11:15:25networks. Now the outcome for text
11:15:28mining is you'll calculate the frequency
11:15:30of words, you'll understand the patterns
11:15:32between different words, you'll
11:15:33understand the correlations between two
11:15:35different words and you'll see how these
11:15:37two words occur together more frequently
11:15:39and why they occur together more
11:15:41frequently. So text mining basically
11:15:43will give you a more understanding about
11:15:45the words that are used in a document.
11:15:48Whereas in NLP, you'll understand the
11:15:50grammar behind the text. You'll
11:15:52understand in more depth about the
11:15:54language that is used in the document or
11:15:57in whatever you're trying to analyze. So
11:15:59that is the difference between NLP and
11:16:01text mining. NLP is a little more
11:16:03advanced field because you use deep
11:16:05neural networks to perform this. Text
11:16:07mining on the other hand makes use of
11:16:09NLP. Next question is what are the
11:16:12different components of NLP? Now there
11:16:14are two components of natural language
11:16:16processing, which is natural language
11:16:18understanding and natural language
11:16:20generation. In natural language
11:16:22understanding, you'll basically map your
11:16:24input to some useful representation.
11:16:26This means that you'll try to understand
11:16:28the correlations in your language and
11:16:30it'll also include analyzing different
11:16:32aspects of the language. All right, so
11:16:34this is majorly about understanding your
11:16:36text. When it comes to natural language
11:16:38generation, here you'll understand how
11:16:40to generate text by having a brief plan
11:16:43about the text. You'll have sentence
11:16:45planning and you'll have text
11:16:46realization. Now, natural language
11:16:48generation will basically break down
11:16:50sentences or will break down text in
11:16:52order to understand it better. Okay,
11:16:54that's what natural language generation
11:16:56is. Natural language understanding is
11:16:58more about analyzing your language or
11:17:00analyzing the text that you have at hand
11:17:02and predicting some useful outcome out
11:17:04of it. Generation is more focused on the
11:17:07planning aspect of your text. So, these
11:17:10are the different components of natural
11:17:11language processing. Now, let's look at
11:17:13what is stemming and lemmatization in
11:17:16natural language processing. Now, what
11:17:18is stemming? It is an algorithm which
11:17:20works by cutting off the end or the
11:17:22beginning of the word and only taking
11:17:25into account a list of common prefixes
11:17:27and suffixes that can be found in
11:17:29inflicted words. Now, for example, on
11:17:32the screen you can see that there is a
11:17:34detections, detected, detection, and
11:17:36detecting.
11:17:38Now, if you apply stemming on these four
11:17:40words, it will lead to detect. Okay,
11:17:43because at the end of the day,
11:17:44detections, detected, detection, and
11:17:46detecting is the same thing as detect.
11:17:48So, stemming will help you remove all of
11:17:50these unwanted prefixes and suffixes.
11:17:53This way you can analyze the importance
11:17:55of the word. All right, you don't have
11:17:57to have extra suffix or prefix before
11:17:59the word. Now, sometimes during
11:18:01stemming, cutting off the ends of the
11:18:03words will form an inaccurate result.
11:18:05Okay, that's why we have lemmatization.
11:18:08In lemmatization, the most important
11:18:10thing is the morphological analysis of
11:18:12the word. Okay, so here, in order to
11:18:14perform lemmatization, you have to have
11:18:17a detailed dictionaries which the
11:18:18algorithm can look through and it can
11:18:21form back to its lemma.
11:18:22So, the main difference between stemming
11:18:24and lemmatization is that stemming will
11:18:26just crop the prefix and the suffix,
11:18:28whereas lemmatization will try to
11:18:30understand the word in a grammatical way
11:18:33and give you an actual word as the
11:18:34output. Next is to explain the fuzzy
11:18:37logic architecture. All right, so the
11:18:39fuzzy logic architecture looks like what
11:18:42is shown on the screen. Okay, so
11:18:44basically the input is fed into
11:18:46something known as the fuzzifier. Okay,
11:18:48the fuzzifier or the fuzzification
11:18:50module will transform the system's input
11:18:53into a number of fuzzy sets. Okay, after
11:18:55that it's fed to the controller. Now,
11:18:57the controller will have knowledge base
11:18:59and the inference engine. Knowledge base
11:19:01is basically a set of rules or you can
11:19:03say it's an algorithm which is provided
11:19:06by experts. Inference engine, like the
11:19:08name suggests, will basically infer
11:19:10meaning out of these rules. Okay, so
11:19:12once you've applied the rules to your
11:19:14input, you'll have to draw some useful
11:19:16insights or you'll have to infer these
11:19:18inputs. Okay, for that you use the
11:19:20inference engine. After that, whatever
11:19:23inferences and analysis you've formed
11:19:25from your inference engine is passed on
11:19:26to the defuzzification module. Now, the
11:19:29defuzzification will just give you a
11:19:31crisp output. All right, it'll give you
11:19:33a clear and cut output. That is the
11:19:36whole fuzzy logic architecture.
11:19:38Now, let's understand the components of
11:19:40an expert system. Now, there are three
11:19:42important components in an expert
11:19:44system, which is knowledge base,
11:19:46inference engine, and user interface.
11:19:48Now, like I mentioned in fuzzy logic,
11:19:50the knowledge base and inference engine
11:19:52will play the same part. The user
11:19:54interface is basically to provide
11:19:56interaction between the users of the
11:19:58expert system and the expert system.
11:20:01Okay, the expert system is basically a
11:20:03program that helps in decision-making
11:20:05process. Okay, so here the knowledge
11:20:07base will contain some high-quality
11:20:09knowledge or it contain rules and
11:20:11algorithms. The inference engine will
11:20:13acquire all the knowledge that is needed
11:20:16to solve the problem. And the user
11:20:18interface is just for the users to
11:20:19interact with the expert system. Okay,
11:20:21this is the whole expert system
11:20:23component. Now, obviously this This a
11:20:25little more complex than this, but uh
11:20:27stick to how this works. All right, I'm
11:20:29just going to tell you the working of
11:20:30expert systems and fuzzy logic. If I
11:20:33start to explain each and everything,
11:20:35it's going to take a lot of time. All
11:20:36right, so let's move on to our next
11:20:38question, which is how is computer
11:20:40vision and AI related? Now, computer
11:20:43vision is a field of artificial
11:20:45intelligence that is used to obtain
11:20:47information from images or
11:20:49multi-dimensional data. Now, computer
11:20:51vision is basically the concept behind
11:20:53the self-driving cars that you see these
11:20:55days. All right, computer vision
11:20:57involves a lot of image processing. So,
11:20:59machine learning algorithms like K-means
11:21:01can be used in image segmentation.
11:21:03Support vector machines can be used for
11:21:05image classification. Okay, that's how
11:21:07computer vision and AI are related.
11:21:09Because most of the things that happen
11:21:11in computer vision like image processing
11:21:13and segmentation make use of machine
11:21:15learning algorithms like K-means and
11:21:17support vector machines. So, to sum it
11:21:19up, computer vision makes use of
11:21:21artificial intelligence technologies to
11:21:24solve complex problems such as object
11:21:26detection, image processing, and so on.
11:21:28That is the relationship between
11:21:30computer vision and AI. Now, question
11:21:32number 30 is which is better for image
11:21:35classification? Is it supervised or
11:21:38unsupervised classification? So, guys,
11:21:40earlier in the session we discussed what
11:21:42supervised learning is and what
11:21:43unsupervised learning is. In supervised
11:21:46learning, the images are interpreted
11:21:48manually by the machine learning expert
11:21:50to create feature classes. Now, what
11:21:52this means is you're manually going to
11:21:54feed a labeled set of data to the
11:21:56supervised learning model. All right,
11:21:58that's how supervised learning works.
11:22:00You're manually going to feed a set of
11:22:02images which are labeled to the
11:22:04classifier. In unsupervised learning,
11:22:06the machine learning software creates
11:22:08feature classes based on image pixel
11:22:10values. So, basically in unsupervised
11:22:12classification, the model itself has to
11:22:15figure out what to do and what not to
11:22:17do. Okay, so it'll create a own feature
11:22:19class based on some values such as image
11:22:22pixels or it can also use the image
11:22:24color or it can use intensity factors in
11:22:27order to classify. So, if you ask me it
11:22:29is better to opt for supervised
11:22:31classification because you're manually
11:22:33inputting images with a lot more
11:22:35information. Okay, whereas in
11:22:37unsupervised learning you're totally
11:22:38letting the model perform everything.
11:22:40Okay, so in image classification, I
11:22:42think it's better to go for supervised
11:22:44learning. Now, let's look at question
11:22:46number 31. The next question is finite
11:22:50difference filters in image processing
11:22:51are very susceptible to noise. To cope
11:22:54up with this, which method can you use
11:22:56so that there would be minimal
11:22:58distortions by noise? Now, the noise in
11:23:01an image can be due to high intensity or
11:23:03high contrast. Okay, so if you increase
11:23:06the contrast and increase the intensity
11:23:08of an image, you won't be able to
11:23:10understand each pixel. Okay, so each
11:23:12pixel will have a value associated to it
11:23:15and if the intensity and the contrast of
11:23:17that pixel is a little too much, it'll
11:23:19be hard for us to understand the image
11:23:21properly. It'll be hard to perform image
11:23:24analysis because we don't have a clear
11:23:26image. Contrast and intensity will just
11:23:28cause noise in an image. So, the best
11:23:31method to remove this is image
11:23:32smoothing. Okay, it is used for reducing
11:23:35noise by forcing pixels to be more like
11:23:38their neighbors. Okay, this way you'll
11:23:40have a faded image or you'll have a more
11:23:42equalized image. Now, the next question
11:23:45is how is game theory and AI related? So
11:23:48guys, AI is actually applied in a vast
11:23:51number of fields. Okay, so a lot of
11:23:53fields from computer vision to game
11:23:55theory to machine learning, AI is always
11:23:58a concept behind these fields. Most of
11:24:01the game examples that we see make use
11:24:03of reinforcement learning or deep neural
11:24:05networks. Now, deep neural networks and
11:24:07reinforcement learning are very closely
11:24:09related to AI because they are branches
11:24:11of machine learning. So, machine
11:24:13learning is majorly involved in game
11:24:15theory. An example of this is in Dota 2
11:24:18also they make use of machine learning.
11:24:20So, game theory is just a very logical
11:24:23approach to solving a problem. And
11:24:25machine learning is the best way to
11:24:27implement game theory. Now, question
11:24:29number three is what is the minimax
11:24:31algorithm? Explain the terminologies
11:24:33involved in the problem.
11:24:35Now guys, minimax is one of the main
11:24:37algorithms which is used in game theory.
11:24:39All right, it is used to choose an
11:24:41optimal move for a player assuming that
11:24:43the other player is also playing
11:24:45optimally. Meaning that both of these
11:24:47players are playing in order to win and
11:24:50you're going to use the minimax
11:24:51algorithm on one of these players so
11:24:53that they choose the optimal move. In
11:24:56order to understand the minimax
11:24:57algorithm, you need to know what are the
11:24:59components in a game. Okay, there's
11:25:01something known as game tree. It is
11:25:03basically a tree structure which
11:25:05contains all the possible moves in a
11:25:07game. If it's up, down, right, left, any
11:25:09strategy, everything is mentioned in the
11:25:11game tree. Now, initial state is
11:25:13obviously the initial position of the
11:25:15player on the board. All right, the
11:25:17successor function it defines all the
11:25:19possible moves that a player can make.
11:25:22We'll understand this in the next
11:25:23question itself, so don't worry if you
11:25:25haven't understood this properly.
11:25:27Terminal state is obviously the end of
11:25:29the game. It's basically the state which
11:25:31will lead to the end game or it will
11:25:33lead to your destination. Utility
11:25:35function is a numerical value for the
11:25:37output of the game. So guys, these were
11:25:39the terminologies and this is what the
11:25:41minimax algorithm is. It is basically a
11:25:44game theory algorithm which helps a
11:25:46player choose the best optimal policy in
11:25:49order to win a game. I'll explain this
11:25:51in more depth in the upcoming slides.
11:25:54So, let's move on. Now, the next couple
11:25:56of questions are going to be
11:25:57scenario-based questions. Now, such
11:25:59questions are very important in an
11:26:01interview because this is where the
11:26:03interviewer will understand how well you
11:26:05know the concepts. So, the first
11:26:07question is show the working of the
11:26:09minimax algorithm using the tic-tac-toe
11:26:11game. Now, one of the major applications
11:26:14of the minimax algorithm is the
11:26:16tic-tac-toe game. Okay, you can
11:26:18understand and analyze all the possible
11:26:20outcomes of the tic-tac-toe game by
11:26:22using the minimax algorithm. Let's see
11:26:24how this happens. Now, first of all, in
11:26:27a minimax algorithm or in a game, there
11:26:29are two players involved. Okay, the max
11:26:32is the player that tries to get the
11:26:33highest possible score, and min is the
11:26:36player that tries to get the lowest
11:26:37possible score. So, this algorithm is
11:26:40designed in such a way that assuming
11:26:42that there going to be two players, and
11:26:44obviously one player is going to win the
11:26:45game, and that is the max player, and
11:26:48min is the player which loses the game
11:26:50and has the lowest possible score. Now,
11:26:52the first step in the minimax algorithm
11:26:54is to generate the entire game tree.
11:26:57Okay, the game tree is all the possible
11:26:59outcomes that can happen in tic-tac-toe.
11:27:01Okay, in the figure you can see that
11:27:02first X is aligned in the first box,
11:27:04then in the second box, third box, and
11:27:06so on. All the possible actions that you
11:27:09can take in a tic-tac-toe game are put
11:27:11in this game tree. And then, step number
11:27:14two is to apply the utility function to
11:27:16get the utility values from all the
11:27:18terminal states. Getting utility value
11:27:21is important because this is how you'll
11:27:23understand your outcome. Okay, you'll
11:27:24understand if you're going to win or
11:27:25lose. Now, in the terminal states,
11:27:28whatever numbers you see over here,
11:27:30these are the utility values. Now, step
11:27:32three is determine the utilities of the
11:27:34higher nodes with the help of utilities
11:27:36of the terminal nodes. Now, in this
11:27:39diagram, you can see that in the
11:27:40terminal nodes, we have the utility
11:27:42values. The step three is to get utility
11:27:45values in the higher stages, which is
11:27:48the min stage. All right, these two
11:27:50circles, you need to fill in the utility
11:27:51values by using the utility values which
11:27:54are in the terminal state.
11:27:55Now, how do you calculate the utility
11:27:57value? Let's start by calculating the
11:28:00utility value of the left node. Okay,
11:28:02this red color node, we'll start by
11:28:05calculating this.
11:28:06Now, you calculate that by finding the
11:28:08minimum of the three nodes that it's
11:28:11leading to. Now, this red node is
11:28:13leading to three, five, and 10. And the
11:28:15minimum out of three, five, 10 is three.
11:28:17So, the utility value for this red node
11:28:20is going to be three. Okay, similarly
11:28:22for this green node, it's going to be
11:28:23two because the minimum value between
11:28:25two and two is still two. Now, step four
11:28:27is to fill in these utility values that
11:28:29you've calculated. So, now we have a
11:28:32minimax algorithm which has all the
11:28:34utility values filled in. Now, the only
11:28:36utility value which isn't filled is the
11:28:38one with max. Okay, the one on the root
11:28:41node. Here, we haven't filled the
11:28:43utility value. Again, to fill this
11:28:45value, you're going to check the nodes
11:28:46which are directly connected to it,
11:28:48which is three and two. You'll find the
11:28:50maximum between these two because this
11:28:52is the max function. All right, so here
11:28:54you'll get a value of three. So, that's
11:28:57why the best opening move for max is the
11:28:59left node. Okay, you can make use of the
11:29:02left node in order to win the game. This
11:29:04is the first step that the max player
11:29:06has to take in order to get to the path
11:29:08of winning the game. So guys, by doing
11:29:11this for each and every step, you can
11:29:13win the game. Okay, so you'll have to
11:29:15calculate the utility value at the
11:29:17terminal nodes. You'll have to move up
11:29:19to the other hierarchical nodes above
11:29:21it, calculate the utility values there
11:29:23until you reach the root node. Okay,
11:29:25once you reach the root node, you'll get
11:29:26a utility value and that utility value
11:29:29will be connected to some move or some
11:29:31node. You'll have to take that node or
11:29:34you'll have to take that move in the
11:29:36game in order to win the game. So, this
11:29:38way you'll have to calculate the utility
11:29:40value for each and every move that the
11:29:42player makes so that the player will win
11:29:44the game.
11:29:45So guys, minimax algorithm is quite easy
11:29:47and it's very understandable. All you
11:29:49need to know is a little bit of math in
11:29:51order to solve this problem. Question
11:29:52number 35 is which method is used for
11:29:56optimizing a minimax based game? Now,
11:29:58this is not a scenario-based question,
11:30:00but this question is usually asked if an
11:30:02interviewer asks you about a minimax
11:30:05game.
11:30:05Now, the best way to optimize a minimax
11:30:08game is by using something known as
11:30:10alpha-beta pruning. Now, the main thing
11:30:12about alpha-beta pruning is that it'll
11:30:14remove all the nodes that are not
11:30:16affecting the final decision. It's just
11:30:18a faster way to reach your outcome.
11:30:21That's what alpha-beta pruning is all
11:30:23about. So, let's look at an example to
11:30:26understand this. Okay, let's say there
11:30:27was another node over here. Okay, here
11:30:29you can see that this is going down to a
11:30:31terminal state with utility value two.
11:30:34Okay, now you don't know the value of
11:30:36the other two nodes, but if you use
11:30:38minimax to calculate the utility of the
11:30:41other two nodes, you'll get a value of
11:30:43three. So, in this example again, we'll
11:30:45start at the terminal nodes. So, three,
11:30:47five, 10 are the utility values here.
11:30:50So, this will give us a value of three
11:30:52because we're calculating the minimum
11:30:53over here. Now, here you have two and
11:30:56you have two unknown values. You have A
11:30:58or B. Okay, I've named them as A and B.
11:31:01Okay, let's leave this for now. Let's go
11:31:02to the next node. Okay, here the
11:31:04possibilities are two, seven, and three.
11:31:07So, the minimum between two, seven,
11:31:08three is two.
11:31:10Okay, so here there's going to be three.
11:31:11There's going to be a value, let's say
11:31:13C, and here there's going to be a value,
11:31:15let's say two. Now, we know that the
11:31:18maximum between three, C, and something
11:31:20else will be three. Okay, that's because
11:31:23two is the minimum value over here, and
11:31:25the maximum will obviously be three. So,
11:31:27the hint here is in the two AB node. We
11:31:30know that the value or the utility value
11:31:32will obviously be equal to two or it'll
11:31:35be less than two because you're
11:31:36calculating the minimum in this step.
11:31:38Now, if you calculate the max out of
11:31:40these three values, we'll obviously get
11:31:42the answer as three.
11:31:44So, this way this entire node itself is
11:31:46removed because you don't need it to get
11:31:48to the final answer. Okay, that's what
11:31:50alpha-beta pruning is all about. It'll
11:31:52identify the nodes which are not going
11:31:54to affect the final decisions, and it'll
11:31:56just remove those nodes. So guys, this
11:31:58is how the optimization for a minimax
11:32:01game is done. It's done using the
11:32:03alpha-beta pruning.
11:32:04The next question is which algorithm
11:32:06does Facebook use for face verification?
11:32:09Now guys, even though this might seem
11:32:11like a general knowledge question, this
11:32:14is actually a very important sort of
11:32:16question in artificial intelligence.
11:32:18Okay, even if you don't know the answer
11:32:20to this, you should have an idea of how
11:32:22the algorithm might work. Okay, that's
11:32:24exactly what the interviewer wants to
11:32:26know. He wants to know whether you know
11:32:28how the algorithm works step-by-step.
11:32:30You might not know the final answer or
11:32:32you might not know the exact algorithm
11:32:34which Facebook uses because obviously
11:32:35Facebook uses more than one algorithm to
11:32:38achieve this, but you must know the
11:32:40steps in which the face verification
11:32:42works. Okay, that's the main goal behind
11:32:44this question. Now anyway, the algorithm
11:32:47used by Facebook is the deep face. Okay,
11:32:49deep face makes use of a lot of neural
11:32:51networks and a lot of algorithms. Okay,
11:32:53so it works on artificial intelligence
11:32:56techniques, like I mentioned earlier.
11:32:58Now how would a face verification work?
11:33:00How do you think it works? Now it starts
11:33:02by an input. So the idea here is you
11:33:05have to scan a huge number of photos and
11:33:07you'll have to feed it to the algorithm.
11:33:10Okay, now these photos can have a lot of
11:33:12disturbance, a lot of distortions and it
11:33:14can have different angles or anything
11:33:16like that. Okay, you have to feed any
11:33:18sort of photos that are possible. Okay,
11:33:20even they are complex to understand, but
11:33:22you have to still feed the model with
11:33:24all the possible photos that you can
11:33:25get. Now the next step is the main
11:33:27process. Here there are a few important
11:33:30things which is detect, align, represent
11:33:32and classify. Detect is basically you'll
11:33:34detect facial features. All right,
11:33:37you'll try to understand the distance
11:33:38between the eyes and the nose of a
11:33:40person, the way the lips is aligned or
11:33:43anything like that. That's what aligning
11:33:45is about. You'll align and compare the
11:33:46various features in the face in order to
11:33:49understand the facial features. You'll
11:33:51represent the key patterns by using some
11:33:533D graphs or 3D models. Okay, it's very
11:33:55important to visualize whatever you get
11:33:58because visualization will help you
11:34:00understand the correlation. It'll help
11:34:02you understand that okay, the eyes are
11:34:03at this distance, the nose is at this
11:34:05distance, and so on. Finally, you'll
11:34:07classify the images based on the
11:34:09similarity. All right, that's how the
11:34:11output comes out. And basically, the
11:34:13output is you need to detect whether two
11:34:15images represent the same person or not.
11:34:18Okay, so by studying the facial features
11:34:20and by using image processing and by
11:34:22using computer vision, Facebook's
11:34:24achieves face verification. You need to
11:34:26know the basic concept behind face
11:34:28verification. You need to know that it
11:34:30starts with image collection or data
11:34:32acquisition. After that, you're going to
11:34:34perform image processing or
11:34:36pre-processing. All right, and this
11:34:38might involve performing conversions
11:34:40from RGB to any other state like YCbCr.
11:34:44Okay, I'm not going to go in depth of
11:34:45this because the video will get to about
11:34:472-3 hours. So, there are a lot of ways
11:34:49in which you can convert an image and
11:34:51you know, you can understand the image
11:34:52more properly. Also, an important thing
11:34:54in image processing is it's not done
11:34:57just based on the image. All right.
11:34:59You're going to take the image, you're
11:35:00going to form a matrix, and you're going
11:35:02to have pixel values in these matrix.
11:35:04So, it is a very in-depth approach. All
11:35:06right, it's not a very simple approach.
11:35:08When I'm speaking about it, it might
11:35:10seem simple, but image analysis is very
11:35:13in-depth. After image analysis, you can
11:35:15perform image segmentation. All right,
11:35:17image segmentation is basically dividing
11:35:19the image into different segments and
11:35:21studying each image segment separately.
11:35:24Then after that, you can do feature
11:35:25extraction. Here, you'll try to
11:35:27understand the features and how they are
11:35:29related to each other. Finally, you'll
11:35:31classify the images and see whether two
11:35:34images represent the same person or not.
11:35:37So, the main idea behind the Facebook
11:35:39algorithm is image processing, neural
11:35:41networks, machine learning, and computer
11:35:43vision. All right, and all of this comes
11:35:45down to artificial intelligence. Next,
11:35:48we have explain the logic behind
11:35:50targeted marketing and how can machine
11:35:52learning help with this? Now, target
11:35:54marketing is something that we see very
11:35:56often. All right, let's say that you
11:35:58were looking for some shoe on Amazon. In
11:36:01a day or two, you just open up YouTube
11:36:03and Facebook. You'll see that you'll get
11:36:05ads of shoes from Amazon. Okay, this is
11:36:08targeted marketing. So, basically Amazon
11:36:10knows that we've been looking for a
11:36:12particular type of shoe, so it's going
11:36:14to target you with that particular ad.
11:36:16Okay, this is what targeted marketing is
11:36:18in short. Targeted marketing can be done
11:36:20in different ways. For example, it can
11:36:23be done depending on your geography or
11:36:25it can be done depending on your social
11:36:27economic profile. Okay, let's say that
11:36:30Amazon has details about your age, it
11:36:33has details about what sport you like to
11:36:35play. Let's say that you've been
11:36:36browsing through a lot of sports. Okay,
11:36:39you've been browsing through a lot of
11:36:40sport equipments or something like that.
11:36:43Amazon will know that you're interested
11:36:44in this by using machine learning, of
11:36:46course, and it will send you ads based
11:36:49on what you're interested in. Okay, this
11:36:51is what target marketing really is.
11:36:53Now, how does machine learning come into
11:36:55target marketing? Okay, so there's
11:36:57something known as text analytics
11:36:59systems. Now, the applications for text
11:37:01analytics ranges from search
11:37:03applications, text classification, named
11:37:06entity recognition, or pattern search.
11:37:09Okay, so it's basically a way to
11:37:10understand what you're looking for.
11:37:12Okay, they'll try to understand your
11:37:14search history and they'll try to target
11:37:16you by using your interests. Clustering
11:37:19is another way of targeted marketing.
11:37:21All right, you'll cluster customers who
11:37:23have similar interests and you'll send
11:37:25them similar ads or you'll send them
11:37:26similar offers. Classification is
11:37:28another method used for targeted
11:37:30marketing. Now, here you'll make use of
11:37:33algorithms like decision trees and
11:37:34neural networks. Now, recommender
11:37:36systems is what I spoke about earlier.
11:37:38When it comes to Amazon, they recommend
11:37:40items to you based on your interest or
11:37:43based on people who have similar
11:37:44interests like you. Market basket
11:37:47analysis is another method that machine
11:37:49learning uses for marketing. Okay, here
11:37:51basically you'll understand a
11:37:53combination of products that are
11:37:54frequently bought. Okay, by
11:37:56understanding what two products are
11:37:58frequently bought, you can give some
11:37:59offers or you can give some discounts on
11:38:01those products so that people buy more
11:38:03and more. Okay, that's how market basket
11:38:05analysis also works. So guys, this is
11:38:07what targeted marketing is.
11:38:10Now next is how can AI be used to detect
11:38:13fraud? AI is used in a lot of ways in
11:38:16credit card fraud detection. It's used
11:38:18in detecting anomalies and all of that.
11:38:20It basically makes use of machine
11:38:22learning algorithms to do this. Now
11:38:24let's try to understand how this process
11:38:26works. Okay, first it begins with data
11:38:28extraction or data collection. So at
11:38:30this stage data is either collected
11:38:32through a survey or through web
11:38:34scraping. Okay, if you're trying to
11:38:35detect credit card fraud then
11:38:37information about the customer's
11:38:39collected. All right, this includes any
11:38:41transactional or any shopping and
11:38:43personal details. Next is data cleaning.
11:38:45So at this stage the redundant data must
11:38:48be removed. Any inconsistencies or any
11:38:50missing values that you have in your
11:38:52data, it has to be removed because they
11:38:54lead to wrongful prediction. Okay, so
11:38:56therefore you have to get rid of any
11:38:58inconsistencies in this stage. Next we
11:39:01have data exploration and analysis. Now
11:39:03this is the most important step in AI.
11:39:06Okay, here you study the relationship
11:39:08between various predictable variables.
11:39:10For example, let's say that a person has
11:39:12spent an unusual sum of money on a
11:39:15particular day. Now the chances for a
11:39:17fraudulent occurrence is very high
11:39:19because usually the person is not used
11:39:21to spending this much money. So such
11:39:23patterns have to be detected and
11:39:24understood in data exploration and
11:39:26analysis. This is followed by building a
11:39:29machine learning model. Now here there
11:39:31are any machine learning algorithms that
11:39:33can be used for fraud detection or
11:39:35anomaly detection. One such example is
11:39:37logistic regression, okay, which is a
11:39:39classification algorithm and it can be
11:39:42used to classify events into two
11:39:44classes. Okay, you can use them to
11:39:46classify a person or classify an event
11:39:49as either fraudulent and non-fraudulent.
11:39:52Then comes model evaluation. Here you'll
11:39:54basically test the efficiency of the
11:39:56machine learning model. Okay, so if
11:39:58there's any room for improvement, then
11:40:00you can perform parameter tuning and you
11:40:02can improve the model. This will just
11:40:04improve the accuracy of the model. So
11:40:06guys, all of these complex problems like
11:40:09fraud detection or object detection, all
11:40:12of this is done through a process. Okay,
11:40:14and in general the process is data
11:40:16collection, data cleaning, exploration
11:40:18and analysis, building a model and model
11:40:21evaluation. Most of these complex
11:40:23problems can be solved by using this
11:40:25approach. Now let's look at our next
11:40:27question.
11:40:28Okay, a bank manager is given a data set
11:40:31containing records of thousands of
11:40:33applicants who have applied for loan.
11:40:35How can AI help the manager understand
11:40:37which loans he can approve?
11:40:39To be more specific, this problem
11:40:41statement can easily be solved by using
11:40:43the KNN algorithm. Okay, KNN is
11:40:46basically stands for K nearest neighbor.
11:40:49All right, this is a classification and
11:40:51a regression algorithm.
11:40:53So if you use a KNN algorithm, it'll
11:40:55form two classes. One is the loan is
11:40:57approved and the other is applicants
11:40:59whose loan is not been approved. So like
11:41:02I said, K nearest neighbor is a
11:41:03supervised learning algorithm that
11:41:05classifies a new data point into the
11:41:08target class depending on the features
11:41:10of its neighboring data points. All
11:41:12right, so KNN basically focuses on the
11:41:14neighbors and it understands that if a
11:41:16new data point is similar to one of its
11:41:18neighbors, then it has to classify that
11:41:20new data point into that neighbor's
11:41:22class. Now again, the methodology for
11:41:24solving this problem is same. You start
11:41:27by data collection, data cleaning,
11:41:29exploration and analysis, building a
11:41:31model and model evaluation. All right.
11:41:34So, while data collection, you can
11:41:35collect data like account balance,
11:41:37credit amount, age, occupation, loan
11:41:40records, and all of that. So, by using
11:41:42this data, you can predict whether or
11:41:43not to approve the loan of an applicant.
11:41:46Data cleaning, again, you have to remove
11:41:48any variables which will not help the
11:41:50model. Okay, any variables which will
11:41:52just increase the complexity of the
11:41:54model. Okay, so you'll remove such
11:41:55variables at this stage. In data
11:41:58exploration and analysis, you will
11:42:00understand the patterns in your data.
11:42:02Okay, let's see that a person has a
11:42:04history of unpaid loans. Okay, if any
11:42:07person or any applicant has a history of
11:42:09unpaid loans, then the chances are that
11:42:11he might not get approval on his loan
11:42:13application. Okay, this is obvious
11:42:15because the manager is going to see that
11:42:17his previous loans are still due. So,
11:42:20that's why he won't be able to approve
11:42:21the application. So, these are the kind
11:42:23of patterns that are detected in
11:42:25exploration and analysis. Now, building
11:42:28a machine learning model, you can use n
11:42:30number of models when it comes to
11:42:31predicting whether an applicant loan
11:42:34request is approved or not. Now, like I
11:42:36mentioned, one of the easy algorithms
11:42:38that you can implement is the K nearest
11:42:39algorithm. Okay, it can be used for both
11:42:42classification and regression. It
11:42:44classify the applicant's loan request
11:42:46into approved or disapproved based on
11:42:49the socio-economic profile of the
11:42:50applicant. Okay, based on variables like
11:42:53loan based on variables like the salary,
11:42:55the occupation of the applicant. Now,
11:42:58model evaluation, again, is the same
11:43:00thing. You'll basically evaluate the
11:43:02efficiency of the model. You'll try to
11:43:03improve the accuracy of the model by
11:43:05using parameter tuning or cross
11:43:07validation. So, guys, this is how a bank
11:43:10manager can understand whether a loan
11:43:12can be approved or not. Again, in this
11:43:14question, they are just trying to test
11:43:16if you know how the flow of the problem
11:43:18will go. If you know how this problem
11:43:20can be solved. You don't have to know
11:43:22the exact details, but you have to know
11:43:24how you can approach the problem.
11:43:26Okay, now let's move on and look at
11:43:28question number 40. Now the question
11:43:30here is place an agent in any one of the
11:43:32rooms and the goal is to reach outside
11:43:35the building. Can this be achieved
11:43:37through AI? If yes, explain how it can
11:43:39be done. Now in this question there is a
11:43:42diagram along with a small explanation.
11:43:44Okay, so basically there are four rooms
11:43:47in this diagram. Basically 0 1 2 3 and 4
11:43:50represent rooms and this 5 represents
11:43:54outside the building. Okay, now the goal
11:43:56is to place an agent in any one of these
11:43:59rooms in such a way that he has to reach
11:44:01room number five or he has to reach
11:44:03outside the building. Now they've also
11:44:05mentioned that a room number one and
11:44:07room number four directly lead outside
11:44:10the building. That's correct because
11:44:11room number one is directly connected to
11:44:13five and four is also directly connected
11:44:16to five. Right? Four leads outside the
11:44:18building. Now if you look at room number
11:44:21zero, if you want to go from zero to
11:44:23five, first from zero you'll have to go
11:44:25to four and then only you can go to
11:44:26five. Similarly, if you look at room
11:44:28number three, if you want to go from
11:44:30three to five, you'll have to take
11:44:32three, then you'll have to go to one and
11:44:34then only you'll have to go to five.
11:44:36These are not directly connected to
11:44:38outside the building, whereas room
11:44:39number one and four are directly
11:44:41connected to outside the building. All
11:44:43right, I hope the question is clear. Now
11:44:45as soon as you read the question, you
11:44:47must know that this is a reinforcement
11:44:49learning question. All right, it's
11:44:51pretty clear because they have mentioned
11:44:53that there is an agent which is going to
11:44:55be placed in any one of the rooms and he
11:44:57has to basically explore the environment
11:45:00and reach five, which is basically
11:45:02outside the building. So as soon as you
11:45:03read the question, the first thing that
11:45:05should come into your head is that this
11:45:06is a reinforcement learning problem. Now
11:45:09this problem can be solved by using the
11:45:11Q-learning algorithm. Now if you
11:45:13remember earlier in the session, I
11:45:15discussed what exactly Q-learning is and
11:45:17how it works. So Q-learning is basically
11:45:20a reinforcement learning algorithm which
11:45:22is used to solve reward-based problems.
11:45:24So, now let's look at how we'll solve
11:45:26the problem. First step would be to
11:45:29represent the rooms on a graph. All
11:45:31right, so each room over here you'll
11:45:33represent it as a node and each door
11:45:35will represent a link. So, if you look
11:45:37at this figure over here, this is our
11:45:39original figure that was given in the
11:45:41question and now this is the graph that
11:45:43we draw from this figure. So, we have
11:45:46node one. Let's look at how node one is
11:45:49connected to node three. So, basically
11:45:51you can go from node one to node three
11:45:53and you can go from node three to node
11:45:55one. If you look at the diagram, there
11:45:57is a direct connection from one to three
11:45:59and three to one. Okay, let's look at
11:46:01one and two. Now, there's no link
11:46:03between one and two because if you look
11:46:05over here, if you want to go from room
11:46:07number one to room number two, you
11:46:08cannot directly go. You'll have to go
11:46:11from room number one to room number
11:46:12three and only then you'll reach room
11:46:14number two. All right, that's why
11:46:15there's no link between room number one
11:46:18and node number two.
11:46:19Similarly, if you look at node one and
11:46:22four, they are directly connected to
11:46:24five. This is because room number one
11:46:26and four directly lead to this goal. Our
11:46:29goal is to reach room number five. So,
11:46:32one and four are directly connected to
11:46:34five, whereas the others are connected
11:46:36just like how they're shown in this
11:46:38figure. Okay, it's pretty
11:46:39understandable, guys. This is just
11:46:41logic. Now, let's look at the next step.
11:46:43Now, the next step is to associate a
11:46:46reward value to each door. What we're
11:46:48going to do here is we're going to build
11:46:50a reward matrix. Now, I'll tell you what
11:46:52that exactly means. So, for the doors
11:46:55that lead directly to the end goal,
11:46:57which is room number five, you'll assign
11:46:59a reward of 100 to those doors. So, if
11:47:01you're traversing from node number one
11:47:03to node number five, you'll get a reward
11:47:05of 100. Similarly, if you're traversing
11:47:08from four to five, you'll get a reward
11:47:09of 100. Five to five also you'll get a
11:47:12reward of 100 because your end goal is
11:47:14five, right? So, basically any link that
11:47:16leads directly to our end goal, for that
11:47:19link, we're going to assign a reward of
11:47:20100. Now, doors that are not directly
11:47:23connected to the target room will have a
11:47:25reward zero. This is because if you take
11:47:28room number two or if you take room
11:47:29number three, you won't reach room
11:47:31number five directly. And our goal here
11:47:34is to reach room number five. That's why
11:47:36for the other doors, we've given a
11:47:38reward of zero. For door number one and
11:47:40door number four, however, we have
11:47:42rewards of 100. Similarly, for door
11:47:44number five or room number five, also we
11:47:46have a reward of 100. So, basically,
11:47:49each action or each link will represent
11:47:52a reward. So, let's say that you're
11:47:54traversing from room number one to room
11:47:56number four. If you go from room number
11:47:59one to room number three and then three
11:48:01to four, your reward is going to remain
11:48:03zero because you're not reaching the end
11:48:05goal here. Only if you traverse from one
11:48:07to five or if you traverse from four to
11:48:09five or five to five, you'll get a
11:48:11reward of 100. All right, it's as simple
11:48:14as that. Now, let's see how the
11:48:15Q-learning algorithm works in this
11:48:18particular problem statement. Now, there
11:48:20are two main components in this
11:48:22algorithm. All right, there is state and
11:48:24there is action. Now, basically, all
11:48:26these rooms will represent the state and
11:48:28the agent's movement from one room to
11:48:30the other will represent an action. So,
11:48:33basically, 0 1 2 3 4 and 5 represent the
11:48:36state and let's say you're traversing
11:48:38from two to three. This two to three
11:48:40will basically represent an action. In
11:48:42order to make you understand how this
11:48:44works, let's say that you're traversing
11:48:46from room number two to room number
11:48:47five. Your initial state is going to be
11:48:50room number two. Your next state is
11:48:52going to be room number three. All
11:48:53right, so you're moving from two to
11:48:55three and you're getting a reward of
11:48:56zero. Remember that. Now, from state
11:48:58three, you'll either be going to state
11:49:00two or you can go to state one or state
11:49:03four. All right, if you choose four, you
11:49:06can uh directly go to five and here
11:49:08you'll get a reward of 100. If you
11:49:10choose one, again, you'll go directly to
11:49:12five. You'll get a reward of 100. But if
11:49:15you go back to two, you'll get a reward
11:49:16zero. So guys, let me tell you that the
11:49:18agent is going to explore over here,
11:49:20okay? He has no idea about the
11:49:22environment, so he's not going to go
11:49:23from two, three, one, five directly, all
11:49:26right? He's not going to know that this
11:49:28will lead to the output. He has to
11:49:30explore, he has to make mistakes, he has
11:49:32to learn, and he has to find out the
11:49:34best path to reach room number five. Now
11:49:36next is our reward matrix. So guys,
11:49:39there are two main matrices in
11:49:40Q-learning algorithm. One is a reward
11:49:42matrix, and the other is going to be the
11:49:44Q matrix or the memory matrix. All
11:49:47right, I'll be discussing the memory
11:49:48matrix in a while, but for now let's
11:49:50look at the reward matrix. So what I'm
11:49:53doing here is I'm basically putting all
11:49:55the reward values for traversing from
11:49:57one node to the other node in a matrix
11:49:59known as the reward matrix. Now the
11:50:01minus one will basically represent the
11:50:03null values. What I'm trying to say is
11:50:06there is no connection from zero to
11:50:07zero, all right? That's why I'm giving a
11:50:09value of minus one. But you might say
11:50:11that why is the reward from five to five
11:50:14hundred? Now this is because if you go
11:50:16from room number five to room number
11:50:18five, you're still reaching the end
11:50:20goal, all right? That's why there's an a
11:50:22reward of 100 over here. Let's look at
11:50:24zero {comma} four, all right? There is a
11:50:26link from zero to four, but the reward
11:50:28is going to be zero because four is not
11:50:31the end goal, all right? Four is not
11:50:33your goal room or anything, that's why
11:50:35your reward is going to be zero. Now the
11:50:37reward is 100 only for one {comma} five
11:50:39that is if you traverse from one to
11:50:41five, for four {comma} five, which is if
11:50:44you traverse from four to five, and five
11:50:46{comma} five. Now this is because
11:50:48through all these three actions you're
11:50:49going back to room number five, all
11:50:51right? Which is your end goal, that's
11:50:53why we have a reward of 100 for these
11:50:55three actions, all right? So like I
11:50:57mentioned earlier, we're going to have
11:50:59another matrix known as a Q matrix,
11:51:01which will basically represent the
11:51:03memory of what the agent has learned
11:51:05through experience. Okay, that's the
11:51:07only way the agent will actually learn
11:51:09further. If the agent forgets everything
11:51:11that he's learned, then there's no point
11:51:13because he'll have to redo everything
11:51:14from scratch and again he'll forget
11:51:16everything. That's why we have a matrix
11:51:18known as a Q matrix, which represents
11:51:20the memory of the agent. Okay, so if a
11:51:22agent has traveled from room number two
11:51:24to three, he's going to remember the
11:51:26reward. Okay, that reward is going to be
11:51:28stored in the Q matrix. And the rows of
11:51:31the Q matrix will represent the current
11:51:32state and the columns will represent the
11:51:35possible actions which lead to the next
11:51:37state. Now, the formula to calculate the
11:51:39Q matrix is the following. You have Q
11:51:42state {comma} action is equals to R
11:51:44state {comma} action. Okay, let's say
11:51:46the state is your S state number one and
11:51:48you're moving to state number two. Okay,
11:51:50so Q 1 {comma} 2 is going to be so on. R
11:51:53state {comma} action will represent the
11:51:55reward of 1 {comma} 2. Okay, let's try
11:51:57to understand the reward. So, the reward
11:52:00of 1 {comma} 2 is minus 1. Okay, so here
11:52:03you'll get a value of minus 1. And then
11:52:05you'll have plus the gamma parameter
11:52:08into the maximum of the next state and
11:52:10the next possible action that you can
11:52:12take. Okay, now what is a gamma
11:52:14parameter? So, the gamma parameter has a
11:52:17range of 0 to 1. Okay, so it can be
11:52:19between the value of 0 and 1. Now, if
11:52:22gamma is closer to 0, then the agent
11:52:24will tend to consider only immediate
11:52:26rewards. But if the gamma parameter is
11:52:29close to 1, the agent will consider
11:52:31future rewards with greater weight. Now,
11:52:33I don't know if this reminds you of
11:52:35something, but earlier in the session, I
11:52:37discussed two important concepts of
11:52:39reinforcement learning, which was
11:52:41exploration and exploitation trade-off.
11:52:44Now, if you're exploring, then the gamma
11:52:46parameter is going to be closer to 1,
11:52:48but if you're exploiting, then the gamma
11:52:50parameter is going to be closer to 0.
11:52:53It's better if the gamma parameter is
11:52:54closer to 1 because it means that you're
11:52:56exploring the entire environment and
11:52:59you're trying to get future rewards.
11:53:00You're going to get greater weightage
11:53:02and more rewards.
11:53:03So guys, to sum up the entire thing,
11:53:05let's look at how the Q-learning
11:53:07algorithm will solve the problem. So you
11:53:09begin by setting the gamma parameter and
11:53:12the environment rewards in reward matrix
11:53:14R. We already did that. After that, you
11:53:17set the matrix Q to zero because
11:53:19initially the agent will start with no
11:53:21knowledge of the environment. As the
11:53:24agent explores the environment, the Q
11:53:26matrix will start filling up. After
11:53:28that, the next step will be select a
11:53:30random initial state. Like I mentioned
11:53:33earlier, initially you'll randomly
11:53:34select any state because the agent has
11:53:37no idea about the environment. So you'll
11:53:39randomly select a state at step number
11:53:41three. Then step number four is you set
11:53:44the initial state as current state. Step
11:53:47number five, select one among all
11:53:49possible actions for the current state.
11:53:51Okay, so if you've chosen the current
11:53:52state as one, let's say that all the
11:53:55possible states that you can traverse to
11:53:57from one is two, three, and so on. So
11:54:00these are all the possible actions for
11:54:02the current state. Now use the possible
11:54:04actions and consider going to the next
11:54:06state. Once you know what are the
11:54:08possible actions from the current state,
11:54:10you're going to go to one of those
11:54:12possible actions and then you move to
11:54:14the next state. After that, get maximum
11:54:16Q value for this next state based on all
11:54:19possible actions. If you get the maximum
11:54:21Q value, it means that you've chosen an
11:54:24optimal policy in order to reach your
11:54:26end state. Finally, you'll compute the Q
11:54:29value by using the formula that we
11:54:30discussed earlier. And the last step is
11:54:33you have to repeat all of these states
11:54:35until your current state is equal to
11:54:37your goal state. And our goal state is
11:54:39room number five. Now guys, this is a
11:54:41very logical solution because you can
11:54:43easily understand what is happening over
11:54:45here. You have an agent, he has to
11:54:47explore through all the states in such a
11:54:49way that he reaches the end goal using
11:54:51the optimum policy. Okay, that's the end
11:54:54goal of Q-learning algorithm and that's
11:54:56exactly how you're going to solve this
11:54:58problem.
11:54:59And