Full transcript
Course Introduction
0:00hey everyone and welcome back to the
0:01channel and today we've got a very
0:03exciting video because we're going to be
0:05starting a brand new series here on the
0:08channel we're going to be looking at how
0:10you can become a Microsoft certified
0:13fabric analytics engineer bit of a
0:15mouthful you might know it better as the
0:17dp600 exam that you need to pass if you
0:20want to get this certification now a lot
0:22of people have been asking me about
0:24dp600 how do I get certified what do I
0:27need to learn and so in this course what
0:29I'm going to be doing is bringing
0:31together everything that you need to
0:33know for this certification so that you
0:35can hopefully pass first time so what
0:38are we going to be doing in this video
0:39well this is a bit of an introduction
0:41okay so we're going to be looking at
0:42what's in the exam what are some of the
0:45details of the exam so how long is it
0:47how how can you pass do you do it online
0:49or in person that kind of thing we're
0:51going to look at an overview of this
0:52course so specifically what are we going
0:55to learn in this course in YouTube and
0:57how are you going to learn as well then
0:59we're going to finish by looking at why
1:01you might want to take this course and
1:02why you might want to get fabric
1:04analytics engineering certification in
1:06general so what's in the exam so the
1:09first section is around planning
1:11implementing and managing solutions for
1:14data analytics now within that we've got
1:16various kind of sub modules okay so you
1:20need to understand the requirements how
1:22do we identify requirements for
1:24Solutions how do we do things like
1:26security how do you know the difference
1:28between different data gateways how do
1:30you set up access control and workspaces
1:34and capacities and how do you modify the
1:36settings of all these things as well
1:38another important section of the exam
1:40and being an analytics engineer in
1:42general is Version Control you'll need
1:44to understand how to set up Version
1:47Control for Azure devops understand some
1:50of the settings and the configuration
1:51options that that entails as well as
1:54some of the deployment pipeline
1:56functionality as well now this is not a
1:58completely exhaustive list of what's in
2:00that section of the exam we'll go
2:02through each of those in a bit more
2:03detail it's just the high level kind of
2:05areas that are covered in that section
2:07and this section is worth 10 to 15% of
2:10the exam so it's a small section but
2:12it's a very important section I think
2:14anyway in terms of becoming an analytics
2:16engineer these are really important
2:18topics that you need to understand so up
2:19next we've got 40 to 45% of the exam so
2:23this is really the core of the exam that
2:25you need should probably spend most of
2:27your time studying okay and it's a
2:29around preparation and serving of data
2:33so this is one of the core tasks that
2:35you'll be asked to carry out as an
2:38analytics engineer and within that we've
2:40got quite a lot of really important
2:42topics okay so you've got understanding
2:44The Lakehouse how do we set one up how
2:47do we create tables difference between
2:49tables and files the warehouse so the
2:51tsql experience creating shortcuts
2:54ingesting data from external locations
2:57into our warehouse or our Lakehouse what
3:01are the different methods that we can
3:02choose here and when do we choose
3:04specific ones okay then we've got data
3:06transformation so once we've got our
3:08data within fabric we might want to do
3:10some Transformations on it now you can
3:12do that with tsql spark and you will
3:15need to know at least the basics of tsql
3:17and Spark and also Dax in the next
3:19section as well so there's quite a lot
3:21of languages that you need to know to
3:23kind of an intermediate level I would
3:24say for this exam next we've got
3:26performance so how do we optimize
3:29performance both in terms of the getting
3:32data into fabric but also in terms of
3:34the transformation piece so if you've
3:36got spark jobs that are running really
3:39long or you've got some tsql scripts
3:41that are really not very performing very
3:43well they're taking a long time to kind
3:45of return your data what can you do to
3:47monitor performance then optimize it and
3:49as I mentioned you will need to know P
3:51spark and T SQL to a kind of
3:53intermediate level there so that's
3:55something to bear in mind for this exam
3:57so the next section of the exam is
3:58around implementing and managing
4:00semantic models and this is worth 20 to
4:0325% of the exam and within this category
4:06you've got the different storage modes
4:09so direct query import mode and direct
4:12late mode and kind of understanding how
4:14these things work how to set up direct
4:16Lake mode when to use it maybe when not
4:18to use it you're going to need to have a
4:20good understanding of Dax here there's
4:22quite a few questions around Dax and Dax
4:25studio and tabul editor 2 as well so
4:28these are kind of things that you need
4:29to be be aware of because you'll
4:31probably get some questions around that
4:32as well also within the section you've
4:34got data modeling so things like Star
4:36schemas Bridge tables how do we deal
4:39with many to many relationships and also
4:41we've got things around security how do
4:43we set up roow level security within
4:45your santic model object level security
4:47and how do we validate that that's
4:49actually working correctly so in the
4:50final section of the exam which is worth
4:52around 20 to 25% again is explore and
4:56analyze data so here we're really
4:58looking at the analysis part of being an
5:01analytics engineer because most of your
5:03role might be in around the engineering
5:06piece but really you need to know the
5:08analytics side of thing as well because
5:10that's going to make you a much better
5:11analytics engineer so in this section
5:13you're going to be asked questions about
5:14data analysis specifically using tsql
5:17okay so analyzing your lake house SQL
5:20endpoint analyzing your data warehouse
5:23coming up with insights from data using
5:25tsql you'll also be asked about data
5:27profiling so understanding the pro
5:29profile different tables based on some
5:31of the profiling metrics that you get
5:33and also analyzing data via the xmla
5:35endpoint so that's another part of this
5:38section of the exam so these are the
5:39four sections there's quite a lot to go
5:42through and it covers quite a broad
5:44range of skills that you need to know
5:47from ppar tsql Dax data modeling getting
5:51data in data modeling serving data in
5:53semantic models as well so quite a broad
5:55range endtoend exam so in the exam there
5:58is between 40 and 60 question and the
6:02results are scaled and you get given a
6:04result between 0 and 1,000 so the pass
6:07Mark here is 700 out of 1,000 and that's
6:10the scaled score so that's something to
6:12bear in mind and the questions can be of
6:14different types okay so some of them are
6:16multiple choice some of them might be a
6:18case study and the case studies normally
6:20take two three four questions so you
6:22really need to understand what's going
6:23on and you get multiple questions about
6:25the same case study you also have things
6:27like drag and drop and ordering lists
6:29and for a full list of the question
6:31types recommend you go to this resource
6:33here it's on the Microsoft learn the
6:36exam question section of Microsoft learn
6:39I'll leave a link to that in the
6:40description below you can also use the
6:42exam sandbox that provides you with a
6:44Sandbox experience of the exact question
6:47types that you can expect in the exam
6:49you'll get given 100 minutes to actually
6:51carry out the exam you should set aside
6:53at least 2 hours though so 120 minutes
6:56for the exam just so that you can kind
6:57of get in get settled as I mentioned
7:00it's 700 out of 1,000 to pass this exam
7:03so 70% you can take the exam either in
7:06person in Pon view sensors or you can do
7:09it online so my personal recommendation
7:11would be to take the exam in person if
7:14you can um it kind of eliminates a lot
7:16of the doubts and the problems around
7:19like Wi-Fi and worrying about your desk
7:21setup and your room has to be a very
7:23specific layout and so if you can visit
7:26a center in person I think it's a lot
7:28better cuz you can just walk in take the
7:30exam and walk out whereas online you
7:31have to think about lots of different
7:33things for me it's less stressful to do
7:36it in person if you can so there is an
7:39exam fee which is
7:40$165 for people in the USA it varies for
7:44different countries but if you look on
7:46the right hand side we got a free exam
7:48so if you're very quick and you go to
7:51the link in the description and you
7:53complete the fabric AI skills challenge
7:55training course before the 19th of April
7:58so you've only got a few days you can
8:00get a voucher to take the DP 600 exam
8:04for free so you can do that skills
8:06challenge training very quickly get your
8:09free voucher and then watch all of this
8:11series on YouTube and as long as you
8:13take that exam I think it's before June
8:15the 22nd or something so you've got two
8:17or three months to kind of go through
8:19the material at your own pace and then
8:21you get to save yourself
8:23$165 or the equivalent in your currency
8:26in your country so this is what this
8:28series is going to look like here on
8:30YouTube we've got the first video which
8:31is an introduction to the course then
8:33I'm going to be covering all 11 chapters
8:36of the exam and the content is going to
8:37be delivered through various real world
8:40scenarios CU I want to make this content
8:42interesting for you and also make it
8:44relevant for you if you want to be an
8:47analytics engineer or if you are an
8:48analytics engineer currently in your
8:50career and to do this I'm going to be
8:52combining Theory so I think there is
8:54some theory in some of these modules you
8:56do have to understand how things work
8:58but then practice as well how to
8:59actually implement this stuff in fabric
9:02at the same time throughout the course
9:03I'm going to be reinforcing that
9:05knowledge as we go through and as asking
9:07rhetorical questions as well as sample
9:10questions and at the end I do plan to go
9:12through a full video kind of like a
9:13practice paper let's say designing lots
9:16of questions that you can expect within
9:18the exam as with all of my courses all
9:21of the resources and module notes and
9:23scripts and notebooks and all this kind
9:25of thing I'll be posting that in the
9:27school community so if you're not
9:28already a member there I'll leave a link
9:30in the description below it's completely
9:32for free so make sure you sign up and
9:34yeah you'll get access to all of that so
9:36there's a few reasons that I just want
9:37to cover quickly here around why you
9:39might want to take this course and then
9:41go on to become a certified analytics
9:44engineer well the course gives structure
9:46to your learning Microsoft have given us
9:49a study guide and they've said these are
9:51what we think is important to learn for
9:54an analytics engineer so it gives you a
9:55really good Pathway to follow might also
9:58help you get a new job having that
10:00certification on your CV is definitely
10:02not going to hinder your chances it
10:04would also be good to go into kind of
10:06promotion talks or payiz talks with your
10:08boss and say yeah well last year I did
10:10the certification and this is especially
10:12true if you work in a consultancy right
10:15because consultant here they're kind of
10:16selling your skills and your experience
10:18so if you have certification then that
10:21can help them win work in the future so
10:23in this lesson we've looked at what's in
10:25the exam we've looked at some of the
10:26exam details we've also looked at the
10:29overview of this course so what are the
10:31different chapters that we're going to
10:32be looking through and why I think you
10:34should take this course join us in the
10:37next lesson where we'll be starting the
10:40course properly and we'll be looking at
10:43how to plan and Implement a data
10:45analytics solution hello and welcome to
Plan a data analytics environment
10:48this first chapter in this dp600 exam
10:52preparation course we're going to be
10:54looking at how to plan a data analytics
10:57environment in fabric now this is the
10:59first chapter in 11 chapters that we're
11:01going to be going through teaching you
11:03everything you need to know to hopefully
11:04pass the dp600 exam in this chapter
11:08we're going to be covering exactly what
11:09you need to know if you look at the
11:11study guide in Microsoft learn these are
11:14the elements that we're going to be
11:15covering how do we identify requirements
11:17for a solution so the various components
11:20features performance capacity skus that
11:23kind of thing how do we make decisions
11:24about that we're also going to be
11:26looking at how to recommend settings in
11:27the fabric admin portal
11:29how do we choose data Gateway types and
11:32also creating custom powerbi report
11:35theme towards the end of the lesson
11:36we'll be testing your knowledge with
11:38five sample questions and just as a
11:41reminder all of the lesson notes and key
11:44points and link to further learning
11:46resources they're going to be published
11:47on the school community so if you're not
11:49already a member I'll leave a link in
11:50the description now you play the main
11:53character in a scenario and this
11:55scenario is going to walk you through
11:57everything you need to know for those
11:58four ele Els of the study guide are you
12:00ready let's begin so you are a
12:02consultant and you're starting your
12:04first day on a new project and this is
12:07Camila she is your client for the
12:10project and on the phone before the
12:11meeting Camila had mentioned that she
12:13wants to implement fabric but she
12:15doesn't know really where to start and
12:17that's where you come in you're going to
12:19start with a requirements Gathering
12:21Workshop so you organize a full day
12:23workshop with Camila the client to truly
12:26understand their business and their
12:27requirements now your goals for this
12:29Workshop are to extract a set of
12:32requirements from the client to help you
12:34build a plan for their new fabric
12:37environment and another goal is to do
12:40such a great job in planning their
12:42environment that the client is going to
12:43give you a new contract by the end of it
12:45to build the thing okay so this
12:47requirements Gathering Workshop what are
12:49you going to ask Camila what do you need
12:52to know when you're identifying the
12:53requirements you should think about
12:55focusing on these three elements to
12:57begin with the capacities so how many do
12:59we need what sizing do the capacities
13:01need to be in this new environment then
13:03we're going to look at data ingestion
13:05methods so there's lots of different
13:06ways that we can ingest data into fabric
13:09you're going to ask a set of questions
13:11that's going to kind of deduce the best
13:13method for getting data into fabric
13:15based on the requirements similarly
13:17we've got data storage so we've got
13:19three different options for storing data
13:21in fabric how do you ask the right
13:23questions and identify the requirements
13:25to choose the right one so let's start
13:27off thinking about capacity requ
13:29requirements now the requirements that
13:30we need here are really the number of
13:32capacities that are required and the
13:34sizing so the SKU the stock keeping
13:37units you probably know by now that in
13:39fabric we have capacities of varying
13:42sizes so what determines the number of
13:44capacities required so from previous
13:46videos you've probably understood that
13:48one of the things that impacts the
13:50number of capacities required is
13:52compliance with data residency
13:53regulations so the capacity dictates
13:56where your data is stored so if you have
13:59regulations that dictate that your data
14:02must reside in the EU for example for
14:04gdpr that's going to be one capacity in
14:06your business if you have other
14:08requirements that say these data sets
14:09need to be stored in the US you're going
14:12to have to have a separate capacity for
14:13that as well another thing that can
14:15impact the number of capacities is the
14:17billing preference so the capacity is
14:20how you get build in fabric so some
14:22organizations might want to separate the
14:24billing between different departments in
14:27their organization so they might have
14:28one capacity for the finance department
14:31one capacity for your Consulting
14:32division one capacity for your marketing
14:34department for example another thing
14:36that could determine the number of
14:37capacities that you need is segregating
14:40by workload type so if you have a lot of
14:42heavy intensive data engineering
14:44workloads then you might want to put
14:46those in a separate capacity and give it
14:48enough resource to allow you to do that
14:51in a confined capacity then you're
14:53serving of business intelligence you
14:55might want to do that in a separate
14:56capacity so that the read performance on
14:58those kind of dashboards is not impacted
15:00by the heavy data engineering stuff
15:02maybe machine learning stuff that's
15:04being done in other capacities you might
15:06also want to segregate by department
15:08just through business preference as well
15:09aligned with that billing preference so
15:11some companies like to have their
15:13capacity aligned to various dep
15:14departments within their business so
15:16these are the things you need to extract
15:18in terms of requirements when you're
15:19talking with this client and what about
15:21the sizing well we've touched on that
15:23already but some of the things that
15:24impacts the sizing of a capacity are the
15:27intensity of the expected workloads so
15:30are you going to be doing High volumes
15:31of data ingestion are you going to be
15:33getting gigabytes of fresh data into
15:36Fabric or even terabytes of data into
15:38fabric every day these are going to use
15:40a lot of your resources and to go
15:42through them quickly it helps if you
15:44have a higher capacity similarly heavy
15:46data transformation so if you're doing a
15:48lot of heavy transformations in spark
15:51that's going to use a lot of resources
15:52so if that's something you're going to
15:53be doing regularly in your business you
15:56want to be choosing a high capacity for
15:57that again machine learning training can
16:00you be very resource intensive going to
16:02take hours or sometimes even days to
16:04train a machine learning model if that's
16:06something you're going to doing
16:07regularly you want to be having that on
16:08a high capacity the budget of your
16:11client also dictates the capacity the
16:14sizing of the capacity that you're going
16:16to choose obviously the more resources
16:18the higher that SKU that you decide the
16:21more expensive it's going to be and some
16:23clients might be very sensitive around
16:25the cost and related to that is can you
16:27afford to wait or can the client afford
16:29to wait because if you procure a F2 skew
16:33it's probably going to go through your
16:34data but it might take a very long time
16:37and in some business that might not be a
16:38problem maybe you're just doing data
16:39ingestion once per day you ingest all of
16:42your fresh data and it might take a lot
16:45longer on an F2 capacity but that's not
16:47necessarily a problem maybe you can do
16:48it overnight and by the time people come
16:50in in the morning all of that data has
16:52been ingested or transformed and it's
16:54ready for consumption in the morning so
16:56what's your propensity to wait now some
16:59other companies might have regular data
17:01coming in every hour like gigabytes of
17:04data every hour and in that scenario you
17:06really need a high capacity to be able
17:08to churn through all of that stuff and
17:10get it processed before the next hourly
17:14load for example another thing that can
17:16determine the sizing of the capacity is
17:19does the client want access to f64
17:22Features so there's quite a lot actually
17:24of features that open up when you get to
17:27f64 so co-pilot being a good example
17:30currently and there's many many more
17:32I'll list them on the screen here these
17:34are features that only really are
17:35available if you choose f64 capacity or
17:39above so that's something to bear in
17:40mind if you want to use any of these
17:42features you need an f64 plus so what
17:44about the data ingestion requirements
17:46well here what we really need to know is
17:49what are the fabric items and or
17:51features that you need to get data into
17:53Fabric and how are you going to
17:54configure these items once you've built
17:56them now some of the options here and
17:58this is not an exhaustive list there's
18:00lots of different options here we've got
18:02the shortcut database mirroring ETL via
18:05data flow ETL via data Pipeline and a
18:07notebook and the event stream so these
18:09are some of the options that you might
18:11want to consider so what are the
18:13questions that you need to ask of a
18:15client when you're identifying the
18:17requirements to help you make the
18:19decision here well these are some of the
18:20deciding factors the main one really is
18:22where is the external data stored if
18:25it's in ADLs Gen 2 Amazon S3 or S3
18:29compatible storage location like Cloud
18:31flare for example Google Cloud Storage
18:33or the data verse but then these are the
18:35ones that are going to be available for
18:37you to shortcut into fabric so if you
18:40get any questions in the exam around you
18:42know my data rest stored in ADLs Gen 2
18:45well obviously the shortcut is a good
18:47option for that now it's not necessarily
18:48the only option you can still do ETL via
18:52any of these storage locations but it
18:53does open up that shortcut possibility
18:56now if you see Azure SQL Azure Cosmos DB
18:59or snowflake mentioned then immediately
19:01you should start thinking okay this
19:02could be database mirrored so you can
19:05use database mirroring to create that
19:07kind of live link to the database and
19:09it's going to maintain a mirror inside a
19:12fabric is it on premises now if you're
19:14data stored on premises then you're
19:17going to be probably want to be using
19:18the ETL via data flows or data pipelines
19:22because these two activities these two
19:24items allow you to create that
19:26on-premise data Gateway on your on
19:29premise server and then connect to that
19:30via the data flow or the data Pipeline
19:32and if you got realtime events realtime
19:34streaming data obviously you probably
19:36want to using the event stream to get
19:38that data into fabric anything else
19:40really you're going to be looking at ETL
19:42by either the data flow the data
19:44Pipeline and the notebook and when to
19:46choose which one well I've done a very
19:48long video I'll leave a link in the
19:49description or you can click here to
19:51make that decision about which of these
19:53is best for that particular organization
19:56so related to that is also what skills
19:58exist in the team because you don't want
20:00to build a solution that can't be
20:02maintained managed by the company or
20:04your company or your client's company so
20:06if you're looking for a predominantly no
20:08and low code experience then you're
20:10going to want to be focusing on the ETL
20:13via data flows and data pipelines both
20:15of these are fairly low and no code
20:18experiences help you get data into
20:19fabric if You' got a lot of SQL
20:21experience in your team then here you
20:23can be using the data pipeline you can
20:25use the script activities to do
20:27Transformations on your data as is
20:28coming in and if you have people that
20:30are familiar with spark python Scala
20:33that kind of thing then you can use the
20:36ETL notebook if you're you know perhaps
20:39you've got data coming from a rest API
20:41and you want to be using python
20:42libraries to get that in that's a good
20:44option for you there so whil we're on
20:46the topic of data ingestion there's a
20:48few other features you need to be aware
20:50of that might come up in the exam that
20:53can help you identify different
20:54requirements for getting data into
20:55fabric these are the on premise daily
20:58Gateway which we've mentioned the v-net
21:00the virtual network data Gateway fast
21:02copy and staging so you might be asking
21:05some questions about these things in the
21:06exam as well so when do we decide on
21:09these sorts of things well you need to
21:11ask how the data in the external system
21:13is being secured right so if it's on
21:16premise if it's an on premise SQL Server
21:18you have to be using the on premise data
21:20Gateway if your data is living in Azure
21:23behind some sort of virtual Network or
21:25private endpoint that kind of thing then
21:28you want to be setting up the v-net data
21:30gateway to access that and in terms of
21:31the volume of data this is also going to
21:33have an impact on the items that you
21:36choose for doing your data ingestion and
21:38also some of the features available so
21:40if you've got low or medium data per day
21:43well if it's low then you probably don't
21:44need any of these specific features like
21:46the out of the box Solutions will be
21:47good enough but if you've got quite a
21:49lot of data gigabytes per day in that
21:52kind of range you want to be using some
21:54of the features like Fast copy and
21:56staging similarly if You' got very high
21:58amounts of data these are going to be
22:00one of using the fast copy and the
22:01staging if you're using data flows
22:03alternatively you can use data pipelines
22:06and if you can get data in bya a fabric
22:08notebook then that's another option as
22:10well so before we move on I just want to
22:12mention a bit more detail around the
22:14data gateways now as you probably know
22:17already there are two types of data
22:19gateways that we can configure in
22:21Microsoft fabric number one is the
22:23on-premise data Gateway and number two
22:25is the virtual network data Gateway and
22:28a data Gateway in essence helps us
22:30access data that's otherwise secured so
22:34if his data is on an on-premise SQL
22:36server for example it gives us a secure
22:38way to access that data and bring it
22:40into fabric likewise if you've got data
22:43behind a virtual Network secured in
22:45Azure in like blob storage or ADLs Gen 2
22:49it provides us with a secure mechanism
22:51to access that data so I'm not going to
22:53show you step by- step how to set up a
22:55data Gateway in this lesson but what it
22:58done is linked to two other videos by
23:00other creators that show the process in
23:02detail if you want to go and have a look
23:04I'll leave that in the school Community
23:05but I do want to just cover kind of the
23:07high level process for each of them just
23:09so you understand a bit more about what
23:11that looks like if you've never set one
23:13up before so for the on- premise data
23:15Gateway there's a few high level steps
23:18number one we need to install the data
23:19Gateway on the on premis server and if
23:22you've already got an on- premise data
23:23Gateway set up on your on premise server
23:26perhaps you're using it in traditional
23:28powerbi data flows for example then
23:30you're going to need to update it to the
23:32latest version cuz that's going to be
23:33compatible with Microsoft fabric the
23:36next step is to in fabric create a new
23:39on premise data Gateway connection and
23:41then from that you can connect to that
23:43data Gateway from either a data flow and
23:46now also a data pipeline so the data
23:48pipeline was recently added in the last
23:49few weeks I think it's still in preview
23:52that connection so you might not get
23:54asked about it in the exam but it's good
23:55to know that now it's actually possible
23:57via the data flow and the data pipeline
23:59to set up the v-net data Gateway we're
24:01going to start in Azure there's a few
24:03settings that you need to configure in
24:05your Azure environment before you can
24:07set up the v-net data Gateway connection
24:10so you're going to need to register a
24:12Power Platform resource provider within
24:14your Azure subscription and then within
24:16the item that you want to share or you
24:18want to access for example in your Azure
24:21blob storage item in Azure you need to
24:23create a private endpoint in the
24:24networking settings then create a subnet
24:27and then we're going to use that in
24:28fabric to create a new virtual network
24:31data Gateway connection and then again
24:33from that you can connect to it via your
24:34data flow to be able to access that data
24:37that is behind that virtual Network in
24:40Azure so next let's look at the data
24:42storage requirements and when we're
24:44talking to our client here identifying
24:46requirements really what we're trying to
24:48extract is okay what fabric data stores
24:52are going to be best for these
24:54requirements and what overall
24:56architectural pattern are we going to be
24:58aiming for with this solution now the
25:00options here are obviously The Lakehouse
25:02the data warehouse and the kql database
25:05and some of the deciding factors to
25:07choose between these are well what's the
25:09data type okay so is it structured or
25:12semi structured or even unstructured so
25:15are you going to be getting raw files
25:18CSV Json maybe from AR rest API is it
25:21unstructured is it image data video is
25:24it audio data for example these are all
25:27going to be wanted to store in The
25:28Lakehouse because this is kind of the
25:30only place in fabric where you can store
25:32a variety of different file formats if
25:35your data is relational and structured
25:37then obviously you can keep that in
25:39either the lake house or the data
25:41warehouse and if it's real time and
25:42streaming you're going to be want to
25:43streaming that into your kql database
25:46next up another important consideration
25:48when choosing a data store is what
25:51skills exist in the team so if you're
25:53predominantly tsql based then you're
25:55going to be want to be using the data
25:57warehouse experience
25:59if you're predominantly spark and python
26:01Scara that kind of thing then you're
26:03going to be wanting to storing your data
26:04predominantly in the lakeh house and if
26:06you're predominantly using kql in your
26:08organization that's going to be one to
26:10using kql database for your data storage
26:14congratulations you've completed your
26:16first engagement for Camila you've
26:18convinced her to set up a proof of
26:20concept project in her organization so
26:23she's already created the fabric free
26:25trial she's set up her environment but
26:28immediately she's hit a bit of a hurdle
26:30so this is your next mission she Rings
26:32you up and she says hey I need some help
26:34I open the fabric admin portal and
26:37nearly had a heart attack please can you
26:39help me understand all of these settings
26:41so you set up a call with Camila to help
26:43her understand the fabric admin portal
26:46how are you going to teach her and what
26:47are you going to teach her about the
26:49admin portal what are the most important
26:51settings that she needs to know about
26:52okay just before we get into the fabric
26:54admin portal and look at some of the
26:56settings available to us in there there
26:58it's important to note that to be able
27:00to access the admin portal of course
27:02first you need a fabric license but then
27:04you need to have one of the following
27:07roles you need to be either a global
27:09administrator a Power Platform
27:11administrator or a fabric administrator
27:14So within the admin portal here in
27:17fabric you'll see this menu on the left
27:20hand side so these are some of the
27:21important settings in tenant settings
27:24here you can allow users to create
27:25fabric items so if you just set up
27:27fabric in your ization you need to allow
27:29people to actually create fabric items
27:31without that you can't really get very
27:32far enable preview features so every
27:35time Microsoft release new features
27:37normally they put them in the admin
27:39portal and you can allow or disallow
27:42users in your organization to use them
27:44you can also allow users to create
27:46workspaces there's a whole host of
27:48security related features that you can
27:51manage and get control over in your
27:53tenant so for example how do you manage
27:55guest users allowing single sign on
27:57options for things like snowflake big
27:59query red shift accounts that kind of
28:01thing how do you block public internet
28:03access so that's really important to
28:05know enabling other features like Azure
28:08private link for example allowing the
28:10service principal access to the fabric
28:12apis so if you're going to be doing some
28:13automation you need to allow access to
28:17service principles to the API there's
28:19also options in there for allowing git
28:21integration so if you're setting up
28:22Version Control that needs to be enabled
28:25there and there's also some features
28:27like allowing cop pilot within the
28:28organization as well now in general some
28:31of the settings can be one of three
28:33things it could be enabled for the
28:35entire organization it can be enabled
28:37for specific security groups so say you
28:39only want super users to be able to use
28:43this feature or admins within your
28:45fabric environment to use a specific
28:47feature then you can enable it for
28:48specific security groups or you can
28:50enable it for all except certain
28:52security groups so everyone in your
28:54organization gets access apart from
28:57these people perhaps guest users is a
29:00good example now other settings in the
29:02fabric tenant settings are kind of
29:04binary you either enable them or you
29:06disable them for the entire organization
29:08another important point in the fabric
29:10admin portal are the capacity settings
29:13so this section here and in here you can
29:16create new capacities delete capacities
29:18manage the capacity commissions and also
29:21change the size of a capacity so these
29:23are some important capacity settings
29:25that you need to be aware of understand
29:27how they work and how to manage them
29:29within your fabric environment so great
29:31you've talk Camila about the fabric
29:33admin portal and she's very grateful but
29:36before the meeting ends she has one more
29:38thing she wants to ask you about she
29:40says one final thing before you go and
29:42it might seem a bit random but when we
29:44migrate to fabric I want our bi team to
29:47create more consistent reports have you
29:50got any ideas about how we can achieve
29:51that and of course the first thing you
29:53think of are custom powerbi report
29:55themes now there are many ways to create
29:58a custom report theme in powerbi you can
30:01either update the current theme if
30:03you're in powerbi desktop or you can
30:05write a kind of Json template yourself
30:07using the documentation if you're
30:08feeling a bit Brave you can do that
30:10yourself or you can also use a third
30:12party online tool there's quite a few
30:14report theme generator tools that exist
30:16online but it's unlikely you're going to
30:18be tested on that in the exam so your
30:20task is to show Cilla how to create a
30:22custom report theme so let's have a look
30:25at how you can do that within powerbi
30:27desktop top so here I've got a report
30:30and what I'm going to do to access the
30:31report themes you need to go to the view
30:33tab then you can see these themes here
30:35and obviously these are the preset
30:37themes so you can just click and update
30:39the current theme very simply like that
30:41but to do most of the customization you
30:43need to click on this button here and
30:45you can access the current theme and all
30:47accessible themes are currently
30:49installed on this machine then there's a
30:51few settings down here that quite
30:52important to know so browse for themes
30:55if you click on that it's going to allow
30:56you to import a powerbi report theme so
31:00if you've already got a theme Here For
31:01example this one here then you can
31:03select that and install it into your
31:05environment like so if you want to
31:07customize the current theme you can do
31:10that like this and it's going to bring
31:11you through to this UI environment just
31:13to you know change some colors change
31:15some text change some visuals what you
31:18have to think about for this section of
31:20the exam is what could they ask you you
31:22have to think about how could you
31:23possibly be tested on this so in terms
31:26of the powerbi report theme stuff you're
31:28likely to be tested on these buttons
31:30here and what they do plus they could
31:32ask you about a Json theme so they could
31:35show you a Json theme and maybe ask you
31:38about okay how can you edit this theme
31:40what doesn't look right in this theme
31:42that kind of thing so it's good to have
31:44a bit of familiarity about the different
31:47sections in these Json files so the name
31:50of it how you can store your data colors
31:52as a list some of these different
31:54settings here you're probably not
31:55expected to memorize all of the
31:57different settings in Json format but
31:59you might get shown a theme in Json
32:01format and asked to modify it or asked
32:04to comment on it in some way to export
32:06the current theme you can also use this
32:08save current theme and that's going to
32:09allow you to export a Json file that you
32:11can share within your organization and
32:13you've also got here access to the theme
32:15Gallery so this is going to bring you
32:17through to the theme Gallery website
32:19where you can download other people's
32:21themes for your report to finish up this
32:23video and this lesson we're going to go
32:26through five practice questions just to
32:28kind of solidify that knowledge make
32:29sure you're understanding some of the
32:31key Concepts within a context of a
32:33scenario so the first question is you're
32:36running an F2 capacity and you regularly
32:39experience throttling with that capacity
32:41now there's a number of long running
32:43spark jobs that take on average 3 hours
32:45to complete and you need these to
32:47complete in under 1 hour so you plan to
32:49increase the SKU of the capacity where
32:52would you go to make this change would
32:53you go to the workspace settings and
32:55configure spark settings would you go to
32:57to the admin portal and the capacity
32:59settings section and then click through
33:01to Azure to update your capacity would
33:03you go to the monitoring Hub and look at
33:05the Run history or would you use the
33:07capacity metrics app so pause the video
33:09here and have a bit of a think and then
33:12I'll move on so the answer is 2b so you
33:15can manage your capacity settings within
33:16admin portal and then capacity settings
33:19and then you can actually click through
33:21to Azure it gives you a link to the
33:22Azure portal and that's where you're
33:24going to change the capacity within the
33:26Azure portal OB so you can't be within
33:28the spark settings that's for managing
33:30the configuration of your spark cluster
33:32within a workspace and in the monitoring
33:34Hub we can't get anything there to do
33:36with capacity settings that's just going
33:38to tell you how your jobs are running
33:39and in the capacity metrics app that's
33:41just a readon app for having a look at
33:44how your capacity is being used so that
33:46wouldn't also be suitable either
33:48question number two your data governance
33:50team would like to certify a semantic
33:52model to make it discoverable in your
33:55organization now only the data
33:56governance team should be able to do
33:58this in what order should you complete
34:00the following tasks to certify a
34:02semantic model so have a look at the
34:04five actions here and what you're going
34:07to have to do is put these in an order
34:10so these are this is an ordered list it
34:12should be so one of the things you'll
34:13have to do first second third fourth and
34:16fifth so once you've got these in order
34:18we'll move on so let's look at the
34:19answer now so the correct order looks a
34:21bit like this so we start by creating a
34:25security group for the data governance
34:27team the clue in the question was that
34:30only the data governance team should be
34:32able to do this so when you see that you
34:35think okay well they need to be within a
34:37security group to enable this number two
34:40and you could argue that one and two
34:41could be interchangeable but these are
34:43the first two items anyway but enable
34:45the make certified content discoverable
34:48So within the admin portal the tenant
34:51settings as a section for Discovery
34:54you're going to need to enable that for
34:55the organization and then after that
34:57going to have to make sure that that
34:59settings is applied only to the data
35:01governance security group that you set
35:03up then you're going to need to ask the
35:05data governance team to go into the
35:08semantic model settings and then
35:09endorsement and Discovery and click
35:12certify for that semantic model and then
35:14you want to validate that that has been
35:16set up correctly and your business user
35:18can see that certified semantic model
35:22within the one leg data Hub three you
35:24join a new company and you're given a
35:26powerbi report theme as a Json file to
35:28use for all new projects how do you
35:30apply this Json file theme to the report
35:33that you're currently developing is it a
35:35in powerb desktop go to view themes and
35:38customize current theme B go to the
35:40fabric admin portal click on custom
35:43branding and then set the default report
35:45theme C use tabulate editor 2 to update
35:48the theme or D in power by desktop go to
35:51the view themes and then browse for
35:53themes so the answer here is D in power
35:57desktop go to the view themes and then
35:59browse for themes so d and a are quite
36:02similar but a is for customizing a
36:05current theme so that's not going to be
36:06allowing you to import a Json file
36:09that's going to allow you to use the
36:11user interface to update the current
36:13theme so that's not what we want to do
36:15we want to import ad Json file as our
36:17report theme which is possible using D B
36:20that functionality doesn't actually
36:21exist custom branding does exist but
36:23that allows you just to update the
36:25colors and the icons within fabric not a
36:28default report theme and tabul editor 2
36:31is also the incorrect answer question
36:33four you have 1,000 Json files stored in
36:36Azure data Lake storage ADLs Gen 2 that
36:39you want to bring into fabric the ADLs
36:42Gen 2 storage account is secured using a
36:45virtual Network which of these actions
36:48would you need to perform first is it a
36:51in fabric go to manage connections and
36:53gateways and then click on create a new
36:56virtual network data Gateway B create a
36:58shortcut to the ADLs Gen 2 storage
37:01account C in Azure register a new
37:03resource provider and create a private
37:05endpoint and subnet or D install an on
37:08premise data Gateway on an Azure virtual
37:10machine in the same virtual Network or E
37:13enable Public Access in the storage
37:15account network settings so for this one
37:17you'll remember that the answer is C so
37:20the first step in setting up a virtual
37:22network data gateways well we need to go
37:24into Azure we need to perform some
37:26network configuration okay so you need
37:28to register that new resource provider
37:31that Microsoft Power Platform resource
37:33provider within your subscription and
37:35then on the item create a private
37:36endpoint and a subnet all of the other
37:38options some of them are steps in the
37:41process but not the first step so the
37:43question was which of these actions
37:45would you need to perform first so yes
37:48we do need to do a but it's not going to
37:50be the first thing that you're going to
37:51do B is kind of a bit of a a red herring
37:54here cuz you might have seen ads Gen 2
37:56and thought ah shortcut but actually you
37:58need to configure the virtual Network
38:00daily Gateway before you can even think
38:03about kind of connecting to it D
38:04installing the on- premise data Gateway
38:07well you know that we're looking at a
38:08virtual Network here so you're going to
38:10be choosing the virtual network data
38:12Gateway rather than an on- premise Ste
38:13goway and E enable Public Access in the
38:16storage account network settings while
38:18that's going to expose your data to the
38:20public internet so not advisable
38:23question five you have data stored in
38:25tables in Snowflake which of the
38:27following cannot be used to bring the
38:29data into fabric a use the data pipeline
38:32copy data activity B create a shortcut
38:34to the snowflake tables from your Lake
38:37housee C use the data flow Gen 2 with
38:40the snowflake connector D use database
38:42mirroring to create a mirrored snowflake
38:45database in fabric so the answer here is
38:47B to create a shortcut to the snowflake
38:50tables from your lake house as you'll
38:52know you can only shortcut to ADLs Gen 2
38:55or Amazon S3 or Google Cloud Storage so
38:59the ability to shortcut is generally on
39:02files when we're talking about tables in
39:05databases whether there all of the other
39:07three we can use so you can do a copy
39:09data activity from a data pipeline to
39:11bring that data in if you want to copy
39:13it in or you can use a data flow Gen 2
39:16or you can use database mirroring
39:18because snowflake is one of the
39:20databases where database mirroring is
39:22possible Camila says thanks she's
39:24seriously impressed with your knowledge
39:26well done in this lesson we've looked at
39:28how you can identify requirements for a
39:30fabric solution we've looked at the
39:32different types of data gateways that
39:34are available to us in fabric we've
39:36looked at the settings in the admin
39:38portal and we've also looked at how to
39:40create custom powerbi report themes and
39:43the good news is you want an extension
39:44to the contract Camila would like you to
39:47implement and manage her data analytics
39:49environment so you've got the next stage
39:51of the contract in the next lesson we'll
39:53look at how you can do that how you can
39:55set up Access Control sensitivity
39:58labeling workspaces capacities all that
40:01kind of stuff how do we set these things
40:03up inside fabric so make sure you click
40:05here for the next lesson hey everyone
Implement and manage a data analytics environment
40:07welcome back to the channel today we're
40:09going to be continuing our dp600 series
40:12and we're going to be looking at
40:13implementing and managing a data
40:16analytics environment this is the second
40:18part of the dp600 syllabus or the study
40:21guide that we're going through on the
40:23channel and today we're going to be
40:24going through these particular elements
40:26and these are coming straight from the
40:28study guide so number one we're going to
40:30be looking at implementing workspace and
40:32item level Access Control implementing
40:34data sharing for workspaces warehouses
40:37and lake houses managing sensitivity
40:39labels configuring fabric enabled
40:42workspace settings and managing fabric
40:45capacity we've got five sample questions
40:47so at the end of the video we'll be
40:48going through some sample questions to
40:50test your knowledge and as with the last
40:52video I'll be posting the key points and
40:55links to further resources in the school
40:57Community available for free I'll leave
40:59a link in the description for that we'll
41:01continue the scenario that we were
41:03developing in the last lesson you're the
41:05main character again are you ready let's
41:07begin so just to recap you are a
41:09consultant and you're working with your
41:10client who is called Camila and in the
41:12last engagement you successfully planned
41:15their data analytics environment now
41:17you've won the contract to support
41:18Camila in implementing that solution so
41:21Camila is busy doing her resource
41:24planning she's thinking about who she's
41:26going to need to support support this
41:27environment in fabric she's asking for
41:30your assistance to help her structure
41:32her thoughts and also her team so let's
41:34have a look at what that looks like in
41:36fabric so this is a high level structure
41:39of a fabric implementation now you
41:42notice that it's hierarchical at the top
41:45we have tenant level so this is kind of
41:47the one tenant that you're going to have
41:49in your organization then below that you
41:51might have one or many capacities and we
41:54talked about capacities in the previous
41:56lesson we're going to going to be doing
41:57a bit more on capacities in this lesson
41:59as well then in each capacity you might
42:01have one or multiple workspaces then in
42:04each workspace we go down to the item
42:07level so you might have a data warehouse
42:09and a Lakehouse in your workspace then
42:11we can actually go one level deeper than
42:13that which is called the object level So
42:15within the data warehouse you have dbo
42:17do customer that might be a table in
42:19your data warehouse or a view in your
42:21data warehouse that's at the object
42:23level now when we're administering
42:25fabric we need to be aware of these
42:27different levels because at each of
42:28these different levels Administration
42:31happens in a different way in the last
42:33lesson we looked at the tenant level
42:35admin settings so that's mostly things
42:37in the admin portal under the tenant
42:40settings section today we will explore
42:42item level a little bit later on but
42:45first I just want to look at how we can
42:47administer each of these three top
42:49levels so here we have a table now the
42:52table isn't complete yet we're going to
42:54walk through it together so we have the
42:55three top levels we got tenant level the
42:58capacity level and the workspace level
42:59and then on the right hand side we're
43:01going to go through what the
43:02administrator or who the administrator
43:04is what role they require and also where
43:07the admin happens right so where are
43:09they going to be working where are the
43:11settings that they need to administer at
43:13each level so starting at the top level
43:15we've got the tenant admin now to get
43:18the rights to be able to be a tenant
43:20admin we're actually going to go higher
43:21than fabric we need an entra ID role of
43:24global administrator Power Platform
43:26administrator or fabric administrator
43:28and we looked at what that looks like in
43:30the last lesson and they're going to be
43:32working predominantly in the fabric
43:34admin portal so if you have any of these
43:36three roles that's going to be available
43:38to you and you can configure your tenant
43:40settings in there one level down at the
43:43capacity level well the capacity admin
43:45is assigned when you create a new fabric
43:48capacity in Azure now where the
43:50administration happens so any sort of
43:53administration at that capacity level is
43:55going to be done either in the Azure
43:57portal as we mentioned before or in the
44:00fabric admin portal there's a section on
44:02capacity settings we're going to have a
44:04look at both in a minute at the
44:05workspace level so we're going one level
44:08down now you're going to have a
44:09workspace admin they're going to be the
44:11person that is kind of in charge of that
44:13admin level and the role required here
44:16is a workspace role so it's going to be
44:18a person or a group with the workspace
44:21role of admin and again we'll have a
44:23look at what that means in more detail a
44:26bit later on now there going to be doing
44:27most of their Administration within the
44:30workspace settings and also within the
44:32manage access so these are the two areas
44:34that they're going to be focusing most
44:36of their time on now we looked at the
44:38tenant level admin settings previously
44:40in the last video in this lesson we're
44:41going to focus on the capacity level
44:44settings and then the workspace level
44:46settings so let's just focus in on the
44:48capacity administrator settings for a
44:50moment and as I mentioned there's two
44:52really places where capacity
44:55Administration gets done number one one
44:57is in Azure because we need to use Azure
45:00portal to purchase capacity right so
45:04that's going to be where you go to
45:05create a new capacity delete existing
45:07capacities changing the size of a
45:10capacity so if you've got an F2 skew and
45:13you want to go up to an F4 that's going
45:15to be done in the admin portal and also
45:17changing the capacity administrator so
45:18you can do that within the Azure portal
45:20as well as well as that within fabric we
45:24can change some of the settings for a
45:26particular Capac
45:27so we can do things like enabling
45:29Disaster Recovery viewing the capacity
45:31usage report so how much is our capacity
45:34being used we can Define who can create
45:36workspaces within that capacity we can
45:38Define who is a capacity administrator
45:41we can update the powerbi connection
45:43settings so who can connect to this
45:46capacity or items within this capacity
45:49from powerbi and how does that look like
45:51we can permit workspace admins to size
45:54their own custom spark palls so this is
45:57quite important right because you might
45:58want to set some sort of limits on the
46:00sizing of the spark poles that workspace
46:03owners underneath the capacity or in
46:06this capacity you might want to limit
46:08how high they can go with their spark
46:09poles because that's going to have quite
46:10a big impact on the overall capacity
46:13usage so if you're sitting at the
46:14capacity level you might want to add
46:16some restrictions on you know the custom
46:18spark configurations that happen in Your
46:21Capacity and we can assign workspaces to
46:24the capacity in this section as well
46:26okay so so here we are in the portal and
46:28I just wanted to show you some of the
46:30capacity settings how to administer a
46:33fabric capacity and we're going to start
46:34right from the beginning so how do you
46:36actually set up and buy a capacity well
46:38you go to Microsoft fabric if it's not
46:40already there then you can just search
46:42for it here Microsoft fabric this is
46:43going to bring you through to the the
46:45resource creation tool we're going to
46:47click on Create and we're going to walk
46:49through these steps to create a fabric
46:50capacity so you need a subscription and
46:53a resource Group and you can enter the
46:55capacity name so give it a region change
46:58the size I'm just going to do an F2
47:01select and here is where you sign the
47:02capacity administrator and again we can
47:04change that afterwards but that's you
47:06need at least one to set it up so it's
47:08going to give you the estimated cost per
47:10month here and then you press create and
47:12it's going to deploy that capacity okay
47:15so now my capacity has been created we
47:18can click on go to Resource and this is
47:20where we're going to do some of the
47:21administration tasks within the Azure
47:24portal right so here we can have a look
47:26at well firstly we can pause it so if
47:28you want to pause the capacity you can
47:30do that here delete it as well down the
47:32left hand side we've got some useful
47:34things here so capacity administrators
47:36that's where you're going to change your
47:38capacity administrator we've also got
47:40change size so if you're finding that F2
47:44is not enough for your workloads that
47:47you're running you can change it to F4
47:49or f256 if you've got a spare 40 Grand a
47:52month to be using on fabric I'm going to
47:55keep it as an F2 that's just something
47:57to bear in mind there okay now I just
47:58want to have a look at what that
47:59capacity setting looks like inside
48:02fabric so if we go to the admin portal
48:05and capacity settings if we go over to
48:07the fabric capacity tab here we can see
48:10that fabric capacity that we've just set
48:11up so it's an F2 it's in UK South and
48:14it's active so we've got some actions
48:16here we can change the name we can have
48:17a look at the admin you can't really do
48:19much there there's a link through back
48:21into Azure so if you want to make any
48:23changes to it from within fabric you
48:25need to click on this link here but if
48:26you click on the actual capacity name we
48:28go through to the capacity settings
48:31right so this is where you're going to
48:32be doing things like enabling Disaster
48:35Recovery having a look at the usage
48:37report for that capacity turning on
48:39notifications so it's going to give you
48:41a notification when you've used x amount
48:44of percent of your capacity updating who
48:46is the administrator for that capacity
48:49can also be done here changing how we
48:51can access powerbi and how powerbi can
48:54access data in the capacity we've also
48:57got data engineering settings it's
48:59mainly spark settings and this is where
49:02we're going to have that permission to
49:04permit people to change the custom spark
49:07poing right so either on or off and we
49:10can assign certain workspaces to this
49:12capacity so if you just create a new
49:13capacity it's going to be empty and you
49:15can move existing workspaces onto that
49:18capacity so as a workspace administrator
49:20we've got a few different options so
49:22here we're stepping down level into the
49:24workspace level and a workspace
49:26administrator as we mentioned deals
49:28primarily in the workspace settings and
49:31here you can edit the license for the
49:33workspace so change it from like for
49:36example a trial capacity to a fabric
49:38capacity or from Pro to PPU premium per
49:41user we can also configure connections
49:43to Azure as well as well as configuring
49:46Azure devops connections so if you want
49:48to use Git Version Control for this
49:51workspace that's going to be done in the
49:53workspace settings another thing we can
49:54do in the workspace settings is set up
49:57what's called a workspace identity now
49:59this is basically having a managed
50:02service principle dedicated for your
50:04workspace and it basically means you can
50:06connect to things like ADLs Gen 2 for
50:09things like shortcuts and you can do
50:11that in a kind of trusted workspace
50:13access manner basically this is another
50:15pretty new security feature that they've
50:17added quite recently and I'll be going
50:18through more of the security principles
50:20in more detail probably in a separate
50:22video at the end of this series CU
50:24they're quite important we can also edit
50:26some of the power settings in the
50:27workspace settings and also the spark
50:30settings so particular default
50:32environments that we might want to set
50:34up within this workspace things like
50:36that just note here that managing access
50:38is done through the the managing access
50:41section so it's slightly set it's in the
50:43same kind of area but it's not
50:44necessarily in workspace settings where
50:46we add users and add groups into our
50:49workspace we'll have a look at that in a
50:51bit more detail okay next I just wanted
50:52to walk through some of the workspace
50:55settings in a bit more detail so here we
50:57are in a workspace it's called Share Hub
50:59what we're going to be doing is clicking
51:01on this Dot and you can see that we've
51:03got two here that are useful for
51:05workspace administrators number one is
51:07managing access so this is how we're
51:08going to give people access to our
51:10workspace either person or a group we
51:13can add people in here and we can give
51:15them admin member contributor or viewer
51:17if we click on these dots again and then
51:19go through to the workspace settings
51:21this is where we're going to be able to
51:23edit some of the settings for our
51:24workspace General is you just do the
51:27image and the description and also
51:28domains if you're using domains license
51:30info so this is where you're going to
51:32change the potential capacity and the
51:34license that's being used in that
51:36workspace so this one is a trial
51:37workspace so maybe we want to actually
51:39change that to an F2 maybe we've got
51:41some fabric capacity we can select which
51:44one we want to use I'm going to be using
51:46this fabric F2 learn capacity and that's
51:48going to change the license for the
51:50workspace we've also got connecting to
51:53Azure connecting to git downloading
51:55things like the file explor
51:57as well and enabling caching for
51:59shortcuts is another workspace setting
52:01we've got here managed identities so if
52:03you're on an f64 capacity or higher you
52:06can make use of workspace identities and
52:09that's going to basically allow you to
52:10create kind of like a managed service
52:12principle just for this workspace so
52:14give your workspace an identity and
52:16allow it to connect to ADLs Gen 2 create
52:20your shortcuts things like that in a
52:22secure manner kind of trusted workspace
52:24access is what it's called and if you
52:26want to learn more more about that I'll
52:27leave a link to the workspace identity
52:29section in the school Community we can
52:31also do things like adding private
52:33endpoints and that kind of thing for
52:34connecting via spark to things in Azure
52:37you've got your spark settings down here
52:40for configuring something about the pool
52:43that you're using the spark pool that is
52:45being used in this workspace you can
52:47change the default environment that's
52:49where you're going to be going to add
52:51libraries and things like that so if you
52:53want to pre-install python packages onto
52:55your spark cluster so that every time
52:58you run a notebook or start a new
53:00notebook you have those libraries there
53:02ready to go that's where you do this
53:03change some settings for high con
53:05currency and that kind of thing there so
53:07Camila says okay great I now have some
53:09clarity on administering fabric at the
53:12tenant capacity in the workspace level
53:14what I'm not sure about is giving access
53:16to the people on my team so that's
53:18important right we build all these
53:20things in fabric but how can we give
53:22people access the right amount of access
53:25to these items that's a look at that in
53:27a bit more detail so if we go back to
53:28our structure of a fabric implementation
53:32generally when we're sharing items with
53:35people that's really done at these
53:37bottom three levels of our fabric
53:40architecture right so sharing things in
53:43fabric is normally done at these three
53:45levels now object level sharing is
53:48possible for the data warehouse and the
53:50seal endpoint in the lakeh house but I
53:52don't think it's assessed as part of
53:53this exam so it's not part of the study
53:56guide in in anyway so we're not going to
53:57be covering that in this lesson there's
53:59some documentation on the Microsoft
54:01learn website and I'll leave a link to
54:02that if you're interested in object
54:04level sharing if you want to bit learn a
54:06bit more about that we're going to be
54:08focusing on the workspace level sharing
54:10and item level sharing so let's just
54:12start with workspace level sharing
54:15people or groups can be given workspace
54:18level access and when sharing the
54:21personal group is assigned a workspace
54:23role as you can see on the right hand
54:25side there we've got admin member
54:27contributor and viewer now this role
54:30applies to all items in the workspace
54:33for example a viewer in the workspace
54:35will be able to view all of the items in
54:37the workspace let's just take a bit of a
54:39moment cuz roles are really important
54:42and the role that you assign someone
54:44dictates basically what they can do in
54:46your workspace this image here comes
54:48from the Microsoft documentation again
54:51I'll leave a link to this in the school
54:53community and I definitely recommend you
54:55take some time to study it this is what
54:57we're going to be doing here so the
54:58first thing to note with this diagram is
55:01let's just start with the admin so if
55:03you give someone an admin permission
55:05what can they do the first thing is that
55:06they can update and delete the workspace
55:08so this is really high level permissions
55:11that only maybe one or two people really
55:13should have people that you trust in
55:14your organization they can also add and
55:17remove people including other admins so
55:20it's the only role that allows you to
55:22add an admin add another admin next we
55:24move down to the member and the member
55:27can do similar things to an admin but
55:30they can't add an admin okay so a member
55:33cannot add an admin they can only add
55:35people with lower permissions or other
55:38members okay the other permission that
55:40is unique at the member level is you can
55:43give other people the permission to
55:45share items so being able to share items
55:48is a fairly high level thing to do and
55:50you're giving people that permission to
55:51share okay so that's something to bear
55:54in mind as well then we move down to the
55:56contributor level now contributors can
55:59do pretty much everything in the
56:01workspace other than as we see here
56:03deleting the workspace adding other
56:06people into the workspace and allowing
56:08other people to share but they can do
56:09when you're talking about contributing
56:12to fabric items anything around lake
56:15houses or warehouses or data pipelines
56:18they have read and write access to all
56:20of these things so the viewer has a
56:22unique set of missions those six green
56:25ticks
56:26and if we go kind of from top to bottom
56:28they can view and read content in a data
56:31pipeline a notebook spark job definition
56:33machine learning model so they can view
56:35kind of the outputs of these things they
56:37can also View and read the content of
56:39kql databases query sets and real time
56:42dashboards they can connect to the SQL
56:44analytics endpoint of a Lakehouse or a
56:47data warehouse and they can read
56:49Lakehouse and data warehouse data and
56:51shortcuts with tsql so the viewer can
56:54basically use SQL to analyze data in
56:57either The Lakehouse or the data
57:00warehouse what they can't do is access
57:03any of the one lake apis or spark so
57:06they can't run spark jobs or notebooks
57:10or anything like that now one unique
57:11thing about the viewer permission is in
57:14the data pipelines now they can't edit
57:16or update any of the activities in a
57:19data pipeline but they can execute and
57:23cancel the execution of a data pipeline
57:25run so that's an important kind of edge
57:28case to remember for the exam and they
57:30can also finally view the output of data
57:33pipelines notebooks and machine learning
57:35models so that's kind of a high level
57:37overview of all of these workspace roles
57:40and what they can do at each level again
57:42this is really important to understand
57:44for the exam so I definitely recommend
57:46going into the documentation taking some
57:48time to understand these different
57:49things because you'll probably be tested
57:51quite a lot on these so let's just have
57:52a bit of a workspace level access
57:55example this is John he is a business
57:58analyst working in Camila's team and
58:01Camila has asked you to give him
58:03contributor access to workspace one this
58:06is workspace one this is the
58:08architecture that they've got here now
58:09this is what John's access currently
58:11looks like where the red box is
58:13basically no access at all and a green
58:15box if there's any access and you can
58:17see everything's red So currently has no
58:18access to anything you are an admin in
58:21the workspace now what steps would you
58:23take to give John this access have a
58:25look little think about that and then
58:27we'll talk about it in a second okay so
58:29what steps would you take well
58:31personally what I would be asking is
58:33does John fit into an existing security
58:36group that has contributed access to the
58:39workspace because best practice here is
58:41to add people into groups rather than
58:43adding them individually just makes
58:44maintenance in the future a lot easier
58:47where possible we always want to add
58:49people into groups before we add them
58:51individually now if a security group
58:54doesn't exist then you might want to
58:55create one for John maybe you want to
58:57create an analyst security group so that
58:59in the future when another analyst wants
59:01to join the team or join the workpace
59:04you can just add that person into the
59:06group rather than having this long list
59:09of individual contributors in that
59:11workspace so you create an analyst
59:12Security Group add John to the security
59:15group and give the group contributor
59:17access to workspace one so in the
59:19picture how does that update well it
59:21looks a bit like this right so John now
59:23has access to workspace one and
59:26everything within it because we've given
59:28him access the security group access at
59:31that workspace level you'll notice that
59:33workspace 2 he still has no access to
59:36that he can't even see that so that's
59:37something to bear in mind when you're
59:39giving workspace level access Camila
59:42suddenly Rings you she realizes that JN
59:44shouldn't have access to everything in
59:46the workspace instead she wants you to
59:49give him access to the data warehouse
59:51only not the semantic model not the data
59:54pipeline so how would you check change
59:56what we've just done to reflect this so
59:59this is what we're we're looking at here
1:00:00we want to go from this which is the
1:00:02arrangement that we've just done for
1:00:04John at the workspace level to this at
1:00:07the item level now this might be
1:00:08important because it kind of reflects
1:00:11quite an important principle when it
1:00:13comes to giving people access which is
1:00:15the principle of least privilege now in
1:00:18general in Data Systems information
1:00:21security we want to give people the
1:00:23amount of access that they need to
1:00:25perform their roles and nothing more
1:00:27right so if you don't technically need
1:00:29access to the semantic model or the data
1:00:31pipeline then one way of kind of getting
1:00:34around that is to give people item level
1:00:36access giving people access to only what
1:00:38they
1:00:40need okay just to recap on some of the
1:00:44additional permissions so when you share
1:00:46a data warehouse you get these three
1:00:49additional permissions we have read all
1:00:51data using SQL and what that means is it
1:00:55allows people to read all objects within
1:00:57the warehouse using tsql we also have
1:01:00the read all one L data and with this
1:01:03you're allowing that person to read the
1:01:05underlying oneel files using spark
1:01:08pipelines anything else basically so in
1:01:10the top one they can only use SQL if you
1:01:13give them the second permission it
1:01:14allows them to basically do anything
1:01:15with that data and the third permission
1:01:18allows the user to build reports on the
1:01:20default semantic model not any custom
1:01:22models just the default semantic model
1:01:24when it comes to the lak house these
1:01:25permissions are similar but they're just
1:01:27worded a little bit differently so again
1:01:29if you give them the read all SQL
1:01:31endpoint data it allows them to perform
1:01:34tsql on the tsql endpoint if you give
1:01:37them read all Apache spark then again
1:01:41it's going to allow them to run
1:01:42notebooks and Spark code on top of that
1:01:44data and again the build reports on the
1:01:47default semantic model does exactly what
1:01:49it says on the tin now one point that I
1:01:51just did want to make here is around one
1:01:53Lake data access model now this is a
1:01:56very newly announced feature so it might
1:01:59not have made its way into the exam yet
1:02:01but I do think it's going to have a very
1:02:02big impact on how we manage security in
1:02:06fabric going forward so I did want to at
1:02:08least mention it here I'm not going to
1:02:10be going through it in detail I did just
1:02:12want to flag it you might want to have a
1:02:13look at the documentation page just so
1:02:16that you become aware of it now this
1:02:18feature is not really something I looked
1:02:20at yet much in detail but from just from
1:02:22looking at the documentation from how I
1:02:24understand it is it's going to allow you
1:02:25to perform rback so Ro based access
1:02:29control on things like folders so now
1:02:32that they've implemented folders within
1:02:33a workspace it's going to allow you to
1:02:35define a specific role or give a
1:02:38specific role access control over that
1:02:41folder and then the permissions are
1:02:42going to be inherited for every item in
1:02:44that folder but like I mentioned it's a
1:02:46preview feature and it's relatively new
1:02:48so I'd be surprised if they ask you
1:02:50about this in the exam but it is very
1:02:52important I do think it will change
1:02:53quite a lot in fabric so I wanted to
1:02:55mention it and if you look at the study
1:02:57guide it does say for the dp600 exam it
1:03:00does say that most questions cover
1:03:02features that are generally available
1:03:03the exam may contain questions on
1:03:05preview features if those features are
1:03:07commonly used I think at the moment this
1:03:09isn't commonly used because it's only
1:03:10been released a few weeks ago so that's
1:03:12something to bear in mind Camila says
1:03:14thank you now I understand workspace
1:03:16level and item level sharing in more
1:03:18detail One Last Thing Before You Go
1:03:20we've been working on this government
1:03:22project and I need to apply sensitivity
1:03:24labeling in a workspace can you walk me
1:03:27through it so what even is a sensitivity
1:03:29label well sensitivity labels are a data
1:03:32governance feature and they're created
1:03:35and managed in Microsoft purview so
1:03:38fabric items such as a semantic model
1:03:40can be given a sensitivity label such as
1:03:43confidential right and it's for
1:03:45information protection purposes now in
1:03:47some Industries labeling data and
1:03:49information with a sensitivity label is
1:03:52necessary for compliance with
1:03:54information protection regulations now
1:03:56to apply a sensitivity label in fabric
1:03:59really there's two main methods if we go
1:04:02into that item for example this
1:04:04Lakehouse here what you have in the top
1:04:06tool bar you've got the sensitivity
1:04:08label and you can just click on that
1:04:09drop- down change sensitivity label in
1:04:12there the other option is to go into the
1:04:15settings of that particular fabric item
1:04:17and you can see that in the left hand
1:04:19tool bar there you've got sensitivity
1:04:21label and you can change the sensitivity
1:04:22label in there now one of the options
1:04:25that you can give it if you go through
1:04:26the settings method is to apply to
1:04:28Downstream items so again we've got that
1:04:31notion or that concept of inheritance of
1:04:34the label that you give it here also
1:04:36applies to everything Downstream okay so
1:04:38now we are going to test some of your
1:04:39knowledge for everything that we've
1:04:41learned in this section of the study
1:04:44guide and we're going to start with a
1:04:46case study style question so we're going
1:04:48to going through a bit of a case study
1:04:50and then going to be asked three maybe
1:04:52four questions on this particular case
1:04:54study let's again Toby creates a new
1:04:57workspace with some fabric items to be
1:04:59used by data analysts Toby creates a new
1:05:03security group called Data analysts he
1:05:05includes himself as a member of this
1:05:07Security Group Toby gives the data
1:05:09analyst Security Group a viewer role in
1:05:13the workspace what workspace role does
1:05:15Toby have is it a viewer B member C
1:05:18admin or D contributor pause the video
1:05:21here have a think and then we'll move
1:05:23forward to the answer so the answer here
1:05:25is see now this combines two pretty
1:05:27important Concepts to understand when
1:05:29we're looking at workspace level sharing
1:05:32number one is that the creator of a
1:05:35workspace is always given admin
1:05:37permissions in that workspace now we
1:05:39also have Toby with the viewer role in
1:05:41that workspace CU he's in the security
1:05:43group with viewer role and this is
1:05:46another concept if you have more than
1:05:47one level of permission within the
1:05:49workspace you're always given the higher
1:05:51level so he's got admin role cuz he
1:05:53created the workspace and he's got
1:05:55viewer role because he's in that
1:05:56Security Group Well the admin
1:05:58permissions is always going to be
1:05:59prioritized he's always going to take
1:06:01that role over his viewer role so let's
1:06:03continue this case study Sarah is also a
1:06:07member of that data analyst Security
1:06:09Group she has no other role in the
1:06:11workspace which of the following can
1:06:13Sarah not do in the workspace a execute
1:06:16a data pipeline run SQL scripts in the
1:06:19data warehouse run spark notebook or
1:06:22review the evaluation metrics of a
1:06:24machine learning model now the answer
1:06:25here is C run the spark notebook now
1:06:28when we're looking at the workspace
1:06:30level roles and the permissions for each
1:06:32role we know that Sarah is a viewer in
1:06:34the workspace that's the highest level
1:06:36of permission and you remember that a
1:06:38viewer role can actually execute a data
1:06:41pipeline in a workspace they can also
1:06:44run tsql scripts in data warehouse or a
1:06:47SQL endpoint of a lake house what they
1:06:49can't do is run a spark notebook okay so
1:06:52anything in a notebook they can have a
1:06:54look at the notebook but they can't
1:06:56actually execute any code so C is the
1:06:58right answer cuz what we're looking for
1:07:00is what can she not do in the workspace
1:07:03and D is review the evaluation metrics
1:07:05of a machine learning model which we
1:07:07know we can do because she's just
1:07:08reading the output of that model to
1:07:11continue this case study again Toby
1:07:13wants to delegate some of the management
1:07:15responsibility in the workspace he wants
1:07:17to give this person the ability to share
1:07:20content within the workspace invite new
1:07:23contributors to the workspace but not
1:07:25add new admins to the workspace what
1:07:27role should Toby give this person a
1:07:29admin B member C contributor or D viewer
1:07:33so the answer here is B now the the key
1:07:36point in the question was but not add
1:07:38new admins to the workspace so we know
1:07:41that to be able to add another admin
1:07:42into a workspace you need to have admin
1:07:44permissions yourself so Toby doesn't
1:07:46want to give that person this ability
1:07:48basically so we know it can't be admin
1:07:50it's not going to be viewer it's not
1:07:52going to be contributor we know that the
1:07:54member is kind of one down from admin
1:07:57and that's going to allow you to do all
1:07:58of these three things they can share
1:08:00content they can invite other
1:08:02contributors because a member can add
1:08:04new people either members or
1:08:06contributors or viewers but they can't
1:08:08add other admins so it' be B member the
1:08:11next question is completely separate you
1:08:13have admin role in a workspace Sheila is
1:08:16a data engineer in your team she
1:08:18currently has no access to this
1:08:20workspace at all now Sheila needs to
1:08:22update a data transformation script in a
1:08:24pisb notebook and the script gets data
1:08:27from a Lakehouse table cleans it and
1:08:29then writes it to a table in the same
1:08:30Lakehouse now you want to adhere to the
1:08:33principle of leas privilege what actions
1:08:35should you take to enable this is it a
1:08:38you're going to give Sheila the
1:08:39contributor role in the workspace b
1:08:42share the Lakehouse item with read or
1:08:44spark data permission C give Sheila the
1:08:47admin role in the workspace or D share
1:08:50the lake house item with read all spark
1:08:52data permissions and share the notebook
1:08:54with edit permissions so the answer here
1:08:56is D so one of the clues in this
1:08:58question was the line where it says you
1:09:01want to adhere to the principle of leas
1:09:02privilege so immediately when you see
1:09:04that giving people workspace level
1:09:07access is not really good enough so A
1:09:09and C is giving a role in the workspace
1:09:13so it's going to enable her to
1:09:14contribute and change and edit
1:09:16everything in the workspace but it
1:09:18doesn't adhere to the principal of least
1:09:19privilege so we can immediately rule out
1:09:21a and C so another really important
1:09:25point in the question here was Sheila
1:09:27needs to update a data transformation
1:09:30script in a notebook so she needs to
1:09:32edit the code in a notebook She's Not
1:09:34Just executing an existing notebook she
1:09:36needs to actually make changes to a
1:09:38notebook and so for B you wouldn't have
1:09:41that permission you've got read all data
1:09:43for the spark so you can actually
1:09:45execute a notebook but you can't edit a
1:09:47notebook so be able to make these
1:09:49changes really need access to the
1:09:51notebook and The Lakehouse that that
1:09:54notebook is interfacing with that that
1:09:56notebook is reading from because you
1:09:58can't just share the notebook because
1:09:59then you won't have access to the
1:10:00underlying data and we can't just share
1:10:02the lake housee because you won't have
1:10:03access to the notebook that she needs to
1:10:05edit so the answer is D share the
1:10:07Lakehouse item we're giving spark
1:10:09permissions and we're also giving edit
1:10:11permissions on the notebook next
1:10:13question you have admin role in a
1:10:15workspace you want to pre-install some
1:10:17useful python packages to be used across
1:10:20all notebooks in the workspace how do
1:10:22you achieve this a in the fabric ad
1:10:25admin portal go to spark settings and
1:10:27install the libraries B go to workspace
1:10:29settings spark settings and then Library
1:10:32management C create an environment
1:10:34install the packages in the environment
1:10:36go to the workspace settings spark
1:10:38settings and set it as the default
1:10:40environment or D go to capacity settings
1:10:43and then default libraries so the answer
1:10:45here is C creating an environment and
1:10:49then going into your workspace settings
1:10:50and setting it as the default
1:10:53environment for for spark now now this
1:10:55is a bit of a a naughty question because
1:10:57B is the old way so it used to be you go
1:11:00to workspace settings spark settings and
1:11:02there was a section for Library
1:11:03management but that's actually not
1:11:05possible anymore the way to do it as I
1:11:07mentioned is to create an environment
1:11:09and then in your spark settings make it
1:11:10the default environment A and D don't
1:11:13actually exist these capabilities so
1:11:16these kind of red herrings so the answer
1:11:18is C Camila says thanks she's seriously
1:11:20impressed with your knowledge again in
1:11:22this lesson we covered all five of these
1:11:25elements of the dp600 study guide from
1:11:28workspace and item level sharing data
1:11:30sharing for data warehouses and lake
1:11:33houses sensitivity labeling and then
1:11:36workspace and capacity level settings
1:11:39and the good news is again you've won an
1:11:40extension to the contract Camila would
1:11:42like you to implement control over the
1:11:45entire analytics development life cycle
1:11:48in her organization so for this we're
1:11:50talking Version Control deployment
1:11:52pipelines powerbi projects all that good
1:11:55stuff that's what we going be looking at
1:11:56in the next lesson so click here to
1:11:59continue that lesson hey everyone
Manage the analytics development lifecycle
1:12:01welcome back to the channel today we're
1:12:03continuing the DP 600 series looking at
1:12:06what it's going to take to hopefully P
1:12:08that exam and become a fabric certified
1:12:11analytics engineer Today Is video 4
1:12:14we're going to be looking at managing
1:12:16the analytics development life cycle and
1:12:18in this section exam we're going to be
1:12:20focusing on implementing Version Control
1:12:23creating and managing powerbi projects
1:12:26planning and implementing deployment
1:12:27Solutions performing impact analysis on
1:12:31Downstream activities deploying and
1:12:33managing semantic models through the
1:12:35xmla endpoint and creating reusable
1:12:38assets so powerbi template files powerbi
1:12:41data source files all these kind of
1:12:42things so that's what we've got in store
1:12:44for you today as ever there's going to
1:12:46be some sample questions at the end I'll
1:12:48also be posting all of the lesson notes
1:12:52and links for further resources in the
1:12:54school community go and grab them there
1:12:56you will play the main character in the
1:12:58scenario again we're going to be
1:12:59continuing this theme this scenario that
1:13:02we've been developing throughout this
1:13:03course you're the main character are you
1:13:05ready let's begin so as you know already
1:13:07we are a consultant working with the
1:13:10client called Camila and you've already
1:13:12helped Camila plan and Implement her
1:13:14fabric implementation her environment
1:13:17now the time has come to implement
1:13:19analytics development life cycle now she
1:13:21wants to focus on Version Control at
1:13:23least initially but to be honest he's
1:13:25not really familiar with Git Version
1:13:28Control never really used that in an
1:13:29analytics environment before so what
1:13:31you're going to have to do is set up a
1:13:34call with Camila and walk her through
1:13:36the basics of git first and Version
1:13:38Control what is it why does it even
1:13:40exist why is it now a part of fabric and
1:13:44what are some of the key terms and
1:13:45terminology that we need to understand
1:13:47when we're implementing Version Control
1:13:49just bear in mind that all of the git
1:13:52integration features are currently in
1:13:54preview and some of the other features
1:13:56we're going to be looking at today like
1:13:57the deployment pipeline functionality a
1:14:00lot in this kind of analytics life cycle
1:14:02stuff is still in the previous stages so
1:14:04just bear that in mind when we're
1:14:06walking through these examples I'll show
1:14:07you what's possible today but if you're
1:14:09coming from using like GitHub or git a
1:14:12lot in the past and a lot of the
1:14:14features that you might be used to are
1:14:16currently not available so let's move on
1:14:19okay so I just wanted to start this
1:14:21video with a bit of a primer on git and
1:14:24Version Control in general and I think
1:14:26the best way to do that is to show you
1:14:28what this looks like and along the way
1:14:30we can explain some of the key Concepts
1:14:32in git and Version Control I realize
1:14:35that probably a lot of people who have
1:14:36been coming from maybe a powerbi
1:14:38background don't have experience with
1:14:40Version Control and get in general so
1:14:42let's set up a project and start at a
1:14:45basic level introduce more and more
1:14:47Concepts as we go through this demo so
1:14:50in fabric the way you're going to start
1:14:52is with Azure devops now AZ devops is a
1:14:55service that's used primarily for
1:14:57software development and managing
1:14:59infrastructure and the devops life cycle
1:15:02of that infrastructure and software
1:15:04development but it's also where we're
1:15:06going to use to store our code and our
1:15:09artifacts and our powerbi project files
1:15:12because they come with repositories so
1:15:14you need to set up an account just go to
1:15:16this website here and I'll leave a link
1:15:18in the description or in the school
1:15:19Community now once you've set up Azure
1:15:21devops for your organization you'll come
1:15:24through to this page here I've set up an
1:15:27organization called fabric University
1:15:29and then you have this concept of a
1:15:31project so a project is going to be
1:15:33where you store your repositories your
1:15:35code has lots of other features as well
1:15:37but we won't be going into too many of
1:15:39the other features in this short video
1:15:41so we're going to be creating a new
1:15:42project for this demo let's just call it
1:15:45dp600 practice we're going to make it
1:15:48private and then just click on create
1:15:50okay so this is an Azure devops project
1:15:53on the left hand side you can see here
1:15:55make that a little bit bigger on the
1:15:56left hand side we've got some useful
1:15:59things to know about right so this is
1:16:01just the overview where you can see an
1:16:02overview of the project probably the
1:16:04most important ones to bear in mind are
1:16:06boards so this is like a Work Management
1:16:08tool for doing development work we're
1:16:11not going to be focusing much on that in
1:16:12this video what we're interested in is
1:16:14this Repose repositories right so
1:16:17repositories is kind of the core thing
1:16:19that we need to implement Version
1:16:22Control now repositories are also
1:16:24available in GitHub but currently that
1:16:26integration is not possible between
1:16:28Fabric and GitHub you can only use Azure
1:16:31devops repos so that's why we're using
1:16:34this currently so a repository is just
1:16:36where we're going to store our code in
1:16:39the cloud so when you're developing
1:16:41locally like a powerbi report for
1:16:43example really what we want to do is
1:16:46store that report in the cloud in our
1:16:50repository and that's where we're going
1:16:51to be controlling tracking changes
1:16:54managing who can make changes to that
1:16:56project in the repository so let's just
1:16:59do a bit of configuration here to set up
1:17:00this repository so that we can use it
1:17:02for Version Control so the first thing
1:17:04I'm going to do is right at the bottom
1:17:06here initialize a main branch with a
1:17:08readme and a readme file is just a
1:17:11markdown file it's like a text file just
1:17:13to initialize our branch and we'll talk
1:17:16a bit more about branches in a bit more
1:17:18detail here currently I have one branch
1:17:20and it's the main branch here now
1:17:22branching is kind of like a whole field
1:17:24in itself there's lots of different
1:17:26strategies for how we can manage
1:17:28different branches when we developing
1:17:30our code or developing our fabric
1:17:33artifacts but we'll get into that in a
1:17:34bit more detail a bit later on for now
1:17:37what we're interested in really is
1:17:39creating a copy of this repository on
1:17:43our local machine right so that's going
1:17:45to be the first sync we want to get
1:17:47these two in sync because the way that
1:17:50version control works in general is you
1:17:53develop locally on your local machine
1:17:55for powerbi and then you sync those
1:17:57changes up to the repository in the
1:18:00cloud and if you have a big team of
1:18:02people maybe you have 10 people doing
1:18:03this they're all syncing their changes
1:18:06into this main branch by default so what
1:18:09we want to do is clone this repository
1:18:12and we get this URL here for this git
1:18:14repository we can copy the URL and what
1:18:17we want to do is clone that on our local
1:18:19machine so there's various different
1:18:20ways to do that you might see in the
1:18:22documentation Microsoft use a tool
1:18:25called vs code I personally prefer using
1:18:28GitHub desktop this is another tool it's
1:18:30completely free you can download it just
1:18:33Google GitHub desktop download it's
1:18:35built by obviously GitHub but it's also
1:18:38possible to do Azure devops repos in
1:18:41here as well and it just makes the whole
1:18:43process of Version Control and git a lot
1:18:47simpler it basically builds a UI on top
1:18:49of a lot of the git functionality so the
1:18:52first thing we want to do is click on
1:18:53file and clone repository because we're
1:18:56trying to drag things down from the
1:18:59internet and if you click over to URL we
1:19:01can just copy the URL that we've got
1:19:03here that we got from our azid Devo
1:19:05repository and we can clone it now it's
1:19:07going to ask for a username and password
1:19:09so if you go back to a devops we've got
1:19:12generate this git credentials and we can
1:19:15just copy the username in there copy the
1:19:18password in there save and retry so
1:19:20that's just going to authenticate with
1:19:22our Azure devops account so we know that
1:19:24okay this is the actual person that owns
1:19:26that repository it's creating that
1:19:28authentication between the two okay so
1:19:30what has that actually done well now
1:19:32what we've got is a copy of what we have
1:19:35in Azure devops in our local machine you
1:19:38can see the file path there is C users
1:19:41learn documents git and what we can
1:19:44actually do is open this in Explorer
1:19:47okay so now I have a copy of those files
1:19:50and folders from Azure devops on my
1:19:53local machine so that is the first first
1:19:54thing that we need to do really is clone
1:19:57the repository get that copy local and
1:19:59what's going to happen now we've got
1:20:01this all set up and synced so that
1:20:04GitHub desktop is Now tracking this
1:20:06folder so any changes you make locally
1:20:09to this folder it's going to pick those
1:20:11up right so for example if I do a new
1:20:14text file call it my file now if we go
1:20:16back to get up desktop you can see that
1:20:18it's picked up that change already it's
1:20:20saying you've added a new file in there
1:20:23it's called my file . text and if I open
1:20:26that with notepad for example just call
1:20:28it my amazing file and I'll save that
1:20:31here we can see again it's tracked to
1:20:32the changes so git works by tracking the
1:20:35changes in text based documents okay so
1:20:40traditionally it's used for code so a
1:20:43python file a SQL script or a Javascript
1:20:46file or C all these kind of things are
1:20:49textural based right so they work really
1:20:52well with Git so now we've made some
1:20:54Chang changes to our local repository
1:20:57right we've added this text file but
1:20:59what you'll notice is that you know here
1:21:01in Azure devops nothing has changed yet
1:21:03because what we need to do is push those
1:21:05changes into the cloud environment maybe
1:21:09you've got a team you want other people
1:21:10to see the changes that you've made and
1:21:12we'll start by using this text file
1:21:15right and then we'll slowly move up to
1:21:16more complicated stuff we'll look at
1:21:18powerbi projects as well but for now
1:21:21let's just push these changes so you see
1:21:23down here you can do created a new file
1:21:26so created a new file and then what
1:21:28we're going to do is click this commit
1:21:29to main button so this is going to
1:21:32commit our changes onto that main branch
1:21:35and up here we click on push to origin
1:21:38so we're going to commit and then push
1:21:40the changes up into the cloud right into
1:21:42Azure devops and so now what's going to
1:21:44happen if we refresh our Azure devops we
1:21:47can see that file there we can see my
1:21:49file. text my amazing file that's in
1:21:52there in this repository in azure devops
1:21:55now Okay then if we want to make more
1:21:57changes maybe we want to change this a
1:21:59bit more even more amazing file and we
1:22:01save that again it's going to show you
1:22:03oh well your original file was this okay
1:22:06this is what we've got stored in Azure
1:22:08devops this is what the the main branch
1:22:11is currently saying now you've changed
1:22:13that file right and we've noticed that
1:22:15okay so it's tracking the changes
1:22:17between every change you make in that
1:22:19file so that's really important thing to
1:22:21bear in mind so let's just update this
1:22:24text file obviously in practice in
1:22:26proper environments you add a more
1:22:28meaningful commit message because these
1:22:29are really important and we'll push that
1:22:31again to the origin using this button
1:22:33here so again if we just go back to the
1:22:36Azure devops you can see that that has
1:22:38now updated in here as well great so
1:22:40that is the very very basic
1:22:42implementation of git and Version
1:22:45Control and syncing our local repository
1:22:48with our Azure devops repository in the
1:22:50cloud so at this stage I think it's
1:22:52worthwhile just noting what we we've
1:22:54done here right so at the top we've
1:22:56created an A devops account we've added
1:22:59that readme file we've copied the git
1:23:01URL and cloned it locally and then we've
1:23:04modified that local folder create a new
1:23:06file modified the file committed those
1:23:09changes back pushed the origin and
1:23:11observe the changes in as devops now
1:23:13this is great it's a good first step
1:23:14right we can track changes now between
1:23:16files but one of the core benefits of
1:23:19git and Version Control in general is
1:23:21that instead of just allowing anyone to
1:23:23update that main branch what we can do
1:23:26is we can protect that main branch so
1:23:29that we can control who has access to
1:23:32and who has the ability to overwrite
1:23:35those changes so let's make a few
1:23:37changes to the figuration of our
1:23:40repository in Azure devops what we're
1:23:42going to do is we're going to protect
1:23:44that main branch so that nobody can just
1:23:46overwrite that file anymore we want to
1:23:49add in a review and approval phase and
1:23:52this is going to help a lot with our
1:23:54quality and control over who and how
1:23:58that code base is being changed so let's
1:24:00have a look at that in a bit more detail
1:24:02okay so in our Azure devops project here
1:24:05we can go to Project settings and then
1:24:07we can go down to repositories and then
1:24:09what we're going to do is add in a
1:24:12policy for branch policies protect
1:24:15important branches Nam spaces so we're
1:24:17going to add something in here protect
1:24:19the default Branch create automatically
1:24:21included reviewers so I'm going to add
1:24:23myself in here as a reviewer just
1:24:26because there's only one person in this
1:24:28project in reality you might have a team
1:24:30of maybe senior developers or lead
1:24:32developers that responsible for
1:24:35reviewing code and approving changes to
1:24:38the the code base so now we've set that
1:24:40up let's have a look at doing something
1:24:42a bit more so we've got our file here so
1:24:45if I save this I've just added another
1:24:47line in here new changes and if we go
1:24:49back to our GitHub here we can see that
1:24:52oh it's picked up that new changes again
1:24:54but now what I'm going to do is update
1:24:58this so added some new changes and I'm
1:25:01going to commit that and when I try to
1:25:02push this to the origin is going to say
1:25:04error okay because now pushes to this
1:25:07Branch are not permitted you must use a
1:25:09poll request to update this Branch so
1:25:11that's going to introduce a few new
1:25:13Concepts that we need to get around this
1:25:17really number one is the concept of
1:25:19branches and number two is the concept
1:25:21of pool requests let's have a look at
1:25:23both of those now now so now that we've
1:25:25implemented some protection over that
1:25:28main branch we have to change how we
1:25:30develop okay so now we have to use
1:25:34branching So currently all of my files
1:25:37that text file and the readme file is on
1:25:40this main branch so if I now want to
1:25:42update what's in that text file I'm
1:25:45going to have to create a new Branch
1:25:46because nobody can update and edit that
1:25:49main branch directly so what we can do
1:25:51in GitHub desktop it's very easy so what
1:25:54we're going to do is click on Branch New
1:25:57Branch now we can change the name to
1:25:59change text file create this branch and
1:26:02now if we publish this branch and we can
1:26:03add in some changes here so now if we go
1:26:06back to Azure devops let's just go back
1:26:08to our project here and back to Repose
1:26:11and if we go down to this change text
1:26:14file you can see that it has actually
1:26:16found this new Branch this is a branch
1:26:19and if we go over to the pull requests
1:26:21section we'll explain what that means in
1:26:22a minute but you can see that it's
1:26:24registered the fact that we have now
1:26:26published a new Branch so we've made
1:26:29some changes added that new changes line
1:26:31in the text file and it's popped up
1:26:33saying oh do you want to create a pool
1:26:35request here and so the pool request is
1:26:37when we want to make a change or merge
1:26:40some changes into the main branch okay
1:26:43and you can only do this through a
1:26:45review because that's the policy that
1:26:46I've set up on that main branch so what
1:26:49we can do is create a PO request you can
1:26:51say oh I've added some new changes in
1:26:53here you can add in a reviewer which is
1:26:55going to be me and you can say oh please
1:26:57review this and you can create a pull
1:27:00request so what that's going to do is
1:27:01going to notify the reviewer and say oh
1:27:05this person wants you to review their
1:27:07changes to the code base on this change
1:27:09text file branch and now as the reviewer
1:27:12because I'm kind of playing both roles
1:27:14here I can have a look at the file
1:27:16changes I can review the changes like oh
1:27:19yep added new changes that looks good to
1:27:21me I'm going to approve that change and
1:27:23now what you can do is complete so what
1:27:25complete is going to do is merge those
1:27:29changes into the main branch so your
1:27:31protected Branch because we've gone
1:27:33through the reviewer process the text
1:27:36file has been reviewed now we can merge
1:27:38it safely into the main branch there's a
1:27:41few options here that are quite
1:27:43important to bear in mind and we'll look
1:27:46at those when we move into fabric we'll
1:27:48look at those in a bit more detail but
1:27:50generally kind of the default settings
1:27:52here complete Associated work items
1:27:54after merging it's fine we don't want to
1:27:57delete change to text file after merging
1:28:01generally in Fabric and we'll get to why
1:28:03that is in a minute we'll just do on
1:28:05complete merge Okay so we've now merged
1:28:07our poll request let's have a look at
1:28:10the repository so now we have my text
1:28:13file in the main branch and it's got new
1:28:15changes Okay so we've merged our changes
1:28:18into this main branch great that's the
1:28:21first section of this demo done okay and
1:28:24this second half there what we did was
1:28:25protect the main branch Okay so we've
1:28:28added an approv in the repository we've
1:28:31updated the repo approv policy we tried
1:28:35committing a new change and it was
1:28:36rejected we've added a new feature
1:28:38Branch we brought the changes onto that
1:28:40Branch we've committed that branch and
1:28:43we've added a p request then our
1:28:45approver has approved it and we've
1:28:47merged it into the main branch and just
1:28:49to reiterate the benefits of that is
1:28:51that now we're checking the changes
1:28:53between files and we're also controlling
1:28:55who can make changes so we're protecting
1:28:58that main branch and saying okay if you
1:28:59want to update this it has to go through
1:29:01an approval process now up until this
1:29:03point we've just been using one text
1:29:06file actually git can track the changes
1:29:09between any text based file format text
1:29:13base that's the the problem that existed
1:29:16with PB files PX files kind of a bit of
1:29:19a black box right if you can't open a
1:29:22file in notepad and understand what's
1:29:24going on and it's going to struggle with
1:29:26Source control Version Control in git
1:29:29now that is one of the core reasons why
1:29:31Microsoft developed the powerbi project
1:29:35file format because it represents a
1:29:37powerbi project in text based files
1:29:41right so let's just have a look at an
1:29:43example of a pbip first and then we'll
1:29:46look at how we can integrate that into
1:29:48Version Control Systems okay so here we
1:29:49are in powerbi desktop and I've just got
1:29:51one of the sample reports here from
1:29:54Microsoft the competitive marketing
1:29:56analysis report now what we're going to
1:29:58do is this is just a PB file at the
1:30:00moment what we're going to do is save
1:30:03this as a pbip a powerbi project file
1:30:07then we're going to explore that and
1:30:08have a look at what that looks like so
1:30:10it's fairly easy to save your powerbi
1:30:13file as a powerbi project file so if we
1:30:16go to file and then save as and what we
1:30:20can do is we can find our git reposit so
1:30:24that's in document git dp600 practice
1:30:28that's the repository that we just
1:30:29cloned from Azure devops and it's
1:30:31currently empty for this kind of thing
1:30:33we can change the save as type to a
1:30:35powerbi project and we can just save it
1:30:37like that click on Save and now we have
1:30:39our powerbi project saved in pbip format
1:30:43so let's just have a look at what that
1:30:45looks like okay so this is our git
1:30:47repository dp600 practice and we've
1:30:50saved our powerbi project as a BP file
1:30:55it's actually a number of different
1:30:57files and folders right what we've got
1:30:59here is we've got the read me that was
1:31:01already in the repository and my file is
1:31:03that text file that we've been working
1:31:05with but now what we've got is two
1:31:06folders a git ignore file and this pbip
1:31:11file format so what it's done is it's
1:31:14basically decomposed our powerbi project
1:31:17into a series of files and folders that
1:31:20describe what's going on in that report
1:31:23Okay so let's move back to uh GitHub
1:31:26desktop here and what you'll notice is
1:31:29well the first thing is I've created a
1:31:30new Branch okay so now that we're going
1:31:32to be updating this pbip file we'll
1:31:36notice that all of the files are now
1:31:37listed here CU we created that pbip
1:31:40they're all text based they're all
1:31:42trackable using Version Control first
1:31:44commit of the pbip again we can commit
1:31:47that up into Azure devops push that to
1:31:51the origin in Azure devops now if we go
1:31:54through into azid devops again poll
1:31:56requests we can see that here we can
1:31:58create a new poll request and then we
1:32:00can commit that in there the reviewer
1:32:03will be myself again create that approve
1:32:05it complete merge so now what we've got
1:32:09is our powerbi report in our Azure
1:32:13devops environment and it's been
1:32:15approved and it's now trackable through
1:32:17Version Control now imagine I'm a
1:32:19powerbi developer and I want to make
1:32:21some changes to that report what is that
1:32:24going to look like well we're going to
1:32:25create a new Branch so anytime you're
1:32:27making changes to your powerbi file
1:32:29you're going to create a new Branch
1:32:31powerbi changes obviously you'd make it
1:32:33a bit more descriptive than that but
1:32:35that's good enough for this purpose here
1:32:37then we're going to go through and
1:32:40actually this isn't a marketing report
1:32:41it's sales and marketing so we're going
1:32:43to update the title here just as an
1:32:45example we're going to save that file
1:32:47and now this is the real power of
1:32:50Version Control for our power VI assets
1:32:53but also all of the stuff in fabric that
1:32:55we'll have a look in a minute any
1:32:56changes that we make to that powerbi
1:33:00file are now being tracked in our
1:33:03Version Control System okay and then
1:33:05when we're done we can push those
1:33:07changes up into Azure devops go through
1:33:10the review process merge them into the
1:33:13the main branch okay so you're probably
1:33:14thinking great but up until this point
1:33:17you haven't even mentioned fabric I
1:33:19thought this was a channel talking about
1:33:20fabric well we just built up the
1:33:21concepts of git vers control we've
1:33:24looked at the local case of powerbi
1:33:27report development but everything else
1:33:28in fabric happens in the fabric Cloud so
1:33:31now let's look at where fabric fits into
1:33:33all this okay how do we set up a
1:33:35repository in our workspace how do we
1:33:38link a repository to our workspace let's
1:33:40have a look at how fabric fits into this
1:33:42picture now okay so here we are in
1:33:44Microsoft fabric what I've done is I've
1:33:46set up a workpace what we're going to do
1:33:49is link it to our Azure devops
1:33:52repository our Azure devops project
1:33:54so to do this you will need to be a
1:33:56workspace administrator go to workspace
1:33:58settings so the first time that you
1:34:01click on get integration you'll need to
1:34:03actually sync it with your account I've
1:34:05already done that with mine and it's
1:34:07going to allow you to select the
1:34:08organization the project the repository
1:34:11the P600 practice and a specific branch
1:34:14that you might want to connect to now
1:34:15one thing I would say is that the way
1:34:17that we've done that protecting of the
1:34:19branching in the last section of the
1:34:22video what I found in practice is that
1:34:24this doesn't work particularly well
1:34:25currently with the way that git
1:34:27integration is set up in fabric but
1:34:29we'll have a look at what it looks like
1:34:30just click on connect and sync okay so
1:34:33now our first Sync has started and it's
1:34:36been done successfully what you'll
1:34:37notice is this Source control section
1:34:40here it's got zero in it so our source
1:34:42control so our things in AZ devops
1:34:45repository and what's in fabric is
1:34:47perfectly in sync currently and we can
1:34:49see that as well with the green sync and
1:34:52green synced here so now that we've got
1:34:54uh what's in our Azure devops repost
1:34:56synced with what's in fabric now we can
1:35:00add in fabric items into this mix right
1:35:03so we can create a new data pipeline for
1:35:06example create a data pipeline at the
1:35:08moment I'll just keep it as an empty one
1:35:09and we can just go back to our
1:35:11repository and we can see that it has
1:35:14been uncommitted right and we can see
1:35:16the source control button up here is now
1:35:19saying one so if we click on that we can
1:35:20see that we got one change here so comp
1:35:23compared to what's in our main branch in
1:35:26AZ devops it's noticed that we've got a
1:35:28change so it's similar to how we saw in
1:35:31GitHub desktop which tracks local
1:35:33changes this is tracking changes in
1:35:36fabric as well so what we can do is
1:35:40commit this change this data pipeline
1:35:42into our source control in AZ devops now
1:35:45to do this because we've got a
1:35:47protection on that main branch let's
1:35:49just try committing and see what it says
1:35:51first commit data pipeline but we know
1:35:53know that we've got protections on that
1:35:55main branch so when we commit I don't
1:35:57think it's going to allow us to do that
1:35:59no forbidden due to the branch policy
1:36:01and this is where it gets a bit more
1:36:02complicated within fabric what we're
1:36:04going to have to do is check out a new
1:36:06Branch so maybe we want to do adding
1:36:10data pipeline that could be our new
1:36:12branch and then that's going to allow us
1:36:14to commit added a new data pipeline
1:36:16we're going to click on the item that we
1:36:18want to merge which is our data pipeline
1:36:20or commit not merge and now that's going
1:36:22to go through into Azure devops okay so
1:36:26if we open up our poll requests section
1:36:28boom we can see adding data pipeline so
1:36:31now we've got these two things in sync
1:36:33between Fabric and Azure devops and our
1:36:36local repository for powerbi development
1:36:38and Azure devops and again if we want to
1:36:40merge that into the code base into the
1:36:43main branch we can go through a familiar
1:36:45process of choosing a reviewer creating
1:36:48that PLL request and if you're the
1:36:50approver you can then approve that PLL
1:36:52request and complete and merge it into
1:36:55the main branch so then if we go back to
1:36:57our repository now we can see that in
1:37:00our main branch we have this my data
1:37:01pipeline data pipeline so that is now
1:37:04being tracked with Version Control in
1:37:07our repository and if we go back to
1:37:09fabric we can say that we've now got
1:37:12three synced items here so it's
1:37:15perfectly in sync with what's in our
1:37:17Branch now the problem here or the
1:37:20slight limitation that I found is
1:37:22changing back to the main branch cuz
1:37:25currently we're on this adding data
1:37:27pipeline Branch okay and that was just a
1:37:30feature Branch we want to move back onto
1:37:32the main branch normally so that we can
1:37:34maybe create another another Branch but
1:37:35we don't want to be branching from this
1:37:37added data pipeline we want to be
1:37:39branching from the main branch now that
1:37:41is possible we can just go back to
1:37:43workspace settings go through to the get
1:37:46integration again and we can change our
1:37:48Branch back to the main so we're going
1:37:50to switch and override and now we're
1:37:51going back to the the main branch and
1:37:54all three items are synced because we
1:37:55know that we've merged into that main
1:37:57branch but the problem here is that only
1:37:59the workspace administrator can change
1:38:01that Branch back so if you found any
1:38:03better ways of working with protected
1:38:05branches in git and fabric please let me
1:38:08know okay so we've covered quite a lot
1:38:10of ground there from the basics of git
1:38:13and Version Control in general right
1:38:15through to syncing our changes or
1:38:18tracking changes in a powerbi project
1:38:20into AZ devops and then the same from a
1:38:23fabric environment so tracking changes
1:38:24that we make within fabric also to the
1:38:27same Azure devops repository let's just
1:38:29do a bit of a summary there and focus
1:38:31back on the exam because obviously not
1:38:33all of that is going to be tested some
1:38:34of it was just for your background
1:38:36knowledge so in general Version Control
1:38:38with Git allows you to track changes
1:38:41made to fabric items we can also revert
1:38:44back to older versions of an item as
1:38:46well I didn't show you how to do that
1:38:48but that is also possible within git
1:38:50item management Git Version Control now
1:38:52one of the benefits of this is that
1:38:53multiple users can collaborate on the
1:38:55same fabric item or the same powerbi
1:38:58report okay so no longer are you sending
1:39:00around a PBX file to your colleague
1:39:03everyone can work on the same pbip file
1:39:07and changes are tracked using Version
1:39:10Control and you can also update the same
1:39:12report at the same time as long as
1:39:14there's no conflicts when you merge into
1:39:16your branch two different changes from
1:39:18two different people can both be merged
1:39:20into the same kind of central branch as
1:39:23long as there's no conflict as I
1:39:24mentioned that's absolutely fine so it's
1:39:26a good way that enables collaboration on
1:39:28powerbi reports we've also looked at how
1:39:30you can Implement a check and approval
1:39:32process for approving changes made to
1:39:35fabric items so if you don't want your
1:39:38Junior developer to be updating your
1:39:41fabric notebooks and just pushing those
1:39:43into your Git Version controlled
1:39:45repository Without You approving them
1:39:47you want to set up protections on that
1:39:49on your main branch or your equivalent
1:39:51of a main branch for they get pushed
1:39:53into to production now currently the
1:39:55following items are supported I would
1:39:57argue that not all of these are probably
1:39:58fully supported like as as fully
1:40:01supported as we would like but it is
1:40:03possible to at least check them into a
1:40:05version control system so data pipelines
1:40:08Lakehouse notebooks pageat reports
1:40:11reports apart from ones that are
1:40:13connected to Azure analysis services and
1:40:15also semantic models except for these
1:40:18exceptions here so that is git
1:40:20integration and Version Control now the
1:40:22next section of the exam is is related
1:40:24to this but it's around deployment
1:40:26Solutions so let's just look at what do
1:40:28we mean by deployment Solutions well
1:40:31actually let's start with deployment
1:40:32what do we mean by deployment because
1:40:34for many people coming from the world of
1:40:36analytics deployment is probably a bit
1:40:38of a New Concept okay so rather than
1:40:41having one copy of a powerbi report for
1:40:45example which is your production copy
1:40:47instead we're going to have multiple
1:40:48copies typically three sometimes four
1:40:51we're going to have a development
1:40:52version and that's the version that
1:40:53you're using when you're making changes
1:40:55to the reports we're going to have a
1:40:57test version of that report so that is
1:40:59the report that you send to your client
1:41:02or your colleagues to review or maybe
1:41:04you've got automated testing in place
1:41:07that's done at that stage and the test
1:41:09version is sometimes also called the
1:41:10staging version in databases we've also
1:41:13got the production report right so that
1:41:15is the the public facing report that
1:41:18gets given to your client or it gets
1:41:20shared within your organization now
1:41:22Microsoft have released a feature called
1:41:26deployment pipelines to help manage
1:41:28these three environments or more in a
1:41:31bit more detail so let's have a look at
1:41:33deployment pipelines in Microsoft Fabric
1:41:35in a bit more detail okay so now let's
1:41:36look at how we can set up deployment
1:41:39pipelines in Microsoft fabric so what
1:41:41you're going to need here is three
1:41:43workspaces and each workspace is going
1:41:46to represent a different environment a
1:41:48different stage in our deployment
1:41:49pipeline so if we click through to
1:41:51workspaces you'll see that I've set up
1:41:54dp600 Dev test and production so these
1:41:57are three separate workspaces and
1:41:59currently they're completely empty so
1:42:01what we're going to do is set up a
1:42:02deployment pipeline so that we can
1:42:04manage that deployment process from
1:42:06development test through to production
1:42:08because what we want to be doing is make
1:42:10doing a lot of our development work
1:42:11obviously in the development workspace
1:42:13and then when it's ready we want to push
1:42:15those changes into our test workspace so
1:42:18that we can do testing know share it
1:42:20with colleagues who might want to test a
1:42:22report or a notebook run test scripts so
1:42:26integration tests unit tests data
1:42:29validation checks something like that in
1:42:31this test environment and then the
1:42:32production is going to be where we're
1:42:34going to be sharing making our public
1:42:36our production reports or notebooks and
1:42:39things like that so what we're going to
1:42:40do to set up a deployment pipeline is
1:42:42click on the workspaces and then
1:42:44deployment pipelines and we're going to
1:42:45set up a new Pipeline and we're going to
1:42:47give it a name so we can call it dp600
1:42:50Pipeline and it's going to ask you to
1:42:52customize your stages now you can add
1:42:55multiple stages you can add as many as
1:42:56you like here for our example we're just
1:42:58going to do three development test and
1:43:00production and click create and now
1:43:02we're going to get through to this kind
1:43:03of wizard this UI interface it's going
1:43:06to allow us to specify which workspace
1:43:08we want to connect to which stage in our
1:43:11deployment pipeline process so we want
1:43:13to find our dp600 Dev workspace assign
1:43:17that here then for the test stage dp600
1:43:22test assign that here then for the
1:43:24production do that as well so now we've
1:43:26synced our three workspaces with our
1:43:29three stages in the deployment Pipeline
1:43:31and currently we can see that this
1:43:32completely empty right so we haven't got
1:43:34any items in any of these three
1:43:35workspaces so let's go through to our
1:43:38dp600 Dev workspace and we can see that
1:43:40it's here we synced it with this
1:43:42deployment pipeline so it knows it's
1:43:44part of a deployment pipeline here so if
1:43:46we create a new notebook so this is our
1:43:49new ETL notebook let's just call it ETL
1:43:52notebook could be doing whatever you
1:43:54want in there doesn't really matter for
1:43:56these for the purposes of this demo and
1:43:57we're going to go back into our
1:43:59workspace and now we can see we've got
1:44:01this ETL notebook say we're happy with
1:44:04that we've done our development work and
1:44:05we're happy with it now we want to push
1:44:07that into our test environment so we'll
1:44:11go into our deployment Pipeline and here
1:44:14what you can see is that now we can see
1:44:15this deploy right so we can select any
1:44:18items that we've got in our Dev
1:44:20environment and we can deploy them we
1:44:22can add a note in there
1:44:23if you want and we can deploy that into
1:44:25our test environment so it's going to
1:44:27move one stage to the right into our
1:44:29test environment what it's going to do
1:44:30is copy exactly the file that you've got
1:44:33there and it's going to copy it and
1:44:34paste it into our test environment so
1:44:37now we can see we've got one notebook in
1:44:39this test environment okay so let's just
1:44:41have a look at that there so workspaces
1:44:43dp600 test so now we have ETL notebook
1:44:46in our test environment now one thing to
1:44:48bear in mind is that we have a history
1:44:50of our deployments so we've seen that
1:44:52we've deployed to test at this time of
1:44:54date by this person me and we can see
1:44:57the number of items that have changed so
1:44:59we've got one new item in that
1:45:01deployment now another thing that we
1:45:02need to look at is this button here so
1:45:05this is quite important it's called
1:45:06deployment rules so if we click on there
1:45:09we can look at a deployment Rule and
1:45:12deployment rules basically allow you to
1:45:15change different parameters and settings
1:45:19at different stages in the pipeline and
1:45:21they differ depending on the the fabric
1:45:23item that you're looking to assign these
1:45:26deployment rules so for the notebook our
1:45:28deployment rule that we can set is
1:45:30changing the default lake house okay so
1:45:33we can add a rule in there whereby we
1:45:35can change the default Lakehouse so
1:45:36maybe you have a development Lakehouse
1:45:39and a test Lakehouse and when we push to
1:45:41the test environment actually we want
1:45:43that notebook to be reading from the
1:45:45test Lakehouse so that's what we're
1:45:47going to be doing here in this
1:45:48deployment rules section so that's
1:45:51something important to note there now if
1:45:54we want to actually deploy this all the
1:45:56way into production we can click on
1:45:57deployment at this test stage and now
1:46:00it's going to copy that notebook into
1:46:01our production environment now if we
1:46:03click on the settings of this stage the
1:46:05production stage we can see that we've
1:46:07got this make stage public okay and what
1:46:10that means is that people can actually
1:46:12access the output of this stage so if
1:46:15they're in that workpace they're going
1:46:17to be able to see the output they're
1:46:19going to be able to see the items that
1:46:21are in there with the other stages like
1:46:22this test environment by default that is
1:46:25not a public stage so so that's
1:46:27something to bear in mind around public
1:46:29stages and non-public stages so that's
1:46:32the very basics of deployment pipelines
1:46:35now currently the functionality I would
1:46:37say is quite limited as to what you can
1:46:39do in deployment pipelines in fabric I
1:46:42think this is a feature that they're
1:46:43going to be adding a lot more to
1:46:45currently it's a very manual process
1:46:46right and normally when we're deploying
1:46:49stuff in in the real world in software
1:46:51development World a lot of this is
1:46:53automated Okay so we've looked at the
1:46:55basics of deployment pipelines what that
1:46:58looks like in fabric let's just do a bit
1:46:59of a summary of what we've just learned
1:47:01there again focusing back on the exam so
1:47:04the overall goal is to add layers of
1:47:07control when we're developing and
1:47:09deploying new fabric items okay or
1:47:11making changes to existing items in
1:47:13fabric ultimately we trying to ensure
1:47:15that new things that you develop are not
1:47:17going to break your existing analytic
1:47:20Solutions right so you're adding in that
1:47:22test
1:47:23so we're not just pushing straight into
1:47:25production and risking ruining any sort
1:47:27of analytics in your environment now
1:47:29normally this includes three stages
1:47:32development test staging test St staging
1:47:35and production sometimes you had a
1:47:36fourth one in there called like pre-prod
1:47:38it just depends on your strategy now
1:47:40deployment using deployment pipelines
1:47:42involves copying items from workspace to
1:47:44another and by default in fabric that's
1:47:46a manual process and deployment rules
1:47:49can be implemented to change things like
1:47:51the default Lakehouse for a notebook or
1:47:54the data sets and data sources that your
1:47:57semantic model reads from for example
1:48:00now as well as the deployment pipelines
1:48:02functionality there's a number of other
1:48:04ways to manage deployment in fabric now
1:48:07that could be done through branching in
1:48:10Azure devops you can also use in Azure
1:48:12devops functionality called pipelines
1:48:15right so you create a yaml template
1:48:17we're not going to go through what that
1:48:18looks like but just for the exam know
1:48:20that it is possible and you can also do
1:48:22deployment of semantic models via the
1:48:25xmla endpoint we're going to be looking
1:48:27at that in a bit more detail a bit later
1:48:29on in this lesson okay so as I mentioned
1:48:31there are a few other ways that we can
1:48:34deploy things in fabric now a really
1:48:36good resource here to check out is Kevin
1:48:39Chance's blog and I'll leave a link to
1:48:40that in the school Community
1:48:42specifically there's two blog posts here
1:48:44around cicd for the data warehouse now
1:48:47the data warehouse is not natively
1:48:48supported yet in fabric but you can set
1:48:51up a SE project and then use things like
1:48:55yaml pipelines and Azure devops so if
1:48:58you want to understand other ways that
1:48:59we can deploy items in fabric I
1:49:02recommend checking out these blogs here
1:49:04now if we're looking specifically at the
1:49:06semantic model and how we can deploy
1:49:08these we also have the option of the
1:49:10xmla endpoint so broadly speaking
1:49:13there's two ways to create and manage
1:49:16semantic models number one is to create
1:49:18and manage your semantic model within
1:49:20your workspace so create it from a lake
1:49:23housee or a data warehouse within your
1:49:25workspace but method two is to create
1:49:28your semantic model in a third party
1:49:30tool for example TBL editor and then you
1:49:32can deploy that via What's called the
1:49:35xmla endpoint into your workspace now to
1:49:38grab that xmla end point you need to go
1:49:41to your workspace settings and then you
1:49:42get this palbi URL kind of thing and you
1:49:46can connect that to TBL editor or SS SMS
1:49:50or DAC Studio as well to deploy your
1:49:54models into fabric using this xmla
1:49:58endpoint okay so the next section of the
1:50:00study guide that we're going to look at
1:50:01is these three different file types or
1:50:03three different items that we can create
1:50:06in pobi and Microsoft fabric that you
1:50:08need to know for the exam number one is
1:50:10the powerbi template file then we got
1:50:12the powerbi data source file and then
1:50:15shared semantic models so we're going to
1:50:17look at each of these in a bit more
1:50:18detail starting with the powerbi
1:50:20template file okay so the powerbi
1:50:22template file is a reusable asset that
1:50:25can improve the efficiency and
1:50:27consistency when you're creating power
1:50:29reports now you can easily save a PBX
1:50:32file as a powerbi template file and then
1:50:34you can use that to generate new reports
1:50:37in a given style or with a given layout
1:50:40already specified in that template file
1:50:43if you got parameters in that powerbi
1:50:45file then when you create a new project
1:50:48or a new file from the template you'll
1:50:50be asked to set your parameter
1:50:53in there if there's any parameters in
1:50:55that report let's just have a quick look
1:50:56at how you can create a PBI template
1:50:59file from a pbix file okay so here we
1:51:03are back in pobi desktop and here we've
1:51:06got a pobi project that we're working on
1:51:08our sales and marketing analysis report
1:51:10say we want to create multiple versions
1:51:13of this report all with the same format
1:51:16the same structure and the same layout
1:51:18here so the same pages in this report
1:51:21what we can do simply is go to file save
1:51:23as and then at the bottom here rather
1:51:25than saving as a PBX or a pbip we're
1:51:29going to save it as a PB a powerbi
1:51:31template file so we just select a folder
1:51:33to save it to and then click on Save we
1:51:36can give it a description sales template
1:51:39okay and that's going to save your
1:51:40report and the next time you go to
1:51:42create a new report we can then import
1:51:44that template and that can be your
1:51:45starting point for your new report so
1:51:48the powerbi data source file or PB IDs
1:51:52is a another file type that basically
1:51:54represents a data source file so it's a
1:51:57reusable asset and it can help us
1:51:59quickly transfer all of the data
1:52:01connections that you create in one
1:52:02powerbi file transfer them over to
1:52:05another report okay in this section of
1:52:07the exam we're going to be looking at
1:52:10impact analysis and the lineage tool in
1:52:13Microsoft fabric so here we have a
1:52:15workspace and it's got quite a lot of
1:52:17different items and if we click on this
1:52:20button here in the top right hand corner
1:52:21we can change from this list View to a
1:52:24lineage View and this is what this looks
1:52:26like here we can see all of our fabric
1:52:28items within that workspace and we can
1:52:31have a look at how data is Flowing from
1:52:34Source through to different lake houses
1:52:37here we've got some notebooks here the
1:52:39semantic model that's being created from
1:52:41that Lakehouse now for each of the main
1:52:43items in our lineage view we've got this
1:52:46button here which is the impact analysis
1:52:48button so for this Bronze Lake housee we
1:52:51can click on the impact analysis and we
1:52:54can look at what the downstream items of
1:52:58that Lakehouse are so if you're going to
1:53:00be making some changes to that Lake
1:53:02housee that might break some of the
1:53:05notebooks or the semantic model for
1:53:07example the impact analysis basically
1:53:09allows us to see what the downstream
1:53:11items are now in our example we've only
1:53:13got three Downstream items but you could
1:53:15have 10 or 20s or hundreds of different
1:53:18Downstream items so it's really
1:53:20important to know if you make a change
1:53:22Upstream what the impact is going to be
1:53:25another piece of functionality that you
1:53:26have here is to notify people so you can
1:53:29notify people that are listening for
1:53:32notifications on any of those Downstream
1:53:34items and you can let them know before
1:53:36you make a change so we're going to add
1:53:37a new column to this table in the lake
1:53:40house for example we're going to change
1:53:41add some new tables FYI just to make
1:53:44them aware of the changes that you're
1:53:46going to make before you make them so
1:53:48that's basically all the lineage tool
1:53:50and the impact analysis tool tool in
1:53:52fabric look like at the moment again I
1:53:54think they're adding a lot more
1:53:55functionality to these features in the
1:53:58future that's basically all you can do
1:53:59with this at the moment just bearing
1:54:01that in mind for the exam because you
1:54:02might get asked questions around
1:54:04notification about how do you find
1:54:07Downstream items of a particular fabric
1:54:09item so just have a quick look at some
1:54:11of your workspaces perform some impact
1:54:13analysis just by looking at the
1:54:16downstream items for the exam Okay so
1:54:18we've covered a lot of ground there in
1:54:21that video Let's just round the video up
1:54:24with some practice questions to test
1:54:26your knowledge of this section of the
1:54:28exam question one you are looking to
1:54:30improve the efficiency and consistency
1:54:32of your powerbi development team you
1:54:34want each report created by the team to
1:54:37always consist of three pages the intro
1:54:40the context and the analysis the reports
1:54:42should always align to the company
1:54:44branding which of the following would
1:54:46help you achieve this number one would
1:54:48you create a pbip file two create a PBX
1:54:52file three create a pbit file four
1:54:55create a PB IDs file or number five
1:54:58using a Json custom report theme pause
1:55:02the video here have a think and I'll
1:55:03reveal the answer shortly so the answer
1:55:06here is the pbit file because that is
1:55:08the powerbi template file now the
1:55:11important part of this question was you
1:55:13want each report created by the team to
1:55:16consist of the following pages right so
1:55:18you might have thought oh we're talking
1:55:20about branding talking about following a
1:55:21star guide e might be an answer there
1:55:24using the Json report theme that we
1:55:26looked at in the first video in this
1:55:28series but if we want to include Pages
1:55:31kind of template pages and that's going
1:55:33to be the powerbi template file PBP the
1:55:36project file the PBX that's not going to
1:55:38do the job and the PBS is just for data
1:55:41sources not for giving us a template
1:55:44structure to follow for our report which
1:55:46of the following most accurately
1:55:48describes git a to use Git You must be
1:55:51using GitHub B git is a Microsoft
1:55:54product for tracking changes made to
1:55:55fabric items C git is an open-source
1:55:59version control system that tracks
1:56:01changes in any set of text based files D
1:56:04git allows us to add deployment rules to
1:56:07fabric deployment pipelines so the
1:56:09answer here is C git is an open-source
1:56:12version control system that tracks
1:56:14changes in any set of text based files
1:56:17right so git the underlying technology
1:56:20is open source and it's used in a wide
1:56:23variety of Version Control Systems one
1:56:25of which is azure devops the repo is in
1:56:28there GitHub is another example bit
1:56:30bucket is another example there's lots
1:56:32of these different ones so you don't
1:56:33have to be using GitHub to be using git
1:56:36git is not a Microsoft product for using
1:56:39specifically in fabric it's kind of like
1:56:40a generic tool that's used right across
1:56:42software development industry and it's
1:56:44nothing to do really with deployment
1:56:46rules in fabric deployment pipelines
1:56:49although you can set up git to be the
1:56:52the version control system in different
1:56:54stages different workspaces in your
1:56:57deployment pipelines but it's not really
1:56:59best describing git question three in an
1:57:01Azure devops repo the main branch is
1:57:04protected so it needs approval before
1:57:06any changes are merged into it the repo
1:57:08contains one pbip file you have to
1:57:11update the title in the report merging
1:57:13these changes to the main branch in
1:57:15which order should you carry out the
1:57:17following tasks to achieve this now this
1:57:19is an unordered list your job is to
1:57:22order this list so the tasks here are
1:57:25commit and push the feature Branch wait
1:57:27for approval then merge into the main
1:57:29branch clone the repository to your
1:57:31local machine make the required changes
1:57:33to the report check out a new feature
1:57:35Branch from the main branch and then
1:57:38open a pull request in Azure repos or
1:57:41Azure devops so order this list and then
1:57:43we'll show you the correct order shortly
1:57:46so the correct ordering here starts with
1:57:48cloning the repository to your local
1:57:50machine so if you want to make any
1:57:51changes to report you need that
1:57:53repository to be in your local
1:57:54environment first then you're going to
1:57:56check out a branch because the main
1:57:59branch is protected so we can't edit
1:58:01that directly need to check out a new
1:58:03Branch from the main branch then we're
1:58:04going to make the changes to the report
1:58:06these are going to be tracked via our
1:58:08Version Control System we're going to
1:58:10commit and push those changes that we
1:58:12made on that feature Branch we're going
1:58:14to open a poll request in Azure repos or
1:58:17as devops and we're going to wait for
1:58:18approval and then merge it into the main
1:58:21branch so that is the correct order here
1:58:23question four you want to deploy a
1:58:25semantic model using the xmla endpoint
1:58:28where can you find the xmla endpoint to
1:58:30set up a connection with a third party
1:58:32tool is it a go to the workspace
1:58:34settings for the workspace you want to
1:58:36deploy your model to B go to the fabric
1:58:39admin portal and then capacity settings
1:58:41C in your workspace find your semantic
1:58:43model then click on the settings to get
1:58:45the xmla endpoint address in the Azure
1:58:48portal in your fabric capacity go to the
1:58:50xmla endpoint connect ction string
1:58:53settings so the answer here is a go to
1:58:55the workspace settings for your
1:58:57workspace you want to deploy your model
1:59:00to when we're creating that xmla
1:59:02endpoint connection that's going to link
1:59:04to our workspace in fabric that's where
1:59:06you're going to go to get the address
1:59:09not in capacity settings not in the
1:59:11Azure portal and see in your workspace
1:59:13find your sematic model where the
1:59:15sematic model doesn't actually exist in
1:59:16that workspace because we want to deploy
1:59:18it into there and even if it did exist
1:59:20it doesn't have the xmla end point in
1:59:23the settings anyway so congratulations
1:59:25that is the first section of the exam
1:59:28study guide complete next up we move
1:59:31into the biggest section of the exam
1:59:33which is worth 40 to 45% Camila has been
1:59:36seriously impressed with your skills and
1:59:37knowledge so far in the next lesson
1:59:39we'll be looking at how to create
1:59:41objects in a Lakehouse and a data
1:59:43warehouse so click here for the next
1:59:46lesson in this series hey everyone
Getting data into Fabric
1:59:49welcome back to the channel and this is
1:59:51going to be video five in our dp600 exam
1:59:54preparation course and in this video the
1:59:57focus is going to be on getting data
1:59:59into fabric now this is the first part
2:00:02of the second section in the exam which
2:00:04is all around preparing and serving data
2:00:06and it's worth 40 to 45% of the exam so
2:00:10there's going to be a lot of questions
2:00:11in this so let's get into it in the
2:00:13video we're going to be covering
2:00:14ingesting data using data pipeline data
2:00:16flow notebooks copying data which is
2:00:19basically the same thing but they've
2:00:20included it twice in the study choosing
2:00:22an appropriate method for copying data
2:00:24so not just understanding what the tools
2:00:27available to us but making decisions
2:00:28about the best method given a particular
2:00:31problem or a particular scenario and
2:00:32also creating and managing shortcuts now
2:00:34you'll notice that these are slightly
2:00:36deviation from the study guide what I've
2:00:38done is I've had a look at the whole of
2:00:40section two and I've changed some of the
2:00:42order of things don't worry we'll be
2:00:43going through all of the skills in the
2:00:45study guide but we're just going to be
2:00:46going through them in a slightly
2:00:47different order so these are the four
2:00:49that we're going to focus on today now I
2:00:50have released a video in the last month
2:00:53this one here data pipelines versus data
2:00:54flow shortcuts notebooks and this is a
2:00:56more comprehensive video so I definitely
2:00:59recommend watching this either before or
2:01:01after this video so as such for this
2:01:03video we won't be starting from zero
2:01:06more revising what we went through in
2:01:08the last video updating it a little bit
2:01:10because there have been some changes
2:01:11revising some of the Core Concepts and
2:01:14the distinctions between these tools
2:01:16that you need to know for the exam as
2:01:17ever there will be five sample questions
2:01:19at the end of this video and we've got
2:01:21the school Community with quite
2:01:23extensive notes and links to further
2:01:26resources if you want to go into a bit
2:01:28more detail about any of the topics that
2:01:29we cover in this lesson so let's start
2:01:32with a bit of a framing of all the
2:01:33different options that we have available
2:01:35to us when we're talking about ingesting
2:01:37data and getting data into fabric so one
2:01:40group of tools is the data ingestion the
2:01:43El or the ETL tools right so extract
2:01:46transform and load these are going to be
2:01:48copying data from external systems
2:01:51bringing them into Fabric and saving
2:01:52them in some sort of data store in
2:01:54fabric so we've got the data flow data
2:01:56pipeline the notebook and the event
2:01:58stream now the event stream I don't
2:01:59think is actually covered in the DP 600
2:02:01exam so we're not going to be talking
2:02:03about that much today we also have
2:02:04shortcuts so that's another method that
2:02:06we need to be aware of that we're going
2:02:08to go through in this video and that
2:02:09involves bringing data in from Amazon S3
2:02:12ADLs Gen 2 the data verse and also
2:02:15Google Cloud Storage as well we've also
2:02:17got internal shortcuts so connecting our
2:02:20different data sets Within fabric so
2:02:23creating references to other Lakehouse
2:02:25tables from a lake house for example
2:02:27then at the bottom we've got a new
2:02:29feature that's in preview at the moment
2:02:30which is database mirroring so you can
2:02:32create a mirror of your database within
2:02:35fabric from either snowflake Cosmos DB
2:02:37or as a SQL at the moment and again I
2:02:40don't think this is actually part of the
2:02:41exam study guide currently so we're not
2:02:44going to be talking about it in this
2:02:45section of the study guide if you don't
2:02:47want to learn a bit more about databas
2:02:48Maring I did mention it in that video
2:02:51that I mentioned previously so in this
2:02:53video we're just going to be going
2:02:54through the data flow the data pipeline
2:02:56the notebook and shortcuts in a bit more
2:02:59detail cuz these are the things that you
2:03:00need to study for the exam starting with
2:03:03the data flow so as you know the data
2:03:05flow comes with about 150 connectors to
2:03:08external systems and you can bring that
2:03:10data in using power query like that no
2:03:13low code interface and you can do data
2:03:15transformation on that data before
2:03:17writing it to one of the fabric data
2:03:19stores so when should you use it and
2:03:21maybe when should you not use it well if
2:03:23you want to use any of those 150
2:03:24connectors then it's definitely good to
2:03:27use data flows it's a no and low code
2:03:29solution so it's quite maintainable if
2:03:32you have people who perhaps don't have
2:03:34SQL or python skills in your
2:03:36organization this is a good method and
2:03:38as I mentioned we can do extract
2:03:40transform and loading all in one tool
2:03:42now the data flow is one of the tools
2:03:44that we can use to access on premise
2:03:46data via the on- premise data Gateway
2:03:49and it's also quite useful when you want
2:03:51to ingest more than one data set and
2:03:54maybe combine them in the same data flow
2:03:56although if you're following a kind of
2:03:58Medallion architecture maybe that's not
2:04:00what you want to do but it's just an
2:04:02option that is possible with the data
2:04:03flow it's also the only tool that we can
2:04:05use to upload raw local files so maybe
2:04:09you have a CSV file you know this file
2:04:12won't ever have to really be updated you
2:04:14just want to get some data into fabric
2:04:17you can do that in the data flow so
2:04:19maybe when you shouldn't be using this
2:04:21well traditionally they've been
2:04:23struggling with large data sets but
2:04:25there's been a number of features
2:04:26released to try and speed that up so
2:04:29recently they released fast copy so
2:04:31that's one feature that they released to
2:04:32try and speed up data flows now the data
2:04:35flow uses the same backend
2:04:37infrastructure that the data pipeline
2:04:38uses in the copy data activity so the
2:04:41performance between these two should now
2:04:42be a lot more similar another aspect
2:04:45with power queries it's quite difficult
2:04:46to implement data validation right so
2:04:49because all of our logic is being locked
2:04:52up in those power query routines it's
2:04:54difficult to kind of validate those
2:04:55steps as you're going through them so
2:04:57that's maybe one reason why you might
2:04:58want to go for the elt routine rather
2:05:01than the ETL extract load transform
2:05:04rather than extract transform load now
2:05:07it's also the only tool out of the data
2:05:10flow the data Pipeline and the notebook
2:05:12where you can't pass in external
2:05:14parameters so it's difficult to build
2:05:16kind of metadata driven architectures
2:05:18with the data flow at the moment now
2:05:20there is a kind way around this where
2:05:22you can create another query within your
2:05:25power query engine of the the data set
2:05:28or the the metadata that you might want
2:05:30to use to parameterize a solution so
2:05:33there is kind of a workaround but
2:05:35there's no native functionality for
2:05:36passing parameters into a data flow from
2:05:39a data pipeline for example next up
2:05:41we're going to be talking about
2:05:42ingesting data with a data Pipeline and
2:05:45the data pipeline is primarily an
2:05:46orchestration tool right but it can also
2:05:49be used to get data into fabric using
2:05:52the copy data activity and also some
2:05:54other activities as well you can use for
2:05:56this but the copy data activity is the
2:05:58main one now one of the main pros of a
2:05:59data pipeline is that where it performs
2:06:01well on large data sets and now as I
2:06:04mentioned before the data flow now has
2:06:06fast copy so the for performance should
2:06:09be comparable between the two it has
2:06:11many connections to cloud data sources
2:06:13especially in Azure so it's good for
2:06:15using if you've got data in Azure for
2:06:17example we need some sort of control
2:06:19flow logic so maybe looping through
2:06:21through different tables for example
2:06:23there's a lot of functionality for
2:06:25building metadata driven or
2:06:27parameterized data ingestion methods so
2:06:31that's something to bear in mind with
2:06:32the data Pipeline and as well as the
2:06:33copy data activity you can also use it
2:06:36to trigger a wide variety of other
2:06:38actions in fabric for example a stored
2:06:41procedure and a stored procedure can
2:06:43also be used for ingesting data into
2:06:47fabric for example the copy into
2:06:49statement in t s call can be used to
2:06:51ingest a CSV file into a data warehouse
2:06:55directly for example now some of the
2:06:56cons in the data pipeline where it can't
2:06:58do the transform piece natively so
2:07:01there's no real data transformation
2:07:03activities but what you can do is embed
2:07:05notebooks and data flows into a data
2:07:08pipeline if you want to do that
2:07:09transformation has no ability to upload
2:07:12local files so that's not possible with
2:07:14a data pipeline at the moment you do
2:07:15need to be careful with any sort of
2:07:17crossw workpace data pipeline usage now
2:07:20we did submit an idea on ideas. fabric.
2:07:24microsoft.com and it is actually going
2:07:26to be planned now so that's good news so
2:07:28they've planned this feature we don't
2:07:29know when it's going to be released yet
2:07:31but they are working on support for
2:07:34crossw workpace data pipelines and what
2:07:37we mean by that is maybe you want to
2:07:38bring data in from a data source and you
2:07:42want to write it to a destination that's
2:07:45in a different workspace to the data
2:07:47pipeline so that currently isn't
2:07:48possible but they're working on that
2:07:50feature as we speak next up we've got
2:07:52ingesting data with a notebook and a
2:07:55notebook is just a general purpose
2:07:57coding notebook which can be used to
2:07:59well for a wide variety of things but
2:08:01one of the things is to bring data into
2:08:03fabric now we can do this via either
2:08:06connecting to apis using something like
2:08:09the requests library in python or
2:08:11something similar or by using client
2:08:14python libraries so for example if you
2:08:16have a thirdparty SAS product like
2:08:19HubSpot a lot of the big ones have
2:08:21python libraries that you can use to
2:08:23bring data in as well as a lot of the
2:08:25Azure tooling so Azure data Lakes for
2:08:28example have a python client that you
2:08:30can use to bring data in that's another
2:08:33option with the notebook so some of the
2:08:34pros here is again it's really good for
2:08:37extraction from apis because you know if
2:08:40you know how to code with python you can
2:08:42do quite customized logic around things
2:08:45like authentication imagination and that
2:08:47becomes really simple in a notebook if
2:08:49you want to be using any of those client
2:08:51libraries so anything from Azure is also
2:08:54really good or HubSpot as we mentioned
2:08:55previously now it's good for code reuse
2:08:58so a notebook can be parameterized and
2:09:01then used in lots of different
2:09:02situations you can also embed data
2:09:04validation and data quality testing into
2:09:07the incoming data so we done quite a lot
2:09:09on this channel around data quality and
2:09:11validating incoming data so that becomes
2:09:13a lot easier in a notebook and in terms
2:09:16of the performance well a few people
2:09:19have been testing the different
2:09:20performance of these different meth
2:09:21methods and the notebook always comes
2:09:23out on top really in terms of speed and
2:09:26therefore also capacity usage so if
2:09:29you're really sensitive around the
2:09:31amount of capacity uh units that you're
2:09:33using then notebook is going to be the
2:09:35most efficient and I'll leave a link to
2:09:37a bit of an analysis done on the Lucid
2:09:40bi blog and it goes through some of the
2:09:41investigation work that they've been
2:09:43doing there it's a really good blog if
2:09:44you want to learn more about that now
2:09:46some of the cons well when you don't
2:09:48have a python capability in your
2:09:49organization that might sound a bit of
2:09:51an obvious one but if you don't have a
2:09:53team to write and then support these ETL
2:09:57notebooks then it's not going to be a
2:09:58good choice for you and secondly one of
2:10:00the limitations with the notebook is you
2:10:03can't actually currently write into a
2:10:05data warehouse so if that's your
2:10:07destination then you're going to be wan
2:10:10to using other tools for data ingestion
2:10:12not the notebook okay the other method
2:10:14that we can use to bring data into
2:10:16Fabric or at least make data accessible
2:10:19from within fabric is the shortcut and
2:10:21we've talked quite a lot about shortcuts
2:10:23on this channel so far so what we're
2:10:25going to be doing is just a bit of a
2:10:26review of what's possible with a
2:10:28shortcut and some of the things that you
2:10:29need to bear in mind for the exam so the
2:10:31first thing is kind of a bit of an
2:10:32overview really a shortcut enables you
2:10:34to create a live link to data stored in
2:10:37another part of fabric which is an
2:10:40internal shortcut or in the following
2:10:42external storage locations so ADLs gen2
2:10:45Azure data Lake storage Amazon S3 other
2:10:48services that use Amazon S3 for storage
2:10:51for example Cloud flare buckets also use
2:10:54Amazon S3 and that's quite a big section
2:10:57of the market there's a lot of tools
2:10:58that use Amazon S3 for their storage and
2:11:01that is now opened up for shortcuts
2:11:03Google Cloud Storage is another one and
2:11:05also tables in the data verse now a
2:11:07shortcut can be set up for individual
2:11:09files but also for a folder so if you
2:11:12set up a shortcut to a folder it's
2:11:14basically going to bring in all of the
2:11:16files that are in that folder and sync
2:11:19them now one thing to be careful of is
2:11:20the cross region egress fees so if your
2:11:23fabric capacity is in UK South Region
2:11:26for example but your ADLs storage
2:11:29account is in West us then you're going
2:11:31to be charged cross region ESS fees by
2:11:35Azure basically and that's 1 cent per
2:11:38gigabyte of data that's transferred now
2:11:40you can also create shortcuts now via
2:11:42the fabric rest API as well so if you
2:11:45want to do some sort of programmatic
2:11:46creation of shortcuts maybe you want to
2:11:49do hundreds of tables shortcuts cutting
2:11:51in one go then that's an option for you
2:11:53there now something that might come up
2:11:54in the exam is around permissions for
2:11:57shortcuts so what I've done is I've just
2:11:59copied the documentation piece here and
2:12:01I'll link to this in the school
2:12:02Community as always so what this table
2:12:04is showing is the shortcut related
2:12:06permissions for each workspace role so
2:12:10starting at the top there we've got
2:12:11creating a shortcut well to be able to
2:12:13create a shortcut the user needs write
2:12:15permission in the place that they're
2:12:17creating the shortcut and also read
2:12:19permission of the file or the that
2:12:21they're shortcutting too okay so these
2:12:23are the two permissions that you need
2:12:24side by side to create a new shortcut
2:12:27secondly if you just want to read the
2:12:28file contents of a shortcut you're going
2:12:31to need read permissions in both the
2:12:33place where the shortcut lives but also
2:12:35where the reference file is living as
2:12:38well so you need at least read
2:12:40permission in both of these locations if
2:12:41we want to write new files or new data
2:12:43to a shortcut Target location you're
2:12:45going to need write permissions for both
2:12:48of those locations both where you're
2:12:50writing the shortcut data to and where
2:12:52that shortcut data is being read into as
2:12:54well so let's just talk now about
2:12:56deciding when to use which method
2:12:59because and we did touch on this in the
2:13:01first lesson in this series we talked
2:13:03about some of the deciding factors when
2:13:05we're planning out our fabric
2:13:06implementation right so talking about
2:13:08the storage where is it stored and also
2:13:11what skills exist in the team and that's
2:13:13a really good start but there's also
2:13:14some other factors that we need to bear
2:13:16in mind when we're thinking about
2:13:17choosing a data ingestion method so if
2:13:20you have a requirements for real time or
2:13:23near realtime data you're going to be
2:13:25wanted to prefer the options like the
2:13:28shortcut if it's a file or folders or
2:13:30database mirroring if it's in a table in
2:13:33any of the three database types that are
2:13:35supported by database mirroring because
2:13:37these are live links to those locations
2:13:40so when you query it it's going to have
2:13:42near realtime data coming back we talked
2:13:44about these skills in the team so if
2:13:45you've got predominantly no and low code
2:13:47users then the data flow in the day
2:13:49pipeline if you've got SQL based then
2:13:52you can use the data Pipeline and stored
2:13:55procedure activity or script activity
2:13:57and we'll look at some of that in the
2:13:59next lesson we'll talking more about
2:14:00store procedures and that kind of thing
2:14:02but it is possible to use for example
2:14:04the copy into statement for ingesting
2:14:07data from files like parket files or
2:14:10from CSV files into a data warehouse and
2:14:13if you got python or Scala skills then
2:14:16you can be using notebooks the next
2:14:18thing is around crossworks bace
2:14:19limitations as we mentioned before the
2:14:21data pipeline must be in the same
2:14:24workspace as your destination store when
2:14:27we're talking about data ingestion right
2:14:29so it needs to be in the same Works spay
2:14:30and the other two methods the data flow
2:14:33and the notebook they don't have those
2:14:35limitations just to be clear on that
2:14:37next talk about the scalability and the
2:14:39size of your data and cost and capacity
2:14:41usage and all these things are pretty
2:14:44tightly linked right so in general The
2:14:47Notebook from quite a few people's
2:14:50analyses that I've seen is the most
2:14:52efficient method I'm not sure if you'll
2:14:53be tested on this in the exam but just
2:14:55something to bear in mind for when
2:14:56you're actually designing Solutions okay
2:14:58so now let's test some of our knowledge
2:15:01in this part of the exam when we're
2:15:03talking about getting data into fabric
2:15:05question one you're trying to create a
2:15:07shortcut to a folder of CSV files in
2:15:10Azure data L Storage Gen 2 which of the
2:15:12following is a valid connection string
2:15:14you can connect to A B C or D now I'll
2:15:18pause the video here have a bit of a
2:15:19think and I'll reveal the answer to you
2:15:21shortly so the answer here is D so in
2:15:24Azure there's a number of different end
2:15:25points that we can connect to for a
2:15:27storage account and the one that we need
2:15:29for a shortcut is the DFS the
2:15:32distributed file system path which is D
2:15:35the database. windows.net well it's not
2:15:37a database so it's not going to be A and
2:15:39B and C are both end points of a storage
2:15:42account but it's not the ones that you
2:15:43need to create a shortcut question two
2:15:45you're implementing the first stage in a
2:15:47medallion architecture your goal is to
2:15:49retrieve data from a rest API using a
2:15:52get request and you're going to save the
2:15:53raw Json response in the files area of a
2:15:56bronze Lakehouse which of the following
2:15:57methods can you use to achieve this now
2:16:00there's three correct answers here is it
2:16:02a the data pipeline copy data activity B
2:16:04the data flow with the web API connector
2:16:06C the event stream B the data pipeline
2:16:08web activity or E the fabric notebook so
2:16:11there's three methods here a the data
2:16:13pipeline copy data activity D the data
2:16:16pipeline web activity and E the fabric
2:16:19notebook now one of the important an
2:16:21parts of this question is that we want
2:16:22to save the raw Json response in the
2:16:25files area of a bronze Lakehouse so
2:16:29we're not going to be ingesting it
2:16:30directly into a table we want to save
2:16:32the raw Json and these three are the
2:16:34only three where that's possible the
2:16:36data flow we have to Output it into a
2:16:38data store so into a lake house table or
2:16:41into a data warehouse table for example
2:16:43so we need to perform some sort of
2:16:44transformation on that Json we can't
2:16:46just write out the raw file same with
2:16:48the event stream and the other three
2:16:49does allow us to do that so the data
2:16:51pipeline you can output in various
2:16:53different file formats and with the
2:16:55fabric notebook you can do that as well
2:16:57question three your goal is to extract a
2:16:59CSV file in Azure blob storage and write
2:17:03it to a fabric data warehouse table
2:17:06which of the following methods can you
2:17:07not use to achieve this data pipeline
2:17:09copy data activity B data flow with a
2:17:12blob storage connector C in a fabric
2:17:14notebook use the Azure blob storage
2:17:16client library for python get the file
2:17:18and write the data into the data we
2:17:21table D use the copy into statement in
2:17:23tsql from within your fabric data
2:17:25warehouse so the answer here is C which
2:17:28of the following methods can you not use
2:17:30to achieve this so you can't as we
2:17:32mentioned in the lesson today you can't
2:17:34use a notebook to write data directly
2:17:38into a fabric data warehouse so that was
2:17:41the clue in this question all of the
2:17:43other three methods so the data pipeline
2:17:45the data flow and the copy into
2:17:47statement in a tsql script they can be
2:17:49used to ingest data into a fabric data
2:17:53wouse table question four workspace a
2:17:55contains lake house a workspace B
2:17:58contains lakeh house B you want to
2:18:00create an internal shortcut in Lakehouse
2:18:03a pointing to a table in lakeh house B
2:18:06what's the minimum level of workspace
2:18:08permissions you need to achieve this a
2:18:10contributor in workspace a and viewer in
2:18:13workspace b b contributor in workspace a
2:18:17and contributor in workspace b c member
2:18:20in workspace a and contributor in
2:18:22workspace b or d viewer in workspace a
2:18:26and viewer in workspace B so the answer
2:18:28here is contributor in workspace a and
2:18:30viewer in workspace B so when we were
2:18:33talking about the workspace roles
2:18:35permission required to create a new
2:18:38shortcut where you need some sort of
2:18:39write permission in the workspace where
2:18:41you're creating the shortcut and read
2:18:44permissions in the lake house that
2:18:45you're actually referencing so the
2:18:47important part of the question here is
2:18:48what's the minimum leval of workspace
2:18:51permissions that you need to achieve
2:18:52this so the others or B and C at least
2:18:55would allow you to do this but it's not
2:18:57the minimum level of permissions that's
2:18:59required and D viewer in workspace a if
2:19:02we're a viewer in workspace a then you
2:19:03won't have the permissions to create the
2:19:05shortcut in Lake housee a number five
2:19:08one of your team is a superstar using
2:19:10mcode for data extraction in which of
2:19:12the following data extraction tools can
2:19:14you write M code is it a the fabric
2:19:16notebook b in a tsql script in the data
2:19:19warehouse C in a data flow or D in a
2:19:23data pipeline mapping data flow so the
2:19:25answer here is C the data flow gen to
2:19:29that's the one with the power query
2:19:31interface and you can write in the
2:19:33advanced editor you can write M query
2:19:35obviously not in the fabric notebook or
2:19:37a tsql script and the data pipeline
2:19:39doesn't actually have a mapping data
2:19:41flow activity so that's a bit of a red
2:19:43herring that's coming from Azure data
2:19:45Factory where that did exist but not in
2:19:47fabric congratulations you've completed
2:19:49the first part in section two of the
2:19:52study guide preparing and serving data
2:19:54in the next lesson we're going to be
2:19:55looking at the fabric Data Warehouse in
2:19:58more detail and we're going to be
2:19:59looking at how you can schedule all of
2:20:01your ETL workloads whether that be a
2:20:04data flow a notebook or a data pipeline
2:20:06so click here to continue your dp600
2:20:09Learning Journey I'll see you there hey
SQL, Data Warehouse and scheduling
2:20:11everyone welcome back to the channel
2:20:12today we're continuing the dp600 exam
2:20:15preparation course and we're going to be
2:20:17looking at SQL the data warehouse and
2:20:20how we can schedule things to run in
2:20:23Microsoft fabric this is video six in
2:20:26our Series so we're nearly halfway we've
2:20:28covered a lot of ground already but
2:20:29there's some way to go still until we
2:20:31get to the end of the course in this
2:20:32video we're going to be covering
2:20:33creating views functions and stored
2:20:36procedures in that data warehouse
2:20:38experience how we can add stored
2:20:40procedures notebooks data flows to a
2:20:43data pipeline then how we can schedule
2:20:45data pipelines and also schedule things
2:20:47like data flows and notebooks and
2:20:50throughout this section of the course
2:20:51obviously we're going to be focusing on
2:20:53what you need to know to prepare for
2:20:55this dp600 exam a lot of these topics
2:20:58can go really deep but we're just going
2:20:59to go through what I think would be
2:21:01sensible to learn for the exam in case
2:21:03they come up now this lesson will be
2:21:05pretty much 100% practical so we're
2:21:07going to be going into fabric having a
2:21:09look and creating all of these things
2:21:11ourselves at the end of the lesson we'll
2:21:13be doing four sample questions and as
2:21:15ever I've got some key points and links
2:21:18to further learning resources if you
2:21:20want to learn more about a particular
2:21:21topic and really brush up on your skills
2:21:23in a particular area for the exam okay
2:21:25so let's start off by looking at
2:21:27creating functions and stored procedures
2:21:30and Views in the data warehouse
2:21:32experience what all these things are
2:21:34when you should be using them how you
2:21:36should be using them and things you need
2:21:38to bear in mind for the exam so we're
2:21:39starting off in a workspace here and
2:21:41I've created some fabric items the most
2:21:44important one being this data warehouse
2:21:46dp600 data warehouse if we open that up
2:21:48and have a quick look around just
2:21:50created some really simple tables we're
2:21:52going to be using one or maybe two of
2:21:53these tables for this tutorial
2:21:56specifically this dbo do employees table
2:21:59so only got four rows but that's good
2:22:00enough just to show you the
2:22:02functionality of a function and a store
2:22:04procedure and a view so what we're going
2:22:06to be doing is we're not actually going
2:22:07to be working in the fabric online
2:22:10experience we're going to be using SQL
2:22:12Server management studio so what we're
2:22:14going to be needing to do is go into the
2:22:15settings of the data warehouse collect
2:22:17this SQL connection string copy that and
2:22:20then then go over to SQL Server
2:22:22management Studio you can download that
2:22:24for free I'll leave a link in the
2:22:25description or in the school community
2:22:27and then we can connect to our fabric
2:22:29Data Warehouse from within SQL Server
2:22:32management studio now I've already set
2:22:34up the connection here but if it's your
2:22:36first time using SS SMS and you haven't
2:22:38connected it before just click on
2:22:39connect to a new database Engine add in
2:22:42your server name here use the Microsoft
2:22:45entra multiactor authentication it will
2:22:48ask you to authenticate with fabric and
2:22:50the on line experience and then you
2:22:51should see your databases listed here
2:22:54we've got this dp600 data warehouse that
2:22:56we were looking at previously and you
2:22:58can see that we've got some tables we've
2:22:59got the employees table the gold table
2:23:01and some other tables in here as well so
2:23:03let's start off just by exploring what
2:23:06we've got here right so let's just do a
2:23:09very simple select star from db.
2:23:11employees and that's just going to
2:23:13return all of that data just so we can
2:23:14have a quick look at it make sure it's
2:23:15all connected correctly yeah so we've
2:23:17got our four rows there three columns
2:23:19employee ID D name and age perfect so
2:23:23say we've got this table here db.
2:23:24employees and for a powerbi report that
2:23:27we're wanting to create we actually want
2:23:29to do a bit of transformation on this
2:23:31table we don't want the raw table just
2:23:34as it looks like here we want to do some
2:23:36sort of transformation it doesn't really
2:23:37matter what that transformation is for
2:23:38this demo maybe we want to do a where
2:23:41statement so where name is like Jack for
2:23:45example this is just going to bring us
2:23:46back all of the employees that are
2:23:49called Jack and we're using this
2:23:51percentage Wild Card operator here and
2:23:53what that means is if the first four
2:23:55characters are j a c k anything after
2:23:58that is going to return in our results
2:24:00set here as you can see we've only got
2:24:02one Jack in the data set so that's
2:24:03coming back correctly now in our powerbi
2:24:05imagine we want to query just this
2:24:08results set but what we can do is we can
2:24:11save this query as a view now this is a
2:24:14bit of a toy example but sometimes you
2:24:16have a view that might have really
2:24:17complex transformation right it's not
2:24:19stored in the database but whenever we
2:24:22query it we want the transformations to
2:24:24be done at query time so all you can do
2:24:27to create a view of this is add in
2:24:29create view give the viewer name dbo do
2:24:32employee get Jack and then use the
2:24:35keyword as right then everything that
2:24:37follows that is going to be part of your
2:24:39view and if we just execute that we can
2:24:42see that that's been successfully
2:24:44completed and then we can just do select
2:24:46star from dbo do view employee get Jack
2:24:50and if we execute that we get the same
2:24:52result as before so now when you're
2:24:54creating your power report you can just
2:24:56query this view rather than the
2:24:58underlying table to get the transformed
2:25:00data back okay so that's the basic
2:25:03Syntax for creating a view what I've
2:25:05done is I've just written down some key
2:25:07points to understand just in general but
2:25:09also for the exam as well so for aiew
2:25:12the transform data is not stored right
2:25:15it's just the transformation logic and
2:25:17the code the SQL code is stored and then
2:25:19every time you query it maybe in a
2:25:21powerr report is going to perform that
2:25:23transformation at query time now you can
2:25:25create a view from another view so you
2:25:28can query another view from for example
2:25:31this view employee get Jack we can query
2:25:33that if we create another view that's
2:25:35possible we do have to be careful with
2:25:37performance sometimes when you string
2:25:40multiple views together and they're
2:25:41doing lots of heavy processing then it
2:25:43can have an impact on performance and it
2:25:46also can be difficult to understand and
2:25:47maintain if you're constantly querying
2:25:49other views from other views that's
2:25:51something to bear in mind now with a
2:25:52view we can't specify any sort of
2:25:54parameters okay it's just simply select
2:25:57statements you're just going to be able
2:25:58to query it like so and whatever the
2:26:00results is you get those back now this
2:26:03is possible with functions and stored
2:26:05procedures as we're going to see in a
2:26:06minute as we mentioned reading a SQL
2:26:08view into semantic model will fall back
2:26:10to direct query mode so direct like mode
2:26:13is not possible and that makes sense
2:26:15right because direct Lake works by
2:26:18reading the underlying Delta tables and
2:26:20in a view those Delta tables don't exist
2:26:23right so that's something to bear in
2:26:24mind with with a view so if that's a
2:26:26view now let's look at a function so say
2:26:30we want to parameterize that view that
2:26:33we had right we want to pass in a
2:26:36parameter maybe first name so that we're
2:26:38not just getting back Jacks we can
2:26:40potentially use it to get a list of
2:26:42employees given any first name let's
2:26:44have a look at what that might look like
2:26:46so this is a declaration of a function
2:26:49now you notice it's similar in some ways
2:26:51to the view but there are some key
2:26:53differences we're going to start by
2:26:54calling create function we give the
2:26:56function a name this one's function
2:26:58employee first name search then we're
2:27:01going to open some brackets and the
2:27:03brackets are really important in a
2:27:04function it's a bit like a function in
2:27:07Python for example you're going to open
2:27:09those brackets and pass in a parameter
2:27:11now the parameters are optional you
2:27:13don't have to pass in a parameter but it
2:27:15is an option we're defining our
2:27:17parameter with this at symbol so at
2:27:19first name and we give it a data type
2:27:21just for our char2 and this is a default
2:27:23value this is saying we have one
2:27:25parameter and the parameter name is
2:27:26first name the parameter type is far
2:27:29Char 20 and the default value is an
2:27:31empty string now the next really
2:27:33important piece of syntax here is this
2:27:35returns table so if you've used any sort
2:27:37of SQL functions in other flavors of SQL
2:27:40you know there's a few different types
2:27:42of functions now in fabric we're talking
2:27:45about table functions which means that
2:27:47it returns a table right we're not
2:27:49talking about Scala functions because
2:27:51that's currently not possible in
2:27:53Microsoft fabric version of tsql We're
2:27:56going to be calling a function and it's
2:27:57always going to return a table so in our
2:27:59syntax we have to Define that right we
2:28:00say returning a table as return Open
2:28:04brackets and then we can pass in
2:28:07whatever we want to declare in our
2:28:09select statement now this is very
2:28:11similar to our view declaration but here
2:28:13we're making it Dynamic right we're
2:28:15parameterizing it and we're using our
2:28:17first name parameter so let's just run
2:28:19this okay so so we've created our
2:28:20function let's just have a look at what
2:28:22that looks like so if we go to
2:28:24programmability in SQL Server management
2:28:26studio and in our functions you can see
2:28:28that in SQL Server management Studio
2:28:30they do actually make the
2:28:32differentiation between scalar valued
2:28:34functions and table valued functions so
2:28:36as we know R1 is going to be in this
2:28:38table valued functions here it is here
2:28:40dbo function employee first name search
2:28:43and we can call it like this select star
2:28:46from dbo our function name Open brackets
2:28:50with the parameter so if we call this
2:28:53execute and you see it Returns the same
2:28:56as what we had before but is
2:28:57parameterized so now instead of just
2:28:59Jack we've created a bit of a
2:29:01parameterized function here so we can
2:29:03call Jack we can also call Sarah you
2:29:06know whatever that needs to be we're
2:29:08just encapsulating all of that logic
2:29:10into a function and we parameterized it
2:29:12so that you can use it in multiple
2:29:14different places in your tsql code base
2:29:16and it just packages up that logic into
2:29:19a nice reusable function fun so what are
2:29:20some of the key points with a function
2:29:22as I mentioned it's useful for packaging
2:29:24up logic that you might want to reuse in
2:29:27different places and it can be
2:29:28parameterized now with a function you
2:29:30can only do select statements so you
2:29:31can't actually update any rows you can't
2:29:34do any sort of insert statements it's
2:29:36only ddl flavors of SQL only select
2:29:40statements so to call a function we can
2:29:42only really use a select statement like
2:29:45so or we can call it from within a
2:29:47stored procedure or from within a view
2:29:50which is kind of underneath this ddl as
2:29:52well if you've got a ddl view and by ddl
2:29:55I just mean select statements basically
2:29:57it can't be orchestrated directly right
2:30:00so we can't use it in a data pipeline
2:30:03directly although we can embed it in a
2:30:05stored procedure as we mentioned
2:30:07previously it allows one or more
2:30:08parameters specifically input parameters
2:30:11and the output type should always be a
2:30:13table okay so in fabric we're only going
2:30:16to using table valued functions and it's
2:30:18always going to return a table okay so
2:30:20that's the function now let's move on to
2:30:22the stored procedure so this is a very
2:30:26basic definition of a stored procedure
2:30:28we've got create procedure we give the
2:30:31procedure a name as and here we're just
2:30:32doing a select statement right and all
2:30:34this is going to do there's no
2:30:36parameters you'll notice when we execute
2:30:38a stored procedure we're just going to
2:30:39use this exec which is execute so that's
2:30:41one of the differences between a stored
2:30:43procedure and a function and a view the
2:30:45way in which we call it right so at the
2:30:47moment we're not really doing much with
2:30:48this store procedure it's just the most
2:30:50basic store procedure possible we're
2:30:52just passing in a select statement and
2:30:54we're executing it and we're getting
2:30:55back the results set the full results
2:30:57set now let's step it Upp a gear and add
2:30:59in a parameter so we've got a very
2:31:01similar thing to what we were looking at
2:31:03before create store procedure now we've
2:31:05got dbo Spore employee uncore get by
2:31:09first name and we're passing in we're
2:31:12declaring a parameter the parameter is
2:31:14called first name again it's going to be
2:31:16of type varar 20 underneath that we're
2:31:19declaring what what we actually want to
2:31:20do in this stored procedure so we want
2:31:22to select all of the employees where the
2:31:25name is like first name which is going
2:31:27to be passed in as the parameter and
2:31:29we're adding this kind of wild card
2:31:31operator onto the end so that anything
2:31:33that goes after that whatever the
2:31:34surname we're going to return that as
2:31:37well so let's just execute this to
2:31:39create the stored procedure ah yeah so
2:31:40let's just drop it drop procedure if
2:31:43exists dbo dot okay so now we've dropped
2:31:45it we can recreate it again just to show
2:31:47that it works completed successfully and
2:31:49then if we want to execute that again
2:31:51we're going to use the exec command pass
2:31:53in the name of our store procedure and
2:31:56our parameter which is Jack if we
2:31:57execute that we get this so that's the
2:31:59basic definition of a stored procedure
2:32:03what it looks like now let's think about
2:32:05some of the key points here so the store
2:32:06procedure we have the ability again to
2:32:08Define input parameters but also output
2:32:11parameters as well I didn't show you
2:32:12that in this lesson maybe that's one for
2:32:15another lesson but we can actually
2:32:16Define output parameters which are
2:32:18useful in the data Pipelines scenario
2:32:20maybe we'll look at that another time on
2:32:22the channel probably all you need to
2:32:24know for the exam is that it is possible
2:32:25now they're called using this EXA
2:32:28Command right it's not as part of a
2:32:31select statement you can call other
2:32:33store procedures from a store procedure
2:32:35so you can create three stored
2:32:37procedures and then create kind of like
2:32:39a master stored procedure that calls
2:32:41each of these other stored procedures in
2:32:43series now one of the key use cases of a
2:32:45store procedure is to give people access
2:32:49just to the stored procedure and not to
2:32:51the underlying data so you can use it as
2:32:54a security mechanism just by giving
2:32:56people access to the stored procedure
2:32:58and you know they're only getting access
2:33:00to the results of that store procedure
2:33:02and not any of the underlying data now
2:33:04one of the key points in fabric to
2:33:06understand is that the stored procedure
2:33:09is the well it's one of the only things
2:33:10that we can embed into a data Pipeline
2:33:13and we can pass the parameters that
2:33:15we've seen here from other notebook
2:33:17activities so that's something to really
2:33:20important to understand with store
2:33:21procedures they become a lot more useful
2:33:24because they can be orchestrated as part
2:33:27of a data pipeline then they become
2:33:29really useful for data transformation
2:33:31and all of these kinds of things in your
2:33:33data warehouse and in your architectures
2:33:35in general now the one I'm missing here
2:33:37is the ability to do inserts updates
2:33:41deletes now this is what's unique to the
2:33:43stored procedure at least in this list
2:33:45here is that here we're just using a
2:33:47select statement but we can use stored
2:33:49proced procedures for updating and
2:33:52inserting the underlying data sets so
2:33:54they can be really powerful tools that
2:33:57we can use for any sorts of data
2:33:59transformation data loading all these
2:34:01kind of things are possible with a store
2:34:03procedure okay so let's just focus now
2:34:05on the stored procedure and we mentioned
2:34:07here it can be embedded in a data
2:34:09pipeline so let's go back into fabric
2:34:11now and have a look at what that looks
2:34:13like okay so here we are in a data
2:34:16Pipeline and you'll notice if we go over
2:34:18to the activities tab there's a lot of
2:34:20different options for activities now the
2:34:22three that they mention for the dp600
2:34:24study guide are the stored procedure
2:34:26Activity The Notebook activity and the
2:34:29data flow activity here we're going to
2:34:31go through each one in a little bit of
2:34:33detail have a look at some of the
2:34:34settings that we can apply when we're
2:34:36creating these activities in a data
2:34:38pipeline then we're going to go on to
2:34:40look at how we can schedule these things
2:34:42starting with the store procedure so
2:34:43when we drop the store procedure
2:34:44activity onto the data pipeline canvas
2:34:47we can have a look at some of the
2:34:48settings that we get of the box right so
2:34:51most of the configuration happens within
2:34:53the settings tab you can navigate to
2:34:55your specific data warehouse that you
2:34:57want to connect to obviously noting that
2:34:59the data warehouse you connect to
2:35:01currently has to be in the same
2:35:02workspace as your data pipeline we've
2:35:05mentioned that quite a lot on the
2:35:06channel around the some of the crossw
2:35:08workpace limitations with the data
2:35:10pipeline so just be wary of that then we
2:35:12connect to the warehouse and then we can
2:35:14choose from a number of different stored
2:35:16procedures that are in that data
2:35:18warehouse we you can Define any sort of
2:35:20parameters that we've got here there is
2:35:22this option to automatically import
2:35:24parameters so it's going to look into
2:35:26your store procedure it's going to pick
2:35:27out any parameters that you've got there
2:35:30here it's found one called first name
2:35:32it's of type string it knows that and
2:35:34currently we're just hardcoding in a
2:35:35value here called Jack but one of the
2:35:38benefits of the store procedure is that
2:35:39you can add Dynamic content right so
2:35:41you're going to be able to parameterize
2:35:43that stored procedure activity and pass
2:35:46something into this value here using
2:35:48data pipeline parameters on the general
2:35:50tab we've got just the name of it so you
2:35:52can update the name you can make it
2:35:54active or deactive we can add a timeout
2:35:56so this might be quite useful for some
2:35:58data pipeline activities maybe you got a
2:36:00really long-standing stored procedure it
2:36:02might take half an hour to run you might
2:36:04want to add a timeout here at 1 hour
2:36:07because if it runs for 1 hour or longer
2:36:10than 1 hour you know that probably
2:36:11something has gone wrong and you don't
2:36:13want it to lock up your database your
2:36:15data warehouse so that's something to
2:36:16bear in mind here the the timeout
2:36:18another one is the retry so this retry
2:36:22setting is available on quite a few of
2:36:24the data pipeline activities and as the
2:36:26name suggests it's going to try it and
2:36:28if it fails that activity it will retry
2:36:31the number of times that you specify in
2:36:33this box here we can give it an a retry
2:36:35interval so we're going to say okay
2:36:36we're going to try it once then we're
2:36:38going to wait for 30 seconds and then
2:36:39try it again secure output and secure
2:36:42input as well this is just going to
2:36:44specify whether you want to Output into
2:36:46a log the results of that activity or in
2:36:50the input as well so that's the stored
2:36:52procedure activity now let's look at
2:36:53this notebook activity so we've got a
2:36:56notebook here and again if you go
2:36:58through to the settings you can Define
2:37:00your workspace and your notebook here
2:37:02I've got notebook load to Silver
2:37:05similarly with the stored procedure we
2:37:07can declare any parameters so if you've
2:37:09got a parameter cell in that notebook
2:37:12you can pass parameters into the
2:37:14notebook from other activities in your
2:37:16data Pipeline and again we've got very
2:37:18similar settings on the The Notebook as
2:37:20we had with the store procedure activity
2:37:22we've got a number of retries so with
2:37:24the notebook a retry is probably more
2:37:27important or more likely you're going to
2:37:28be using this because with a notebook
2:37:30you're obviously using the spark cluster
2:37:32and also potentially querying rest apis
2:37:35so there's a lot more that can go wrong
2:37:37I would say in a notebook than in a
2:37:38store procedure so the retry
2:37:40functionality here I think Microsoft
2:37:42recommends that you set the retry to two
2:37:44or three just so that if your spark
2:37:47cluster is busy and you're doing lots of
2:37:49computation on it and the session
2:37:51times's out or something goes wrong with
2:37:54your notebook execution it's going to
2:37:55retry it so always recommended to add in
2:37:58one or two retries into a notebook
2:38:00activity again we can specify the retry
2:38:02interval down here and secure input
2:38:04output so these are the same as the
2:38:06store procedure activity now the final
2:38:08one we're going to look at is the data
2:38:09flow activity here similarly you're
2:38:11going to go over to the settings table
2:38:13find your workspace and your data flow
2:38:15select a particular data flow that you
2:38:17want to run and that's basically if we
2:38:19go back to the general tab you'll see
2:38:21that the settings here are exactly the
2:38:23same we give it a name description we
2:38:24can set the activity state to active or
2:38:27deactive we can give it a timeout and a
2:38:29number of retries and the retry interval
2:38:32so that's a little bit about the stored
2:38:34procedure activity the data flow
2:38:36activity and the notebook activity in
2:38:39data pipelines now let's look at
2:38:40scheduling a number of these different
2:38:42items in Fabric and the different
2:38:43options that are available to us there
2:38:45so here we are back in our workspace and
2:38:47let's look at how we can do scheduling
2:38:49of different items within Microsoft
2:38:51fabric now there's a few different ways
2:38:54that we can schedule things to work in
2:38:56fabric these ETL jobs mainly we can
2:38:59schedule them to run maybe every hour or
2:39:02something like that depending on your
2:39:03use case we're going to be looking at
2:39:05scheduling data pipelines data flows and
2:39:07notebooks in a bit more detail here so
2:39:09with the data flow we can actually
2:39:10schedule from a workspace so we can
2:39:13click on the settings of a particular
2:39:16data flow just clicking on these
2:39:18ellipses three dotts here here clicking
2:39:20on the settings of that workflow go down
2:39:21to the refresh settings configure a
2:39:24scheduled refresh and we can turn that
2:39:26on and we can change the update
2:39:27frequency to daily or weekly or if you
2:39:31want more fine grained refreshes than
2:39:33that we can add specific times 1:00 a.m.
2:39:36maybe 2 a.m. 3:00 a.m. that kind of
2:39:39thing here now with the data flow the
2:39:40maximum number of Refreshers per day
2:39:43that you can do is 48 so that's in line
2:39:46with what you could do previously in
2:39:48palbi premium we can also send refresh
2:39:50failure notifications here to specific
2:39:52people or the owner or both let's just
2:39:55hop back to our workspace now and look
2:39:59at the notebook so with the notebook
2:40:01it's a similar story we can click on
2:40:04either schedule here or settings both
2:40:07take you through to the same thing here
2:40:09opens up this sidebar menu where we can
2:40:11specify the schedule and again we can
2:40:13repeat it every hour every day weekly or
2:40:17by the minute so every 5 minutes for
2:40:19example example we can specify an start
2:40:21time and an end time and the time zone
2:40:24that you want that schedule to be
2:40:25running on so this is one of the key
2:40:26differentiations between the notebook
2:40:28and the data flow obviously we can
2:40:31specify a frequency that's a lot more
2:40:33than 48 refreshes per hour if we're
2:40:36using this but bear in mind obviously
2:40:38it's going to use a lot more capacity
2:40:39units so if you don't need to refresh
2:40:41your data at this Cadence at this
2:40:43interval then undo it so that's the
2:40:45scheduling of notebooks now the data
2:40:48pipeline is obviously a orchestration
2:40:50tool so another way that we can do it is
2:40:52by putting our notebook and our data
2:40:55flows within a data Pipeline and then
2:40:58scheduling the data pipeline so you'll
2:41:00see here within the data pipeline we've
2:41:02got this schedule button again it's
2:41:03going to open up a sidebar menu we can
2:41:06click on here and specify by the minute
2:41:08hourly daily weekly again like so start
2:41:11and end time and the time zone so it's
2:41:13very similar to The Notebook scheduling
2:41:15functionality now if you're coming from
2:41:17ADF as a data Factory currently one
2:41:19thing to bear in mind is the only way of
2:41:21triggering a data pipeline is with
2:41:23scheduling right to schedule it to
2:41:25trigger a data pipeline using a schedule
2:41:27we don't currently have event based
2:41:29triggers or window triggering or HTTP
2:41:32request triggering all those different
2:41:35really useful functionality for
2:41:37triggering a data pipeline that does
2:41:39exist in ADF as a data Factory currently
2:41:42doesn't exist in fabric the only way we
2:41:44can use is this schedule trigger now you
2:41:46might be thinking okay when should I use
2:41:49the inbuilt scheduling for a data flow
2:41:51for example and when should I use a data
2:41:54pipeline for scheduling well it just
2:41:55gives you a bit more flexibility and
2:41:57functionality for handling different
2:41:59events if you use the day pipeline say
2:42:02for example you wanted to add some sort
2:42:03of activity on fail maybe you want
2:42:06custom notification or some sort of
2:42:09other activity or logging to be done
2:42:11that's possible obviously if you
2:42:12schedule it using an a pipeline if you
2:42:14schedule a data flow to refresh within
2:42:17the actual data flow you don't have
2:42:18access access to all this other
2:42:20functionality that you get in a data
2:42:21pipeline so that's why you might want to
2:42:23think about you know embedding your data
2:42:26flows into a data Pipeline and then
2:42:28triggering them from a data pipeline
2:42:30okay so that rounds up the content for
2:42:31this lesson now let's go back to the
2:42:33slides and test some of the knowledge of
2:42:35the things that we've learned in this
2:42:36section of the study guide okay so let's
2:42:39just round off this video with some
2:42:41practice questions number one the
2:42:43maximum number of scheduled refreshes
2:42:46allowed per day with the data flow Gen 2
2:42:49is a 12 B 24 C 48 D 96 or E unlimited
2:42:56refreshes per day pause the video here
2:42:59have a think and then I'll reveal the
2:43:00answer to you shortly so the answer here
2:43:02is 48 refreshes per day so you can get a
2:43:06data flow Gen 2 to refresh every half an
2:43:08hour if that's what you want to do now
2:43:11if you're coming from the powerbi
2:43:12premium World you'll be used to this the
2:43:14other figures are just incorrect really
2:43:16question two which of the following can
2:43:18you use used to update a row in a data
2:43:21warehouse table is it a a stored
2:43:23procedure B A View C A create table
2:43:26statement or d a function so the correct
2:43:28answer here is a stored procedure that's
2:43:31the only one that allows you to actually
2:43:33update data in a data warehouse table a
2:43:36view as the name suggests is just read
2:43:38only it creates a view on top of the
2:43:40data it doesn't actually modify the
2:43:41underlying data a create table statement
2:43:43is not going to be able to update a row
2:43:45and a function also cannot actually
2:43:48update the underlying data so the answer
2:43:50here is a a stored procedure in a stored
2:43:52procedure we have a lot of flexibility
2:43:55to insert into to update rows to delete
2:43:58rows all of that kind of DML is exposed
2:44:02in the stored procedure and available to
2:44:04us and by DML I mean data manipulation
2:44:06language it's a subset of SQL it's part
2:44:08of the SQL language question three which
2:44:11of the following statements is false
2:44:13when talking about operations in a
2:44:15fabric data warehouse a you can call a
2:44:17function from a stored procedure B you
2:44:20can call a function from A View C you
2:44:22can query A View From Another view D you
2:44:25can call a store procedure from a
2:44:27function so the answer here is D you can
2:44:29call a stored procedure from a function
2:44:32this is the only one that you actually
2:44:34can't do so you remember that with the
2:44:35stored procedure we have to use that
2:44:37exec the ex execute command to execute
2:44:40the stored procedure and you can't do
2:44:42that from within a function now you can
2:44:44call a stored procedure from another
2:44:46stor procedure but that's pretty much
2:44:47the only time when we can call another
2:44:49stored procedure from another item the
2:44:51others you can call a function from a
2:44:52stored procedure well we know we can do
2:44:54that you can call a function from A View
2:44:56yes you can do that and you can query A
2:44:58View From Another view you can also do
2:45:00that so the correct answer here is D
2:45:02that statement is false question four
2:45:04you orchestrating many spark notebooks
2:45:07to perform data transformation
2:45:09activities at the same time now you
2:45:10notice that sometimes the notebook
2:45:12execution is failing because your
2:45:14cluster is busy that's the error message
2:45:16that's it's giving you now what
2:45:17modification can you make to the data
2:45:19pipeline notebook activity to give the
2:45:21pipeline more chances to run
2:45:22successfully is it a deactivate and
2:45:25reactivate the activity B set the number
2:45:27of retries to two or three C change the
2:45:30parameters you're passing into the
2:45:31notebook or d add a failure activity to
2:45:34handle the execution failure so the
2:45:36answer here is B set the number of
2:45:38retries to two or three you'll remember
2:45:41that the the retry option in a notebook
2:45:45activity it allows the activity to retry
2:45:48if it fails a number of times so if your
2:45:51spark cluster is busy and it can't
2:45:53execute the first time it's going to
2:45:55wait and then retry it two three or any
2:45:58amount of times that you specify in that
2:46:00retry parameter in that retry setting
2:46:03deactivating and reactivating the
2:46:04activity that's not going to make much
2:46:06difference changing the parameters well
2:46:08you probably don't want to do that cuz
2:46:09it's going to change the output and
2:46:10adding a failure activity to handle the
2:46:12execution failure so that might be a
2:46:14good idea but it's not actually going to
2:46:16impact the result of your activity right
2:46:20it's not going to allow it to rerun and
2:46:22get a successful execution so the answer
2:46:24here is B congratulations you've
2:46:26completed the second part of section two
2:46:29preparing and serving data in the next
2:46:31lesson we're going to be deep diving
2:46:33into the exciting world of data
2:46:35Transformations within Microsoft fabric
2:46:38make sure you click this video here to
2:46:39join us in the next lesson hey everyone
Transforming data with Dataflows, PySpark, T-SQL
2:46:41welcome back and this is the next video
2:46:43in our dp600 exam preparation course
2:46:47this is video 7 we're going to be
2:46:49looking at transforming data and there's
2:46:51a lot to get through with these and it's
2:46:53a little bit overwhelming when you look
2:46:54at the things that we're going to be
2:46:56covering don't worry I'm going to break
2:46:57them down into a number of different
2:46:58sections hopefully to make it a little
2:47:00bit more digestible so we're going to be
2:47:02looking at First Data cleansing how can
2:47:05we Implement a data cleansing process
2:47:07we're going to look at resolving some
2:47:08common issues that we get with data so
2:47:11duplicates missing data null values
2:47:14conversion of data types and filtering
2:47:16data and we're going to be looking at
2:47:18how we can do that those things using
2:47:20the data flow tsql and Spark as well
2:47:23then we're going to move on to data
2:47:24enrichment so under this category we're
2:47:26going to be looking at merging and
2:47:28joining different data sets together and
2:47:30also enriching the data that we've
2:47:31already got so adding new columns new
2:47:33tables based on our existing data
2:47:36finally we're going to take a look at
2:47:37data modeling we're going to look at the
2:47:38star schema what that is we're going to
2:47:40look at type one and type two slowly
2:47:42changing Dimensions we're going to look
2:47:44at the bridge table and what problem
2:47:46that that solves and how we can
2:47:48implement the solution solution using
2:47:49tsql we're going to look at data
2:47:51denormalization Aggregate and
2:47:53deaggregating data as well as ever we're
2:47:55going to be testing your knowledge at
2:47:57the end of this video so we've got some
2:47:58sample questions and all of the key
2:48:01points for this section of the exam plus
2:48:03links to any further learning resources
2:48:05if you want to dig into a bit more
2:48:07detail into any of these topics they're
2:48:08going to be posted on the school
2:48:10Community I'll leave a link in the
2:48:11description below so first let's take a
2:48:13look at the data cleansing process and
2:48:16to give a bit of structure as to what we
2:48:18looking at here I'm going to be talking
2:48:20through the lens of The Medallion
2:48:22architecture whereby we have bronze
2:48:24silver and gold areas in our data
2:48:27processing pipeline now when we talk
2:48:29about data cleansing normally this takes
2:48:31place anywhere really between the bronze
2:48:34and the silver and maybe in the silver
2:48:36as well this data cleaning CU we're
2:48:38going to get some raw data in bronze but
2:48:41it's going to be messy so we want to be
2:48:42doing all of our data cleaning steps
2:48:44normally between bronze and silver so
2:48:47the way that I'm going to do this is to
2:48:48walk you through how we can do common
2:48:51data cleansing routines and operations
2:48:54within all three of the tools listed
2:48:56there so starting with the data flow
2:48:58then we're going to look at how we can
2:48:59do all these things with tsql in the
2:49:01data warehouse and also in a spark
2:49:03notebook as well so let's jump into
2:49:05Fabric and start with data cleansing in
2:49:08a data flow okay so for this part of the
2:49:10lesson we're going to be jumping into
2:49:12Fabric and we're going to be
2:49:13transforming a particular data set here
2:49:16and it's that car dealership sales data
2:49:19set now the data set itself comes from
2:49:21this blog here Tableau server Guru so
2:49:24thanks very much to this person who has
2:49:25some sample data sets on their website
2:49:28don't worry we're not going to be
2:49:28talking about Tableau servers or
2:49:30anything like that but they do have this
2:49:31nice car sales data set that I
2:49:33downloaded that comes in a pretty good
2:49:35star schema and when you download this
2:49:37it comes actually in
2:49:39xlsx now this data set is actually quite
2:49:42a clean data set so have actually made
2:49:44some modifications to make it a bit more
2:49:46dirty so that we can perform some data
2:49:48cleaning on this data set and I'll leave
2:49:50a link to the dirtified data sets if you
2:49:52can call them that I'll leave them in
2:49:54the school community so go over to there
2:49:56to get the source data files that I used
2:49:58here and to get us started what I've
2:50:00done is I've just put them into a
2:50:02Lakehouse files area so it's number of
2:50:04these csvs and then I've just loaded
2:50:06them simply into a number of tables so
2:50:08these are going to be our bronze
2:50:09Lakehouse tables that we're going to be
2:50:11using for this part of the tutorial so
2:50:13here we are in the data flow Gen 2 and
2:50:16I'm just going to show you how I can do
2:50:17some data cleansing in within the data
2:50:20flow itself I just pulled in one of
2:50:22those tables from our Lakehouse area
2:50:24which this is the source here it's in
2:50:26our lake house and I pulled in this
2:50:27Revenue table so this is what it looks
2:50:29like basically completely untransformed
2:50:33now you notice I've introduced some null
2:50:34values here I've introduced a few
2:50:36duplicate values as well so what we're
2:50:38going to do is just step through the
2:50:40different transformation steps to clean
2:50:42this data set within the power query
2:50:45experience within the data flow so the
2:50:47first thing that we might want to do is
2:50:50remove any duplicate values that we've
2:50:52got so I know because I've introduced
2:50:54some duplicates into this data set there
2:50:56are duplicates in here so to remove any
2:50:58duplicates in this power query engine
2:51:01what we're going to do we're just going
2:51:02to highlight all of these columns
2:51:04clicking on the left hand column holding
2:51:06shift and then clicking on the right
2:51:07hand column that's going to select all
2:51:09of our columns then we can right click
2:51:11on the top here and then remove any
2:51:13duplicate rows so this is important
2:51:15because we want to check that the whole
2:51:17row is a dup so we need to select all of
2:51:19the different columns here and then
2:51:21remove the duplicates like that and you
2:51:22can see it's been added into this
2:51:24applied steps here so if we take a look
2:51:26at our Revenue column here you can see
2:51:28that we do actually have some null
2:51:29values in this data set now obviously it
2:51:32depends on your use case for data
2:51:34cleansing but if this Revenue value is
2:51:37actually really important and this is
2:51:39the only thing that you really care
2:51:40about in that data set then you might
2:51:41want to remove these NS all together
2:51:44from the data model again it depends
2:51:46very much on your use case so to remove
2:51:48the values in this Revenue column we
2:51:49just click on the column itself and then
2:51:52deselect this null value press okay
2:51:54that's going to remove those null values
2:51:57from the data set now another thing we
2:51:58can do is to change the type so if you
2:52:01want to change type in the data flow
2:52:03simply click here on change type now it
2:52:06might not make sense for this particular
2:52:08column because this is already looks
2:52:09good looks like an integer value and
2:52:11it's got an integer type here whole
2:52:12number but say for example this was a
2:52:15text value and you know that deep down
2:52:17is actually an integer then you can just
2:52:19right click on that change type to any
2:52:21of these types here now another thing we
2:52:23can do with the data flow Gen 2 is add
2:52:27in new columns so you can see here add
2:52:29column and what we can do is add a
2:52:31custom column and maybe we want to do
2:52:33some sort of Revenue bin uh maybe you
2:52:36want to add in some logic here to kind
2:52:37of group these revenues into something a
2:52:40little bit different I'm just going to
2:52:42do a simple one here revenue is greater
2:52:44than 1 million and we're going to give
2:52:46it the data type as a true or false and
2:52:48then if we click okay that should give
2:52:50us this True Value just to check some of
2:52:53them are false so yeah that's happening
2:52:55correctly so now maybe we want to change
2:52:57the name of that something a bit more
2:52:58descriptive Revenue over 1 million maybe
2:53:01that's probably a better description for
2:53:03what that new column is you know another
2:53:06way we can filter these data sets maybe
2:53:07you don't want actually want all of
2:53:09these units sold maybe perhaps you're
2:53:11just for this particular piece of
2:53:12analysis that you're doing or this data
2:53:14set you want to actually remove the unit
2:53:16sold over one and again just showing the
2:53:19functionality here really depends on
2:53:22what your use case is in your business
2:53:24as to which data cleansing steps are
2:53:25going to make sense so that's a bit of
2:53:27an overview of power query and how we
2:53:29can do data cleansing in the data flow
2:53:31Gen 2 now let's jump over to the data
2:53:33warehouse and look at how you can
2:53:34Implement similar data cleansing process
2:53:36using tsql okay so I've just jumped over
2:53:38to the SQL end point here in that same
2:53:42bronze Lakehouse because here we've got
2:53:44the same tables here but now we're just
2:53:46in the SQL endpoint experience so we can
2:53:48write some tsql and to begin with let's
2:53:50just have a look at our data set here
2:53:52just make sure that that is coming
2:53:53through okay yep looks exactly the same
2:53:56as we had in the data flow so that's
2:53:58good here so the first step we're going
2:53:59to look at here is identifying
2:54:01duplicates in this table now there's
2:54:04many different ways that we can use to
2:54:06identify and remove duplicates from a
2:54:10table using tsql now this method here
2:54:12I'm using is Group by so what we're
2:54:15going to be doing is grouping by
2:54:16something that I know is unique or
2:54:19should be unique actually for every Row
2:54:22in this fact table and we're going to be
2:54:24using having where a count of more than
2:54:26one so what this is going to mean is
2:54:29that for each of these groups there
2:54:30should be exactly one row so if this
2:54:34returns a count of more than one for
2:54:37this group then we'll see that that is
2:54:40in fact a duplicate row so let's just
2:54:42have a look at these so we can see that
2:54:44here for dealer these two dealer IDs and
2:54:46these two dates we can see there's
2:54:48actually three rows here so if we just
2:54:51give this bit of number of rows just to
2:54:54make that a bit clearer and then we
2:54:55rerun that yeah so now we can see our
2:54:57number of rows is three and four so how
2:55:00do we go about removing those duplicates
2:55:02well again what way of doing that is by
2:55:05using the group buy again we can Group
2:55:07by the same two columns here the DLo ID
2:55:10and the date ID and we can just bring
2:55:11back a aggregate function this is an
2:55:13aggregate function the max of the
2:55:15revenue and it's worth checking before
2:55:17you do this that the revenue figures for
2:55:19each of these rows are actually the same
2:55:21so the max or the Min doesn't actually
2:55:23make a difference we're just going to
2:55:24bring through whatever that value is and
2:55:26the result of that is going to be the
2:55:28same data set but with the duplicates
2:55:30removed which is this one here and then
2:55:32you can save that into another table or
2:55:35whatever Downstream activities you want
2:55:36to do with it so next let's look at
2:55:38missing data nulls removing nulls
2:55:41filtering that kind of thing here so
2:55:43there are some null values in this
2:55:45Revenue column so we're just going to
2:55:46use where revenue is null to identify
2:55:49those first things first so you can see
2:55:51here we've got six rows here where the
2:55:54revenue is null so again you might want
2:55:55to remove those and obviously to remove
2:55:58those what it's pretty simple we can
2:55:59just use is not null and that will
2:56:01return you all where the revenue is not
2:56:04null so one way can we can just quickly
2:56:06verify that that is actually the case is
2:56:08if we just bring this into a bit of a
2:56:10CTE with remove nulls as this we do a
2:56:14select count star remove nulls we'll
2:56:17also do a select
2:56:18count star from the original table let's
2:56:22just compare these two values oh yeah I
2:56:24have to actually call it so we have
2:56:25result one is
2:56:281855 result two is 1861 so we can see
2:56:31that the row count Has Changed by six
2:56:34and those are those six null values that
2:56:36we saw when we did the where revenue is
2:56:38null so that's worked correctly so
2:56:40another thing that we need to bear in
2:56:41mind for the exam is tsql type
2:56:44conversion so here you can see what
2:56:46we're using is the cast function
2:56:48function and here we're going to cast
2:56:49the values in the revenue column as a
2:56:52float I think currently it is yeah I
2:56:54think currently it's an integer and so
2:56:56what we're doing here is we're casting
2:56:58it as a float so basically a decimal
2:57:01number so by doing that you can see here
2:57:02that when we bring both of these columns
2:57:04we've got the original Revenue here in
2:57:06this column and that's an integer and
2:57:07here it's changed the data type into a
2:57:09float using this cast functionality
2:57:12another thing we've done here is to add
2:57:14another column so when you're using ttal
2:57:16you can just add in another row here
2:57:18into your SQL scripts and you can do
2:57:20whatever you want here this is just a
2:57:22bit of an example here divided the
2:57:24revenue by two and I've given it a name
2:57:26of half the revenue so that's just an
2:57:28example of adding new columns that we
2:57:31can do in tsql okay so just to round off
2:57:34this part of the tutorial next we're
2:57:35going to look at transforming data using
2:57:38pypar notebooks and specifically we're
2:57:40going to be looking at some of the data
2:57:42cleansing routines that we've been
2:57:43looking at in the data flow and the tsql
2:57:46engine now we're going to look at how we
2:57:47can implement them in a spark notebook
2:57:50now I did do a 3.5 hour tutorial which
2:57:53goes into a lot more depth about
2:57:55different data cleansing operations you
2:57:58can have a look at that video here I'll
2:58:00leave a link in the school Community if
2:58:01you want to have a look at that in a bit
2:58:03more detail but here we're just going to
2:58:04go on a bit of a quick Deep dive into
2:58:06some of the most common operations and
2:58:08I'll also leave this notebook on our
2:58:11school community so if you want to play
2:58:13along at home then you can do that as
2:58:14well so we're using this same bronze
2:58:18lake house we've got our tables here I'm
2:58:20just going to be having a look at this
2:58:21Revenue table which is our fact table in
2:58:23a bit more detail and I've begun by just
2:58:26reading it into a spark data frame and
2:58:28displaying the results here just to
2:58:30check that everything's loading okay and
2:58:31got everything logged in here now you
2:58:33notied that I haven't actually committed
2:58:35or haven't actually written any of those
2:58:38changes that we made in the data flow
2:58:39and the tcq engine so this is still
2:58:42reading from the raw data here in that
2:58:44bronze layer so let's start by looking
2:58:46at duplicate data again to identify
2:58:49duplicates we can use a similar kind of
2:58:50pattern that we used in tsql by looking
2:58:53at group by or by using Group by and
2:58:55then looking at count of more than one
2:58:58so here you can see we've identified the
2:59:00rows where we have some duplicates in
2:59:03that data set and luckily it's the same
2:59:06as what we were seeing in our tsql
2:59:09engine so we can see that these Branch
2:59:11IDs and date IDs have a count of four
2:59:13and three respectively how do we go
2:59:15about dropping those duplicate values
2:59:18from our data set well in spark we have
2:59:21this drop duplicates method so what I'm
2:59:23going to do here is just start by
2:59:25counting the rows in the original data
2:59:28set performing this drop duplicates
2:59:30saving it into D duped which is a new
2:59:33data frame and then counting the rows in
2:59:35that new duped data frame okay so I've
2:59:37just printed out the result here we can
2:59:39see that the process has removed if I
2:59:41could spell process correctly it's
2:59:42removed five rows from the data set so
2:59:44I've just taken the difference really
2:59:46between our original data set and that
2:59:48that duped data set so that makes sense
2:59:50CU we've got seven count here originally
2:59:53and obviously two of those we want to
2:59:55keep right because they are actually
2:59:56valid data so we've removed five rows so
2:59:59five of them are duplicates so that
3:00:01looks like it's worked correctly and if
3:00:03we were just going to verify again that
3:00:05that's worked we can perform the same
3:00:07operation that we did before grouping by
3:00:09these two Fields looking at where the
3:00:12count is more than one and we've got
3:00:13this empty data frame here so that looks
3:00:15like it's worked correctly next we're
3:00:17going to take a look at missing data and
3:00:19how we handle nulls and that kind of
3:00:21thing in spark so first off we're going
3:00:23to look at identifying missing values
3:00:26you might want to do a bit of like
3:00:27inspection of your data before you go
3:00:29ahead and drop things so a good idea to
3:00:32interrogate your data a bit have a look
3:00:34at what you might be deleting so let's
3:00:37just run this and then we'll talk
3:00:38through it so we've got our original
3:00:39data frame here and we're filtering on
3:00:42this specific column here so we're
3:00:44passing in DF do revenue. isnull so so
3:00:48this is a method that we get on a spark
3:00:51column and we can see that it's
3:00:52returning these six rows that we know
3:00:54are null now another way we can do that
3:00:57is I've changed two things in this
3:00:59second example here but it's basically
3:01:00using DF do Weare so DF do filter and DF
3:01:03do Weare are basically synonymous in
3:01:05spark this time we've passed in call
3:01:08which is one of the spark SQL functions
3:01:11it's just another way of writing and
3:01:13obtaining that column data and again
3:01:15we're calling is null on that column
3:01:17data and we're saving it into this nulls
3:01:192 variable and we're displaying that and
3:01:21so we can see that these two are exactly
3:01:23the same so both methods have returned
3:01:25the same null values which is good so
3:01:28now let's look at dropping some of those
3:01:30null values using the drop na method so
3:01:33this method has a few different
3:01:35parameters these are two of them how
3:01:37threshold and also subset as we're using
3:01:40in this example here so if you pass in a
3:01:43value in the how parameter it can be
3:01:45either any or all so if it's any it's
3:01:47going to drop the row if any of the
3:01:49values in any column is null if you pass
3:01:52in a how value of all it's going to drop
3:01:55the row only if all of the values are
3:01:58null now threshold is going to give you
3:02:00a threshold for the number of columns
3:02:02that need to be null for that row to be
3:02:05dropped and obviously if you use this
3:02:06value it's going to overwrite that how
3:02:08parameter now the other one that we're
3:02:09going to be using is subset so you can
3:02:11pass in subset and it's going to limit
3:02:13The Columns that it looks for for your
3:02:17drop so example we only want to drop it
3:02:19if there's a null value in this Revenue
3:02:23column and you can pass this in either
3:02:24as a list or as a string as well in our
3:02:27example we've only got one column in
3:02:29that list so we're going to save the
3:02:30result as no Nas and then we're going to
3:02:33do the same thing we're going to print
3:02:35out the result here just to check that
3:02:37we have actually removed some rows so
3:02:39we've removed those six rows the null
3:02:41values from a data set so that looks
3:02:43like it's worked correctly so next let's
3:02:45look at type conversion and also adding
3:02:48new columns into a spark data frame so
3:02:51we can call DF do print schema and it
3:02:54gives you a bit of a look at what our
3:02:56schema is initially for this data set
3:02:59including the data types for each column
3:03:02what we're going to do is get the column
3:03:04using DF do unit sold so that's one of
3:03:07our columns here currently it's an
3:03:08integer and for the purposes of this
3:03:10demo we want to use thecast method and
3:03:13we're going to give it the data type of
3:03:15string then what we've done is we've
3:03:16printed the schema again just to check
3:03:18what that schema looks like and we can
3:03:20see here this unit sold converted so
3:03:23we've passed in this units sold
3:03:25converted which is the new column name
3:03:27that we get with this with column
3:03:29function and we can see that from the
3:03:31print schema it's showing a data type of
3:03:34string so we've successfully casted that
3:03:36value from an integer into a string
3:03:39value so next just take a quick look at
3:03:41filtering and we've already had a look
3:03:42at some filtering previously in this
3:03:45lesson but let's just go over it again
3:03:46so in a filtered data frame here we're
3:03:49getting the original data frame and
3:03:51we're calling DF do where we're passing
3:03:54in the column name and the column here
3:03:56is revenue and we want to filter only
3:03:58for data that is more than 10 million so
3:04:02we can see that this has actually
3:04:03removed it's filtered out
3:04:071,19 rows from this data set as I
3:04:10mentioned previously we can use DF wear
3:04:12or DF filter if you want to do filtering
3:04:14if I change this to DF do filter should
3:04:17should work exactly the same there you
3:04:19go so now let's look at data enrichment
3:04:21so adding new columns so we've already
3:04:24seen this one as well we're going to be
3:04:26using DF dowi column and you can also
3:04:29use with columns if you want to do more
3:04:31than one of these at a time but we're
3:04:33just going to be showing you one column
3:04:34here what we're doing we're taking the
3:04:35original data frame we're calling with
3:04:37column to add a new column we're giving
3:04:39the column a name half revenue and we're
3:04:42giving the function to apply to that new
3:04:45column and in our example we're just
3:04:47going to be in the revenue so DF Revenue
3:04:49divid two and then we're displaying the
3:04:51results in our enriched data frame here
3:04:54so if we just call that that's going to
3:04:55look like this so we've got our new
3:04:57column which is half Revenue here which
3:04:59is the revenue divid by two finally
3:05:01we're just going to look at joining and
3:05:02merging data frames in spark so for this
3:05:05one we're going to be using some
3:05:07different tables I've just pulled in the
3:05:09Dealer's data frame so the de dealers
3:05:12table which is this one here and also
3:05:14the countries data frame and you'll
3:05:16notice that these do actually have a
3:05:17joining key so they both have country ID
3:05:20so each dealer has a country ID what we
3:05:23want to be doing is pulling through the
3:05:25country name maybe we've got a
3:05:28normalized data model and we want to be
3:05:30denormalizing it and we're going to be
3:05:31looking at what that means in a bit more
3:05:33detail a bit later in this tutorial but
3:05:35for now let's just look at joining and
3:05:37merging these two data sets together so
3:05:39the basic Syntax for a join in spark is
3:05:42well we're going to get one data set
3:05:44here which is the dealer data set we're
3:05:47going to call dealers DF do jooin that's
3:05:50going to be the first part of our join
3:05:53we're going to pass in the second data
3:05:55frame that we want to join it to in our
3:05:57case countries DF then we're going to
3:05:59specify on what column we're going to be
3:06:01joining on so we're going to be joining
3:06:03on dealers DF country ID is equal to
3:06:07countries DF do countryid in this next
3:06:09row we're just going to specify some
3:06:11select statements so what we've done is
3:06:13we've used this brackets here because
3:06:15we're going to be chaining more than one
3:06:17command
3:06:18together and in this second row we're
3:06:19just selecting a few different columns
3:06:22to be returned in this joined data frame
3:06:25we just want to specify that we want
3:06:26returned the dealer ID the country ID
3:06:28and the country name this is all we care
3:06:30about in that resultant data frame so
3:06:33dealers DF is not defined that's cuz we
3:06:34haven't defined it let's just read that
3:06:37data into a spark data frame first and
3:06:39then we can run this one here okay so
3:06:40now we've got dealer ID country ID and
3:06:43country name in the same data frame now
3:06:45one thing you notice here is that the
3:06:47country name has actually pulled through
3:06:49a few interesting characters so that's
3:06:52something you might want to change later
3:06:53on in your data processing workflow
3:06:56might be that I've saved the CSV file in
3:06:59an incorrect format maybe that's
3:07:00something we want to clean as well in
3:07:03our data cleansing process but for now
3:07:05we're just going to leave it like this
3:07:06so finally we're going to be taking a
3:07:08look at data modeling and typically this
3:07:10takes place within that gold layer
3:07:13because these are going to be the
3:07:14analytical models and the data models
3:07:16that we're going to be bringing into to
3:07:17our semantic layer so let's start off
3:07:19with the star schema so this is an
3:07:21example of a star schema what you can
3:07:24see is in the middle we've got a fact
3:07:26table and this fact table represents
3:07:29sales so it's revenue for a particular
3:07:31company here this example is using car
3:07:34sales so the amount of cars sold from
3:07:36different dealerships so we've also got
3:07:39a dealership Dimension table a dim date
3:07:42dim model so the type of car that's sold
3:07:44and the branch that that was actually
3:07:46sold at so called a star schema because
3:07:49we have our fact table in the middle and
3:07:50then multiple Dimensions all linking to
3:07:53that fact table via some sort of primary
3:07:55key now the reason we prefer star
3:07:57schemas is if you want to be doing bi
3:08:00powerbi basically because the powerbi
3:08:02engine is most efficient in this star
3:08:04schema so a pretty common piece of data
3:08:06modeling that you might have to do as an
3:08:07analytics engineer is prepare this data
3:08:11model in your gold layer of a data
3:08:13warehouse for example so that your
3:08:15powerbi developers can just pick up this
3:08:17data model and you know write their
3:08:19measures on top of it now normally when
3:08:21we're modeling in this star schema data
3:08:23model the fact table is going to be aend
3:08:26only so we're not really going to be
3:08:28updating many values in that fact table
3:08:31it's just going to be new sales added
3:08:32onto the end of that fact table normally
3:08:35it's going to be a very long list your
3:08:37fact table and you're going to create
3:08:38connections from that fact table into
3:08:41the other dimensions so if you want to
3:08:43know more information about that
3:08:45particular sale in the fact table then
3:08:47you can going to find those in the
3:08:48dimensions and the dimensions tables can
3:08:51change over time right take for example
3:08:54dim branch in that Dimension table you'd
3:08:57expect to see details about the
3:08:59different branches that exist for this
3:09:01car manufacturer and some of those
3:09:02things can change over time some of the
3:09:04details about a particular Branch might
3:09:07change right they might change address
3:09:09they might change name certain details
3:09:11about each Dimension might change over
3:09:14time and one way of dealing with that in
3:09:16data model well we need to introduce
3:09:18this concept of slowly changing
3:09:21Dimensions because with our fact table
3:09:22as we mentioned we're not really
3:09:23updating anything over time we're just
3:09:26appending new rows with our Dimensions
3:09:29we need to be able to handle different
3:09:31changes to that Dimension table over
3:09:33time and it's what we call slowly
3:09:35changing Dimensions because these are
3:09:37not going to be big updates might happen
3:09:39once a month or every week a lot of
3:09:41slower frequency than the data is going
3:09:43to be added into that fact table and in
3:09:45data modeling there's a lot of different
3:09:47ways
3:09:47that we can model these slowly changing
3:09:50Dimensions here's an example here take a
3:09:53different example we're looking at
3:09:54employee table we've got employee ID on
3:09:57the left hand side employee name and the
3:09:59department now it's not uncommon for an
3:10:01employee to change departments so on the
3:10:04right hand side you can see that this
3:10:05Dimension has actually changed so Danny
3:10:08Walker who employee ID number two is
3:10:11moved from the marketing department into
3:10:14the sales department and we can model
3:10:15this in a number of different ways ways
3:10:18now the first way in which you need to
3:10:19really be aware of for the exam is the
3:10:22type one SCD or slowly changing
3:10:24Dimension so in the type one SCD we're
3:10:27going to be overwriting any new data so
3:10:30we're going to get the employee data and
3:10:32every time we query that data set that
3:10:35Source data again we're just going to
3:10:37overwrite whatever's in that table we're
3:10:39not going to store any sort of History
3:10:41so we're going to be implementing it
3:10:42with the overwrite writing mode and you
3:10:45can do that either in the data pip line
3:10:47the data flow or in py spark as well now
3:10:50if you're using a Lakehouse as your data
3:10:53store it's important to bear in mind
3:10:54that the history can still actually be
3:10:57retrieved at a point in time using the
3:10:59Delta log so in a lake house
3:11:01architecture in the fabric Lakehouse
3:11:03because you can access those Delta logs
3:11:05just because you overwrite it doesn't
3:11:07necessarily mean that that data isn't
3:11:08stored so the second way that we can
3:11:10deal with this sort of change in a
3:11:13dimension is the type two slowly
3:11:15changing Dimension so in this example
3:11:18we're going to need a few extra columns
3:11:19you can notice that at the top there
3:11:21we've added valid from and valid to and
3:11:24optionally an is current as well which
3:11:26is also quite useful for bi purposes as
3:11:29well so you notice that when we first
3:11:30write a record into this table we're
3:11:33going to populate the valid from that's
3:11:35when this row is valid in that data set
3:11:38and when you first write it the valid to
3:11:40is going to be sometime long into the
3:11:42future normally it's the year
3:11:4499,999 and we also set the is current
3:11:46flag to one or true now when we get an
3:11:50update to that data set so we get a
3:11:52brand new set of data well in the type
3:11:54two SCD we add new rows based on any
3:11:58incoming data so you see in this type
3:12:01two SCD we're going to be adding a third
3:12:03row and it's going to be Daniel Walker
3:12:04but this time the department is sales we
3:12:07have to do a few things here with the
3:12:08valid 2 and the valid from dates so when
3:12:11you write that third row into the data
3:12:13set you need to update row two to
3:12:16populate that valid 2 column and set it
3:12:18is current to zero or false the new row
3:12:21when we write that third row into the
3:12:24data set obviously we're going to have
3:12:26department is sales we're going to add
3:12:27the valid from date as the day that
3:12:29you're making the change the update with
3:12:31the valid two of the year 99,999 and
3:12:34then is current of true now you can see
3:12:37that by doing this we're actually
3:12:39storing a bit of a history about how our
3:12:42Dimension is evolving over time and so
3:12:45you can on the back of this create some
3:12:47quite sophisticated queries about what
3:12:49the state of that Dimension table is at
3:12:52any given point in time now when we're
3:12:54implementing the type two slowly
3:12:56changing Dimension there's a few things
3:12:58that you need to bear in mind so if
3:12:59you're using the Lakehouse and Spark
3:13:02engine then bear in mind that as we
3:13:03mentioned previously the Delta log
3:13:05actually stores a history of all the
3:13:07rights to a particular table so this
3:13:10might be a simple option and depending
3:13:12on how you want to use the history
3:13:14Downstream in your analysis that might
3:13:16be good enough for some use cases now if
3:13:18not then you can use the merge into so
3:13:21we can use that as part of the spark SQL
3:13:23library and we can use that to basically
3:13:25update the underline Delta tables now if
3:13:28you're using the data warehouse and the
3:13:30tcq experience then one method is to
3:13:34load your new data into a staging table
3:13:37and then use a stored procedure to
3:13:40perform the checks around the updates
3:13:42and the inserts and the valid to and
3:13:44valid from calculations now
3:13:46unfortunately the merge operation in
3:13:48tsql the tsql surface area is not
3:13:50currently supported so you have to
3:13:52actually implement this manually if you
3:13:54want to do that in the data warehouse
3:13:56currently now one kind of implementation
3:13:58trick is to use row hashing here so say
3:14:02for example you want to check which rows
3:14:05have changed from One update to the
3:14:07other and one method that you can use to
3:14:10do this in a bit more of an efficient
3:14:11manner is to implement row hashing so
3:14:14you can hash the value of an entire row
3:14:17both in your existing data set and in
3:14:19the new data set that you're checking
3:14:21and then you can compare the two hashes
3:14:23and obviously if the hash values match
3:14:25then you know that that record hasn't
3:14:28changed but if the two hash values are
3:14:30different then you can go ahead with the
3:14:32update logic that you need to update
3:14:35that table next we're going to talk
3:14:36about Bridge tables so imagine a company
3:14:39has many different projects running and
3:14:42the employees assigned to one or many
3:14:45projects what does this look like if you
3:14:47want to do a bit of a data model and
3:14:49maybe build a power VI report off that
3:14:51well in the question here or in the
3:14:53scenario we can see that actually many
3:14:55employees can work for many different
3:14:58projects at the same time so that is a
3:15:00many to many relationship which exists
3:15:02between our dim projects and our dim
3:15:05employees table and this can cause quite
3:15:07a lot of issues when it comes to powerbi
3:15:09and efficiency in very large data models
3:15:12as well so one thing that we can do to
3:15:14resolve this is to implement a bridge
3:15:17table and what that looks like is
3:15:19basically a onetoone mapping of all
3:15:22projects and all participants now this
3:15:25is really useful because it turns our
3:15:27many to many relationship into two on to
3:15:30many relationships let's have a look at
3:15:32how you can Implement that in a tsql
3:15:35data warehouse just to kind of show you
3:15:37what that looks like in real life okay
3:15:39let's explore Bridge tables in a bit
3:15:41more detail and here we're in the data
3:15:44warehouse experience and I've just built
3:15:47a gold data warehouse and we've got
3:15:48these two tables here we got a projects
3:15:50table and an employees table now it's
3:15:53just a simple demo just to show you what
3:15:54a bridge table might look like and how
3:15:56you can implement it in SQL our tables
3:15:59here are the projects we've got the
3:16:00project ID and the project name and
3:16:02we've also got these project participant
3:16:04and if we have a look at the data here
3:16:05for our projects table we've got a bit
3:16:08of a messy string separated values in
3:16:12this project participants column so we
3:16:15can see that each project has multiple
3:16:18project participants and some of them
3:16:20are overlapping right so these are
3:16:21employee numbers that come from this
3:16:23employees table so employee 101 who is
3:16:27John Smith he's actually working on
3:16:29multiple projects right so in our data
3:16:31model here we've actually got a many
3:16:32many relationship we can't really
3:16:34implement it at the moment because of
3:16:35that string concatenation in this
3:16:38project participants column so how do we
3:16:41get around this well and I have actually
3:16:43got the tsql here and I'll leave this in
3:16:46in the school Community as well if you
3:16:48want to have a play around with this
3:16:50yourself I've just created the projects
3:16:52table I've created the employees table
3:16:54and I've inserted some values here so
3:16:56one way that we can use to resolve that
3:16:59many to many relationship is to create a
3:17:01bridge table between those two tables
3:17:05between the projects and the employees
3:17:06table and in this example I'm
3:17:08implementing that as a SQL View and what
3:17:11I'm doing is I'm using the cross apply
3:17:13function here in tsql along with string
3:17:16split
3:17:17so string split is basically going to
3:17:19look in that project participants column
3:17:22within dbo do projects so that's the a
3:17:25comma separated value column which is a
3:17:28bit messy but we can use cross apply and
3:17:31string split together and it's basically
3:17:33going to separate all those values based
3:17:35on this separator here the the comma
3:17:37separation and what this is going to do
3:17:39is it's going to create a one to one
3:17:40mapping for every single project and
3:17:43every single project participant so if
3:17:45we just run this here here let's just
3:17:47have a look at what that looks like here
3:17:48so you can see if we look at the
3:17:50original table just to remind ourselves
3:17:52of what that project's table looks like
3:17:55so initially it looked like this it had
3:17:57two rows and it had project one with two
3:18:00participants and project two with three
3:18:04participants and what we've done in this
3:18:07view if we just recalculate that it's
3:18:09basically separated out all of these
3:18:12comma separated values and it's made one
3:18:14row for each one so now we have five
3:18:17rows because these are all the different
3:18:18combinations of projects and employees
3:18:21what we can do is we can create a view
3:18:23with this logic like so and I've just
3:18:26called it dbo view Bridge Project
3:18:29participants and that's going to make
3:18:30that available in our data model now so
3:18:33let's just have a look at this so now
3:18:35we've got this view bridge table in our
3:18:38data model and what we can do is we can
3:18:40now connect the project ID to the
3:18:42project ID in this instance it's
3:18:44actually going to be one too many
3:18:46because this is a dimension table this
3:18:48is our Bridge table so it's going to be
3:18:49one to many and we can do the same on
3:18:51this side here so this time it is going
3:18:53to be many to one so what we've done is
3:18:54we've transformed a many to many
3:18:56relationship into two one to many
3:18:59relationships and this is going to make
3:19:00it a lot easier when you implement this
3:19:02stuff in powerbi and that's one way that
3:19:05you can resolve a many to many
3:19:06relationships using a bridge table now
3:19:09we've implemented this in a SQL view
3:19:12just because it's quite easy in reality
3:19:14if you wanted to use direct Lake mode
3:19:17obviously you can't use a view with
3:19:19direct late mode so if you want to use
3:19:22direct late mode then you might want to
3:19:24use a stored procedure to materialize
3:19:27out the data in this bridge table so
3:19:29that it doesn't fall back to direct
3:19:31query mode but that's Bridge tables in
3:19:34tsql okay so the next thing we want to
3:19:36talk about here is normalized and
3:19:38denormalized data so if we go back to
3:19:41our data model here where we were
3:19:43looking at car sales so we have some
3:19:45sort of fact Revenue in the middle and
3:19:47some Dimensions that give us more
3:19:49information about that particular sale
3:19:52so what branch it was at what model of
3:19:54car was sold the date that it was sold
3:19:57and also the dealership and what we've
3:19:58done is we've added in another dimension
3:20:00here onto that dim dealers so there a
3:20:03relationship between the dim dealer and
3:20:05the dim cities so it's basically telling
3:20:08us the city of that dealership now this
3:20:11is what's called a snowflake
3:20:12architecture now this can be a very
3:20:15efficient way of storing really large
3:20:17data models because you'll notice in the
3:20:19dim dealers table we're only storing the
3:20:21city ID we're not bringing through any
3:20:24other data into that dim dealers now
3:20:26this can be beneficial in some instances
3:20:29for example as we mentioned for
3:20:30efficiency but when we move this kind of
3:20:32model into the powerbi world it can
3:20:34bring some limitations as well so one
3:20:37way to get around this is with the
3:20:39denormalized data model and this is what
3:20:42this looks like here and it's
3:20:43effectively turning our snowflake model
3:20:46back back into our star schema model
3:20:49which we know can be a lot more
3:20:50efficient when we've got very large data
3:20:52sets now this does introduce some
3:20:55redundancy into that dim dealers
3:20:57Dimension because instead of just
3:20:59storing the city ID for every dealership
3:21:02we're going to be repeating the city ID
3:21:04and the country ID and the region for
3:21:06all of the rows in our dim dealerships
3:21:10but by D normalizing our data model we
3:21:13actually get a number of other benefits
3:21:15so now we have everything in that dim
3:21:17dealer's table which makes it a lot
3:21:19easier to build things like filters on
3:21:21top of that Dimension table we can also
3:21:23have hierarchical filters so we can look
3:21:26at the dealerships by City Country and
3:21:29region and build a bit of a hierarchy
3:21:31there which isn't possible when you've
3:21:33got that split across two Dimension
3:21:35tables finally we're going to take a
3:21:36look at data aggregation and the word
3:21:39aggregation can mean a few different
3:21:41things in data analysis and data
3:21:43modeling and I'm not 100% sure what
3:21:45Microsoft expects for the data
3:21:47aggregation and deaggregation that
3:21:49they've mentioned in the study guide but
3:21:50I'm going to talk through both examples
3:21:52so that you know both of them and if you
3:21:54have any insight as to what Microsoft
3:21:56mean from the study guide when they talk
3:21:57about aggregation then let us know in
3:22:00the comments so the first possible
3:22:01meaning when we talk about data
3:22:02aggregation is when we have different
3:22:04slices of the same data set but they're
3:22:07spread across different files so this
3:22:09one we have one file with the UK data
3:22:12one file with the USA Data and one file
3:22:15with the Canadian data data and in this
3:22:17context data aggregation can mean
3:22:19basically combining these data sets into
3:22:22one long table of transactions in this
3:22:25case showing you Revenue across all your
3:22:27different Source data sets right and we
3:22:29can implement this in a number of
3:22:31different ways really in Fabric in the
3:22:34data flow we can use the append
3:22:36functionality and in the tsql experience
3:22:38and also spark SQL we can use Union and
3:22:41Union all and Union is also a method in
3:22:44pypar as well now the difference here
3:22:46between Union and Union all with a union
3:22:49it's actually going to remove any
3:22:51duplicate rows in the resultant data set
3:22:54whereas Union all is just going to
3:22:56basically append all of the different
3:22:58data sets on top of each other and it's
3:23:00not going to check for duplicates so
3:23:02also mentioned in the study guide is
3:23:04data deaggregation and again there can
3:23:06be many different meanings to the word
3:23:08deaggregation so assuming we mean the
3:23:10first meaning of aggregation then
3:23:12deaggregation is going to be the
3:23:13opposite right it's going to be
3:23:14splitting one large file into multiple
3:23:17different categories of data so another
3:23:19possible definition of data aggregation
3:23:23is when we transform a data set from a
3:23:26more granular data set into a less
3:23:28granular data set so on the left hand
3:23:30side for example we have all of our
3:23:32transactions and the revenue for that
3:23:34transaction and the different country
3:23:37that that transaction was made in now if
3:23:39we were to get aggregate that data
3:23:42perhaps by country then we get some sort
3:23:45of aggregation met tric for each country
3:23:48so here we're showing the total revenue
3:23:50for each country so the UK here has 600
3:23:53which is 100 plus 500 the US has 200 and
3:23:57Canada has 800 now this is typically
3:24:00implemented using the group by statement
3:24:03and the group by functionality and this
3:24:05can be done in the data flow tsql or in
3:24:07spark and when we build this kind of
3:24:09aggregate transformation we also need an
3:24:11aggregation function and it's normally
3:24:14one of count sum Max Min average so it's
3:24:19how are you combining all of the values
3:24:21within that group buy statement so in
3:24:23our example in the top right hand corner
3:24:25we used a sum so we just summed up all
3:24:27of the different Revenue numbers for
3:24:29each country but you could do average
3:24:32revenue you could do the max Revenue it
3:24:34depends on your use case here okay let's
3:24:36just round up everything that we've gone
3:24:38through in this lesson and test some of
3:24:41your knowledge question one when using
3:24:43DF dojin in pisar Notebook the default
3:24:47join type is a full outer join B inner
3:24:51join C left join d right join or E anti
3:24:55join pause the video here take a moment
3:24:57to think about the answer and I'll
3:24:58reveal the answer to you shortly so the
3:25:00default join type in a p spark or in any
3:25:04spark notebook is the inner join not
3:25:06much to say about that one that's just
3:25:08something you need to know all the
3:25:09others are incorrect question two you're
3:25:11looking to migrate a data transformation
3:25:14workload that is currently done using a
3:25:17data flow Gen 2 and convert it into a
3:25:20tsql script now the data flow appends
3:25:23two data sets together and removes any
3:25:26duplicate rows which tsql command can
3:25:29you use to implement this transformation
3:25:31a a left joint B Union or C concat D
3:25:35append or E union so the answer here is
3:25:39the union now we mentioned when we were
3:25:41talking about unions that the union
3:25:43removes duplicates and in our question
3:25:46here obviously the important sentence to
3:25:48pick out was that the source data flow
3:25:50the thing that we're trying to convert
3:25:52into a tsql script well currently that
3:25:54appens two data sets together and
3:25:56removes any duplicate rows so the tsql
3:25:58equivalent of that is the union it's not
3:26:01going to be the union all because that's
3:26:02not going to remove the duplicate rows
3:26:04it's not going to be the left join
3:26:05because that won't do what we want to do
3:26:07now aend is not a tcq function conat is
3:26:12how you could achieve this in pandas but
3:26:14not in tsql so the answer here is Union
3:26:17question three you have a spark data
3:26:19frame called DF your goal is to remove
3:26:22rows that contain a null value in the
3:26:24transaction date column which of the
3:26:26following will help you achieve this a
3:26:29DF do drop duplicates B DF do dropna
3:26:32with how equal to all DF filter
3:26:35transaction date do is null D DF drop na
3:26:39how equals any or E DF do dropna with
3:26:42subset equal to transaction date so the
3:26:45correct answer here is is e we want to
3:26:47be using the drop Na and passing in
3:26:50subset equal to transaction date now the
3:26:53question here is asking us to remove
3:26:55rows that contain null values in a
3:26:58specific column so that specific column
3:27:00is the important part of the question so
3:27:02we want to be using drop na because it's
3:27:04going to remove the rows and by passing
3:27:05in subset equals transaction date that's
3:27:08going to specify only to look in the
3:27:10transaction date column so maybe that's
3:27:12a really important column in our data
3:27:15set we want to be abs Ely sure there's
3:27:17no na values CU if there's any na values
3:27:19then maybe that's going to make our
3:27:20analysis completely redundant or it's
3:27:22going to ruin all of our Downstream
3:27:23analysis so you might want to remove
3:27:25those rows entirely it's not going to be
3:27:27B because in B and D we're not actually
3:27:30specifying that subset now if you to
3:27:32implement D then it would drop the rows
3:27:36where there is a null value in the
3:27:37transaction date column but it would
3:27:39also remove rows with null values in any
3:27:42other column so that might be too much
3:27:44based on your requirements that's
3:27:46probably not what you want to be doing
3:27:47filter that's going to just filter the
3:27:50data set for null values cuz that's
3:27:52going to return all of the rows which
3:27:54are null so we want to be doing the
3:27:55opposite of that basically and a drop C
3:27:58duplicates well that's not what we're
3:27:59trying to achieve here so that's not
3:28:01going to be the answer question four a
3:28:02classical star schema data model
3:28:05consists of the following is it a one
3:28:07Central fact table and multiple
3:28:09Dimension tables B one dimension table
3:28:11and one or more fact tables c one fact
3:28:14table multiple dimens di tables some
3:28:17with Dimension to Dimension
3:28:18relationships or D one big fact table
3:28:21fully denormalized without any Dimension
3:28:23tables so the answer here is one Central
3:28:25fact table with multiple Dimension
3:28:27tables so B is obviously the wrong
3:28:29answer here because it's got one
3:28:31dimension table and many fact tables
3:28:33it's not going to be the star schema C
3:28:36is a snowflake one fact table with
3:28:38multiple Dimension tables some with
3:28:40Dimension Dimension relationships so
3:28:42that's not going to be what we're
3:28:43looking for in a star schema and one big
3:28:46fact table fully denormalized without
3:28:48any Dimension tables again that's not
3:28:50really a classical star schema so the
3:28:52answer here is a you inherit a data
3:28:54project and you're inspecting the tables
3:28:57in the data warehouse one of the tables
3:28:59is a dimension table dim contacts with
3:29:02the following columns contact ID contact
3:29:04name contact address effective date and
3:29:06effective until make an assumption about
3:29:08the type of data modeling that being
3:29:10implemented in this Dimension table is
3:29:12it a type zero SCD slowly change di
3:29:15mention a type one SCD a type 2 SCD or a
3:29:18Type 3 SCD so here the answer is Type 2
3:29:22SCD so we can see from inspecting the
3:29:25columns here that they've got two date
3:29:27fields or we can at least assume their
3:29:29dates so effective date is probably when
3:29:33that row when that contact data point
3:29:36was entered into the system and
3:29:38effective until is basically the same as
3:29:40that valid two now we can't actually see
3:29:42the data here but we can assume that
3:29:44that's what those two columns are doing
3:29:46a type zero SCD is a fixed Dimension so
3:29:49something that never changes type one is
3:29:51obviously you're not going to be
3:29:52tracking that effective date and
3:29:54effective until those two dates it's
3:29:57just going to be overwritten in a type
3:29:58one and a type three is another slowly
3:30:00changing Dimension type that is not
3:30:02actually asked about in the exam but
3:30:04it's basically going to store previous
3:30:06values for each of your columns or at
3:30:08least the columns that you're interested
3:30:10in storing so say for example you'd have
3:30:12contact name there you might also have
3:30:14previous contact name and then when your
3:30:16data gets updated you're going to update
3:30:18the previous contact name and the new
3:30:19contact name as well so that's a Type 3
3:30:22SCD don't think that's in the exam but
3:30:24just something to bear in mind so the
3:30:25answer here is c a type two slowly
3:30:28changing Dimension congratulations
3:30:29you've now completed the third part of
3:30:32section two preparing and serving data
3:30:35in the next lesson we're going to be
3:30:36looking at performance monitoring and
3:30:39optimization of all of our data
3:30:41processing workloads in fabric so make
3:30:44sure you click here to join us in the
3:30:46next lesson I'll see you there hey
Optimizing performance
3:30:48everyone welcome back to the channel
3:30:50today we're continuing our dp600 exam
3:30:53preparation course and we're up to video
3:30:55eight we're making a very good progress
3:30:58here on the course plan today we're
3:31:00going to be looking at optimizing
3:31:02performance and specifically we're going
3:31:04to be covering these bullet points in
3:31:06the dp600 study guide so we're going to
3:31:09be looking at mainly identifying and
3:31:11resolving performance issues right so
3:31:13when you're loading data or when you're
3:31:15querying data or transforming data and
3:31:17specifically we're going to be looking
3:31:18within the data flow The Notebook so The
3:31:21Spark engine and also SQL queries as
3:31:24well then we're going to look at Delta
3:31:25tables in a bit more detail how we can
3:31:27identify and resolve issues within our
3:31:29Delta tables as you know fabric is built
3:31:31on top of the Delta file format so
3:31:33that's a really important topic to
3:31:35understand and as part of that we're
3:31:36going to be looking at file partitioning
3:31:39as well so what that is what that looks
3:31:41like why you might want to implement
3:31:43file partitioning in your Lake housee so
3:31:46as ever at the end of the video we'll be
3:31:48testing some of your knowledge from the
3:31:50topics that we cover in this lesson and
3:31:53as ever I've got some quite detailed
3:31:55notes that you can use to enhance your
3:31:58vision available in our school Community
3:32:00I'll leave a link to that in the
3:32:02description box below so most of this
3:32:03video I'm going to be diving into Fabric
3:32:05and going through performance
3:32:07optimization in a number of different
3:32:09places in fabric but I just wanted to
3:32:11start by Framing what we mean really by
3:32:14performance optimization in fabric now
3:32:16as you know fabric is a very diverse
3:32:18tool so when we talk about performance
3:32:21really we need to get a bit more
3:32:22specific about well what are we talking
3:32:24about it could be data flow performance
3:32:27could be a SQL script in your data
3:32:29warehouse or we could talk about the
3:32:30Delta files and optimizing how they get
3:32:33written and read in our one Lake in our
3:32:36lake houses as well and I would make the
3:32:38distinction here between identifying
3:32:40performance issues and then resolving
3:32:42them it's kind of like a two-step
3:32:43process first we need to know how to
3:32:45identify performance issues in each of
3:32:48these tools and then we need to think
3:32:50about how we can possibly resolve these
3:32:51issues after we identify them and you'll
3:32:53notice for each of the fabric items how
3:32:56we identify performance issues is going
3:32:58to be a little bit different right so
3:33:00for the data flow we're going to be
3:33:01looking at the refresh history and the
3:33:03monitoring Hub and the capacity metrics
3:33:04app and you notice that some of these
3:33:06actually repeat so the monitoring Hub
3:33:08and the capacity metrics app is kind of
3:33:09like a generic place where you can do
3:33:11lots of performance monitoring across
3:33:13Fabric in the data warehouse we have
3:33:15query insights and DMVs Dynamic
3:33:18management views with the Spark engine
3:33:20we have the spark history server and we
3:33:22also have access to quite detailed
3:33:24monitoring in the monitoring Hub as well
3:33:26and when we're identifying performance
3:33:27issues in Delta files there's a number
3:33:29of places you can do that one of them
3:33:31that we're going to be looking at is
3:33:32describe then when it comes to resolving
3:33:34some of these issues well that's where
3:33:36it gets a bit more difficult to Define
3:33:38right because it's normally going to
3:33:39involve some element of refactoring so
3:33:42using different operations in your data
3:33:44flow for example or refactoring your SQL
3:33:47code or refactoring your spark jobs as
3:33:50well now in the data flow we have some
3:33:52specific performance optimization
3:33:54features that is worth going through so
3:33:56we'll be talking a little bit about
3:33:57staging and fast copy as well then when
3:34:00it comes to Delta file optimization
3:34:02we're going to go into a bit more detail
3:34:03about V order optimization file
3:34:06partitioning and also the vacuum and
3:34:08optimize which are two Delta table
3:34:11functions that we can Implement to
3:34:13improve performance with Delta files so
3:34:15that's a bit of an overview of what
3:34:16we're going to be discussing in this
3:34:18lesson now let's dive into Fabric and
3:34:21we're going to begin by looking at some
3:34:22of the generic tools like the monitoring
3:34:24Hub and the capacity metrics app before
3:34:26diving into the data flow data warehouse
3:34:29and the spark notebook in more detail so
3:34:31let's begin okay so just before we jump
3:34:33into fabric for this tutorial I'm just
3:34:36going to start in the school Community
3:34:37here and talk through some of the notes
3:34:39that we have for optimizing performance
3:34:41so here is video 8 optimizing
3:34:44performance we've got this framing
3:34:45performance optimization chart that we
3:34:47spoke about previously but to get us
3:34:48started I want to speak generally about
3:34:50performance monitoring there's a few
3:34:52tools that are quite General they apply
3:34:55across different workloads that are
3:34:57useful to know for the exam and also in
3:34:59fabric generally so we're going to be
3:35:00talking about the monitoring Hub and the
3:35:02capacity metrics app and these are two
3:35:04tools that can be used to monitor
3:35:06performance for a wide variety of
3:35:08operations in fabric so the monitoring
3:35:10Hub is the first one so let's start by
3:35:12looking at the monitoring Hub and if I
3:35:14just flick over to pobi here obviously
3:35:17to access the monitoring Hub you've
3:35:19probably seen it here it's in the left
3:35:20hand toolbar here you got this big
3:35:22button for monitoring Hub and within
3:35:24here we can basically have a look at the
3:35:26runs of a lot of different item types
3:35:29right so you can see semantic model
3:35:30refreshes notebooks so this is going to
3:35:32be a spark session and for each item
3:35:34type we have different logging that gets
3:35:37exposed in this monitoring Hub now the
3:35:39notebook we're going to have a look at
3:35:40in a bit more detail because that's
3:35:41exposing spark log information you can
3:35:44also see like data flow Gen 2 we can
3:35:46have a look at whether Those runs have
3:35:48succeeded or not table loading
3:35:51information in a lake house for example
3:35:53and we can obviously click through into
3:35:55specific items that we care about and we
3:35:57get more information right so this is
3:35:59showing a table load into the bronze
3:36:02Lakehouse and it's showing you the
3:36:03different jobs because this is a lake
3:36:05house is these spark jobs right so it's
3:36:07also going to tell us in the monitoring
3:36:09Hub whether runs have been successful or
3:36:11failed so this particular data flow run
3:36:14we can see that it actually failed right
3:36:16so on the 1st of May at 12:46 I tried to
3:36:21refresh this data flow and actually
3:36:23failed so you can click on view detail
3:36:24and you can get a few more details about
3:36:27what happened here you can't really
3:36:28diagnose what went wrong with that
3:36:30particular data flow to actually get the
3:36:32details of the data flow you have to
3:36:33actually go into the data flow itself
3:36:36which is this one here and then we can
3:36:38click on these three dots here click on
3:36:39the refresh history so for data flows if
3:36:41you want to actually debug what went
3:36:43wrong this is obviously pretty poor data
3:36:45this but you can click on the individual
3:36:47runs and get more detailed information
3:36:49about what's going wrong here so here we
3:36:51can see it's actually this activity here
3:36:54that's failed we can click on that and
3:36:55then get more information about why
3:36:57specifically that column or that data
3:37:00set can't be refreshed so the next
3:37:02general tool that we can use to monitor
3:37:05performance and resource consumption
3:37:08within our fabric capaces is the
3:37:10capacity metrics app and the capacity
3:37:13metrics app and I'll leave a link you
3:37:15can you can obviously get the install
3:37:16instructions here if you have never
3:37:18installed this before this is a powerbi
3:37:21app that you install within your fabric
3:37:23environment you give it your capacity ID
3:37:26and again the capacity settings is where
3:37:28you'll find that capacity ID and then
3:37:30it's going to bring you through to this
3:37:31kind of capacity metrics app now the
3:37:34capacity metrics app is split into two
3:37:36sections we have compute and storage and
3:37:40this obviously lines up with how fabric
3:37:42is build right you're build partly on
3:37:44storage so the amount of storage you
3:37:46have in fabric plus the resources that
3:37:49you consume during compute so if you've
3:37:51already installed the capacity metric
3:37:53app you can find it in the powerbi
3:37:55experience go to apps and then you
3:37:58should see it there the Microsoft fabric
3:38:00capacity metrics app now as a mentioned
3:38:01there's two tabs to this report it's
3:38:03compute and storage on the compute page
3:38:06so starting at the top left we can see
3:38:08the capacity unit spend for particular
3:38:12item types in Fabric and we can also
3:38:14break it down by duration
3:38:16different operations and by user as well
3:38:18we also get this time series of capacity
3:38:21usage over time as a percentage of the
3:38:24total capacity that you have available
3:38:26based on your skew So currently I'm on
3:38:29this trial capacity so this is going to
3:38:30be an X f64 and as you can see I'm not
3:38:33really using barely any of this capacity
3:38:35we've also got these other interesting
3:38:37graphs around throttling so if you're
3:38:39using more than 100% of your capacity
3:38:42usage on that particular capacity this
3:38:45is going to show you where you're
3:38:46throttling and it's also going to show
3:38:48you rejections so if you've got
3:38:50workloads that are being rejected On
3:38:52Your Capacity because you're again over
3:38:55100% And it's currently rejecting
3:38:57workloads then that's going to be
3:38:59exposed here we've also got a graph on
3:39:01overages so again if your capacity is
3:39:04throttled an overage is basically you
3:39:07repaying that capacity usage from your
3:39:09future spend right and for each of these
3:39:11obviously I haven't actually been
3:39:13throttled or you know there's no overage
3:39:16on my actual capacity but there is this
3:39:18explore button that you can drill
3:39:19through to specific events that you want
3:39:21to explore in more detail if that
3:39:23something that's happening on your
3:39:24capacity down below you've got a table
3:39:27of all of the different items in your
3:39:29fabric capacity and the capacity unit
3:39:31seconds that are being used by that
3:39:34specific resource so this is a synapse
3:39:36notebook and we can see that that is the
3:39:38most resource intensive it's used up the
3:39:40most of our capacity unit seconds and
3:39:42it's also got this tool tip where you
3:39:43can look at specific activities and runs
3:39:46of that notebook to dig into a bit more
3:39:48detail there on the storage tab
3:39:50obviously this focuses on the amount of
3:39:52gigabytes of storage in this fabric
3:39:55capacity we can see how it's changing
3:39:57over time we can look at the specific
3:39:59storage by date and we can also look at
3:40:01the top 10 workspaces by bable storage
3:40:04once that's loaded that's what that
3:40:06brings you there here we go okay so next
3:40:08I just wanted to talk about data flows
3:40:10and just to summarize what we looked at
3:40:11previously well if you want to monitor
3:40:14the performance of a data flow well at a
3:40:16high level we can do that within the
3:40:17monitoring Hub but if you want a bit of
3:40:19a lower level data and to understand
3:40:21what's happening within a particular
3:40:22data flow you're going to be wanting to
3:40:24look at the refresh history as I showed
3:40:27you previously for a particular data
3:40:29flow here you can inspect the error
3:40:30messages you get breakdown of the the
3:40:32different sub activities in that load
3:40:35for data flow so that's going to be
3:40:36really important for you to diagnose
3:40:38what's going wrong in a particular data
3:40:40flow if it's not refreshing correctly
3:40:41that's when you where you're going to go
3:40:43to have a look there now there's a
3:40:44couple of features that you need to be
3:40:46aware about in terms of optimizing the
3:40:49performance specifically when we're
3:40:51talking about data flows right so the
3:40:53main one is staging now staging is
3:40:55probably best described using this
3:40:57diagram here so this diagram actually
3:40:59comes from this link here it's the
3:41:01spotlight blog on data flows and it's
3:41:03got some top tips for improving the
3:41:06performance in your data flows so I
3:41:08think to understand what's going on with
3:41:09staging this diagram gives you a pretty
3:41:12good idea so let's start by talking
3:41:14about when staging is disabled so this
3:41:16bottom diagram here right so if staging
3:41:19is disabled all of your transformation
3:41:22in a data flow is going to be done by
3:41:24the data flow engine it's otherwise
3:41:26known as the mashup engine right and if
3:41:28you got a really big data set or you're
3:41:30doing lots of transformation that might
3:41:32not be the most efficient way of doing
3:41:35it so a feature we have available to us
3:41:37to try and improve the speed of doing
3:41:39all these Transformations within a data
3:41:40flow is staging so at the top here when
3:41:43staging is enabled what is going to
3:41:45going to do is it's going to read in the
3:41:46data from the data source and then it's
3:41:48going to immediately write that data
3:41:50into a Lakehouse staging table then it's
3:41:52going to use that Lake housee to perform
3:41:55the transformation right so leveraging
3:41:58the Spark engine rather than the mashup
3:42:00engine to do your transformation then
3:42:03it's going to read the data back into
3:42:04the mashup engine and write it into the
3:42:06destination wherever that might be so a
3:42:08few things to bear in mind here if
3:42:10you've got a lot of data Transformations
3:42:12or you've got very large amounts of data
3:42:14staging is probably going to be a lot
3:42:16more efficient now if you got small data
3:42:18sets it's probably going to be less
3:42:19efficient right because you're going to
3:42:21have to write the data into a lake house
3:42:23transform it into the lake house then
3:42:25write it then the mashup engine picks it
3:42:27up again and writes it to your output
3:42:29destination so there's lots of kind of
3:42:31reading and writing here so on small
3:42:33data sets you probably want to disable
3:42:35staging or not enable staging it's only
3:42:38really when the data set becomes large
3:42:40or you're doing lots of transformation
3:42:41on that data set then we want to enable
3:42:44staging that's what I've of summarized
3:42:45with this sentence here there's a bit of
3:42:47an overhead when you're performing
3:42:49staging and so it doesn't work in all
3:42:51cases only really when your data set is
3:42:53large or you're doing lots of
3:42:54Transformations or both now another
3:42:56feature that they recently announced and
3:42:58therefore might not actually be in the
3:42:59exam yet but it's good to kind of
3:43:01understand know that it exists is fast
3:43:04copy and the way that I think fast copy
3:43:06works is that under the hood it uses the
3:43:08same technology is the data pipeline
3:43:10copy data activity rather than the data
3:43:13flow technology basically as I'm I
3:43:15mentioned it's still a preview feature
3:43:16and it's relatively recent so it might
3:43:17not actually be in the exam yet but you
3:43:19know if you're using data flows in the
3:43:21real world and you're struggling with
3:43:23performance it's worthwhile enabling
3:43:25fast copy just to give it a go see how
3:43:27it impacts the performance in your data
3:43:29flows okay so next up we're going to
3:43:31move on from data flows and now we're
3:43:33going to focus on SQL so we have the SQL
3:43:37engine within the data warehouse and
3:43:39also the SQL endpoint of the lake house
3:43:41as well and we're going to have a look
3:43:42at how we can diagnose and then optimize
3:43:46the performance of SQL scripts and as I
3:43:49mentioned before the capacity metrics
3:43:51app can give you a good kind of high
3:43:52level overview of the resource
3:43:54consumption of specific operations that
3:43:57you're doing within your data warehouse
3:43:59but we have a lot more functionality
3:44:01within the data warehouse to actually
3:44:03explore and identify things like long
3:44:05running queries frequently used queries
3:44:08next we're going to talk about Dynamic
3:44:10management views or DMVs and if you take
3:44:13a look in the data warehouse under the
3:44:15CIS schema there's obviously a lot of
3:44:18different views in there that we can use
3:44:20for database management in general and
3:44:22there's three main ones really for
3:44:25understanding the live SQL query life
3:44:27cycle okay so things that are currently
3:44:29going on in your data warehouse that you
3:44:31need to be aware of and getting some
3:44:33insights about what's happening there
3:44:35and these are exact connections exact
3:44:37sessions and exact requests and these
3:44:39are related in this way here so we've
3:44:41got a bit of data model here so whenever
3:44:43you start a query execution in the data
3:44:45warehouse it's going to start up a
3:44:47session the session is going to have a
3:44:49one toone relationship normally with the
3:44:52connections right so it's going to
3:44:53create a connection between your data
3:44:55warehouse and the underlying seal engine
3:44:57so that's what a connection is going to
3:44:59show you and then you're going to have
3:45:00many requests normally for each of these
3:45:03connections right so using these three
3:45:06commands these three dmbs we can begin
3:45:09to build a bit of a picture about who
3:45:11and how your data warehouse is being
3:45:13queried and these DMV are going to help
3:45:15you answer questions like who is the
3:45:18user running the current session when
3:45:20was the session started by the user
3:45:22what's the IDE of the connection to the
3:45:24data warehouse that is running a
3:45:25particular request how many queries are
3:45:28actually currently active and which
3:45:30queries are long running so you begin to
3:45:32build a bit of a picture about how your
3:45:35data warehouse is being queried who's
3:45:37querying it what they're doing and you
3:45:39know the performance of those queries
3:45:41now the DMV is quite a lowlevel view
3:45:44right we can do lot of information here
3:45:46can merge these tables in different ways
3:45:48to get more and more information so
3:45:50alongside the DMVs Microsoft also expose
3:45:53query insights so query insights is in a
3:45:56different schema so if we have a look
3:45:58here at this particular example of data
3:46:01warehouse we've got the Cy which is our
3:46:02DMVs what contains our DMVs as well as
3:46:05other database management system
3:46:07generated views we've also got query
3:46:09insights so in here we've got these four
3:46:12views that give us basically more
3:46:15userfriendly abstractions over the DMVs
3:46:18right so it's going to expose things
3:46:19like frequently run queries long running
3:46:22queries and you don't have to actually
3:46:24perform those joins of the underlying
3:46:26DMVs to get this information it just
3:46:28exposes them right here and so we're
3:46:30just going to focus on three of these
3:46:32query insights views so exact requests
3:46:35history it's going to return information
3:46:37about each completed SQL request on that
3:46:40particular dat Warehouse frequently run
3:46:42queries is obviously going to give you
3:46:44information about the most frequently
3:46:45run queries and long running queries is
3:46:47basically going to return you
3:46:48information about queries by execution
3:46:51time so this long running queries is
3:46:52going to be really useful to as the name
3:46:54suggests identify queries that are
3:46:56running for a long time and it might be
3:46:57causing performance issues in your data
3:47:00warehouse now one thing to note here is
3:47:02that if you're coming from a SQL Server
3:47:05background when we're talking about
3:47:06performance optimization a really
3:47:08important tool there is the query plan
3:47:10right so currently I don't think it's
3:47:11possible to expose the query plan for
3:47:14particular SQL query but I do think they
3:47:16are planning to support that in the
3:47:18future next up I want to talk about
3:47:20identifying performance issues with the
3:47:22Spark engine and you'll notice that
3:47:23we're back in the monitoring Hub and I
3:47:26just want to look at one of the item
3:47:28details here for this specific notebook
3:47:30so I've been running this notebook it's
3:47:32called Delta optimization and when we
3:47:34click through on this item we get a
3:47:35really detailed analysis of the
3:47:38different jobs that have been run in
3:47:40this notebook we can see that all of
3:47:42these have succeeded you can see which
3:47:44the duration of particular job the data
3:47:47that's been read and written for that
3:47:49particular job so this is a really good
3:47:51place to go if you want to understand
3:47:53what's actually happening when you click
3:47:55run in a spark notebook what's happening
3:47:58under the hood and the success or
3:48:00failure of each of the individual spark
3:48:01jobs now within the monitoring Hub we've
3:48:03also got this link through to the spark
3:48:06history server and as you can see from
3:48:07the UI here we've actually switched from
3:48:09a fabric tool to a generic spark tool
3:48:12right so here you're going to get a lot
3:48:14more detailed information about specific
3:48:16sparkk jobs that you're running you can
3:48:18look at graph so if you do want a bit of
3:48:20a query plan look at the different
3:48:22stages in execution of a particular
3:48:24spark job you can have a look at that
3:48:26here so if you're looking for really
3:48:27fine grain control and Analysis of
3:48:30what's going on on your Spark engine
3:48:32you're going to come to the spark
3:48:33history server now to actually interpret
3:48:36and understand what's going on here
3:48:38would be a whole series in itself I
3:48:40don't think you need to know the the
3:48:41nitty-gritty details of actually what's
3:48:42going on in the spark history server for
3:48:44the dp600 exam just to understand you
3:48:47know what's possible in the spark
3:48:49history server what does it log what can
3:48:51you monitor there I think that's good
3:48:52enough for the dp600 exam so finally in
3:48:55this tutorial I just want to focus on
3:48:57Delta table optimization now as you
3:48:59probably know fabric is built on top of
3:49:01Delta tables so this is a really
3:49:03important topic to understand and the
3:49:05Delta file format is great but it can
3:49:07lead to poor performance and Bloated
3:49:09storage sizes if we're not managing
3:49:12those Delta files correctly now this is
3:49:14a topic that can run very deep like a
3:49:16lot of the topics that I mentioned today
3:49:18we're just going to go through some of
3:49:18the basics of what you need to know for
3:49:20the exam and to do that we're going to
3:49:22be going through a spark notebook so
3:49:25this is the spark notebook that we're
3:49:26going to talk through now I'm going to
3:49:27start by just exploring the problem in a
3:49:29little bit of detail and I think a
3:49:31really good way of understanding what
3:49:32the problem or a problem that can arise
3:49:35with Delta files is this visual here and
3:49:37this visual comes from this blog post
3:49:39here by Sid daba and it's around
3:49:41efficient data partitioning with
3:49:43Microsoft fabric best practices and
3:49:45implementation guide and I've left a
3:49:47link to that in this notebook The
3:49:49Notebook is obviously available in the
3:49:51community here as well and I've also
3:49:53left a link to it here efficient data
3:49:55partitioning here as well so if you have
3:49:58one big file that's 10 gigabytes one
3:50:01paret file then the Spark engine is
3:50:03going to struggle to process that right
3:50:05because as you know spark is a
3:50:07distributed processing engine which
3:50:09means it works best when it splits your
3:50:11file your paret files into smaller
3:50:14chunks then processes these chunks in
3:50:16parallel but to make that possible we
3:50:18need to partition our data and
3:50:20partitioning is basically the process of
3:50:22converting one really big or several
3:50:24really big files into more manageable
3:50:26chunks so that our data can be
3:50:28transformed in parallel basically by The
3:50:30Spark engine so in this notebook we're
3:50:32going to start by looking at file
3:50:33partitioning and then look at some other
3:50:35methods for optimizing Delta tables in
3:50:37doing so this can help improve the read
3:50:40and retrieval performance so by doing so
3:50:43there's no need to scan through millions
3:50:45and millions of rows in your really big
3:50:46paret file if you got a good
3:50:48partitioning system in place it's going
3:50:50to speed up the read performance of your
3:50:52queries and it's also going to improve
3:50:54your transformation performance right as
3:50:56we mentioned before because partitions
3:50:58can be transformed in parallel so file
3:51:00partitioning I'm going to walk through a
3:51:01bit of a demo here and I've used a demo
3:51:04file and it's available on this website
3:51:06here but it's just a parket file what
3:51:07I've done is I've just put it in the
3:51:09files location this one here is called
3:51:11Flights 1M parket and before we get
3:51:14started it's just bear in mind the best
3:51:16practices on Delta Lake partitioning
3:51:19right and I've left a link here this is
3:51:21on the Delta Lake website where they
3:51:23give some general best practices about
3:51:25managing Delta files okay and one of
3:51:27them is around choosing the right
3:51:29partition column now most commonly it's
3:51:32done by date so if you've got time
3:51:34series data you've got some dates in
3:51:36your data set and it's common to use
3:51:37your date for partitioning but there's
3:51:40two kind of rules of thumb to bear in
3:51:41mind when we're talking about
3:51:42partitioning and deciding what col
3:51:44colums you want to partition on so in
3:51:46general we don't want to choose our
3:51:48partitioning column to be something of
3:51:50really high cardinality so say for
3:51:51example you have a column of user ID
3:51:54well that's going to be really high
3:51:55cardinality right every row is basically
3:51:57going to be unique so if you've got a
3:51:59million rows that's going to be a really
3:52:01bad partitioning strategy right because
3:52:03you're going to get a million different
3:52:04partitions and they're all going to be
3:52:05really small and as a general rule of
3:52:07thumb the amount of data that should be
3:52:09in each partition should be around one
3:52:11gigabyte that's kind of like the the
3:52:12good balance of what you should be
3:52:15aiming for with each partition so let's
3:52:17take a look at file partitioning and how
3:52:18to actually implement it in Fabric in a
3:52:21spark notebook so I'm going to begin by
3:52:23just reading in that parquet file our
3:52:25flights 1M parket file I'm just reading
3:52:28it into a data frame I'm just displaying
3:52:30it here and you notice we've got this
3:52:31date column here it's going to be useful
3:52:33for our partitioning strategy we're
3:52:35going to be partitioning on date and
3:52:37it's got some other items here that we
3:52:39don't really care about for this
3:52:40tutorial so before we do some
3:52:41partitioning and we write this file into
3:52:44into our Lakehouse using partitions
3:52:46we're going to do a bit of preparation
3:52:48and specifically we're going to add some
3:52:50columns into our data frame we're going
3:52:52to use DF with columns to add more than
3:52:54one column at a time and we're going to
3:52:56pass in this dictionary object here and
3:52:58each uh part of that dictionary each key
3:53:01is going to be the new column name and
3:53:03the value is going to be how we're
3:53:04actually Computing that value so here
3:53:07I'm just reading in some P spark SQL
3:53:09functions to extract the year the month
3:53:12and the day of month of that date field
3:53:14right that we looked at previously so if
3:53:15we just run that and if we just display
3:53:17the results here okay so now we've got
3:53:18this transformed data frame object right
3:53:21and so you can see here there added in
3:53:23this year month and day column and we
3:53:27can use these in our partitioning
3:53:29strategy now so you can see here from
3:53:31our code cell that we've actually
3:53:33written three different write modes here
3:53:36so the first one is going to write
3:53:37without partitions so this is just the
3:53:39normal saving of the table into a
3:53:42Lakehouse table from our parket file
3:53:44we're going to put it in tables and
3:53:45we're going to call the table flights
3:53:47not partitioned the second one we're
3:53:49going to do is we're going to call
3:53:51Partition by and we're going to
3:53:53Partition by the year and the month and
3:53:55we're going to save this one into
3:53:56another table called Flights partitioned
3:53:58and then as a third example we're going
3:54:00to write with some more partitions so
3:54:02we're going to be partitioning our data
3:54:05into smaller partitions here and again
3:54:06we're calling Partition by but this time
3:54:08we're passing in the year the month and
3:54:10the day so these are going to be more
3:54:11fine grained and I'm going to save that
3:54:13into a table called flights partitioned
3:54:15daily let's just run those okay so now
3:54:18our cell has been executed all of our
3:54:20spark jobs have concluded let's just
3:54:22have a look at this so yeah now you can
3:54:24see in our lake house tables we've got
3:54:26flights not partitioned flights
3:54:28partitioned and flights partitioned
3:54:30daily so we got three different tables
3:54:31here all with different partitioning
3:54:33strategies so let's inspect those and
3:54:36see what's going on okay so I've got
3:54:37three different cells here and you'll
3:54:39notice we're using the SQL so we're
3:54:41using spark SQL here and we're calling
3:54:43describe detail on that particular table
3:54:46it's going to inspect this table it's
3:54:48going to describe what's going on there
3:54:50and some of the results give us a bit of
3:54:51a picture as to how this table has been
3:54:54written into one L so if we inspect the
3:54:57results here and this is our flights
3:54:59partitioned we can see that the
3:55:00partition columns here are year and
3:55:03month which makes sense because this is
3:55:04the year and month one in flights
3:55:06partition and we can see the number of
3:55:07files it's created so the number of
3:55:09paret files we've got here is two next
3:55:11if we compare that to what we get when
3:55:13we look at flight not partitioned we can
3:55:16see that we got zero partition columns
3:55:17and we got one file so everything's just
3:55:19been written into one file in that
3:55:21instance here and for the daily one
3:55:23again if we inspect what's going on here
3:55:25we can see that we've got year month and
3:55:27day partition columns and here we've got
3:55:3059 files so here it's been partitioned
3:55:33we've broken up that data into 59
3:55:36smaller chunks of data now that might be
3:55:38too fine grained because if we scroll
3:55:40back up to the top here one thing that I
3:55:42did mention is that it's a bit of a
3:55:43balancing
3:55:44this because if your files are too big
3:55:47the spark engine's going to have
3:55:48performance issues but there's also the
3:55:50small file problem if your file sizes
3:55:52are under that gigabyte then you know
3:55:55that's also going to cause a lot of
3:55:56issues it's going to have to work harder
3:55:58you're going to have to go through the
3:55:59operation 59 times rather than two so
3:56:02again it's a bit of a balancing act
3:56:04trying to get a good partitioning
3:56:06strategy for your Delta tables but for
3:56:08the purposes of the exam I think it's
3:56:10worthwhile understanding the Syntax for
3:56:13creating part partions like so and then
3:56:16analyzing different partitions using
3:56:18this describe method spark SQL method
3:56:20okay so next I just want you to talk
3:56:21about V order optimization now V
3:56:23ordering is a Microsoft proprietary
3:56:26algorithm and it basically changes the
3:56:28structure of your parket file and what
3:56:31I've put here is is kind of providing a
3:56:32bit of special source so it does some
3:56:34special sorting compaction compression
3:56:37of that parquet files right ultimately
3:56:39to improve the read performance of these
3:56:42parket files across all of the different
3:56:44engines in fabric now whilst the actual
3:56:47algorithm is proprietary it's only used
3:56:49by Microsoft the output of like the
3:56:52parket file is fully kind of Open Source
3:56:55aligns to the traditional parket
3:56:56standards so you can actually read
3:56:58vorded parket files wherever you can
3:57:01read normal parket files that's not a
3:57:03problem at all Now by default V ordering
3:57:06so this algorithm that you used to write
3:57:08parket files it's enabled by default in
3:57:11the fabric spark runtime and you can
3:57:13check that by running this spark comp
3:57:16get so we're looking at the
3:57:17configuration and you can see that here
3:57:19is actually returning true now it can
3:57:21actually be manually disabled if you
3:57:23want it to so we can set the spark
3:57:25configuration by passing in this
3:57:27specific property here spark SQL paret V
3:57:30order enabled to false and then we can
3:57:32reenable it by doing the opposite right
3:57:34putting it to True again so for the exam
3:57:36you might be asked about how do you know
3:57:38whether a particular notebook or a
3:57:41particular spark environment has v order
3:57:43enable well that's this one or how do
3:57:45you enable it or disable it in a spark
3:57:48notebook as well that could be a common
3:57:49question that you might get asked okay
3:57:51just finally I just want to mention a
3:57:52few more Delta table maintenance and
3:57:55optimization techniques so there's a few
3:57:58that come from the actual Delta format
3:58:00itself so we have this function called
3:58:02optimize which is a Delta Lake method
3:58:04that performs bin compaction it it can
3:58:06basically improve the speed of your read
3:58:09queries so if you got multiple small
3:58:11files it's basically going to coals
3:58:14basically mean joining small files into
3:58:16larger files vacuum is another function
3:58:18that we can run and it basically
3:58:20involves removing files that are no
3:58:22longer referenced by a Delta table and
3:58:25then on the spark side there's two they
3:58:27might at least want to be familiar with
3:58:28is coales so as we mentioned when we're
3:58:30talking about optimize optimize is
3:58:32basically the what that's doing under
3:58:34the hood is calling coales and coales is
3:58:36a spark method or basically reducing the
3:58:39amount of partitions in your Delta table
3:58:41so if you've got 100 partitions for a
3:58:43particular file you can coals that Delta
3:58:45table into 10 partitions so coals is a
3:58:48pretty efficient way of grouping
3:58:50partitions into a smaller number of
3:58:53partitions right and I say it's quite
3:58:54efficient because it doesn't require a
3:58:57shuffle of the data it's just grouping
3:58:59partitions together doesn't actually
3:59:01reorganize within particular partitions
3:59:03your data now repartition is similar to
3:59:06coales but it's actually less efficient
3:59:08because it involves breaking up your
3:59:10existing partitions and then creating
3:59:11new partitions and because of this you
3:59:14create either more or less partition
3:59:16it's basically just restructuring how
3:59:18your partitions are created and it does
3:59:20involve some shuffling involves breaking
3:59:23up of your existing partitions and
3:59:24repartitioning them now you might be
3:59:26thinking what's the difference between
3:59:27the spark functions and the the Delta
3:59:28optimize well the Delta optimize has a
3:59:30few kind of things working under the
3:59:32hood that makes it more efficient for
3:59:34number one it's item potent so if you
3:59:36run it repeatedly it's not going to
3:59:38reoptimize files that have already been
3:59:40optimized whereas repartition is going
3:59:42to always repart partition your files
3:59:45you can keep on running this again and
3:59:46again and it's never going to get more
3:59:48efficient it's always going to
3:59:49repartition the files that's one
3:59:51difference between repartition and
3:59:53optimize and you can also run optimize
3:59:55on specific partitions in your data set
3:59:57whereas repartition that's kind of All
4:00:00or Nothing approach you have to
4:00:01repartition your whole table in one go
4:00:04so if you're looking for a bit more fine
4:00:05grained optimization you're going to be
4:00:07want to using the Delta optimize okay so
4:00:10let's just round off the video here by
4:00:12testing some of your knowledge of the
4:00:14things the topics that we've covered in
4:00:15this lesson question one a client you're
4:00:17working with wants to reduce the SKU of
4:00:19their fabric capacity from an F-16 to an
4:00:22f8 to save some money they want to find
4:00:24the most resource intensive workloads
4:00:27and optimize them to use less capacity
4:00:29unit seconds where should they look to
4:00:31find this information is it a the
4:00:32monitoring Hub B capacity metrics app C
4:00:35query insights D spark history server or
4:00:39e the one Lake Hub pause the video here
4:00:42have a little think and I'll reveal the
4:00:43answer to you shortly okay so the answer
4:00:45here is B the capacity metrics app we're
4:00:48talking about resource intensive
4:00:50workloads and our capacity is where
4:00:52we're going to get those resources and
4:00:54specifically it's going to tell you
4:00:56which workloads are the most resource
4:00:58intensive are the workloads that are
4:00:59going to use more of your capacity units
4:01:02seconds right so the answer is going to
4:01:04be your capacity metrics app we can look
4:01:07at all of our spark jobs our data
4:01:09warehouse operations and our data flows
4:01:12and we can come to conclusions about
4:01:14which of these are good candidates for
4:01:16refactoring or optimization now all of
4:01:19the others they might be useful for
4:01:21understanding performance of specific
4:01:23workloads within fabric but the capacity
4:01:25metrics app is the only one here that
4:01:27converts that into capacity units right
4:01:30and that's the important part of the
4:01:31question to understand question two you
4:01:33noticed one of your data flow Gen 2 runs
4:01:35failed to refresh last night where would
4:01:37you go to find out why a particular data
4:01:39flow might have failed a particular Run
4:01:41is it a the capacity metrics app B the
4:01:44monitoring Hub C power query error Hub D
4:01:47data flow refresh history or E the data
4:01:50pipeline run history so the answer here
4:01:52is D the data flow refresh history is
4:01:55where you're going to go to analyze
4:01:57error messages and debug particular runs
4:02:01of a data flow now the power query error
4:02:03Hub that doesn't actually exist I made
4:02:05that up the data pipeline run history
4:02:08but we're not talking about data
4:02:09pipeline here so it's not going to be
4:02:10that the monitor the monitoring Hub will
4:02:12give you some information so it will
4:02:14tell you whether a particular run has
4:02:16failed or succeeded but it doesn't give
4:02:18you more detailed information about
4:02:20error messages and things like that and
4:02:22the capacity metrics app is not going to
4:02:24tell you that answer either so the
4:02:25answer here is D data flow refresh
4:02:28history question three when talking
4:02:30about Delta table optimization which of
4:02:32the following operation removes old
4:02:35files no longer referenced by a Delta
4:02:37table log is it a v order optimization B
4:02:40Zed ordering C vacuum d optim or E bin
4:02:45compaction so the correct answer here is
4:02:47C vacuum so as we mentioned previously
4:02:50the vacuum command does exactly as it's
4:02:53mentioned there basically removes old
4:02:55files that no longer referenced by a
4:02:57Delta table log so the correct answer
4:02:59here is C question four which of the
4:03:02following statements about V order
4:03:04optimization is false a v order
4:03:07optimization is enabled by default in
4:03:10the fabric spark runtime b v order can
4:03:13be enabled during table creation using
4:03:15table properties c a table can be both V
4:03:18ordered and Zed ordered d v order
4:03:20improves the read performance for parket
4:03:22files e v order speeds up the right time
4:03:25of a parket file so the correct answer
4:03:27here is e so the question asked which of
4:03:30the following statements is false and so
4:03:32e is actually false V order does not
4:03:35speed up the right time it actually
4:03:36increases the right time of a parket
4:03:39file the benefit of V ordering comes in
4:03:42the read performance right that's why we
4:03:43do it takes a bit longer to write these
4:03:46files but it massively improves the read
4:03:48performance across any of the engines
4:03:50that you might want to use it in fabric
4:03:52all of the other options are true so it
4:03:54is enabled by default B if it's not
4:03:57already enabled in your spark
4:03:59environment it can be enabled for
4:04:00specific tables during table creation
4:04:03using table properties you can actually
4:04:04optimize both V ordering and Zed
4:04:07ordering for a particular table or PAR
4:04:09paret file within that table at the same
4:04:11time and D is also true because we
4:04:14mentioned that V ordering does improve
4:04:16the read performance for parket files
4:04:19that's why we do it question five you
4:04:21want to analyze long runn queries in a
4:04:23fabric data warehouse what's the minimum
4:04:26workspace role you need to run the
4:04:28following query select start from query
4:04:31insights. longrun inqueries is it a
4:04:34admin B member C contributor or D viewer
4:04:38so the answer here is C contributor to
4:04:40be able to run a query insights query
4:04:43you know autogenerated views that give
4:04:45us information about long running
4:04:47queries in this example you need to have
4:04:49a workspace role of contributor
4:04:51obviously the question asked for the
4:04:52minimum workspace role you can also run
4:04:55these queries with admin or member but
4:04:57the minimum workspace role would be
4:04:59contributor so if you have a viewer role
4:05:01you can't run these queries in a data
4:05:04warehouse congratulations you've
4:05:05completed the biggest section of the
4:05:08exam we're well over halfway now so in
4:05:10the next lesson we're going to be
4:05:12starting the third section section of
4:05:14the exam which is all about building
4:05:16semantic models well done and I'll see
4:05:18you in the next video hello and welcome
Design and build semantic models
4:05:20back to the channel today we're
4:05:22continuing our dp600 exam preparation
4:05:25course and we're up to video 9 designing
4:05:28and building semantic models now as you
4:05:31can see from the course plan we're
4:05:32making really good progress we just got
4:05:34a few more important sections to look at
4:05:36and today we're going to be starting
4:05:37powerbi and semantic modeling part of
4:05:40the exam specifically we're going to be
4:05:41looking at the different storage mod
4:05:43modes import mode direct query direct
4:05:46Lake we're also going to be looking at
4:05:47composite models and what we mean by
4:05:49that including aggregations as well
4:05:52we're going to be looking at the large
4:05:53format data set and then we're going to
4:05:55be digging into a bit of a practical
4:05:57example in palbi desktop defining
4:06:00different Dax measures we're going to be
4:06:02looking at functions iterators table
4:06:04filtering windowing information
4:06:06functions and some of the other more
4:06:08advanced or more recent features as well
4:06:10including calculation groups Dynamic
4:06:12strings field parameters that kind of
4:06:13thing as well as ever we're going to
4:06:15have five sample questions at the end of
4:06:18this video to test your knowledge now
4:06:20bear in mind that I wouldn't classify
4:06:21myself as a powerbi developer I used
4:06:23powerbi quite a lot about six seven
4:06:26years ago recently I've been more
4:06:27focused around data engineering data
4:06:29science so I'll try and explain these
4:06:31Concepts as best possible but I
4:06:32definitely recommend doing your own
4:06:33research I'll leave a link to a lot of
4:06:35really good resources for this kind of
4:06:37thing in the school community so that
4:06:40you can go in a bit more detail there so
4:06:42first up we're going to be looking at
4:06:43storage modes now I've mentioned this
4:06:46fair amount on the channel and lots of
4:06:48people talk about storage modes within
4:06:50powerbi and fabric so we're going to do
4:06:52a bit of a revision what we mean by
4:06:54different storage modes some of the
4:06:55advantages and disadvantages and how you
4:06:58can choose between them so this is the
4:07:00diagram that exists in the documentation
4:07:02I think it does a pretty good job at
4:07:04framing out three different storage
4:07:07modes connection modes you might also
4:07:09hear it called as well now if you've
4:07:11worked in powerbi for a while you're
4:07:13definitely going to be familiar with
4:07:15both the import mode this one in the
4:07:17middle plus the direct query mode
4:07:20potentially you might have used that as
4:07:22well at the top there so for those not
4:07:24coming from the PBI background the the
4:07:26import mode is basically going to take a
4:07:29copy of your data from source and load
4:07:31it into powerbi so it's storing a a copy
4:07:35of all your data in the actual powerbi
4:07:37data model itself now that makes it
4:07:38really fast when you're building ports
4:07:41cuz your data is right there now it does
4:07:42have some limit ations in that because
4:07:45you're copying your data into powerbi
4:07:46there are some limitations around the
4:07:47size right that's one of the main
4:07:49limitations on import mode and also
4:07:51because you're having to do this lift
4:07:53and shift import on a schedule normally
4:07:55your data in your powerbi data model is
4:07:58not going to be always up to date
4:07:59because you're going to have to do it
4:08:00every hour potentially there is the
4:08:02chance that your data will become a
4:08:04little bit stale on the other hand we
4:08:05have direct query so moving up to this
4:08:07top one here in direct query whenever
4:08:10the user views a particular visual
4:08:12particular report page power actually
4:08:15sends a query back to your Source it's
4:08:17going to perform a query of that source
4:08:20and then get back fresh data so in doing
4:08:23so it's near real time so that's one of
4:08:25the benefits of direct query some of the
4:08:27downsides of that mean that we can't
4:08:29actually perform much transformation on
4:08:31that data it has to be transformed in
4:08:32the source right you have to create
4:08:34views and tables that already are
4:08:36transformed in practice direct query can
4:08:38be really slow because you've got to do
4:08:41that query every time you go back back
4:08:43to the original data source and for the
4:08:45user they're sitting and waiting for
4:08:47their visuals to up update every time
4:08:49you click a filter or you change the
4:08:52page it's going to take time it's going
4:08:53to be pretty bad user experience so the
4:08:55final one we've got here is the new one
4:08:57that came with powerbi it's called
4:08:58direct Lake and direct Lake creates a
4:09:01connection between what you create in
4:09:04your P report and the underlying parket
4:09:06files so it's going to read the parket
4:09:09files in your one Lake environment
4:09:12directly so let's just have a look at
4:09:14potential reasons why you might want to
4:09:16choose import mode or direct late mode
4:09:18or direct careering mode so some of the
4:09:20key considerations well as we mentioned
4:09:23for import mode it's going to be really
4:09:25good when your data is small enough to
4:09:28fit within a palbi data model including
4:09:30large format semantic models which we're
4:09:32going to talk about in a short while
4:09:34another good use case for import mode is
4:09:36when you want really good read
4:09:38performance in your dashboards you know
4:09:40like interactivity and a good user
4:09:42experience for your dashboard users when
4:09:44you don't have requirements for near
4:09:46real time updates and if you want to use
4:09:49calculated columns or calculated tables
4:09:52if you want to be doing that kind of
4:09:53stuff then you're going to be want to
4:09:54using import mode and if you want to
4:09:56combine data from multiple different
4:09:58data sources that's another good use
4:10:00case for import mode which me leads me
4:10:02nicely into when you would choose direct
4:10:05Lake mode well the first limitation
4:10:07really is that your data has to be
4:10:08stored in one fabric data store so in a
4:10:12lake house or a data warehouse for
4:10:13example you can't use direct Lake mode
4:10:16to access data across lots of different
4:10:18data stores so that's one limitation the
4:10:21prime use cases when your data set is
4:10:23really really big we're talking tens or
4:10:25hundreds of gigabytes here now obviously
4:10:27that would be too big for most import
4:10:29mode models but that use case is really
4:10:32really good in direct late mode now with
4:10:35direct late mode it will require a
4:10:37little bit of a different skill set
4:10:39within your team because you're going to
4:10:40have to do a lot of the data modeling
4:10:42more up dream in your Lake housee in
4:10:44your data warehouse right because you
4:10:46need to materialize parquet files that
4:10:49can be read by the direct Lake mode
4:10:51connection and typically what that means
4:10:53is your data transformation your data
4:10:55modeling is going to have to be done in
4:10:56your lake house so either using spark or
4:10:59TC call and that might be a slightly
4:11:01different skill set to what you might
4:11:03have in a an import mode powerbi team
4:11:06for example okay so finally let's just
4:11:08talk about when you might want to choose
4:11:10direct query mode for your semantic
4:11:12models well again if you need near real
4:11:15time updates that's going to be a really
4:11:16important one again you're going to be
4:11:18needing to do your trans data
4:11:20Transformations more Upstream so in your
4:11:22data source wherever that might be and
4:11:24direct query is also important part of
4:11:26what we call composite models which
4:11:27we're going to look at in more detail
4:11:29shortly so let's look at composite
4:11:31models then so a composite model
4:11:33combines one or more of these different
4:11:35connection modes that we just discussed
4:11:37previously now commonly this is a direct
4:11:40query fact table and import mode
4:11:43Dimension tables because if we think
4:11:45about the common characteristics of a
4:11:49fact table versus a dimension table well
4:11:52in our fact table we're going to have a
4:11:54lot of rows normally a lot of data could
4:11:56be millions or even billions of rows and
4:11:59it's likely to be updated very often
4:12:02maybe every minute or every second even
4:12:04in some oltp transaction processing type
4:12:07fact tables they could be hundreds and
4:12:09hundreds of records every second now
4:12:11because of those characteristics direct
4:12:13query can be a good match for that type
4:12:16of data set right because you get near
4:12:19real time updates so on data sets that
4:12:21are changing very often and very fast
4:12:24direct query gives you that near real
4:12:25time access to fresh data right now with
4:12:28the dimensions they might be changing a
4:12:30lot slower so it makes sense to use
4:12:33import mode for those Dimensions you
4:12:35know your product table might be updated
4:12:38once per day for example so a direct
4:12:41query connection mode in that example
4:12:43wouldn't really make too much sense
4:12:44because you're going to lead to user
4:12:45experience issues on the front end for
4:12:48that table and you're not going to be
4:12:49getting much benefit because the
4:12:50underline data isn't really changing
4:12:52very often now another benefit of
4:12:54composite models is that they provide a
4:12:56way to model many to many relationships
4:12:59without the need for bridge tables as
4:13:01well so next up we're going to look at
4:13:03aggregations now in this context we're
4:13:05talking about a specific powerbi feature
4:13:07for managing aggregation and
4:13:10specifically what this feature does in
4:13:12powerbi is it takes a really large data
4:13:15set it creates a aggregation either
4:13:17automatically or user generated you can
4:13:20assign the aggregation that you want to
4:13:21build and then it caches the actual
4:13:23aggregation so when you have really
4:13:25large data models it can improve the
4:13:27performance because you're caching the
4:13:29aggregation rather than loading in you
4:13:30know a really long fact table for
4:13:32example and because of that they are
4:13:34often used in conjunction with composite
4:13:37models now either you can create the
4:13:39actual aggregation itself in your data
4:13:41source and then just pull into your
4:13:43powerbi data model or you can bring it
4:13:45into powerbi and then use the power
4:13:47query engine to create an aggregation
4:13:50which then gets loaded into your powerbi
4:13:53engine now as I mentioned there's two
4:13:54really types of aggregation we have
4:13:56userdefined aggregations and that's
4:13:58using the the manage aggregations
4:14:00dialogue in P your desktop to Define
4:14:02these aggregations for a specific
4:14:04aggregation column and then you choose
4:14:06how you want to summarize do you want a
4:14:08Min a Max that kind of thing plus the
4:14:10detail table and detail column
4:14:12properties now if you have access to a
4:14:14premium subscription powerbi then you
4:14:16also get automatic aggregations and
4:14:18these are basically going to use machine
4:14:20learning to try and optimize direct
4:14:22query semantic models and they going to
4:14:24look for the best aggregation to improve
4:14:26performance another thing we need to be
4:14:28aware of for this part of the exam is
4:14:29the large format semantic model now
4:14:32large format semantic models provide a
4:14:34highly compressed inmemory cache for
4:14:38optimized query performance enabling
4:14:40fast user interactivity so if you've got
4:14:42a semantic model model that's perhaps
4:14:44bigger than 10 20 30 GB what you can do
4:14:47is you can convert that small format
4:14:49semantic model into a large format
4:14:51semantic model it's going to apply this
4:14:53compression this in-memory caching
4:14:56that's going to improve the performance
4:14:57of those models now it's not just models
4:15:00that are over 10 GB in size where you
4:15:03might want to think about converting to
4:15:05a large format semantic model in many
4:15:07cases even below that kind of 10 GB
4:15:10threshold there's some benefits in
4:15:12converting to a large format scientic
4:15:14model firstly you're going to get the
4:15:15performance benefits anyway secondly
4:15:17it's commonly used when connecting to
4:15:19thirdparty tools via the xmla endpoint
4:15:22now another feature of large semantic
4:15:24models is ond demand loading now if you
4:15:26watched some of my previous videos you
4:15:28know that direct L connection mode also
4:15:30uses this on demand loading and what
4:15:33that means is that when a user is
4:15:35viewing a particular page in a report
4:15:37they don't need the full data set to be
4:15:39loaded into memory on demand loading has
4:15:42a look at the online paret files and
4:15:44only loads the required data that is
4:15:46needed to visualize the data that's
4:15:48being requested at that specific time
4:15:50now that can really improve performance
4:15:52again but it's a feature that's shared
4:15:53between large format semantic models and
4:15:56the direct L connection mode as well
4:15:57okay so for the next part of this video
4:16:00we're going to switch over to powerbi
4:16:02desktop and we're going to be using this
4:16:04to explain some of the key things that
4:16:07we need to know for this module in the
4:16:09exam so we're going to be looking
4:16:10specifically at variables if iterators
4:16:13table filtering window functions
4:16:15information functions calculation groups
4:16:17Dynamic strings and field parameters and
4:16:20each of these different features and Dax
4:16:22Expressions I've got a little bit of an
4:16:24example just to talk you through the
4:16:26implementation what that looks like so
4:16:28let's start off with variables now
4:16:30variables are very common in pretty much
4:16:32every programming language that exists
4:16:35and ax variables can help us avoid code
4:16:38repetition and also potentially improve
4:16:40the performance of your Dax code too so
4:16:43here what we've done is we've created
4:16:44two variables one called total revenue
4:16:46and one called total days you notice the
4:16:48syntax here is to use V to declare it as
4:16:52a variable and then it stores that
4:16:54result locally and then you can use it
4:16:57later on in your Dax expression
4:16:59typically you'll need to use a return
4:17:00statement as well when we specify these
4:17:03variables and so in this example we're
4:17:04declaring total revenue and total days
4:17:06as variables then we're using that in
4:17:09this return statement to give us the
4:17:11overall average revenue per day next
4:17:13we're going to look at iterator
4:17:15functions and iterator functions in
4:17:18powerbi basically enumerate through all
4:17:21of the rows in a table and they perform
4:17:24some calculation depending on the
4:17:26specific iterator function that you
4:17:28choose and then it's going to aggregate
4:17:30the result now examples include sum X
4:17:34count X average X most of them have this
4:17:37x afterwards and here we've got a bit of
4:17:39an example here to Showcase iterator
4:17:41functions and we using it in this
4:17:44cumulative measure so we've got this
4:17:45cumulative revenue and we're doing sum X
4:17:48okay and we're using it on this filter
4:17:50so what we're basically saying is for
4:17:52each of the rows in our date table we're
4:17:56going to go one by one and for each of
4:17:58them we're going to increase the number
4:18:00of rows that are being filtered right so
4:18:02the first row here we're basically
4:18:04comparing the the current date or the
4:18:06the date in the row that we're
4:18:07interested in with the max date which on
4:18:10the first pass when we're enumerating
4:18:12through this table there's only going to
4:18:14be one row the top row right and then
4:18:16we're calculating the sum of the revenue
4:18:18on that particular date now on the
4:18:21second row obviously this is going to
4:18:23increase the two rows so then we're
4:18:24going to do a cumulative revenue for the
4:18:27revenue on the first and the second row
4:18:29then we're going to go down to the third
4:18:30row in that table and it's going to give
4:18:32us the sum of the revenue on the first
4:18:34second and third rows so when we do this
4:18:36through the whole table the result is
4:18:38this cumulative chart of Revenue
4:18:41basically from the first dat here all
4:18:43the way through to the last date which
4:18:45is about 19 billion in revenue on the
4:18:49last date in our data set which is 31st
4:18:51of May 2020 next up we have table
4:18:54filtering now table filtering uses the
4:18:56function filter and it Bally returns a
4:18:59table that represents a subset of
4:19:01another table or expression that you're
4:19:03using so in this measure we're
4:19:05calculating total revenue by category
4:19:08and so we're using the calculate
4:19:10function and we're passing in the sum of
4:19:13the revenue and then we're filtering it
4:19:15by specific product names so what this
4:19:17is going to do is create us this chart
4:19:19right so we get Revenue figures for each
4:19:22particular category because we're doing
4:19:24this filtering of particular product
4:19:27names so like Audi tataa Hyundai these
4:19:31are all product names and our filter
4:19:33expression here is basically filtering
4:19:35out the product name where is equal to
4:19:38the selected value of product name next
4:19:40up we're going to look at window
4:19:42function
4:19:43and there's actually three functions in
4:19:45Dax that a class as window functions you
4:19:47have the window function called window
4:19:50index and offset we're going to be
4:19:51focusing on the window function in this
4:19:54example and some of the use cases where
4:19:56you might want to use a window function
4:19:57well if you've used window functions
4:19:59maybe in SQL the result is quite similar
4:20:02the way that you implement it is quite a
4:20:03bit different actually so a use case
4:20:05might be for things like rolling
4:20:07averages so if you want to create a 3mon
4:20:10moving average of revenue for example
4:20:14you might want to use window functions
4:20:15obviously there's lots of ways to do
4:20:16moving averages in powerbi in DAC window
4:20:19function is one of them or our window
4:20:21can actually be a cat variable so say
4:20:24for example the average revenue for each
4:20:25department in a company so if we take a
4:20:27look at the documentation for the window
4:20:30function at a really high level
4:20:31basically what it's going to do is
4:20:33return multiple rows which are
4:20:34positioned within a given interval for
4:20:37example we're going to give it a from
4:20:39and a to window basically and it's going
4:20:42to return
4:20:43the rows that meet that criteria right
4:20:45between the from and the two dat now
4:20:47there's a few other parameters maybe to
4:20:49be aware of here including from type and
4:20:51to type so here we can specify either
4:20:53absolute or relative values so absolute
4:20:57is just going to take the whole table
4:20:58from top to bottom and pick the absolute
4:21:01value in your from uh parameter here or
4:21:04do you want it to be relative right so
4:21:06relative so if you pick a from type of
4:21:08relative rather than looking at all the
4:21:11values in this table and picking the
4:21:14first or the second or 90th value from
4:21:16top to bottom the relative is going to
4:21:18look at the current date and then look
4:21:20at maybe minus 10 might be a relative
4:21:23from parameter now on its own the window
4:21:26function is not particularly useful
4:21:27normally how it's used if we look scroll
4:21:30down to an example here it's going to be
4:21:31used in combination with some other sort
4:21:33of measure typically as you can see here
4:21:36it's used in this iterator function so
4:21:38we're using the window function to get a
4:21:41window of data and then applying average
4:21:43X on through that window and that's to
4:21:46return the 3-day average price so that's
4:21:49a bit of an example of how you can use
4:21:50the window function in practice so
4:21:52information functions are another class
4:21:55of functions that exist within the Dax
4:21:57language and and there's a lot of
4:21:58examples of information function you've
4:22:00probably used them if you've used Dax
4:22:02before like contains or contain string
4:22:05or has one value is blank is error
4:22:07selected measure user principal name as
4:22:10well it's basically going to check a
4:22:12particular value in your table and
4:22:15return depending on the information
4:22:17function that you use normally it's a
4:22:19Boolean value but sometimes it's
4:22:20something else like user principal name
4:22:22it's not actually looking in a table in
4:22:23that point it's looking at your actual
4:22:25PBI file and looking at the logged in
4:22:27user so in this example here we've got a
4:22:29table very simple table of simple
4:22:32transactions we've got transaction IDs
4:22:33and some revenues and we're using this
4:22:36function here is blank which is an
4:22:38information function and we're basically
4:22:39using it in an if statement so if this
4:22:42is blank returns true then obviously
4:22:45we're going to use no Revenue recorded
4:22:47if it returns false so I.E there is some
4:22:50value in that column then we're just
4:22:52going to write out Revenue recorded and
4:22:54then we get this sort of table here next
4:22:55up we have calculation groups and it
4:22:57provides a simple way to reduce the
4:22:59number of measures in a model or at
4:23:02least the maintenance of those measures
4:23:05in a model so say you want to create
4:23:07daily average monthly average and then
4:23:11yearly average measures and you might
4:23:13want to do this for revenue and you
4:23:15might want to do this for cost and you
4:23:17might want to do this for salaries
4:23:20there's a lot of repetition that you're
4:23:21going to have to do in this Dax code so
4:23:23what we can do in calculation groups is
4:23:26basically parameterize those Dax
4:23:28measures so that the actual maintainable
4:23:30code that you're writing is a lot less
4:23:32now you can create these now in Pia
4:23:34desktop as of a few months ago and also
4:23:37in table editor as well got an example
4:23:39here where I've created this calculation
4:23:41group and it's called aages now to have
4:23:43a look at our calculation group let's
4:23:45just go over to the modeling Tab and
4:23:48then you can see here we've got
4:23:49calculation groups averages and we've
4:23:51got some different calculation items so
4:23:54obviously to create a new one you can
4:23:55create a new calculation group here now
4:23:57in this calculation group I've actually
4:23:58got three calculation items the first
4:24:01one is just the total which is just
4:24:02selected measure that's kind of like
4:24:04your your Baseline measure and then
4:24:06we're going to reuse that or we're going
4:24:07to call that within other measures so
4:24:10like the daily average is going to call
4:24:11Select measure here the monthly measure
4:24:14is also going to call selected measure
4:24:15but this time on the month next up we're
4:24:17going to look at Dynamic string
4:24:19formatting and this is a pretty cool one
4:24:21it basically allows you to apply string
4:24:23formatting on numerical measures without
4:24:26updating the underlying data type
4:24:28underneath so it can remain as a numeric
4:24:30measure in this example here we've
4:24:32created this measure called Dynamic
4:24:34format measure we're using this some
4:24:36Revenue just as an example here in our
4:24:39example what we're doing is giving a few
4:24:41different options in our switch
4:24:43statement so if it's less than a th000
4:24:45we're not going to do any formatting at
4:24:47all if it's between a th000 and a
4:24:50million we're going to add in this K so
4:24:52thousand we're going to add in an M if
4:24:54it's in the million ranges and billion
4:24:58if it's in the billion ranges basically
4:25:00now the benefit of this is that well in
4:25:02our tables and in our this is just a
4:25:05card it's going to show a much nicer
4:25:08presentation of that number 5246
4:25:12million rather than lots and lots of
4:25:14numbers and the big benefit is that the
4:25:16difference between this and the modeling
4:25:18tab format string so if we were to click
4:25:21on a specific Revenue number here for
4:25:25example so we can maintain our format as
4:25:28whole number but we're also formatting
4:25:30the presentation right and this is
4:25:32important maintaining this format as
4:25:34whole number because we might want to do
4:25:37something like visualize this data
4:25:40within a chart so the final feature we
4:25:42are going to look at today in powerbi is
4:25:45these field parameters and field
4:25:46parameters allow the report user so the
4:25:50end user here is going to be coming into
4:25:51your report to select different
4:25:54categoric variables and also measures as
4:25:56well kind of dynamically so depending on
4:25:59the way in which they want to slice the
4:26:00data you can set up field parameters to
4:26:03give them this flexibility so in this
4:26:05example here we've got a number of
4:26:06different parameters and we're looking
4:26:08at the revenue by different parameters
4:26:11So currently we're looking at by
4:26:12location but you might also want to look
4:26:14at by dealer or by country or by model
4:26:18or by product and you're giving the end
4:26:19user the flex ability here to decide now
4:26:22to set up field parameters you can go to
4:26:25the modeling tab new parameters fields
4:26:27and in that way you can set up a new
4:26:29field parameter just give it a name and
4:26:31then pass in the different parameters
4:26:34that you want you also rename them if
4:26:35you want here as well that's going to
4:26:36set up your field parameters and then in
4:26:38the actual visuals it's going to create
4:26:40this parameter maybe I could to prove
4:26:42the naming here but it's basically going
4:26:43to look like this so it's just going to
4:26:46have this object here and it's going to
4:26:47say location is name of this it's going
4:26:50to give it an index here 0 1 2 3 so that
4:26:53when you click on the filter it's going
4:26:55to update the visual and on the visual
4:26:57side again we've got on the y- axis just
4:26:59this parameter that we created and our
4:27:01measure that we want to visualize so sum
4:27:04of Revenue now in this instance we've
4:27:06got a field parameter in the y- AIS but
4:27:08we could also have a field parameter for
4:27:10the revenue as well maybe you wanted to
4:27:11do some of the Revenue average revenue
4:27:13median Revenue whatever you would want
4:27:15and you want to give the user
4:27:16flexibility to change these dynamically
4:27:19that's how you would do that there okay
4:27:21so we've been through a lot in this
4:27:22video now let's test your knowledge of
4:27:25what we've been through and some of the
4:27:27questions that you might face in this
4:27:29section of the exam question one the Dax
4:27:32expression average X is an example of a
4:27:35and information function b a calculation
4:27:38Group C table filtering d a window
4:27:41function or E an iterator function pause
4:27:44the video here have a little think and
4:27:46then I'll reveal the answer to you
4:27:47shortly so the answer here is obviously
4:27:49an iterator function now the big clue
4:27:51here is the X at the end of average X
4:27:54which generally denotes an iterator
4:27:55function of course an iterator function
4:27:57is those ones where we're going to be
4:27:59enumerating through every row in a
4:28:00particular table and Performing some
4:28:02calculation before combining the results
4:28:05in some way depending on how you set up
4:28:07your iterator function it's not going to
4:28:09be information functions it's not going
4:28:11to be a calculation group group table
4:28:13filtering here well you might actually
4:28:15do some table filtering within your
4:28:17iterator function but average X itself
4:28:19is not table filtering function
4:28:21similarly with window functioning again
4:28:23you might use an iterator function
4:28:24average X within a window function but
4:28:27average X itself is not actually a
4:28:28windowing function question two on
4:28:30demand loading I loading only the data
4:28:33is needed for a particular query is a
4:28:35feature of which two are the following a
4:28:38import mode B direct Lake mode C direct
4:28:41query mode d large format semantic
4:28:43models or E the xmla endpoint the answer
4:28:46here is B direct Lake mode and D large
4:28:49format semantic models so we mentioned
4:28:51this when we were going through the
4:28:52slides this feature on Dem M loading is
4:28:55actually shared by two of these modes
4:28:57here so direct late mode and also it's a
4:29:00feature included in large format Mantic
4:29:02models as well on demand loading doesn't
4:29:04really make sense in import mode CU in
4:29:07import mode we've got the full data set
4:29:09there for us to query anyway now you
4:29:11could argue that direct query mode what
4:29:13it's actually doing is very similar to
4:29:16On Demand loading whenever you get a
4:29:18request for a query you're actually
4:29:20going back to the data source and you're
4:29:22querying that data source directly and
4:29:24then loading in only the data that's
4:29:26needed whatever comes back from that
4:29:28database query but I do think there is a
4:29:29distinction between what direct query
4:29:31mode is doing and specific feature
4:29:34called On Demand loading and I do think
4:29:36these two are slightly separate so
4:29:37although direct query does a similar job
4:29:39it's not actually leveraging on demand
4:29:41loading and the xmla end point e is just
4:29:44not the correct answer question three
4:29:46Dynamic format strings overcome which
4:29:49significant limitation that comes from
4:29:51using the Dax format function is it a
4:29:55the format function is slow on large
4:29:57data sets B the format function returns
4:30:00a string value so the values can't be
4:30:02used in chart visuals is it C the format
4:30:05function can't handle date local
4:30:08conversion whilst formatting or is it D
4:30:10the format function can't be used with
4:30:13field parameters so the answer here is B
4:30:17the format function returns a string
4:30:20value so the values can't be used in
4:30:22chart visuals one of the major benefits
4:30:25of dynamic format strings is that you
4:30:27actually retain the original data type
4:30:29for that particular column so if you've
4:30:31got a numeric data type maybe you want
4:30:33to format some millions or billions in
4:30:36that numeric data type you can create a
4:30:38dynamic format string that's going to
4:30:40visually format the string but the
4:30:42underlying data type Still Remains
4:30:43numeric so you can use that field within
4:30:46charts right you maybe you want a Time
4:30:48series chart that wouldn't be possible
4:30:50if you used the format string because
4:30:52that's going to return a string value
4:30:53and you can't visualize a string value
4:30:55in a chart like that so the answer here
4:30:57is B question four the Dax expression
4:31:01selected measure is most likely found in
4:31:03the construction of which of the
4:31:04following is it a a calculation item in
4:31:07a calculation Group B field parameters C
4:31:11an iterator function D large format
4:31:13semantic models or e a window function
4:31:16now the answer here is a a calculation
4:31:19item as part of calculation groups so as
4:31:21you remember you need to have that
4:31:23selected measure that's what makes it
4:31:24Dynamic and parameterizable let's say is
4:31:27that selected measure expression it's
4:31:29not part of field parameters or iterator
4:31:31functions or window functions large
4:31:34format sematic models but of course you
4:31:36could have a calculation group within
4:31:38your large format model but the most
4:31:40likely place that you're going to find
4:31:41this because it's pretty much necessary
4:31:43is in that calculation item when you're
4:31:45creating calculation groups question
4:31:47five which of the following is an
4:31:50irreversible operation which means it
4:31:52can't be changed afterwards where you
4:31:53can't go backwards is it a changing the
4:31:56cross filtering of a relationship to
4:31:57bidirectional B changing the storage
4:32:00mode of a table to import C naming a
4:32:04calculation Group D converting a
4:32:06semantic model into a large format
4:32:08semantic model C creating a window
4:32:11function so the here is B changing the
4:32:13storage mode of a table to import mode
4:32:16now this assumes that the original
4:32:17storage mode of that table was direct
4:32:19query and if we're moving it back to
4:32:21import mode that is an irreversal
4:32:23operation I'll leave link to the
4:32:24documentation there that kind of
4:32:26specifies where that's the case
4:32:28obviously with a the cross filtering of
4:32:30a relationship to bidirectional we can
4:32:32change the cross filtering that's not a
4:32:33problem of a particular relationship we
4:32:35can also rename calculation groups
4:32:37creating a window function doesn't
4:32:38particularly make sense because of
4:32:39course you can just delete the window
4:32:41function or delete the the measure now
4:32:43converting a semantic model into a large
4:32:44format semantic model is an interesting
4:32:46one now I was under the impression that
4:32:48this also was an irreversible operation
4:32:51but then when I was actually going to
4:32:52research it and test it out I could
4:32:54actually convert a large format semantic
4:32:56model back into a small sematic model so
4:32:58for that reason I've included it as
4:33:00false but let me know if you think that
4:33:02D is also a correct answer for this
4:33:04question congratulations that's the
4:33:05first part of the semantic modeling
4:33:08section of this exam complete in the
4:33:11next lesson we're going to be looking at
4:33:13model optimization and security so make
4:33:16sure you click here to join us for the
4:33:18next lesson I'll see you there hi
Secure and optimize semantic models
4:33:20welcome to video 10 out of 12 in this
4:33:23dp600 exam preparation course today
4:33:27we're going to be looking at securing
4:33:29and optimizing semantic models so this
4:33:31is the second part in our semantic
4:33:33modeling section of the study guide and
4:33:36as you can see here we're very close to
4:33:37the end of the course so we got two more
4:33:39modules after this one but today let's
4:33:41focus on semantic modeling again
4:33:43specifically we're going to be looking
4:33:45at implementing Dynamic row level
4:33:47security and object level security
4:33:50implementing incremental refresh
4:33:52implementing performance improvements in
4:33:54queries and Report visuals and to do
4:33:57that we're going to dive into the use
4:33:59cases for external tools like Dax Studio
4:34:01tabular editor 2 and then we're going to
4:34:03go into a bit more detail about okay
4:34:05what can we actually do in terms of
4:34:07improving Dax performance in Dax studio
4:34:10and also optimizing semantic model
4:34:11models using tablet editor as ever at
4:34:14the end of this video I'll be asking you
4:34:16five sample questions to test your
4:34:18knowledge of the things that we go over
4:34:20in this part of the study guide now as
4:34:22with the last video I've left links to
4:34:24really good further learning resources
4:34:26from people that are much more
4:34:27experienced in power be development and
4:34:30semantic modeling so I will caveat this
4:34:32lesson by saying you know definitely go
4:34:33and check out the further Learning
4:34:35Resources by MVPs Microsoft mvvs for the
4:34:38powerbi side a lot of this content is
4:34:40powerbi and also third party tools that
4:34:42connect to powerbi as well so let's
4:34:45start by looking at Dynamic roow LEL
4:34:47security so in general when we talk
4:34:50about row level security we're talking
4:34:52about restricting who can see what data
4:34:55at the row level in specific tables in a
4:34:58powerbi report now Dynamic Road level
4:35:01security is kind of an extension to Road
4:35:03level security by applying Road level
4:35:06security using the usable principle name
4:35:08now there are other information
4:35:10functions that we can use but user
4:35:12principal name basically gives us the
4:35:14email address of the logged in user so
4:35:17when the user logs into the powerbi
4:35:19service then behind the scenes we're
4:35:21going to get access to their email
4:35:23address and we can use that to apply
4:35:25filters to the data in our power report
4:35:28so that that user only sees the
4:35:30information that you have configured
4:35:32that they should be seeing now you can
4:35:34configure Road level security in the
4:35:36semantic model the data warehouse and
4:35:38the tsql endpoint of The Lakehouse in
4:35:41fabric but if you're using direct Lake
4:35:43mode then you want to be configuring
4:35:45Road level security in the semantic
4:35:46model otherwise you're going to be
4:35:47falling back to direct query mode and in
4:35:50general the whole One Security model in
4:35:53fabric is still kind of under
4:35:54construction so I definitely recommend
4:35:56where we are currently I would recommend
4:35:58just using the conventional semantic
4:36:00model Road level security Now one thing
4:36:01to bear in mind is that road Lev
4:36:03security only really works when the
4:36:05users that you're trying to give Road
4:36:07level security to have viewer
4:36:09permissions in a particular workspace so
4:36:11if they have more than viewer so admin
4:36:12or member or contributor technically
4:36:15this Ro level security is not going to
4:36:17be enforced because they'll have lots of
4:36:19other ways to access that data so that's
4:36:20one thing to bear in mind there need to
4:36:22be viewers in the workspace your users
4:36:24or your viewers of reports so at a high
4:36:26level these are the steps that are
4:36:28required to implement Road level
4:36:29security so the steps for implementing
4:36:32Road level security in P desktop first
4:36:34we're going to have to create a role now
4:36:37after you've created a role you want to
4:36:38select the table that you want to apply
4:36:40Road level security two you're going to
4:36:42enter a table filter Dax expression to
4:36:45configure when and who the road level
4:36:47security is applied to and we're going
4:36:49to validate that roow level security has
4:36:52been applied correctly now as well as
4:36:54roow level security we can also apply
4:36:57object level security to our semantic
4:37:00models but the way that we're going to
4:37:01be doing that is different because
4:37:04object level security can only be
4:37:05configured via third party tools such as
4:37:08tabular editor and when we talk about
4:37:10object level secur security what we're
4:37:12talking about is restricting access to a
4:37:15particular table in a semantic model or
4:37:18a specific column so you might have a
4:37:21sensitive column within your powerbi
4:37:24report and you want to restrict who can
4:37:26see that particular column of data in
4:37:28your spany model as I mentioned to
4:37:30configure object level security you need
4:37:33to use an external tool such as tabulate
4:37:35editor and similarly to row level
4:37:37security object level security only
4:37:39restricts data access for users with
4:37:42viewer permissions so we can't give them
4:37:44admin member or contributor roles in a
4:37:47workspace now the high level steps to
4:37:49implement object level security in tabul
4:37:51editor well again we're going to need to
4:37:53start by creating a role in power your
4:37:56desktop or you can also create it in tab
4:37:58editor as well then within Tabet editor
4:38:00you're going to want to find the role
4:38:01that you've created click on the table
4:38:04properties for that role set the
4:38:06permissions for the particular table
4:38:08that you want to apply object level
4:38:10security on set that permission to
4:38:12either none or read okay so obviously
4:38:15none will be if you don't want to give
4:38:16them access to that table and read will
4:38:19be if you do want to give them access to
4:38:20that table similarly we can do the same
4:38:22for a particular column as well then
4:38:24we're going to be wanting to publish the
4:38:26report to the service and add the people
4:38:28uh and groups to the particular role in
4:38:31the service next we're going to talk
4:38:32about incremental refresh in pobi now
4:38:36one thing to bear in mind here is that
4:38:37incremental refresh is a developing
4:38:39field in the world of fabric but but for
4:38:41the purposes of this exam I think what
4:38:44they are looking for is incremental
4:38:46refresh which is a feature in powerbi
4:38:49not talking about incremental refresh in
4:38:51data flow gen 2s or data pipelines or
4:38:54anything like that so I think we're
4:38:55focused here on the specific feature
4:38:57within powerbi called incremental
4:38:59refresh and typically this is used on
4:39:02large fact tables because incremental
4:39:04refresh allows you to pull in only the
4:39:06data that is changed within a given
4:39:08range right so perhaps in the last 24
4:39:12hours or the last hour you might only
4:39:15want to bring in the new data that's
4:39:17changed within that period rather than
4:39:19you know loading in all the data that's
4:39:22in that Source database and obviously
4:39:24that's going to have many different
4:39:25benefits if we can do this incrementally
4:39:27rather than allinone for starters we're
4:39:29going to need fewer refreshes the
4:39:32refreshes are going to be a lot quicker
4:39:33because you're only pulling in what's
4:39:34new your resource consumption could be
4:39:37lower as well because again you're only
4:39:39bringing in what's changed you're not
4:39:41bringing in everything every time now
4:39:42the refreshes can be more reliable
4:39:45because rather than pulling in hundreds
4:39:47of thousands or millions of rows every
4:39:49time you refresh the data set this could
4:39:51create open connections that are going
4:39:53to run really long on your database and
4:39:56have the potential to be timed out as
4:39:57well obviously when you move to an
4:39:59incremental refresh that problem is
4:40:01likely to go away because your refresh
4:40:03is going to be a lot quicker as I
4:40:04mentioned currently incremental refresh
4:40:06is only possible within the powerbi side
4:40:08so they are working on incremental
4:40:11refresh features for ETL items like data
4:40:15Factory items like the data flow Gen 2
4:40:17and the data pipeline but for the
4:40:18purposes of this exam currently when we
4:40:21talk about incremental refresh we're
4:40:22talking about powerbi and incremental
4:40:24refresh in powerbi is available for
4:40:27powerbi premium licenses only so PPU or
4:40:30premium capacity subscriptions and the
4:40:32incremental refresh policies are defined
4:40:34within power desktop so just have a look
4:40:37at how we can Implement incremental
4:40:39refresh so it starts by creating two
4:40:42parameters called range start and range
4:40:45end they must be called range start and
4:40:46range end these are reserved keywords
4:40:48for these parameters and then in your
4:40:50power query you're going to be wanting
4:40:52to apply some custom date filters to
4:40:55filter the data based on that tables
4:40:57date column to keep only the range
4:41:00between the range start and the range
4:41:01end then you're going to want to find
4:41:03your incremental refresh policy and this
4:41:06is what this looks like you're going to
4:41:07select the table you're going to set
4:41:09your import and refresh ranges and
4:41:12there's a few optional settings there
4:41:14like only refreshing on complete days
4:41:16and detecting data changes as well and
4:41:18then you're going to review and apply
4:41:19your policy and then when you publish
4:41:22your report into the service that's when
4:41:24your incremental refresh is going to
4:41:26kick in based on the policy that you've
4:41:28set and the range that you've set so
4:41:29next we're going to switch our attention
4:41:32and we're going to talk about semantic
4:41:33model performance and whenever we talk
4:41:35about performance of anything really
4:41:37it's useful to think of it in two steps
4:41:40number one is around monitoring and
4:41:42observation and Gathering data about
4:41:45what's happening in our semantic model
4:41:47in this case and then we want to talk
4:41:49about optimization so once we understand
4:41:51what's going wrong how can we optimize
4:41:53it to improve performance ultimately So
4:41:55within powerbi and the external tools
4:41:58that you can connect to powerbi there's
4:42:00quite a few different ways that you can
4:42:01monitor semantic model performance on
4:42:04this slide we'll just do a high level
4:42:05summary of all of them before digging
4:42:07into each of them in a bit more detail
4:42:09so when we're talking about power query
4:42:10perform performance there's a tool
4:42:12called the query analyzer tool when
4:42:14we're talking about analyzing the visual
4:42:16and query performance well we can use
4:42:18the performance analyzer in powerbi
4:42:20desktop we can also export that data
4:42:23into other tools like Dax Studio as well
4:42:25for a bit more fine grained analysis
4:42:27we'll take a look at that shortly so
4:42:29when we're talking about Dax performance
4:42:31we can again use Dax Studio we can bring
4:42:34in data from performance analyzer we can
4:42:36also use things like traces to monitor
4:42:39all of the different events both on the
4:42:40client side on the server side as well
4:42:42and for semantic model performance we
4:42:44can use things like the best practice
4:42:46analyzer which is a tool within tabular
4:42:48editor so for the purposes of the exam
4:42:51you're going to need a bit of a
4:42:52knowledge about what you can do in
4:42:53powerbi desktop the limits of
4:42:56performance analysis in P desktop and
4:42:58then knowing when to switch to tools
4:43:00like DAC studio and tabular editor as
4:43:03well when you want a bit more advanced
4:43:05analysis or deeper diving into the
4:43:07things that are going wrong in your Dax
4:43:09and your semantic models let's take a
4:43:10look at these three in a bit more detail
4:43:12so first let's look at Dax studio so one
4:43:15of the core use cases that we can use
4:43:17Dax studio for is loading the powerbi
4:43:20performance analyzer data in Dax studio
4:43:22for further analysis So within powerbi
4:43:25we can record different actions when
4:43:28we're using the report like clicking on
4:43:30different filters refreshing different
4:43:32pages and it's going to log the query
4:43:35time and the visual load time for every
4:43:37visual on a specific page but the UI in
4:43:41pobi for analyzing this performance
4:43:42analyzer data it's a little bit limited
4:43:44so what we can actually do is export
4:43:47that data and we can import it into Dax
4:43:49studio for a bit of a deeper dive
4:43:51analysis on that we get better filtering
4:43:53and sorting than in powerbi desktop and
4:43:55you can also view the queries behind
4:43:57each visual load now since powerbi has
4:44:00introduced the Dax query view this has
4:44:03become a bit of a less of an advantage
4:44:05for Dax Studio because now you can also
4:44:07run the underlying Dax query for a
4:44:10specific visual within the Dax query
4:44:11view in power desktop but it's something
4:44:13to bear in mind another really core use
4:44:15case of Dax studio is using the view
4:44:18metrics to look at the vertac analyzer
4:44:21so the verti PAC is the engine the
4:44:23analysis Services engine that basically
4:44:25power runs on and so we can analyze
4:44:28what's going on in there using this view
4:44:30metric we can take a look at table and
4:44:34column sizes and obviously the size of
4:44:36your table and your column has a really
4:44:38big impact on performance right and then
4:44:40we start to look a bit deeper into
4:44:43reasons why things might be slow like
4:44:46the cardinality of a column the
4:44:48different data types that you're using
4:44:49for a specific column and much much more
4:44:52we can also use the verti PAC analyzer
4:44:54to look for referential integrity
4:44:57violations and what we mean by that is a
4:45:00mismatch in the unique keys on two sides
4:45:03of a relationship you know there's lots
4:45:04more information in the verti pack
4:45:07analyzer for the purposes of the exam I
4:45:09think it's good just just to have a look
4:45:11at the verti pack analyzer understand
4:45:13what you can do there and the insights
4:45:15that you can gather from that another
4:45:16really useful feature in Dax studio is
4:45:19the trace analysis so for Trace analysis
4:45:21there's three traces really we need to
4:45:23be aware of the all queries Trace is
4:45:27going to capture different query events
4:45:29from client side tools such as power
4:45:30desktop so this is the only one of these
4:45:32traces that's going to able to capture
4:45:34both on the client side and on the
4:45:36server side your queries as well the
4:45:38query plan Trace there is only really
4:45:40going to capture
4:45:41the query plan Trace events from the
4:45:43analysis Services tabular server right
4:45:45so if you're using the Dax editor within
4:45:49Dax studio and you're using that to run
4:45:51queries and kind of analyze their
4:45:53performance that's where you're going to
4:45:54be using the query plan because it's
4:45:56going to give you the plan of that
4:45:57specific query on the analysis Services
4:45:59engine and the server timing Trace is
4:46:01going to give us the query timing from a
4:46:03server perspective so next let's switch
4:46:05over to tabul editor and look at some of
4:46:08the use cases that we might want to be
4:46:09using tabul editor for and the main one
4:46:11when it comes to optimizing semantic
4:46:14model performance is the best practice
4:46:16analyzer and this is a tool that
4:46:18basically performs a scan of your
4:46:20semantic model and it checks it for
4:46:22common issues now there's a list of
4:46:25rules that can be downloaded from GitHub
4:46:27and you can also create your own custom
4:46:29rules as well but there's a predefined
4:46:31list of rules that you can download from
4:46:33GitHub and they're organized into these
4:46:35categories we have ones that talk about
4:46:37performance ones that talk about your
4:46:38actual Dax Expressions that you're using
4:46:40error prevention formatting and
4:46:43maintenance so these are different
4:46:44categories of those rules in the best
4:46:47practice analyzer rule set now these
4:46:49checks can also be run from the tabulate
4:46:52editor CLI as part of a cicd process so
4:46:56when you deploy a semantic model from a
4:46:59development environment to a testing
4:47:02environment you might want to run the
4:47:04best practice analyzer rule sets against
4:47:06your semantic model to give you an
4:47:08automated way of analyzing ing the
4:47:11quality of that model to surface any
4:47:13particular errors that you might get
4:47:15with that model and things you need to
4:47:16be aware of from a quality and
4:47:18maintenance perspective in that model so
4:47:21I've left this use cases for Dax studio
4:47:23and Tabet editor at the end here just to
4:47:25kind of review everything we've talked
4:47:26about when it comes to Dax studio and
4:47:29tablet editor so on the Dax Studio side
4:47:31we're going to be wanting to use Dax
4:47:32Studio when we want to write and execute
4:47:35and debug Dax queries but they do
4:47:38actually need to be manually copied over
4:47:40to power desktop as Dax studo is readon
4:47:44now since as I mentioned previously
4:47:46powerbi desktop now has the Dax query
4:47:49view you can actually write and execute
4:47:51Dax queries and view the results
4:47:54similarly to how you can do in Dax
4:47:55Studio but that's a relatively new
4:47:57feature that didn't used to exist in
4:47:58powerbi we can use the verti PAC
4:48:01analyzer to understand the size of your
4:48:03semantic model as well as individual
4:48:05tables and columns within the model we
4:48:07can bring powerbi desktop performance
4:48:09analyzer data into act Studio to analyze
4:48:12it further and we can use those Trace
4:48:14analysis functionalities to analyze
4:48:17query events both on the client side and
4:48:20the server side depending on the trace
4:48:21that you select and when we're talking
4:48:22about the use cases for tabular editor
4:48:24we're going to be able to quickly edit
4:48:26data models so we can create measures
4:48:28perspectives calculation groups from the
4:48:31Dax editor within tabular editor and we
4:48:33can publish them directly into the
4:48:35semantic model so that's a bit of a
4:48:36distinction between the functionalities
4:48:38of tabular editor and D Studio in tab
4:48:41editor we can actually update our
4:48:42semantic models in the tool itself
4:48:45there's also functionality for
4:48:46automating repetitive tasks using
4:48:48scripting and as we've mentioned we can
4:48:50incorporate devops into the kind of
4:48:52tabular mod model life cycle using that
4:48:55cicd functionality in the tabular editor
4:48:58CLI and we can use the best practice
4:49:00analyzer to identify common issues in
4:49:03your powerbi static model finally a good
4:49:05use case for tabular editor is if you
4:49:07want to implement object level security
4:49:09as we mentioned before that's not
4:49:11possible currently within power desktop
4:49:13so if you want to be defining object
4:49:15level security that's going to be done
4:49:17in tabular editor okay we've covered a
4:49:18lot of ground again there so let's just
4:49:20wrap up this video with some practice
4:49:23questions to test your knowledge of this
4:49:25section of the exam question one your
4:49:27goal is to analyze performance analyzer
4:49:30data in Dax Studio to find the visual in
4:49:33your report with the longest total load
4:49:35time put the following steps in the
4:49:37correct order to achieve this so what
4:49:39you've got here is five steps in a
4:49:42process your goal is to reorder this
4:49:44list so that it makes sense for this
4:49:46particular goal that we're trying to do
4:49:48analyzing performance analyzer data in
4:49:50Dax studio so take a moment here get a
4:49:52bit of paper maybe write these down in
4:49:55the correct order and I'll show you the
4:49:56answer shortly so the answer here is
4:49:59like this so first we're going to be
4:50:01starting a recording in powerbi
4:50:04performance analizer we need to actually
4:50:05record it in powerbi to begin with then
4:50:08you're going to be wanting to click
4:50:09refresh visual just to update the
4:50:11visuals on the page or you can interact
4:50:14with the report as well if that's what
4:50:15you want to be analyzing then clicking
4:50:17stop recording then we can export that
4:50:19performance data Json and then import it
4:50:22into Dax Studio then we can go over to
4:50:24the powerbi performance Tab and sort by
4:50:27the total milliseconds descending to
4:50:29find the longest refresh time for a
4:50:31particular visual on the page question
4:50:33two you want to use the best practice
4:50:35analyzer all within tabul Editor to
4:50:38assess your Dax performance which of the
4:50:40following severity codes for best
4:50:43practice analyzer rule violations
4:50:45indicates an error is it a level zero B
4:50:49level one C Level Two and above D level
4:50:52two e level three and above so the
4:50:54answer here is level three and above so
4:50:58in the best practice analyzer within
4:51:00tabular editor we have many different
4:51:02levels of severity of the different rule
4:51:05violations level one is just for
4:51:07information only a level two is a
4:51:10warning and level three and above is an
4:51:12error so level three and above e is the
4:51:15correct answer to this question question
4:51:17three when implementing Dynamic Ro level
4:51:19security which information function
4:51:21should you use to filter the data in a
4:51:24specific table based on the logged in
4:51:26user's email address is it a user B user
4:51:29object ID C user email D user principal
4:51:33name or E username so the answer here is
4:51:37D user principle name so as you recall
4:51:39when we were talking about Dynamic role
4:51:41of security and specifically in the
4:51:43question is talking about using the
4:51:45logged in users email address and the
4:51:48information function to give us that is
4:51:50the user principal name now the user
4:51:53function doesn't exist the user object
4:51:56ID and the username does exist as
4:51:58information functions in Dax but these
4:52:01are not the correct answer they won't
4:52:02give us the email address the username
4:52:04will just give you the domain and the
4:52:06the user's name rather than the actual
4:52:09email address and and user email C is
4:52:11also incorrect that one is is made up so
4:52:14that's not what the function is called
4:52:15it's user principal name d question four
4:52:18in Dax Studio which of the following
4:52:20records queries are generated by a
4:52:23client tool like powerbi desktop a query
4:52:26plan Trace B SQL profiler C all queries
4:52:30Trace D server timings trace or E the
4:52:33verti PAC analyzer so the answer here is
4:52:35see the all queries Trace so the clue in
4:52:39the question here was talking about the
4:52:41client tool and recording queries that
4:52:43generated specifically within the client
4:52:45tool itself within powerbi desktop so
4:52:47the all queries Trace is the only one
4:52:49that's going to give you that
4:52:50information the verti pack analyzer well
4:52:52that's just going to analyze things like
4:52:54referential integrity and the sizing of
4:52:57different tables and columns within your
4:52:59semantic model not necessarily the load
4:53:01time directly within the client tool
4:53:03power VI and the SQL profiler the server
4:53:06timings and the query plan these are all
4:53:08kind of serers side backend tools
4:53:11they're not going to give you
4:53:11information about queries generated in
4:53:14powerbi so the answer here is C the all
4:53:16queries Trace question five the first
4:53:18step in implementing incremental refresh
4:53:21in powerbi is to a add a range start and
4:53:25a range end column to your data set B
4:53:27add refresh start and refresh end
4:53:29parameters to your powerbi desktop
4:53:32project C add a refresh start and a
4:53:34refresh end column to your data set or
4:53:37add range start and range end parameters
4:53:40to your powerbi desktop project so the
4:53:42answer here is D adding range start and
4:53:45range end parameters to your powerbi
4:53:47desktop project now as we mentioned
4:53:48previously range start and range end
4:53:51they're reserved keywords when we're
4:53:52talking about parameters so you do need
4:53:54to specific when you're adding in these
4:53:56parameters should be range start and
4:53:57range end answers A and C talk about
4:54:00adding columns to your data set that's
4:54:02not going to do anything or at least
4:54:04anything useful when it comes to
4:54:05incremental refresh we need those as
4:54:08parameters cuz we're going to use those
4:54:09to filter date column in the table that
4:54:12you're interested in so we're going to
4:54:13use those parameters to filter that
4:54:15table and as I mentioned previously B is
4:54:17wrong because we're using refresh start
4:54:19and refresh end rather than range start
4:54:21and range end we need to be specific
4:54:23when we using inal refresh to be using
4:54:26range start range congratulations you've
4:54:28now completed the first three sections
4:54:31of the exam study guide we only have two
4:54:34quite short sections to go so make sure
4:54:36you click here to join me in the next
4:54:38video where we'll be looking at
4:54:40exploratory analytics in a bit more
4:54:42detail see you there hello and welcome
Perform exploratory analytics
4:54:44back to the channel today we're going to
4:54:46be continuing our dp600 exam preparation
4:54:50course and we've made it to video 11 of
4:54:5312 in this series and today we're going
4:54:55to be looking at performing exploratory
4:54:58data analytics now there's this video
4:55:01and then one more video to go so we're
4:55:02very close to the end so keep going
4:55:04we're almost there and within this video
4:55:06we're going to be looking at descriptive
4:55:09diagnostic predictive and prescriptive
4:55:12analytics and specifically we're looking
4:55:14at how to implement those things within
4:55:17powerbi we'll also be taking a look at
4:55:19the data profiling tool which is part of
4:55:21the power query experience so you'll see
4:55:23that within the data flow Gen 2 and also
4:55:26within the power query engine in powerbi
4:55:28as well towards the end of the lesson
4:55:30we'll be finishing with five sample
4:55:32questions just to test your knowledge of
4:55:34this section of the study guide and as
4:55:36ever I'll be leaving links to further
4:55:39learning resources if if you want to go
4:55:40into more detail about anything I
4:55:42mention in this video and I'll leave a
4:55:45link to the school community in the
4:55:46description below if you want to grab
4:55:49those learning notes okay so just to
4:55:51kick us off I think it's worthwhile just
4:55:53going through those four types of
4:55:56analytics and looking at what they mean
4:55:58in Microsoft's own words so we'll start
4:56:01by looking at descriptive analytics so
4:56:04when we talk about descriptive analytics
4:56:06we're talking about analytics that
4:56:07interpret past data and kpi to look for
4:56:11Trends and patterns so we're looking at
4:56:13what happened in the past we're
4:56:15describing what happened in the past
4:56:17when we talk about descriptive analytics
4:56:19and we'll be going into some examples of
4:56:21descriptive analytics and some visuals
4:56:23you can use to perform descriptive
4:56:25analytics but for now let's just look at
4:56:27these definitions the second one to know
4:56:29is diagnostic analytics so diagnostic
4:56:33analytics varies from descriptive
4:56:35analytics because we're not just worried
4:56:37about what happened in the past but we
4:56:40look looking at analytics that describe
4:56:42which data element will influence
4:56:45specific Trends and the possibility of
4:56:47future events so when we talk about
4:56:50diagnostic analytics we're not just
4:56:51talking about what happened but we're
4:56:53inferring why something happened so we
4:56:56don't just care about okay what happened
4:56:58in the last 12 months but we're looking
4:56:59for more detail around particularly why
4:57:02certain events or certain Trends might
4:57:04have occurred now typically this uses
4:57:06techniques like correlation analysis and
4:57:09data mining at at least in the words of
4:57:11Microsoft and there's various techniques
4:57:12that we can use to kind of Infuse our
4:57:15visual reports in powerbi to try and
4:57:17expose some of this diagnostic
4:57:19information number three is Predictive
4:57:22Analytics so with Predictive Analytics
4:57:25we're going to be using statistics or
4:57:27machine learning as well to forecast
4:57:30future outcomes with statistical models
4:57:32and machine learning techniques as I
4:57:34mentioned now these analytics provide
4:57:36context and Clarity for future decisions
4:57:40so in Predictive Analytics we're not
4:57:41just worried about what's happened in
4:57:43the past but we're using that data our
4:57:45historic data to think about what might
4:57:48happen in the future with some amount of
4:57:50certainty and finally the next level of
4:57:54analytics is prescriptive analytics so
4:57:57in prescriptive analytics we're not just
4:57:59predicting what's going to happen in the
4:58:01future but we're providing
4:58:03recommendations and we're recommending
4:58:04actions that might be the best course of
4:58:07action to either prevent a particular
4:58:10scenario or to increase the chances of a
4:58:14particular scenario particular outcome
4:58:15for your business now with all of these
4:58:17four the clue really is in the name so
4:58:20if you're struggling to remember what
4:58:22each of these types of analytics is and
4:58:24it helps just to go back to the first
4:58:26word in each of them and really
4:58:28understand what each of them means so in
4:58:30the first one we're describing so it's
4:58:32descriptive analytics we're describing
4:58:35what happened in the past with
4:58:36diagnostic analytics we're diagnosing so
4:58:39we're not just describing what happens
4:58:42but there's some sort of causality there
4:58:44we're thinking why something happened in
4:58:46the past so we're diagnosing a
4:58:48particular Trend obviously with
4:58:50Predictive Analytics we're going to be
4:58:52predicting what happens in the future
4:58:54and with prescriptive analytics
4:58:55obviously the key word there is
4:58:57prescribe we're not just predicting
4:58:58what's going to happen in the future
4:59:00we're also prescribing a suitable course
4:59:03of action that you should follow to
4:59:05optimize a particular outcome in the
4:59:07future now one thing I'll say before we
4:59:09start here is that most of the content
4:59:12for this section of the exam comes from
4:59:15or at least is inspired by the pl300
4:59:18exam for the PBI data analyst so that is
4:59:21what we'll focus on for this section of
4:59:23the study guide for me it's a bit less
4:59:25about descriptive diagnostic predictive
4:59:28and prescriptive analytics although you
4:59:30could argue that the visuals that
4:59:31they're talking about here align to one
4:59:34of these four categories really it's
4:59:36more about powerbi and the visuals in
4:59:38powerbi when you should use which Visual
4:59:41and when you should use specific powerbi
4:59:43features so that's what we're going to
4:59:44be focusing on in this lesson and when
4:59:46it comes to visuals and choosing which
4:59:49ones you should use when well visual
4:59:52selection is somewhat subjective I would
4:59:54argue for the exam recommend that don't
4:59:57try to be too clever when you're
4:59:59thinking about visual selection so they
5:00:00might ask you when should you use this
5:00:03particular visual or given this scenario
5:00:06which visual would you choose so instead
5:00:08of trying to be too clever in instead I
5:00:10would recommend thinking about what
5:00:12Microsoft see as the main use case for a
5:00:14particular visual you know get inside
5:00:16the heads of the examiner and of
5:00:17Microsoft and think about how they want
5:00:19you to use the tools keep that in mind
5:00:21as we go through the following examples
5:00:23okay so let's go through some of the
5:00:25commonly used visuals in powerbi and
5:00:28when you might consider choosing them or
5:00:29maybe not choosing them as well so let's
5:00:31start with the table and the Matrix
5:00:34visualization so these are really good
5:00:35for visualizing fine grain details in
5:00:38your data allowing your user the report
5:00:41user to explore the data themselves
5:00:44right and these are commonly used in
5:00:45drill through functionality so when you
5:00:47want to provide the user the option to
5:00:49drill through to find more detail about
5:00:51a particular metric you might allow them
5:00:53to drill through and we'll be talking
5:00:55through that functionality in more
5:00:56detail a little bit later on they can
5:00:58also be used to display aggregate
5:01:00information so for example the revenue
5:01:02broken down by month for example one of
5:01:05the drawbacks of the table and Matrix
5:01:07visualizations that obviously it's
5:01:09difficult to to spot long-term trends at
5:01:12least visually next we have the bar and
5:01:15the column chart so when you have one
5:01:17categoric variable and one numeric
5:01:20variable for example you might be
5:01:21looking at the revenue by region these
5:01:24are particularly useful they can also be
5:01:26stacked if you want to add in another
5:01:28categoric variable into your analysis so
5:01:31maybe Revenue by region but then also by
5:01:34product type as well just to give it an
5:01:36extra layer into your analysis again bar
5:01:39and column charts they can be used to
5:01:41visualize time series information but it
5:01:43does get a bit messy if you've got lots
5:01:45and lots of time periods along your
5:01:47xaxis generally it's better to present
5:01:50time series information on a line chart
5:01:53which brings us nicely into the next one
5:01:55which is the line and the area chart as
5:01:56I mentioned this are really good for
5:01:58visualizing time series information you
5:02:01can fit a lot of information into one
5:02:03chart over large time ranges we can also
5:02:06use the legend to kind of break up a
5:02:08single line into multiple lines to
5:02:10compare how that metric is changing
5:02:13within each category over time now
5:02:15another variation of this is the area
5:02:17chart on the right hand side there my
5:02:19personal opinion is I don't really like
5:02:20the area chart it's difficult to
5:02:22interpret all these different areas
5:02:25because you know the color of each area
5:02:28is kind of impacted by the color
5:02:29underneath it so it can be difficult to
5:02:31interpret in my opinion from a user
5:02:32perspective but again this is not about
5:02:34my personal opinion it's about what
5:02:36Microsoft sees as the core use cases for
5:02:39each of these charts s next we have the
5:02:41card visualization now these are
5:02:42obviously really good for kpi metrics
5:02:46and you can also include percentage
5:02:48change metrics you can add a bit more
5:02:50context about what's happened is that
5:02:52sales amount an increase or a decrease
5:02:55since the last sales amount as well now
5:02:57out of the box by default it doesn't
5:02:59really give you that longer term Trend
5:03:02but it can be coupled with things like a
5:03:04spark line so if you want to show the
5:03:06momentum of a particular metric over
5:03:08time adding that spark line can be
5:03:10really useful to give your users a bit
5:03:12of context about how that's changing
5:03:14over time and the card visual is
5:03:16something that has been developed quite
5:03:17a lot by Microsoft over the last few
5:03:19months and years and it's actually quite
5:03:21feature Rich now so you can really add a
5:03:23lot of information into these card
5:03:25visuals the pie chart donut chart and
5:03:27tree maps are a little bit controversial
5:03:29but their main goal is to show ratios
5:03:32between different categories now the
5:03:35reason why pie charts and donut charts
5:03:37these kind of ratio charts can be
5:03:40controversial is that they're more
5:03:41difficult for the human brain to
5:03:43interpret the difference between an area
5:03:46which is what we're showing in a pie
5:03:48chart or a tree map for example and a
5:03:50more linear comparison that you get with
5:03:53like a bar chart another potential
5:03:54downside of this type of visual is that
5:03:57by presenting ratios it does somewhat
5:04:00obscure the overall numbers so you might
5:04:02be comparing the sales in USA to the
5:04:05sales in the UK as a ratio we don't
5:04:08really by default while get a view of
5:04:10the overall sales in each category now
5:04:13you can add that to the labeling and
5:04:16also potentially a tool tip to add that
5:04:17information in but I would argue if
5:04:18you're going to do that there's better
5:04:20visualizations to choose another
5:04:21downside is that we can't really see the
5:04:23trends over time so simil with the card
5:04:26visual we're only getting the point in
5:04:28time metric for that particular metric
5:04:31right we're not seeing how that metric
5:04:34has changed over time which is normally
5:04:36more useful for the user next we come to
5:04:38combo charts now these are typically
5:04:41used to visualize more than one metric
5:04:44on the Y AIS so you can see here we've
5:04:46got the bar chart showing one particular
5:04:49metric and then we've got a line chart
5:04:51showing another metric on top of the
5:04:53same visual right now this can be useful
5:04:56in some scenarios but you do have to be
5:04:58careful because sometimes this allows
5:05:00the user to come to a certain conclusion
5:05:02about correlation between these two
5:05:04metrics which you might not actually be
5:05:06correct that correlation and it can also
5:05:08be difficult for a user to interpret
5:05:12which axis is showing which metric you
5:05:14know often you need a really good Legend
5:05:16to explain the differences between these
5:05:18two axes these two metrics as well next
5:05:21we come to the funnel visualization now
5:05:23these are really good for showing some
5:05:25sort of movement through a linear
5:05:27process an example here would be
5:05:29tracking website conversion so at the
5:05:31top of your funnel you might have a
5:05:33website viewer visiting your website
5:05:36then the next stage in that linear
5:05:38process might be okay they're going to a
5:05:40product page then the next stage might
5:05:42be okay they've clicked on buy they've
5:05:44added something to their cart for
5:05:46example and then the next stage in the
5:05:48process might be they've actually bought
5:05:49the product and so on and so on so when
5:05:51you've got linear processes and you're
5:05:53trying to visualize some sort of
5:05:55movement through that process funnel
5:05:57visualizations are really useful now if
5:05:59I was to be picky I think you could
5:06:00argue that is difficult for some users
5:06:02to grasp the scale of the difference
5:06:04between each of the levels in this
5:06:07funnel but again that's probably not
5:06:08something for the exam that's just
5:06:09personal preference next is the gauge
5:06:12chart now the gauge chart is useful when
5:06:14we want to show progress of a particular
5:06:16metric towards a goal so say you have an
5:06:20annual revenue Target and you want to
5:06:22show halfway through the year that oh
5:06:24we're actually at 53% of our Target so
5:06:27we're on track for example next up is
5:06:29the waterfall visualization so the
5:06:32waterfall visualization is useful for
5:06:34showing a running total over either a
5:06:37time period or within specific
5:06:40categories right so the goal here is
5:06:42really we want to understand which
5:06:44periods or categories contribute most to
5:06:47that change to that overall figure right
5:06:49and again from personal experience the
5:06:51waterfall chart I think should be used
5:06:53carefully because it can be in my
5:06:55opinion somewhat difficult for users to
5:06:56interpret but that's just based on my
5:06:59experience next we have the scatter
5:07:01chart which is generally used when you
5:07:03want to visualize two numeric or
5:07:06continuous variables now again you can
5:07:07add more information to the legend of
5:07:09these types of visuals if you want to
5:07:12you know add a further layer of analysis
5:07:14so you might be visualizing someone's
5:07:16age versus their height and then your
5:07:19third variable that you want to add into
5:07:21that analysis might be the country that
5:07:24they grew up in or the country that they
5:07:26live that might be an extra variable
5:07:28that you want to add into this analysis
5:07:29and you can do that by color coding the
5:07:31dots also changing the shape of those
5:07:35dots as well now obviously you need to
5:07:36be careful when you do that adding a
5:07:37third variable because you know having a
5:07:40scatter chart with lots and lots of
5:07:41different categories of different colors
5:07:43can be quite difficult to interpret from
5:07:45a user perspective another kind of
5:07:47warning with these types of charts is it
5:07:48can sometimes lead the user to come to
5:07:51conclusions about correlation between
5:07:53these two variables and as we know
5:07:54correlation doesn't always equal
5:07:56causation so that's something to Bear In
5:07:58Mind as a report author you might want
5:08:01to compare these two variables when in
5:08:02reality they might not actually have a
5:08:04causative relationship next we've got
5:08:06custom visuals so custom visuals help
5:08:09you go beyond the out thebox visuals
5:08:12that come with powerbi and there are
5:08:14lots and lots of custom visuals that you
5:08:17can explore the app Source now these can
5:08:19be really useful if you want to use
5:08:22slightly more Niche visual types that
5:08:24maybe haven't made their way into the
5:08:27core powerbi visuals set yet and of
5:08:30course bear in mind that if you use a
5:08:31custom visual these are normally built
5:08:33by Third parties and sometimes they
5:08:36involve licensing and things like pay
5:08:38walls you might be able to use it for a
5:08:40short period and then you have to pay
5:08:42but that's something just to bear in
5:08:43mind around custom visuals now the final
5:08:44visual type we're going to look at is
5:08:46the Q&A visual in powerbi which allows
5:08:49users to ask natural language questions
5:08:52about the data in your underlying data
5:08:54models now this sounds good but in
5:08:56practice I think it's difficult to
5:08:57implement well obviously the questions
5:09:00currently need to be carefully
5:09:01articulated to kind of match your data
5:09:03model so it requires the user to
5:09:06understand the columns and the table in
5:09:09your data set to craft good questions
5:09:13that can give good results basically
5:09:15currently it's only available in English
5:09:17and Spanish as well if you enable it in
5:09:19the admin settings so that's a summary
5:09:21of some commonly used visuals now let's
5:09:25look at some powerbi features that can
5:09:28help us as report developers to go kind
5:09:30of an extra layer in our analysis to
5:09:33help with maybe some of the diagnostic
5:09:35elements or the predictive elements that
5:09:37we're looking for in this part of the
5:09:38study starting off with the drill down
5:09:41functionality so drill down basically
5:09:44allows your user to explore your data
5:09:46through layers of a hierarchy and that
5:09:50hierarchy can either be explicit so as a
5:09:53a powerbi hierarchy or it can be
5:09:55implicit so you maybe you haven't
5:09:56actually defined it as a hierarchy but
5:09:58there is some implicit hierarchical
5:10:01relationship between these variables now
5:10:04a classic example of a drill down might
5:10:06be to visualize at the top level of your
5:10:09drill down Revenue in a particular
5:10:11country and then you might want to drill
5:10:13down and look at Revenue by specific
5:10:16State and then within that state you
5:10:18might want to look at Revenue by store
5:10:20in a specific state for example now
5:10:22alongside drill down and drill up
5:10:24functionality you also have drill
5:10:25through and drill through is a powerbi
5:10:28feature that allows users to drill
5:10:30through from one report page into
5:10:33another report page and it carries
5:10:34through that information they click on
5:10:36so an example here you might want to
5:10:38drill through this SharePoint category
5:10:41and take them through a different page
5:10:44in the analysis pre-filtered based on
5:10:46whatever you drill through on obviously
5:10:48with this drill through you need to
5:10:49think carefully about the user journey
5:10:51in your report and the navigation
5:10:53Journey that they're going through you
5:10:55might want to add in things like back
5:10:56buttons to make sure that they don't get
5:10:59lost in your report now with both of
5:11:01these features again this is just from
5:11:03personal experience it requires your end
5:11:05user to have knowledge of drill down and
5:11:08drill through so I would argue this is a
5:11:10bit of a downside for these features but
5:11:12something to bear in mind when you're
5:11:13authoring these reports next up grouping
5:11:16so grouping allows report authors to
5:11:19group two or more categories within a
5:11:22visual we can create groups of months in
5:11:25this example or anything really that
5:11:27makes sense for the particular visual
5:11:29that you're creating now when your data
5:11:31is continuous so it's numeric you might
5:11:34want to use binning and binning
5:11:36basically allows you to create different
5:11:38bins for for these continuous variables
5:11:41now an example of this might be to
5:11:43create a salary bin so rather than just
5:11:46have salary as a number you might want
5:11:47to create salary ranges so from 0 to 30k
5:11:5130k to 60k for example and obviously
5:11:53that opens up different visual types
5:11:56that you might want to explore and use
5:11:58for this type of data some other
5:12:00features to be aware of are reference
5:12:02lines so this allows report authors to
5:12:06provide a static reference line across
5:12:08the ex for the Y AIS to give the user
5:12:11some more context so it might be the
5:12:14average sales amount for sales people so
5:12:17maybe you're you're visualizing how all
5:12:19the different sales people have
5:12:20performed and you can see all of the
5:12:21data it might be useful to show report
5:12:24users what the average is so they can
5:12:26make a comparison give them some context
5:12:28for that comparison another feature here
5:12:30is around visualizing errors so if your
5:12:32data set contains some sort of errors in
5:12:35them normally it's around prediction
5:12:37uncertainty or measurement uncertainty
5:12:41then we can use error bars to show that
5:12:44uncertainty to the user now these can
5:12:46either be markers or they can be lines
5:12:49or they can be shaded areas maybe if
5:12:51you're doing like a Time series forecast
5:12:53that kind of thing you can have a shaded
5:12:54area for the uncertainty of a particular
5:12:57prediction obviously this requires data
5:12:59on errors or prediction uncertainty next
5:13:02we have what if parameters so we can use
5:13:05what if parameters to perform some what
5:13:08I would say is kind of rudimentary
5:13:09scenario analysis so giving your report
5:13:13users the option to change a particular
5:13:16variable particular parameter in this
5:13:18case we've got discount percentage and
5:13:21they can see what impact that has on a
5:13:23particular metric in the chart below now
5:13:25to actually get whatif parameters to
5:13:27work you need to do quite a lot of data
5:13:30engineering and potentially prediction
5:13:32algorithms as well so it can be quite
5:13:35difficult to set up a good what if
5:13:37analysis finally we've got time series
5:13:39forecasting and this can be accessed in
5:13:42that analytics pane so the third icon
5:13:45when you're setting up a visual is the
5:13:47analytics Pane and here if you have a
5:13:49Time series data set you've got revenue
5:13:51numbers for the last 24 months for
5:13:54example powerbi gives you the
5:13:56functionality to create a Time series
5:14:01analysis obviously this is quite limited
5:14:03and I would argue you should be careful
5:14:05here you know if you want to be doing
5:14:06any sort of serious time series analysis
5:14:08I would argue that it shouldn't be done
5:14:09solely within powerbi but the
5:14:11functionality is there and you might get
5:14:13asked about it great so let's finish off
5:14:15this lesson with a look at the power
5:14:18query data profiling tool okay so let's
5:14:20just explore the data profiling tool
5:14:23which comes with the power query engine
5:14:25so we can use this in a data flow Gen 2
5:14:28or also in the power query engine in
5:14:30powerbi powerbi desktop if you're using
5:14:32that as well now what I've got here is
5:14:34just a query on this Revenue data set
5:14:37and I want to explore some some
5:14:39potential data quality issues in this
5:14:41data set now to do this we can go to the
5:14:43view Tab and have a look at the data
5:14:45view enable column profile and then we
5:14:47got a few different options here so show
5:14:49column quality details if we just do
5:14:51that one to begin with and let's just
5:14:53explore what that gives us so we can see
5:14:54we've got these three kind of categories
5:14:57so it goes through each column and it
5:14:58gives us a percentage of how many values
5:15:02in that column are valid how many are
5:15:04errors and how many are empty so it
5:15:06gives you a really quick indication of
5:15:08column quality Now by default you can
5:15:11see down here that the column profiling
5:15:13is based on the top 1,000 rows So
5:15:16currently is looking at the top 1,000
5:15:18rows and for this dealer ID column
5:15:20saying that none of them are empty
5:15:21there's no errors and 100% are valid now
5:15:24if we change this to the entire data set
5:15:27obviously going to take a bit longer to
5:15:29calculate but then it's going to look at
5:15:30your entire data set every value in this
5:15:33column and it's going to perform that
5:15:34same categorization as you can see this
5:15:36data set looks pretty good we've got
5:15:38100% valid for all columns no empty and
5:15:40no errors so that's good if we go back
5:15:42to the data view now we click on the
5:15:45column value distribution let's have a
5:15:47look at what that gives us so now we've
5:15:49got this quite high level view on the
5:15:53distribution of values within this
5:15:55column and if we make that a bit bigger
5:15:57here we can begin to see how many of
5:15:59these values are distinct now to explain
5:16:02the difference between distinct and
5:16:05unique I think it helps to look at this
5:16:07column here so we've got true and false
5:16:10values in every Row in this column now
5:16:13it's showing as a distinct count of two
5:16:16which is obviously referring to true and
5:16:18false so it's kind of like a distinct
5:16:20count of the values in that column
5:16:22unique means values for which there is
5:16:25only one row so here there's no unique
5:16:29values because there's true in more than
5:16:31one row you know these are not unique
5:16:33it's included in multiple rows and false
5:16:36there are actually multiple falses so if
5:16:38in our sample size here we only had one
5:16:42false value then it would be in fact
5:16:44unique so we'd get that one in the
5:16:46unique column but currently because we
5:16:47have more than one false value in this
5:16:50column it's not unique so zero unique
5:16:52values in this column so next up to
5:16:54explore we have the details pane so if
5:16:56we enable the details pane here and then
5:17:00we click on a specific kind of
5:17:02distribution we've got here make this a
5:17:03bit smaller so we get a bit more
5:17:06information about the distribution in a
5:17:08particular column so we can see the
5:17:10counts the error counts the null counts
5:17:12we can also see the distinct count
5:17:13unique count empty string counts and the
5:17:15minimum and the maximum values now
5:17:17obviously this is a minimum maximum of a
5:17:19string value but it would also work with
5:17:22numeric numbers as well I think this
5:17:24column only has unit sold one so the Min
5:17:27and the max is also one so it's not a
5:17:29particularly good example here and we've
5:17:30also got things like the average
5:17:32standard deviation number of odds number
5:17:34of evens so it give you a bit more
5:17:36detail about what there is in that
5:17:38column now once you've got a pretty good
5:17:40idea about the kind of distribution of a
5:17:42column you might be able to spot things
5:17:44like duplicate values in here what we
5:17:46can do is we can right click on it and
5:17:48we can remove duplicates or you might
5:17:50want to remove errors as well you can do
5:17:52that from just right clicking on this
5:17:54top section here and it gives you a
5:17:56quick way to remove or replace these
5:17:58errors or duplicates as well okay let's
5:18:01just round off this video by going
5:18:03through some practice questions to test
5:18:05the knowledge of what you've learned
5:18:07during this section of of the exam of
5:18:10the study guide question one a company
5:18:12annual report shows a net profit of $34
5:18:15million now the company has 12 business
5:18:18units each with their own net profit or
5:18:21loss amount which of the following
5:18:22visual types could best be used to
5:18:25visually show how each business unit
5:18:27contributed to the overall net profit
5:18:29metric is it a a line chart b a scatter
5:18:33chart c a matrix chart d a waterfall
5:18:37chart or e a question answer visual
5:18:39pause the video here have a think and
5:18:41I'll reveal the answer to you shortly so
5:18:43the answer here was the waterfall chart
5:18:46as you remember when we were going
5:18:47through the waterfall chart one of the
5:18:48core use cases of that waterfall chart
5:18:51is to break down one top level metric
5:18:54into its constituent parts to allow the
5:18:57user to explore how each in this case
5:19:00business unit contributes to the overall
5:19:02net profit metric now the other visuals
5:19:04that we've listed here you know you
5:19:06might be able to glean some of that
5:19:08information but as I mentioned
5:19:10previously what we're after here is what
5:19:12would be the best visual to visualize
5:19:14this information and for me that's the
5:19:16waterfall chart question two you want to
5:19:17add measurement error bars on a Time
5:19:19series line chart to show the potential
5:19:21error in each measurement where would
5:19:23you go to add this information to your
5:19:26visual is it a the analytics pane B the
5:19:29format visual pane C view Tab and show
5:19:32error bars D the build visual pain or e
5:19:35in the model view so the answer here is
5:19:37the analytics pain when you're
5:19:39configuring the settings for a
5:19:40particular visual there's obviously
5:19:42three panes or three different windows
5:19:44that we can use to declare different
5:19:46settings and the error bars
5:19:48functionality is included in the
5:19:50analytics pane it's not going to be the
5:19:52build visual pane that's where you add
5:19:53your different variables different
5:19:55columns to a particular visual format
5:19:57visual that's where you change things
5:19:59like the text and the border and all
5:20:00that kind of thing it's not going to be
5:20:01the model view CU that's where you
5:20:03define relationships and things like
5:20:04that and it's not going to be in the
5:20:06view tab show ER bars that doesn't
5:20:08actually exist that functionality I just
5:20:09made it up question three you want to
5:20:11visually compare two continuous
5:20:13variables age and height of survey
5:20:16respondents in one chart which of the
5:20:17following visual types could best
5:20:20represent this data is it a a line chart
5:20:23b a stacked bar visual c a scatter
5:20:26visual D question answer visual or e a
5:20:29matrix visual so the answer here is C
5:20:32the scatter visual obviously when you're
5:20:35visualizing two continuous variables so
5:20:38age and height then probably the best
5:20:40visual to use for that would be the
5:20:41scatter chart put your height on the y-
5:20:44Axis or your age on the AIS and then you
5:20:47can plot different dots for each of your
5:20:50survey respondents again the other
5:20:52visual types not really going to help
5:20:53with that kind of analysis You could
5:20:55argue that the The Matrix would
5:20:57potentially show you all of that but
5:20:58it's going to be a really big chart if
5:20:59you've got lots of respondents so the
5:21:01answer here is C the scatter visual
5:21:03question four by default the data
5:21:06profiling tool reviews the top n rows of
5:21:09your data set to show you potential data
5:21:11quality issues so what here is n so how
5:21:14many rows is it a 10 B 100 C 500 D 1,000
5:21:20or E 10,000 so the answer here is D
5:21:231,000 rows as you can see here from that
5:21:26visual by default the column profiling
5:21:29is based on the top 1,000 rows obviously
5:21:32we can change it to include the entire
5:21:34data set as well but by default it's
5:21:361,000 rows question five which the
5:21:38following features of the data profiling
5:21:40tool can help you identify duplicate
5:21:42values in a column that you plan to use
5:21:44as a key to join on is it a column
5:21:47quality details B column value
5:21:49distribution C column duplicate analysis
5:21:52D column key constraints or E group but
5:21:55so the answer here is B the column value
5:21:58distribution so as we mentioned one of
5:22:00the key use cases for the the value
5:22:02distribution is to look for duplicate
5:22:05values because if you've got a key that
5:22:07you plan to join on you want to be
5:22:09looking for duplicate values in that key
5:22:11and in the column with no duplicate
5:22:13values is going to look like this with a
5:22:15just a flat line so every value is going
5:22:17to be distinct here on the left hand
5:22:19side this would indicate that there are
5:22:21some duplicate values there's some
5:22:22values here that have a count of more
5:22:24than the others right so this would be a
5:22:27potentially problematic key on which to
5:22:29join on but this date ID column all of
5:22:31the values here are unique at least in
5:22:33the sample of data that you're using for
5:22:35profiling the column quality details a
5:22:38doesn't give us this information the
5:22:40column duplicate analysis C doesn't
5:22:42actually exist neither does d the column
5:22:44key constraints and E the group buy okay
5:22:47you could potentially use a group buy to
5:22:49look for duplicate values but it's not a
5:22:51feature of the data profiling tool
5:22:53itself congratulations we're nearly
5:22:55there only one section of the study
5:22:57guide to go in the next lesson we'll be
5:22:59looking at the final section of the
5:23:01study guide we'll be looking at how to
5:23:03use SQL to analyze data via the
5:23:06Lakehouse SQL endpoint in the to
5:23:08warehouse and also via the xmla endpoint
5:23:11as well so make sure you click here to
5:23:13join us in the last lesson in this
5:23:16series I'll see you there hello and
Query data using T-SQL
5:23:18welcome to video 12 in this dp600 exam
5:23:22preparation course and we've made it
5:23:24this is the final video in the study
5:23:27guide we're going to be looking at
5:23:28querying data by using tsql now you'll
5:23:32noticed that I've added an extra video
5:23:34there cuz I really wanted to give you
5:23:35one more video just to really help you
5:23:38prepare
5:23:38for the exam so we'll be going through
5:23:41how to prepare for the exam some other
5:23:43resources that I think you should take a
5:23:44look at before you take the exam and
5:23:47some advice for when you're actually
5:23:48sitting the exam as well but we'll be
5:23:51going through that in the next video for
5:23:52this video we're going to be focusing on
5:23:55these three sections of the study guide
5:23:57so we're going to be looking at querying
5:23:59the Lakehouse and the data warehouse
5:24:03using tsql and we're also going to be
5:24:05looking at the visual query editor which
5:24:07is a feature of both of those two SQL
5:24:10endpoints as well finally we'll be
5:24:11mentioning how to connect to and query
5:24:14data sets using the xmla endpoint as
5:24:17ever I've released study notes to help
5:24:20you go a little bit deeper just to make
5:24:22sure that you're covering off all the
5:24:24right points and links to other further
5:24:26resources if you want to go a bit deeper
5:24:27on whatever I've mentioned in this
5:24:30section of the study guide as however
5:24:31I've also got five sample questions to
5:24:33test your knowledge at the end of this
5:24:36video so let's just start by looking at
5:24:38the different ways that we can access
5:24:41the tcq engine within fabric has a few
5:24:44different ways to be aware of so one of
5:24:46them is the Lakehouse TC endpoint now
5:24:50one thing to bear in mind here as we've
5:24:51already mentioned quite a lot already is
5:24:53that this is read only so all you can
5:24:55really do here is Select statements ddl
5:24:58that kind of thing you can't do any sort
5:25:00of inserts updates deletes all that kind
5:25:03of stuff from the TC queno in The
5:25:06Lakehouse now the more obvious place to
5:25:07do tsql is within the fabric data
5:25:11warehouse here you're going to have the
5:25:12opportunity to write tsql both ddl DML
5:25:17inserts updates deletes select
5:25:19statements all of that stuff so the data
5:25:22warehouse is going to give you the
5:25:23ultimate flexibility really to write
5:25:27tsql scripts within fabric now on top of
5:25:30the tsql query editor you can actually
5:25:34also create tsql like queries using the
5:25:38visual query editor and this is quite
5:25:40similar to the data flow Gen 2 if you've
5:25:43ever used the visual editor in a data
5:25:46flow so we can do things like merging
5:25:48different tables and filtering tables
5:25:52and adding additional columns that kind
5:25:54of thing but we can do it through a no
5:25:56code visual interface and we'll be
5:25:59taking a look at that in a bit more
5:26:00detail shortly so the other option for
5:26:02writing SQL is via the xmla endpoint and
5:26:06as we've mentioned previously in this
5:26:07course to connect to that xmla endpoint
5:26:11we need to go into our workspace
5:26:12settings as you can see on the left here
5:26:13grab the connection string go into SS
5:26:17SMS in this example connect via the
5:26:20analysis Services server type pass in
5:26:23your xmla endpoint in there that's going
5:26:26to bring through all of our lake houses
5:26:28and our warehouses within that workspace
5:26:31and then we can write different queries
5:26:33depending on your use case from that
5:26:35xmla endpoint so now that we have a good
5:26:37understanding of where we can write tsql
5:26:40for most of this exam you need to have a
5:26:42pretty good level of tsql at least in
5:26:45appreciation for what a lot of different
5:26:47tsql functions do and be able to at
5:26:50least read tsql quite well and I don't
5:26:53think there's kind of a definitive list
5:26:55of which tsql functions you need to be
5:26:57familiar with but I definitely recommend
5:26:59being comfortable with the following so
5:27:02the difference between where and having
5:27:04group by summarizations Union and Union
5:27:07all different joins and when to use them
5:27:09Common Table Expressions things like
5:27:11lead and lag row number partitioning
5:27:14that kind of thing subqueries and cross
5:27:17Warehouse queries as well so rather than
5:27:19just describing all these things I think
5:27:21it would be better to jump into fabric
5:27:24open up a data warehouse and show you
5:27:26some of these functions in action okay
5:27:29so here we are in SQL Server management
5:27:31studio and I've connected to one of my
5:27:35gold data warehouses here called d gold
5:27:39using the SQL connection string and you
5:27:41can see that currently we don't actually
5:27:42have any data in this data warehouse the
5:27:45tables are empty so the first thing that
5:27:47we're going to do is just to bring some
5:27:49data from some other tables that we've
5:27:51got into this data warehouse so to do
5:27:53that I'm just going to be using this
5:27:55Seas so create table as select so it's
5:27:58going to enable us to create a new table
5:28:01in this data warehouse using existing
5:28:04data in another data warehouse or in
5:28:07this example it's it's actually a lake
5:28:08house so we're connecting to the SQL
5:28:10endpoint here so this is an example of
5:28:13cross database querying because what
5:28:15we're doing here is we're creating a
5:28:17table from another Lake housee Al
5:28:20together so we're getting the all of the
5:28:22data from this db. Revenue table in our
5:28:25Lakehouse bronze and we're creating a
5:28:27new table called dboa Revenue so then if
5:28:31we refresh these tables we should now
5:28:33have the first one which is dbo Factor
5:28:35Revenue this one here and we can do the
5:28:38same for dim date and dim Branch we're
5:28:41going to be using these data sets just
5:28:43to show some of the functions that you
5:28:45need to be familiar with for the exam
5:28:47we're not going to cover all of them
5:28:49because to be honest I'm not sure
5:28:50exactly the full breadth and depth of
5:28:52what is expected for the exam for tsql
5:28:55but we're going to go over some common
5:28:56types of problems that you might see in
5:28:58the exam so now if we refresh our tables
5:29:01again so now we've got our dbo fact
5:29:03Revenue our dim dates and our dim Branch
5:29:06so let's just start by visualizing our
5:29:09data here so specifically I'm going to
5:29:11be looking at this fact revenue and
5:29:13these data sets come from the same data
5:29:16set that we used actually in a different
5:29:18part of this exam preparation course is
5:29:20relating to Car Sales and car
5:29:23dealerships so you can see in our fact
5:29:26Revenue table we've got a revenue figure
5:29:28and we've got a date ID column and a
5:29:30branch ID column here and that's what
5:29:33we're going to be using for the rest of
5:29:35this analysis so what I'm going to be
5:29:37doing is Pres presenting you with a
5:29:38series of problems then we're going to
5:29:40walk through how you might solve that in
5:29:43SE we'll be starting off quite simple
5:29:44and then we'll be adding in more and
5:29:46more functionality as we go through so
5:29:48what if we wanted to calculate the top
5:29:50five branches by total revenue so to do
5:29:54that we're going to be needing to
5:29:56perform some sort of aggregate so this
5:29:57is what our fact table looks like
5:29:59currently we've got Branch ID and
5:30:01revenue and what we want to be doing is
5:30:03calculating the top five branches so
5:30:06what we could do is just select the top
5:30:09five here Branch IDs some of the revenue
5:30:11so what we're doing here is doing a
5:30:13simple group Buy on the branch ID
5:30:16because we're looking for the top five
5:30:17branches and we're going to order it by
5:30:19the sum of the revenue so this is the
5:30:22aggregate calculation that we're running
5:30:24on this aggregate here it's a sum of the
5:30:26revenue and we're ordering it by the sum
5:30:28of the revenue descending so we're
5:30:30getting the top five now let's change
5:30:33the problem a little bit and maybe we
5:30:35want to return only the top five
5:30:38branches in Spain or maybe the top three
5:30:41branches now Spain is a country name
5:30:45that comes from a different table in our
5:30:47data set so what we've added in this
5:30:49example is an inner join on dim branch
5:30:52and we're joining on the branch ID
5:30:54because we have the branch ID in both
5:30:56data sets we're using an inner join here
5:30:58because we only want to return data for
5:31:00which we have both keys again we're
5:31:02aggregating by the sum of the revenue
5:31:05this time we've actually brought through
5:31:07a few different columns from that Branch
5:31:10Dimension table so let's just run it all
5:31:12and see what we get here okay so this
5:31:15query has returned the top three and
5:31:17what we've done is we've filtered this
5:31:20results for only country names that are
5:31:22equal to Spain now we've had to add in
5:31:25some new things into this group by
5:31:27selection because we've also got country
5:31:29name and Branch name mentioned here so
5:31:32we've also brought brought through the
5:31:33branch name so that we can just you know
5:31:35we've got more than the branch ID maybe
5:31:37you don't not familiar with the branch
5:31:38IDs you want the actual Branch names in
5:31:41this example and we can see that that
5:31:43has brought through the top three
5:31:46branches that are in Spain by the total
5:31:49revenue so that has sorted that problem
5:31:52but what if we wanted to filter after
5:31:54the aggregate so for example give me all
5:31:56of the branches that have a revenue of
5:32:00greater than some amount so here we're
5:32:03not going to be doing the wear statement
5:32:05because we want to be using having so
5:32:07you want to be doing filtering after the
5:32:09aggregate or on that aggregate value
5:32:12then we're going to be wanting to use
5:32:13having and that's going to come after
5:32:15our group by statement so previously the
5:32:17wear statement here because we were kind
5:32:19of pre-filtering on this country name
5:32:21now we're looking for the results of an
5:32:24aggregation that are greater than in
5:32:26this example 5 million or 50 million so
5:32:28we've got a very similar setup we've
5:32:30removed the wear statement because now
5:32:32we're not just interested in Spain we
5:32:34want all the branches and we're going to
5:32:36use this having some of the Reven Vue
5:32:38greater than 50 million so here you can
5:32:40see it's returned only two branches
5:32:43which makes sense what if we change this
5:32:46having statement we removed one of the
5:32:48zeros yeah so here we're just looking at
5:32:495 million so when we've got 50 million
5:32:51there's only actually two branches that
5:32:53have more than 50 million Revenue over
5:32:56this time period or over the full data
5:32:58set in that fact table now in this
5:33:00example we've been using the inner join
5:33:02because that's what we wanted for this
5:33:04specific use case but for the exam
5:33:07you're going to be need to be familiar
5:33:09with left join right join inner join F
5:33:12outer join all of these different join
5:33:14types and when they are useful for
5:33:16different scenarios now we're not going
5:33:18to be going through all of these here
5:33:20because there's quite a lot to go
5:33:21through but I definitely recommend you
5:33:22know if you're not familiar with these
5:33:23things learning the differences between
5:33:25these and when you might want to use one
5:33:28or the other so another thing that you
5:33:30might need to be familiar with for the
5:33:32exam is commentable Expressions now
5:33:35these are really heavily used in the
5:33:38world of SQL when you're creating views
5:33:40or things like that so you need to be
5:33:41familiar with how you construct a Common
5:33:45Table expression what the the key words
5:33:47are what the general structure is and
5:33:49that kind of thing here in general the
5:33:51Comon table expression allows us to
5:33:53Define kind of like variables that we
5:33:56can use later on in our script so here
5:33:59we've got these three lines are the
5:34:02query we're selecting the top five
5:34:04Branch IDs and summing the revenue we're
5:34:06grouping by that branch ID so it's
5:34:08similar to what we saw before and so we
5:34:10can actually just select these three
5:34:11rows and have a look at what that
5:34:13returns us and then what we're doing is
5:34:15we're kind of saving that in this
5:34:16variable or this commentable expression
5:34:19called top five rev and the syntax here
5:34:21is always going to start with with which
5:34:24is the keyword for a commentable
5:34:25expression you're going to give it a
5:34:26name and then as Open brackets put in
5:34:30your expression within those brackets
5:34:32and then we can use this top five rev
5:34:34later on in our query now one of the
5:34:36other benefits of table Expressions is
5:34:38that we can create more than one of
5:34:40these so for any subsequent Expressions
5:34:43that we declare firstly we're going to
5:34:45need a comma there and then we can call
5:34:47a second one branches so for subsequent
5:34:50ones we don't need the with keyword we
5:34:52only need that on the first one and so
5:34:53here I've defined branches as select
5:34:56Branch ID and Branch name from dbo dim
5:35:00Branch so this is what this one looks
5:35:02like so we're just getting the branch ID
5:35:03and the branch name and we're storing it
5:35:05as this branches and then the final
5:35:08kind of section of a CTE is the select
5:35:11statement so you're always going to need
5:35:12to return something from this comment
5:35:14table expression and we can do that with
5:35:16a select statement so as we can see here
5:35:18we're selecting the branch ID the branch
5:35:22name and the total revenue and these are
5:35:24coming from top five rev so that's our
5:35:27keyword for our first expression that we
5:35:30defined up here and we're joining it
5:35:32with our branches which is our second
5:35:33one here we're giving it these aliases
5:35:35and to run this we have to select all
5:35:37all of the different sections and then
5:35:38press execute and we get this result
5:35:41here so we've got the top five branches
5:35:43by revenue and then we've joined that on
5:35:47the last section of our CTE back to the
5:35:49branch data set now obviously this could
5:35:52also be written in a different way you
5:35:53probably don't especially need a CTE to
5:35:56return that result but I just want to
5:35:58show you the structure of a commentable
5:36:00expression because you might get asked
5:36:02about that in the exam or you might get
5:36:04shown some code that is in this format
5:36:06and you need to understand and how it
5:36:08works another thing that you might need
5:36:09to be aware of for the exam are the lag
5:36:11and the lead function so this is a way
5:36:15of creating new columns that reference
5:36:19other columns but with some specified
5:36:22offset so probably worthwhile just
5:36:24taking a look at a bit of an example
5:36:25here so what I've done is I've defined
5:36:27another Common Table expression here so
5:36:30in the first part we getting the top
5:36:32five branches and then we're dividing
5:36:34this Revenue by 1 million just to give
5:36:38us you know some a bit easier values to
5:36:40comprehend and we're getting the floor
5:36:43of that division so it's going to round
5:36:44down to single digits in this case and
5:36:47we're defining that column as Revenue in
5:36:49millions just to make it easier to
5:36:52understand what's going on with this lag
5:36:54and Lead function that we're going to
5:36:55introduce shortly so this is what our
5:36:57data set looks like we've got five
5:36:59branches and we've got rev which are all
5:37:01singled digit whole numbers now so we've
5:37:03stored that as rev T and next we're
5:37:05going to introduce a lag function over
5:37:09this data set so we got select Branch ID
5:37:11and rev M which is looks like this and
5:37:14we're going to add on another column
5:37:16into this using this lag function so
5:37:19what the lag function is going to do if
5:37:21I just run this let's just start with an
5:37:23offset of one and we rerun this so the
5:37:25lag function is going to look at the
5:37:29column which you pass in which is
5:37:31revenue M and it's going to look at the
5:37:34value of the row number minus the offset
5:37:37so for row number one here it's going to
5:37:40look at row zero which doesn't exist so
5:37:43that will return null but for row two
5:37:45it's going to look for the value in row
5:37:48one of the column that we pass into the
5:37:51function now another thing that we've
5:37:52done here is we've ordered it by this
5:37:54rev M so the order here has changed it's
5:37:57now in ascending order from 1 2 2 3 6
5:38:00like so and we can change this lag
5:38:02function to anything we want so here
5:38:05we're doing it with a lag of two so the
5:38:08offset parameter here is going to be two
5:38:10so it's going to introduce another null
5:38:11because now we're offsetting by two and
5:38:13it looks like this now hand in hand with
5:38:15the lag function we also have the lead
5:38:18function which looks in the opposite
5:38:21direction really so rather than looking
5:38:23at the lag so Looking Back In Time the
5:38:27lead function is going to look forward
5:38:28in time or at least forward in your row
5:38:31index so it's actually going to start on
5:38:33this row here now let's just change this
5:38:36back to a offset of one so we have and
5:38:39let's just change this to lead so when
5:38:41we run the lead function here on the
5:38:44same column Revenue M with an offset of
5:38:47one this is going to be the result here
5:38:49so now our first value in this lead call
5:38:53column is actually going to be the value
5:38:54here similarly this value is going to be
5:38:56here this value is three six comes from
5:38:59here and now we have an old value at the
5:39:01end of our data set because we don't
5:39:02have you know a sixth Row from which to
5:39:05pull that data from now one thing to
5:39:07bear in mind is that we can't do
5:39:09something like this so we can't do a
5:39:11lead or a lag with an offset of minus
5:39:13one the offset parameter cannot be a
5:39:15negative value so bear that in mind now
5:39:16another function to be aware of is the
5:39:19row number and partitioning so row
5:39:22number as a name suggest basically runs
5:39:24through your data set and sequentially
5:39:26numbers your rows okay so let's just run
5:39:30this first section of code just to have
5:39:32a look at what this is doing here so
5:39:34here we've defined our row number and
5:39:36we've given at this name row num and
5:39:39we've also passed in a partition value
5:39:42so what it's going to do is it's going
5:39:43to group all of our dealer IDs so here
5:39:47you can see in the output here all of
5:39:49these dealer IDs so dealer ID DLR 001
5:39:53this first one all of these values are
5:39:55within that partition okay they're all
5:39:58that same dealer ID and the row number
5:40:00is going to assign a row number based on
5:40:03the order so the revenue so these are
5:40:05all ordered in Revenue sending order and
5:40:08then we're creating our row number 1 to8
5:40:11within this partition then you'll notice
5:40:13for the next dealer ID DLR 002 now we've
5:40:17got 10 values within this partition
5:40:19ordered in the same way and we've got
5:40:21row numbers defined here from 1 through
5:40:2410 so that's this section within the CTE
5:40:27what I've done is I've just selected
5:40:28star from Parts which is the name of our
5:40:31CTE and I've just done a bit of a wear
5:40:33statement just to only get two of the
5:40:36dealer IDs back so here you can see this
5:40:39example here we've got DLR 001 which is
5:40:42this one and we got DLR 001 7 which is
5:40:46this one so make sure you understand row
5:40:48numbering and partitioning and ordering
5:40:51by for the exam because this is
5:40:53something that could easily come up now
5:40:55the last section I wanted to mention or
5:40:57the last topic that I wanted to cover is
5:40:59subqueries so for all of the other
5:41:01examples within this tutorial or this
5:41:04video so far we've been using select
5:41:06star from table right so if we go back
5:41:10up here we' got select all of this stuff
5:41:13from dbo fact Revenue now a subquery
5:41:16basically doesn't have that structure of
5:41:18Select star from dbo doable instead we
5:41:22can actually open up a brackets here and
5:41:26rather than doing a whole table or
5:41:27getting the data from the whole table we
5:41:29can create a subquery that's going to
5:41:31basically prefilter or do some sort of
5:41:34tsql query return the results of that
5:41:36sub query back up to the top level here
5:41:39and we've got to give it this Alias of
5:41:42sub doesn't have to be sub but that's
5:41:43just what I've called it here and we'll
5:41:45see the results there so what this is
5:41:47doing is first going to calculate this
5:41:49so whatever is in our brackets we're
5:41:51going to get the revenue figures just
5:41:53for these two dealer IDs and then we're
5:41:55going to get select star from that
5:41:57result so again this is just a bit of a
5:41:59toy example just to show you subqueries
5:42:01in a bit of action at least the
5:42:02structure of a subquery just so that if
5:42:04you come across it in the exam you know
5:42:06what that is okay so here we are in a
5:42:09data warehouse and it's the same data
5:42:11warehouse that we were using for the
5:42:13first part of this video the DW gold so
5:42:15we got our fact Revenue table our dim
5:42:18date and dim branch in here now
5:42:20specifically I want to show you the
5:42:22visual query editor so we can access it
5:42:24by clicking on a new visual query and
5:42:26immediately we're going to get this
5:42:28introduction here to build a visual
5:42:30query need to drag some tables onto our
5:42:33canvas here so let's just drag on the
5:42:35fact revenue and we'll see what we get
5:42:37so the visual query engine basically
5:42:39tries to make it as simple as possible
5:42:41to transform your data that is in your
5:42:45data warehouse or also in the Lakehouse
5:42:47tsql endpoint as well so we've got our
5:42:50fact Revenue table so what can we
5:42:52actually do with it here well you can
5:42:54see along the top here we've got some
5:42:55kind of quick options for choosing
5:42:58specific columns removing columns that
5:43:00kind of thing filtering so removing
5:43:03certain rows sorting rows transforming
5:43:06we can do group by
5:43:07we can also do merging and appending
5:43:09different data sets together so if we
5:43:11want to drag another one of these onto
5:43:13the canvas as well then we can kind of
5:43:15combine those using either merge or
5:43:19append but we're not going to do that in
5:43:21this example now another thing that you
5:43:22can do within the visual query engine is
5:43:24to click on this plus button here and we
5:43:27get access to a few more commands so
5:43:29we've got all of the ones that are
5:43:31available in that top menu plus we've
5:43:33also got some transformation so we can
5:43:36do some text Transformations on a
5:43:38specific column we can do length finding
5:43:40the First characters adding columns
5:43:43using for example a conditional column
5:43:45or column from examples so if you're
5:43:47used to using the data flow power query
5:43:50engine you might be familiar with
5:43:52creating a new column from examples and
5:43:54we can also do some adding column from
5:43:56text using these methods here so say for
5:43:58example maybe I want to do a group Buy
5:44:01on this fact Revenue table and I want to
5:44:05get the median value just going to
5:44:08change this for the median and I want it
5:44:10on the revenue column and I want to
5:44:12group by our branches so we also have
5:44:15this fuzzy grouping as well so if this
5:44:18column is not particularly good quality
5:44:20you might want to add in some fuzzy
5:44:22grouping which is basically going to
5:44:24look at likeness or similarity between
5:44:27different values in that column and if
5:44:29it's above a certain threshhold so
5:44:31they're very similar so maybe there's
5:44:33just one character that's different it's
5:44:35going to add it into the same group so
5:44:37that's what fuzzy grouping would be we
5:44:38don't want to enable that there cuz
5:44:40we've already managed our data quality
5:44:42so if we do an okay here you can see
5:44:43it's add in this step here so now along
5:44:47the bottom you can see the result has
5:44:49actually updated so we've got this group
5:44:51by the branch ID and we've got our new
5:44:53column which is the median value within
5:44:55that aggregate basically so say I wanted
5:44:58to I just remove this one because we
5:44:59don't don't want to be doing anything
5:45:00with our dim date we got our fat Revenue
5:45:02now and what we can do is either save as
5:45:06table so this is going to save the
5:45:09results as a new table or we can save it
5:45:12as a view so you can see that this is
5:45:15actually gr out here cuz it's you can't
5:45:17save this as a view because the query
5:45:19fact revenue is not supported as a
5:45:21warehouse view since it cannot be fully
5:45:23translated to SQL and the reason for
5:45:25that is because we've used this median
5:45:28and median is not actually a function in
5:45:30tsql so what we can instead do is rather
5:45:32than grouping by and aggregating on that
5:45:35median value we change this to sum I
5:45:37expect that yeah so now that error
5:45:39message or that warning message has now
5:45:41gone away we might want to also change
5:45:44this to some of the revenue and then we
5:45:46can save it as a SQL view so maybe we
5:45:50want to do sum of the revenue
5:45:52aggregation give our viewer name it also
5:45:54gives you the SQL statement for that
5:45:56view which we can copy to the clipboard
5:45:58if we want and we just save it as a view
5:46:00which you can query from a powerbi
5:46:03semantic model obviously acknowledging
5:46:05the fact that with a view it's going to
5:46:07fall back to direct query mode if you're
5:46:09using direct Lake mode to access this
5:46:12data so that was just a tour of the
5:46:15visual query engine in the data
5:46:17warehouse think for the exam just
5:46:19understand what it's capable of have a
5:46:21look at the different functionality here
5:46:23understand what you can do here because
5:46:25you might get a question or two about
5:46:27the visual query engine okay so let's
5:46:29just round up this video by looking at
5:46:31some practice questions that you could
5:46:33expect for this section of the exam
5:46:35question one the SQL scripts creates the
5:46:38results shown below what is function in
5:46:42this example so we're doing select
5:46:44Branch ID column one and then your
5:46:46function which you have to work out what
5:46:48that function is from sales and
5:46:50returning that table below so take a
5:46:52look at the answers on the right hand
5:46:55side have a little think about what that
5:46:57function would be to produce the results
5:47:00in the table at the bottom of the page
5:47:02pause the video here have a think and
5:47:04then I'll reveal the answer to you
5:47:05shortly so the answer here is a we want
5:47:08to be using a lag function because we
5:47:12can see that what we're trying to do is
5:47:14make a transformation from column 1 to
5:47:17column 2 and we can see the pattern
5:47:19there that on Row three in column 2
5:47:22there's obviously an offset going on
5:47:24there and we can see the offset is two
5:47:27because Row three in column two maps to
5:47:31row one in column one and we know there
5:47:32going to be a lag function because
5:47:35column 3 is actually two rows behind
5:47:38what's on column one and you can see
5:47:40that replicates below with rows four and
5:47:43five as well and another clue here is
5:47:45that we' got two null values at the Top
5:47:47If You Got null values at the top it's
5:47:49going to be a lag function because it
5:47:50doesn't have a value for that Row one in
5:47:53column 2 and row two in column two so
5:47:56that one's always going to be null so we
5:47:57know it's got to be a lag function and
5:47:58we know that the offset is two so it
5:48:00can't be B which is a lead function the
5:48:03lead function is going to look ahead of
5:48:05time rather than looking back in time
5:48:07and the bottom two there have an offset
5:48:09of minus two so the offset in a lag a
5:48:13lead is always positive so those two
5:48:15would not be correct either question two
5:48:19you have the following query analyzing
5:48:21sales data for various products your
5:48:24goal is to analyze the sales data by
5:48:26product name and year but only for the
5:48:28products that have a yearly sales amount
5:48:31of more than $50,000 how would you
5:48:33complete this query so take a look at
5:48:36the different options there have a think
5:48:38about how you would come to that
5:48:40conclusion that this question is asking
5:48:42for and I'll reveal the answer to you
5:48:44shortly so the answer here that we were
5:48:46looking for is B so the question here is
5:48:49asking you to analyze sales data by
5:48:53product name and by year so that is the
5:48:56first clue here in our group by we need
5:48:59to be grouping by the product name and
5:49:02the year not the product key and the
5:49:05date key cuz that wouldn't be the right
5:49:07aggregation for what the question is
5:49:09asking here now the second part of the
5:49:11question is we're asking for products
5:49:14that have a yearly sales amount of more
5:49:16than 50,000 so the yearly sales amount
5:49:18of more than 50,000 that's going to come
5:49:21from our sum of the sales amount and the
5:49:23sum of the sales amount is going to be
5:49:24calculated during that aggregation so we
5:49:26need to be using having here what we're
5:49:28trying to do here is filter out the
5:49:31result of the aggregate we're not
5:49:32filtering out before the aggregate cuz
5:49:34that would just filter out individual
5:49:36rows in this table we want to be looking
5:49:39at the result of an aggregate function
5:49:42where the sum of the sales amount not
5:49:44just individual sales amounts the sum of
5:49:46the sales amount is more than $50,000 so
5:49:49the answer here is B we need the group
5:49:50buy product name and year and then
5:49:53having we need to use the having
5:49:54statement to only return the product
5:49:57names with a yearly sales of more than
5:49:5950,000 question three you're trying to
5:50:01inspect a join between two tables to
5:50:04spot referential integ violations which
5:50:07of the following tsql join types would
5:50:10be easiest to identify keys on both
5:50:13sides of the join that do not have a
5:50:15match on the other side of the join it's
5:50:17a left join B right join C inner join D
5:50:21full outer join e cross join so the
5:50:24answer here is going to be the full
5:50:25alter join so again there's two parts to
5:50:28this question really firstly we need to
5:50:30understand what a referential Integrity
5:50:32violation is so referential Integrity is
5:50:35when you're joining two tables together
5:50:38obviously you're going to be joining on
5:50:39a specific key now a referential
5:50:42Integrity violation would occur when one
5:50:45of the keys in the left hand table is
5:50:48not in the right hand data set and vice
5:50:52versa as well so to be able to identify
5:50:54that using a SQL join we're going to be
5:50:57needing to use the full outer join
5:51:00because the full out to join is going to
5:51:02bring back firstly the instances where
5:51:05those keys do match and then secondly
5:51:08all of the other results that don't
5:51:10match from both tables so when we get
5:51:13this back we're going to bring back some
5:51:15null values where there isn't an
5:51:17appropriate join or is an appropriate
5:51:19match on that join for the table on the
5:51:21left and equally the same on the table
5:51:24on the right so that's going to help us
5:51:25identify referential Integrity
5:51:27violations by investigating where there
5:51:30are null values in that output the left
5:51:32joint and the right joint they're not
5:51:34going to help us here CU we need to
5:51:35identify VI Rel ation on both sides of
5:51:38that join the inner join is only going
5:51:40to return the the matching keys from
5:51:42both sides and the cross join is just
5:51:44going to return us all the different
5:51:45combinations of the keys that exist in
5:51:47these tables so that's not going to help
5:51:49us really identify referential Integrity
5:51:52violations question four you have two
5:51:54warehouses Warehouse 1 and Warehouse 2
5:51:57you want to create a SQL view in
5:51:58Warehouse 2 that combines data from both
5:52:01warehouses Your solution should minimize
5:52:03dwell effort which solution do you
5:52:05recommend is it a a user data pipeline
5:52:07copy activity to copy the table in
5:52:10Warehouse 1 to Warehouse 2 B create a
5:52:13shortcut from Warehouse 1 to Warehouse 2
5:52:16to perform the query C use cross
5:52:18database querying between Warehouse 1
5:52:21and Warehouse 2 or D use a data flow Gen
5:52:242 to read the table in Warehouse 1 and
5:52:27set Warehouse 2 as the destination so
5:52:29the answer here is C use cross database
5:52:32querying between Warehouse 1 and
5:52:34Warehouse 2 now within a warehouse we
5:52:36can query any other Warehouse within the
5:52:39same workspace so this would be the most
5:52:42economical or require the least amount
5:52:44of effort because we can just query
5:52:46directly the data the table Warehouse 1
5:52:48from Warehouse 2 B would be incorrect
5:52:51create a shortcut because we can't
5:52:52actually shortcut from a warehouse one
5:52:54to Warehouse 2 functionality doesn't
5:52:56exist at the moment and the data
5:52:58pipeline copy activity and the data flow
5:53:00Gen 2 wouldn't minimize the development
5:53:04effort so that would technically meet
5:53:06the requirements apart from the
5:53:08requirement that says your solution
5:53:10should minimize development effort and
5:53:12this is something that you might see on
5:53:14quite a few questions within the exam so
5:53:16when you see that it's a bit of a flag
5:53:18that you should always look for the most
5:53:20efficient way of doing things or
5:53:22normally there's a method that requires
5:53:24little to no effort against some that
5:53:26require more development effort so the
5:53:28answer here is C using cross database
5:53:31query question five you're using the
5:53:33tcal query editor and your goal is to
5:53:36add a new column to your data set called
5:53:39salary bins and this column is going to
5:53:41bin a continuous salary variable into
5:53:45three different bins less than 30,000
5:53:48between 30,000 and 60,000 and more than
5:53:5160,000 which functionality should you
5:53:53use to add this new column A add
5:53:56additional column B add column from
5:53:59examples from all columns C add column
5:54:02from examples from selection D duplicate
5:54:04the column e add column from text so the
5:54:07answer here is a add a conditional
5:54:09column so when we're in the TC cor
5:54:12visual query editor there's a few
5:54:14different options there to help us add
5:54:16columns and for this specific use case
5:54:18creating a new column which is going to
5:54:20bin the values in a specific column
5:54:23we're going to be wanting to use the
5:54:24conditional column because that's where
5:54:26we can add in the logic for less than
5:54:29less than or equal to and more than and
5:54:31we can add more than one conditions on
5:54:32that column which would help us achieve
5:54:35our goal of creating the bins in this
5:54:37salary bins column now add columns from
5:54:39examples so these two are functionality
5:54:42within the tsql visual editor but it's
5:54:44going to be very difficult to implement
5:54:47that logic from a columns from examples
5:54:51it's not really a good use case for that
5:54:53type of adding column functionality and
5:54:55D duplicating the column well that's not
5:54:57going to do much it's just going to
5:54:58duplicate the column and similarly e
5:55:00adding a column from text that's also
5:55:03not going to achieve what we're looking
5:55:04for so the answer here is a adding a
5:55:07conditional column we did it
5:55:08congratulations we made it to the end of
5:55:10the study guide this is video 12 out of
5:55:1412 and we've covered an awful lot of
5:55:15ground over the last 6 weeks so well
5:55:18done for sticking with it it's a very
5:55:19tough exam we've covered a lot of
5:55:21different things this exam covers a very
5:55:24wide range of topic from tsql py spark
5:55:28Dax all of the different planning and
5:55:31all that sort of things so to get this
5:55:32far is a really good job now as a bonus
5:55:35I'll be recording one more video in this
5:55:38series basically to help you prepare to
5:55:41take the exam so things you need to know
5:55:44before the exam when you're booking your
5:55:46exam and whil you're in the exam to try
5:55:49and help you get as good a score as
5:55:51possible so thank you very much for
5:55:53joining us in this series I'll see you
5:55:56in the next video which will be the
5:55:57final video I look forward to seeing you
5:55:59all there hey everyone you thought the
TOP TIPS for the exam
5:56:02series was over but I've just got one
5:56:04more bonus video in this dp600 exam
5:56:08preparation course I really wanted to
5:56:10bring you some top tips for the exam so
5:56:13we're not going to be covering any of
5:56:14the technical content but I just have
5:56:16some words of advice or some tips that
5:56:19might help you when you're actually
5:56:21doing the exam itself so this is the
5:56:23bonus round of our course plan if you're
5:56:25just joining us here in this video I
5:56:27recommend you go back through all of the
5:56:28last 12 videos CU that's where we cover
5:56:30most of the content and I wanted to talk
5:56:31to you about booking your exam preparing
5:56:35for the exam some some advice for during
5:56:37the exam and then also what you should
5:56:39do after the exam as well so when you're
5:56:41booking the exam if you've used the AI
5:56:44skills challenge voucher then make sure
5:56:47you schedule your exam before the
5:56:49expiration date and that means you have
5:56:51to schedule the exam to be conducted
5:56:54before that expiration date now I think
5:56:56that was the 24th of June 2024 but
5:56:59double check that in your own emails
5:57:01that you have that date there now if you
5:57:03don't have that AI skills voucher that
5:57:06was given away maybe 2 months ago now
5:57:08then you can get a 50% off voucher for
5:57:12the exam by completing the cloud skills
5:57:15challenge I'll leave a link to that in
5:57:16the description so if you want 50% of
5:57:18the dp600 exam and you still haven't
5:57:21booked your exam yet or you still
5:57:22haven't got a voucher for it you can use
5:57:24that there and I've said this before but
5:57:26I always recommend if possible to take
5:57:28the exam in person now you might have
5:57:31heard there's a lot of horror stories
5:57:33for people that take it online you know
5:57:34have an online proor experience and
5:57:37they're very very busy and there's
5:57:38always problems with technology and you
5:57:41know you have to clear your desk and
5:57:43clear your room and there's all these
5:57:45kind of hurdles that if you can avoid by
5:57:48going to an inperson test center I would
5:57:50very much recommend doing that
5:57:52appreciate not everyone lives near a
5:57:54test center but if you do or you can
5:57:56commute to one just for the exam I would
5:57:58definitely recommend that reduces a lot
5:58:00of stress you just walk in take the exam
5:58:03and walk out so preparing for the exam
5:58:05well obviously you can go back through
5:58:07these videos and look in the school
5:58:10Community as well I've got lots of notes
5:58:12in there the reason I kind of created
5:58:14those was to help people who are you
5:58:15know just about to take the exam they
5:58:17can go through the key points and just
5:58:19make sure they remember everything as
5:58:21they go through each of the different
5:58:23sections in that school community and if
5:58:25you're still not sure about anything in
5:58:27any section of the exam you can dig into
5:58:29the further learning resources that I've
5:58:31linked in every section there now one of
5:58:33the key resources that I would
5:58:34definitely use when you're preparing is
5:58:36the official practice assessment for the
5:58:38dp600 exam and again I'll leave a link
5:58:40to that in the description now you can
5:58:42go through this multiple times I think
5:58:44it's a 50 question assessment but you
5:58:46can refresh the page and you get fresh
5:58:48questions when you refresh I don't know
5:58:50exactly how many different questions
5:58:51there are in that exam set but there's
5:58:54definitely more than 50 so go through
5:58:56that practice assessment numerous times
5:58:58to get a really good idea of the types
5:59:00of questions that you're going to be
5:59:01asked and to highlight any gaps in your
5:59:04knowledge another really good resource
5:59:06is the Microsoft learn together series
5:59:08that's kind of been running side by side
5:59:10with this series I think we started
5:59:12around the same time I think they
5:59:13finished maybe a week or two ago now
5:59:15this is a really good series delivered
5:59:17by Microsoft MVPs so there's lots and
5:59:20lots of content that you can go through
5:59:22there to dig in a bit deeper about any
5:59:24of the aspects in the dp600 study guide
5:59:27again I'll leave a link to that in the
5:59:28description as well as the technical
5:59:30content for Microsoft MVPs they've also
5:59:33got a few lectures or a few videos that
5:59:36describe to you what the exam entails so
5:59:39if you've never taken a Microsoft exam
5:59:42before I would definitely recommend
5:59:44watching some of those videos I think
5:59:45they're at the end of that Series where
5:59:47they walk through types of questions you
5:59:49might get how to prepare for it how much
5:59:51time you have all of that kind of stuff
5:59:53just so that you can enter that exam as
5:59:55prepared as possible now as well as that
5:59:57there's lots of other great content from
5:59:59other people in the community definitely
6:00:02recommend data Mozart Nicola's blog
6:00:05there and also you you Channel as well
6:00:07now the data guy newsletter on LinkedIn
6:00:10something managed by Abu Baka he
6:00:12basically collated a lot of learning
6:00:14resources for each section of the study
6:00:16guide definitely recommend that
6:00:18similarly there's a post on serverless
6:00:20squl by Andy Cutler which is basically a
6:00:22guide to the exam as well and in there
6:00:25he highlights a lot of resources for the
6:00:27exam as well so during the exam I would
6:00:30say don't spend too long on one question
6:00:33if you get stuck then you can flag it
6:00:35for VI and you can come back to it at
6:00:37the end if you have time now a lot of
6:00:39people are saying that there's not much
6:00:42time in this exam you have 100 minutes
6:00:45to actually do the exam and there's
6:00:47normally around 55 to 60 Questions and
6:00:50some of the questions are very wordy so
6:00:52it might take you 30 seconds or even a
6:00:54minute just to try and understand what
6:00:56they're asking you so if you get stuck
6:00:58and you don't know the answer don't
6:01:00waste too long on one question just flag
6:01:02it for review and come back to it at the
6:01:04end now you will receive a at least one
6:01:06case study question where you'll be
6:01:08given a lot of context about a business
6:01:11and you'll be asked a series of
6:01:13questions afterwards now I definitely
6:01:15recommend that you use a pen and paper
6:01:17to draw the case study architecture as
6:01:20you're reading that question because
6:01:22normally these case studies they're very
6:01:24complex and they'll describe a scenario
6:01:27to you for example oh company has
6:01:30capacity one and in capacity one there's
6:01:32workspace one and workspace 2 and in
6:01:34workspace one there's Lake housee one
6:01:36and Lake housee when you're reading this
6:01:38you don't really take any of it in so I
6:01:40definitely recommend drawing it as
6:01:42you're reading it this can really help
6:01:44you when you go forward to answering the
6:01:46questions because you don't have to keep
6:01:48on going back to the the case study the
6:01:50context when you're answering the
6:01:52questions you can just refer to your
6:01:54your diagram if you don't know an answer
6:01:56guess you don't lose marks for an
6:01:58incorrect guess so you might as well and
6:02:01bear in mind that the questions were
6:02:03created many months ago so I don't think
6:02:06they've updated the questions since it
6:02:08was first came out so bear that in mind
6:02:11so a lot of the more recent features
6:02:14they are not going to be the correct
6:02:15answer because they weren't in general
6:02:17availability when the questions were
6:02:19created so bear that in mind now a lot
6:02:21of the questions I would argue are
6:02:23ambiguous or at least the answers so
6:02:25what you have to bear in mind or try and
6:02:27keep asking yourself is what do I think
6:02:29that the Microsoft examiner is expecting
6:02:32to see here so don't try and be too
6:02:34clever always think about what AIC
6:02:35Microsoft trying to teach us about
6:02:37Microsoft fabric how do they want us to
6:02:39use the platform and always answer your
6:02:41questions with that in mind so after the
6:02:43exam so you will know if you pass or
6:02:46fail directly after the exam you get the
6:02:48results straight after now if you fail I
6:02:51would say don't be too hard on yourself
6:02:52honestly it's a very very tough exam it
6:02:55expects you to be familiar with a very
6:02:57wide range of topics right for starters
6:03:01SQL py spark M Dax plus all of the
6:03:04planning stuff and all the lake houses
6:03:07data pipelines data flows all of this
6:03:10stuff is a very wide ranging exam that
6:03:12you need to be at least familiar with so
6:03:14if you do fail don't be too hard on
6:03:16yourself at all and if you pass then
6:03:18well congratulations you can be very
6:03:20very proud of your achievement so that
6:03:22is all I have thank you so much for
6:03:25joining me in this series and the best
6:03:27of luck for all of you that are taking
6:03:29the exam in the future thank you so much
6:03:31for joining me in this series I'm going
6:03:34to be taking a short break now and be
6:03:36back on the Channel with more videos
6:03:38teaching Fabric in the future thank you