Free YouTube Transcribe

Video transcript

DP-600 Exam Full Course (6+ hours) | Microsoft Fabric Analytics Engineer

Learn Microsoft Fabric with Will · 70,834 words · 322 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Course Introduction

0:00hey everyone and welcome back to the

0:01channel and today we've got a very

0:03exciting video because we're going to be

0:05starting a brand new series here on the

0:08channel we're going to be looking at how

0:10you can become a Microsoft certified

0:13fabric analytics engineer bit of a

0:15mouthful you might know it better as the

0:17dp600 exam that you need to pass if you

0:20want to get this certification now a lot

0:22of people have been asking me about

0:24dp600 how do I get certified what do I

0:27need to learn and so in this course what

0:29I'm going to be doing is bringing

0:31together everything that you need to

0:33know for this certification so that you

0:35can hopefully pass first time so what

0:38are we going to be doing in this video

0:39well this is a bit of an introduction

0:41okay so we're going to be looking at

0:42what's in the exam what are some of the

0:45details of the exam so how long is it

0:47how how can you pass do you do it online

0:49or in person that kind of thing we're

0:51going to look at an overview of this

0:52course so specifically what are we going

0:55to learn in this course in YouTube and

0:57how are you going to learn as well then

0:59we're going to finish by looking at why

1:01you might want to take this course and

1:02why you might want to get fabric

1:04analytics engineering certification in

1:06general so what's in the exam so the

1:09first section is around planning

1:11implementing and managing solutions for

1:14data analytics now within that we've got

1:16various kind of sub modules okay so you

1:20need to understand the requirements how

1:22do we identify requirements for

1:24Solutions how do we do things like

1:26security how do you know the difference

1:28between different data gateways how do

1:30you set up access control and workspaces

1:34and capacities and how do you modify the

1:36settings of all these things as well

1:38another important section of the exam

1:40and being an analytics engineer in

1:42general is Version Control you'll need

1:44to understand how to set up Version

1:47Control for Azure devops understand some

1:50of the settings and the configuration

1:51options that that entails as well as

1:54some of the deployment pipeline

1:56functionality as well now this is not a

1:58completely exhaustive list of what's in

2:00that section of the exam we'll go

2:02through each of those in a bit more

2:03detail it's just the high level kind of

2:05areas that are covered in that section

2:07and this section is worth 10 to 15% of

2:10the exam so it's a small section but

2:12it's a very important section I think

2:14anyway in terms of becoming an analytics

2:16engineer these are really important

2:18topics that you need to understand so up

2:19next we've got 40 to 45% of the exam so

2:23this is really the core of the exam that

2:25you need should probably spend most of

2:27your time studying okay and it's a

2:29around preparation and serving of data

2:33so this is one of the core tasks that

2:35you'll be asked to carry out as an

2:38analytics engineer and within that we've

2:40got quite a lot of really important

2:42topics okay so you've got understanding

2:44The Lakehouse how do we set one up how

2:47do we create tables difference between

2:49tables and files the warehouse so the

2:51tsql experience creating shortcuts

2:54ingesting data from external locations

2:57into our warehouse or our Lakehouse what

3:01are the different methods that we can

3:02choose here and when do we choose

3:04specific ones okay then we've got data

3:06transformation so once we've got our

3:08data within fabric we might want to do

3:10some Transformations on it now you can

3:12do that with tsql spark and you will

3:15need to know at least the basics of tsql

3:17and Spark and also Dax in the next

3:19section as well so there's quite a lot

3:21of languages that you need to know to

3:23kind of an intermediate level I would

3:24say for this exam next we've got

3:26performance so how do we optimize

3:29performance both in terms of the getting

3:32data into fabric but also in terms of

3:34the transformation piece so if you've

3:36got spark jobs that are running really

3:39long or you've got some tsql scripts

3:41that are really not very performing very

3:43well they're taking a long time to kind

3:45of return your data what can you do to

3:47monitor performance then optimize it and

3:49as I mentioned you will need to know P

3:51spark and T SQL to a kind of

3:53intermediate level there so that's

3:55something to bear in mind for this exam

3:57so the next section of the exam is

3:58around implementing and managing

4:00semantic models and this is worth 20 to

4:0325% of the exam and within this category

4:06you've got the different storage modes

4:09so direct query import mode and direct

4:12late mode and kind of understanding how

4:14these things work how to set up direct

4:16Lake mode when to use it maybe when not

4:18to use it you're going to need to have a

4:20good understanding of Dax here there's

4:22quite a few questions around Dax and Dax

4:25studio and tabul editor 2 as well so

4:28these are kind of things that you need

4:29to be be aware of because you'll

4:31probably get some questions around that

4:32as well also within the section you've

4:34got data modeling so things like Star

4:36schemas Bridge tables how do we deal

4:39with many to many relationships and also

4:41we've got things around security how do

4:43we set up roow level security within

4:45your santic model object level security

4:47and how do we validate that that's

4:49actually working correctly so in the

4:50final section of the exam which is worth

4:52around 20 to 25% again is explore and

4:56analyze data so here we're really

4:58looking at the analysis part of being an

5:01analytics engineer because most of your

5:03role might be in around the engineering

5:06piece but really you need to know the

5:08analytics side of thing as well because

5:10that's going to make you a much better

5:11analytics engineer so in this section

5:13you're going to be asked questions about

5:14data analysis specifically using tsql

5:17okay so analyzing your lake house SQL

5:20endpoint analyzing your data warehouse

5:23coming up with insights from data using

5:25tsql you'll also be asked about data

5:27profiling so understanding the pro

5:29profile different tables based on some

5:31of the profiling metrics that you get

5:33and also analyzing data via the xmla

5:35endpoint so that's another part of this

5:38section of the exam so these are the

5:39four sections there's quite a lot to go

5:42through and it covers quite a broad

5:44range of skills that you need to know

5:47from ppar tsql Dax data modeling getting

5:51data in data modeling serving data in

5:53semantic models as well so quite a broad

5:55range endtoend exam so in the exam there

5:58is between 40 and 60 question and the

6:02results are scaled and you get given a

6:04result between 0 and 1,000 so the pass

6:07Mark here is 700 out of 1,000 and that's

6:10the scaled score so that's something to

6:12bear in mind and the questions can be of

6:14different types okay so some of them are

6:16multiple choice some of them might be a

6:18case study and the case studies normally

6:20take two three four questions so you

6:22really need to understand what's going

6:23on and you get multiple questions about

6:25the same case study you also have things

6:27like drag and drop and ordering lists

6:29and for a full list of the question

6:31types recommend you go to this resource

6:33here it's on the Microsoft learn the

6:36exam question section of Microsoft learn

6:39I'll leave a link to that in the

6:40description below you can also use the

6:42exam sandbox that provides you with a

6:44Sandbox experience of the exact question

6:47types that you can expect in the exam

6:49you'll get given 100 minutes to actually

6:51carry out the exam you should set aside

6:53at least 2 hours though so 120 minutes

6:56for the exam just so that you can kind

6:57of get in get settled as I mentioned

7:00it's 700 out of 1,000 to pass this exam

7:03so 70% you can take the exam either in

7:06person in Pon view sensors or you can do

7:09it online so my personal recommendation

7:11would be to take the exam in person if

7:14you can um it kind of eliminates a lot

7:16of the doubts and the problems around

7:19like Wi-Fi and worrying about your desk

7:21setup and your room has to be a very

7:23specific layout and so if you can visit

7:26a center in person I think it's a lot

7:28better cuz you can just walk in take the

7:30exam and walk out whereas online you

7:31have to think about lots of different

7:33things for me it's less stressful to do

7:36it in person if you can so there is an

7:39exam fee which is

7:40$165 for people in the USA it varies for

7:44different countries but if you look on

7:46the right hand side we got a free exam

7:48so if you're very quick and you go to

7:51the link in the description and you

7:53complete the fabric AI skills challenge

7:55training course before the 19th of April

7:58so you've only got a few days you can

8:00get a voucher to take the DP 600 exam

8:04for free so you can do that skills

8:06challenge training very quickly get your

8:09free voucher and then watch all of this

8:11series on YouTube and as long as you

8:13take that exam I think it's before June

8:15the 22nd or something so you've got two

8:17or three months to kind of go through

8:19the material at your own pace and then

8:21you get to save yourself

8:23$165 or the equivalent in your currency

8:26in your country so this is what this

8:28series is going to look like here on

8:30YouTube we've got the first video which

8:31is an introduction to the course then

8:33I'm going to be covering all 11 chapters

8:36of the exam and the content is going to

8:37be delivered through various real world

8:40scenarios CU I want to make this content

8:42interesting for you and also make it

8:44relevant for you if you want to be an

8:47analytics engineer or if you are an

8:48analytics engineer currently in your

8:50career and to do this I'm going to be

8:52combining Theory so I think there is

8:54some theory in some of these modules you

8:56do have to understand how things work

8:58but then practice as well how to

8:59actually implement this stuff in fabric

9:02at the same time throughout the course

9:03I'm going to be reinforcing that

9:05knowledge as we go through and as asking

9:07rhetorical questions as well as sample

9:10questions and at the end I do plan to go

9:12through a full video kind of like a

9:13practice paper let's say designing lots

9:16of questions that you can expect within

9:18the exam as with all of my courses all

9:21of the resources and module notes and

9:23scripts and notebooks and all this kind

9:25of thing I'll be posting that in the

9:27school community so if you're not

9:28already a member there I'll leave a link

9:30in the description below it's completely

9:32for free so make sure you sign up and

9:34yeah you'll get access to all of that so

9:36there's a few reasons that I just want

9:37to cover quickly here around why you

9:39might want to take this course and then

9:41go on to become a certified analytics

9:44engineer well the course gives structure

9:46to your learning Microsoft have given us

9:49a study guide and they've said these are

9:51what we think is important to learn for

9:54an analytics engineer so it gives you a

9:55really good Pathway to follow might also

9:58help you get a new job having that

10:00certification on your CV is definitely

10:02not going to hinder your chances it

10:04would also be good to go into kind of

10:06promotion talks or payiz talks with your

10:08boss and say yeah well last year I did

10:10the certification and this is especially

10:12true if you work in a consultancy right

10:15because consultant here they're kind of

10:16selling your skills and your experience

10:18so if you have certification then that

10:21can help them win work in the future so

10:23in this lesson we've looked at what's in

10:25the exam we've looked at some of the

10:26exam details we've also looked at the

10:29overview of this course so what are the

10:31different chapters that we're going to

10:32be looking through and why I think you

10:34should take this course join us in the

10:37next lesson where we'll be starting the

10:40course properly and we'll be looking at

10:43how to plan and Implement a data

10:45analytics solution hello and welcome to

Plan a data analytics environment

10:48this first chapter in this dp600 exam

10:52preparation course we're going to be

10:54looking at how to plan a data analytics

10:57environment in fabric now this is the

10:59first chapter in 11 chapters that we're

11:01going to be going through teaching you

11:03everything you need to know to hopefully

11:04pass the dp600 exam in this chapter

11:08we're going to be covering exactly what

11:09you need to know if you look at the

11:11study guide in Microsoft learn these are

11:14the elements that we're going to be

11:15covering how do we identify requirements

11:17for a solution so the various components

11:20features performance capacity skus that

11:23kind of thing how do we make decisions

11:24about that we're also going to be

11:26looking at how to recommend settings in

11:27the fabric admin portal

11:29how do we choose data Gateway types and

11:32also creating custom powerbi report

11:35theme towards the end of the lesson

11:36we'll be testing your knowledge with

11:38five sample questions and just as a

11:41reminder all of the lesson notes and key

11:44points and link to further learning

11:46resources they're going to be published

11:47on the school community so if you're not

11:49already a member I'll leave a link in

11:50the description now you play the main

11:53character in a scenario and this

11:55scenario is going to walk you through

11:57everything you need to know for those

11:58four ele Els of the study guide are you

12:00ready let's begin so you are a

12:02consultant and you're starting your

12:04first day on a new project and this is

12:07Camila she is your client for the

12:10project and on the phone before the

12:11meeting Camila had mentioned that she

12:13wants to implement fabric but she

12:15doesn't know really where to start and

12:17that's where you come in you're going to

12:19start with a requirements Gathering

12:21Workshop so you organize a full day

12:23workshop with Camila the client to truly

12:26understand their business and their

12:27requirements now your goals for this

12:29Workshop are to extract a set of

12:32requirements from the client to help you

12:34build a plan for their new fabric

12:37environment and another goal is to do

12:40such a great job in planning their

12:42environment that the client is going to

12:43give you a new contract by the end of it

12:45to build the thing okay so this

12:47requirements Gathering Workshop what are

12:49you going to ask Camila what do you need

12:52to know when you're identifying the

12:53requirements you should think about

12:55focusing on these three elements to

12:57begin with the capacities so how many do

12:59we need what sizing do the capacities

13:01need to be in this new environment then

13:03we're going to look at data ingestion

13:05methods so there's lots of different

13:06ways that we can ingest data into fabric

13:09you're going to ask a set of questions

13:11that's going to kind of deduce the best

13:13method for getting data into fabric

13:15based on the requirements similarly

13:17we've got data storage so we've got

13:19three different options for storing data

13:21in fabric how do you ask the right

13:23questions and identify the requirements

13:25to choose the right one so let's start

13:27off thinking about capacity requ

13:29requirements now the requirements that

13:30we need here are really the number of

13:32capacities that are required and the

13:34sizing so the SKU the stock keeping

13:37units you probably know by now that in

13:39fabric we have capacities of varying

13:42sizes so what determines the number of

13:44capacities required so from previous

13:46videos you've probably understood that

13:48one of the things that impacts the

13:50number of capacities required is

13:52compliance with data residency

13:53regulations so the capacity dictates

13:56where your data is stored so if you have

13:59regulations that dictate that your data

14:02must reside in the EU for example for

14:04gdpr that's going to be one capacity in

14:06your business if you have other

14:08requirements that say these data sets

14:09need to be stored in the US you're going

14:12to have to have a separate capacity for

14:13that as well another thing that can

14:15impact the number of capacities is the

14:17billing preference so the capacity is

14:20how you get build in fabric so some

14:22organizations might want to separate the

14:24billing between different departments in

14:27their organization so they might have

14:28one capacity for the finance department

14:31one capacity for your Consulting

14:32division one capacity for your marketing

14:34department for example another thing

14:36that could determine the number of

14:37capacities that you need is segregating

14:40by workload type so if you have a lot of

14:42heavy intensive data engineering

14:44workloads then you might want to put

14:46those in a separate capacity and give it

14:48enough resource to allow you to do that

14:51in a confined capacity then you're

14:53serving of business intelligence you

14:55might want to do that in a separate

14:56capacity so that the read performance on

14:58those kind of dashboards is not impacted

15:00by the heavy data engineering stuff

15:02maybe machine learning stuff that's

15:04being done in other capacities you might

15:06also want to segregate by department

15:08just through business preference as well

15:09aligned with that billing preference so

15:11some companies like to have their

15:13capacity aligned to various dep

15:14departments within their business so

15:16these are the things you need to extract

15:18in terms of requirements when you're

15:19talking with this client and what about

15:21the sizing well we've touched on that

15:23already but some of the things that

15:24impacts the sizing of a capacity are the

15:27intensity of the expected workloads so

15:30are you going to be doing High volumes

15:31of data ingestion are you going to be

15:33getting gigabytes of fresh data into

15:36Fabric or even terabytes of data into

15:38fabric every day these are going to use

15:40a lot of your resources and to go

15:42through them quickly it helps if you

15:44have a higher capacity similarly heavy

15:46data transformation so if you're doing a

15:48lot of heavy transformations in spark

15:51that's going to use a lot of resources

15:52so if that's something you're going to

15:53be doing regularly in your business you

15:56want to be choosing a high capacity for

15:57that again machine learning training can

16:00you be very resource intensive going to

16:02take hours or sometimes even days to

16:04train a machine learning model if that's

16:06something you're going to doing

16:07regularly you want to be having that on

16:08a high capacity the budget of your

16:11client also dictates the capacity the

16:14sizing of the capacity that you're going

16:16to choose obviously the more resources

16:18the higher that SKU that you decide the

16:21more expensive it's going to be and some

16:23clients might be very sensitive around

16:25the cost and related to that is can you

16:27afford to wait or can the client afford

16:29to wait because if you procure a F2 skew

16:33it's probably going to go through your

16:34data but it might take a very long time

16:37and in some business that might not be a

16:38problem maybe you're just doing data

16:39ingestion once per day you ingest all of

16:42your fresh data and it might take a lot

16:45longer on an F2 capacity but that's not

16:47necessarily a problem maybe you can do

16:48it overnight and by the time people come

16:50in in the morning all of that data has

16:52been ingested or transformed and it's

16:54ready for consumption in the morning so

16:56what's your propensity to wait now some

16:59other companies might have regular data

17:01coming in every hour like gigabytes of

17:04data every hour and in that scenario you

17:06really need a high capacity to be able

17:08to churn through all of that stuff and

17:10get it processed before the next hourly

17:14load for example another thing that can

17:16determine the sizing of the capacity is

17:19does the client want access to f64

17:22Features so there's quite a lot actually

17:24of features that open up when you get to

17:27f64 so co-pilot being a good example

17:30currently and there's many many more

17:32I'll list them on the screen here these

17:34are features that only really are

17:35available if you choose f64 capacity or

17:39above so that's something to bear in

17:40mind if you want to use any of these

17:42features you need an f64 plus so what

17:44about the data ingestion requirements

17:46well here what we really need to know is

17:49what are the fabric items and or

17:51features that you need to get data into

17:53Fabric and how are you going to

17:54configure these items once you've built

17:56them now some of the options here and

17:58this is not an exhaustive list there's

18:00lots of different options here we've got

18:02the shortcut database mirroring ETL via

18:05data flow ETL via data Pipeline and a

18:07notebook and the event stream so these

18:09are some of the options that you might

18:11want to consider so what are the

18:13questions that you need to ask of a

18:15client when you're identifying the

18:17requirements to help you make the

18:19decision here well these are some of the

18:20deciding factors the main one really is

18:22where is the external data stored if

18:25it's in ADLs Gen 2 Amazon S3 or S3

18:29compatible storage location like Cloud

18:31flare for example Google Cloud Storage

18:33or the data verse but then these are the

18:35ones that are going to be available for

18:37you to shortcut into fabric so if you

18:40get any questions in the exam around you

18:42know my data rest stored in ADLs Gen 2

18:45well obviously the shortcut is a good

18:47option for that now it's not necessarily

18:48the only option you can still do ETL via

18:52any of these storage locations but it

18:53does open up that shortcut possibility

18:56now if you see Azure SQL Azure Cosmos DB

18:59or snowflake mentioned then immediately

19:01you should start thinking okay this

19:02could be database mirrored so you can

19:05use database mirroring to create that

19:07kind of live link to the database and

19:09it's going to maintain a mirror inside a

19:12fabric is it on premises now if you're

19:14data stored on premises then you're

19:17going to be probably want to be using

19:18the ETL via data flows or data pipelines

19:22because these two activities these two

19:24items allow you to create that

19:26on-premise data Gateway on your on

19:29premise server and then connect to that

19:30via the data flow or the data Pipeline

19:32and if you got realtime events realtime

19:34streaming data obviously you probably

19:36want to using the event stream to get

19:38that data into fabric anything else

19:40really you're going to be looking at ETL

19:42by either the data flow the data

19:44Pipeline and the notebook and when to

19:46choose which one well I've done a very

19:48long video I'll leave a link in the

19:49description or you can click here to

19:51make that decision about which of these

19:53is best for that particular organization

19:56so related to that is also what skills

19:58exist in the team because you don't want

20:00to build a solution that can't be

20:02maintained managed by the company or

20:04your company or your client's company so

20:06if you're looking for a predominantly no

20:08and low code experience then you're

20:10going to want to be focusing on the ETL

20:13via data flows and data pipelines both

20:15of these are fairly low and no code

20:18experiences help you get data into

20:19fabric if You' got a lot of SQL

20:21experience in your team then here you

20:23can be using the data pipeline you can

20:25use the script activities to do

20:27Transformations on your data as is

20:28coming in and if you have people that

20:30are familiar with spark python Scala

20:33that kind of thing then you can use the

20:36ETL notebook if you're you know perhaps

20:39you've got data coming from a rest API

20:41and you want to be using python

20:42libraries to get that in that's a good

20:44option for you there so whil we're on

20:46the topic of data ingestion there's a

20:48few other features you need to be aware

20:50of that might come up in the exam that

20:53can help you identify different

20:54requirements for getting data into

20:55fabric these are the on premise daily

20:58Gateway which we've mentioned the v-net

21:00the virtual network data Gateway fast

21:02copy and staging so you might be asking

21:05some questions about these things in the

21:06exam as well so when do we decide on

21:09these sorts of things well you need to

21:11ask how the data in the external system

21:13is being secured right so if it's on

21:16premise if it's an on premise SQL Server

21:18you have to be using the on premise data

21:20Gateway if your data is living in Azure

21:23behind some sort of virtual Network or

21:25private endpoint that kind of thing then

21:28you want to be setting up the v-net data

21:30gateway to access that and in terms of

21:31the volume of data this is also going to

21:33have an impact on the items that you

21:36choose for doing your data ingestion and

21:38also some of the features available so

21:40if you've got low or medium data per day

21:43well if it's low then you probably don't

21:44need any of these specific features like

21:46the out of the box Solutions will be

21:47good enough but if you've got quite a

21:49lot of data gigabytes per day in that

21:52kind of range you want to be using some

21:54of the features like Fast copy and

21:56staging similarly if You' got very high

21:58amounts of data these are going to be

22:00one of using the fast copy and the

22:01staging if you're using data flows

22:03alternatively you can use data pipelines

22:06and if you can get data in bya a fabric

22:08notebook then that's another option as

22:10well so before we move on I just want to

22:12mention a bit more detail around the

22:14data gateways now as you probably know

22:17already there are two types of data

22:19gateways that we can configure in

22:21Microsoft fabric number one is the

22:23on-premise data Gateway and number two

22:25is the virtual network data Gateway and

22:28a data Gateway in essence helps us

22:30access data that's otherwise secured so

22:34if his data is on an on-premise SQL

22:36server for example it gives us a secure

22:38way to access that data and bring it

22:40into fabric likewise if you've got data

22:43behind a virtual Network secured in

22:45Azure in like blob storage or ADLs Gen 2

22:49it provides us with a secure mechanism

22:51to access that data so I'm not going to

22:53show you step by- step how to set up a

22:55data Gateway in this lesson but what it

22:58done is linked to two other videos by

23:00other creators that show the process in

23:02detail if you want to go and have a look

23:04I'll leave that in the school Community

23:05but I do want to just cover kind of the

23:07high level process for each of them just

23:09so you understand a bit more about what

23:11that looks like if you've never set one

23:13up before so for the on- premise data

23:15Gateway there's a few high level steps

23:18number one we need to install the data

23:19Gateway on the on premis server and if

23:22you've already got an on- premise data

23:23Gateway set up on your on premise server

23:26perhaps you're using it in traditional

23:28powerbi data flows for example then

23:30you're going to need to update it to the

23:32latest version cuz that's going to be

23:33compatible with Microsoft fabric the

23:36next step is to in fabric create a new

23:39on premise data Gateway connection and

23:41then from that you can connect to that

23:43data Gateway from either a data flow and

23:46now also a data pipeline so the data

23:48pipeline was recently added in the last

23:49few weeks I think it's still in preview

23:52that connection so you might not get

23:54asked about it in the exam but it's good

23:55to know that now it's actually possible

23:57via the data flow and the data pipeline

23:59to set up the v-net data Gateway we're

24:01going to start in Azure there's a few

24:03settings that you need to configure in

24:05your Azure environment before you can

24:07set up the v-net data Gateway connection

24:10so you're going to need to register a

24:12Power Platform resource provider within

24:14your Azure subscription and then within

24:16the item that you want to share or you

24:18want to access for example in your Azure

24:21blob storage item in Azure you need to

24:23create a private endpoint in the

24:24networking settings then create a subnet

24:27and then we're going to use that in

24:28fabric to create a new virtual network

24:31data Gateway connection and then again

24:33from that you can connect to it via your

24:34data flow to be able to access that data

24:37that is behind that virtual Network in

24:40Azure so next let's look at the data

24:42storage requirements and when we're

24:44talking to our client here identifying

24:46requirements really what we're trying to

24:48extract is okay what fabric data stores

24:52are going to be best for these

24:54requirements and what overall

24:56architectural pattern are we going to be

24:58aiming for with this solution now the

25:00options here are obviously The Lakehouse

25:02the data warehouse and the kql database

25:05and some of the deciding factors to

25:07choose between these are well what's the

25:09data type okay so is it structured or

25:12semi structured or even unstructured so

25:15are you going to be getting raw files

25:18CSV Json maybe from AR rest API is it

25:21unstructured is it image data video is

25:24it audio data for example these are all

25:27going to be wanted to store in The

25:28Lakehouse because this is kind of the

25:30only place in fabric where you can store

25:32a variety of different file formats if

25:35your data is relational and structured

25:37then obviously you can keep that in

25:39either the lake house or the data

25:41warehouse and if it's real time and

25:42streaming you're going to be want to

25:43streaming that into your kql database

25:46next up another important consideration

25:48when choosing a data store is what

25:51skills exist in the team so if you're

25:53predominantly tsql based then you're

25:55going to be want to be using the data

25:57warehouse experience

25:59if you're predominantly spark and python

26:01Scara that kind of thing then you're

26:03going to be wanting to storing your data

26:04predominantly in the lakeh house and if

26:06you're predominantly using kql in your

26:08organization that's going to be one to

26:10using kql database for your data storage

26:14congratulations you've completed your

26:16first engagement for Camila you've

26:18convinced her to set up a proof of

26:20concept project in her organization so

26:23she's already created the fabric free

26:25trial she's set up her environment but

26:28immediately she's hit a bit of a hurdle

26:30so this is your next mission she Rings

26:32you up and she says hey I need some help

26:34I open the fabric admin portal and

26:37nearly had a heart attack please can you

26:39help me understand all of these settings

26:41so you set up a call with Camila to help

26:43her understand the fabric admin portal

26:46how are you going to teach her and what

26:47are you going to teach her about the

26:49admin portal what are the most important

26:51settings that she needs to know about

26:52okay just before we get into the fabric

26:54admin portal and look at some of the

26:56settings available to us in there there

26:58it's important to note that to be able

27:00to access the admin portal of course

27:02first you need a fabric license but then

27:04you need to have one of the following

27:07roles you need to be either a global

27:09administrator a Power Platform

27:11administrator or a fabric administrator

27:14So within the admin portal here in

27:17fabric you'll see this menu on the left

27:20hand side so these are some of the

27:21important settings in tenant settings

27:24here you can allow users to create

27:25fabric items so if you just set up

27:27fabric in your ization you need to allow

27:29people to actually create fabric items

27:31without that you can't really get very

27:32far enable preview features so every

27:35time Microsoft release new features

27:37normally they put them in the admin

27:39portal and you can allow or disallow

27:42users in your organization to use them

27:44you can also allow users to create

27:46workspaces there's a whole host of

27:48security related features that you can

27:51manage and get control over in your

27:53tenant so for example how do you manage

27:55guest users allowing single sign on

27:57options for things like snowflake big

27:59query red shift accounts that kind of

28:01thing how do you block public internet

28:03access so that's really important to

28:05know enabling other features like Azure

28:08private link for example allowing the

28:10service principal access to the fabric

28:12apis so if you're going to be doing some

28:13automation you need to allow access to

28:17service principles to the API there's

28:19also options in there for allowing git

28:21integration so if you're setting up

28:22Version Control that needs to be enabled

28:25there and there's also some features

28:27like allowing cop pilot within the

28:28organization as well now in general some

28:31of the settings can be one of three

28:33things it could be enabled for the

28:35entire organization it can be enabled

28:37for specific security groups so say you

28:39only want super users to be able to use

28:43this feature or admins within your

28:45fabric environment to use a specific

28:47feature then you can enable it for

28:48specific security groups or you can

28:50enable it for all except certain

28:52security groups so everyone in your

28:54organization gets access apart from

28:57these people perhaps guest users is a

29:00good example now other settings in the

29:02fabric tenant settings are kind of

29:04binary you either enable them or you

29:06disable them for the entire organization

29:08another important point in the fabric

29:10admin portal are the capacity settings

29:13so this section here and in here you can

29:16create new capacities delete capacities

29:18manage the capacity commissions and also

29:21change the size of a capacity so these

29:23are some important capacity settings

29:25that you need to be aware of understand

29:27how they work and how to manage them

29:29within your fabric environment so great

29:31you've talk Camila about the fabric

29:33admin portal and she's very grateful but

29:36before the meeting ends she has one more

29:38thing she wants to ask you about she

29:40says one final thing before you go and

29:42it might seem a bit random but when we

29:44migrate to fabric I want our bi team to

29:47create more consistent reports have you

29:50got any ideas about how we can achieve

29:51that and of course the first thing you

29:53think of are custom powerbi report

29:55themes now there are many ways to create

29:58a custom report theme in powerbi you can

30:01either update the current theme if

30:03you're in powerbi desktop or you can

30:05write a kind of Json template yourself

30:07using the documentation if you're

30:08feeling a bit Brave you can do that

30:10yourself or you can also use a third

30:12party online tool there's quite a few

30:14report theme generator tools that exist

30:16online but it's unlikely you're going to

30:18be tested on that in the exam so your

30:20task is to show Cilla how to create a

30:22custom report theme so let's have a look

30:25at how you can do that within powerbi

30:27desktop top so here I've got a report

30:30and what I'm going to do to access the

30:31report themes you need to go to the view

30:33tab then you can see these themes here

30:35and obviously these are the preset

30:37themes so you can just click and update

30:39the current theme very simply like that

30:41but to do most of the customization you

30:43need to click on this button here and

30:45you can access the current theme and all

30:47accessible themes are currently

30:49installed on this machine then there's a

30:51few settings down here that quite

30:52important to know so browse for themes

30:55if you click on that it's going to allow

30:56you to import a powerbi report theme so

31:00if you've already got a theme Here For

31:01example this one here then you can

31:03select that and install it into your

31:05environment like so if you want to

31:07customize the current theme you can do

31:10that like this and it's going to bring

31:11you through to this UI environment just

31:13to you know change some colors change

31:15some text change some visuals what you

31:18have to think about for this section of

31:20the exam is what could they ask you you

31:22have to think about how could you

31:23possibly be tested on this so in terms

31:26of the powerbi report theme stuff you're

31:28likely to be tested on these buttons

31:30here and what they do plus they could

31:32ask you about a Json theme so they could

31:35show you a Json theme and maybe ask you

31:38about okay how can you edit this theme

31:40what doesn't look right in this theme

31:42that kind of thing so it's good to have

31:44a bit of familiarity about the different

31:47sections in these Json files so the name

31:50of it how you can store your data colors

31:52as a list some of these different

31:54settings here you're probably not

31:55expected to memorize all of the

31:57different settings in Json format but

31:59you might get shown a theme in Json

32:01format and asked to modify it or asked

32:04to comment on it in some way to export

32:06the current theme you can also use this

32:08save current theme and that's going to

32:09allow you to export a Json file that you

32:11can share within your organization and

32:13you've also got here access to the theme

32:15Gallery so this is going to bring you

32:17through to the theme Gallery website

32:19where you can download other people's

32:21themes for your report to finish up this

32:23video and this lesson we're going to go

32:26through five practice questions just to

32:28kind of solidify that knowledge make

32:29sure you're understanding some of the

32:31key Concepts within a context of a

32:33scenario so the first question is you're

32:36running an F2 capacity and you regularly

32:39experience throttling with that capacity

32:41now there's a number of long running

32:43spark jobs that take on average 3 hours

32:45to complete and you need these to

32:47complete in under 1 hour so you plan to

32:49increase the SKU of the capacity where

32:52would you go to make this change would

32:53you go to the workspace settings and

32:55configure spark settings would you go to

32:57to the admin portal and the capacity

32:59settings section and then click through

33:01to Azure to update your capacity would

33:03you go to the monitoring Hub and look at

33:05the Run history or would you use the

33:07capacity metrics app so pause the video

33:09here and have a bit of a think and then

33:12I'll move on so the answer is 2b so you

33:15can manage your capacity settings within

33:16admin portal and then capacity settings

33:19and then you can actually click through

33:21to Azure it gives you a link to the

33:22Azure portal and that's where you're

33:24going to change the capacity within the

33:26Azure portal OB so you can't be within

33:28the spark settings that's for managing

33:30the configuration of your spark cluster

33:32within a workspace and in the monitoring

33:34Hub we can't get anything there to do

33:36with capacity settings that's just going

33:38to tell you how your jobs are running

33:39and in the capacity metrics app that's

33:41just a readon app for having a look at

33:44how your capacity is being used so that

33:46wouldn't also be suitable either

33:48question number two your data governance

33:50team would like to certify a semantic

33:52model to make it discoverable in your

33:55organization now only the data

33:56governance team should be able to do

33:58this in what order should you complete

34:00the following tasks to certify a

34:02semantic model so have a look at the

34:04five actions here and what you're going

34:07to have to do is put these in an order

34:10so these are this is an ordered list it

34:12should be so one of the things you'll

34:13have to do first second third fourth and

34:16fifth so once you've got these in order

34:18we'll move on so let's look at the

34:19answer now so the correct order looks a

34:21bit like this so we start by creating a

34:25security group for the data governance

34:27team the clue in the question was that

34:30only the data governance team should be

34:32able to do this so when you see that you

34:35think okay well they need to be within a

34:37security group to enable this number two

34:40and you could argue that one and two

34:41could be interchangeable but these are

34:43the first two items anyway but enable

34:45the make certified content discoverable

34:48So within the admin portal the tenant

34:51settings as a section for Discovery

34:54you're going to need to enable that for

34:55the organization and then after that

34:57going to have to make sure that that

34:59settings is applied only to the data

35:01governance security group that you set

35:03up then you're going to need to ask the

35:05data governance team to go into the

35:08semantic model settings and then

35:09endorsement and Discovery and click

35:12certify for that semantic model and then

35:14you want to validate that that has been

35:16set up correctly and your business user

35:18can see that certified semantic model

35:22within the one leg data Hub three you

35:24join a new company and you're given a

35:26powerbi report theme as a Json file to

35:28use for all new projects how do you

35:30apply this Json file theme to the report

35:33that you're currently developing is it a

35:35in powerb desktop go to view themes and

35:38customize current theme B go to the

35:40fabric admin portal click on custom

35:43branding and then set the default report

35:45theme C use tabulate editor 2 to update

35:48the theme or D in power by desktop go to

35:51the view themes and then browse for

35:53themes so the answer here is D in power

35:57desktop go to the view themes and then

35:59browse for themes so d and a are quite

36:02similar but a is for customizing a

36:05current theme so that's not going to be

36:06allowing you to import a Json file

36:09that's going to allow you to use the

36:11user interface to update the current

36:13theme so that's not what we want to do

36:15we want to import ad Json file as our

36:17report theme which is possible using D B

36:20that functionality doesn't actually

36:21exist custom branding does exist but

36:23that allows you just to update the

36:25colors and the icons within fabric not a

36:28default report theme and tabul editor 2

36:31is also the incorrect answer question

36:33four you have 1,000 Json files stored in

36:36Azure data Lake storage ADLs Gen 2 that

36:39you want to bring into fabric the ADLs

36:42Gen 2 storage account is secured using a

36:45virtual Network which of these actions

36:48would you need to perform first is it a

36:51in fabric go to manage connections and

36:53gateways and then click on create a new

36:56virtual network data Gateway B create a

36:58shortcut to the ADLs Gen 2 storage

37:01account C in Azure register a new

37:03resource provider and create a private

37:05endpoint and subnet or D install an on

37:08premise data Gateway on an Azure virtual

37:10machine in the same virtual Network or E

37:13enable Public Access in the storage

37:15account network settings so for this one

37:17you'll remember that the answer is C so

37:20the first step in setting up a virtual

37:22network data gateways well we need to go

37:24into Azure we need to perform some

37:26network configuration okay so you need

37:28to register that new resource provider

37:31that Microsoft Power Platform resource

37:33provider within your subscription and

37:35then on the item create a private

37:36endpoint and a subnet all of the other

37:38options some of them are steps in the

37:41process but not the first step so the

37:43question was which of these actions

37:45would you need to perform first so yes

37:48we do need to do a but it's not going to

37:50be the first thing that you're going to

37:51do B is kind of a bit of a a red herring

37:54here cuz you might have seen ads Gen 2

37:56and thought ah shortcut but actually you

37:58need to configure the virtual Network

38:00daily Gateway before you can even think

38:03about kind of connecting to it D

38:04installing the on- premise data Gateway

38:07well you know that we're looking at a

38:08virtual Network here so you're going to

38:10be choosing the virtual network data

38:12Gateway rather than an on- premise Ste

38:13goway and E enable Public Access in the

38:16storage account network settings while

38:18that's going to expose your data to the

38:20public internet so not advisable

38:23question five you have data stored in

38:25tables in Snowflake which of the

38:27following cannot be used to bring the

38:29data into fabric a use the data pipeline

38:32copy data activity B create a shortcut

38:34to the snowflake tables from your Lake

38:37housee C use the data flow Gen 2 with

38:40the snowflake connector D use database

38:42mirroring to create a mirrored snowflake

38:45database in fabric so the answer here is

38:47B to create a shortcut to the snowflake

38:50tables from your lake house as you'll

38:52know you can only shortcut to ADLs Gen 2

38:55or Amazon S3 or Google Cloud Storage so

38:59the ability to shortcut is generally on

39:02files when we're talking about tables in

39:05databases whether there all of the other

39:07three we can use so you can do a copy

39:09data activity from a data pipeline to

39:11bring that data in if you want to copy

39:13it in or you can use a data flow Gen 2

39:16or you can use database mirroring

39:18because snowflake is one of the

39:20databases where database mirroring is

39:22possible Camila says thanks she's

39:24seriously impressed with your knowledge

39:26well done in this lesson we've looked at

39:28how you can identify requirements for a

39:30fabric solution we've looked at the

39:32different types of data gateways that

39:34are available to us in fabric we've

39:36looked at the settings in the admin

39:38portal and we've also looked at how to

39:40create custom powerbi report themes and

39:43the good news is you want an extension

39:44to the contract Camila would like you to

39:47implement and manage her data analytics

39:49environment so you've got the next stage

39:51of the contract in the next lesson we'll

39:53look at how you can do that how you can

39:55set up Access Control sensitivity

39:58labeling workspaces capacities all that

40:01kind of stuff how do we set these things

40:03up inside fabric so make sure you click

40:05here for the next lesson hey everyone

Implement and manage a data analytics environment

40:07welcome back to the channel today we're

40:09going to be continuing our dp600 series

40:12and we're going to be looking at

40:13implementing and managing a data

40:16analytics environment this is the second

40:18part of the dp600 syllabus or the study

40:21guide that we're going through on the

40:23channel and today we're going to be

40:24going through these particular elements

40:26and these are coming straight from the

40:28study guide so number one we're going to

40:30be looking at implementing workspace and

40:32item level Access Control implementing

40:34data sharing for workspaces warehouses

40:37and lake houses managing sensitivity

40:39labels configuring fabric enabled

40:42workspace settings and managing fabric

40:45capacity we've got five sample questions

40:47so at the end of the video we'll be

40:48going through some sample questions to

40:50test your knowledge and as with the last

40:52video I'll be posting the key points and

40:55links to further resources in the school

40:57Community available for free I'll leave

40:59a link in the description for that we'll

41:01continue the scenario that we were

41:03developing in the last lesson you're the

41:05main character again are you ready let's

41:07begin so just to recap you are a

41:09consultant and you're working with your

41:10client who is called Camila and in the

41:12last engagement you successfully planned

41:15their data analytics environment now

41:17you've won the contract to support

41:18Camila in implementing that solution so

41:21Camila is busy doing her resource

41:24planning she's thinking about who she's

41:26going to need to support support this

41:27environment in fabric she's asking for

41:30your assistance to help her structure

41:32her thoughts and also her team so let's

41:34have a look at what that looks like in

41:36fabric so this is a high level structure

41:39of a fabric implementation now you

41:42notice that it's hierarchical at the top

41:45we have tenant level so this is kind of

41:47the one tenant that you're going to have

41:49in your organization then below that you

41:51might have one or many capacities and we

41:54talked about capacities in the previous

41:56lesson we're going to going to be doing

41:57a bit more on capacities in this lesson

41:59as well then in each capacity you might

42:01have one or multiple workspaces then in

42:04each workspace we go down to the item

42:07level so you might have a data warehouse

42:09and a Lakehouse in your workspace then

42:11we can actually go one level deeper than

42:13that which is called the object level So

42:15within the data warehouse you have dbo

42:17do customer that might be a table in

42:19your data warehouse or a view in your

42:21data warehouse that's at the object

42:23level now when we're administering

42:25fabric we need to be aware of these

42:27different levels because at each of

42:28these different levels Administration

42:31happens in a different way in the last

42:33lesson we looked at the tenant level

42:35admin settings so that's mostly things

42:37in the admin portal under the tenant

42:40settings section today we will explore

42:42item level a little bit later on but

42:45first I just want to look at how we can

42:47administer each of these three top

42:49levels so here we have a table now the

42:52table isn't complete yet we're going to

42:54walk through it together so we have the

42:55three top levels we got tenant level the

42:58capacity level and the workspace level

42:59and then on the right hand side we're

43:01going to go through what the

43:02administrator or who the administrator

43:04is what role they require and also where

43:07the admin happens right so where are

43:09they going to be working where are the

43:11settings that they need to administer at

43:13each level so starting at the top level

43:15we've got the tenant admin now to get

43:18the rights to be able to be a tenant

43:20admin we're actually going to go higher

43:21than fabric we need an entra ID role of

43:24global administrator Power Platform

43:26administrator or fabric administrator

43:28and we looked at what that looks like in

43:30the last lesson and they're going to be

43:32working predominantly in the fabric

43:34admin portal so if you have any of these

43:36three roles that's going to be available

43:38to you and you can configure your tenant

43:40settings in there one level down at the

43:43capacity level well the capacity admin

43:45is assigned when you create a new fabric

43:48capacity in Azure now where the

43:50administration happens so any sort of

43:53administration at that capacity level is

43:55going to be done either in the Azure

43:57portal as we mentioned before or in the

44:00fabric admin portal there's a section on

44:02capacity settings we're going to have a

44:04look at both in a minute at the

44:05workspace level so we're going one level

44:08down now you're going to have a

44:09workspace admin they're going to be the

44:11person that is kind of in charge of that

44:13admin level and the role required here

44:16is a workspace role so it's going to be

44:18a person or a group with the workspace

44:21role of admin and again we'll have a

44:23look at what that means in more detail a

44:26bit later on now there going to be doing

44:27most of their Administration within the

44:30workspace settings and also within the

44:32manage access so these are the two areas

44:34that they're going to be focusing most

44:36of their time on now we looked at the

44:38tenant level admin settings previously

44:40in the last video in this lesson we're

44:41going to focus on the capacity level

44:44settings and then the workspace level

44:46settings so let's just focus in on the

44:48capacity administrator settings for a

44:50moment and as I mentioned there's two

44:52really places where capacity

44:55Administration gets done number one one

44:57is in Azure because we need to use Azure

45:00portal to purchase capacity right so

45:04that's going to be where you go to

45:05create a new capacity delete existing

45:07capacities changing the size of a

45:10capacity so if you've got an F2 skew and

45:13you want to go up to an F4 that's going

45:15to be done in the admin portal and also

45:17changing the capacity administrator so

45:18you can do that within the Azure portal

45:20as well as well as that within fabric we

45:24can change some of the settings for a

45:26particular Capac

45:27so we can do things like enabling

45:29Disaster Recovery viewing the capacity

45:31usage report so how much is our capacity

45:34being used we can Define who can create

45:36workspaces within that capacity we can

45:38Define who is a capacity administrator

45:41we can update the powerbi connection

45:43settings so who can connect to this

45:46capacity or items within this capacity

45:49from powerbi and how does that look like

45:51we can permit workspace admins to size

45:54their own custom spark palls so this is

45:57quite important right because you might

45:58want to set some sort of limits on the

46:00sizing of the spark poles that workspace

46:03owners underneath the capacity or in

46:06this capacity you might want to limit

46:08how high they can go with their spark

46:09poles because that's going to have quite

46:10a big impact on the overall capacity

46:13usage so if you're sitting at the

46:14capacity level you might want to add

46:16some restrictions on you know the custom

46:18spark configurations that happen in Your

46:21Capacity and we can assign workspaces to

46:24the capacity in this section as well

46:26okay so so here we are in the portal and

46:28I just wanted to show you some of the

46:30capacity settings how to administer a

46:33fabric capacity and we're going to start

46:34right from the beginning so how do you

46:36actually set up and buy a capacity well

46:38you go to Microsoft fabric if it's not

46:40already there then you can just search

46:42for it here Microsoft fabric this is

46:43going to bring you through to the the

46:45resource creation tool we're going to

46:47click on Create and we're going to walk

46:49through these steps to create a fabric

46:50capacity so you need a subscription and

46:53a resource Group and you can enter the

46:55capacity name so give it a region change

46:58the size I'm just going to do an F2

47:01select and here is where you sign the

47:02capacity administrator and again we can

47:04change that afterwards but that's you

47:06need at least one to set it up so it's

47:08going to give you the estimated cost per

47:10month here and then you press create and

47:12it's going to deploy that capacity okay

47:15so now my capacity has been created we

47:18can click on go to Resource and this is

47:20where we're going to do some of the

47:21administration tasks within the Azure

47:24portal right so here we can have a look

47:26at well firstly we can pause it so if

47:28you want to pause the capacity you can

47:30do that here delete it as well down the

47:32left hand side we've got some useful

47:34things here so capacity administrators

47:36that's where you're going to change your

47:38capacity administrator we've also got

47:40change size so if you're finding that F2

47:44is not enough for your workloads that

47:47you're running you can change it to F4

47:49or f256 if you've got a spare 40 Grand a

47:52month to be using on fabric I'm going to

47:55keep it as an F2 that's just something

47:57to bear in mind there okay now I just

47:58want to have a look at what that

47:59capacity setting looks like inside

48:02fabric so if we go to the admin portal

48:05and capacity settings if we go over to

48:07the fabric capacity tab here we can see

48:10that fabric capacity that we've just set

48:11up so it's an F2 it's in UK South and

48:14it's active so we've got some actions

48:16here we can change the name we can have

48:17a look at the admin you can't really do

48:19much there there's a link through back

48:21into Azure so if you want to make any

48:23changes to it from within fabric you

48:25need to click on this link here but if

48:26you click on the actual capacity name we

48:28go through to the capacity settings

48:31right so this is where you're going to

48:32be doing things like enabling Disaster

48:35Recovery having a look at the usage

48:37report for that capacity turning on

48:39notifications so it's going to give you

48:41a notification when you've used x amount

48:44of percent of your capacity updating who

48:46is the administrator for that capacity

48:49can also be done here changing how we

48:51can access powerbi and how powerbi can

48:54access data in the capacity we've also

48:57got data engineering settings it's

48:59mainly spark settings and this is where

49:02we're going to have that permission to

49:04permit people to change the custom spark

49:07poing right so either on or off and we

49:10can assign certain workspaces to this

49:12capacity so if you just create a new

49:13capacity it's going to be empty and you

49:15can move existing workspaces onto that

49:18capacity so as a workspace administrator

49:20we've got a few different options so

49:22here we're stepping down level into the

49:24workspace level and a workspace

49:26administrator as we mentioned deals

49:28primarily in the workspace settings and

49:31here you can edit the license for the

49:33workspace so change it from like for

49:36example a trial capacity to a fabric

49:38capacity or from Pro to PPU premium per

49:41user we can also configure connections

49:43to Azure as well as well as configuring

49:46Azure devops connections so if you want

49:48to use Git Version Control for this

49:51workspace that's going to be done in the

49:53workspace settings another thing we can

49:54do in the workspace settings is set up

49:57what's called a workspace identity now

49:59this is basically having a managed

50:02service principle dedicated for your

50:04workspace and it basically means you can

50:06connect to things like ADLs Gen 2 for

50:09things like shortcuts and you can do

50:11that in a kind of trusted workspace

50:13access manner basically this is another

50:15pretty new security feature that they've

50:17added quite recently and I'll be going

50:18through more of the security principles

50:20in more detail probably in a separate

50:22video at the end of this series CU

50:24they're quite important we can also edit

50:26some of the power settings in the

50:27workspace settings and also the spark

50:30settings so particular default

50:32environments that we might want to set

50:34up within this workspace things like

50:36that just note here that managing access

50:38is done through the the managing access

50:41section so it's slightly set it's in the

50:43same kind of area but it's not

50:44necessarily in workspace settings where

50:46we add users and add groups into our

50:49workspace we'll have a look at that in a

50:51bit more detail okay next I just wanted

50:52to walk through some of the workspace

50:55settings in a bit more detail so here we

50:57are in a workspace it's called Share Hub

50:59what we're going to be doing is clicking

51:01on this Dot and you can see that we've

51:03got two here that are useful for

51:05workspace administrators number one is

51:07managing access so this is how we're

51:08going to give people access to our

51:10workspace either person or a group we

51:13can add people in here and we can give

51:15them admin member contributor or viewer

51:17if we click on these dots again and then

51:19go through to the workspace settings

51:21this is where we're going to be able to

51:23edit some of the settings for our

51:24workspace General is you just do the

51:27image and the description and also

51:28domains if you're using domains license

51:30info so this is where you're going to

51:32change the potential capacity and the

51:34license that's being used in that

51:36workspace so this one is a trial

51:37workspace so maybe we want to actually

51:39change that to an F2 maybe we've got

51:41some fabric capacity we can select which

51:44one we want to use I'm going to be using

51:46this fabric F2 learn capacity and that's

51:48going to change the license for the

51:50workspace we've also got connecting to

51:53Azure connecting to git downloading

51:55things like the file explor

51:57as well and enabling caching for

51:59shortcuts is another workspace setting

52:01we've got here managed identities so if

52:03you're on an f64 capacity or higher you

52:06can make use of workspace identities and

52:09that's going to basically allow you to

52:10create kind of like a managed service

52:12principle just for this workspace so

52:14give your workspace an identity and

52:16allow it to connect to ADLs Gen 2 create

52:20your shortcuts things like that in a

52:22secure manner kind of trusted workspace

52:24access is what it's called and if you

52:26want to learn more more about that I'll

52:27leave a link to the workspace identity

52:29section in the school Community we can

52:31also do things like adding private

52:33endpoints and that kind of thing for

52:34connecting via spark to things in Azure

52:37you've got your spark settings down here

52:40for configuring something about the pool

52:43that you're using the spark pool that is

52:45being used in this workspace you can

52:47change the default environment that's

52:49where you're going to be going to add

52:51libraries and things like that so if you

52:53want to pre-install python packages onto

52:55your spark cluster so that every time

52:58you run a notebook or start a new

53:00notebook you have those libraries there

53:02ready to go that's where you do this

53:03change some settings for high con

53:05currency and that kind of thing there so

53:07Camila says okay great I now have some

53:09clarity on administering fabric at the

53:12tenant capacity in the workspace level

53:14what I'm not sure about is giving access

53:16to the people on my team so that's

53:18important right we build all these

53:20things in fabric but how can we give

53:22people access the right amount of access

53:25to these items that's a look at that in

53:27a bit more detail so if we go back to

53:28our structure of a fabric implementation

53:32generally when we're sharing items with

53:35people that's really done at these

53:37bottom three levels of our fabric

53:40architecture right so sharing things in

53:43fabric is normally done at these three

53:45levels now object level sharing is

53:48possible for the data warehouse and the

53:50seal endpoint in the lakeh house but I

53:52don't think it's assessed as part of

53:53this exam so it's not part of the study

53:56guide in in anyway so we're not going to

53:57be covering that in this lesson there's

53:59some documentation on the Microsoft

54:01learn website and I'll leave a link to

54:02that if you're interested in object

54:04level sharing if you want to bit learn a

54:06bit more about that we're going to be

54:08focusing on the workspace level sharing

54:10and item level sharing so let's just

54:12start with workspace level sharing

54:15people or groups can be given workspace

54:18level access and when sharing the

54:21personal group is assigned a workspace

54:23role as you can see on the right hand

54:25side there we've got admin member

54:27contributor and viewer now this role

54:30applies to all items in the workspace

54:33for example a viewer in the workspace

54:35will be able to view all of the items in

54:37the workspace let's just take a bit of a

54:39moment cuz roles are really important

54:42and the role that you assign someone

54:44dictates basically what they can do in

54:46your workspace this image here comes

54:48from the Microsoft documentation again

54:51I'll leave a link to this in the school

54:53community and I definitely recommend you

54:55take some time to study it this is what

54:57we're going to be doing here so the

54:58first thing to note with this diagram is

55:01let's just start with the admin so if

55:03you give someone an admin permission

55:05what can they do the first thing is that

55:06they can update and delete the workspace

55:08so this is really high level permissions

55:11that only maybe one or two people really

55:13should have people that you trust in

55:14your organization they can also add and

55:17remove people including other admins so

55:20it's the only role that allows you to

55:22add an admin add another admin next we

55:24move down to the member and the member

55:27can do similar things to an admin but

55:30they can't add an admin okay so a member

55:33cannot add an admin they can only add

55:35people with lower permissions or other

55:38members okay the other permission that

55:40is unique at the member level is you can

55:43give other people the permission to

55:45share items so being able to share items

55:48is a fairly high level thing to do and

55:50you're giving people that permission to

55:51share okay so that's something to bear

55:54in mind as well then we move down to the

55:56contributor level now contributors can

55:59do pretty much everything in the

56:01workspace other than as we see here

56:03deleting the workspace adding other

56:06people into the workspace and allowing

56:08other people to share but they can do

56:09when you're talking about contributing

56:12to fabric items anything around lake

56:15houses or warehouses or data pipelines

56:18they have read and write access to all

56:20of these things so the viewer has a

56:22unique set of missions those six green

56:25ticks

56:26and if we go kind of from top to bottom

56:28they can view and read content in a data

56:31pipeline a notebook spark job definition

56:33machine learning model so they can view

56:35kind of the outputs of these things they

56:37can also View and read the content of

56:39kql databases query sets and real time

56:42dashboards they can connect to the SQL

56:44analytics endpoint of a Lakehouse or a

56:47data warehouse and they can read

56:49Lakehouse and data warehouse data and

56:51shortcuts with tsql so the viewer can

56:54basically use SQL to analyze data in

56:57either The Lakehouse or the data

57:00warehouse what they can't do is access

57:03any of the one lake apis or spark so

57:06they can't run spark jobs or notebooks

57:10or anything like that now one unique

57:11thing about the viewer permission is in

57:14the data pipelines now they can't edit

57:16or update any of the activities in a

57:19data pipeline but they can execute and

57:23cancel the execution of a data pipeline

57:25run so that's an important kind of edge

57:28case to remember for the exam and they

57:30can also finally view the output of data

57:33pipelines notebooks and machine learning

57:35models so that's kind of a high level

57:37overview of all of these workspace roles

57:40and what they can do at each level again

57:42this is really important to understand

57:44for the exam so I definitely recommend

57:46going into the documentation taking some

57:48time to understand these different

57:49things because you'll probably be tested

57:51quite a lot on these so let's just have

57:52a bit of a workspace level access

57:55example this is John he is a business

57:58analyst working in Camila's team and

58:01Camila has asked you to give him

58:03contributor access to workspace one this

58:06is workspace one this is the

58:08architecture that they've got here now

58:09this is what John's access currently

58:11looks like where the red box is

58:13basically no access at all and a green

58:15box if there's any access and you can

58:17see everything's red So currently has no

58:18access to anything you are an admin in

58:21the workspace now what steps would you

58:23take to give John this access have a

58:25look little think about that and then

58:27we'll talk about it in a second okay so

58:29what steps would you take well

58:31personally what I would be asking is

58:33does John fit into an existing security

58:36group that has contributed access to the

58:39workspace because best practice here is

58:41to add people into groups rather than

58:43adding them individually just makes

58:44maintenance in the future a lot easier

58:47where possible we always want to add

58:49people into groups before we add them

58:51individually now if a security group

58:54doesn't exist then you might want to

58:55create one for John maybe you want to

58:57create an analyst security group so that

58:59in the future when another analyst wants

59:01to join the team or join the workpace

59:04you can just add that person into the

59:06group rather than having this long list

59:09of individual contributors in that

59:11workspace so you create an analyst

59:12Security Group add John to the security

59:15group and give the group contributor

59:17access to workspace one so in the

59:19picture how does that update well it

59:21looks a bit like this right so John now

59:23has access to workspace one and

59:26everything within it because we've given

59:28him access the security group access at

59:31that workspace level you'll notice that

59:33workspace 2 he still has no access to

59:36that he can't even see that so that's

59:37something to bear in mind when you're

59:39giving workspace level access Camila

59:42suddenly Rings you she realizes that JN

59:44shouldn't have access to everything in

59:46the workspace instead she wants you to

59:49give him access to the data warehouse

59:51only not the semantic model not the data

59:54pipeline so how would you check change

59:56what we've just done to reflect this so

59:59this is what we're we're looking at here

1:00:00we want to go from this which is the

1:00:02arrangement that we've just done for

1:00:04John at the workspace level to this at

1:00:07the item level now this might be

1:00:08important because it kind of reflects

1:00:11quite an important principle when it

1:00:13comes to giving people access which is

1:00:15the principle of least privilege now in

1:00:18general in Data Systems information

1:00:21security we want to give people the

1:00:23amount of access that they need to

1:00:25perform their roles and nothing more

1:00:27right so if you don't technically need

1:00:29access to the semantic model or the data

1:00:31pipeline then one way of kind of getting

1:00:34around that is to give people item level

1:00:36access giving people access to only what

1:00:38they

1:00:40need okay just to recap on some of the

1:00:44additional permissions so when you share

1:00:46a data warehouse you get these three

1:00:49additional permissions we have read all

1:00:51data using SQL and what that means is it

1:00:55allows people to read all objects within

1:00:57the warehouse using tsql we also have

1:01:00the read all one L data and with this

1:01:03you're allowing that person to read the

1:01:05underlying oneel files using spark

1:01:08pipelines anything else basically so in

1:01:10the top one they can only use SQL if you

1:01:13give them the second permission it

1:01:14allows them to basically do anything

1:01:15with that data and the third permission

1:01:18allows the user to build reports on the

1:01:20default semantic model not any custom

1:01:22models just the default semantic model

1:01:24when it comes to the lak house these

1:01:25permissions are similar but they're just

1:01:27worded a little bit differently so again

1:01:29if you give them the read all SQL

1:01:31endpoint data it allows them to perform

1:01:34tsql on the tsql endpoint if you give

1:01:37them read all Apache spark then again

1:01:41it's going to allow them to run

1:01:42notebooks and Spark code on top of that

1:01:44data and again the build reports on the

1:01:47default semantic model does exactly what

1:01:49it says on the tin now one point that I

1:01:51just did want to make here is around one

1:01:53Lake data access model now this is a

1:01:56very newly announced feature so it might

1:01:59not have made its way into the exam yet

1:02:01but I do think it's going to have a very

1:02:02big impact on how we manage security in

1:02:06fabric going forward so I did want to at

1:02:08least mention it here I'm not going to

1:02:10be going through it in detail I did just

1:02:12want to flag it you might want to have a

1:02:13look at the documentation page just so

1:02:16that you become aware of it now this

1:02:18feature is not really something I looked

1:02:20at yet much in detail but from just from

1:02:22looking at the documentation from how I

1:02:24understand it is it's going to allow you

1:02:25to perform rback so Ro based access

1:02:29control on things like folders so now

1:02:32that they've implemented folders within

1:02:33a workspace it's going to allow you to

1:02:35define a specific role or give a

1:02:38specific role access control over that

1:02:41folder and then the permissions are

1:02:42going to be inherited for every item in

1:02:44that folder but like I mentioned it's a

1:02:46preview feature and it's relatively new

1:02:48so I'd be surprised if they ask you

1:02:50about this in the exam but it is very

1:02:52important I do think it will change

1:02:53quite a lot in fabric so I wanted to

1:02:55mention it and if you look at the study

1:02:57guide it does say for the dp600 exam it

1:03:00does say that most questions cover

1:03:02features that are generally available

1:03:03the exam may contain questions on

1:03:05preview features if those features are

1:03:07commonly used I think at the moment this

1:03:09isn't commonly used because it's only

1:03:10been released a few weeks ago so that's

1:03:12something to bear in mind Camila says

1:03:14thank you now I understand workspace

1:03:16level and item level sharing in more

1:03:18detail One Last Thing Before You Go

1:03:20we've been working on this government

1:03:22project and I need to apply sensitivity

1:03:24labeling in a workspace can you walk me

1:03:27through it so what even is a sensitivity

1:03:29label well sensitivity labels are a data

1:03:32governance feature and they're created

1:03:35and managed in Microsoft purview so

1:03:38fabric items such as a semantic model

1:03:40can be given a sensitivity label such as

1:03:43confidential right and it's for

1:03:45information protection purposes now in

1:03:47some Industries labeling data and

1:03:49information with a sensitivity label is

1:03:52necessary for compliance with

1:03:54information protection regulations now

1:03:56to apply a sensitivity label in fabric

1:03:59really there's two main methods if we go

1:04:02into that item for example this

1:04:04Lakehouse here what you have in the top

1:04:06tool bar you've got the sensitivity

1:04:08label and you can just click on that

1:04:09drop- down change sensitivity label in

1:04:12there the other option is to go into the

1:04:15settings of that particular fabric item

1:04:17and you can see that in the left hand

1:04:19tool bar there you've got sensitivity

1:04:21label and you can change the sensitivity

1:04:22label in there now one of the options

1:04:25that you can give it if you go through

1:04:26the settings method is to apply to

1:04:28Downstream items so again we've got that

1:04:31notion or that concept of inheritance of

1:04:34the label that you give it here also

1:04:36applies to everything Downstream okay so

1:04:38now we are going to test some of your

1:04:39knowledge for everything that we've

1:04:41learned in this section of the study

1:04:44guide and we're going to start with a

1:04:46case study style question so we're going

1:04:48to going through a bit of a case study

1:04:50and then going to be asked three maybe

1:04:52four questions on this particular case

1:04:54study let's again Toby creates a new

1:04:57workspace with some fabric items to be

1:04:59used by data analysts Toby creates a new

1:05:03security group called Data analysts he

1:05:05includes himself as a member of this

1:05:07Security Group Toby gives the data

1:05:09analyst Security Group a viewer role in

1:05:13the workspace what workspace role does

1:05:15Toby have is it a viewer B member C

1:05:18admin or D contributor pause the video

1:05:21here have a think and then we'll move

1:05:23forward to the answer so the answer here

1:05:25is see now this combines two pretty

1:05:27important Concepts to understand when

1:05:29we're looking at workspace level sharing

1:05:32number one is that the creator of a

1:05:35workspace is always given admin

1:05:37permissions in that workspace now we

1:05:39also have Toby with the viewer role in

1:05:41that workspace CU he's in the security

1:05:43group with viewer role and this is

1:05:46another concept if you have more than

1:05:47one level of permission within the

1:05:49workspace you're always given the higher

1:05:51level so he's got admin role cuz he

1:05:53created the workspace and he's got

1:05:55viewer role because he's in that

1:05:56Security Group Well the admin

1:05:58permissions is always going to be

1:05:59prioritized he's always going to take

1:06:01that role over his viewer role so let's

1:06:03continue this case study Sarah is also a

1:06:07member of that data analyst Security

1:06:09Group she has no other role in the

1:06:11workspace which of the following can

1:06:13Sarah not do in the workspace a execute

1:06:16a data pipeline run SQL scripts in the

1:06:19data warehouse run spark notebook or

1:06:22review the evaluation metrics of a

1:06:24machine learning model now the answer

1:06:25here is C run the spark notebook now

1:06:28when we're looking at the workspace

1:06:30level roles and the permissions for each

1:06:32role we know that Sarah is a viewer in

1:06:34the workspace that's the highest level

1:06:36of permission and you remember that a

1:06:38viewer role can actually execute a data

1:06:41pipeline in a workspace they can also

1:06:44run tsql scripts in data warehouse or a

1:06:47SQL endpoint of a lake house what they

1:06:49can't do is run a spark notebook okay so

1:06:52anything in a notebook they can have a

1:06:54look at the notebook but they can't

1:06:56actually execute any code so C is the

1:06:58right answer cuz what we're looking for

1:07:00is what can she not do in the workspace

1:07:03and D is review the evaluation metrics

1:07:05of a machine learning model which we

1:07:07know we can do because she's just

1:07:08reading the output of that model to

1:07:11continue this case study again Toby

1:07:13wants to delegate some of the management

1:07:15responsibility in the workspace he wants

1:07:17to give this person the ability to share

1:07:20content within the workspace invite new

1:07:23contributors to the workspace but not

1:07:25add new admins to the workspace what

1:07:27role should Toby give this person a

1:07:29admin B member C contributor or D viewer

1:07:33so the answer here is B now the the key

1:07:36point in the question was but not add

1:07:38new admins to the workspace so we know

1:07:41that to be able to add another admin

1:07:42into a workspace you need to have admin

1:07:44permissions yourself so Toby doesn't

1:07:46want to give that person this ability

1:07:48basically so we know it can't be admin

1:07:50it's not going to be viewer it's not

1:07:52going to be contributor we know that the

1:07:54member is kind of one down from admin

1:07:57and that's going to allow you to do all

1:07:58of these three things they can share

1:08:00content they can invite other

1:08:02contributors because a member can add

1:08:04new people either members or

1:08:06contributors or viewers but they can't

1:08:08add other admins so it' be B member the

1:08:11next question is completely separate you

1:08:13have admin role in a workspace Sheila is

1:08:16a data engineer in your team she

1:08:18currently has no access to this

1:08:20workspace at all now Sheila needs to

1:08:22update a data transformation script in a

1:08:24pisb notebook and the script gets data

1:08:27from a Lakehouse table cleans it and

1:08:29then writes it to a table in the same

1:08:30Lakehouse now you want to adhere to the

1:08:33principle of leas privilege what actions

1:08:35should you take to enable this is it a

1:08:38you're going to give Sheila the

1:08:39contributor role in the workspace b

1:08:42share the Lakehouse item with read or

1:08:44spark data permission C give Sheila the

1:08:47admin role in the workspace or D share

1:08:50the lake house item with read all spark

1:08:52data permissions and share the notebook

1:08:54with edit permissions so the answer here

1:08:56is D so one of the clues in this

1:08:58question was the line where it says you

1:09:01want to adhere to the principle of leas

1:09:02privilege so immediately when you see

1:09:04that giving people workspace level

1:09:07access is not really good enough so A

1:09:09and C is giving a role in the workspace

1:09:13so it's going to enable her to

1:09:14contribute and change and edit

1:09:16everything in the workspace but it

1:09:18doesn't adhere to the principal of least

1:09:19privilege so we can immediately rule out

1:09:21a and C so another really important

1:09:25point in the question here was Sheila

1:09:27needs to update a data transformation

1:09:30script in a notebook so she needs to

1:09:32edit the code in a notebook She's Not

1:09:34Just executing an existing notebook she

1:09:36needs to actually make changes to a

1:09:38notebook and so for B you wouldn't have

1:09:41that permission you've got read all data

1:09:43for the spark so you can actually

1:09:45execute a notebook but you can't edit a

1:09:47notebook so be able to make these

1:09:49changes really need access to the

1:09:51notebook and The Lakehouse that that

1:09:54notebook is interfacing with that that

1:09:56notebook is reading from because you

1:09:58can't just share the notebook because

1:09:59then you won't have access to the

1:10:00underlying data and we can't just share

1:10:02the lake housee because you won't have

1:10:03access to the notebook that she needs to

1:10:05edit so the answer is D share the

1:10:07Lakehouse item we're giving spark

1:10:09permissions and we're also giving edit

1:10:11permissions on the notebook next

1:10:13question you have admin role in a

1:10:15workspace you want to pre-install some

1:10:17useful python packages to be used across

1:10:20all notebooks in the workspace how do

1:10:22you achieve this a in the fabric ad

1:10:25admin portal go to spark settings and

1:10:27install the libraries B go to workspace

1:10:29settings spark settings and then Library

1:10:32management C create an environment

1:10:34install the packages in the environment

1:10:36go to the workspace settings spark

1:10:38settings and set it as the default

1:10:40environment or D go to capacity settings

1:10:43and then default libraries so the answer

1:10:45here is C creating an environment and

1:10:49then going into your workspace settings

1:10:50and setting it as the default

1:10:53environment for for spark now now this

1:10:55is a bit of a a naughty question because

1:10:57B is the old way so it used to be you go

1:11:00to workspace settings spark settings and

1:11:02there was a section for Library

1:11:03management but that's actually not

1:11:05possible anymore the way to do it as I

1:11:07mentioned is to create an environment

1:11:09and then in your spark settings make it

1:11:10the default environment A and D don't

1:11:13actually exist these capabilities so

1:11:16these kind of red herrings so the answer

1:11:18is C Camila says thanks she's seriously

1:11:20impressed with your knowledge again in

1:11:22this lesson we covered all five of these

1:11:25elements of the dp600 study guide from

1:11:28workspace and item level sharing data

1:11:30sharing for data warehouses and lake

1:11:33houses sensitivity labeling and then

1:11:36workspace and capacity level settings

1:11:39and the good news is again you've won an

1:11:40extension to the contract Camila would

1:11:42like you to implement control over the

1:11:45entire analytics development life cycle

1:11:48in her organization so for this we're

1:11:50talking Version Control deployment

1:11:52pipelines powerbi projects all that good

1:11:55stuff that's what we going be looking at

1:11:56in the next lesson so click here to

1:11:59continue that lesson hey everyone

Manage the analytics development lifecycle

1:12:01welcome back to the channel today we're

1:12:03continuing the DP 600 series looking at

1:12:06what it's going to take to hopefully P

1:12:08that exam and become a fabric certified

1:12:11analytics engineer Today Is video 4

1:12:14we're going to be looking at managing

1:12:16the analytics development life cycle and

1:12:18in this section exam we're going to be

1:12:20focusing on implementing Version Control

1:12:23creating and managing powerbi projects

1:12:26planning and implementing deployment

1:12:27Solutions performing impact analysis on

1:12:31Downstream activities deploying and

1:12:33managing semantic models through the

1:12:35xmla endpoint and creating reusable

1:12:38assets so powerbi template files powerbi

1:12:41data source files all these kind of

1:12:42things so that's what we've got in store

1:12:44for you today as ever there's going to

1:12:46be some sample questions at the end I'll

1:12:48also be posting all of the lesson notes

1:12:52and links for further resources in the

1:12:54school community go and grab them there

1:12:56you will play the main character in the

1:12:58scenario again we're going to be

1:12:59continuing this theme this scenario that

1:13:02we've been developing throughout this

1:13:03course you're the main character are you

1:13:05ready let's begin so as you know already

1:13:07we are a consultant working with the

1:13:10client called Camila and you've already

1:13:12helped Camila plan and Implement her

1:13:14fabric implementation her environment

1:13:17now the time has come to implement

1:13:19analytics development life cycle now she

1:13:21wants to focus on Version Control at

1:13:23least initially but to be honest he's

1:13:25not really familiar with Git Version

1:13:28Control never really used that in an

1:13:29analytics environment before so what

1:13:31you're going to have to do is set up a

1:13:34call with Camila and walk her through

1:13:36the basics of git first and Version

1:13:38Control what is it why does it even

1:13:40exist why is it now a part of fabric and

1:13:44what are some of the key terms and

1:13:45terminology that we need to understand

1:13:47when we're implementing Version Control

1:13:49just bear in mind that all of the git

1:13:52integration features are currently in

1:13:54preview and some of the other features

1:13:56we're going to be looking at today like

1:13:57the deployment pipeline functionality a

1:14:00lot in this kind of analytics life cycle

1:14:02stuff is still in the previous stages so

1:14:04just bear that in mind when we're

1:14:06walking through these examples I'll show

1:14:07you what's possible today but if you're

1:14:09coming from using like GitHub or git a

1:14:12lot in the past and a lot of the

1:14:14features that you might be used to are

1:14:16currently not available so let's move on

1:14:19okay so I just wanted to start this

1:14:21video with a bit of a primer on git and

1:14:24Version Control in general and I think

1:14:26the best way to do that is to show you

1:14:28what this looks like and along the way

1:14:30we can explain some of the key Concepts

1:14:32in git and Version Control I realize

1:14:35that probably a lot of people who have

1:14:36been coming from maybe a powerbi

1:14:38background don't have experience with

1:14:40Version Control and get in general so

1:14:42let's set up a project and start at a

1:14:45basic level introduce more and more

1:14:47Concepts as we go through this demo so

1:14:50in fabric the way you're going to start

1:14:52is with Azure devops now AZ devops is a

1:14:55service that's used primarily for

1:14:57software development and managing

1:14:59infrastructure and the devops life cycle

1:15:02of that infrastructure and software

1:15:04development but it's also where we're

1:15:06going to use to store our code and our

1:15:09artifacts and our powerbi project files

1:15:12because they come with repositories so

1:15:14you need to set up an account just go to

1:15:16this website here and I'll leave a link

1:15:18in the description or in the school

1:15:19Community now once you've set up Azure

1:15:21devops for your organization you'll come

1:15:24through to this page here I've set up an

1:15:27organization called fabric University

1:15:29and then you have this concept of a

1:15:31project so a project is going to be

1:15:33where you store your repositories your

1:15:35code has lots of other features as well

1:15:37but we won't be going into too many of

1:15:39the other features in this short video

1:15:41so we're going to be creating a new

1:15:42project for this demo let's just call it

1:15:45dp600 practice we're going to make it

1:15:48private and then just click on create

1:15:50okay so this is an Azure devops project

1:15:53on the left hand side you can see here

1:15:55make that a little bit bigger on the

1:15:56left hand side we've got some useful

1:15:59things to know about right so this is

1:16:01just the overview where you can see an

1:16:02overview of the project probably the

1:16:04most important ones to bear in mind are

1:16:06boards so this is like a Work Management

1:16:08tool for doing development work we're

1:16:11not going to be focusing much on that in

1:16:12this video what we're interested in is

1:16:14this Repose repositories right so

1:16:17repositories is kind of the core thing

1:16:19that we need to implement Version

1:16:22Control now repositories are also

1:16:24available in GitHub but currently that

1:16:26integration is not possible between

1:16:28Fabric and GitHub you can only use Azure

1:16:31devops repos so that's why we're using

1:16:34this currently so a repository is just

1:16:36where we're going to store our code in

1:16:39the cloud so when you're developing

1:16:41locally like a powerbi report for

1:16:43example really what we want to do is

1:16:46store that report in the cloud in our

1:16:50repository and that's where we're going

1:16:51to be controlling tracking changes

1:16:54managing who can make changes to that

1:16:56project in the repository so let's just

1:16:59do a bit of configuration here to set up

1:17:00this repository so that we can use it

1:17:02for Version Control so the first thing

1:17:04I'm going to do is right at the bottom

1:17:06here initialize a main branch with a

1:17:08readme and a readme file is just a

1:17:11markdown file it's like a text file just

1:17:13to initialize our branch and we'll talk

1:17:16a bit more about branches in a bit more

1:17:18detail here currently I have one branch

1:17:20and it's the main branch here now

1:17:22branching is kind of like a whole field

1:17:24in itself there's lots of different

1:17:26strategies for how we can manage

1:17:28different branches when we developing

1:17:30our code or developing our fabric

1:17:33artifacts but we'll get into that in a

1:17:34bit more detail a bit later on for now

1:17:37what we're interested in really is

1:17:39creating a copy of this repository on

1:17:43our local machine right so that's going

1:17:45to be the first sync we want to get

1:17:47these two in sync because the way that

1:17:50version control works in general is you

1:17:53develop locally on your local machine

1:17:55for powerbi and then you sync those

1:17:57changes up to the repository in the

1:18:00cloud and if you have a big team of

1:18:02people maybe you have 10 people doing

1:18:03this they're all syncing their changes

1:18:06into this main branch by default so what

1:18:09we want to do is clone this repository

1:18:12and we get this URL here for this git

1:18:14repository we can copy the URL and what

1:18:17we want to do is clone that on our local

1:18:19machine so there's various different

1:18:20ways to do that you might see in the

1:18:22documentation Microsoft use a tool

1:18:25called vs code I personally prefer using

1:18:28GitHub desktop this is another tool it's

1:18:30completely free you can download it just

1:18:33Google GitHub desktop download it's

1:18:35built by obviously GitHub but it's also

1:18:38possible to do Azure devops repos in

1:18:41here as well and it just makes the whole

1:18:43process of Version Control and git a lot

1:18:47simpler it basically builds a UI on top

1:18:49of a lot of the git functionality so the

1:18:52first thing we want to do is click on

1:18:53file and clone repository because we're

1:18:56trying to drag things down from the

1:18:59internet and if you click over to URL we

1:19:01can just copy the URL that we've got

1:19:03here that we got from our azid Devo

1:19:05repository and we can clone it now it's

1:19:07going to ask for a username and password

1:19:09so if you go back to a devops we've got

1:19:12generate this git credentials and we can

1:19:15just copy the username in there copy the

1:19:18password in there save and retry so

1:19:20that's just going to authenticate with

1:19:22our Azure devops account so we know that

1:19:24okay this is the actual person that owns

1:19:26that repository it's creating that

1:19:28authentication between the two okay so

1:19:30what has that actually done well now

1:19:32what we've got is a copy of what we have

1:19:35in Azure devops in our local machine you

1:19:38can see the file path there is C users

1:19:41learn documents git and what we can

1:19:44actually do is open this in Explorer

1:19:47okay so now I have a copy of those files

1:19:50and folders from Azure devops on my

1:19:53local machine so that is the first first

1:19:54thing that we need to do really is clone

1:19:57the repository get that copy local and

1:19:59what's going to happen now we've got

1:20:01this all set up and synced so that

1:20:04GitHub desktop is Now tracking this

1:20:06folder so any changes you make locally

1:20:09to this folder it's going to pick those

1:20:11up right so for example if I do a new

1:20:14text file call it my file now if we go

1:20:16back to get up desktop you can see that

1:20:18it's picked up that change already it's

1:20:20saying you've added a new file in there

1:20:23it's called my file . text and if I open

1:20:26that with notepad for example just call

1:20:28it my amazing file and I'll save that

1:20:31here we can see again it's tracked to

1:20:32the changes so git works by tracking the

1:20:35changes in text based documents okay so

1:20:40traditionally it's used for code so a

1:20:43python file a SQL script or a Javascript

1:20:46file or C all these kind of things are

1:20:49textural based right so they work really

1:20:52well with Git so now we've made some

1:20:54Chang changes to our local repository

1:20:57right we've added this text file but

1:20:59what you'll notice is that you know here

1:21:01in Azure devops nothing has changed yet

1:21:03because what we need to do is push those

1:21:05changes into the cloud environment maybe

1:21:09you've got a team you want other people

1:21:10to see the changes that you've made and

1:21:12we'll start by using this text file

1:21:15right and then we'll slowly move up to

1:21:16more complicated stuff we'll look at

1:21:18powerbi projects as well but for now

1:21:21let's just push these changes so you see

1:21:23down here you can do created a new file

1:21:26so created a new file and then what

1:21:28we're going to do is click this commit

1:21:29to main button so this is going to

1:21:32commit our changes onto that main branch

1:21:35and up here we click on push to origin

1:21:38so we're going to commit and then push

1:21:40the changes up into the cloud right into

1:21:42Azure devops and so now what's going to

1:21:44happen if we refresh our Azure devops we

1:21:47can see that file there we can see my

1:21:49file. text my amazing file that's in

1:21:52there in this repository in azure devops

1:21:55now Okay then if we want to make more

1:21:57changes maybe we want to change this a

1:21:59bit more even more amazing file and we

1:22:01save that again it's going to show you

1:22:03oh well your original file was this okay

1:22:06this is what we've got stored in Azure

1:22:08devops this is what the the main branch

1:22:11is currently saying now you've changed

1:22:13that file right and we've noticed that

1:22:15okay so it's tracking the changes

1:22:17between every change you make in that

1:22:19file so that's really important thing to

1:22:21bear in mind so let's just update this

1:22:24text file obviously in practice in

1:22:26proper environments you add a more

1:22:28meaningful commit message because these

1:22:29are really important and we'll push that

1:22:31again to the origin using this button

1:22:33here so again if we just go back to the

1:22:36Azure devops you can see that that has

1:22:38now updated in here as well great so

1:22:40that is the very very basic

1:22:42implementation of git and Version

1:22:45Control and syncing our local repository

1:22:48with our Azure devops repository in the

1:22:50cloud so at this stage I think it's

1:22:52worthwhile just noting what we we've

1:22:54done here right so at the top we've

1:22:56created an A devops account we've added

1:22:59that readme file we've copied the git

1:23:01URL and cloned it locally and then we've

1:23:04modified that local folder create a new

1:23:06file modified the file committed those

1:23:09changes back pushed the origin and

1:23:11observe the changes in as devops now

1:23:13this is great it's a good first step

1:23:14right we can track changes now between

1:23:16files but one of the core benefits of

1:23:19git and Version Control in general is

1:23:21that instead of just allowing anyone to

1:23:23update that main branch what we can do

1:23:26is we can protect that main branch so

1:23:29that we can control who has access to

1:23:32and who has the ability to overwrite

1:23:35those changes so let's make a few

1:23:37changes to the figuration of our

1:23:40repository in Azure devops what we're

1:23:42going to do is we're going to protect

1:23:44that main branch so that nobody can just

1:23:46overwrite that file anymore we want to

1:23:49add in a review and approval phase and

1:23:52this is going to help a lot with our

1:23:54quality and control over who and how

1:23:58that code base is being changed so let's

1:24:00have a look at that in a bit more detail

1:24:02okay so in our Azure devops project here

1:24:05we can go to Project settings and then

1:24:07we can go down to repositories and then

1:24:09what we're going to do is add in a

1:24:12policy for branch policies protect

1:24:15important branches Nam spaces so we're

1:24:17going to add something in here protect

1:24:19the default Branch create automatically

1:24:21included reviewers so I'm going to add

1:24:23myself in here as a reviewer just

1:24:26because there's only one person in this

1:24:28project in reality you might have a team

1:24:30of maybe senior developers or lead

1:24:32developers that responsible for

1:24:35reviewing code and approving changes to

1:24:38the the code base so now we've set that

1:24:40up let's have a look at doing something

1:24:42a bit more so we've got our file here so

1:24:45if I save this I've just added another

1:24:47line in here new changes and if we go

1:24:49back to our GitHub here we can see that

1:24:52oh it's picked up that new changes again

1:24:54but now what I'm going to do is update

1:24:58this so added some new changes and I'm

1:25:01going to commit that and when I try to

1:25:02push this to the origin is going to say

1:25:04error okay because now pushes to this

1:25:07Branch are not permitted you must use a

1:25:09poll request to update this Branch so

1:25:11that's going to introduce a few new

1:25:13Concepts that we need to get around this

1:25:17really number one is the concept of

1:25:19branches and number two is the concept

1:25:21of pool requests let's have a look at

1:25:23both of those now now so now that we've

1:25:25implemented some protection over that

1:25:28main branch we have to change how we

1:25:30develop okay so now we have to use

1:25:34branching So currently all of my files

1:25:37that text file and the readme file is on

1:25:40this main branch so if I now want to

1:25:42update what's in that text file I'm

1:25:45going to have to create a new Branch

1:25:46because nobody can update and edit that

1:25:49main branch directly so what we can do

1:25:51in GitHub desktop it's very easy so what

1:25:54we're going to do is click on Branch New

1:25:57Branch now we can change the name to

1:25:59change text file create this branch and

1:26:02now if we publish this branch and we can

1:26:03add in some changes here so now if we go

1:26:06back to Azure devops let's just go back

1:26:08to our project here and back to Repose

1:26:11and if we go down to this change text

1:26:14file you can see that it has actually

1:26:16found this new Branch this is a branch

1:26:19and if we go over to the pull requests

1:26:21section we'll explain what that means in

1:26:22a minute but you can see that it's

1:26:24registered the fact that we have now

1:26:26published a new Branch so we've made

1:26:29some changes added that new changes line

1:26:31in the text file and it's popped up

1:26:33saying oh do you want to create a pool

1:26:35request here and so the pool request is

1:26:37when we want to make a change or merge

1:26:40some changes into the main branch okay

1:26:43and you can only do this through a

1:26:45review because that's the policy that

1:26:46I've set up on that main branch so what

1:26:49we can do is create a PO request you can

1:26:51say oh I've added some new changes in

1:26:53here you can add in a reviewer which is

1:26:55going to be me and you can say oh please

1:26:57review this and you can create a pull

1:27:00request so what that's going to do is

1:27:01going to notify the reviewer and say oh

1:27:05this person wants you to review their

1:27:07changes to the code base on this change

1:27:09text file branch and now as the reviewer

1:27:12because I'm kind of playing both roles

1:27:14here I can have a look at the file

1:27:16changes I can review the changes like oh

1:27:19yep added new changes that looks good to

1:27:21me I'm going to approve that change and

1:27:23now what you can do is complete so what

1:27:25complete is going to do is merge those

1:27:29changes into the main branch so your

1:27:31protected Branch because we've gone

1:27:33through the reviewer process the text

1:27:36file has been reviewed now we can merge

1:27:38it safely into the main branch there's a

1:27:41few options here that are quite

1:27:43important to bear in mind and we'll look

1:27:46at those when we move into fabric we'll

1:27:48look at those in a bit more detail but

1:27:50generally kind of the default settings

1:27:52here complete Associated work items

1:27:54after merging it's fine we don't want to

1:27:57delete change to text file after merging

1:28:01generally in Fabric and we'll get to why

1:28:03that is in a minute we'll just do on

1:28:05complete merge Okay so we've now merged

1:28:07our poll request let's have a look at

1:28:10the repository so now we have my text

1:28:13file in the main branch and it's got new

1:28:15changes Okay so we've merged our changes

1:28:18into this main branch great that's the

1:28:21first section of this demo done okay and

1:28:24this second half there what we did was

1:28:25protect the main branch Okay so we've

1:28:28added an approv in the repository we've

1:28:31updated the repo approv policy we tried

1:28:35committing a new change and it was

1:28:36rejected we've added a new feature

1:28:38Branch we brought the changes onto that

1:28:40Branch we've committed that branch and

1:28:43we've added a p request then our

1:28:45approver has approved it and we've

1:28:47merged it into the main branch and just

1:28:49to reiterate the benefits of that is

1:28:51that now we're checking the changes

1:28:53between files and we're also controlling

1:28:55who can make changes so we're protecting

1:28:58that main branch and saying okay if you

1:28:59want to update this it has to go through

1:29:01an approval process now up until this

1:29:03point we've just been using one text

1:29:06file actually git can track the changes

1:29:09between any text based file format text

1:29:13base that's the the problem that existed

1:29:16with PB files PX files kind of a bit of

1:29:19a black box right if you can't open a

1:29:22file in notepad and understand what's

1:29:24going on and it's going to struggle with

1:29:26Source control Version Control in git

1:29:29now that is one of the core reasons why

1:29:31Microsoft developed the powerbi project

1:29:35file format because it represents a

1:29:37powerbi project in text based files

1:29:41right so let's just have a look at an

1:29:43example of a pbip first and then we'll

1:29:46look at how we can integrate that into

1:29:48Version Control Systems okay so here we

1:29:49are in powerbi desktop and I've just got

1:29:51one of the sample reports here from

1:29:54Microsoft the competitive marketing

1:29:56analysis report now what we're going to

1:29:58do is this is just a PB file at the

1:30:00moment what we're going to do is save

1:30:03this as a pbip a powerbi project file

1:30:07then we're going to explore that and

1:30:08have a look at what that looks like so

1:30:10it's fairly easy to save your powerbi

1:30:13file as a powerbi project file so if we

1:30:16go to file and then save as and what we

1:30:20can do is we can find our git reposit so

1:30:24that's in document git dp600 practice

1:30:28that's the repository that we just

1:30:29cloned from Azure devops and it's

1:30:31currently empty for this kind of thing

1:30:33we can change the save as type to a

1:30:35powerbi project and we can just save it

1:30:37like that click on Save and now we have

1:30:39our powerbi project saved in pbip format

1:30:43so let's just have a look at what that

1:30:45looks like okay so this is our git

1:30:47repository dp600 practice and we've

1:30:50saved our powerbi project as a BP file

1:30:55it's actually a number of different

1:30:57files and folders right what we've got

1:30:59here is we've got the read me that was

1:31:01already in the repository and my file is

1:31:03that text file that we've been working

1:31:05with but now what we've got is two

1:31:06folders a git ignore file and this pbip

1:31:11file format so what it's done is it's

1:31:14basically decomposed our powerbi project

1:31:17into a series of files and folders that

1:31:20describe what's going on in that report

1:31:23Okay so let's move back to uh GitHub

1:31:26desktop here and what you'll notice is

1:31:29well the first thing is I've created a

1:31:30new Branch okay so now that we're going

1:31:32to be updating this pbip file we'll

1:31:36notice that all of the files are now

1:31:37listed here CU we created that pbip

1:31:40they're all text based they're all

1:31:42trackable using Version Control first

1:31:44commit of the pbip again we can commit

1:31:47that up into Azure devops push that to

1:31:51the origin in Azure devops now if we go

1:31:54through into azid devops again poll

1:31:56requests we can see that here we can

1:31:58create a new poll request and then we

1:32:00can commit that in there the reviewer

1:32:03will be myself again create that approve

1:32:05it complete merge so now what we've got

1:32:09is our powerbi report in our Azure

1:32:13devops environment and it's been

1:32:15approved and it's now trackable through

1:32:17Version Control now imagine I'm a

1:32:19powerbi developer and I want to make

1:32:21some changes to that report what is that

1:32:24going to look like well we're going to

1:32:25create a new Branch so anytime you're

1:32:27making changes to your powerbi file

1:32:29you're going to create a new Branch

1:32:31powerbi changes obviously you'd make it

1:32:33a bit more descriptive than that but

1:32:35that's good enough for this purpose here

1:32:37then we're going to go through and

1:32:40actually this isn't a marketing report

1:32:41it's sales and marketing so we're going

1:32:43to update the title here just as an

1:32:45example we're going to save that file

1:32:47and now this is the real power of

1:32:50Version Control for our power VI assets

1:32:53but also all of the stuff in fabric that

1:32:55we'll have a look in a minute any

1:32:56changes that we make to that powerbi

1:33:00file are now being tracked in our

1:33:03Version Control System okay and then

1:33:05when we're done we can push those

1:33:07changes up into Azure devops go through

1:33:10the review process merge them into the

1:33:13the main branch okay so you're probably

1:33:14thinking great but up until this point

1:33:17you haven't even mentioned fabric I

1:33:19thought this was a channel talking about

1:33:20fabric well we just built up the

1:33:21concepts of git vers control we've

1:33:24looked at the local case of powerbi

1:33:27report development but everything else

1:33:28in fabric happens in the fabric Cloud so

1:33:31now let's look at where fabric fits into

1:33:33all this okay how do we set up a

1:33:35repository in our workspace how do we

1:33:38link a repository to our workspace let's

1:33:40have a look at how fabric fits into this

1:33:42picture now okay so here we are in

1:33:44Microsoft fabric what I've done is I've

1:33:46set up a workpace what we're going to do

1:33:49is link it to our Azure devops

1:33:52repository our Azure devops project

1:33:54so to do this you will need to be a

1:33:56workspace administrator go to workspace

1:33:58settings so the first time that you

1:34:01click on get integration you'll need to

1:34:03actually sync it with your account I've

1:34:05already done that with mine and it's

1:34:07going to allow you to select the

1:34:08organization the project the repository

1:34:11the P600 practice and a specific branch

1:34:14that you might want to connect to now

1:34:15one thing I would say is that the way

1:34:17that we've done that protecting of the

1:34:19branching in the last section of the

1:34:22video what I found in practice is that

1:34:24this doesn't work particularly well

1:34:25currently with the way that git

1:34:27integration is set up in fabric but

1:34:29we'll have a look at what it looks like

1:34:30just click on connect and sync okay so

1:34:33now our first Sync has started and it's

1:34:36been done successfully what you'll

1:34:37notice is this Source control section

1:34:40here it's got zero in it so our source

1:34:42control so our things in AZ devops

1:34:45repository and what's in fabric is

1:34:47perfectly in sync currently and we can

1:34:49see that as well with the green sync and

1:34:52green synced here so now that we've got

1:34:54uh what's in our Azure devops repost

1:34:56synced with what's in fabric now we can

1:35:00add in fabric items into this mix right

1:35:03so we can create a new data pipeline for

1:35:06example create a data pipeline at the

1:35:08moment I'll just keep it as an empty one

1:35:09and we can just go back to our

1:35:11repository and we can see that it has

1:35:14been uncommitted right and we can see

1:35:16the source control button up here is now

1:35:19saying one so if we click on that we can

1:35:20see that we got one change here so comp

1:35:23compared to what's in our main branch in

1:35:26AZ devops it's noticed that we've got a

1:35:28change so it's similar to how we saw in

1:35:31GitHub desktop which tracks local

1:35:33changes this is tracking changes in

1:35:36fabric as well so what we can do is

1:35:40commit this change this data pipeline

1:35:42into our source control in AZ devops now

1:35:45to do this because we've got a

1:35:47protection on that main branch let's

1:35:49just try committing and see what it says

1:35:51first commit data pipeline but we know

1:35:53know that we've got protections on that

1:35:55main branch so when we commit I don't

1:35:57think it's going to allow us to do that

1:35:59no forbidden due to the branch policy

1:36:01and this is where it gets a bit more

1:36:02complicated within fabric what we're

1:36:04going to have to do is check out a new

1:36:06Branch so maybe we want to do adding

1:36:10data pipeline that could be our new

1:36:12branch and then that's going to allow us

1:36:14to commit added a new data pipeline

1:36:16we're going to click on the item that we

1:36:18want to merge which is our data pipeline

1:36:20or commit not merge and now that's going

1:36:22to go through into Azure devops okay so

1:36:26if we open up our poll requests section

1:36:28boom we can see adding data pipeline so

1:36:31now we've got these two things in sync

1:36:33between Fabric and Azure devops and our

1:36:36local repository for powerbi development

1:36:38and Azure devops and again if we want to

1:36:40merge that into the code base into the

1:36:43main branch we can go through a familiar

1:36:45process of choosing a reviewer creating

1:36:48that PLL request and if you're the

1:36:50approver you can then approve that PLL

1:36:52request and complete and merge it into

1:36:55the main branch so then if we go back to

1:36:57our repository now we can see that in

1:37:00our main branch we have this my data

1:37:01pipeline data pipeline so that is now

1:37:04being tracked with Version Control in

1:37:07our repository and if we go back to

1:37:09fabric we can say that we've now got

1:37:12three synced items here so it's

1:37:15perfectly in sync with what's in our

1:37:17Branch now the problem here or the

1:37:20slight limitation that I found is

1:37:22changing back to the main branch cuz

1:37:25currently we're on this adding data

1:37:27pipeline Branch okay and that was just a

1:37:30feature Branch we want to move back onto

1:37:32the main branch normally so that we can

1:37:34maybe create another another Branch but

1:37:35we don't want to be branching from this

1:37:37added data pipeline we want to be

1:37:39branching from the main branch now that

1:37:41is possible we can just go back to

1:37:43workspace settings go through to the get

1:37:46integration again and we can change our

1:37:48Branch back to the main so we're going

1:37:50to switch and override and now we're

1:37:51going back to the the main branch and

1:37:54all three items are synced because we

1:37:55know that we've merged into that main

1:37:57branch but the problem here is that only

1:37:59the workspace administrator can change

1:38:01that Branch back so if you found any

1:38:03better ways of working with protected

1:38:05branches in git and fabric please let me

1:38:08know okay so we've covered quite a lot

1:38:10of ground there from the basics of git

1:38:13and Version Control in general right

1:38:15through to syncing our changes or

1:38:18tracking changes in a powerbi project

1:38:20into AZ devops and then the same from a

1:38:23fabric environment so tracking changes

1:38:24that we make within fabric also to the

1:38:27same Azure devops repository let's just

1:38:29do a bit of a summary there and focus

1:38:31back on the exam because obviously not

1:38:33all of that is going to be tested some

1:38:34of it was just for your background

1:38:36knowledge so in general Version Control

1:38:38with Git allows you to track changes

1:38:41made to fabric items we can also revert

1:38:44back to older versions of an item as

1:38:46well I didn't show you how to do that

1:38:48but that is also possible within git

1:38:50item management Git Version Control now

1:38:52one of the benefits of this is that

1:38:53multiple users can collaborate on the

1:38:55same fabric item or the same powerbi

1:38:58report okay so no longer are you sending

1:39:00around a PBX file to your colleague

1:39:03everyone can work on the same pbip file

1:39:07and changes are tracked using Version

1:39:10Control and you can also update the same

1:39:12report at the same time as long as

1:39:14there's no conflicts when you merge into

1:39:16your branch two different changes from

1:39:18two different people can both be merged

1:39:20into the same kind of central branch as

1:39:23long as there's no conflict as I

1:39:24mentioned that's absolutely fine so it's

1:39:26a good way that enables collaboration on

1:39:28powerbi reports we've also looked at how

1:39:30you can Implement a check and approval

1:39:32process for approving changes made to

1:39:35fabric items so if you don't want your

1:39:38Junior developer to be updating your

1:39:41fabric notebooks and just pushing those

1:39:43into your Git Version controlled

1:39:45repository Without You approving them

1:39:47you want to set up protections on that

1:39:49on your main branch or your equivalent

1:39:51of a main branch for they get pushed

1:39:53into to production now currently the

1:39:55following items are supported I would

1:39:57argue that not all of these are probably

1:39:58fully supported like as as fully

1:40:01supported as we would like but it is

1:40:03possible to at least check them into a

1:40:05version control system so data pipelines

1:40:08Lakehouse notebooks pageat reports

1:40:11reports apart from ones that are

1:40:13connected to Azure analysis services and

1:40:15also semantic models except for these

1:40:18exceptions here so that is git

1:40:20integration and Version Control now the

1:40:22next section of the exam is is related

1:40:24to this but it's around deployment

1:40:26Solutions so let's just look at what do

1:40:28we mean by deployment Solutions well

1:40:31actually let's start with deployment

1:40:32what do we mean by deployment because

1:40:34for many people coming from the world of

1:40:36analytics deployment is probably a bit

1:40:38of a New Concept okay so rather than

1:40:41having one copy of a powerbi report for

1:40:45example which is your production copy

1:40:47instead we're going to have multiple

1:40:48copies typically three sometimes four

1:40:51we're going to have a development

1:40:52version and that's the version that

1:40:53you're using when you're making changes

1:40:55to the reports we're going to have a

1:40:57test version of that report so that is

1:40:59the report that you send to your client

1:41:02or your colleagues to review or maybe

1:41:04you've got automated testing in place

1:41:07that's done at that stage and the test

1:41:09version is sometimes also called the

1:41:10staging version in databases we've also

1:41:13got the production report right so that

1:41:15is the the public facing report that

1:41:18gets given to your client or it gets

1:41:20shared within your organization now

1:41:22Microsoft have released a feature called

1:41:26deployment pipelines to help manage

1:41:28these three environments or more in a

1:41:31bit more detail so let's have a look at

1:41:33deployment pipelines in Microsoft Fabric

1:41:35in a bit more detail okay so now let's

1:41:36look at how we can set up deployment

1:41:39pipelines in Microsoft fabric so what

1:41:41you're going to need here is three

1:41:43workspaces and each workspace is going

1:41:46to represent a different environment a

1:41:48different stage in our deployment

1:41:49pipeline so if we click through to

1:41:51workspaces you'll see that I've set up

1:41:54dp600 Dev test and production so these

1:41:57are three separate workspaces and

1:41:59currently they're completely empty so

1:42:01what we're going to do is set up a

1:42:02deployment pipeline so that we can

1:42:04manage that deployment process from

1:42:06development test through to production

1:42:08because what we want to be doing is make

1:42:10doing a lot of our development work

1:42:11obviously in the development workspace

1:42:13and then when it's ready we want to push

1:42:15those changes into our test workspace so

1:42:18that we can do testing know share it

1:42:20with colleagues who might want to test a

1:42:22report or a notebook run test scripts so

1:42:26integration tests unit tests data

1:42:29validation checks something like that in

1:42:31this test environment and then the

1:42:32production is going to be where we're

1:42:34going to be sharing making our public

1:42:36our production reports or notebooks and

1:42:39things like that so what we're going to

1:42:40do to set up a deployment pipeline is

1:42:42click on the workspaces and then

1:42:44deployment pipelines and we're going to

1:42:45set up a new Pipeline and we're going to

1:42:47give it a name so we can call it dp600

1:42:50Pipeline and it's going to ask you to

1:42:52customize your stages now you can add

1:42:55multiple stages you can add as many as

1:42:56you like here for our example we're just

1:42:58going to do three development test and

1:43:00production and click create and now

1:43:02we're going to get through to this kind

1:43:03of wizard this UI interface it's going

1:43:06to allow us to specify which workspace

1:43:08we want to connect to which stage in our

1:43:11deployment pipeline process so we want

1:43:13to find our dp600 Dev workspace assign

1:43:17that here then for the test stage dp600

1:43:22test assign that here then for the

1:43:24production do that as well so now we've

1:43:26synced our three workspaces with our

1:43:29three stages in the deployment Pipeline

1:43:31and currently we can see that this

1:43:32completely empty right so we haven't got

1:43:34any items in any of these three

1:43:35workspaces so let's go through to our

1:43:38dp600 Dev workspace and we can see that

1:43:40it's here we synced it with this

1:43:42deployment pipeline so it knows it's

1:43:44part of a deployment pipeline here so if

1:43:46we create a new notebook so this is our

1:43:49new ETL notebook let's just call it ETL

1:43:52notebook could be doing whatever you

1:43:54want in there doesn't really matter for

1:43:56these for the purposes of this demo and

1:43:57we're going to go back into our

1:43:59workspace and now we can see we've got

1:44:01this ETL notebook say we're happy with

1:44:04that we've done our development work and

1:44:05we're happy with it now we want to push

1:44:07that into our test environment so we'll

1:44:11go into our deployment Pipeline and here

1:44:14what you can see is that now we can see

1:44:15this deploy right so we can select any

1:44:18items that we've got in our Dev

1:44:20environment and we can deploy them we

1:44:22can add a note in there

1:44:23if you want and we can deploy that into

1:44:25our test environment so it's going to

1:44:27move one stage to the right into our

1:44:29test environment what it's going to do

1:44:30is copy exactly the file that you've got

1:44:33there and it's going to copy it and

1:44:34paste it into our test environment so

1:44:37now we can see we've got one notebook in

1:44:39this test environment okay so let's just

1:44:41have a look at that there so workspaces

1:44:43dp600 test so now we have ETL notebook

1:44:46in our test environment now one thing to

1:44:48bear in mind is that we have a history

1:44:50of our deployments so we've seen that

1:44:52we've deployed to test at this time of

1:44:54date by this person me and we can see

1:44:57the number of items that have changed so

1:44:59we've got one new item in that

1:45:01deployment now another thing that we

1:45:02need to look at is this button here so

1:45:05this is quite important it's called

1:45:06deployment rules so if we click on there

1:45:09we can look at a deployment Rule and

1:45:12deployment rules basically allow you to

1:45:15change different parameters and settings

1:45:19at different stages in the pipeline and

1:45:21they differ depending on the the fabric

1:45:23item that you're looking to assign these

1:45:26deployment rules so for the notebook our

1:45:28deployment rule that we can set is

1:45:30changing the default lake house okay so

1:45:33we can add a rule in there whereby we

1:45:35can change the default Lakehouse so

1:45:36maybe you have a development Lakehouse

1:45:39and a test Lakehouse and when we push to

1:45:41the test environment actually we want

1:45:43that notebook to be reading from the

1:45:45test Lakehouse so that's what we're

1:45:47going to be doing here in this

1:45:48deployment rules section so that's

1:45:51something important to note there now if

1:45:54we want to actually deploy this all the

1:45:56way into production we can click on

1:45:57deployment at this test stage and now

1:46:00it's going to copy that notebook into

1:46:01our production environment now if we

1:46:03click on the settings of this stage the

1:46:05production stage we can see that we've

1:46:07got this make stage public okay and what

1:46:10that means is that people can actually

1:46:12access the output of this stage so if

1:46:15they're in that workpace they're going

1:46:17to be able to see the output they're

1:46:19going to be able to see the items that

1:46:21are in there with the other stages like

1:46:22this test environment by default that is

1:46:25not a public stage so so that's

1:46:27something to bear in mind around public

1:46:29stages and non-public stages so that's

1:46:32the very basics of deployment pipelines

1:46:35now currently the functionality I would

1:46:37say is quite limited as to what you can

1:46:39do in deployment pipelines in fabric I

1:46:42think this is a feature that they're

1:46:43going to be adding a lot more to

1:46:45currently it's a very manual process

1:46:46right and normally when we're deploying

1:46:49stuff in in the real world in software

1:46:51development World a lot of this is

1:46:53automated Okay so we've looked at the

1:46:55basics of deployment pipelines what that

1:46:58looks like in fabric let's just do a bit

1:46:59of a summary of what we've just learned

1:47:01there again focusing back on the exam so

1:47:04the overall goal is to add layers of

1:47:07control when we're developing and

1:47:09deploying new fabric items okay or

1:47:11making changes to existing items in

1:47:13fabric ultimately we trying to ensure

1:47:15that new things that you develop are not

1:47:17going to break your existing analytic

1:47:20Solutions right so you're adding in that

1:47:22test

1:47:23so we're not just pushing straight into

1:47:25production and risking ruining any sort

1:47:27of analytics in your environment now

1:47:29normally this includes three stages

1:47:32development test staging test St staging

1:47:35and production sometimes you had a

1:47:36fourth one in there called like pre-prod

1:47:38it just depends on your strategy now

1:47:40deployment using deployment pipelines

1:47:42involves copying items from workspace to

1:47:44another and by default in fabric that's

1:47:46a manual process and deployment rules

1:47:49can be implemented to change things like

1:47:51the default Lakehouse for a notebook or

1:47:54the data sets and data sources that your

1:47:57semantic model reads from for example

1:48:00now as well as the deployment pipelines

1:48:02functionality there's a number of other

1:48:04ways to manage deployment in fabric now

1:48:07that could be done through branching in

1:48:10Azure devops you can also use in Azure

1:48:12devops functionality called pipelines

1:48:15right so you create a yaml template

1:48:17we're not going to go through what that

1:48:18looks like but just for the exam know

1:48:20that it is possible and you can also do

1:48:22deployment of semantic models via the

1:48:25xmla endpoint we're going to be looking

1:48:27at that in a bit more detail a bit later

1:48:29on in this lesson okay so as I mentioned

1:48:31there are a few other ways that we can

1:48:34deploy things in fabric now a really

1:48:36good resource here to check out is Kevin

1:48:39Chance's blog and I'll leave a link to

1:48:40that in the school Community

1:48:42specifically there's two blog posts here

1:48:44around cicd for the data warehouse now

1:48:47the data warehouse is not natively

1:48:48supported yet in fabric but you can set

1:48:51up a SE project and then use things like

1:48:55yaml pipelines and Azure devops so if

1:48:58you want to understand other ways that

1:48:59we can deploy items in fabric I

1:49:02recommend checking out these blogs here

1:49:04now if we're looking specifically at the

1:49:06semantic model and how we can deploy

1:49:08these we also have the option of the

1:49:10xmla endpoint so broadly speaking

1:49:13there's two ways to create and manage

1:49:16semantic models number one is to create

1:49:18and manage your semantic model within

1:49:20your workspace so create it from a lake

1:49:23housee or a data warehouse within your

1:49:25workspace but method two is to create

1:49:28your semantic model in a third party

1:49:30tool for example TBL editor and then you

1:49:32can deploy that via What's called the

1:49:35xmla endpoint into your workspace now to

1:49:38grab that xmla end point you need to go

1:49:41to your workspace settings and then you

1:49:42get this palbi URL kind of thing and you

1:49:46can connect that to TBL editor or SS SMS

1:49:50or DAC Studio as well to deploy your

1:49:54models into fabric using this xmla

1:49:58endpoint okay so the next section of the

1:50:00study guide that we're going to look at

1:50:01is these three different file types or

1:50:03three different items that we can create

1:50:06in pobi and Microsoft fabric that you

1:50:08need to know for the exam number one is

1:50:10the powerbi template file then we got

1:50:12the powerbi data source file and then

1:50:15shared semantic models so we're going to

1:50:17look at each of these in a bit more

1:50:18detail starting with the powerbi

1:50:20template file okay so the powerbi

1:50:22template file is a reusable asset that

1:50:25can improve the efficiency and

1:50:27consistency when you're creating power

1:50:29reports now you can easily save a PBX

1:50:32file as a powerbi template file and then

1:50:34you can use that to generate new reports

1:50:37in a given style or with a given layout

1:50:40already specified in that template file

1:50:43if you got parameters in that powerbi

1:50:45file then when you create a new project

1:50:48or a new file from the template you'll

1:50:50be asked to set your parameter

1:50:53in there if there's any parameters in

1:50:55that report let's just have a quick look

1:50:56at how you can create a PBI template

1:50:59file from a pbix file okay so here we

1:51:03are back in pobi desktop and here we've

1:51:06got a pobi project that we're working on

1:51:08our sales and marketing analysis report

1:51:10say we want to create multiple versions

1:51:13of this report all with the same format

1:51:16the same structure and the same layout

1:51:18here so the same pages in this report

1:51:21what we can do simply is go to file save

1:51:23as and then at the bottom here rather

1:51:25than saving as a PBX or a pbip we're

1:51:29going to save it as a PB a powerbi

1:51:31template file so we just select a folder

1:51:33to save it to and then click on Save we

1:51:36can give it a description sales template

1:51:39okay and that's going to save your

1:51:40report and the next time you go to

1:51:42create a new report we can then import

1:51:44that template and that can be your

1:51:45starting point for your new report so

1:51:48the powerbi data source file or PB IDs

1:51:52is a another file type that basically

1:51:54represents a data source file so it's a

1:51:57reusable asset and it can help us

1:51:59quickly transfer all of the data

1:52:01connections that you create in one

1:52:02powerbi file transfer them over to

1:52:05another report okay in this section of

1:52:07the exam we're going to be looking at

1:52:10impact analysis and the lineage tool in

1:52:13Microsoft fabric so here we have a

1:52:15workspace and it's got quite a lot of

1:52:17different items and if we click on this

1:52:20button here in the top right hand corner

1:52:21we can change from this list View to a

1:52:24lineage View and this is what this looks

1:52:26like here we can see all of our fabric

1:52:28items within that workspace and we can

1:52:31have a look at how data is Flowing from

1:52:34Source through to different lake houses

1:52:37here we've got some notebooks here the

1:52:39semantic model that's being created from

1:52:41that Lakehouse now for each of the main

1:52:43items in our lineage view we've got this

1:52:46button here which is the impact analysis

1:52:48button so for this Bronze Lake housee we

1:52:51can click on the impact analysis and we

1:52:54can look at what the downstream items of

1:52:58that Lakehouse are so if you're going to

1:53:00be making some changes to that Lake

1:53:02housee that might break some of the

1:53:05notebooks or the semantic model for

1:53:07example the impact analysis basically

1:53:09allows us to see what the downstream

1:53:11items are now in our example we've only

1:53:13got three Downstream items but you could

1:53:15have 10 or 20s or hundreds of different

1:53:18Downstream items so it's really

1:53:20important to know if you make a change

1:53:22Upstream what the impact is going to be

1:53:25another piece of functionality that you

1:53:26have here is to notify people so you can

1:53:29notify people that are listening for

1:53:32notifications on any of those Downstream

1:53:34items and you can let them know before

1:53:36you make a change so we're going to add

1:53:37a new column to this table in the lake

1:53:40house for example we're going to change

1:53:41add some new tables FYI just to make

1:53:44them aware of the changes that you're

1:53:46going to make before you make them so

1:53:48that's basically all the lineage tool

1:53:50and the impact analysis tool tool in

1:53:52fabric look like at the moment again I

1:53:54think they're adding a lot more

1:53:55functionality to these features in the

1:53:58future that's basically all you can do

1:53:59with this at the moment just bearing

1:54:01that in mind for the exam because you

1:54:02might get asked questions around

1:54:04notification about how do you find

1:54:07Downstream items of a particular fabric

1:54:09item so just have a quick look at some

1:54:11of your workspaces perform some impact

1:54:13analysis just by looking at the

1:54:16downstream items for the exam Okay so

1:54:18we've covered a lot of ground there in

1:54:21that video Let's just round the video up

1:54:24with some practice questions to test

1:54:26your knowledge of this section of the

1:54:28exam question one you are looking to

1:54:30improve the efficiency and consistency

1:54:32of your powerbi development team you

1:54:34want each report created by the team to

1:54:37always consist of three pages the intro

1:54:40the context and the analysis the reports

1:54:42should always align to the company

1:54:44branding which of the following would

1:54:46help you achieve this number one would

1:54:48you create a pbip file two create a PBX

1:54:52file three create a pbit file four

1:54:55create a PB IDs file or number five

1:54:58using a Json custom report theme pause

1:55:02the video here have a think and I'll

1:55:03reveal the answer shortly so the answer

1:55:06here is the pbit file because that is

1:55:08the powerbi template file now the

1:55:11important part of this question was you

1:55:13want each report created by the team to

1:55:16consist of the following pages right so

1:55:18you might have thought oh we're talking

1:55:20about branding talking about following a

1:55:21star guide e might be an answer there

1:55:24using the Json report theme that we

1:55:26looked at in the first video in this

1:55:28series but if we want to include Pages

1:55:31kind of template pages and that's going

1:55:33to be the powerbi template file PBP the

1:55:36project file the PBX that's not going to

1:55:38do the job and the PBS is just for data

1:55:41sources not for giving us a template

1:55:44structure to follow for our report which

1:55:46of the following most accurately

1:55:48describes git a to use Git You must be

1:55:51using GitHub B git is a Microsoft

1:55:54product for tracking changes made to

1:55:55fabric items C git is an open-source

1:55:59version control system that tracks

1:56:01changes in any set of text based files D

1:56:04git allows us to add deployment rules to

1:56:07fabric deployment pipelines so the

1:56:09answer here is C git is an open-source

1:56:12version control system that tracks

1:56:14changes in any set of text based files

1:56:17right so git the underlying technology

1:56:20is open source and it's used in a wide

1:56:23variety of Version Control Systems one

1:56:25of which is azure devops the repo is in

1:56:28there GitHub is another example bit

1:56:30bucket is another example there's lots

1:56:32of these different ones so you don't

1:56:33have to be using GitHub to be using git

1:56:36git is not a Microsoft product for using

1:56:39specifically in fabric it's kind of like

1:56:40a generic tool that's used right across

1:56:42software development industry and it's

1:56:44nothing to do really with deployment

1:56:46rules in fabric deployment pipelines

1:56:49although you can set up git to be the

1:56:52the version control system in different

1:56:54stages different workspaces in your

1:56:57deployment pipelines but it's not really

1:56:59best describing git question three in an

1:57:01Azure devops repo the main branch is

1:57:04protected so it needs approval before

1:57:06any changes are merged into it the repo

1:57:08contains one pbip file you have to

1:57:11update the title in the report merging

1:57:13these changes to the main branch in

1:57:15which order should you carry out the

1:57:17following tasks to achieve this now this

1:57:19is an unordered list your job is to

1:57:22order this list so the tasks here are

1:57:25commit and push the feature Branch wait

1:57:27for approval then merge into the main

1:57:29branch clone the repository to your

1:57:31local machine make the required changes

1:57:33to the report check out a new feature

1:57:35Branch from the main branch and then

1:57:38open a pull request in Azure repos or

1:57:41Azure devops so order this list and then

1:57:43we'll show you the correct order shortly

1:57:46so the correct ordering here starts with

1:57:48cloning the repository to your local

1:57:50machine so if you want to make any

1:57:51changes to report you need that

1:57:53repository to be in your local

1:57:54environment first then you're going to

1:57:56check out a branch because the main

1:57:59branch is protected so we can't edit

1:58:01that directly need to check out a new

1:58:03Branch from the main branch then we're

1:58:04going to make the changes to the report

1:58:06these are going to be tracked via our

1:58:08Version Control System we're going to

1:58:10commit and push those changes that we

1:58:12made on that feature Branch we're going

1:58:14to open a poll request in Azure repos or

1:58:17as devops and we're going to wait for

1:58:18approval and then merge it into the main

1:58:21branch so that is the correct order here

1:58:23question four you want to deploy a

1:58:25semantic model using the xmla endpoint

1:58:28where can you find the xmla endpoint to

1:58:30set up a connection with a third party

1:58:32tool is it a go to the workspace

1:58:34settings for the workspace you want to

1:58:36deploy your model to B go to the fabric

1:58:39admin portal and then capacity settings

1:58:41C in your workspace find your semantic

1:58:43model then click on the settings to get

1:58:45the xmla endpoint address in the Azure

1:58:48portal in your fabric capacity go to the

1:58:50xmla endpoint connect ction string

1:58:53settings so the answer here is a go to

1:58:55the workspace settings for your

1:58:57workspace you want to deploy your model

1:59:00to when we're creating that xmla

1:59:02endpoint connection that's going to link

1:59:04to our workspace in fabric that's where

1:59:06you're going to go to get the address

1:59:09not in capacity settings not in the

1:59:11Azure portal and see in your workspace

1:59:13find your sematic model where the

1:59:15sematic model doesn't actually exist in

1:59:16that workspace because we want to deploy

1:59:18it into there and even if it did exist

1:59:20it doesn't have the xmla end point in

1:59:23the settings anyway so congratulations

1:59:25that is the first section of the exam

1:59:28study guide complete next up we move

1:59:31into the biggest section of the exam

1:59:33which is worth 40 to 45% Camila has been

1:59:36seriously impressed with your skills and

1:59:37knowledge so far in the next lesson

1:59:39we'll be looking at how to create

1:59:41objects in a Lakehouse and a data

1:59:43warehouse so click here for the next

1:59:46lesson in this series hey everyone

Getting data into Fabric

1:59:49welcome back to the channel and this is

1:59:51going to be video five in our dp600 exam

1:59:54preparation course and in this video the

1:59:57focus is going to be on getting data

1:59:59into fabric now this is the first part

2:00:02of the second section in the exam which

2:00:04is all around preparing and serving data

2:00:06and it's worth 40 to 45% of the exam so

2:00:10there's going to be a lot of questions

2:00:11in this so let's get into it in the

2:00:13video we're going to be covering

2:00:14ingesting data using data pipeline data

2:00:16flow notebooks copying data which is

2:00:19basically the same thing but they've

2:00:20included it twice in the study choosing

2:00:22an appropriate method for copying data

2:00:24so not just understanding what the tools

2:00:27available to us but making decisions

2:00:28about the best method given a particular

2:00:31problem or a particular scenario and

2:00:32also creating and managing shortcuts now

2:00:34you'll notice that these are slightly

2:00:36deviation from the study guide what I've

2:00:38done is I've had a look at the whole of

2:00:40section two and I've changed some of the

2:00:42order of things don't worry we'll be

2:00:43going through all of the skills in the

2:00:45study guide but we're just going to be

2:00:46going through them in a slightly

2:00:47different order so these are the four

2:00:49that we're going to focus on today now I

2:00:50have released a video in the last month

2:00:53this one here data pipelines versus data

2:00:54flow shortcuts notebooks and this is a

2:00:56more comprehensive video so I definitely

2:00:59recommend watching this either before or

2:01:01after this video so as such for this

2:01:03video we won't be starting from zero

2:01:06more revising what we went through in

2:01:08the last video updating it a little bit

2:01:10because there have been some changes

2:01:11revising some of the Core Concepts and

2:01:14the distinctions between these tools

2:01:16that you need to know for the exam as

2:01:17ever there will be five sample questions

2:01:19at the end of this video and we've got

2:01:21the school Community with quite

2:01:23extensive notes and links to further

2:01:26resources if you want to go into a bit

2:01:28more detail about any of the topics that

2:01:29we cover in this lesson so let's start

2:01:32with a bit of a framing of all the

2:01:33different options that we have available

2:01:35to us when we're talking about ingesting

2:01:37data and getting data into fabric so one

2:01:40group of tools is the data ingestion the

2:01:43El or the ETL tools right so extract

2:01:46transform and load these are going to be

2:01:48copying data from external systems

2:01:51bringing them into Fabric and saving

2:01:52them in some sort of data store in

2:01:54fabric so we've got the data flow data

2:01:56pipeline the notebook and the event

2:01:58stream now the event stream I don't

2:01:59think is actually covered in the DP 600

2:02:01exam so we're not going to be talking

2:02:03about that much today we also have

2:02:04shortcuts so that's another method that

2:02:06we need to be aware of that we're going

2:02:08to go through in this video and that

2:02:09involves bringing data in from Amazon S3

2:02:12ADLs Gen 2 the data verse and also

2:02:15Google Cloud Storage as well we've also

2:02:17got internal shortcuts so connecting our

2:02:20different data sets Within fabric so

2:02:23creating references to other Lakehouse

2:02:25tables from a lake house for example

2:02:27then at the bottom we've got a new

2:02:29feature that's in preview at the moment

2:02:30which is database mirroring so you can

2:02:32create a mirror of your database within

2:02:35fabric from either snowflake Cosmos DB

2:02:37or as a SQL at the moment and again I

2:02:40don't think this is actually part of the

2:02:41exam study guide currently so we're not

2:02:44going to be talking about it in this

2:02:45section of the study guide if you don't

2:02:47want to learn a bit more about databas

2:02:48Maring I did mention it in that video

2:02:51that I mentioned previously so in this

2:02:53video we're just going to be going

2:02:54through the data flow the data pipeline

2:02:56the notebook and shortcuts in a bit more

2:02:59detail cuz these are the things that you

2:03:00need to study for the exam starting with

2:03:03the data flow so as you know the data

2:03:05flow comes with about 150 connectors to

2:03:08external systems and you can bring that

2:03:10data in using power query like that no

2:03:13low code interface and you can do data

2:03:15transformation on that data before

2:03:17writing it to one of the fabric data

2:03:19stores so when should you use it and

2:03:21maybe when should you not use it well if

2:03:23you want to use any of those 150

2:03:24connectors then it's definitely good to

2:03:27use data flows it's a no and low code

2:03:29solution so it's quite maintainable if

2:03:32you have people who perhaps don't have

2:03:34SQL or python skills in your

2:03:36organization this is a good method and

2:03:38as I mentioned we can do extract

2:03:40transform and loading all in one tool

2:03:42now the data flow is one of the tools

2:03:44that we can use to access on premise

2:03:46data via the on- premise data Gateway

2:03:49and it's also quite useful when you want

2:03:51to ingest more than one data set and

2:03:54maybe combine them in the same data flow

2:03:56although if you're following a kind of

2:03:58Medallion architecture maybe that's not

2:04:00what you want to do but it's just an

2:04:02option that is possible with the data

2:04:03flow it's also the only tool that we can

2:04:05use to upload raw local files so maybe

2:04:09you have a CSV file you know this file

2:04:12won't ever have to really be updated you

2:04:14just want to get some data into fabric

2:04:17you can do that in the data flow so

2:04:19maybe when you shouldn't be using this

2:04:21well traditionally they've been

2:04:23struggling with large data sets but

2:04:25there's been a number of features

2:04:26released to try and speed that up so

2:04:29recently they released fast copy so

2:04:31that's one feature that they released to

2:04:32try and speed up data flows now the data

2:04:35flow uses the same backend

2:04:37infrastructure that the data pipeline

2:04:38uses in the copy data activity so the

2:04:41performance between these two should now

2:04:42be a lot more similar another aspect

2:04:45with power queries it's quite difficult

2:04:46to implement data validation right so

2:04:49because all of our logic is being locked

2:04:52up in those power query routines it's

2:04:54difficult to kind of validate those

2:04:55steps as you're going through them so

2:04:57that's maybe one reason why you might

2:04:58want to go for the elt routine rather

2:05:01than the ETL extract load transform

2:05:04rather than extract transform load now

2:05:07it's also the only tool out of the data

2:05:10flow the data Pipeline and the notebook

2:05:12where you can't pass in external

2:05:14parameters so it's difficult to build

2:05:16kind of metadata driven architectures

2:05:18with the data flow at the moment now

2:05:20there is a kind way around this where

2:05:22you can create another query within your

2:05:25power query engine of the the data set

2:05:28or the the metadata that you might want

2:05:30to use to parameterize a solution so

2:05:33there is kind of a workaround but

2:05:35there's no native functionality for

2:05:36passing parameters into a data flow from

2:05:39a data pipeline for example next up

2:05:41we're going to be talking about

2:05:42ingesting data with a data Pipeline and

2:05:45the data pipeline is primarily an

2:05:46orchestration tool right but it can also

2:05:49be used to get data into fabric using

2:05:52the copy data activity and also some

2:05:54other activities as well you can use for

2:05:56this but the copy data activity is the

2:05:58main one now one of the main pros of a

2:05:59data pipeline is that where it performs

2:06:01well on large data sets and now as I

2:06:04mentioned before the data flow now has

2:06:06fast copy so the for performance should

2:06:09be comparable between the two it has

2:06:11many connections to cloud data sources

2:06:13especially in Azure so it's good for

2:06:15using if you've got data in Azure for

2:06:17example we need some sort of control

2:06:19flow logic so maybe looping through

2:06:21through different tables for example

2:06:23there's a lot of functionality for

2:06:25building metadata driven or

2:06:27parameterized data ingestion methods so

2:06:31that's something to bear in mind with

2:06:32the data Pipeline and as well as the

2:06:33copy data activity you can also use it

2:06:36to trigger a wide variety of other

2:06:38actions in fabric for example a stored

2:06:41procedure and a stored procedure can

2:06:43also be used for ingesting data into

2:06:47fabric for example the copy into

2:06:49statement in t s call can be used to

2:06:51ingest a CSV file into a data warehouse

2:06:55directly for example now some of the

2:06:56cons in the data pipeline where it can't

2:06:58do the transform piece natively so

2:07:01there's no real data transformation

2:07:03activities but what you can do is embed

2:07:05notebooks and data flows into a data

2:07:08pipeline if you want to do that

2:07:09transformation has no ability to upload

2:07:12local files so that's not possible with

2:07:14a data pipeline at the moment you do

2:07:15need to be careful with any sort of

2:07:17crossw workpace data pipeline usage now

2:07:20we did submit an idea on ideas. fabric.

2:07:24microsoft.com and it is actually going

2:07:26to be planned now so that's good news so

2:07:28they've planned this feature we don't

2:07:29know when it's going to be released yet

2:07:31but they are working on support for

2:07:34crossw workpace data pipelines and what

2:07:37we mean by that is maybe you want to

2:07:38bring data in from a data source and you

2:07:42want to write it to a destination that's

2:07:45in a different workspace to the data

2:07:47pipeline so that currently isn't

2:07:48possible but they're working on that

2:07:50feature as we speak next up we've got

2:07:52ingesting data with a notebook and a

2:07:55notebook is just a general purpose

2:07:57coding notebook which can be used to

2:07:59well for a wide variety of things but

2:08:01one of the things is to bring data into

2:08:03fabric now we can do this via either

2:08:06connecting to apis using something like

2:08:09the requests library in python or

2:08:11something similar or by using client

2:08:14python libraries so for example if you

2:08:16have a thirdparty SAS product like

2:08:19HubSpot a lot of the big ones have

2:08:21python libraries that you can use to

2:08:23bring data in as well as a lot of the

2:08:25Azure tooling so Azure data Lakes for

2:08:28example have a python client that you

2:08:30can use to bring data in that's another

2:08:33option with the notebook so some of the

2:08:34pros here is again it's really good for

2:08:37extraction from apis because you know if

2:08:40you know how to code with python you can

2:08:42do quite customized logic around things

2:08:45like authentication imagination and that

2:08:47becomes really simple in a notebook if

2:08:49you want to be using any of those client

2:08:51libraries so anything from Azure is also

2:08:54really good or HubSpot as we mentioned

2:08:55previously now it's good for code reuse

2:08:58so a notebook can be parameterized and

2:09:01then used in lots of different

2:09:02situations you can also embed data

2:09:04validation and data quality testing into

2:09:07the incoming data so we done quite a lot

2:09:09on this channel around data quality and

2:09:11validating incoming data so that becomes

2:09:13a lot easier in a notebook and in terms

2:09:16of the performance well a few people

2:09:19have been testing the different

2:09:20performance of these different meth

2:09:21methods and the notebook always comes

2:09:23out on top really in terms of speed and

2:09:26therefore also capacity usage so if

2:09:29you're really sensitive around the

2:09:31amount of capacity uh units that you're

2:09:33using then notebook is going to be the

2:09:35most efficient and I'll leave a link to

2:09:37a bit of an analysis done on the Lucid

2:09:40bi blog and it goes through some of the

2:09:41investigation work that they've been

2:09:43doing there it's a really good blog if

2:09:44you want to learn more about that now

2:09:46some of the cons well when you don't

2:09:48have a python capability in your

2:09:49organization that might sound a bit of

2:09:51an obvious one but if you don't have a

2:09:53team to write and then support these ETL

2:09:57notebooks then it's not going to be a

2:09:58good choice for you and secondly one of

2:10:00the limitations with the notebook is you

2:10:03can't actually currently write into a

2:10:05data warehouse so if that's your

2:10:07destination then you're going to be wan

2:10:10to using other tools for data ingestion

2:10:12not the notebook okay the other method

2:10:14that we can use to bring data into

2:10:16Fabric or at least make data accessible

2:10:19from within fabric is the shortcut and

2:10:21we've talked quite a lot about shortcuts

2:10:23on this channel so far so what we're

2:10:25going to be doing is just a bit of a

2:10:26review of what's possible with a

2:10:28shortcut and some of the things that you

2:10:29need to bear in mind for the exam so the

2:10:31first thing is kind of a bit of an

2:10:32overview really a shortcut enables you

2:10:34to create a live link to data stored in

2:10:37another part of fabric which is an

2:10:40internal shortcut or in the following

2:10:42external storage locations so ADLs gen2

2:10:45Azure data Lake storage Amazon S3 other

2:10:48services that use Amazon S3 for storage

2:10:51for example Cloud flare buckets also use

2:10:54Amazon S3 and that's quite a big section

2:10:57of the market there's a lot of tools

2:10:58that use Amazon S3 for their storage and

2:11:01that is now opened up for shortcuts

2:11:03Google Cloud Storage is another one and

2:11:05also tables in the data verse now a

2:11:07shortcut can be set up for individual

2:11:09files but also for a folder so if you

2:11:12set up a shortcut to a folder it's

2:11:14basically going to bring in all of the

2:11:16files that are in that folder and sync

2:11:19them now one thing to be careful of is

2:11:20the cross region egress fees so if your

2:11:23fabric capacity is in UK South Region

2:11:26for example but your ADLs storage

2:11:29account is in West us then you're going

2:11:31to be charged cross region ESS fees by

2:11:35Azure basically and that's 1 cent per

2:11:38gigabyte of data that's transferred now

2:11:40you can also create shortcuts now via

2:11:42the fabric rest API as well so if you

2:11:45want to do some sort of programmatic

2:11:46creation of shortcuts maybe you want to

2:11:49do hundreds of tables shortcuts cutting

2:11:51in one go then that's an option for you

2:11:53there now something that might come up

2:11:54in the exam is around permissions for

2:11:57shortcuts so what I've done is I've just

2:11:59copied the documentation piece here and

2:12:01I'll link to this in the school

2:12:02Community as always so what this table

2:12:04is showing is the shortcut related

2:12:06permissions for each workspace role so

2:12:10starting at the top there we've got

2:12:11creating a shortcut well to be able to

2:12:13create a shortcut the user needs write

2:12:15permission in the place that they're

2:12:17creating the shortcut and also read

2:12:19permission of the file or the that

2:12:21they're shortcutting too okay so these

2:12:23are the two permissions that you need

2:12:24side by side to create a new shortcut

2:12:27secondly if you just want to read the

2:12:28file contents of a shortcut you're going

2:12:31to need read permissions in both the

2:12:33place where the shortcut lives but also

2:12:35where the reference file is living as

2:12:38well so you need at least read

2:12:40permission in both of these locations if

2:12:41we want to write new files or new data

2:12:43to a shortcut Target location you're

2:12:45going to need write permissions for both

2:12:48of those locations both where you're

2:12:50writing the shortcut data to and where

2:12:52that shortcut data is being read into as

2:12:54well so let's just talk now about

2:12:56deciding when to use which method

2:12:59because and we did touch on this in the

2:13:01first lesson in this series we talked

2:13:03about some of the deciding factors when

2:13:05we're planning out our fabric

2:13:06implementation right so talking about

2:13:08the storage where is it stored and also

2:13:11what skills exist in the team and that's

2:13:13a really good start but there's also

2:13:14some other factors that we need to bear

2:13:16in mind when we're thinking about

2:13:17choosing a data ingestion method so if

2:13:20you have a requirements for real time or

2:13:23near realtime data you're going to be

2:13:25wanted to prefer the options like the

2:13:28shortcut if it's a file or folders or

2:13:30database mirroring if it's in a table in

2:13:33any of the three database types that are

2:13:35supported by database mirroring because

2:13:37these are live links to those locations

2:13:40so when you query it it's going to have

2:13:42near realtime data coming back we talked

2:13:44about these skills in the team so if

2:13:45you've got predominantly no and low code

2:13:47users then the data flow in the day

2:13:49pipeline if you've got SQL based then

2:13:52you can use the data Pipeline and stored

2:13:55procedure activity or script activity

2:13:57and we'll look at some of that in the

2:13:59next lesson we'll talking more about

2:14:00store procedures and that kind of thing

2:14:02but it is possible to use for example

2:14:04the copy into statement for ingesting

2:14:07data from files like parket files or

2:14:10from CSV files into a data warehouse and

2:14:13if you got python or Scala skills then

2:14:16you can be using notebooks the next

2:14:18thing is around crossworks bace

2:14:19limitations as we mentioned before the

2:14:21data pipeline must be in the same

2:14:24workspace as your destination store when

2:14:27we're talking about data ingestion right

2:14:29so it needs to be in the same Works spay

2:14:30and the other two methods the data flow

2:14:33and the notebook they don't have those

2:14:35limitations just to be clear on that

2:14:37next talk about the scalability and the

2:14:39size of your data and cost and capacity

2:14:41usage and all these things are pretty

2:14:44tightly linked right so in general The

2:14:47Notebook from quite a few people's

2:14:50analyses that I've seen is the most

2:14:52efficient method I'm not sure if you'll

2:14:53be tested on this in the exam but just

2:14:55something to bear in mind for when

2:14:56you're actually designing Solutions okay

2:14:58so now let's test some of our knowledge

2:15:01in this part of the exam when we're

2:15:03talking about getting data into fabric

2:15:05question one you're trying to create a

2:15:07shortcut to a folder of CSV files in

2:15:10Azure data L Storage Gen 2 which of the

2:15:12following is a valid connection string

2:15:14you can connect to A B C or D now I'll

2:15:18pause the video here have a bit of a

2:15:19think and I'll reveal the answer to you

2:15:21shortly so the answer here is D so in

2:15:24Azure there's a number of different end

2:15:25points that we can connect to for a

2:15:27storage account and the one that we need

2:15:29for a shortcut is the DFS the

2:15:32distributed file system path which is D

2:15:35the database. windows.net well it's not

2:15:37a database so it's not going to be A and

2:15:39B and C are both end points of a storage

2:15:42account but it's not the ones that you

2:15:43need to create a shortcut question two

2:15:45you're implementing the first stage in a

2:15:47medallion architecture your goal is to

2:15:49retrieve data from a rest API using a

2:15:52get request and you're going to save the

2:15:53raw Json response in the files area of a

2:15:56bronze Lakehouse which of the following

2:15:57methods can you use to achieve this now

2:16:00there's three correct answers here is it

2:16:02a the data pipeline copy data activity B

2:16:04the data flow with the web API connector

2:16:06C the event stream B the data pipeline

2:16:08web activity or E the fabric notebook so

2:16:11there's three methods here a the data

2:16:13pipeline copy data activity D the data

2:16:16pipeline web activity and E the fabric

2:16:19notebook now one of the important an

2:16:21parts of this question is that we want

2:16:22to save the raw Json response in the

2:16:25files area of a bronze Lakehouse so

2:16:29we're not going to be ingesting it

2:16:30directly into a table we want to save

2:16:32the raw Json and these three are the

2:16:34only three where that's possible the

2:16:36data flow we have to Output it into a

2:16:38data store so into a lake house table or

2:16:41into a data warehouse table for example

2:16:43so we need to perform some sort of

2:16:44transformation on that Json we can't

2:16:46just write out the raw file same with

2:16:48the event stream and the other three

2:16:49does allow us to do that so the data

2:16:51pipeline you can output in various

2:16:53different file formats and with the

2:16:55fabric notebook you can do that as well

2:16:57question three your goal is to extract a

2:16:59CSV file in Azure blob storage and write

2:17:03it to a fabric data warehouse table

2:17:06which of the following methods can you

2:17:07not use to achieve this data pipeline

2:17:09copy data activity B data flow with a

2:17:12blob storage connector C in a fabric

2:17:14notebook use the Azure blob storage

2:17:16client library for python get the file

2:17:18and write the data into the data we

2:17:21table D use the copy into statement in

2:17:23tsql from within your fabric data

2:17:25warehouse so the answer here is C which

2:17:28of the following methods can you not use

2:17:30to achieve this so you can't as we

2:17:32mentioned in the lesson today you can't

2:17:34use a notebook to write data directly

2:17:38into a fabric data warehouse so that was

2:17:41the clue in this question all of the

2:17:43other three methods so the data pipeline

2:17:45the data flow and the copy into

2:17:47statement in a tsql script they can be

2:17:49used to ingest data into a fabric data

2:17:53wouse table question four workspace a

2:17:55contains lake house a workspace B

2:17:58contains lakeh house B you want to

2:18:00create an internal shortcut in Lakehouse

2:18:03a pointing to a table in lakeh house B

2:18:06what's the minimum level of workspace

2:18:08permissions you need to achieve this a

2:18:10contributor in workspace a and viewer in

2:18:13workspace b b contributor in workspace a

2:18:17and contributor in workspace b c member

2:18:20in workspace a and contributor in

2:18:22workspace b or d viewer in workspace a

2:18:26and viewer in workspace B so the answer

2:18:28here is contributor in workspace a and

2:18:30viewer in workspace B so when we were

2:18:33talking about the workspace roles

2:18:35permission required to create a new

2:18:38shortcut where you need some sort of

2:18:39write permission in the workspace where

2:18:41you're creating the shortcut and read

2:18:44permissions in the lake house that

2:18:45you're actually referencing so the

2:18:47important part of the question here is

2:18:48what's the minimum leval of workspace

2:18:51permissions that you need to achieve

2:18:52this so the others or B and C at least

2:18:55would allow you to do this but it's not

2:18:57the minimum level of permissions that's

2:18:59required and D viewer in workspace a if

2:19:02we're a viewer in workspace a then you

2:19:03won't have the permissions to create the

2:19:05shortcut in Lake housee a number five

2:19:08one of your team is a superstar using

2:19:10mcode for data extraction in which of

2:19:12the following data extraction tools can

2:19:14you write M code is it a the fabric

2:19:16notebook b in a tsql script in the data

2:19:19warehouse C in a data flow or D in a

2:19:23data pipeline mapping data flow so the

2:19:25answer here is C the data flow gen to

2:19:29that's the one with the power query

2:19:31interface and you can write in the

2:19:33advanced editor you can write M query

2:19:35obviously not in the fabric notebook or

2:19:37a tsql script and the data pipeline

2:19:39doesn't actually have a mapping data

2:19:41flow activity so that's a bit of a red

2:19:43herring that's coming from Azure data

2:19:45Factory where that did exist but not in

2:19:47fabric congratulations you've completed

2:19:49the first part in section two of the

2:19:52study guide preparing and serving data

2:19:54in the next lesson we're going to be

2:19:55looking at the fabric Data Warehouse in

2:19:58more detail and we're going to be

2:19:59looking at how you can schedule all of

2:20:01your ETL workloads whether that be a

2:20:04data flow a notebook or a data pipeline

2:20:06so click here to continue your dp600

2:20:09Learning Journey I'll see you there hey

SQL, Data Warehouse and scheduling

2:20:11everyone welcome back to the channel

2:20:12today we're continuing the dp600 exam

2:20:15preparation course and we're going to be

2:20:17looking at SQL the data warehouse and

2:20:20how we can schedule things to run in

2:20:23Microsoft fabric this is video six in

2:20:26our Series so we're nearly halfway we've

2:20:28covered a lot of ground already but

2:20:29there's some way to go still until we

2:20:31get to the end of the course in this

2:20:32video we're going to be covering

2:20:33creating views functions and stored

2:20:36procedures in that data warehouse

2:20:38experience how we can add stored

2:20:40procedures notebooks data flows to a

2:20:43data pipeline then how we can schedule

2:20:45data pipelines and also schedule things

2:20:47like data flows and notebooks and

2:20:50throughout this section of the course

2:20:51obviously we're going to be focusing on

2:20:53what you need to know to prepare for

2:20:55this dp600 exam a lot of these topics

2:20:58can go really deep but we're just going

2:20:59to go through what I think would be

2:21:01sensible to learn for the exam in case

2:21:03they come up now this lesson will be

2:21:05pretty much 100% practical so we're

2:21:07going to be going into fabric having a

2:21:09look and creating all of these things

2:21:11ourselves at the end of the lesson we'll

2:21:13be doing four sample questions and as

2:21:15ever I've got some key points and links

2:21:18to further learning resources if you

2:21:20want to learn more about a particular

2:21:21topic and really brush up on your skills

2:21:23in a particular area for the exam okay

2:21:25so let's start off by looking at

2:21:27creating functions and stored procedures

2:21:30and Views in the data warehouse

2:21:32experience what all these things are

2:21:34when you should be using them how you

2:21:36should be using them and things you need

2:21:38to bear in mind for the exam so we're

2:21:39starting off in a workspace here and

2:21:41I've created some fabric items the most

2:21:44important one being this data warehouse

2:21:46dp600 data warehouse if we open that up

2:21:48and have a quick look around just

2:21:50created some really simple tables we're

2:21:52going to be using one or maybe two of

2:21:53these tables for this tutorial

2:21:56specifically this dbo do employees table

2:21:59so only got four rows but that's good

2:22:00enough just to show you the

2:22:02functionality of a function and a store

2:22:04procedure and a view so what we're going

2:22:06to be doing is we're not actually going

2:22:07to be working in the fabric online

2:22:10experience we're going to be using SQL

2:22:12Server management studio so what we're

2:22:14going to be needing to do is go into the

2:22:15settings of the data warehouse collect

2:22:17this SQL connection string copy that and

2:22:20then then go over to SQL Server

2:22:22management Studio you can download that

2:22:24for free I'll leave a link in the

2:22:25description or in the school community

2:22:27and then we can connect to our fabric

2:22:29Data Warehouse from within SQL Server

2:22:32management studio now I've already set

2:22:34up the connection here but if it's your

2:22:36first time using SS SMS and you haven't

2:22:38connected it before just click on

2:22:39connect to a new database Engine add in

2:22:42your server name here use the Microsoft

2:22:45entra multiactor authentication it will

2:22:48ask you to authenticate with fabric and

2:22:50the on line experience and then you

2:22:51should see your databases listed here

2:22:54we've got this dp600 data warehouse that

2:22:56we were looking at previously and you

2:22:58can see that we've got some tables we've

2:22:59got the employees table the gold table

2:23:01and some other tables in here as well so

2:23:03let's start off just by exploring what

2:23:06we've got here right so let's just do a

2:23:09very simple select star from db.

2:23:11employees and that's just going to

2:23:13return all of that data just so we can

2:23:14have a quick look at it make sure it's

2:23:15all connected correctly yeah so we've

2:23:17got our four rows there three columns

2:23:19employee ID D name and age perfect so

2:23:23say we've got this table here db.

2:23:24employees and for a powerbi report that

2:23:27we're wanting to create we actually want

2:23:29to do a bit of transformation on this

2:23:31table we don't want the raw table just

2:23:34as it looks like here we want to do some

2:23:36sort of transformation it doesn't really

2:23:37matter what that transformation is for

2:23:38this demo maybe we want to do a where

2:23:41statement so where name is like Jack for

2:23:45example this is just going to bring us

2:23:46back all of the employees that are

2:23:49called Jack and we're using this

2:23:51percentage Wild Card operator here and

2:23:53what that means is if the first four

2:23:55characters are j a c k anything after

2:23:58that is going to return in our results

2:24:00set here as you can see we've only got

2:24:02one Jack in the data set so that's

2:24:03coming back correctly now in our powerbi

2:24:05imagine we want to query just this

2:24:08results set but what we can do is we can

2:24:11save this query as a view now this is a

2:24:14bit of a toy example but sometimes you

2:24:16have a view that might have really

2:24:17complex transformation right it's not

2:24:19stored in the database but whenever we

2:24:22query it we want the transformations to

2:24:24be done at query time so all you can do

2:24:27to create a view of this is add in

2:24:29create view give the viewer name dbo do

2:24:32employee get Jack and then use the

2:24:35keyword as right then everything that

2:24:37follows that is going to be part of your

2:24:39view and if we just execute that we can

2:24:42see that that's been successfully

2:24:44completed and then we can just do select

2:24:46star from dbo do view employee get Jack

2:24:50and if we execute that we get the same

2:24:52result as before so now when you're

2:24:54creating your power report you can just

2:24:56query this view rather than the

2:24:58underlying table to get the transformed

2:25:00data back okay so that's the basic

2:25:03Syntax for creating a view what I've

2:25:05done is I've just written down some key

2:25:07points to understand just in general but

2:25:09also for the exam as well so for aiew

2:25:12the transform data is not stored right

2:25:15it's just the transformation logic and

2:25:17the code the SQL code is stored and then

2:25:19every time you query it maybe in a

2:25:21powerr report is going to perform that

2:25:23transformation at query time now you can

2:25:25create a view from another view so you

2:25:28can query another view from for example

2:25:31this view employee get Jack we can query

2:25:33that if we create another view that's

2:25:35possible we do have to be careful with

2:25:37performance sometimes when you string

2:25:40multiple views together and they're

2:25:41doing lots of heavy processing then it

2:25:43can have an impact on performance and it

2:25:46also can be difficult to understand and

2:25:47maintain if you're constantly querying

2:25:49other views from other views that's

2:25:51something to bear in mind now with a

2:25:52view we can't specify any sort of

2:25:54parameters okay it's just simply select

2:25:57statements you're just going to be able

2:25:58to query it like so and whatever the

2:26:00results is you get those back now this

2:26:03is possible with functions and stored

2:26:05procedures as we're going to see in a

2:26:06minute as we mentioned reading a SQL

2:26:08view into semantic model will fall back

2:26:10to direct query mode so direct like mode

2:26:13is not possible and that makes sense

2:26:15right because direct Lake works by

2:26:18reading the underlying Delta tables and

2:26:20in a view those Delta tables don't exist

2:26:23right so that's something to bear in

2:26:24mind with with a view so if that's a

2:26:26view now let's look at a function so say

2:26:30we want to parameterize that view that

2:26:33we had right we want to pass in a

2:26:36parameter maybe first name so that we're

2:26:38not just getting back Jacks we can

2:26:40potentially use it to get a list of

2:26:42employees given any first name let's

2:26:44have a look at what that might look like

2:26:46so this is a declaration of a function

2:26:49now you notice it's similar in some ways

2:26:51to the view but there are some key

2:26:53differences we're going to start by

2:26:54calling create function we give the

2:26:56function a name this one's function

2:26:58employee first name search then we're

2:27:01going to open some brackets and the

2:27:03brackets are really important in a

2:27:04function it's a bit like a function in

2:27:07Python for example you're going to open

2:27:09those brackets and pass in a parameter

2:27:11now the parameters are optional you

2:27:13don't have to pass in a parameter but it

2:27:15is an option we're defining our

2:27:17parameter with this at symbol so at

2:27:19first name and we give it a data type

2:27:21just for our char2 and this is a default

2:27:23value this is saying we have one

2:27:25parameter and the parameter name is

2:27:26first name the parameter type is far

2:27:29Char 20 and the default value is an

2:27:31empty string now the next really

2:27:33important piece of syntax here is this

2:27:35returns table so if you've used any sort

2:27:37of SQL functions in other flavors of SQL

2:27:40you know there's a few different types

2:27:42of functions now in fabric we're talking

2:27:45about table functions which means that

2:27:47it returns a table right we're not

2:27:49talking about Scala functions because

2:27:51that's currently not possible in

2:27:53Microsoft fabric version of tsql We're

2:27:56going to be calling a function and it's

2:27:57always going to return a table so in our

2:27:59syntax we have to Define that right we

2:28:00say returning a table as return Open

2:28:04brackets and then we can pass in

2:28:07whatever we want to declare in our

2:28:09select statement now this is very

2:28:11similar to our view declaration but here

2:28:13we're making it Dynamic right we're

2:28:15parameterizing it and we're using our

2:28:17first name parameter so let's just run

2:28:19this okay so so we've created our

2:28:20function let's just have a look at what

2:28:22that looks like so if we go to

2:28:24programmability in SQL Server management

2:28:26studio and in our functions you can see

2:28:28that in SQL Server management Studio

2:28:30they do actually make the

2:28:32differentiation between scalar valued

2:28:34functions and table valued functions so

2:28:36as we know R1 is going to be in this

2:28:38table valued functions here it is here

2:28:40dbo function employee first name search

2:28:43and we can call it like this select star

2:28:46from dbo our function name Open brackets

2:28:50with the parameter so if we call this

2:28:53execute and you see it Returns the same

2:28:56as what we had before but is

2:28:57parameterized so now instead of just

2:28:59Jack we've created a bit of a

2:29:01parameterized function here so we can

2:29:03call Jack we can also call Sarah you

2:29:06know whatever that needs to be we're

2:29:08just encapsulating all of that logic

2:29:10into a function and we parameterized it

2:29:12so that you can use it in multiple

2:29:14different places in your tsql code base

2:29:16and it just packages up that logic into

2:29:19a nice reusable function fun so what are

2:29:20some of the key points with a function

2:29:22as I mentioned it's useful for packaging

2:29:24up logic that you might want to reuse in

2:29:27different places and it can be

2:29:28parameterized now with a function you

2:29:30can only do select statements so you

2:29:31can't actually update any rows you can't

2:29:34do any sort of insert statements it's

2:29:36only ddl flavors of SQL only select

2:29:40statements so to call a function we can

2:29:42only really use a select statement like

2:29:45so or we can call it from within a

2:29:47stored procedure or from within a view

2:29:50which is kind of underneath this ddl as

2:29:52well if you've got a ddl view and by ddl

2:29:55I just mean select statements basically

2:29:57it can't be orchestrated directly right

2:30:00so we can't use it in a data pipeline

2:30:03directly although we can embed it in a

2:30:05stored procedure as we mentioned

2:30:07previously it allows one or more

2:30:08parameters specifically input parameters

2:30:11and the output type should always be a

2:30:13table okay so in fabric we're only going

2:30:16to using table valued functions and it's

2:30:18always going to return a table okay so

2:30:20that's the function now let's move on to

2:30:22the stored procedure so this is a very

2:30:26basic definition of a stored procedure

2:30:28we've got create procedure we give the

2:30:31procedure a name as and here we're just

2:30:32doing a select statement right and all

2:30:34this is going to do there's no

2:30:36parameters you'll notice when we execute

2:30:38a stored procedure we're just going to

2:30:39use this exec which is execute so that's

2:30:41one of the differences between a stored

2:30:43procedure and a function and a view the

2:30:45way in which we call it right so at the

2:30:47moment we're not really doing much with

2:30:48this store procedure it's just the most

2:30:50basic store procedure possible we're

2:30:52just passing in a select statement and

2:30:54we're executing it and we're getting

2:30:55back the results set the full results

2:30:57set now let's step it Upp a gear and add

2:30:59in a parameter so we've got a very

2:31:01similar thing to what we were looking at

2:31:03before create store procedure now we've

2:31:05got dbo Spore employee uncore get by

2:31:09first name and we're passing in we're

2:31:12declaring a parameter the parameter is

2:31:14called first name again it's going to be

2:31:16of type varar 20 underneath that we're

2:31:19declaring what what we actually want to

2:31:20do in this stored procedure so we want

2:31:22to select all of the employees where the

2:31:25name is like first name which is going

2:31:27to be passed in as the parameter and

2:31:29we're adding this kind of wild card

2:31:31operator onto the end so that anything

2:31:33that goes after that whatever the

2:31:34surname we're going to return that as

2:31:37well so let's just execute this to

2:31:39create the stored procedure ah yeah so

2:31:40let's just drop it drop procedure if

2:31:43exists dbo dot okay so now we've dropped

2:31:45it we can recreate it again just to show

2:31:47that it works completed successfully and

2:31:49then if we want to execute that again

2:31:51we're going to use the exec command pass

2:31:53in the name of our store procedure and

2:31:56our parameter which is Jack if we

2:31:57execute that we get this so that's the

2:31:59basic definition of a stored procedure

2:32:03what it looks like now let's think about

2:32:05some of the key points here so the store

2:32:06procedure we have the ability again to

2:32:08Define input parameters but also output

2:32:11parameters as well I didn't show you

2:32:12that in this lesson maybe that's one for

2:32:15another lesson but we can actually

2:32:16Define output parameters which are

2:32:18useful in the data Pipelines scenario

2:32:20maybe we'll look at that another time on

2:32:22the channel probably all you need to

2:32:24know for the exam is that it is possible

2:32:25now they're called using this EXA

2:32:28Command right it's not as part of a

2:32:31select statement you can call other

2:32:33store procedures from a store procedure

2:32:35so you can create three stored

2:32:37procedures and then create kind of like

2:32:39a master stored procedure that calls

2:32:41each of these other stored procedures in

2:32:43series now one of the key use cases of a

2:32:45store procedure is to give people access

2:32:49just to the stored procedure and not to

2:32:51the underlying data so you can use it as

2:32:54a security mechanism just by giving

2:32:56people access to the stored procedure

2:32:58and you know they're only getting access

2:33:00to the results of that store procedure

2:33:02and not any of the underlying data now

2:33:04one of the key points in fabric to

2:33:06understand is that the stored procedure

2:33:09is the well it's one of the only things

2:33:10that we can embed into a data Pipeline

2:33:13and we can pass the parameters that

2:33:15we've seen here from other notebook

2:33:17activities so that's something to really

2:33:20important to understand with store

2:33:21procedures they become a lot more useful

2:33:24because they can be orchestrated as part

2:33:27of a data pipeline then they become

2:33:29really useful for data transformation

2:33:31and all of these kinds of things in your

2:33:33data warehouse and in your architectures

2:33:35in general now the one I'm missing here

2:33:37is the ability to do inserts updates

2:33:41deletes now this is what's unique to the

2:33:43stored procedure at least in this list

2:33:45here is that here we're just using a

2:33:47select statement but we can use stored

2:33:49proced procedures for updating and

2:33:52inserting the underlying data sets so

2:33:54they can be really powerful tools that

2:33:57we can use for any sorts of data

2:33:59transformation data loading all these

2:34:01kind of things are possible with a store

2:34:03procedure okay so let's just focus now

2:34:05on the stored procedure and we mentioned

2:34:07here it can be embedded in a data

2:34:09pipeline so let's go back into fabric

2:34:11now and have a look at what that looks

2:34:13like okay so here we are in a data

2:34:16Pipeline and you'll notice if we go over

2:34:18to the activities tab there's a lot of

2:34:20different options for activities now the

2:34:22three that they mention for the dp600

2:34:24study guide are the stored procedure

2:34:26Activity The Notebook activity and the

2:34:29data flow activity here we're going to

2:34:31go through each one in a little bit of

2:34:33detail have a look at some of the

2:34:34settings that we can apply when we're

2:34:36creating these activities in a data

2:34:38pipeline then we're going to go on to

2:34:40look at how we can schedule these things

2:34:42starting with the store procedure so

2:34:43when we drop the store procedure

2:34:44activity onto the data pipeline canvas

2:34:47we can have a look at some of the

2:34:48settings that we get of the box right so

2:34:51most of the configuration happens within

2:34:53the settings tab you can navigate to

2:34:55your specific data warehouse that you

2:34:57want to connect to obviously noting that

2:34:59the data warehouse you connect to

2:35:01currently has to be in the same

2:35:02workspace as your data pipeline we've

2:35:05mentioned that quite a lot on the

2:35:06channel around the some of the crossw

2:35:08workpace limitations with the data

2:35:10pipeline so just be wary of that then we

2:35:12connect to the warehouse and then we can

2:35:14choose from a number of different stored

2:35:16procedures that are in that data

2:35:18warehouse we you can Define any sort of

2:35:20parameters that we've got here there is

2:35:22this option to automatically import

2:35:24parameters so it's going to look into

2:35:26your store procedure it's going to pick

2:35:27out any parameters that you've got there

2:35:30here it's found one called first name

2:35:32it's of type string it knows that and

2:35:34currently we're just hardcoding in a

2:35:35value here called Jack but one of the

2:35:38benefits of the store procedure is that

2:35:39you can add Dynamic content right so

2:35:41you're going to be able to parameterize

2:35:43that stored procedure activity and pass

2:35:46something into this value here using

2:35:48data pipeline parameters on the general

2:35:50tab we've got just the name of it so you

2:35:52can update the name you can make it

2:35:54active or deactive we can add a timeout

2:35:56so this might be quite useful for some

2:35:58data pipeline activities maybe you got a

2:36:00really long-standing stored procedure it

2:36:02might take half an hour to run you might

2:36:04want to add a timeout here at 1 hour

2:36:07because if it runs for 1 hour or longer

2:36:10than 1 hour you know that probably

2:36:11something has gone wrong and you don't

2:36:13want it to lock up your database your

2:36:15data warehouse so that's something to

2:36:16bear in mind here the the timeout

2:36:18another one is the retry so this retry

2:36:22setting is available on quite a few of

2:36:24the data pipeline activities and as the

2:36:26name suggests it's going to try it and

2:36:28if it fails that activity it will retry

2:36:31the number of times that you specify in

2:36:33this box here we can give it an a retry

2:36:35interval so we're going to say okay

2:36:36we're going to try it once then we're

2:36:38going to wait for 30 seconds and then

2:36:39try it again secure output and secure

2:36:42input as well this is just going to

2:36:44specify whether you want to Output into

2:36:46a log the results of that activity or in

2:36:50the input as well so that's the stored

2:36:52procedure activity now let's look at

2:36:53this notebook activity so we've got a

2:36:56notebook here and again if you go

2:36:58through to the settings you can Define

2:37:00your workspace and your notebook here

2:37:02I've got notebook load to Silver

2:37:05similarly with the stored procedure we

2:37:07can declare any parameters so if you've

2:37:09got a parameter cell in that notebook

2:37:12you can pass parameters into the

2:37:14notebook from other activities in your

2:37:16data Pipeline and again we've got very

2:37:18similar settings on the The Notebook as

2:37:20we had with the store procedure activity

2:37:22we've got a number of retries so with

2:37:24the notebook a retry is probably more

2:37:27important or more likely you're going to

2:37:28be using this because with a notebook

2:37:30you're obviously using the spark cluster

2:37:32and also potentially querying rest apis

2:37:35so there's a lot more that can go wrong

2:37:37I would say in a notebook than in a

2:37:38store procedure so the retry

2:37:40functionality here I think Microsoft

2:37:42recommends that you set the retry to two

2:37:44or three just so that if your spark

2:37:47cluster is busy and you're doing lots of

2:37:49computation on it and the session

2:37:51times's out or something goes wrong with

2:37:54your notebook execution it's going to

2:37:55retry it so always recommended to add in

2:37:58one or two retries into a notebook

2:38:00activity again we can specify the retry

2:38:02interval down here and secure input

2:38:04output so these are the same as the

2:38:06store procedure activity now the final

2:38:08one we're going to look at is the data

2:38:09flow activity here similarly you're

2:38:11going to go over to the settings table

2:38:13find your workspace and your data flow

2:38:15select a particular data flow that you

2:38:17want to run and that's basically if we

2:38:19go back to the general tab you'll see

2:38:21that the settings here are exactly the

2:38:23same we give it a name description we

2:38:24can set the activity state to active or

2:38:27deactive we can give it a timeout and a

2:38:29number of retries and the retry interval

2:38:32so that's a little bit about the stored

2:38:34procedure activity the data flow

2:38:36activity and the notebook activity in

2:38:39data pipelines now let's look at

2:38:40scheduling a number of these different

2:38:42items in Fabric and the different

2:38:43options that are available to us there

2:38:45so here we are back in our workspace and

2:38:47let's look at how we can do scheduling

2:38:49of different items within Microsoft

2:38:51fabric now there's a few different ways

2:38:54that we can schedule things to work in

2:38:56fabric these ETL jobs mainly we can

2:38:59schedule them to run maybe every hour or

2:39:02something like that depending on your

2:39:03use case we're going to be looking at

2:39:05scheduling data pipelines data flows and

2:39:07notebooks in a bit more detail here so

2:39:09with the data flow we can actually

2:39:10schedule from a workspace so we can

2:39:13click on the settings of a particular

2:39:16data flow just clicking on these

2:39:18ellipses three dotts here here clicking

2:39:20on the settings of that workflow go down

2:39:21to the refresh settings configure a

2:39:24scheduled refresh and we can turn that

2:39:26on and we can change the update

2:39:27frequency to daily or weekly or if you

2:39:31want more fine grained refreshes than

2:39:33that we can add specific times 1:00 a.m.

2:39:36maybe 2 a.m. 3:00 a.m. that kind of

2:39:39thing here now with the data flow the

2:39:40maximum number of Refreshers per day

2:39:43that you can do is 48 so that's in line

2:39:46with what you could do previously in

2:39:48palbi premium we can also send refresh

2:39:50failure notifications here to specific

2:39:52people or the owner or both let's just

2:39:55hop back to our workspace now and look

2:39:59at the notebook so with the notebook

2:40:01it's a similar story we can click on

2:40:04either schedule here or settings both

2:40:07take you through to the same thing here

2:40:09opens up this sidebar menu where we can

2:40:11specify the schedule and again we can

2:40:13repeat it every hour every day weekly or

2:40:17by the minute so every 5 minutes for

2:40:19example example we can specify an start

2:40:21time and an end time and the time zone

2:40:24that you want that schedule to be

2:40:25running on so this is one of the key

2:40:26differentiations between the notebook

2:40:28and the data flow obviously we can

2:40:31specify a frequency that's a lot more

2:40:33than 48 refreshes per hour if we're

2:40:36using this but bear in mind obviously

2:40:38it's going to use a lot more capacity

2:40:39units so if you don't need to refresh

2:40:41your data at this Cadence at this

2:40:43interval then undo it so that's the

2:40:45scheduling of notebooks now the data

2:40:48pipeline is obviously a orchestration

2:40:50tool so another way that we can do it is

2:40:52by putting our notebook and our data

2:40:55flows within a data Pipeline and then

2:40:58scheduling the data pipeline so you'll

2:41:00see here within the data pipeline we've

2:41:02got this schedule button again it's

2:41:03going to open up a sidebar menu we can

2:41:06click on here and specify by the minute

2:41:08hourly daily weekly again like so start

2:41:11and end time and the time zone so it's

2:41:13very similar to The Notebook scheduling

2:41:15functionality now if you're coming from

2:41:17ADF as a data Factory currently one

2:41:19thing to bear in mind is the only way of

2:41:21triggering a data pipeline is with

2:41:23scheduling right to schedule it to

2:41:25trigger a data pipeline using a schedule

2:41:27we don't currently have event based

2:41:29triggers or window triggering or HTTP

2:41:32request triggering all those different

2:41:35really useful functionality for

2:41:37triggering a data pipeline that does

2:41:39exist in ADF as a data Factory currently

2:41:42doesn't exist in fabric the only way we

2:41:44can use is this schedule trigger now you

2:41:46might be thinking okay when should I use

2:41:49the inbuilt scheduling for a data flow

2:41:51for example and when should I use a data

2:41:54pipeline for scheduling well it just

2:41:55gives you a bit more flexibility and

2:41:57functionality for handling different

2:41:59events if you use the day pipeline say

2:42:02for example you wanted to add some sort

2:42:03of activity on fail maybe you want

2:42:06custom notification or some sort of

2:42:09other activity or logging to be done

2:42:11that's possible obviously if you

2:42:12schedule it using an a pipeline if you

2:42:14schedule a data flow to refresh within

2:42:17the actual data flow you don't have

2:42:18access access to all this other

2:42:20functionality that you get in a data

2:42:21pipeline so that's why you might want to

2:42:23think about you know embedding your data

2:42:26flows into a data Pipeline and then

2:42:28triggering them from a data pipeline

2:42:30okay so that rounds up the content for

2:42:31this lesson now let's go back to the

2:42:33slides and test some of the knowledge of

2:42:35the things that we've learned in this

2:42:36section of the study guide okay so let's

2:42:39just round off this video with some

2:42:41practice questions number one the

2:42:43maximum number of scheduled refreshes

2:42:46allowed per day with the data flow Gen 2

2:42:49is a 12 B 24 C 48 D 96 or E unlimited

2:42:56refreshes per day pause the video here

2:42:59have a think and then I'll reveal the

2:43:00answer to you shortly so the answer here

2:43:02is 48 refreshes per day so you can get a

2:43:06data flow Gen 2 to refresh every half an

2:43:08hour if that's what you want to do now

2:43:11if you're coming from the powerbi

2:43:12premium World you'll be used to this the

2:43:14other figures are just incorrect really

2:43:16question two which of the following can

2:43:18you use used to update a row in a data

2:43:21warehouse table is it a a stored

2:43:23procedure B A View C A create table

2:43:26statement or d a function so the correct

2:43:28answer here is a stored procedure that's

2:43:31the only one that allows you to actually

2:43:33update data in a data warehouse table a

2:43:36view as the name suggests is just read

2:43:38only it creates a view on top of the

2:43:40data it doesn't actually modify the

2:43:41underlying data a create table statement

2:43:43is not going to be able to update a row

2:43:45and a function also cannot actually

2:43:48update the underlying data so the answer

2:43:50here is a a stored procedure in a stored

2:43:52procedure we have a lot of flexibility

2:43:55to insert into to update rows to delete

2:43:58rows all of that kind of DML is exposed

2:44:02in the stored procedure and available to

2:44:04us and by DML I mean data manipulation

2:44:06language it's a subset of SQL it's part

2:44:08of the SQL language question three which

2:44:11of the following statements is false

2:44:13when talking about operations in a

2:44:15fabric data warehouse a you can call a

2:44:17function from a stored procedure B you

2:44:20can call a function from A View C you

2:44:22can query A View From Another view D you

2:44:25can call a store procedure from a

2:44:27function so the answer here is D you can

2:44:29call a stored procedure from a function

2:44:32this is the only one that you actually

2:44:34can't do so you remember that with the

2:44:35stored procedure we have to use that

2:44:37exec the ex execute command to execute

2:44:40the stored procedure and you can't do

2:44:42that from within a function now you can

2:44:44call a stored procedure from another

2:44:46stor procedure but that's pretty much

2:44:47the only time when we can call another

2:44:49stored procedure from another item the

2:44:51others you can call a function from a

2:44:52stored procedure well we know we can do

2:44:54that you can call a function from A View

2:44:56yes you can do that and you can query A

2:44:58View From Another view you can also do

2:45:00that so the correct answer here is D

2:45:02that statement is false question four

2:45:04you orchestrating many spark notebooks

2:45:07to perform data transformation

2:45:09activities at the same time now you

2:45:10notice that sometimes the notebook

2:45:12execution is failing because your

2:45:14cluster is busy that's the error message

2:45:16that's it's giving you now what

2:45:17modification can you make to the data

2:45:19pipeline notebook activity to give the

2:45:21pipeline more chances to run

2:45:22successfully is it a deactivate and

2:45:25reactivate the activity B set the number

2:45:27of retries to two or three C change the

2:45:30parameters you're passing into the

2:45:31notebook or d add a failure activity to

2:45:34handle the execution failure so the

2:45:36answer here is B set the number of

2:45:38retries to two or three you'll remember

2:45:41that the the retry option in a notebook

2:45:45activity it allows the activity to retry

2:45:48if it fails a number of times so if your

2:45:51spark cluster is busy and it can't

2:45:53execute the first time it's going to

2:45:55wait and then retry it two three or any

2:45:58amount of times that you specify in that

2:46:00retry parameter in that retry setting

2:46:03deactivating and reactivating the

2:46:04activity that's not going to make much

2:46:06difference changing the parameters well

2:46:08you probably don't want to do that cuz

2:46:09it's going to change the output and

2:46:10adding a failure activity to handle the

2:46:12execution failure so that might be a

2:46:14good idea but it's not actually going to

2:46:16impact the result of your activity right

2:46:20it's not going to allow it to rerun and

2:46:22get a successful execution so the answer

2:46:24here is B congratulations you've

2:46:26completed the second part of section two

2:46:29preparing and serving data in the next

2:46:31lesson we're going to be deep diving

2:46:33into the exciting world of data

2:46:35Transformations within Microsoft fabric

2:46:38make sure you click this video here to

2:46:39join us in the next lesson hey everyone

Transforming data with Dataflows, PySpark, T-SQL

2:46:41welcome back and this is the next video

2:46:43in our dp600 exam preparation course

2:46:47this is video 7 we're going to be

2:46:49looking at transforming data and there's

2:46:51a lot to get through with these and it's

2:46:53a little bit overwhelming when you look

2:46:54at the things that we're going to be

2:46:56covering don't worry I'm going to break

2:46:57them down into a number of different

2:46:58sections hopefully to make it a little

2:47:00bit more digestible so we're going to be

2:47:02looking at First Data cleansing how can

2:47:05we Implement a data cleansing process

2:47:07we're going to look at resolving some

2:47:08common issues that we get with data so

2:47:11duplicates missing data null values

2:47:14conversion of data types and filtering

2:47:16data and we're going to be looking at

2:47:18how we can do that those things using

2:47:20the data flow tsql and Spark as well

2:47:23then we're going to move on to data

2:47:24enrichment so under this category we're

2:47:26going to be looking at merging and

2:47:28joining different data sets together and

2:47:30also enriching the data that we've

2:47:31already got so adding new columns new

2:47:33tables based on our existing data

2:47:36finally we're going to take a look at

2:47:37data modeling we're going to look at the

2:47:38star schema what that is we're going to

2:47:40look at type one and type two slowly

2:47:42changing Dimensions we're going to look

2:47:44at the bridge table and what problem

2:47:46that that solves and how we can

2:47:48implement the solution solution using

2:47:49tsql we're going to look at data

2:47:51denormalization Aggregate and

2:47:53deaggregating data as well as ever we're

2:47:55going to be testing your knowledge at

2:47:57the end of this video so we've got some

2:47:58sample questions and all of the key

2:48:01points for this section of the exam plus

2:48:03links to any further learning resources

2:48:05if you want to dig into a bit more

2:48:07detail into any of these topics they're

2:48:08going to be posted on the school

2:48:10Community I'll leave a link in the

2:48:11description below so first let's take a

2:48:13look at the data cleansing process and

2:48:16to give a bit of structure as to what we

2:48:18looking at here I'm going to be talking

2:48:20through the lens of The Medallion

2:48:22architecture whereby we have bronze

2:48:24silver and gold areas in our data

2:48:27processing pipeline now when we talk

2:48:29about data cleansing normally this takes

2:48:31place anywhere really between the bronze

2:48:34and the silver and maybe in the silver

2:48:36as well this data cleaning CU we're

2:48:38going to get some raw data in bronze but

2:48:41it's going to be messy so we want to be

2:48:42doing all of our data cleaning steps

2:48:44normally between bronze and silver so

2:48:47the way that I'm going to do this is to

2:48:48walk you through how we can do common

2:48:51data cleansing routines and operations

2:48:54within all three of the tools listed

2:48:56there so starting with the data flow

2:48:58then we're going to look at how we can

2:48:59do all these things with tsql in the

2:49:01data warehouse and also in a spark

2:49:03notebook as well so let's jump into

2:49:05Fabric and start with data cleansing in

2:49:08a data flow okay so for this part of the

2:49:10lesson we're going to be jumping into

2:49:12Fabric and we're going to be

2:49:13transforming a particular data set here

2:49:16and it's that car dealership sales data

2:49:19set now the data set itself comes from

2:49:21this blog here Tableau server Guru so

2:49:24thanks very much to this person who has

2:49:25some sample data sets on their website

2:49:28don't worry we're not going to be

2:49:28talking about Tableau servers or

2:49:30anything like that but they do have this

2:49:31nice car sales data set that I

2:49:33downloaded that comes in a pretty good

2:49:35star schema and when you download this

2:49:37it comes actually in

2:49:39xlsx now this data set is actually quite

2:49:42a clean data set so have actually made

2:49:44some modifications to make it a bit more

2:49:46dirty so that we can perform some data

2:49:48cleaning on this data set and I'll leave

2:49:50a link to the dirtified data sets if you

2:49:52can call them that I'll leave them in

2:49:54the school community so go over to there

2:49:56to get the source data files that I used

2:49:58here and to get us started what I've

2:50:00done is I've just put them into a

2:50:02Lakehouse files area so it's number of

2:50:04these csvs and then I've just loaded

2:50:06them simply into a number of tables so

2:50:08these are going to be our bronze

2:50:09Lakehouse tables that we're going to be

2:50:11using for this part of the tutorial so

2:50:13here we are in the data flow Gen 2 and

2:50:16I'm just going to show you how I can do

2:50:17some data cleansing in within the data

2:50:20flow itself I just pulled in one of

2:50:22those tables from our Lakehouse area

2:50:24which this is the source here it's in

2:50:26our lake house and I pulled in this

2:50:27Revenue table so this is what it looks

2:50:29like basically completely untransformed

2:50:33now you notice I've introduced some null

2:50:34values here I've introduced a few

2:50:36duplicate values as well so what we're

2:50:38going to do is just step through the

2:50:40different transformation steps to clean

2:50:42this data set within the power query

2:50:45experience within the data flow so the

2:50:47first thing that we might want to do is

2:50:50remove any duplicate values that we've

2:50:52got so I know because I've introduced

2:50:54some duplicates into this data set there

2:50:56are duplicates in here so to remove any

2:50:58duplicates in this power query engine

2:51:01what we're going to do we're just going

2:51:02to highlight all of these columns

2:51:04clicking on the left hand column holding

2:51:06shift and then clicking on the right

2:51:07hand column that's going to select all

2:51:09of our columns then we can right click

2:51:11on the top here and then remove any

2:51:13duplicate rows so this is important

2:51:15because we want to check that the whole

2:51:17row is a dup so we need to select all of

2:51:19the different columns here and then

2:51:21remove the duplicates like that and you

2:51:22can see it's been added into this

2:51:24applied steps here so if we take a look

2:51:26at our Revenue column here you can see

2:51:28that we do actually have some null

2:51:29values in this data set now obviously it

2:51:32depends on your use case for data

2:51:34cleansing but if this Revenue value is

2:51:37actually really important and this is

2:51:39the only thing that you really care

2:51:40about in that data set then you might

2:51:41want to remove these NS all together

2:51:44from the data model again it depends

2:51:46very much on your use case so to remove

2:51:48the values in this Revenue column we

2:51:49just click on the column itself and then

2:51:52deselect this null value press okay

2:51:54that's going to remove those null values

2:51:57from the data set now another thing we

2:51:58can do is to change the type so if you

2:52:01want to change type in the data flow

2:52:03simply click here on change type now it

2:52:06might not make sense for this particular

2:52:08column because this is already looks

2:52:09good looks like an integer value and

2:52:11it's got an integer type here whole

2:52:12number but say for example this was a

2:52:15text value and you know that deep down

2:52:17is actually an integer then you can just

2:52:19right click on that change type to any

2:52:21of these types here now another thing we

2:52:23can do with the data flow Gen 2 is add

2:52:27in new columns so you can see here add

2:52:29column and what we can do is add a

2:52:31custom column and maybe we want to do

2:52:33some sort of Revenue bin uh maybe you

2:52:36want to add in some logic here to kind

2:52:37of group these revenues into something a

2:52:40little bit different I'm just going to

2:52:42do a simple one here revenue is greater

2:52:44than 1 million and we're going to give

2:52:46it the data type as a true or false and

2:52:48then if we click okay that should give

2:52:50us this True Value just to check some of

2:52:53them are false so yeah that's happening

2:52:55correctly so now maybe we want to change

2:52:57the name of that something a bit more

2:52:58descriptive Revenue over 1 million maybe

2:53:01that's probably a better description for

2:53:03what that new column is you know another

2:53:06way we can filter these data sets maybe

2:53:07you don't want actually want all of

2:53:09these units sold maybe perhaps you're

2:53:11just for this particular piece of

2:53:12analysis that you're doing or this data

2:53:14set you want to actually remove the unit

2:53:16sold over one and again just showing the

2:53:19functionality here really depends on

2:53:22what your use case is in your business

2:53:24as to which data cleansing steps are

2:53:25going to make sense so that's a bit of

2:53:27an overview of power query and how we

2:53:29can do data cleansing in the data flow

2:53:31Gen 2 now let's jump over to the data

2:53:33warehouse and look at how you can

2:53:34Implement similar data cleansing process

2:53:36using tsql okay so I've just jumped over

2:53:38to the SQL end point here in that same

2:53:42bronze Lakehouse because here we've got

2:53:44the same tables here but now we're just

2:53:46in the SQL endpoint experience so we can

2:53:48write some tsql and to begin with let's

2:53:50just have a look at our data set here

2:53:52just make sure that that is coming

2:53:53through okay yep looks exactly the same

2:53:56as we had in the data flow so that's

2:53:58good here so the first step we're going

2:53:59to look at here is identifying

2:54:01duplicates in this table now there's

2:54:04many different ways that we can use to

2:54:06identify and remove duplicates from a

2:54:10table using tsql now this method here

2:54:12I'm using is Group by so what we're

2:54:15going to be doing is grouping by

2:54:16something that I know is unique or

2:54:19should be unique actually for every Row

2:54:22in this fact table and we're going to be

2:54:24using having where a count of more than

2:54:26one so what this is going to mean is

2:54:29that for each of these groups there

2:54:30should be exactly one row so if this

2:54:34returns a count of more than one for

2:54:37this group then we'll see that that is

2:54:40in fact a duplicate row so let's just

2:54:42have a look at these so we can see that

2:54:44here for dealer these two dealer IDs and

2:54:46these two dates we can see there's

2:54:48actually three rows here so if we just

2:54:51give this bit of number of rows just to

2:54:54make that a bit clearer and then we

2:54:55rerun that yeah so now we can see our

2:54:57number of rows is three and four so how

2:55:00do we go about removing those duplicates

2:55:02well again what way of doing that is by

2:55:05using the group buy again we can Group

2:55:07by the same two columns here the DLo ID

2:55:10and the date ID and we can just bring

2:55:11back a aggregate function this is an

2:55:13aggregate function the max of the

2:55:15revenue and it's worth checking before

2:55:17you do this that the revenue figures for

2:55:19each of these rows are actually the same

2:55:21so the max or the Min doesn't actually

2:55:23make a difference we're just going to

2:55:24bring through whatever that value is and

2:55:26the result of that is going to be the

2:55:28same data set but with the duplicates

2:55:30removed which is this one here and then

2:55:32you can save that into another table or

2:55:35whatever Downstream activities you want

2:55:36to do with it so next let's look at

2:55:38missing data nulls removing nulls

2:55:41filtering that kind of thing here so

2:55:43there are some null values in this

2:55:45Revenue column so we're just going to

2:55:46use where revenue is null to identify

2:55:49those first things first so you can see

2:55:51here we've got six rows here where the

2:55:54revenue is null so again you might want

2:55:55to remove those and obviously to remove

2:55:58those what it's pretty simple we can

2:55:59just use is not null and that will

2:56:01return you all where the revenue is not

2:56:04null so one way can we can just quickly

2:56:06verify that that is actually the case is

2:56:08if we just bring this into a bit of a

2:56:10CTE with remove nulls as this we do a

2:56:14select count star remove nulls we'll

2:56:17also do a select

2:56:18count star from the original table let's

2:56:22just compare these two values oh yeah I

2:56:24have to actually call it so we have

2:56:25result one is

2:56:281855 result two is 1861 so we can see

2:56:31that the row count Has Changed by six

2:56:34and those are those six null values that

2:56:36we saw when we did the where revenue is

2:56:38null so that's worked correctly so

2:56:40another thing that we need to bear in

2:56:41mind for the exam is tsql type

2:56:44conversion so here you can see what

2:56:46we're using is the cast function

2:56:48function and here we're going to cast

2:56:49the values in the revenue column as a

2:56:52float I think currently it is yeah I

2:56:54think currently it's an integer and so

2:56:56what we're doing here is we're casting

2:56:58it as a float so basically a decimal

2:57:01number so by doing that you can see here

2:57:02that when we bring both of these columns

2:57:04we've got the original Revenue here in

2:57:06this column and that's an integer and

2:57:07here it's changed the data type into a

2:57:09float using this cast functionality

2:57:12another thing we've done here is to add

2:57:14another column so when you're using ttal

2:57:16you can just add in another row here

2:57:18into your SQL scripts and you can do

2:57:20whatever you want here this is just a

2:57:22bit of an example here divided the

2:57:24revenue by two and I've given it a name

2:57:26of half the revenue so that's just an

2:57:28example of adding new columns that we

2:57:31can do in tsql okay so just to round off

2:57:34this part of the tutorial next we're

2:57:35going to look at transforming data using

2:57:38pypar notebooks and specifically we're

2:57:40going to be looking at some of the data

2:57:42cleansing routines that we've been

2:57:43looking at in the data flow and the tsql

2:57:46engine now we're going to look at how we

2:57:47can implement them in a spark notebook

2:57:50now I did do a 3.5 hour tutorial which

2:57:53goes into a lot more depth about

2:57:55different data cleansing operations you

2:57:58can have a look at that video here I'll

2:58:00leave a link in the school Community if

2:58:01you want to have a look at that in a bit

2:58:03more detail but here we're just going to

2:58:04go on a bit of a quick Deep dive into

2:58:06some of the most common operations and

2:58:08I'll also leave this notebook on our

2:58:11school community so if you want to play

2:58:13along at home then you can do that as

2:58:14well so we're using this same bronze

2:58:18lake house we've got our tables here I'm

2:58:20just going to be having a look at this

2:58:21Revenue table which is our fact table in

2:58:23a bit more detail and I've begun by just

2:58:26reading it into a spark data frame and

2:58:28displaying the results here just to

2:58:30check that everything's loading okay and

2:58:31got everything logged in here now you

2:58:33notied that I haven't actually committed

2:58:35or haven't actually written any of those

2:58:38changes that we made in the data flow

2:58:39and the tcq engine so this is still

2:58:42reading from the raw data here in that

2:58:44bronze layer so let's start by looking

2:58:46at duplicate data again to identify

2:58:49duplicates we can use a similar kind of

2:58:50pattern that we used in tsql by looking

2:58:53at group by or by using Group by and

2:58:55then looking at count of more than one

2:58:58so here you can see we've identified the

2:59:00rows where we have some duplicates in

2:59:03that data set and luckily it's the same

2:59:06as what we were seeing in our tsql

2:59:09engine so we can see that these Branch

2:59:11IDs and date IDs have a count of four

2:59:13and three respectively how do we go

2:59:15about dropping those duplicate values

2:59:18from our data set well in spark we have

2:59:21this drop duplicates method so what I'm

2:59:23going to do here is just start by

2:59:25counting the rows in the original data

2:59:28set performing this drop duplicates

2:59:30saving it into D duped which is a new

2:59:33data frame and then counting the rows in

2:59:35that new duped data frame okay so I've

2:59:37just printed out the result here we can

2:59:39see that the process has removed if I

2:59:41could spell process correctly it's

2:59:42removed five rows from the data set so

2:59:44I've just taken the difference really

2:59:46between our original data set and that

2:59:48that duped data set so that makes sense

2:59:50CU we've got seven count here originally

2:59:53and obviously two of those we want to

2:59:55keep right because they are actually

2:59:56valid data so we've removed five rows so

2:59:59five of them are duplicates so that

3:00:01looks like it's worked correctly and if

3:00:03we were just going to verify again that

3:00:05that's worked we can perform the same

3:00:07operation that we did before grouping by

3:00:09these two Fields looking at where the

3:00:12count is more than one and we've got

3:00:13this empty data frame here so that looks

3:00:15like it's worked correctly next we're

3:00:17going to take a look at missing data and

3:00:19how we handle nulls and that kind of

3:00:21thing in spark so first off we're going

3:00:23to look at identifying missing values

3:00:26you might want to do a bit of like

3:00:27inspection of your data before you go

3:00:29ahead and drop things so a good idea to

3:00:32interrogate your data a bit have a look

3:00:34at what you might be deleting so let's

3:00:37just run this and then we'll talk

3:00:38through it so we've got our original

3:00:39data frame here and we're filtering on

3:00:42this specific column here so we're

3:00:44passing in DF do revenue. isnull so so

3:00:48this is a method that we get on a spark

3:00:51column and we can see that it's

3:00:52returning these six rows that we know

3:00:54are null now another way we can do that

3:00:57is I've changed two things in this

3:00:59second example here but it's basically

3:01:00using DF do Weare so DF do filter and DF

3:01:03do Weare are basically synonymous in

3:01:05spark this time we've passed in call

3:01:08which is one of the spark SQL functions

3:01:11it's just another way of writing and

3:01:13obtaining that column data and again

3:01:15we're calling is null on that column

3:01:17data and we're saving it into this nulls

3:01:192 variable and we're displaying that and

3:01:21so we can see that these two are exactly

3:01:23the same so both methods have returned

3:01:25the same null values which is good so

3:01:28now let's look at dropping some of those

3:01:30null values using the drop na method so

3:01:33this method has a few different

3:01:35parameters these are two of them how

3:01:37threshold and also subset as we're using

3:01:40in this example here so if you pass in a

3:01:43value in the how parameter it can be

3:01:45either any or all so if it's any it's

3:01:47going to drop the row if any of the

3:01:49values in any column is null if you pass

3:01:52in a how value of all it's going to drop

3:01:55the row only if all of the values are

3:01:58null now threshold is going to give you

3:02:00a threshold for the number of columns

3:02:02that need to be null for that row to be

3:02:05dropped and obviously if you use this

3:02:06value it's going to overwrite that how

3:02:08parameter now the other one that we're

3:02:09going to be using is subset so you can

3:02:11pass in subset and it's going to limit

3:02:13The Columns that it looks for for your

3:02:17drop so example we only want to drop it

3:02:19if there's a null value in this Revenue

3:02:23column and you can pass this in either

3:02:24as a list or as a string as well in our

3:02:27example we've only got one column in

3:02:29that list so we're going to save the

3:02:30result as no Nas and then we're going to

3:02:33do the same thing we're going to print

3:02:35out the result here just to check that

3:02:37we have actually removed some rows so

3:02:39we've removed those six rows the null

3:02:41values from a data set so that looks

3:02:43like it's worked correctly so next let's

3:02:45look at type conversion and also adding

3:02:48new columns into a spark data frame so

3:02:51we can call DF do print schema and it

3:02:54gives you a bit of a look at what our

3:02:56schema is initially for this data set

3:02:59including the data types for each column

3:03:02what we're going to do is get the column

3:03:04using DF do unit sold so that's one of

3:03:07our columns here currently it's an

3:03:08integer and for the purposes of this

3:03:10demo we want to use thecast method and

3:03:13we're going to give it the data type of

3:03:15string then what we've done is we've

3:03:16printed the schema again just to check

3:03:18what that schema looks like and we can

3:03:20see here this unit sold converted so

3:03:23we've passed in this units sold

3:03:25converted which is the new column name

3:03:27that we get with this with column

3:03:29function and we can see that from the

3:03:31print schema it's showing a data type of

3:03:34string so we've successfully casted that

3:03:36value from an integer into a string

3:03:39value so next just take a quick look at

3:03:41filtering and we've already had a look

3:03:42at some filtering previously in this

3:03:45lesson but let's just go over it again

3:03:46so in a filtered data frame here we're

3:03:49getting the original data frame and

3:03:51we're calling DF do where we're passing

3:03:54in the column name and the column here

3:03:56is revenue and we want to filter only

3:03:58for data that is more than 10 million so

3:04:02we can see that this has actually

3:04:03removed it's filtered out

3:04:071,19 rows from this data set as I

3:04:10mentioned previously we can use DF wear

3:04:12or DF filter if you want to do filtering

3:04:14if I change this to DF do filter should

3:04:17should work exactly the same there you

3:04:19go so now let's look at data enrichment

3:04:21so adding new columns so we've already

3:04:24seen this one as well we're going to be

3:04:26using DF dowi column and you can also

3:04:29use with columns if you want to do more

3:04:31than one of these at a time but we're

3:04:33just going to be showing you one column

3:04:34here what we're doing we're taking the

3:04:35original data frame we're calling with

3:04:37column to add a new column we're giving

3:04:39the column a name half revenue and we're

3:04:42giving the function to apply to that new

3:04:45column and in our example we're just

3:04:47going to be in the revenue so DF Revenue

3:04:49divid two and then we're displaying the

3:04:51results in our enriched data frame here

3:04:54so if we just call that that's going to

3:04:55look like this so we've got our new

3:04:57column which is half Revenue here which

3:04:59is the revenue divid by two finally

3:05:01we're just going to look at joining and

3:05:02merging data frames in spark so for this

3:05:05one we're going to be using some

3:05:07different tables I've just pulled in the

3:05:09Dealer's data frame so the de dealers

3:05:12table which is this one here and also

3:05:14the countries data frame and you'll

3:05:16notice that these do actually have a

3:05:17joining key so they both have country ID

3:05:20so each dealer has a country ID what we

3:05:23want to be doing is pulling through the

3:05:25country name maybe we've got a

3:05:28normalized data model and we want to be

3:05:30denormalizing it and we're going to be

3:05:31looking at what that means in a bit more

3:05:33detail a bit later in this tutorial but

3:05:35for now let's just look at joining and

3:05:37merging these two data sets together so

3:05:39the basic Syntax for a join in spark is

3:05:42well we're going to get one data set

3:05:44here which is the dealer data set we're

3:05:47going to call dealers DF do jooin that's

3:05:50going to be the first part of our join

3:05:53we're going to pass in the second data

3:05:55frame that we want to join it to in our

3:05:57case countries DF then we're going to

3:05:59specify on what column we're going to be

3:06:01joining on so we're going to be joining

3:06:03on dealers DF country ID is equal to

3:06:07countries DF do countryid in this next

3:06:09row we're just going to specify some

3:06:11select statements so what we've done is

3:06:13we've used this brackets here because

3:06:15we're going to be chaining more than one

3:06:17command

3:06:18together and in this second row we're

3:06:19just selecting a few different columns

3:06:22to be returned in this joined data frame

3:06:25we just want to specify that we want

3:06:26returned the dealer ID the country ID

3:06:28and the country name this is all we care

3:06:30about in that resultant data frame so

3:06:33dealers DF is not defined that's cuz we

3:06:34haven't defined it let's just read that

3:06:37data into a spark data frame first and

3:06:39then we can run this one here okay so

3:06:40now we've got dealer ID country ID and

3:06:43country name in the same data frame now

3:06:45one thing you notice here is that the

3:06:47country name has actually pulled through

3:06:49a few interesting characters so that's

3:06:52something you might want to change later

3:06:53on in your data processing workflow

3:06:56might be that I've saved the CSV file in

3:06:59an incorrect format maybe that's

3:07:00something we want to clean as well in

3:07:03our data cleansing process but for now

3:07:05we're just going to leave it like this

3:07:06so finally we're going to be taking a

3:07:08look at data modeling and typically this

3:07:10takes place within that gold layer

3:07:13because these are going to be the

3:07:14analytical models and the data models

3:07:16that we're going to be bringing into to

3:07:17our semantic layer so let's start off

3:07:19with the star schema so this is an

3:07:21example of a star schema what you can

3:07:24see is in the middle we've got a fact

3:07:26table and this fact table represents

3:07:29sales so it's revenue for a particular

3:07:31company here this example is using car

3:07:34sales so the amount of cars sold from

3:07:36different dealerships so we've also got

3:07:39a dealership Dimension table a dim date

3:07:42dim model so the type of car that's sold

3:07:44and the branch that that was actually

3:07:46sold at so called a star schema because

3:07:49we have our fact table in the middle and

3:07:50then multiple Dimensions all linking to

3:07:53that fact table via some sort of primary

3:07:55key now the reason we prefer star

3:07:57schemas is if you want to be doing bi

3:08:00powerbi basically because the powerbi

3:08:02engine is most efficient in this star

3:08:04schema so a pretty common piece of data

3:08:06modeling that you might have to do as an

3:08:07analytics engineer is prepare this data

3:08:11model in your gold layer of a data

3:08:13warehouse for example so that your

3:08:15powerbi developers can just pick up this

3:08:17data model and you know write their

3:08:19measures on top of it now normally when

3:08:21we're modeling in this star schema data

3:08:23model the fact table is going to be aend

3:08:26only so we're not really going to be

3:08:28updating many values in that fact table

3:08:31it's just going to be new sales added

3:08:32onto the end of that fact table normally

3:08:35it's going to be a very long list your

3:08:37fact table and you're going to create

3:08:38connections from that fact table into

3:08:41the other dimensions so if you want to

3:08:43know more information about that

3:08:45particular sale in the fact table then

3:08:47you can going to find those in the

3:08:48dimensions and the dimensions tables can

3:08:51change over time right take for example

3:08:54dim branch in that Dimension table you'd

3:08:57expect to see details about the

3:08:59different branches that exist for this

3:09:01car manufacturer and some of those

3:09:02things can change over time some of the

3:09:04details about a particular Branch might

3:09:07change right they might change address

3:09:09they might change name certain details

3:09:11about each Dimension might change over

3:09:14time and one way of dealing with that in

3:09:16data model well we need to introduce

3:09:18this concept of slowly changing

3:09:21Dimensions because with our fact table

3:09:22as we mentioned we're not really

3:09:23updating anything over time we're just

3:09:26appending new rows with our Dimensions

3:09:29we need to be able to handle different

3:09:31changes to that Dimension table over

3:09:33time and it's what we call slowly

3:09:35changing Dimensions because these are

3:09:37not going to be big updates might happen

3:09:39once a month or every week a lot of

3:09:41slower frequency than the data is going

3:09:43to be added into that fact table and in

3:09:45data modeling there's a lot of different

3:09:47ways

3:09:47that we can model these slowly changing

3:09:50Dimensions here's an example here take a

3:09:53different example we're looking at

3:09:54employee table we've got employee ID on

3:09:57the left hand side employee name and the

3:09:59department now it's not uncommon for an

3:10:01employee to change departments so on the

3:10:04right hand side you can see that this

3:10:05Dimension has actually changed so Danny

3:10:08Walker who employee ID number two is

3:10:11moved from the marketing department into

3:10:14the sales department and we can model

3:10:15this in a number of different ways ways

3:10:18now the first way in which you need to

3:10:19really be aware of for the exam is the

3:10:22type one SCD or slowly changing

3:10:24Dimension so in the type one SCD we're

3:10:27going to be overwriting any new data so

3:10:30we're going to get the employee data and

3:10:32every time we query that data set that

3:10:35Source data again we're just going to

3:10:37overwrite whatever's in that table we're

3:10:39not going to store any sort of History

3:10:41so we're going to be implementing it

3:10:42with the overwrite writing mode and you

3:10:45can do that either in the data pip line

3:10:47the data flow or in py spark as well now

3:10:50if you're using a Lakehouse as your data

3:10:53store it's important to bear in mind

3:10:54that the history can still actually be

3:10:57retrieved at a point in time using the

3:10:59Delta log so in a lake house

3:11:01architecture in the fabric Lakehouse

3:11:03because you can access those Delta logs

3:11:05just because you overwrite it doesn't

3:11:07necessarily mean that that data isn't

3:11:08stored so the second way that we can

3:11:10deal with this sort of change in a

3:11:13dimension is the type two slowly

3:11:15changing Dimension so in this example

3:11:18we're going to need a few extra columns

3:11:19you can notice that at the top there

3:11:21we've added valid from and valid to and

3:11:24optionally an is current as well which

3:11:26is also quite useful for bi purposes as

3:11:29well so you notice that when we first

3:11:30write a record into this table we're

3:11:33going to populate the valid from that's

3:11:35when this row is valid in that data set

3:11:38and when you first write it the valid to

3:11:40is going to be sometime long into the

3:11:42future normally it's the year

3:11:4499,999 and we also set the is current

3:11:46flag to one or true now when we get an

3:11:50update to that data set so we get a

3:11:52brand new set of data well in the type

3:11:54two SCD we add new rows based on any

3:11:58incoming data so you see in this type

3:12:01two SCD we're going to be adding a third

3:12:03row and it's going to be Daniel Walker

3:12:04but this time the department is sales we

3:12:07have to do a few things here with the

3:12:08valid 2 and the valid from dates so when

3:12:11you write that third row into the data

3:12:13set you need to update row two to

3:12:16populate that valid 2 column and set it

3:12:18is current to zero or false the new row

3:12:21when we write that third row into the

3:12:24data set obviously we're going to have

3:12:26department is sales we're going to add

3:12:27the valid from date as the day that

3:12:29you're making the change the update with

3:12:31the valid two of the year 99,999 and

3:12:34then is current of true now you can see

3:12:37that by doing this we're actually

3:12:39storing a bit of a history about how our

3:12:42Dimension is evolving over time and so

3:12:45you can on the back of this create some

3:12:47quite sophisticated queries about what

3:12:49the state of that Dimension table is at

3:12:52any given point in time now when we're

3:12:54implementing the type two slowly

3:12:56changing Dimension there's a few things

3:12:58that you need to bear in mind so if

3:12:59you're using the Lakehouse and Spark

3:13:02engine then bear in mind that as we

3:13:03mentioned previously the Delta log

3:13:05actually stores a history of all the

3:13:07rights to a particular table so this

3:13:10might be a simple option and depending

3:13:12on how you want to use the history

3:13:14Downstream in your analysis that might

3:13:16be good enough for some use cases now if

3:13:18not then you can use the merge into so

3:13:21we can use that as part of the spark SQL

3:13:23library and we can use that to basically

3:13:25update the underline Delta tables now if

3:13:28you're using the data warehouse and the

3:13:30tcq experience then one method is to

3:13:34load your new data into a staging table

3:13:37and then use a stored procedure to

3:13:40perform the checks around the updates

3:13:42and the inserts and the valid to and

3:13:44valid from calculations now

3:13:46unfortunately the merge operation in

3:13:48tsql the tsql surface area is not

3:13:50currently supported so you have to

3:13:52actually implement this manually if you

3:13:54want to do that in the data warehouse

3:13:56currently now one kind of implementation

3:13:58trick is to use row hashing here so say

3:14:02for example you want to check which rows

3:14:05have changed from One update to the

3:14:07other and one method that you can use to

3:14:10do this in a bit more of an efficient

3:14:11manner is to implement row hashing so

3:14:14you can hash the value of an entire row

3:14:17both in your existing data set and in

3:14:19the new data set that you're checking

3:14:21and then you can compare the two hashes

3:14:23and obviously if the hash values match

3:14:25then you know that that record hasn't

3:14:28changed but if the two hash values are

3:14:30different then you can go ahead with the

3:14:32update logic that you need to update

3:14:35that table next we're going to talk

3:14:36about Bridge tables so imagine a company

3:14:39has many different projects running and

3:14:42the employees assigned to one or many

3:14:45projects what does this look like if you

3:14:47want to do a bit of a data model and

3:14:49maybe build a power VI report off that

3:14:51well in the question here or in the

3:14:53scenario we can see that actually many

3:14:55employees can work for many different

3:14:58projects at the same time so that is a

3:15:00many to many relationship which exists

3:15:02between our dim projects and our dim

3:15:05employees table and this can cause quite

3:15:07a lot of issues when it comes to powerbi

3:15:09and efficiency in very large data models

3:15:12as well so one thing that we can do to

3:15:14resolve this is to implement a bridge

3:15:17table and what that looks like is

3:15:19basically a onetoone mapping of all

3:15:22projects and all participants now this

3:15:25is really useful because it turns our

3:15:27many to many relationship into two on to

3:15:30many relationships let's have a look at

3:15:32how you can Implement that in a tsql

3:15:35data warehouse just to kind of show you

3:15:37what that looks like in real life okay

3:15:39let's explore Bridge tables in a bit

3:15:41more detail and here we're in the data

3:15:44warehouse experience and I've just built

3:15:47a gold data warehouse and we've got

3:15:48these two tables here we got a projects

3:15:50table and an employees table now it's

3:15:53just a simple demo just to show you what

3:15:54a bridge table might look like and how

3:15:56you can implement it in SQL our tables

3:15:59here are the projects we've got the

3:16:00project ID and the project name and

3:16:02we've also got these project participant

3:16:04and if we have a look at the data here

3:16:05for our projects table we've got a bit

3:16:08of a messy string separated values in

3:16:12this project participants column so we

3:16:15can see that each project has multiple

3:16:18project participants and some of them

3:16:20are overlapping right so these are

3:16:21employee numbers that come from this

3:16:23employees table so employee 101 who is

3:16:27John Smith he's actually working on

3:16:29multiple projects right so in our data

3:16:31model here we've actually got a many

3:16:32many relationship we can't really

3:16:34implement it at the moment because of

3:16:35that string concatenation in this

3:16:38project participants column so how do we

3:16:41get around this well and I have actually

3:16:43got the tsql here and I'll leave this in

3:16:46in the school Community as well if you

3:16:48want to have a play around with this

3:16:50yourself I've just created the projects

3:16:52table I've created the employees table

3:16:54and I've inserted some values here so

3:16:56one way that we can use to resolve that

3:16:59many to many relationship is to create a

3:17:01bridge table between those two tables

3:17:05between the projects and the employees

3:17:06table and in this example I'm

3:17:08implementing that as a SQL View and what

3:17:11I'm doing is I'm using the cross apply

3:17:13function here in tsql along with string

3:17:16split

3:17:17so string split is basically going to

3:17:19look in that project participants column

3:17:22within dbo do projects so that's the a

3:17:25comma separated value column which is a

3:17:28bit messy but we can use cross apply and

3:17:31string split together and it's basically

3:17:33going to separate all those values based

3:17:35on this separator here the the comma

3:17:37separation and what this is going to do

3:17:39is it's going to create a one to one

3:17:40mapping for every single project and

3:17:43every single project participant so if

3:17:45we just run this here here let's just

3:17:47have a look at what that looks like here

3:17:48so you can see if we look at the

3:17:50original table just to remind ourselves

3:17:52of what that project's table looks like

3:17:55so initially it looked like this it had

3:17:57two rows and it had project one with two

3:18:00participants and project two with three

3:18:04participants and what we've done in this

3:18:07view if we just recalculate that it's

3:18:09basically separated out all of these

3:18:12comma separated values and it's made one

3:18:14row for each one so now we have five

3:18:17rows because these are all the different

3:18:18combinations of projects and employees

3:18:21what we can do is we can create a view

3:18:23with this logic like so and I've just

3:18:26called it dbo view Bridge Project

3:18:29participants and that's going to make

3:18:30that available in our data model now so

3:18:33let's just have a look at this so now

3:18:35we've got this view bridge table in our

3:18:38data model and what we can do is we can

3:18:40now connect the project ID to the

3:18:42project ID in this instance it's

3:18:44actually going to be one too many

3:18:46because this is a dimension table this

3:18:48is our Bridge table so it's going to be

3:18:49one to many and we can do the same on

3:18:51this side here so this time it is going

3:18:53to be many to one so what we've done is

3:18:54we've transformed a many to many

3:18:56relationship into two one to many

3:18:59relationships and this is going to make

3:19:00it a lot easier when you implement this

3:19:02stuff in powerbi and that's one way that

3:19:05you can resolve a many to many

3:19:06relationships using a bridge table now

3:19:09we've implemented this in a SQL view

3:19:12just because it's quite easy in reality

3:19:14if you wanted to use direct Lake mode

3:19:17obviously you can't use a view with

3:19:19direct late mode so if you want to use

3:19:22direct late mode then you might want to

3:19:24use a stored procedure to materialize

3:19:27out the data in this bridge table so

3:19:29that it doesn't fall back to direct

3:19:31query mode but that's Bridge tables in

3:19:34tsql okay so the next thing we want to

3:19:36talk about here is normalized and

3:19:38denormalized data so if we go back to

3:19:41our data model here where we were

3:19:43looking at car sales so we have some

3:19:45sort of fact Revenue in the middle and

3:19:47some Dimensions that give us more

3:19:49information about that particular sale

3:19:52so what branch it was at what model of

3:19:54car was sold the date that it was sold

3:19:57and also the dealership and what we've

3:19:58done is we've added in another dimension

3:20:00here onto that dim dealers so there a

3:20:03relationship between the dim dealer and

3:20:05the dim cities so it's basically telling

3:20:08us the city of that dealership now this

3:20:11is what's called a snowflake

3:20:12architecture now this can be a very

3:20:15efficient way of storing really large

3:20:17data models because you'll notice in the

3:20:19dim dealers table we're only storing the

3:20:21city ID we're not bringing through any

3:20:24other data into that dim dealers now

3:20:26this can be beneficial in some instances

3:20:29for example as we mentioned for

3:20:30efficiency but when we move this kind of

3:20:32model into the powerbi world it can

3:20:34bring some limitations as well so one

3:20:37way to get around this is with the

3:20:39denormalized data model and this is what

3:20:42this looks like here and it's

3:20:43effectively turning our snowflake model

3:20:46back back into our star schema model

3:20:49which we know can be a lot more

3:20:50efficient when we've got very large data

3:20:52sets now this does introduce some

3:20:55redundancy into that dim dealers

3:20:57Dimension because instead of just

3:20:59storing the city ID for every dealership

3:21:02we're going to be repeating the city ID

3:21:04and the country ID and the region for

3:21:06all of the rows in our dim dealerships

3:21:10but by D normalizing our data model we

3:21:13actually get a number of other benefits

3:21:15so now we have everything in that dim

3:21:17dealer's table which makes it a lot

3:21:19easier to build things like filters on

3:21:21top of that Dimension table we can also

3:21:23have hierarchical filters so we can look

3:21:26at the dealerships by City Country and

3:21:29region and build a bit of a hierarchy

3:21:31there which isn't possible when you've

3:21:33got that split across two Dimension

3:21:35tables finally we're going to take a

3:21:36look at data aggregation and the word

3:21:39aggregation can mean a few different

3:21:41things in data analysis and data

3:21:43modeling and I'm not 100% sure what

3:21:45Microsoft expects for the data

3:21:47aggregation and deaggregation that

3:21:49they've mentioned in the study guide but

3:21:50I'm going to talk through both examples

3:21:52so that you know both of them and if you

3:21:54have any insight as to what Microsoft

3:21:56mean from the study guide when they talk

3:21:57about aggregation then let us know in

3:22:00the comments so the first possible

3:22:01meaning when we talk about data

3:22:02aggregation is when we have different

3:22:04slices of the same data set but they're

3:22:07spread across different files so this

3:22:09one we have one file with the UK data

3:22:12one file with the USA Data and one file

3:22:15with the Canadian data data and in this

3:22:17context data aggregation can mean

3:22:19basically combining these data sets into

3:22:22one long table of transactions in this

3:22:25case showing you Revenue across all your

3:22:27different Source data sets right and we

3:22:29can implement this in a number of

3:22:31different ways really in Fabric in the

3:22:34data flow we can use the append

3:22:36functionality and in the tsql experience

3:22:38and also spark SQL we can use Union and

3:22:41Union all and Union is also a method in

3:22:44pypar as well now the difference here

3:22:46between Union and Union all with a union

3:22:49it's actually going to remove any

3:22:51duplicate rows in the resultant data set

3:22:54whereas Union all is just going to

3:22:56basically append all of the different

3:22:58data sets on top of each other and it's

3:23:00not going to check for duplicates so

3:23:02also mentioned in the study guide is

3:23:04data deaggregation and again there can

3:23:06be many different meanings to the word

3:23:08deaggregation so assuming we mean the

3:23:10first meaning of aggregation then

3:23:12deaggregation is going to be the

3:23:13opposite right it's going to be

3:23:14splitting one large file into multiple

3:23:17different categories of data so another

3:23:19possible definition of data aggregation

3:23:23is when we transform a data set from a

3:23:26more granular data set into a less

3:23:28granular data set so on the left hand

3:23:30side for example we have all of our

3:23:32transactions and the revenue for that

3:23:34transaction and the different country

3:23:37that that transaction was made in now if

3:23:39we were to get aggregate that data

3:23:42perhaps by country then we get some sort

3:23:45of aggregation met tric for each country

3:23:48so here we're showing the total revenue

3:23:50for each country so the UK here has 600

3:23:53which is 100 plus 500 the US has 200 and

3:23:57Canada has 800 now this is typically

3:24:00implemented using the group by statement

3:24:03and the group by functionality and this

3:24:05can be done in the data flow tsql or in

3:24:07spark and when we build this kind of

3:24:09aggregate transformation we also need an

3:24:11aggregation function and it's normally

3:24:14one of count sum Max Min average so it's

3:24:19how are you combining all of the values

3:24:21within that group buy statement so in

3:24:23our example in the top right hand corner

3:24:25we used a sum so we just summed up all

3:24:27of the different Revenue numbers for

3:24:29each country but you could do average

3:24:32revenue you could do the max Revenue it

3:24:34depends on your use case here okay let's

3:24:36just round up everything that we've gone

3:24:38through in this lesson and test some of

3:24:41your knowledge question one when using

3:24:43DF dojin in pisar Notebook the default

3:24:47join type is a full outer join B inner

3:24:51join C left join d right join or E anti

3:24:55join pause the video here take a moment

3:24:57to think about the answer and I'll

3:24:58reveal the answer to you shortly so the

3:25:00default join type in a p spark or in any

3:25:04spark notebook is the inner join not

3:25:06much to say about that one that's just

3:25:08something you need to know all the

3:25:09others are incorrect question two you're

3:25:11looking to migrate a data transformation

3:25:14workload that is currently done using a

3:25:17data flow Gen 2 and convert it into a

3:25:20tsql script now the data flow appends

3:25:23two data sets together and removes any

3:25:26duplicate rows which tsql command can

3:25:29you use to implement this transformation

3:25:31a a left joint B Union or C concat D

3:25:35append or E union so the answer here is

3:25:39the union now we mentioned when we were

3:25:41talking about unions that the union

3:25:43removes duplicates and in our question

3:25:46here obviously the important sentence to

3:25:48pick out was that the source data flow

3:25:50the thing that we're trying to convert

3:25:52into a tsql script well currently that

3:25:54appens two data sets together and

3:25:56removes any duplicate rows so the tsql

3:25:58equivalent of that is the union it's not

3:26:01going to be the union all because that's

3:26:02not going to remove the duplicate rows

3:26:04it's not going to be the left join

3:26:05because that won't do what we want to do

3:26:07now aend is not a tcq function conat is

3:26:12how you could achieve this in pandas but

3:26:14not in tsql so the answer here is Union

3:26:17question three you have a spark data

3:26:19frame called DF your goal is to remove

3:26:22rows that contain a null value in the

3:26:24transaction date column which of the

3:26:26following will help you achieve this a

3:26:29DF do drop duplicates B DF do dropna

3:26:32with how equal to all DF filter

3:26:35transaction date do is null D DF drop na

3:26:39how equals any or E DF do dropna with

3:26:42subset equal to transaction date so the

3:26:45correct answer here is is e we want to

3:26:47be using the drop Na and passing in

3:26:50subset equal to transaction date now the

3:26:53question here is asking us to remove

3:26:55rows that contain null values in a

3:26:58specific column so that specific column

3:27:00is the important part of the question so

3:27:02we want to be using drop na because it's

3:27:04going to remove the rows and by passing

3:27:05in subset equals transaction date that's

3:27:08going to specify only to look in the

3:27:10transaction date column so maybe that's

3:27:12a really important column in our data

3:27:15set we want to be abs Ely sure there's

3:27:17no na values CU if there's any na values

3:27:19then maybe that's going to make our

3:27:20analysis completely redundant or it's

3:27:22going to ruin all of our Downstream

3:27:23analysis so you might want to remove

3:27:25those rows entirely it's not going to be

3:27:27B because in B and D we're not actually

3:27:30specifying that subset now if you to

3:27:32implement D then it would drop the rows

3:27:36where there is a null value in the

3:27:37transaction date column but it would

3:27:39also remove rows with null values in any

3:27:42other column so that might be too much

3:27:44based on your requirements that's

3:27:46probably not what you want to be doing

3:27:47filter that's going to just filter the

3:27:50data set for null values cuz that's

3:27:52going to return all of the rows which

3:27:54are null so we want to be doing the

3:27:55opposite of that basically and a drop C

3:27:58duplicates well that's not what we're

3:27:59trying to achieve here so that's not

3:28:01going to be the answer question four a

3:28:02classical star schema data model

3:28:05consists of the following is it a one

3:28:07Central fact table and multiple

3:28:09Dimension tables B one dimension table

3:28:11and one or more fact tables c one fact

3:28:14table multiple dimens di tables some

3:28:17with Dimension to Dimension

3:28:18relationships or D one big fact table

3:28:21fully denormalized without any Dimension

3:28:23tables so the answer here is one Central

3:28:25fact table with multiple Dimension

3:28:27tables so B is obviously the wrong

3:28:29answer here because it's got one

3:28:31dimension table and many fact tables

3:28:33it's not going to be the star schema C

3:28:36is a snowflake one fact table with

3:28:38multiple Dimension tables some with

3:28:40Dimension Dimension relationships so

3:28:42that's not going to be what we're

3:28:43looking for in a star schema and one big

3:28:46fact table fully denormalized without

3:28:48any Dimension tables again that's not

3:28:50really a classical star schema so the

3:28:52answer here is a you inherit a data

3:28:54project and you're inspecting the tables

3:28:57in the data warehouse one of the tables

3:28:59is a dimension table dim contacts with

3:29:02the following columns contact ID contact

3:29:04name contact address effective date and

3:29:06effective until make an assumption about

3:29:08the type of data modeling that being

3:29:10implemented in this Dimension table is

3:29:12it a type zero SCD slowly change di

3:29:15mention a type one SCD a type 2 SCD or a

3:29:18Type 3 SCD so here the answer is Type 2

3:29:22SCD so we can see from inspecting the

3:29:25columns here that they've got two date

3:29:27fields or we can at least assume their

3:29:29dates so effective date is probably when

3:29:33that row when that contact data point

3:29:36was entered into the system and

3:29:38effective until is basically the same as

3:29:40that valid two now we can't actually see

3:29:42the data here but we can assume that

3:29:44that's what those two columns are doing

3:29:46a type zero SCD is a fixed Dimension so

3:29:49something that never changes type one is

3:29:51obviously you're not going to be

3:29:52tracking that effective date and

3:29:54effective until those two dates it's

3:29:57just going to be overwritten in a type

3:29:58one and a type three is another slowly

3:30:00changing Dimension type that is not

3:30:02actually asked about in the exam but

3:30:04it's basically going to store previous

3:30:06values for each of your columns or at

3:30:08least the columns that you're interested

3:30:10in storing so say for example you'd have

3:30:12contact name there you might also have

3:30:14previous contact name and then when your

3:30:16data gets updated you're going to update

3:30:18the previous contact name and the new

3:30:19contact name as well so that's a Type 3

3:30:22SCD don't think that's in the exam but

3:30:24just something to bear in mind so the

3:30:25answer here is c a type two slowly

3:30:28changing Dimension congratulations

3:30:29you've now completed the third part of

3:30:32section two preparing and serving data

3:30:35in the next lesson we're going to be

3:30:36looking at performance monitoring and

3:30:39optimization of all of our data

3:30:41processing workloads in fabric so make

3:30:44sure you click here to join us in the

3:30:46next lesson I'll see you there hey

Optimizing performance

3:30:48everyone welcome back to the channel

3:30:50today we're continuing our dp600 exam

3:30:53preparation course and we're up to video

3:30:55eight we're making a very good progress

3:30:58here on the course plan today we're

3:31:00going to be looking at optimizing

3:31:02performance and specifically we're going

3:31:04to be covering these bullet points in

3:31:06the dp600 study guide so we're going to

3:31:09be looking at mainly identifying and

3:31:11resolving performance issues right so

3:31:13when you're loading data or when you're

3:31:15querying data or transforming data and

3:31:17specifically we're going to be looking

3:31:18within the data flow The Notebook so The

3:31:21Spark engine and also SQL queries as

3:31:24well then we're going to look at Delta

3:31:25tables in a bit more detail how we can

3:31:27identify and resolve issues within our

3:31:29Delta tables as you know fabric is built

3:31:31on top of the Delta file format so

3:31:33that's a really important topic to

3:31:35understand and as part of that we're

3:31:36going to be looking at file partitioning

3:31:39as well so what that is what that looks

3:31:41like why you might want to implement

3:31:43file partitioning in your Lake housee so

3:31:46as ever at the end of the video we'll be

3:31:48testing some of your knowledge from the

3:31:50topics that we cover in this lesson and

3:31:53as ever I've got some quite detailed

3:31:55notes that you can use to enhance your

3:31:58vision available in our school Community

3:32:00I'll leave a link to that in the

3:32:02description box below so most of this

3:32:03video I'm going to be diving into Fabric

3:32:05and going through performance

3:32:07optimization in a number of different

3:32:09places in fabric but I just wanted to

3:32:11start by Framing what we mean really by

3:32:14performance optimization in fabric now

3:32:16as you know fabric is a very diverse

3:32:18tool so when we talk about performance

3:32:21really we need to get a bit more

3:32:22specific about well what are we talking

3:32:24about it could be data flow performance

3:32:27could be a SQL script in your data

3:32:29warehouse or we could talk about the

3:32:30Delta files and optimizing how they get

3:32:33written and read in our one Lake in our

3:32:36lake houses as well and I would make the

3:32:38distinction here between identifying

3:32:40performance issues and then resolving

3:32:42them it's kind of like a two-step

3:32:43process first we need to know how to

3:32:45identify performance issues in each of

3:32:48these tools and then we need to think

3:32:50about how we can possibly resolve these

3:32:51issues after we identify them and you'll

3:32:53notice for each of the fabric items how

3:32:56we identify performance issues is going

3:32:58to be a little bit different right so

3:33:00for the data flow we're going to be

3:33:01looking at the refresh history and the

3:33:03monitoring Hub and the capacity metrics

3:33:04app and you notice that some of these

3:33:06actually repeat so the monitoring Hub

3:33:08and the capacity metrics app is kind of

3:33:09like a generic place where you can do

3:33:11lots of performance monitoring across

3:33:13Fabric in the data warehouse we have

3:33:15query insights and DMVs Dynamic

3:33:18management views with the Spark engine

3:33:20we have the spark history server and we

3:33:22also have access to quite detailed

3:33:24monitoring in the monitoring Hub as well

3:33:26and when we're identifying performance

3:33:27issues in Delta files there's a number

3:33:29of places you can do that one of them

3:33:31that we're going to be looking at is

3:33:32describe then when it comes to resolving

3:33:34some of these issues well that's where

3:33:36it gets a bit more difficult to Define

3:33:38right because it's normally going to

3:33:39involve some element of refactoring so

3:33:42using different operations in your data

3:33:44flow for example or refactoring your SQL

3:33:47code or refactoring your spark jobs as

3:33:50well now in the data flow we have some

3:33:52specific performance optimization

3:33:54features that is worth going through so

3:33:56we'll be talking a little bit about

3:33:57staging and fast copy as well then when

3:34:00it comes to Delta file optimization

3:34:02we're going to go into a bit more detail

3:34:03about V order optimization file

3:34:06partitioning and also the vacuum and

3:34:08optimize which are two Delta table

3:34:11functions that we can Implement to

3:34:13improve performance with Delta files so

3:34:15that's a bit of an overview of what

3:34:16we're going to be discussing in this

3:34:18lesson now let's dive into Fabric and

3:34:21we're going to begin by looking at some

3:34:22of the generic tools like the monitoring

3:34:24Hub and the capacity metrics app before

3:34:26diving into the data flow data warehouse

3:34:29and the spark notebook in more detail so

3:34:31let's begin okay so just before we jump

3:34:33into fabric for this tutorial I'm just

3:34:36going to start in the school Community

3:34:37here and talk through some of the notes

3:34:39that we have for optimizing performance

3:34:41so here is video 8 optimizing

3:34:44performance we've got this framing

3:34:45performance optimization chart that we

3:34:47spoke about previously but to get us

3:34:48started I want to speak generally about

3:34:50performance monitoring there's a few

3:34:52tools that are quite General they apply

3:34:55across different workloads that are

3:34:57useful to know for the exam and also in

3:34:59fabric generally so we're going to be

3:35:00talking about the monitoring Hub and the

3:35:02capacity metrics app and these are two

3:35:04tools that can be used to monitor

3:35:06performance for a wide variety of

3:35:08operations in fabric so the monitoring

3:35:10Hub is the first one so let's start by

3:35:12looking at the monitoring Hub and if I

3:35:14just flick over to pobi here obviously

3:35:17to access the monitoring Hub you've

3:35:19probably seen it here it's in the left

3:35:20hand toolbar here you got this big

3:35:22button for monitoring Hub and within

3:35:24here we can basically have a look at the

3:35:26runs of a lot of different item types

3:35:29right so you can see semantic model

3:35:30refreshes notebooks so this is going to

3:35:32be a spark session and for each item

3:35:34type we have different logging that gets

3:35:37exposed in this monitoring Hub now the

3:35:39notebook we're going to have a look at

3:35:40in a bit more detail because that's

3:35:41exposing spark log information you can

3:35:44also see like data flow Gen 2 we can

3:35:46have a look at whether Those runs have

3:35:48succeeded or not table loading

3:35:51information in a lake house for example

3:35:53and we can obviously click through into

3:35:55specific items that we care about and we

3:35:57get more information right so this is

3:35:59showing a table load into the bronze

3:36:02Lakehouse and it's showing you the

3:36:03different jobs because this is a lake

3:36:05house is these spark jobs right so it's

3:36:07also going to tell us in the monitoring

3:36:09Hub whether runs have been successful or

3:36:11failed so this particular data flow run

3:36:14we can see that it actually failed right

3:36:16so on the 1st of May at 12:46 I tried to

3:36:21refresh this data flow and actually

3:36:23failed so you can click on view detail

3:36:24and you can get a few more details about

3:36:27what happened here you can't really

3:36:28diagnose what went wrong with that

3:36:30particular data flow to actually get the

3:36:32details of the data flow you have to

3:36:33actually go into the data flow itself

3:36:36which is this one here and then we can

3:36:38click on these three dots here click on

3:36:39the refresh history so for data flows if

3:36:41you want to actually debug what went

3:36:43wrong this is obviously pretty poor data

3:36:45this but you can click on the individual

3:36:47runs and get more detailed information

3:36:49about what's going wrong here so here we

3:36:51can see it's actually this activity here

3:36:54that's failed we can click on that and

3:36:55then get more information about why

3:36:57specifically that column or that data

3:37:00set can't be refreshed so the next

3:37:02general tool that we can use to monitor

3:37:05performance and resource consumption

3:37:08within our fabric capaces is the

3:37:10capacity metrics app and the capacity

3:37:13metrics app and I'll leave a link you

3:37:15can you can obviously get the install

3:37:16instructions here if you have never

3:37:18installed this before this is a powerbi

3:37:21app that you install within your fabric

3:37:23environment you give it your capacity ID

3:37:26and again the capacity settings is where

3:37:28you'll find that capacity ID and then

3:37:30it's going to bring you through to this

3:37:31kind of capacity metrics app now the

3:37:34capacity metrics app is split into two

3:37:36sections we have compute and storage and

3:37:40this obviously lines up with how fabric

3:37:42is build right you're build partly on

3:37:44storage so the amount of storage you

3:37:46have in fabric plus the resources that

3:37:49you consume during compute so if you've

3:37:51already installed the capacity metric

3:37:53app you can find it in the powerbi

3:37:55experience go to apps and then you

3:37:58should see it there the Microsoft fabric

3:38:00capacity metrics app now as a mentioned

3:38:01there's two tabs to this report it's

3:38:03compute and storage on the compute page

3:38:06so starting at the top left we can see

3:38:08the capacity unit spend for particular

3:38:12item types in Fabric and we can also

3:38:14break it down by duration

3:38:16different operations and by user as well

3:38:18we also get this time series of capacity

3:38:21usage over time as a percentage of the

3:38:24total capacity that you have available

3:38:26based on your skew So currently I'm on

3:38:29this trial capacity so this is going to

3:38:30be an X f64 and as you can see I'm not

3:38:33really using barely any of this capacity

3:38:35we've also got these other interesting

3:38:37graphs around throttling so if you're

3:38:39using more than 100% of your capacity

3:38:42usage on that particular capacity this

3:38:45is going to show you where you're

3:38:46throttling and it's also going to show

3:38:48you rejections so if you've got

3:38:50workloads that are being rejected On

3:38:52Your Capacity because you're again over

3:38:55100% And it's currently rejecting

3:38:57workloads then that's going to be

3:38:59exposed here we've also got a graph on

3:39:01overages so again if your capacity is

3:39:04throttled an overage is basically you

3:39:07repaying that capacity usage from your

3:39:09future spend right and for each of these

3:39:11obviously I haven't actually been

3:39:13throttled or you know there's no overage

3:39:16on my actual capacity but there is this

3:39:18explore button that you can drill

3:39:19through to specific events that you want

3:39:21to explore in more detail if that

3:39:23something that's happening on your

3:39:24capacity down below you've got a table

3:39:27of all of the different items in your

3:39:29fabric capacity and the capacity unit

3:39:31seconds that are being used by that

3:39:34specific resource so this is a synapse

3:39:36notebook and we can see that that is the

3:39:38most resource intensive it's used up the

3:39:40most of our capacity unit seconds and

3:39:42it's also got this tool tip where you

3:39:43can look at specific activities and runs

3:39:46of that notebook to dig into a bit more

3:39:48detail there on the storage tab

3:39:50obviously this focuses on the amount of

3:39:52gigabytes of storage in this fabric

3:39:55capacity we can see how it's changing

3:39:57over time we can look at the specific

3:39:59storage by date and we can also look at

3:40:01the top 10 workspaces by bable storage

3:40:04once that's loaded that's what that

3:40:06brings you there here we go okay so next

3:40:08I just wanted to talk about data flows

3:40:10and just to summarize what we looked at

3:40:11previously well if you want to monitor

3:40:14the performance of a data flow well at a

3:40:16high level we can do that within the

3:40:17monitoring Hub but if you want a bit of

3:40:19a lower level data and to understand

3:40:21what's happening within a particular

3:40:22data flow you're going to be wanting to

3:40:24look at the refresh history as I showed

3:40:27you previously for a particular data

3:40:29flow here you can inspect the error

3:40:30messages you get breakdown of the the

3:40:32different sub activities in that load

3:40:35for data flow so that's going to be

3:40:36really important for you to diagnose

3:40:38what's going wrong in a particular data

3:40:40flow if it's not refreshing correctly

3:40:41that's when you where you're going to go

3:40:43to have a look there now there's a

3:40:44couple of features that you need to be

3:40:46aware about in terms of optimizing the

3:40:49performance specifically when we're

3:40:51talking about data flows right so the

3:40:53main one is staging now staging is

3:40:55probably best described using this

3:40:57diagram here so this diagram actually

3:40:59comes from this link here it's the

3:41:01spotlight blog on data flows and it's

3:41:03got some top tips for improving the

3:41:06performance in your data flows so I

3:41:08think to understand what's going on with

3:41:09staging this diagram gives you a pretty

3:41:12good idea so let's start by talking

3:41:14about when staging is disabled so this

3:41:16bottom diagram here right so if staging

3:41:19is disabled all of your transformation

3:41:22in a data flow is going to be done by

3:41:24the data flow engine it's otherwise

3:41:26known as the mashup engine right and if

3:41:28you got a really big data set or you're

3:41:30doing lots of transformation that might

3:41:32not be the most efficient way of doing

3:41:35it so a feature we have available to us

3:41:37to try and improve the speed of doing

3:41:39all these Transformations within a data

3:41:40flow is staging so at the top here when

3:41:43staging is enabled what is going to

3:41:45going to do is it's going to read in the

3:41:46data from the data source and then it's

3:41:48going to immediately write that data

3:41:50into a Lakehouse staging table then it's

3:41:52going to use that Lake housee to perform

3:41:55the transformation right so leveraging

3:41:58the Spark engine rather than the mashup

3:42:00engine to do your transformation then

3:42:03it's going to read the data back into

3:42:04the mashup engine and write it into the

3:42:06destination wherever that might be so a

3:42:08few things to bear in mind here if

3:42:10you've got a lot of data Transformations

3:42:12or you've got very large amounts of data

3:42:14staging is probably going to be a lot

3:42:16more efficient now if you got small data

3:42:18sets it's probably going to be less

3:42:19efficient right because you're going to

3:42:21have to write the data into a lake house

3:42:23transform it into the lake house then

3:42:25write it then the mashup engine picks it

3:42:27up again and writes it to your output

3:42:29destination so there's lots of kind of

3:42:31reading and writing here so on small

3:42:33data sets you probably want to disable

3:42:35staging or not enable staging it's only

3:42:38really when the data set becomes large

3:42:40or you're doing lots of transformation

3:42:41on that data set then we want to enable

3:42:44staging that's what I've of summarized

3:42:45with this sentence here there's a bit of

3:42:47an overhead when you're performing

3:42:49staging and so it doesn't work in all

3:42:51cases only really when your data set is

3:42:53large or you're doing lots of

3:42:54Transformations or both now another

3:42:56feature that they recently announced and

3:42:58therefore might not actually be in the

3:42:59exam yet but it's good to kind of

3:43:01understand know that it exists is fast

3:43:04copy and the way that I think fast copy

3:43:06works is that under the hood it uses the

3:43:08same technology is the data pipeline

3:43:10copy data activity rather than the data

3:43:13flow technology basically as I'm I

3:43:15mentioned it's still a preview feature

3:43:16and it's relatively recent so it might

3:43:17not actually be in the exam yet but you

3:43:19know if you're using data flows in the

3:43:21real world and you're struggling with

3:43:23performance it's worthwhile enabling

3:43:25fast copy just to give it a go see how

3:43:27it impacts the performance in your data

3:43:29flows okay so next up we're going to

3:43:31move on from data flows and now we're

3:43:33going to focus on SQL so we have the SQL

3:43:37engine within the data warehouse and

3:43:39also the SQL endpoint of the lake house

3:43:41as well and we're going to have a look

3:43:42at how we can diagnose and then optimize

3:43:46the performance of SQL scripts and as I

3:43:49mentioned before the capacity metrics

3:43:51app can give you a good kind of high

3:43:52level overview of the resource

3:43:54consumption of specific operations that

3:43:57you're doing within your data warehouse

3:43:59but we have a lot more functionality

3:44:01within the data warehouse to actually

3:44:03explore and identify things like long

3:44:05running queries frequently used queries

3:44:08next we're going to talk about Dynamic

3:44:10management views or DMVs and if you take

3:44:13a look in the data warehouse under the

3:44:15CIS schema there's obviously a lot of

3:44:18different views in there that we can use

3:44:20for database management in general and

3:44:22there's three main ones really for

3:44:25understanding the live SQL query life

3:44:27cycle okay so things that are currently

3:44:29going on in your data warehouse that you

3:44:31need to be aware of and getting some

3:44:33insights about what's happening there

3:44:35and these are exact connections exact

3:44:37sessions and exact requests and these

3:44:39are related in this way here so we've

3:44:41got a bit of data model here so whenever

3:44:43you start a query execution in the data

3:44:45warehouse it's going to start up a

3:44:47session the session is going to have a

3:44:49one toone relationship normally with the

3:44:52connections right so it's going to

3:44:53create a connection between your data

3:44:55warehouse and the underlying seal engine

3:44:57so that's what a connection is going to

3:44:59show you and then you're going to have

3:45:00many requests normally for each of these

3:45:03connections right so using these three

3:45:06commands these three dmbs we can begin

3:45:09to build a bit of a picture about who

3:45:11and how your data warehouse is being

3:45:13queried and these DMV are going to help

3:45:15you answer questions like who is the

3:45:18user running the current session when

3:45:20was the session started by the user

3:45:22what's the IDE of the connection to the

3:45:24data warehouse that is running a

3:45:25particular request how many queries are

3:45:28actually currently active and which

3:45:30queries are long running so you begin to

3:45:32build a bit of a picture about how your

3:45:35data warehouse is being queried who's

3:45:37querying it what they're doing and you

3:45:39know the performance of those queries

3:45:41now the DMV is quite a lowlevel view

3:45:44right we can do lot of information here

3:45:46can merge these tables in different ways

3:45:48to get more and more information so

3:45:50alongside the DMVs Microsoft also expose

3:45:53query insights so query insights is in a

3:45:56different schema so if we have a look

3:45:58here at this particular example of data

3:46:01warehouse we've got the Cy which is our

3:46:02DMVs what contains our DMVs as well as

3:46:05other database management system

3:46:07generated views we've also got query

3:46:09insights so in here we've got these four

3:46:12views that give us basically more

3:46:15userfriendly abstractions over the DMVs

3:46:18right so it's going to expose things

3:46:19like frequently run queries long running

3:46:22queries and you don't have to actually

3:46:24perform those joins of the underlying

3:46:26DMVs to get this information it just

3:46:28exposes them right here and so we're

3:46:30just going to focus on three of these

3:46:32query insights views so exact requests

3:46:35history it's going to return information

3:46:37about each completed SQL request on that

3:46:40particular dat Warehouse frequently run

3:46:42queries is obviously going to give you

3:46:44information about the most frequently

3:46:45run queries and long running queries is

3:46:47basically going to return you

3:46:48information about queries by execution

3:46:51time so this long running queries is

3:46:52going to be really useful to as the name

3:46:54suggests identify queries that are

3:46:56running for a long time and it might be

3:46:57causing performance issues in your data

3:47:00warehouse now one thing to note here is

3:47:02that if you're coming from a SQL Server

3:47:05background when we're talking about

3:47:06performance optimization a really

3:47:08important tool there is the query plan

3:47:10right so currently I don't think it's

3:47:11possible to expose the query plan for

3:47:14particular SQL query but I do think they

3:47:16are planning to support that in the

3:47:18future next up I want to talk about

3:47:20identifying performance issues with the

3:47:22Spark engine and you'll notice that

3:47:23we're back in the monitoring Hub and I

3:47:26just want to look at one of the item

3:47:28details here for this specific notebook

3:47:30so I've been running this notebook it's

3:47:32called Delta optimization and when we

3:47:34click through on this item we get a

3:47:35really detailed analysis of the

3:47:38different jobs that have been run in

3:47:40this notebook we can see that all of

3:47:42these have succeeded you can see which

3:47:44the duration of particular job the data

3:47:47that's been read and written for that

3:47:49particular job so this is a really good

3:47:51place to go if you want to understand

3:47:53what's actually happening when you click

3:47:55run in a spark notebook what's happening

3:47:58under the hood and the success or

3:48:00failure of each of the individual spark

3:48:01jobs now within the monitoring Hub we've

3:48:03also got this link through to the spark

3:48:06history server and as you can see from

3:48:07the UI here we've actually switched from

3:48:09a fabric tool to a generic spark tool

3:48:12right so here you're going to get a lot

3:48:14more detailed information about specific

3:48:16sparkk jobs that you're running you can

3:48:18look at graph so if you do want a bit of

3:48:20a query plan look at the different

3:48:22stages in execution of a particular

3:48:24spark job you can have a look at that

3:48:26here so if you're looking for really

3:48:27fine grain control and Analysis of

3:48:30what's going on on your Spark engine

3:48:32you're going to come to the spark

3:48:33history server now to actually interpret

3:48:36and understand what's going on here

3:48:38would be a whole series in itself I

3:48:40don't think you need to know the the

3:48:41nitty-gritty details of actually what's

3:48:42going on in the spark history server for

3:48:44the dp600 exam just to understand you

3:48:47know what's possible in the spark

3:48:49history server what does it log what can

3:48:51you monitor there I think that's good

3:48:52enough for the dp600 exam so finally in

3:48:55this tutorial I just want to focus on

3:48:57Delta table optimization now as you

3:48:59probably know fabric is built on top of

3:49:01Delta tables so this is a really

3:49:03important topic to understand and the

3:49:05Delta file format is great but it can

3:49:07lead to poor performance and Bloated

3:49:09storage sizes if we're not managing

3:49:12those Delta files correctly now this is

3:49:14a topic that can run very deep like a

3:49:16lot of the topics that I mentioned today

3:49:18we're just going to go through some of

3:49:18the basics of what you need to know for

3:49:20the exam and to do that we're going to

3:49:22be going through a spark notebook so

3:49:25this is the spark notebook that we're

3:49:26going to talk through now I'm going to

3:49:27start by just exploring the problem in a

3:49:29little bit of detail and I think a

3:49:31really good way of understanding what

3:49:32the problem or a problem that can arise

3:49:35with Delta files is this visual here and

3:49:37this visual comes from this blog post

3:49:39here by Sid daba and it's around

3:49:41efficient data partitioning with

3:49:43Microsoft fabric best practices and

3:49:45implementation guide and I've left a

3:49:47link to that in this notebook The

3:49:49Notebook is obviously available in the

3:49:51community here as well and I've also

3:49:53left a link to it here efficient data

3:49:55partitioning here as well so if you have

3:49:58one big file that's 10 gigabytes one

3:50:01paret file then the Spark engine is

3:50:03going to struggle to process that right

3:50:05because as you know spark is a

3:50:07distributed processing engine which

3:50:09means it works best when it splits your

3:50:11file your paret files into smaller

3:50:14chunks then processes these chunks in

3:50:16parallel but to make that possible we

3:50:18need to partition our data and

3:50:20partitioning is basically the process of

3:50:22converting one really big or several

3:50:24really big files into more manageable

3:50:26chunks so that our data can be

3:50:28transformed in parallel basically by The

3:50:30Spark engine so in this notebook we're

3:50:32going to start by looking at file

3:50:33partitioning and then look at some other

3:50:35methods for optimizing Delta tables in

3:50:37doing so this can help improve the read

3:50:40and retrieval performance so by doing so

3:50:43there's no need to scan through millions

3:50:45and millions of rows in your really big

3:50:46paret file if you got a good

3:50:48partitioning system in place it's going

3:50:50to speed up the read performance of your

3:50:52queries and it's also going to improve

3:50:54your transformation performance right as

3:50:56we mentioned before because partitions

3:50:58can be transformed in parallel so file

3:51:00partitioning I'm going to walk through a

3:51:01bit of a demo here and I've used a demo

3:51:04file and it's available on this website

3:51:06here but it's just a parket file what

3:51:07I've done is I've just put it in the

3:51:09files location this one here is called

3:51:11Flights 1M parket and before we get

3:51:14started it's just bear in mind the best

3:51:16practices on Delta Lake partitioning

3:51:19right and I've left a link here this is

3:51:21on the Delta Lake website where they

3:51:23give some general best practices about

3:51:25managing Delta files okay and one of

3:51:27them is around choosing the right

3:51:29partition column now most commonly it's

3:51:32done by date so if you've got time

3:51:34series data you've got some dates in

3:51:36your data set and it's common to use

3:51:37your date for partitioning but there's

3:51:40two kind of rules of thumb to bear in

3:51:41mind when we're talking about

3:51:42partitioning and deciding what col

3:51:44colums you want to partition on so in

3:51:46general we don't want to choose our

3:51:48partitioning column to be something of

3:51:50really high cardinality so say for

3:51:51example you have a column of user ID

3:51:54well that's going to be really high

3:51:55cardinality right every row is basically

3:51:57going to be unique so if you've got a

3:51:59million rows that's going to be a really

3:52:01bad partitioning strategy right because

3:52:03you're going to get a million different

3:52:04partitions and they're all going to be

3:52:05really small and as a general rule of

3:52:07thumb the amount of data that should be

3:52:09in each partition should be around one

3:52:11gigabyte that's kind of like the the

3:52:12good balance of what you should be

3:52:15aiming for with each partition so let's

3:52:17take a look at file partitioning and how

3:52:18to actually implement it in Fabric in a

3:52:21spark notebook so I'm going to begin by

3:52:23just reading in that parquet file our

3:52:25flights 1M parket file I'm just reading

3:52:28it into a data frame I'm just displaying

3:52:30it here and you notice we've got this

3:52:31date column here it's going to be useful

3:52:33for our partitioning strategy we're

3:52:35going to be partitioning on date and

3:52:37it's got some other items here that we

3:52:39don't really care about for this

3:52:40tutorial so before we do some

3:52:41partitioning and we write this file into

3:52:44into our Lakehouse using partitions

3:52:46we're going to do a bit of preparation

3:52:48and specifically we're going to add some

3:52:50columns into our data frame we're going

3:52:52to use DF with columns to add more than

3:52:54one column at a time and we're going to

3:52:56pass in this dictionary object here and

3:52:58each uh part of that dictionary each key

3:53:01is going to be the new column name and

3:53:03the value is going to be how we're

3:53:04actually Computing that value so here

3:53:07I'm just reading in some P spark SQL

3:53:09functions to extract the year the month

3:53:12and the day of month of that date field

3:53:14right that we looked at previously so if

3:53:15we just run that and if we just display

3:53:17the results here okay so now we've got

3:53:18this transformed data frame object right

3:53:21and so you can see here there added in

3:53:23this year month and day column and we

3:53:27can use these in our partitioning

3:53:29strategy now so you can see here from

3:53:31our code cell that we've actually

3:53:33written three different write modes here

3:53:36so the first one is going to write

3:53:37without partitions so this is just the

3:53:39normal saving of the table into a

3:53:42Lakehouse table from our parket file

3:53:44we're going to put it in tables and

3:53:45we're going to call the table flights

3:53:47not partitioned the second one we're

3:53:49going to do is we're going to call

3:53:51Partition by and we're going to

3:53:53Partition by the year and the month and

3:53:55we're going to save this one into

3:53:56another table called Flights partitioned

3:53:58and then as a third example we're going

3:54:00to write with some more partitions so

3:54:02we're going to be partitioning our data

3:54:05into smaller partitions here and again

3:54:06we're calling Partition by but this time

3:54:08we're passing in the year the month and

3:54:10the day so these are going to be more

3:54:11fine grained and I'm going to save that

3:54:13into a table called flights partitioned

3:54:15daily let's just run those okay so now

3:54:18our cell has been executed all of our

3:54:20spark jobs have concluded let's just

3:54:22have a look at this so yeah now you can

3:54:24see in our lake house tables we've got

3:54:26flights not partitioned flights

3:54:28partitioned and flights partitioned

3:54:30daily so we got three different tables

3:54:31here all with different partitioning

3:54:33strategies so let's inspect those and

3:54:36see what's going on okay so I've got

3:54:37three different cells here and you'll

3:54:39notice we're using the SQL so we're

3:54:41using spark SQL here and we're calling

3:54:43describe detail on that particular table

3:54:46it's going to inspect this table it's

3:54:48going to describe what's going on there

3:54:50and some of the results give us a bit of

3:54:51a picture as to how this table has been

3:54:54written into one L so if we inspect the

3:54:57results here and this is our flights

3:54:59partitioned we can see that the

3:55:00partition columns here are year and

3:55:03month which makes sense because this is

3:55:04the year and month one in flights

3:55:06partition and we can see the number of

3:55:07files it's created so the number of

3:55:09paret files we've got here is two next

3:55:11if we compare that to what we get when

3:55:13we look at flight not partitioned we can

3:55:16see that we got zero partition columns

3:55:17and we got one file so everything's just

3:55:19been written into one file in that

3:55:21instance here and for the daily one

3:55:23again if we inspect what's going on here

3:55:25we can see that we've got year month and

3:55:27day partition columns and here we've got

3:55:3059 files so here it's been partitioned

3:55:33we've broken up that data into 59

3:55:36smaller chunks of data now that might be

3:55:38too fine grained because if we scroll

3:55:40back up to the top here one thing that I

3:55:42did mention is that it's a bit of a

3:55:43balancing

3:55:44this because if your files are too big

3:55:47the spark engine's going to have

3:55:48performance issues but there's also the

3:55:50small file problem if your file sizes

3:55:52are under that gigabyte then you know

3:55:55that's also going to cause a lot of

3:55:56issues it's going to have to work harder

3:55:58you're going to have to go through the

3:55:59operation 59 times rather than two so

3:56:02again it's a bit of a balancing act

3:56:04trying to get a good partitioning

3:56:06strategy for your Delta tables but for

3:56:08the purposes of the exam I think it's

3:56:10worthwhile understanding the Syntax for

3:56:13creating part partions like so and then

3:56:16analyzing different partitions using

3:56:18this describe method spark SQL method

3:56:20okay so next I just want you to talk

3:56:21about V order optimization now V

3:56:23ordering is a Microsoft proprietary

3:56:26algorithm and it basically changes the

3:56:28structure of your parket file and what

3:56:31I've put here is is kind of providing a

3:56:32bit of special source so it does some

3:56:34special sorting compaction compression

3:56:37of that parquet files right ultimately

3:56:39to improve the read performance of these

3:56:42parket files across all of the different

3:56:44engines in fabric now whilst the actual

3:56:47algorithm is proprietary it's only used

3:56:49by Microsoft the output of like the

3:56:52parket file is fully kind of Open Source

3:56:55aligns to the traditional parket

3:56:56standards so you can actually read

3:56:58vorded parket files wherever you can

3:57:01read normal parket files that's not a

3:57:03problem at all Now by default V ordering

3:57:06so this algorithm that you used to write

3:57:08parket files it's enabled by default in

3:57:11the fabric spark runtime and you can

3:57:13check that by running this spark comp

3:57:16get so we're looking at the

3:57:17configuration and you can see that here

3:57:19is actually returning true now it can

3:57:21actually be manually disabled if you

3:57:23want it to so we can set the spark

3:57:25configuration by passing in this

3:57:27specific property here spark SQL paret V

3:57:30order enabled to false and then we can

3:57:32reenable it by doing the opposite right

3:57:34putting it to True again so for the exam

3:57:36you might be asked about how do you know

3:57:38whether a particular notebook or a

3:57:41particular spark environment has v order

3:57:43enable well that's this one or how do

3:57:45you enable it or disable it in a spark

3:57:48notebook as well that could be a common

3:57:49question that you might get asked okay

3:57:51just finally I just want to mention a

3:57:52few more Delta table maintenance and

3:57:55optimization techniques so there's a few

3:57:58that come from the actual Delta format

3:58:00itself so we have this function called

3:58:02optimize which is a Delta Lake method

3:58:04that performs bin compaction it it can

3:58:06basically improve the speed of your read

3:58:09queries so if you got multiple small

3:58:11files it's basically going to coals

3:58:14basically mean joining small files into

3:58:16larger files vacuum is another function

3:58:18that we can run and it basically

3:58:20involves removing files that are no

3:58:22longer referenced by a Delta table and

3:58:25then on the spark side there's two they

3:58:27might at least want to be familiar with

3:58:28is coales so as we mentioned when we're

3:58:30talking about optimize optimize is

3:58:32basically the what that's doing under

3:58:34the hood is calling coales and coales is

3:58:36a spark method or basically reducing the

3:58:39amount of partitions in your Delta table

3:58:41so if you've got 100 partitions for a

3:58:43particular file you can coals that Delta

3:58:45table into 10 partitions so coals is a

3:58:48pretty efficient way of grouping

3:58:50partitions into a smaller number of

3:58:53partitions right and I say it's quite

3:58:54efficient because it doesn't require a

3:58:57shuffle of the data it's just grouping

3:58:59partitions together doesn't actually

3:59:01reorganize within particular partitions

3:59:03your data now repartition is similar to

3:59:06coales but it's actually less efficient

3:59:08because it involves breaking up your

3:59:10existing partitions and then creating

3:59:11new partitions and because of this you

3:59:14create either more or less partition

3:59:16it's basically just restructuring how

3:59:18your partitions are created and it does

3:59:20involve some shuffling involves breaking

3:59:23up of your existing partitions and

3:59:24repartitioning them now you might be

3:59:26thinking what's the difference between

3:59:27the spark functions and the the Delta

3:59:28optimize well the Delta optimize has a

3:59:30few kind of things working under the

3:59:32hood that makes it more efficient for

3:59:34number one it's item potent so if you

3:59:36run it repeatedly it's not going to

3:59:38reoptimize files that have already been

3:59:40optimized whereas repartition is going

3:59:42to always repart partition your files

3:59:45you can keep on running this again and

3:59:46again and it's never going to get more

3:59:48efficient it's always going to

3:59:49repartition the files that's one

3:59:51difference between repartition and

3:59:53optimize and you can also run optimize

3:59:55on specific partitions in your data set

3:59:57whereas repartition that's kind of All

4:00:00or Nothing approach you have to

4:00:01repartition your whole table in one go

4:00:04so if you're looking for a bit more fine

4:00:05grained optimization you're going to be

4:00:07want to using the Delta optimize okay so

4:00:10let's just round off the video here by

4:00:12testing some of your knowledge of the

4:00:14things the topics that we've covered in

4:00:15this lesson question one a client you're

4:00:17working with wants to reduce the SKU of

4:00:19their fabric capacity from an F-16 to an

4:00:22f8 to save some money they want to find

4:00:24the most resource intensive workloads

4:00:27and optimize them to use less capacity

4:00:29unit seconds where should they look to

4:00:31find this information is it a the

4:00:32monitoring Hub B capacity metrics app C

4:00:35query insights D spark history server or

4:00:39e the one Lake Hub pause the video here

4:00:42have a little think and I'll reveal the

4:00:43answer to you shortly okay so the answer

4:00:45here is B the capacity metrics app we're

4:00:48talking about resource intensive

4:00:50workloads and our capacity is where

4:00:52we're going to get those resources and

4:00:54specifically it's going to tell you

4:00:56which workloads are the most resource

4:00:58intensive are the workloads that are

4:00:59going to use more of your capacity units

4:01:02seconds right so the answer is going to

4:01:04be your capacity metrics app we can look

4:01:07at all of our spark jobs our data

4:01:09warehouse operations and our data flows

4:01:12and we can come to conclusions about

4:01:14which of these are good candidates for

4:01:16refactoring or optimization now all of

4:01:19the others they might be useful for

4:01:21understanding performance of specific

4:01:23workloads within fabric but the capacity

4:01:25metrics app is the only one here that

4:01:27converts that into capacity units right

4:01:30and that's the important part of the

4:01:31question to understand question two you

4:01:33noticed one of your data flow Gen 2 runs

4:01:35failed to refresh last night where would

4:01:37you go to find out why a particular data

4:01:39flow might have failed a particular Run

4:01:41is it a the capacity metrics app B the

4:01:44monitoring Hub C power query error Hub D

4:01:47data flow refresh history or E the data

4:01:50pipeline run history so the answer here

4:01:52is D the data flow refresh history is

4:01:55where you're going to go to analyze

4:01:57error messages and debug particular runs

4:02:01of a data flow now the power query error

4:02:03Hub that doesn't actually exist I made

4:02:05that up the data pipeline run history

4:02:08but we're not talking about data

4:02:09pipeline here so it's not going to be

4:02:10that the monitor the monitoring Hub will

4:02:12give you some information so it will

4:02:14tell you whether a particular run has

4:02:16failed or succeeded but it doesn't give

4:02:18you more detailed information about

4:02:20error messages and things like that and

4:02:22the capacity metrics app is not going to

4:02:24tell you that answer either so the

4:02:25answer here is D data flow refresh

4:02:28history question three when talking

4:02:30about Delta table optimization which of

4:02:32the following operation removes old

4:02:35files no longer referenced by a Delta

4:02:37table log is it a v order optimization B

4:02:40Zed ordering C vacuum d optim or E bin

4:02:45compaction so the correct answer here is

4:02:47C vacuum so as we mentioned previously

4:02:50the vacuum command does exactly as it's

4:02:53mentioned there basically removes old

4:02:55files that no longer referenced by a

4:02:57Delta table log so the correct answer

4:02:59here is C question four which of the

4:03:02following statements about V order

4:03:04optimization is false a v order

4:03:07optimization is enabled by default in

4:03:10the fabric spark runtime b v order can

4:03:13be enabled during table creation using

4:03:15table properties c a table can be both V

4:03:18ordered and Zed ordered d v order

4:03:20improves the read performance for parket

4:03:22files e v order speeds up the right time

4:03:25of a parket file so the correct answer

4:03:27here is e so the question asked which of

4:03:30the following statements is false and so

4:03:32e is actually false V order does not

4:03:35speed up the right time it actually

4:03:36increases the right time of a parket

4:03:39file the benefit of V ordering comes in

4:03:42the read performance right that's why we

4:03:43do it takes a bit longer to write these

4:03:46files but it massively improves the read

4:03:48performance across any of the engines

4:03:50that you might want to use it in fabric

4:03:52all of the other options are true so it

4:03:54is enabled by default B if it's not

4:03:57already enabled in your spark

4:03:59environment it can be enabled for

4:04:00specific tables during table creation

4:04:03using table properties you can actually

4:04:04optimize both V ordering and Zed

4:04:07ordering for a particular table or PAR

4:04:09paret file within that table at the same

4:04:11time and D is also true because we

4:04:14mentioned that V ordering does improve

4:04:16the read performance for parket files

4:04:19that's why we do it question five you

4:04:21want to analyze long runn queries in a

4:04:23fabric data warehouse what's the minimum

4:04:26workspace role you need to run the

4:04:28following query select start from query

4:04:31insights. longrun inqueries is it a

4:04:34admin B member C contributor or D viewer

4:04:38so the answer here is C contributor to

4:04:40be able to run a query insights query

4:04:43you know autogenerated views that give

4:04:45us information about long running

4:04:47queries in this example you need to have

4:04:49a workspace role of contributor

4:04:51obviously the question asked for the

4:04:52minimum workspace role you can also run

4:04:55these queries with admin or member but

4:04:57the minimum workspace role would be

4:04:59contributor so if you have a viewer role

4:05:01you can't run these queries in a data

4:05:04warehouse congratulations you've

4:05:05completed the biggest section of the

4:05:08exam we're well over halfway now so in

4:05:10the next lesson we're going to be

4:05:12starting the third section section of

4:05:14the exam which is all about building

4:05:16semantic models well done and I'll see

4:05:18you in the next video hello and welcome

Design and build semantic models

4:05:20back to the channel today we're

4:05:22continuing our dp600 exam preparation

4:05:25course and we're up to video 9 designing

4:05:28and building semantic models now as you

4:05:31can see from the course plan we're

4:05:32making really good progress we just got

4:05:34a few more important sections to look at

4:05:36and today we're going to be starting

4:05:37powerbi and semantic modeling part of

4:05:40the exam specifically we're going to be

4:05:41looking at the different storage mod

4:05:43modes import mode direct query direct

4:05:46Lake we're also going to be looking at

4:05:47composite models and what we mean by

4:05:49that including aggregations as well

4:05:52we're going to be looking at the large

4:05:53format data set and then we're going to

4:05:55be digging into a bit of a practical

4:05:57example in palbi desktop defining

4:06:00different Dax measures we're going to be

4:06:02looking at functions iterators table

4:06:04filtering windowing information

4:06:06functions and some of the other more

4:06:08advanced or more recent features as well

4:06:10including calculation groups Dynamic

4:06:12strings field parameters that kind of

4:06:13thing as well as ever we're going to

4:06:15have five sample questions at the end of

4:06:18this video to test your knowledge now

4:06:20bear in mind that I wouldn't classify

4:06:21myself as a powerbi developer I used

4:06:23powerbi quite a lot about six seven

4:06:26years ago recently I've been more

4:06:27focused around data engineering data

4:06:29science so I'll try and explain these

4:06:31Concepts as best possible but I

4:06:32definitely recommend doing your own

4:06:33research I'll leave a link to a lot of

4:06:35really good resources for this kind of

4:06:37thing in the school community so that

4:06:40you can go in a bit more detail there so

4:06:42first up we're going to be looking at

4:06:43storage modes now I've mentioned this

4:06:46fair amount on the channel and lots of

4:06:48people talk about storage modes within

4:06:50powerbi and fabric so we're going to do

4:06:52a bit of a revision what we mean by

4:06:54different storage modes some of the

4:06:55advantages and disadvantages and how you

4:06:58can choose between them so this is the

4:07:00diagram that exists in the documentation

4:07:02I think it does a pretty good job at

4:07:04framing out three different storage

4:07:07modes connection modes you might also

4:07:09hear it called as well now if you've

4:07:11worked in powerbi for a while you're

4:07:13definitely going to be familiar with

4:07:15both the import mode this one in the

4:07:17middle plus the direct query mode

4:07:20potentially you might have used that as

4:07:22well at the top there so for those not

4:07:24coming from the PBI background the the

4:07:26import mode is basically going to take a

4:07:29copy of your data from source and load

4:07:31it into powerbi so it's storing a a copy

4:07:35of all your data in the actual powerbi

4:07:37data model itself now that makes it

4:07:38really fast when you're building ports

4:07:41cuz your data is right there now it does

4:07:42have some limit ations in that because

4:07:45you're copying your data into powerbi

4:07:46there are some limitations around the

4:07:47size right that's one of the main

4:07:49limitations on import mode and also

4:07:51because you're having to do this lift

4:07:53and shift import on a schedule normally

4:07:55your data in your powerbi data model is

4:07:58not going to be always up to date

4:07:59because you're going to have to do it

4:08:00every hour potentially there is the

4:08:02chance that your data will become a

4:08:04little bit stale on the other hand we

4:08:05have direct query so moving up to this

4:08:07top one here in direct query whenever

4:08:10the user views a particular visual

4:08:12particular report page power actually

4:08:15sends a query back to your Source it's

4:08:17going to perform a query of that source

4:08:20and then get back fresh data so in doing

4:08:23so it's near real time so that's one of

4:08:25the benefits of direct query some of the

4:08:27downsides of that mean that we can't

4:08:29actually perform much transformation on

4:08:31that data it has to be transformed in

4:08:32the source right you have to create

4:08:34views and tables that already are

4:08:36transformed in practice direct query can

4:08:38be really slow because you've got to do

4:08:41that query every time you go back back

4:08:43to the original data source and for the

4:08:45user they're sitting and waiting for

4:08:47their visuals to up update every time

4:08:49you click a filter or you change the

4:08:52page it's going to take time it's going

4:08:53to be pretty bad user experience so the

4:08:55final one we've got here is the new one

4:08:57that came with powerbi it's called

4:08:58direct Lake and direct Lake creates a

4:09:01connection between what you create in

4:09:04your P report and the underlying parket

4:09:06files so it's going to read the parket

4:09:09files in your one Lake environment

4:09:12directly so let's just have a look at

4:09:14potential reasons why you might want to

4:09:16choose import mode or direct late mode

4:09:18or direct careering mode so some of the

4:09:20key considerations well as we mentioned

4:09:23for import mode it's going to be really

4:09:25good when your data is small enough to

4:09:28fit within a palbi data model including

4:09:30large format semantic models which we're

4:09:32going to talk about in a short while

4:09:34another good use case for import mode is

4:09:36when you want really good read

4:09:38performance in your dashboards you know

4:09:40like interactivity and a good user

4:09:42experience for your dashboard users when

4:09:44you don't have requirements for near

4:09:46real time updates and if you want to use

4:09:49calculated columns or calculated tables

4:09:52if you want to be doing that kind of

4:09:53stuff then you're going to be want to

4:09:54using import mode and if you want to

4:09:56combine data from multiple different

4:09:58data sources that's another good use

4:10:00case for import mode which me leads me

4:10:02nicely into when you would choose direct

4:10:05Lake mode well the first limitation

4:10:07really is that your data has to be

4:10:08stored in one fabric data store so in a

4:10:12lake house or a data warehouse for

4:10:13example you can't use direct Lake mode

4:10:16to access data across lots of different

4:10:18data stores so that's one limitation the

4:10:21prime use cases when your data set is

4:10:23really really big we're talking tens or

4:10:25hundreds of gigabytes here now obviously

4:10:27that would be too big for most import

4:10:29mode models but that use case is really

4:10:32really good in direct late mode now with

4:10:35direct late mode it will require a

4:10:37little bit of a different skill set

4:10:39within your team because you're going to

4:10:40have to do a lot of the data modeling

4:10:42more up dream in your Lake housee in

4:10:44your data warehouse right because you

4:10:46need to materialize parquet files that

4:10:49can be read by the direct Lake mode

4:10:51connection and typically what that means

4:10:53is your data transformation your data

4:10:55modeling is going to have to be done in

4:10:56your lake house so either using spark or

4:10:59TC call and that might be a slightly

4:11:01different skill set to what you might

4:11:03have in a an import mode powerbi team

4:11:06for example okay so finally let's just

4:11:08talk about when you might want to choose

4:11:10direct query mode for your semantic

4:11:12models well again if you need near real

4:11:15time updates that's going to be a really

4:11:16important one again you're going to be

4:11:18needing to do your trans data

4:11:20Transformations more Upstream so in your

4:11:22data source wherever that might be and

4:11:24direct query is also important part of

4:11:26what we call composite models which

4:11:27we're going to look at in more detail

4:11:29shortly so let's look at composite

4:11:31models then so a composite model

4:11:33combines one or more of these different

4:11:35connection modes that we just discussed

4:11:37previously now commonly this is a direct

4:11:40query fact table and import mode

4:11:43Dimension tables because if we think

4:11:45about the common characteristics of a

4:11:49fact table versus a dimension table well

4:11:52in our fact table we're going to have a

4:11:54lot of rows normally a lot of data could

4:11:56be millions or even billions of rows and

4:11:59it's likely to be updated very often

4:12:02maybe every minute or every second even

4:12:04in some oltp transaction processing type

4:12:07fact tables they could be hundreds and

4:12:09hundreds of records every second now

4:12:11because of those characteristics direct

4:12:13query can be a good match for that type

4:12:16of data set right because you get near

4:12:19real time updates so on data sets that

4:12:21are changing very often and very fast

4:12:24direct query gives you that near real

4:12:25time access to fresh data right now with

4:12:28the dimensions they might be changing a

4:12:30lot slower so it makes sense to use

4:12:33import mode for those Dimensions you

4:12:35know your product table might be updated

4:12:38once per day for example so a direct

4:12:41query connection mode in that example

4:12:43wouldn't really make too much sense

4:12:44because you're going to lead to user

4:12:45experience issues on the front end for

4:12:48that table and you're not going to be

4:12:49getting much benefit because the

4:12:50underline data isn't really changing

4:12:52very often now another benefit of

4:12:54composite models is that they provide a

4:12:56way to model many to many relationships

4:12:59without the need for bridge tables as

4:13:01well so next up we're going to look at

4:13:03aggregations now in this context we're

4:13:05talking about a specific powerbi feature

4:13:07for managing aggregation and

4:13:10specifically what this feature does in

4:13:12powerbi is it takes a really large data

4:13:15set it creates a aggregation either

4:13:17automatically or user generated you can

4:13:20assign the aggregation that you want to

4:13:21build and then it caches the actual

4:13:23aggregation so when you have really

4:13:25large data models it can improve the

4:13:27performance because you're caching the

4:13:29aggregation rather than loading in you

4:13:30know a really long fact table for

4:13:32example and because of that they are

4:13:34often used in conjunction with composite

4:13:37models now either you can create the

4:13:39actual aggregation itself in your data

4:13:41source and then just pull into your

4:13:43powerbi data model or you can bring it

4:13:45into powerbi and then use the power

4:13:47query engine to create an aggregation

4:13:50which then gets loaded into your powerbi

4:13:53engine now as I mentioned there's two

4:13:54really types of aggregation we have

4:13:56userdefined aggregations and that's

4:13:58using the the manage aggregations

4:14:00dialogue in P your desktop to Define

4:14:02these aggregations for a specific

4:14:04aggregation column and then you choose

4:14:06how you want to summarize do you want a

4:14:08Min a Max that kind of thing plus the

4:14:10detail table and detail column

4:14:12properties now if you have access to a

4:14:14premium subscription powerbi then you

4:14:16also get automatic aggregations and

4:14:18these are basically going to use machine

4:14:20learning to try and optimize direct

4:14:22query semantic models and they going to

4:14:24look for the best aggregation to improve

4:14:26performance another thing we need to be

4:14:28aware of for this part of the exam is

4:14:29the large format semantic model now

4:14:32large format semantic models provide a

4:14:34highly compressed inmemory cache for

4:14:38optimized query performance enabling

4:14:40fast user interactivity so if you've got

4:14:42a semantic model model that's perhaps

4:14:44bigger than 10 20 30 GB what you can do

4:14:47is you can convert that small format

4:14:49semantic model into a large format

4:14:51semantic model it's going to apply this

4:14:53compression this in-memory caching

4:14:56that's going to improve the performance

4:14:57of those models now it's not just models

4:15:00that are over 10 GB in size where you

4:15:03might want to think about converting to

4:15:05a large format semantic model in many

4:15:07cases even below that kind of 10 GB

4:15:10threshold there's some benefits in

4:15:12converting to a large format scientic

4:15:14model firstly you're going to get the

4:15:15performance benefits anyway secondly

4:15:17it's commonly used when connecting to

4:15:19thirdparty tools via the xmla endpoint

4:15:22now another feature of large semantic

4:15:24models is ond demand loading now if you

4:15:26watched some of my previous videos you

4:15:28know that direct L connection mode also

4:15:30uses this on demand loading and what

4:15:33that means is that when a user is

4:15:35viewing a particular page in a report

4:15:37they don't need the full data set to be

4:15:39loaded into memory on demand loading has

4:15:42a look at the online paret files and

4:15:44only loads the required data that is

4:15:46needed to visualize the data that's

4:15:48being requested at that specific time

4:15:50now that can really improve performance

4:15:52again but it's a feature that's shared

4:15:53between large format semantic models and

4:15:56the direct L connection mode as well

4:15:57okay so for the next part of this video

4:16:00we're going to switch over to powerbi

4:16:02desktop and we're going to be using this

4:16:04to explain some of the key things that

4:16:07we need to know for this module in the

4:16:09exam so we're going to be looking

4:16:10specifically at variables if iterators

4:16:13table filtering window functions

4:16:15information functions calculation groups

4:16:17Dynamic strings and field parameters and

4:16:20each of these different features and Dax

4:16:22Expressions I've got a little bit of an

4:16:24example just to talk you through the

4:16:26implementation what that looks like so

4:16:28let's start off with variables now

4:16:30variables are very common in pretty much

4:16:32every programming language that exists

4:16:35and ax variables can help us avoid code

4:16:38repetition and also potentially improve

4:16:40the performance of your Dax code too so

4:16:43here what we've done is we've created

4:16:44two variables one called total revenue

4:16:46and one called total days you notice the

4:16:48syntax here is to use V to declare it as

4:16:52a variable and then it stores that

4:16:54result locally and then you can use it

4:16:57later on in your Dax expression

4:16:59typically you'll need to use a return

4:17:00statement as well when we specify these

4:17:03variables and so in this example we're

4:17:04declaring total revenue and total days

4:17:06as variables then we're using that in

4:17:09this return statement to give us the

4:17:11overall average revenue per day next

4:17:13we're going to look at iterator

4:17:15functions and iterator functions in

4:17:18powerbi basically enumerate through all

4:17:21of the rows in a table and they perform

4:17:24some calculation depending on the

4:17:26specific iterator function that you

4:17:28choose and then it's going to aggregate

4:17:30the result now examples include sum X

4:17:34count X average X most of them have this

4:17:37x afterwards and here we've got a bit of

4:17:39an example here to Showcase iterator

4:17:41functions and we using it in this

4:17:44cumulative measure so we've got this

4:17:45cumulative revenue and we're doing sum X

4:17:48okay and we're using it on this filter

4:17:50so what we're basically saying is for

4:17:52each of the rows in our date table we're

4:17:56going to go one by one and for each of

4:17:58them we're going to increase the number

4:18:00of rows that are being filtered right so

4:18:02the first row here we're basically

4:18:04comparing the the current date or the

4:18:06the date in the row that we're

4:18:07interested in with the max date which on

4:18:10the first pass when we're enumerating

4:18:12through this table there's only going to

4:18:14be one row the top row right and then

4:18:16we're calculating the sum of the revenue

4:18:18on that particular date now on the

4:18:21second row obviously this is going to

4:18:23increase the two rows so then we're

4:18:24going to do a cumulative revenue for the

4:18:27revenue on the first and the second row

4:18:29then we're going to go down to the third

4:18:30row in that table and it's going to give

4:18:32us the sum of the revenue on the first

4:18:34second and third rows so when we do this

4:18:36through the whole table the result is

4:18:38this cumulative chart of Revenue

4:18:41basically from the first dat here all

4:18:43the way through to the last date which

4:18:45is about 19 billion in revenue on the

4:18:49last date in our data set which is 31st

4:18:51of May 2020 next up we have table

4:18:54filtering now table filtering uses the

4:18:56function filter and it Bally returns a

4:18:59table that represents a subset of

4:19:01another table or expression that you're

4:19:03using so in this measure we're

4:19:05calculating total revenue by category

4:19:08and so we're using the calculate

4:19:10function and we're passing in the sum of

4:19:13the revenue and then we're filtering it

4:19:15by specific product names so what this

4:19:17is going to do is create us this chart

4:19:19right so we get Revenue figures for each

4:19:22particular category because we're doing

4:19:24this filtering of particular product

4:19:27names so like Audi tataa Hyundai these

4:19:31are all product names and our filter

4:19:33expression here is basically filtering

4:19:35out the product name where is equal to

4:19:38the selected value of product name next

4:19:40up we're going to look at window

4:19:42function

4:19:43and there's actually three functions in

4:19:45Dax that a class as window functions you

4:19:47have the window function called window

4:19:50index and offset we're going to be

4:19:51focusing on the window function in this

4:19:54example and some of the use cases where

4:19:56you might want to use a window function

4:19:57well if you've used window functions

4:19:59maybe in SQL the result is quite similar

4:20:02the way that you implement it is quite a

4:20:03bit different actually so a use case

4:20:05might be for things like rolling

4:20:07averages so if you want to create a 3mon

4:20:10moving average of revenue for example

4:20:14you might want to use window functions

4:20:15obviously there's lots of ways to do

4:20:16moving averages in powerbi in DAC window

4:20:19function is one of them or our window

4:20:21can actually be a cat variable so say

4:20:24for example the average revenue for each

4:20:25department in a company so if we take a

4:20:27look at the documentation for the window

4:20:30function at a really high level

4:20:31basically what it's going to do is

4:20:33return multiple rows which are

4:20:34positioned within a given interval for

4:20:37example we're going to give it a from

4:20:39and a to window basically and it's going

4:20:42to return

4:20:43the rows that meet that criteria right

4:20:45between the from and the two dat now

4:20:47there's a few other parameters maybe to

4:20:49be aware of here including from type and

4:20:51to type so here we can specify either

4:20:53absolute or relative values so absolute

4:20:57is just going to take the whole table

4:20:58from top to bottom and pick the absolute

4:21:01value in your from uh parameter here or

4:21:04do you want it to be relative right so

4:21:06relative so if you pick a from type of

4:21:08relative rather than looking at all the

4:21:11values in this table and picking the

4:21:14first or the second or 90th value from

4:21:16top to bottom the relative is going to

4:21:18look at the current date and then look

4:21:20at maybe minus 10 might be a relative

4:21:23from parameter now on its own the window

4:21:26function is not particularly useful

4:21:27normally how it's used if we look scroll

4:21:30down to an example here it's going to be

4:21:31used in combination with some other sort

4:21:33of measure typically as you can see here

4:21:36it's used in this iterator function so

4:21:38we're using the window function to get a

4:21:41window of data and then applying average

4:21:43X on through that window and that's to

4:21:46return the 3-day average price so that's

4:21:49a bit of an example of how you can use

4:21:50the window function in practice so

4:21:52information functions are another class

4:21:55of functions that exist within the Dax

4:21:57language and and there's a lot of

4:21:58examples of information function you've

4:22:00probably used them if you've used Dax

4:22:02before like contains or contain string

4:22:05or has one value is blank is error

4:22:07selected measure user principal name as

4:22:10well it's basically going to check a

4:22:12particular value in your table and

4:22:15return depending on the information

4:22:17function that you use normally it's a

4:22:19Boolean value but sometimes it's

4:22:20something else like user principal name

4:22:22it's not actually looking in a table in

4:22:23that point it's looking at your actual

4:22:25PBI file and looking at the logged in

4:22:27user so in this example here we've got a

4:22:29table very simple table of simple

4:22:32transactions we've got transaction IDs

4:22:33and some revenues and we're using this

4:22:36function here is blank which is an

4:22:38information function and we're basically

4:22:39using it in an if statement so if this

4:22:42is blank returns true then obviously

4:22:45we're going to use no Revenue recorded

4:22:47if it returns false so I.E there is some

4:22:50value in that column then we're just

4:22:52going to write out Revenue recorded and

4:22:54then we get this sort of table here next

4:22:55up we have calculation groups and it

4:22:57provides a simple way to reduce the

4:22:59number of measures in a model or at

4:23:02least the maintenance of those measures

4:23:05in a model so say you want to create

4:23:07daily average monthly average and then

4:23:11yearly average measures and you might

4:23:13want to do this for revenue and you

4:23:15might want to do this for cost and you

4:23:17might want to do this for salaries

4:23:20there's a lot of repetition that you're

4:23:21going to have to do in this Dax code so

4:23:23what we can do in calculation groups is

4:23:26basically parameterize those Dax

4:23:28measures so that the actual maintainable

4:23:30code that you're writing is a lot less

4:23:32now you can create these now in Pia

4:23:34desktop as of a few months ago and also

4:23:37in table editor as well got an example

4:23:39here where I've created this calculation

4:23:41group and it's called aages now to have

4:23:43a look at our calculation group let's

4:23:45just go over to the modeling Tab and

4:23:48then you can see here we've got

4:23:49calculation groups averages and we've

4:23:51got some different calculation items so

4:23:54obviously to create a new one you can

4:23:55create a new calculation group here now

4:23:57in this calculation group I've actually

4:23:58got three calculation items the first

4:24:01one is just the total which is just

4:24:02selected measure that's kind of like

4:24:04your your Baseline measure and then

4:24:06we're going to reuse that or we're going

4:24:07to call that within other measures so

4:24:10like the daily average is going to call

4:24:11Select measure here the monthly measure

4:24:14is also going to call selected measure

4:24:15but this time on the month next up we're

4:24:17going to look at Dynamic string

4:24:19formatting and this is a pretty cool one

4:24:21it basically allows you to apply string

4:24:23formatting on numerical measures without

4:24:26updating the underlying data type

4:24:28underneath so it can remain as a numeric

4:24:30measure in this example here we've

4:24:32created this measure called Dynamic

4:24:34format measure we're using this some

4:24:36Revenue just as an example here in our

4:24:39example what we're doing is giving a few

4:24:41different options in our switch

4:24:43statement so if it's less than a th000

4:24:45we're not going to do any formatting at

4:24:47all if it's between a th000 and a

4:24:50million we're going to add in this K so

4:24:52thousand we're going to add in an M if

4:24:54it's in the million ranges and billion

4:24:58if it's in the billion ranges basically

4:25:00now the benefit of this is that well in

4:25:02our tables and in our this is just a

4:25:05card it's going to show a much nicer

4:25:08presentation of that number 5246

4:25:12million rather than lots and lots of

4:25:14numbers and the big benefit is that the

4:25:16difference between this and the modeling

4:25:18tab format string so if we were to click

4:25:21on a specific Revenue number here for

4:25:25example so we can maintain our format as

4:25:28whole number but we're also formatting

4:25:30the presentation right and this is

4:25:32important maintaining this format as

4:25:34whole number because we might want to do

4:25:37something like visualize this data

4:25:40within a chart so the final feature we

4:25:42are going to look at today in powerbi is

4:25:45these field parameters and field

4:25:46parameters allow the report user so the

4:25:50end user here is going to be coming into

4:25:51your report to select different

4:25:54categoric variables and also measures as

4:25:56well kind of dynamically so depending on

4:25:59the way in which they want to slice the

4:26:00data you can set up field parameters to

4:26:03give them this flexibility so in this

4:26:05example here we've got a number of

4:26:06different parameters and we're looking

4:26:08at the revenue by different parameters

4:26:11So currently we're looking at by

4:26:12location but you might also want to look

4:26:14at by dealer or by country or by model

4:26:18or by product and you're giving the end

4:26:19user the flex ability here to decide now

4:26:22to set up field parameters you can go to

4:26:25the modeling tab new parameters fields

4:26:27and in that way you can set up a new

4:26:29field parameter just give it a name and

4:26:31then pass in the different parameters

4:26:34that you want you also rename them if

4:26:35you want here as well that's going to

4:26:36set up your field parameters and then in

4:26:38the actual visuals it's going to create

4:26:40this parameter maybe I could to prove

4:26:42the naming here but it's basically going

4:26:43to look like this so it's just going to

4:26:46have this object here and it's going to

4:26:47say location is name of this it's going

4:26:50to give it an index here 0 1 2 3 so that

4:26:53when you click on the filter it's going

4:26:55to update the visual and on the visual

4:26:57side again we've got on the y- axis just

4:26:59this parameter that we created and our

4:27:01measure that we want to visualize so sum

4:27:04of Revenue now in this instance we've

4:27:06got a field parameter in the y- AIS but

4:27:08we could also have a field parameter for

4:27:10the revenue as well maybe you wanted to

4:27:11do some of the Revenue average revenue

4:27:13median Revenue whatever you would want

4:27:15and you want to give the user

4:27:16flexibility to change these dynamically

4:27:19that's how you would do that there okay

4:27:21so we've been through a lot in this

4:27:22video now let's test your knowledge of

4:27:25what we've been through and some of the

4:27:27questions that you might face in this

4:27:29section of the exam question one the Dax

4:27:32expression average X is an example of a

4:27:35and information function b a calculation

4:27:38Group C table filtering d a window

4:27:41function or E an iterator function pause

4:27:44the video here have a little think and

4:27:46then I'll reveal the answer to you

4:27:47shortly so the answer here is obviously

4:27:49an iterator function now the big clue

4:27:51here is the X at the end of average X

4:27:54which generally denotes an iterator

4:27:55function of course an iterator function

4:27:57is those ones where we're going to be

4:27:59enumerating through every row in a

4:28:00particular table and Performing some

4:28:02calculation before combining the results

4:28:05in some way depending on how you set up

4:28:07your iterator function it's not going to

4:28:09be information functions it's not going

4:28:11to be a calculation group group table

4:28:13filtering here well you might actually

4:28:15do some table filtering within your

4:28:17iterator function but average X itself

4:28:19is not table filtering function

4:28:21similarly with window functioning again

4:28:23you might use an iterator function

4:28:24average X within a window function but

4:28:27average X itself is not actually a

4:28:28windowing function question two on

4:28:30demand loading I loading only the data

4:28:33is needed for a particular query is a

4:28:35feature of which two are the following a

4:28:38import mode B direct Lake mode C direct

4:28:41query mode d large format semantic

4:28:43models or E the xmla endpoint the answer

4:28:46here is B direct Lake mode and D large

4:28:49format semantic models so we mentioned

4:28:51this when we were going through the

4:28:52slides this feature on Dem M loading is

4:28:55actually shared by two of these modes

4:28:57here so direct late mode and also it's a

4:29:00feature included in large format Mantic

4:29:02models as well on demand loading doesn't

4:29:04really make sense in import mode CU in

4:29:07import mode we've got the full data set

4:29:09there for us to query anyway now you

4:29:11could argue that direct query mode what

4:29:13it's actually doing is very similar to

4:29:16On Demand loading whenever you get a

4:29:18request for a query you're actually

4:29:20going back to the data source and you're

4:29:22querying that data source directly and

4:29:24then loading in only the data that's

4:29:26needed whatever comes back from that

4:29:28database query but I do think there is a

4:29:29distinction between what direct query

4:29:31mode is doing and specific feature

4:29:34called On Demand loading and I do think

4:29:36these two are slightly separate so

4:29:37although direct query does a similar job

4:29:39it's not actually leveraging on demand

4:29:41loading and the xmla end point e is just

4:29:44not the correct answer question three

4:29:46Dynamic format strings overcome which

4:29:49significant limitation that comes from

4:29:51using the Dax format function is it a

4:29:55the format function is slow on large

4:29:57data sets B the format function returns

4:30:00a string value so the values can't be

4:30:02used in chart visuals is it C the format

4:30:05function can't handle date local

4:30:08conversion whilst formatting or is it D

4:30:10the format function can't be used with

4:30:13field parameters so the answer here is B

4:30:17the format function returns a string

4:30:20value so the values can't be used in

4:30:22chart visuals one of the major benefits

4:30:25of dynamic format strings is that you

4:30:27actually retain the original data type

4:30:29for that particular column so if you've

4:30:31got a numeric data type maybe you want

4:30:33to format some millions or billions in

4:30:36that numeric data type you can create a

4:30:38dynamic format string that's going to

4:30:40visually format the string but the

4:30:42underlying data type Still Remains

4:30:43numeric so you can use that field within

4:30:46charts right you maybe you want a Time

4:30:48series chart that wouldn't be possible

4:30:50if you used the format string because

4:30:52that's going to return a string value

4:30:53and you can't visualize a string value

4:30:55in a chart like that so the answer here

4:30:57is B question four the Dax expression

4:31:01selected measure is most likely found in

4:31:03the construction of which of the

4:31:04following is it a a calculation item in

4:31:07a calculation Group B field parameters C

4:31:11an iterator function D large format

4:31:13semantic models or e a window function

4:31:16now the answer here is a a calculation

4:31:19item as part of calculation groups so as

4:31:21you remember you need to have that

4:31:23selected measure that's what makes it

4:31:24Dynamic and parameterizable let's say is

4:31:27that selected measure expression it's

4:31:29not part of field parameters or iterator

4:31:31functions or window functions large

4:31:34format sematic models but of course you

4:31:36could have a calculation group within

4:31:38your large format model but the most

4:31:40likely place that you're going to find

4:31:41this because it's pretty much necessary

4:31:43is in that calculation item when you're

4:31:45creating calculation groups question

4:31:47five which of the following is an

4:31:50irreversible operation which means it

4:31:52can't be changed afterwards where you

4:31:53can't go backwards is it a changing the

4:31:56cross filtering of a relationship to

4:31:57bidirectional B changing the storage

4:32:00mode of a table to import C naming a

4:32:04calculation Group D converting a

4:32:06semantic model into a large format

4:32:08semantic model C creating a window

4:32:11function so the here is B changing the

4:32:13storage mode of a table to import mode

4:32:16now this assumes that the original

4:32:17storage mode of that table was direct

4:32:19query and if we're moving it back to

4:32:21import mode that is an irreversal

4:32:23operation I'll leave link to the

4:32:24documentation there that kind of

4:32:26specifies where that's the case

4:32:28obviously with a the cross filtering of

4:32:30a relationship to bidirectional we can

4:32:32change the cross filtering that's not a

4:32:33problem of a particular relationship we

4:32:35can also rename calculation groups

4:32:37creating a window function doesn't

4:32:38particularly make sense because of

4:32:39course you can just delete the window

4:32:41function or delete the the measure now

4:32:43converting a semantic model into a large

4:32:44format semantic model is an interesting

4:32:46one now I was under the impression that

4:32:48this also was an irreversible operation

4:32:51but then when I was actually going to

4:32:52research it and test it out I could

4:32:54actually convert a large format semantic

4:32:56model back into a small sematic model so

4:32:58for that reason I've included it as

4:33:00false but let me know if you think that

4:33:02D is also a correct answer for this

4:33:04question congratulations that's the

4:33:05first part of the semantic modeling

4:33:08section of this exam complete in the

4:33:11next lesson we're going to be looking at

4:33:13model optimization and security so make

4:33:16sure you click here to join us for the

4:33:18next lesson I'll see you there hi

Secure and optimize semantic models

4:33:20welcome to video 10 out of 12 in this

4:33:23dp600 exam preparation course today

4:33:27we're going to be looking at securing

4:33:29and optimizing semantic models so this

4:33:31is the second part in our semantic

4:33:33modeling section of the study guide and

4:33:36as you can see here we're very close to

4:33:37the end of the course so we got two more

4:33:39modules after this one but today let's

4:33:41focus on semantic modeling again

4:33:43specifically we're going to be looking

4:33:45at implementing Dynamic row level

4:33:47security and object level security

4:33:50implementing incremental refresh

4:33:52implementing performance improvements in

4:33:54queries and Report visuals and to do

4:33:57that we're going to dive into the use

4:33:59cases for external tools like Dax Studio

4:34:01tabular editor 2 and then we're going to

4:34:03go into a bit more detail about okay

4:34:05what can we actually do in terms of

4:34:07improving Dax performance in Dax studio

4:34:10and also optimizing semantic model

4:34:11models using tablet editor as ever at

4:34:14the end of this video I'll be asking you

4:34:16five sample questions to test your

4:34:18knowledge of the things that we go over

4:34:20in this part of the study guide now as

4:34:22with the last video I've left links to

4:34:24really good further learning resources

4:34:26from people that are much more

4:34:27experienced in power be development and

4:34:30semantic modeling so I will caveat this

4:34:32lesson by saying you know definitely go

4:34:33and check out the further Learning

4:34:35Resources by MVPs Microsoft mvvs for the

4:34:38powerbi side a lot of this content is

4:34:40powerbi and also third party tools that

4:34:42connect to powerbi as well so let's

4:34:45start by looking at Dynamic roow LEL

4:34:47security so in general when we talk

4:34:50about row level security we're talking

4:34:52about restricting who can see what data

4:34:55at the row level in specific tables in a

4:34:58powerbi report now Dynamic Road level

4:35:01security is kind of an extension to Road

4:35:03level security by applying Road level

4:35:06security using the usable principle name

4:35:08now there are other information

4:35:10functions that we can use but user

4:35:12principal name basically gives us the

4:35:14email address of the logged in user so

4:35:17when the user logs into the powerbi

4:35:19service then behind the scenes we're

4:35:21going to get access to their email

4:35:23address and we can use that to apply

4:35:25filters to the data in our power report

4:35:28so that that user only sees the

4:35:30information that you have configured

4:35:32that they should be seeing now you can

4:35:34configure Road level security in the

4:35:36semantic model the data warehouse and

4:35:38the tsql endpoint of The Lakehouse in

4:35:41fabric but if you're using direct Lake

4:35:43mode then you want to be configuring

4:35:45Road level security in the semantic

4:35:46model otherwise you're going to be

4:35:47falling back to direct query mode and in

4:35:50general the whole One Security model in

4:35:53fabric is still kind of under

4:35:54construction so I definitely recommend

4:35:56where we are currently I would recommend

4:35:58just using the conventional semantic

4:36:00model Road level security Now one thing

4:36:01to bear in mind is that road Lev

4:36:03security only really works when the

4:36:05users that you're trying to give Road

4:36:07level security to have viewer

4:36:09permissions in a particular workspace so

4:36:11if they have more than viewer so admin

4:36:12or member or contributor technically

4:36:15this Ro level security is not going to

4:36:17be enforced because they'll have lots of

4:36:19other ways to access that data so that's

4:36:20one thing to bear in mind there need to

4:36:22be viewers in the workspace your users

4:36:24or your viewers of reports so at a high

4:36:26level these are the steps that are

4:36:28required to implement Road level

4:36:29security so the steps for implementing

4:36:32Road level security in P desktop first

4:36:34we're going to have to create a role now

4:36:37after you've created a role you want to

4:36:38select the table that you want to apply

4:36:40Road level security two you're going to

4:36:42enter a table filter Dax expression to

4:36:45configure when and who the road level

4:36:47security is applied to and we're going

4:36:49to validate that roow level security has

4:36:52been applied correctly now as well as

4:36:54roow level security we can also apply

4:36:57object level security to our semantic

4:37:00models but the way that we're going to

4:37:01be doing that is different because

4:37:04object level security can only be

4:37:05configured via third party tools such as

4:37:08tabular editor and when we talk about

4:37:10object level secur security what we're

4:37:12talking about is restricting access to a

4:37:15particular table in a semantic model or

4:37:18a specific column so you might have a

4:37:21sensitive column within your powerbi

4:37:24report and you want to restrict who can

4:37:26see that particular column of data in

4:37:28your spany model as I mentioned to

4:37:30configure object level security you need

4:37:33to use an external tool such as tabulate

4:37:35editor and similarly to row level

4:37:37security object level security only

4:37:39restricts data access for users with

4:37:42viewer permissions so we can't give them

4:37:44admin member or contributor roles in a

4:37:47workspace now the high level steps to

4:37:49implement object level security in tabul

4:37:51editor well again we're going to need to

4:37:53start by creating a role in power your

4:37:56desktop or you can also create it in tab

4:37:58editor as well then within Tabet editor

4:38:00you're going to want to find the role

4:38:01that you've created click on the table

4:38:04properties for that role set the

4:38:06permissions for the particular table

4:38:08that you want to apply object level

4:38:10security on set that permission to

4:38:12either none or read okay so obviously

4:38:15none will be if you don't want to give

4:38:16them access to that table and read will

4:38:19be if you do want to give them access to

4:38:20that table similarly we can do the same

4:38:22for a particular column as well then

4:38:24we're going to be wanting to publish the

4:38:26report to the service and add the people

4:38:28uh and groups to the particular role in

4:38:31the service next we're going to talk

4:38:32about incremental refresh in pobi now

4:38:36one thing to bear in mind here is that

4:38:37incremental refresh is a developing

4:38:39field in the world of fabric but but for

4:38:41the purposes of this exam I think what

4:38:44they are looking for is incremental

4:38:46refresh which is a feature in powerbi

4:38:49not talking about incremental refresh in

4:38:51data flow gen 2s or data pipelines or

4:38:54anything like that so I think we're

4:38:55focused here on the specific feature

4:38:57within powerbi called incremental

4:38:59refresh and typically this is used on

4:39:02large fact tables because incremental

4:39:04refresh allows you to pull in only the

4:39:06data that is changed within a given

4:39:08range right so perhaps in the last 24

4:39:12hours or the last hour you might only

4:39:15want to bring in the new data that's

4:39:17changed within that period rather than

4:39:19you know loading in all the data that's

4:39:22in that Source database and obviously

4:39:24that's going to have many different

4:39:25benefits if we can do this incrementally

4:39:27rather than allinone for starters we're

4:39:29going to need fewer refreshes the

4:39:32refreshes are going to be a lot quicker

4:39:33because you're only pulling in what's

4:39:34new your resource consumption could be

4:39:37lower as well because again you're only

4:39:39bringing in what's changed you're not

4:39:41bringing in everything every time now

4:39:42the refreshes can be more reliable

4:39:45because rather than pulling in hundreds

4:39:47of thousands or millions of rows every

4:39:49time you refresh the data set this could

4:39:51create open connections that are going

4:39:53to run really long on your database and

4:39:56have the potential to be timed out as

4:39:57well obviously when you move to an

4:39:59incremental refresh that problem is

4:40:01likely to go away because your refresh

4:40:03is going to be a lot quicker as I

4:40:04mentioned currently incremental refresh

4:40:06is only possible within the powerbi side

4:40:08so they are working on incremental

4:40:11refresh features for ETL items like data

4:40:15Factory items like the data flow Gen 2

4:40:17and the data pipeline but for the

4:40:18purposes of this exam currently when we

4:40:21talk about incremental refresh we're

4:40:22talking about powerbi and incremental

4:40:24refresh in powerbi is available for

4:40:27powerbi premium licenses only so PPU or

4:40:30premium capacity subscriptions and the

4:40:32incremental refresh policies are defined

4:40:34within power desktop so just have a look

4:40:37at how we can Implement incremental

4:40:39refresh so it starts by creating two

4:40:42parameters called range start and range

4:40:45end they must be called range start and

4:40:46range end these are reserved keywords

4:40:48for these parameters and then in your

4:40:50power query you're going to be wanting

4:40:52to apply some custom date filters to

4:40:55filter the data based on that tables

4:40:57date column to keep only the range

4:41:00between the range start and the range

4:41:01end then you're going to want to find

4:41:03your incremental refresh policy and this

4:41:06is what this looks like you're going to

4:41:07select the table you're going to set

4:41:09your import and refresh ranges and

4:41:12there's a few optional settings there

4:41:14like only refreshing on complete days

4:41:16and detecting data changes as well and

4:41:18then you're going to review and apply

4:41:19your policy and then when you publish

4:41:22your report into the service that's when

4:41:24your incremental refresh is going to

4:41:26kick in based on the policy that you've

4:41:28set and the range that you've set so

4:41:29next we're going to switch our attention

4:41:32and we're going to talk about semantic

4:41:33model performance and whenever we talk

4:41:35about performance of anything really

4:41:37it's useful to think of it in two steps

4:41:40number one is around monitoring and

4:41:42observation and Gathering data about

4:41:45what's happening in our semantic model

4:41:47in this case and then we want to talk

4:41:49about optimization so once we understand

4:41:51what's going wrong how can we optimize

4:41:53it to improve performance ultimately So

4:41:55within powerbi and the external tools

4:41:58that you can connect to powerbi there's

4:42:00quite a few different ways that you can

4:42:01monitor semantic model performance on

4:42:04this slide we'll just do a high level

4:42:05summary of all of them before digging

4:42:07into each of them in a bit more detail

4:42:09so when we're talking about power query

4:42:10perform performance there's a tool

4:42:12called the query analyzer tool when

4:42:14we're talking about analyzing the visual

4:42:16and query performance well we can use

4:42:18the performance analyzer in powerbi

4:42:20desktop we can also export that data

4:42:23into other tools like Dax Studio as well

4:42:25for a bit more fine grained analysis

4:42:27we'll take a look at that shortly so

4:42:29when we're talking about Dax performance

4:42:31we can again use Dax Studio we can bring

4:42:34in data from performance analyzer we can

4:42:36also use things like traces to monitor

4:42:39all of the different events both on the

4:42:40client side on the server side as well

4:42:42and for semantic model performance we

4:42:44can use things like the best practice

4:42:46analyzer which is a tool within tabular

4:42:48editor so for the purposes of the exam

4:42:51you're going to need a bit of a

4:42:52knowledge about what you can do in

4:42:53powerbi desktop the limits of

4:42:56performance analysis in P desktop and

4:42:58then knowing when to switch to tools

4:43:00like DAC studio and tabular editor as

4:43:03well when you want a bit more advanced

4:43:05analysis or deeper diving into the

4:43:07things that are going wrong in your Dax

4:43:09and your semantic models let's take a

4:43:10look at these three in a bit more detail

4:43:12so first let's look at Dax studio so one

4:43:15of the core use cases that we can use

4:43:17Dax studio for is loading the powerbi

4:43:20performance analyzer data in Dax studio

4:43:22for further analysis So within powerbi

4:43:25we can record different actions when

4:43:28we're using the report like clicking on

4:43:30different filters refreshing different

4:43:32pages and it's going to log the query

4:43:35time and the visual load time for every

4:43:37visual on a specific page but the UI in

4:43:41pobi for analyzing this performance

4:43:42analyzer data it's a little bit limited

4:43:44so what we can actually do is export

4:43:47that data and we can import it into Dax

4:43:49studio for a bit of a deeper dive

4:43:51analysis on that we get better filtering

4:43:53and sorting than in powerbi desktop and

4:43:55you can also view the queries behind

4:43:57each visual load now since powerbi has

4:44:00introduced the Dax query view this has

4:44:03become a bit of a less of an advantage

4:44:05for Dax Studio because now you can also

4:44:07run the underlying Dax query for a

4:44:10specific visual within the Dax query

4:44:11view in power desktop but it's something

4:44:13to bear in mind another really core use

4:44:15case of Dax studio is using the view

4:44:18metrics to look at the vertac analyzer

4:44:21so the verti PAC is the engine the

4:44:23analysis Services engine that basically

4:44:25power runs on and so we can analyze

4:44:28what's going on in there using this view

4:44:30metric we can take a look at table and

4:44:34column sizes and obviously the size of

4:44:36your table and your column has a really

4:44:38big impact on performance right and then

4:44:40we start to look a bit deeper into

4:44:43reasons why things might be slow like

4:44:46the cardinality of a column the

4:44:48different data types that you're using

4:44:49for a specific column and much much more

4:44:52we can also use the verti PAC analyzer

4:44:54to look for referential integrity

4:44:57violations and what we mean by that is a

4:45:00mismatch in the unique keys on two sides

4:45:03of a relationship you know there's lots

4:45:04more information in the verti pack

4:45:07analyzer for the purposes of the exam I

4:45:09think it's good just just to have a look

4:45:11at the verti pack analyzer understand

4:45:13what you can do there and the insights

4:45:15that you can gather from that another

4:45:16really useful feature in Dax studio is

4:45:19the trace analysis so for Trace analysis

4:45:21there's three traces really we need to

4:45:23be aware of the all queries Trace is

4:45:27going to capture different query events

4:45:29from client side tools such as power

4:45:30desktop so this is the only one of these

4:45:32traces that's going to able to capture

4:45:34both on the client side and on the

4:45:36server side your queries as well the

4:45:38query plan Trace there is only really

4:45:40going to capture

4:45:41the query plan Trace events from the

4:45:43analysis Services tabular server right

4:45:45so if you're using the Dax editor within

4:45:49Dax studio and you're using that to run

4:45:51queries and kind of analyze their

4:45:53performance that's where you're going to

4:45:54be using the query plan because it's

4:45:56going to give you the plan of that

4:45:57specific query on the analysis Services

4:45:59engine and the server timing Trace is

4:46:01going to give us the query timing from a

4:46:03server perspective so next let's switch

4:46:05over to tabul editor and look at some of

4:46:08the use cases that we might want to be

4:46:09using tabul editor for and the main one

4:46:11when it comes to optimizing semantic

4:46:14model performance is the best practice

4:46:16analyzer and this is a tool that

4:46:18basically performs a scan of your

4:46:20semantic model and it checks it for

4:46:22common issues now there's a list of

4:46:25rules that can be downloaded from GitHub

4:46:27and you can also create your own custom

4:46:29rules as well but there's a predefined

4:46:31list of rules that you can download from

4:46:33GitHub and they're organized into these

4:46:35categories we have ones that talk about

4:46:37performance ones that talk about your

4:46:38actual Dax Expressions that you're using

4:46:40error prevention formatting and

4:46:43maintenance so these are different

4:46:44categories of those rules in the best

4:46:47practice analyzer rule set now these

4:46:49checks can also be run from the tabulate

4:46:52editor CLI as part of a cicd process so

4:46:56when you deploy a semantic model from a

4:46:59development environment to a testing

4:47:02environment you might want to run the

4:47:04best practice analyzer rule sets against

4:47:06your semantic model to give you an

4:47:08automated way of analyzing ing the

4:47:11quality of that model to surface any

4:47:13particular errors that you might get

4:47:15with that model and things you need to

4:47:16be aware of from a quality and

4:47:18maintenance perspective in that model so

4:47:21I've left this use cases for Dax studio

4:47:23and Tabet editor at the end here just to

4:47:25kind of review everything we've talked

4:47:26about when it comes to Dax studio and

4:47:29tablet editor so on the Dax Studio side

4:47:31we're going to be wanting to use Dax

4:47:32Studio when we want to write and execute

4:47:35and debug Dax queries but they do

4:47:38actually need to be manually copied over

4:47:40to power desktop as Dax studo is readon

4:47:44now since as I mentioned previously

4:47:46powerbi desktop now has the Dax query

4:47:49view you can actually write and execute

4:47:51Dax queries and view the results

4:47:54similarly to how you can do in Dax

4:47:55Studio but that's a relatively new

4:47:57feature that didn't used to exist in

4:47:58powerbi we can use the verti PAC

4:48:01analyzer to understand the size of your

4:48:03semantic model as well as individual

4:48:05tables and columns within the model we

4:48:07can bring powerbi desktop performance

4:48:09analyzer data into act Studio to analyze

4:48:12it further and we can use those Trace

4:48:14analysis functionalities to analyze

4:48:17query events both on the client side and

4:48:20the server side depending on the trace

4:48:21that you select and when we're talking

4:48:22about the use cases for tabular editor

4:48:24we're going to be able to quickly edit

4:48:26data models so we can create measures

4:48:28perspectives calculation groups from the

4:48:31Dax editor within tabular editor and we

4:48:33can publish them directly into the

4:48:35semantic model so that's a bit of a

4:48:36distinction between the functionalities

4:48:38of tabular editor and D Studio in tab

4:48:41editor we can actually update our

4:48:42semantic models in the tool itself

4:48:45there's also functionality for

4:48:46automating repetitive tasks using

4:48:48scripting and as we've mentioned we can

4:48:50incorporate devops into the kind of

4:48:52tabular mod model life cycle using that

4:48:55cicd functionality in the tabular editor

4:48:58CLI and we can use the best practice

4:49:00analyzer to identify common issues in

4:49:03your powerbi static model finally a good

4:49:05use case for tabular editor is if you

4:49:07want to implement object level security

4:49:09as we mentioned before that's not

4:49:11possible currently within power desktop

4:49:13so if you want to be defining object

4:49:15level security that's going to be done

4:49:17in tabular editor okay we've covered a

4:49:18lot of ground again there so let's just

4:49:20wrap up this video with some practice

4:49:23questions to test your knowledge of this

4:49:25section of the exam question one your

4:49:27goal is to analyze performance analyzer

4:49:30data in Dax Studio to find the visual in

4:49:33your report with the longest total load

4:49:35time put the following steps in the

4:49:37correct order to achieve this so what

4:49:39you've got here is five steps in a

4:49:42process your goal is to reorder this

4:49:44list so that it makes sense for this

4:49:46particular goal that we're trying to do

4:49:48analyzing performance analyzer data in

4:49:50Dax studio so take a moment here get a

4:49:52bit of paper maybe write these down in

4:49:55the correct order and I'll show you the

4:49:56answer shortly so the answer here is

4:49:59like this so first we're going to be

4:50:01starting a recording in powerbi

4:50:04performance analizer we need to actually

4:50:05record it in powerbi to begin with then

4:50:08you're going to be wanting to click

4:50:09refresh visual just to update the

4:50:11visuals on the page or you can interact

4:50:14with the report as well if that's what

4:50:15you want to be analyzing then clicking

4:50:17stop recording then we can export that

4:50:19performance data Json and then import it

4:50:22into Dax Studio then we can go over to

4:50:24the powerbi performance Tab and sort by

4:50:27the total milliseconds descending to

4:50:29find the longest refresh time for a

4:50:31particular visual on the page question

4:50:33two you want to use the best practice

4:50:35analyzer all within tabul Editor to

4:50:38assess your Dax performance which of the

4:50:40following severity codes for best

4:50:43practice analyzer rule violations

4:50:45indicates an error is it a level zero B

4:50:49level one C Level Two and above D level

4:50:52two e level three and above so the

4:50:54answer here is level three and above so

4:50:58in the best practice analyzer within

4:51:00tabular editor we have many different

4:51:02levels of severity of the different rule

4:51:05violations level one is just for

4:51:07information only a level two is a

4:51:10warning and level three and above is an

4:51:12error so level three and above e is the

4:51:15correct answer to this question question

4:51:17three when implementing Dynamic Ro level

4:51:19security which information function

4:51:21should you use to filter the data in a

4:51:24specific table based on the logged in

4:51:26user's email address is it a user B user

4:51:29object ID C user email D user principal

4:51:33name or E username so the answer here is

4:51:37D user principle name so as you recall

4:51:39when we were talking about Dynamic role

4:51:41of security and specifically in the

4:51:43question is talking about using the

4:51:45logged in users email address and the

4:51:48information function to give us that is

4:51:50the user principal name now the user

4:51:53function doesn't exist the user object

4:51:56ID and the username does exist as

4:51:58information functions in Dax but these

4:52:01are not the correct answer they won't

4:52:02give us the email address the username

4:52:04will just give you the domain and the

4:52:06the user's name rather than the actual

4:52:09email address and and user email C is

4:52:11also incorrect that one is is made up so

4:52:14that's not what the function is called

4:52:15it's user principal name d question four

4:52:18in Dax Studio which of the following

4:52:20records queries are generated by a

4:52:23client tool like powerbi desktop a query

4:52:26plan Trace B SQL profiler C all queries

4:52:30Trace D server timings trace or E the

4:52:33verti PAC analyzer so the answer here is

4:52:35see the all queries Trace so the clue in

4:52:39the question here was talking about the

4:52:41client tool and recording queries that

4:52:43generated specifically within the client

4:52:45tool itself within powerbi desktop so

4:52:47the all queries Trace is the only one

4:52:49that's going to give you that

4:52:50information the verti pack analyzer well

4:52:52that's just going to analyze things like

4:52:54referential integrity and the sizing of

4:52:57different tables and columns within your

4:52:59semantic model not necessarily the load

4:53:01time directly within the client tool

4:53:03power VI and the SQL profiler the server

4:53:06timings and the query plan these are all

4:53:08kind of serers side backend tools

4:53:11they're not going to give you

4:53:11information about queries generated in

4:53:14powerbi so the answer here is C the all

4:53:16queries Trace question five the first

4:53:18step in implementing incremental refresh

4:53:21in powerbi is to a add a range start and

4:53:25a range end column to your data set B

4:53:27add refresh start and refresh end

4:53:29parameters to your powerbi desktop

4:53:32project C add a refresh start and a

4:53:34refresh end column to your data set or

4:53:37add range start and range end parameters

4:53:40to your powerbi desktop project so the

4:53:42answer here is D adding range start and

4:53:45range end parameters to your powerbi

4:53:47desktop project now as we mentioned

4:53:48previously range start and range end

4:53:51they're reserved keywords when we're

4:53:52talking about parameters so you do need

4:53:54to specific when you're adding in these

4:53:56parameters should be range start and

4:53:57range end answers A and C talk about

4:54:00adding columns to your data set that's

4:54:02not going to do anything or at least

4:54:04anything useful when it comes to

4:54:05incremental refresh we need those as

4:54:08parameters cuz we're going to use those

4:54:09to filter date column in the table that

4:54:12you're interested in so we're going to

4:54:13use those parameters to filter that

4:54:15table and as I mentioned previously B is

4:54:17wrong because we're using refresh start

4:54:19and refresh end rather than range start

4:54:21and range end we need to be specific

4:54:23when we using inal refresh to be using

4:54:26range start range congratulations you've

4:54:28now completed the first three sections

4:54:31of the exam study guide we only have two

4:54:34quite short sections to go so make sure

4:54:36you click here to join me in the next

4:54:38video where we'll be looking at

4:54:40exploratory analytics in a bit more

4:54:42detail see you there hello and welcome

Perform exploratory analytics

4:54:44back to the channel today we're going to

4:54:46be continuing our dp600 exam preparation

4:54:50course and we've made it to video 11 of

4:54:5312 in this series and today we're going

4:54:55to be looking at performing exploratory

4:54:58data analytics now there's this video

4:55:01and then one more video to go so we're

4:55:02very close to the end so keep going

4:55:04we're almost there and within this video

4:55:06we're going to be looking at descriptive

4:55:09diagnostic predictive and prescriptive

4:55:12analytics and specifically we're looking

4:55:14at how to implement those things within

4:55:17powerbi we'll also be taking a look at

4:55:19the data profiling tool which is part of

4:55:21the power query experience so you'll see

4:55:23that within the data flow Gen 2 and also

4:55:26within the power query engine in powerbi

4:55:28as well towards the end of the lesson

4:55:30we'll be finishing with five sample

4:55:32questions just to test your knowledge of

4:55:34this section of the study guide and as

4:55:36ever I'll be leaving links to further

4:55:39learning resources if if you want to go

4:55:40into more detail about anything I

4:55:42mention in this video and I'll leave a

4:55:45link to the school community in the

4:55:46description below if you want to grab

4:55:49those learning notes okay so just to

4:55:51kick us off I think it's worthwhile just

4:55:53going through those four types of

4:55:56analytics and looking at what they mean

4:55:58in Microsoft's own words so we'll start

4:56:01by looking at descriptive analytics so

4:56:04when we talk about descriptive analytics

4:56:06we're talking about analytics that

4:56:07interpret past data and kpi to look for

4:56:11Trends and patterns so we're looking at

4:56:13what happened in the past we're

4:56:15describing what happened in the past

4:56:17when we talk about descriptive analytics

4:56:19and we'll be going into some examples of

4:56:21descriptive analytics and some visuals

4:56:23you can use to perform descriptive

4:56:25analytics but for now let's just look at

4:56:27these definitions the second one to know

4:56:29is diagnostic analytics so diagnostic

4:56:33analytics varies from descriptive

4:56:35analytics because we're not just worried

4:56:37about what happened in the past but we

4:56:40look looking at analytics that describe

4:56:42which data element will influence

4:56:45specific Trends and the possibility of

4:56:47future events so when we talk about

4:56:50diagnostic analytics we're not just

4:56:51talking about what happened but we're

4:56:53inferring why something happened so we

4:56:56don't just care about okay what happened

4:56:58in the last 12 months but we're looking

4:56:59for more detail around particularly why

4:57:02certain events or certain Trends might

4:57:04have occurred now typically this uses

4:57:06techniques like correlation analysis and

4:57:09data mining at at least in the words of

4:57:11Microsoft and there's various techniques

4:57:12that we can use to kind of Infuse our

4:57:15visual reports in powerbi to try and

4:57:17expose some of this diagnostic

4:57:19information number three is Predictive

4:57:22Analytics so with Predictive Analytics

4:57:25we're going to be using statistics or

4:57:27machine learning as well to forecast

4:57:30future outcomes with statistical models

4:57:32and machine learning techniques as I

4:57:34mentioned now these analytics provide

4:57:36context and Clarity for future decisions

4:57:40so in Predictive Analytics we're not

4:57:41just worried about what's happened in

4:57:43the past but we're using that data our

4:57:45historic data to think about what might

4:57:48happen in the future with some amount of

4:57:50certainty and finally the next level of

4:57:54analytics is prescriptive analytics so

4:57:57in prescriptive analytics we're not just

4:57:59predicting what's going to happen in the

4:58:01future but we're providing

4:58:03recommendations and we're recommending

4:58:04actions that might be the best course of

4:58:07action to either prevent a particular

4:58:10scenario or to increase the chances of a

4:58:14particular scenario particular outcome

4:58:15for your business now with all of these

4:58:17four the clue really is in the name so

4:58:20if you're struggling to remember what

4:58:22each of these types of analytics is and

4:58:24it helps just to go back to the first

4:58:26word in each of them and really

4:58:28understand what each of them means so in

4:58:30the first one we're describing so it's

4:58:32descriptive analytics we're describing

4:58:35what happened in the past with

4:58:36diagnostic analytics we're diagnosing so

4:58:39we're not just describing what happens

4:58:42but there's some sort of causality there

4:58:44we're thinking why something happened in

4:58:46the past so we're diagnosing a

4:58:48particular Trend obviously with

4:58:50Predictive Analytics we're going to be

4:58:52predicting what happens in the future

4:58:54and with prescriptive analytics

4:58:55obviously the key word there is

4:58:57prescribe we're not just predicting

4:58:58what's going to happen in the future

4:59:00we're also prescribing a suitable course

4:59:03of action that you should follow to

4:59:05optimize a particular outcome in the

4:59:07future now one thing I'll say before we

4:59:09start here is that most of the content

4:59:12for this section of the exam comes from

4:59:15or at least is inspired by the pl300

4:59:18exam for the PBI data analyst so that is

4:59:21what we'll focus on for this section of

4:59:23the study guide for me it's a bit less

4:59:25about descriptive diagnostic predictive

4:59:28and prescriptive analytics although you

4:59:30could argue that the visuals that

4:59:31they're talking about here align to one

4:59:34of these four categories really it's

4:59:36more about powerbi and the visuals in

4:59:38powerbi when you should use which Visual

4:59:41and when you should use specific powerbi

4:59:43features so that's what we're going to

4:59:44be focusing on in this lesson and when

4:59:46it comes to visuals and choosing which

4:59:49ones you should use when well visual

4:59:52selection is somewhat subjective I would

4:59:54argue for the exam recommend that don't

4:59:57try to be too clever when you're

4:59:59thinking about visual selection so they

5:00:00might ask you when should you use this

5:00:03particular visual or given this scenario

5:00:06which visual would you choose so instead

5:00:08of trying to be too clever in instead I

5:00:10would recommend thinking about what

5:00:12Microsoft see as the main use case for a

5:00:14particular visual you know get inside

5:00:16the heads of the examiner and of

5:00:17Microsoft and think about how they want

5:00:19you to use the tools keep that in mind

5:00:21as we go through the following examples

5:00:23okay so let's go through some of the

5:00:25commonly used visuals in powerbi and

5:00:28when you might consider choosing them or

5:00:29maybe not choosing them as well so let's

5:00:31start with the table and the Matrix

5:00:34visualization so these are really good

5:00:35for visualizing fine grain details in

5:00:38your data allowing your user the report

5:00:41user to explore the data themselves

5:00:44right and these are commonly used in

5:00:45drill through functionality so when you

5:00:47want to provide the user the option to

5:00:49drill through to find more detail about

5:00:51a particular metric you might allow them

5:00:53to drill through and we'll be talking

5:00:55through that functionality in more

5:00:56detail a little bit later on they can

5:00:58also be used to display aggregate

5:01:00information so for example the revenue

5:01:02broken down by month for example one of

5:01:05the drawbacks of the table and Matrix

5:01:07visualizations that obviously it's

5:01:09difficult to to spot long-term trends at

5:01:12least visually next we have the bar and

5:01:15the column chart so when you have one

5:01:17categoric variable and one numeric

5:01:20variable for example you might be

5:01:21looking at the revenue by region these

5:01:24are particularly useful they can also be

5:01:26stacked if you want to add in another

5:01:28categoric variable into your analysis so

5:01:31maybe Revenue by region but then also by

5:01:34product type as well just to give it an

5:01:36extra layer into your analysis again bar

5:01:39and column charts they can be used to

5:01:41visualize time series information but it

5:01:43does get a bit messy if you've got lots

5:01:45and lots of time periods along your

5:01:47xaxis generally it's better to present

5:01:50time series information on a line chart

5:01:53which brings us nicely into the next one

5:01:55which is the line and the area chart as

5:01:56I mentioned this are really good for

5:01:58visualizing time series information you

5:02:01can fit a lot of information into one

5:02:03chart over large time ranges we can also

5:02:06use the legend to kind of break up a

5:02:08single line into multiple lines to

5:02:10compare how that metric is changing

5:02:13within each category over time now

5:02:15another variation of this is the area

5:02:17chart on the right hand side there my

5:02:19personal opinion is I don't really like

5:02:20the area chart it's difficult to

5:02:22interpret all these different areas

5:02:25because you know the color of each area

5:02:28is kind of impacted by the color

5:02:29underneath it so it can be difficult to

5:02:31interpret in my opinion from a user

5:02:32perspective but again this is not about

5:02:34my personal opinion it's about what

5:02:36Microsoft sees as the core use cases for

5:02:39each of these charts s next we have the

5:02:41card visualization now these are

5:02:42obviously really good for kpi metrics

5:02:46and you can also include percentage

5:02:48change metrics you can add a bit more

5:02:50context about what's happened is that

5:02:52sales amount an increase or a decrease

5:02:55since the last sales amount as well now

5:02:57out of the box by default it doesn't

5:02:59really give you that longer term Trend

5:03:02but it can be coupled with things like a

5:03:04spark line so if you want to show the

5:03:06momentum of a particular metric over

5:03:08time adding that spark line can be

5:03:10really useful to give your users a bit

5:03:12of context about how that's changing

5:03:14over time and the card visual is

5:03:16something that has been developed quite

5:03:17a lot by Microsoft over the last few

5:03:19months and years and it's actually quite

5:03:21feature Rich now so you can really add a

5:03:23lot of information into these card

5:03:25visuals the pie chart donut chart and

5:03:27tree maps are a little bit controversial

5:03:29but their main goal is to show ratios

5:03:32between different categories now the

5:03:35reason why pie charts and donut charts

5:03:37these kind of ratio charts can be

5:03:40controversial is that they're more

5:03:41difficult for the human brain to

5:03:43interpret the difference between an area

5:03:46which is what we're showing in a pie

5:03:48chart or a tree map for example and a

5:03:50more linear comparison that you get with

5:03:53like a bar chart another potential

5:03:54downside of this type of visual is that

5:03:57by presenting ratios it does somewhat

5:04:00obscure the overall numbers so you might

5:04:02be comparing the sales in USA to the

5:04:05sales in the UK as a ratio we don't

5:04:08really by default while get a view of

5:04:10the overall sales in each category now

5:04:13you can add that to the labeling and

5:04:16also potentially a tool tip to add that

5:04:17information in but I would argue if

5:04:18you're going to do that there's better

5:04:20visualizations to choose another

5:04:21downside is that we can't really see the

5:04:23trends over time so simil with the card

5:04:26visual we're only getting the point in

5:04:28time metric for that particular metric

5:04:31right we're not seeing how that metric

5:04:34has changed over time which is normally

5:04:36more useful for the user next we come to

5:04:38combo charts now these are typically

5:04:41used to visualize more than one metric

5:04:44on the Y AIS so you can see here we've

5:04:46got the bar chart showing one particular

5:04:49metric and then we've got a line chart

5:04:51showing another metric on top of the

5:04:53same visual right now this can be useful

5:04:56in some scenarios but you do have to be

5:04:58careful because sometimes this allows

5:05:00the user to come to a certain conclusion

5:05:02about correlation between these two

5:05:04metrics which you might not actually be

5:05:06correct that correlation and it can also

5:05:08be difficult for a user to interpret

5:05:12which axis is showing which metric you

5:05:14know often you need a really good Legend

5:05:16to explain the differences between these

5:05:18two axes these two metrics as well next

5:05:21we come to the funnel visualization now

5:05:23these are really good for showing some

5:05:25sort of movement through a linear

5:05:27process an example here would be

5:05:29tracking website conversion so at the

5:05:31top of your funnel you might have a

5:05:33website viewer visiting your website

5:05:36then the next stage in that linear

5:05:38process might be okay they're going to a

5:05:40product page then the next stage might

5:05:42be okay they've clicked on buy they've

5:05:44added something to their cart for

5:05:46example and then the next stage in the

5:05:48process might be they've actually bought

5:05:49the product and so on and so on so when

5:05:51you've got linear processes and you're

5:05:53trying to visualize some sort of

5:05:55movement through that process funnel

5:05:57visualizations are really useful now if

5:05:59I was to be picky I think you could

5:06:00argue that is difficult for some users

5:06:02to grasp the scale of the difference

5:06:04between each of the levels in this

5:06:07funnel but again that's probably not

5:06:08something for the exam that's just

5:06:09personal preference next is the gauge

5:06:12chart now the gauge chart is useful when

5:06:14we want to show progress of a particular

5:06:16metric towards a goal so say you have an

5:06:20annual revenue Target and you want to

5:06:22show halfway through the year that oh

5:06:24we're actually at 53% of our Target so

5:06:27we're on track for example next up is

5:06:29the waterfall visualization so the

5:06:32waterfall visualization is useful for

5:06:34showing a running total over either a

5:06:37time period or within specific

5:06:40categories right so the goal here is

5:06:42really we want to understand which

5:06:44periods or categories contribute most to

5:06:47that change to that overall figure right

5:06:49and again from personal experience the

5:06:51waterfall chart I think should be used

5:06:53carefully because it can be in my

5:06:55opinion somewhat difficult for users to

5:06:56interpret but that's just based on my

5:06:59experience next we have the scatter

5:07:01chart which is generally used when you

5:07:03want to visualize two numeric or

5:07:06continuous variables now again you can

5:07:07add more information to the legend of

5:07:09these types of visuals if you want to

5:07:12you know add a further layer of analysis

5:07:14so you might be visualizing someone's

5:07:16age versus their height and then your

5:07:19third variable that you want to add into

5:07:21that analysis might be the country that

5:07:24they grew up in or the country that they

5:07:26live that might be an extra variable

5:07:28that you want to add into this analysis

5:07:29and you can do that by color coding the

5:07:31dots also changing the shape of those

5:07:35dots as well now obviously you need to

5:07:36be careful when you do that adding a

5:07:37third variable because you know having a

5:07:40scatter chart with lots and lots of

5:07:41different categories of different colors

5:07:43can be quite difficult to interpret from

5:07:45a user perspective another kind of

5:07:47warning with these types of charts is it

5:07:48can sometimes lead the user to come to

5:07:51conclusions about correlation between

5:07:53these two variables and as we know

5:07:54correlation doesn't always equal

5:07:56causation so that's something to Bear In

5:07:58Mind as a report author you might want

5:08:01to compare these two variables when in

5:08:02reality they might not actually have a

5:08:04causative relationship next we've got

5:08:06custom visuals so custom visuals help

5:08:09you go beyond the out thebox visuals

5:08:12that come with powerbi and there are

5:08:14lots and lots of custom visuals that you

5:08:17can explore the app Source now these can

5:08:19be really useful if you want to use

5:08:22slightly more Niche visual types that

5:08:24maybe haven't made their way into the

5:08:27core powerbi visuals set yet and of

5:08:30course bear in mind that if you use a

5:08:31custom visual these are normally built

5:08:33by Third parties and sometimes they

5:08:36involve licensing and things like pay

5:08:38walls you might be able to use it for a

5:08:40short period and then you have to pay

5:08:42but that's something just to bear in

5:08:43mind around custom visuals now the final

5:08:44visual type we're going to look at is

5:08:46the Q&A visual in powerbi which allows

5:08:49users to ask natural language questions

5:08:52about the data in your underlying data

5:08:54models now this sounds good but in

5:08:56practice I think it's difficult to

5:08:57implement well obviously the questions

5:09:00currently need to be carefully

5:09:01articulated to kind of match your data

5:09:03model so it requires the user to

5:09:06understand the columns and the table in

5:09:09your data set to craft good questions

5:09:13that can give good results basically

5:09:15currently it's only available in English

5:09:17and Spanish as well if you enable it in

5:09:19the admin settings so that's a summary

5:09:21of some commonly used visuals now let's

5:09:25look at some powerbi features that can

5:09:28help us as report developers to go kind

5:09:30of an extra layer in our analysis to

5:09:33help with maybe some of the diagnostic

5:09:35elements or the predictive elements that

5:09:37we're looking for in this part of the

5:09:38study starting off with the drill down

5:09:41functionality so drill down basically

5:09:44allows your user to explore your data

5:09:46through layers of a hierarchy and that

5:09:50hierarchy can either be explicit so as a

5:09:53a powerbi hierarchy or it can be

5:09:55implicit so you maybe you haven't

5:09:56actually defined it as a hierarchy but

5:09:58there is some implicit hierarchical

5:10:01relationship between these variables now

5:10:04a classic example of a drill down might

5:10:06be to visualize at the top level of your

5:10:09drill down Revenue in a particular

5:10:11country and then you might want to drill

5:10:13down and look at Revenue by specific

5:10:16State and then within that state you

5:10:18might want to look at Revenue by store

5:10:20in a specific state for example now

5:10:22alongside drill down and drill up

5:10:24functionality you also have drill

5:10:25through and drill through is a powerbi

5:10:28feature that allows users to drill

5:10:30through from one report page into

5:10:33another report page and it carries

5:10:34through that information they click on

5:10:36so an example here you might want to

5:10:38drill through this SharePoint category

5:10:41and take them through a different page

5:10:44in the analysis pre-filtered based on

5:10:46whatever you drill through on obviously

5:10:48with this drill through you need to

5:10:49think carefully about the user journey

5:10:51in your report and the navigation

5:10:53Journey that they're going through you

5:10:55might want to add in things like back

5:10:56buttons to make sure that they don't get

5:10:59lost in your report now with both of

5:11:01these features again this is just from

5:11:03personal experience it requires your end

5:11:05user to have knowledge of drill down and

5:11:08drill through so I would argue this is a

5:11:10bit of a downside for these features but

5:11:12something to bear in mind when you're

5:11:13authoring these reports next up grouping

5:11:16so grouping allows report authors to

5:11:19group two or more categories within a

5:11:22visual we can create groups of months in

5:11:25this example or anything really that

5:11:27makes sense for the particular visual

5:11:29that you're creating now when your data

5:11:31is continuous so it's numeric you might

5:11:34want to use binning and binning

5:11:36basically allows you to create different

5:11:38bins for for these continuous variables

5:11:41now an example of this might be to

5:11:43create a salary bin so rather than just

5:11:46have salary as a number you might want

5:11:47to create salary ranges so from 0 to 30k

5:11:5130k to 60k for example and obviously

5:11:53that opens up different visual types

5:11:56that you might want to explore and use

5:11:58for this type of data some other

5:12:00features to be aware of are reference

5:12:02lines so this allows report authors to

5:12:06provide a static reference line across

5:12:08the ex for the Y AIS to give the user

5:12:11some more context so it might be the

5:12:14average sales amount for sales people so

5:12:17maybe you're you're visualizing how all

5:12:19the different sales people have

5:12:20performed and you can see all of the

5:12:21data it might be useful to show report

5:12:24users what the average is so they can

5:12:26make a comparison give them some context

5:12:28for that comparison another feature here

5:12:30is around visualizing errors so if your

5:12:32data set contains some sort of errors in

5:12:35them normally it's around prediction

5:12:37uncertainty or measurement uncertainty

5:12:41then we can use error bars to show that

5:12:44uncertainty to the user now these can

5:12:46either be markers or they can be lines

5:12:49or they can be shaded areas maybe if

5:12:51you're doing like a Time series forecast

5:12:53that kind of thing you can have a shaded

5:12:54area for the uncertainty of a particular

5:12:57prediction obviously this requires data

5:12:59on errors or prediction uncertainty next

5:13:02we have what if parameters so we can use

5:13:05what if parameters to perform some what

5:13:08I would say is kind of rudimentary

5:13:09scenario analysis so giving your report

5:13:13users the option to change a particular

5:13:16variable particular parameter in this

5:13:18case we've got discount percentage and

5:13:21they can see what impact that has on a

5:13:23particular metric in the chart below now

5:13:25to actually get whatif parameters to

5:13:27work you need to do quite a lot of data

5:13:30engineering and potentially prediction

5:13:32algorithms as well so it can be quite

5:13:35difficult to set up a good what if

5:13:37analysis finally we've got time series

5:13:39forecasting and this can be accessed in

5:13:42that analytics pane so the third icon

5:13:45when you're setting up a visual is the

5:13:47analytics Pane and here if you have a

5:13:49Time series data set you've got revenue

5:13:51numbers for the last 24 months for

5:13:54example powerbi gives you the

5:13:56functionality to create a Time series

5:14:01analysis obviously this is quite limited

5:14:03and I would argue you should be careful

5:14:05here you know if you want to be doing

5:14:06any sort of serious time series analysis

5:14:08I would argue that it shouldn't be done

5:14:09solely within powerbi but the

5:14:11functionality is there and you might get

5:14:13asked about it great so let's finish off

5:14:15this lesson with a look at the power

5:14:18query data profiling tool okay so let's

5:14:20just explore the data profiling tool

5:14:23which comes with the power query engine

5:14:25so we can use this in a data flow Gen 2

5:14:28or also in the power query engine in

5:14:30powerbi powerbi desktop if you're using

5:14:32that as well now what I've got here is

5:14:34just a query on this Revenue data set

5:14:37and I want to explore some some

5:14:39potential data quality issues in this

5:14:41data set now to do this we can go to the

5:14:43view Tab and have a look at the data

5:14:45view enable column profile and then we

5:14:47got a few different options here so show

5:14:49column quality details if we just do

5:14:51that one to begin with and let's just

5:14:53explore what that gives us so we can see

5:14:54we've got these three kind of categories

5:14:57so it goes through each column and it

5:14:58gives us a percentage of how many values

5:15:02in that column are valid how many are

5:15:04errors and how many are empty so it

5:15:06gives you a really quick indication of

5:15:08column quality Now by default you can

5:15:11see down here that the column profiling

5:15:13is based on the top 1,000 rows So

5:15:16currently is looking at the top 1,000

5:15:18rows and for this dealer ID column

5:15:20saying that none of them are empty

5:15:21there's no errors and 100% are valid now

5:15:24if we change this to the entire data set

5:15:27obviously going to take a bit longer to

5:15:29calculate but then it's going to look at

5:15:30your entire data set every value in this

5:15:33column and it's going to perform that

5:15:34same categorization as you can see this

5:15:36data set looks pretty good we've got

5:15:38100% valid for all columns no empty and

5:15:40no errors so that's good if we go back

5:15:42to the data view now we click on the

5:15:45column value distribution let's have a

5:15:47look at what that gives us so now we've

5:15:49got this quite high level view on the

5:15:53distribution of values within this

5:15:55column and if we make that a bit bigger

5:15:57here we can begin to see how many of

5:15:59these values are distinct now to explain

5:16:02the difference between distinct and

5:16:05unique I think it helps to look at this

5:16:07column here so we've got true and false

5:16:10values in every Row in this column now

5:16:13it's showing as a distinct count of two

5:16:16which is obviously referring to true and

5:16:18false so it's kind of like a distinct

5:16:20count of the values in that column

5:16:22unique means values for which there is

5:16:25only one row so here there's no unique

5:16:29values because there's true in more than

5:16:31one row you know these are not unique

5:16:33it's included in multiple rows and false

5:16:36there are actually multiple falses so if

5:16:38in our sample size here we only had one

5:16:42false value then it would be in fact

5:16:44unique so we'd get that one in the

5:16:46unique column but currently because we

5:16:47have more than one false value in this

5:16:50column it's not unique so zero unique

5:16:52values in this column so next up to

5:16:54explore we have the details pane so if

5:16:56we enable the details pane here and then

5:17:00we click on a specific kind of

5:17:02distribution we've got here make this a

5:17:03bit smaller so we get a bit more

5:17:06information about the distribution in a

5:17:08particular column so we can see the

5:17:10counts the error counts the null counts

5:17:12we can also see the distinct count

5:17:13unique count empty string counts and the

5:17:15minimum and the maximum values now

5:17:17obviously this is a minimum maximum of a

5:17:19string value but it would also work with

5:17:22numeric numbers as well I think this

5:17:24column only has unit sold one so the Min

5:17:27and the max is also one so it's not a

5:17:29particularly good example here and we've

5:17:30also got things like the average

5:17:32standard deviation number of odds number

5:17:34of evens so it give you a bit more

5:17:36detail about what there is in that

5:17:38column now once you've got a pretty good

5:17:40idea about the kind of distribution of a

5:17:42column you might be able to spot things

5:17:44like duplicate values in here what we

5:17:46can do is we can right click on it and

5:17:48we can remove duplicates or you might

5:17:50want to remove errors as well you can do

5:17:52that from just right clicking on this

5:17:54top section here and it gives you a

5:17:56quick way to remove or replace these

5:17:58errors or duplicates as well okay let's

5:18:01just round off this video by going

5:18:03through some practice questions to test

5:18:05the knowledge of what you've learned

5:18:07during this section of of the exam of

5:18:10the study guide question one a company

5:18:12annual report shows a net profit of $34

5:18:15million now the company has 12 business

5:18:18units each with their own net profit or

5:18:21loss amount which of the following

5:18:22visual types could best be used to

5:18:25visually show how each business unit

5:18:27contributed to the overall net profit

5:18:29metric is it a a line chart b a scatter

5:18:33chart c a matrix chart d a waterfall

5:18:37chart or e a question answer visual

5:18:39pause the video here have a think and

5:18:41I'll reveal the answer to you shortly so

5:18:43the answer here was the waterfall chart

5:18:46as you remember when we were going

5:18:47through the waterfall chart one of the

5:18:48core use cases of that waterfall chart

5:18:51is to break down one top level metric

5:18:54into its constituent parts to allow the

5:18:57user to explore how each in this case

5:19:00business unit contributes to the overall

5:19:02net profit metric now the other visuals

5:19:04that we've listed here you know you

5:19:06might be able to glean some of that

5:19:08information but as I mentioned

5:19:10previously what we're after here is what

5:19:12would be the best visual to visualize

5:19:14this information and for me that's the

5:19:16waterfall chart question two you want to

5:19:17add measurement error bars on a Time

5:19:19series line chart to show the potential

5:19:21error in each measurement where would

5:19:23you go to add this information to your

5:19:26visual is it a the analytics pane B the

5:19:29format visual pane C view Tab and show

5:19:32error bars D the build visual pain or e

5:19:35in the model view so the answer here is

5:19:37the analytics pain when you're

5:19:39configuring the settings for a

5:19:40particular visual there's obviously

5:19:42three panes or three different windows

5:19:44that we can use to declare different

5:19:46settings and the error bars

5:19:48functionality is included in the

5:19:50analytics pane it's not going to be the

5:19:52build visual pane that's where you add

5:19:53your different variables different

5:19:55columns to a particular visual format

5:19:57visual that's where you change things

5:19:59like the text and the border and all

5:20:00that kind of thing it's not going to be

5:20:01the model view CU that's where you

5:20:03define relationships and things like

5:20:04that and it's not going to be in the

5:20:06view tab show ER bars that doesn't

5:20:08actually exist that functionality I just

5:20:09made it up question three you want to

5:20:11visually compare two continuous

5:20:13variables age and height of survey

5:20:16respondents in one chart which of the

5:20:17following visual types could best

5:20:20represent this data is it a a line chart

5:20:23b a stacked bar visual c a scatter

5:20:26visual D question answer visual or e a

5:20:29matrix visual so the answer here is C

5:20:32the scatter visual obviously when you're

5:20:35visualizing two continuous variables so

5:20:38age and height then probably the best

5:20:40visual to use for that would be the

5:20:41scatter chart put your height on the y-

5:20:44Axis or your age on the AIS and then you

5:20:47can plot different dots for each of your

5:20:50survey respondents again the other

5:20:52visual types not really going to help

5:20:53with that kind of analysis You could

5:20:55argue that the The Matrix would

5:20:57potentially show you all of that but

5:20:58it's going to be a really big chart if

5:20:59you've got lots of respondents so the

5:21:01answer here is C the scatter visual

5:21:03question four by default the data

5:21:06profiling tool reviews the top n rows of

5:21:09your data set to show you potential data

5:21:11quality issues so what here is n so how

5:21:14many rows is it a 10 B 100 C 500 D 1,000

5:21:20or E 10,000 so the answer here is D

5:21:231,000 rows as you can see here from that

5:21:26visual by default the column profiling

5:21:29is based on the top 1,000 rows obviously

5:21:32we can change it to include the entire

5:21:34data set as well but by default it's

5:21:361,000 rows question five which the

5:21:38following features of the data profiling

5:21:40tool can help you identify duplicate

5:21:42values in a column that you plan to use

5:21:44as a key to join on is it a column

5:21:47quality details B column value

5:21:49distribution C column duplicate analysis

5:21:52D column key constraints or E group but

5:21:55so the answer here is B the column value

5:21:58distribution so as we mentioned one of

5:22:00the key use cases for the the value

5:22:02distribution is to look for duplicate

5:22:05values because if you've got a key that

5:22:07you plan to join on you want to be

5:22:09looking for duplicate values in that key

5:22:11and in the column with no duplicate

5:22:13values is going to look like this with a

5:22:15just a flat line so every value is going

5:22:17to be distinct here on the left hand

5:22:19side this would indicate that there are

5:22:21some duplicate values there's some

5:22:22values here that have a count of more

5:22:24than the others right so this would be a

5:22:27potentially problematic key on which to

5:22:29join on but this date ID column all of

5:22:31the values here are unique at least in

5:22:33the sample of data that you're using for

5:22:35profiling the column quality details a

5:22:38doesn't give us this information the

5:22:40column duplicate analysis C doesn't

5:22:42actually exist neither does d the column

5:22:44key constraints and E the group buy okay

5:22:47you could potentially use a group buy to

5:22:49look for duplicate values but it's not a

5:22:51feature of the data profiling tool

5:22:53itself congratulations we're nearly

5:22:55there only one section of the study

5:22:57guide to go in the next lesson we'll be

5:22:59looking at the final section of the

5:23:01study guide we'll be looking at how to

5:23:03use SQL to analyze data via the

5:23:06Lakehouse SQL endpoint in the to

5:23:08warehouse and also via the xmla endpoint

5:23:11as well so make sure you click here to

5:23:13join us in the last lesson in this

5:23:16series I'll see you there hello and

Query data using T-SQL

5:23:18welcome to video 12 in this dp600 exam

5:23:22preparation course and we've made it

5:23:24this is the final video in the study

5:23:27guide we're going to be looking at

5:23:28querying data by using tsql now you'll

5:23:32noticed that I've added an extra video

5:23:34there cuz I really wanted to give you

5:23:35one more video just to really help you

5:23:38prepare

5:23:38for the exam so we'll be going through

5:23:41how to prepare for the exam some other

5:23:43resources that I think you should take a

5:23:44look at before you take the exam and

5:23:47some advice for when you're actually

5:23:48sitting the exam as well but we'll be

5:23:51going through that in the next video for

5:23:52this video we're going to be focusing on

5:23:55these three sections of the study guide

5:23:57so we're going to be looking at querying

5:23:59the Lakehouse and the data warehouse

5:24:03using tsql and we're also going to be

5:24:05looking at the visual query editor which

5:24:07is a feature of both of those two SQL

5:24:10endpoints as well finally we'll be

5:24:11mentioning how to connect to and query

5:24:14data sets using the xmla endpoint as

5:24:17ever I've released study notes to help

5:24:20you go a little bit deeper just to make

5:24:22sure that you're covering off all the

5:24:24right points and links to other further

5:24:26resources if you want to go a bit deeper

5:24:27on whatever I've mentioned in this

5:24:30section of the study guide as however

5:24:31I've also got five sample questions to

5:24:33test your knowledge at the end of this

5:24:36video so let's just start by looking at

5:24:38the different ways that we can access

5:24:41the tcq engine within fabric has a few

5:24:44different ways to be aware of so one of

5:24:46them is the Lakehouse TC endpoint now

5:24:50one thing to bear in mind here as we've

5:24:51already mentioned quite a lot already is

5:24:53that this is read only so all you can

5:24:55really do here is Select statements ddl

5:24:58that kind of thing you can't do any sort

5:25:00of inserts updates deletes all that kind

5:25:03of stuff from the TC queno in The

5:25:06Lakehouse now the more obvious place to

5:25:07do tsql is within the fabric data

5:25:11warehouse here you're going to have the

5:25:12opportunity to write tsql both ddl DML

5:25:17inserts updates deletes select

5:25:19statements all of that stuff so the data

5:25:22warehouse is going to give you the

5:25:23ultimate flexibility really to write

5:25:27tsql scripts within fabric now on top of

5:25:30the tsql query editor you can actually

5:25:34also create tsql like queries using the

5:25:38visual query editor and this is quite

5:25:40similar to the data flow Gen 2 if you've

5:25:43ever used the visual editor in a data

5:25:46flow so we can do things like merging

5:25:48different tables and filtering tables

5:25:52and adding additional columns that kind

5:25:54of thing but we can do it through a no

5:25:56code visual interface and we'll be

5:25:59taking a look at that in a bit more

5:26:00detail shortly so the other option for

5:26:02writing SQL is via the xmla endpoint and

5:26:06as we've mentioned previously in this

5:26:07course to connect to that xmla endpoint

5:26:11we need to go into our workspace

5:26:12settings as you can see on the left here

5:26:13grab the connection string go into SS

5:26:17SMS in this example connect via the

5:26:20analysis Services server type pass in

5:26:23your xmla endpoint in there that's going

5:26:26to bring through all of our lake houses

5:26:28and our warehouses within that workspace

5:26:31and then we can write different queries

5:26:33depending on your use case from that

5:26:35xmla endpoint so now that we have a good

5:26:37understanding of where we can write tsql

5:26:40for most of this exam you need to have a

5:26:42pretty good level of tsql at least in

5:26:45appreciation for what a lot of different

5:26:47tsql functions do and be able to at

5:26:50least read tsql quite well and I don't

5:26:53think there's kind of a definitive list

5:26:55of which tsql functions you need to be

5:26:57familiar with but I definitely recommend

5:26:59being comfortable with the following so

5:27:02the difference between where and having

5:27:04group by summarizations Union and Union

5:27:07all different joins and when to use them

5:27:09Common Table Expressions things like

5:27:11lead and lag row number partitioning

5:27:14that kind of thing subqueries and cross

5:27:17Warehouse queries as well so rather than

5:27:19just describing all these things I think

5:27:21it would be better to jump into fabric

5:27:24open up a data warehouse and show you

5:27:26some of these functions in action okay

5:27:29so here we are in SQL Server management

5:27:31studio and I've connected to one of my

5:27:35gold data warehouses here called d gold

5:27:39using the SQL connection string and you

5:27:41can see that currently we don't actually

5:27:42have any data in this data warehouse the

5:27:45tables are empty so the first thing that

5:27:47we're going to do is just to bring some

5:27:49data from some other tables that we've

5:27:51got into this data warehouse so to do

5:27:53that I'm just going to be using this

5:27:55Seas so create table as select so it's

5:27:58going to enable us to create a new table

5:28:01in this data warehouse using existing

5:28:04data in another data warehouse or in

5:28:07this example it's it's actually a lake

5:28:08house so we're connecting to the SQL

5:28:10endpoint here so this is an example of

5:28:13cross database querying because what

5:28:15we're doing here is we're creating a

5:28:17table from another Lake housee Al

5:28:20together so we're getting the all of the

5:28:22data from this db. Revenue table in our

5:28:25Lakehouse bronze and we're creating a

5:28:27new table called dboa Revenue so then if

5:28:31we refresh these tables we should now

5:28:33have the first one which is dbo Factor

5:28:35Revenue this one here and we can do the

5:28:38same for dim date and dim Branch we're

5:28:41going to be using these data sets just

5:28:43to show some of the functions that you

5:28:45need to be familiar with for the exam

5:28:47we're not going to cover all of them

5:28:49because to be honest I'm not sure

5:28:50exactly the full breadth and depth of

5:28:52what is expected for the exam for tsql

5:28:55but we're going to go over some common

5:28:56types of problems that you might see in

5:28:58the exam so now if we refresh our tables

5:29:01again so now we've got our dbo fact

5:29:03Revenue our dim dates and our dim Branch

5:29:06so let's just start by visualizing our

5:29:09data here so specifically I'm going to

5:29:11be looking at this fact revenue and

5:29:13these data sets come from the same data

5:29:16set that we used actually in a different

5:29:18part of this exam preparation course is

5:29:20relating to Car Sales and car

5:29:23dealerships so you can see in our fact

5:29:26Revenue table we've got a revenue figure

5:29:28and we've got a date ID column and a

5:29:30branch ID column here and that's what

5:29:33we're going to be using for the rest of

5:29:35this analysis so what I'm going to be

5:29:37doing is Pres presenting you with a

5:29:38series of problems then we're going to

5:29:40walk through how you might solve that in

5:29:43SE we'll be starting off quite simple

5:29:44and then we'll be adding in more and

5:29:46more functionality as we go through so

5:29:48what if we wanted to calculate the top

5:29:50five branches by total revenue so to do

5:29:54that we're going to be needing to

5:29:56perform some sort of aggregate so this

5:29:57is what our fact table looks like

5:29:59currently we've got Branch ID and

5:30:01revenue and what we want to be doing is

5:30:03calculating the top five branches so

5:30:06what we could do is just select the top

5:30:09five here Branch IDs some of the revenue

5:30:11so what we're doing here is doing a

5:30:13simple group Buy on the branch ID

5:30:16because we're looking for the top five

5:30:17branches and we're going to order it by

5:30:19the sum of the revenue so this is the

5:30:22aggregate calculation that we're running

5:30:24on this aggregate here it's a sum of the

5:30:26revenue and we're ordering it by the sum

5:30:28of the revenue descending so we're

5:30:30getting the top five now let's change

5:30:33the problem a little bit and maybe we

5:30:35want to return only the top five

5:30:38branches in Spain or maybe the top three

5:30:41branches now Spain is a country name

5:30:45that comes from a different table in our

5:30:47data set so what we've added in this

5:30:49example is an inner join on dim branch

5:30:52and we're joining on the branch ID

5:30:54because we have the branch ID in both

5:30:56data sets we're using an inner join here

5:30:58because we only want to return data for

5:31:00which we have both keys again we're

5:31:02aggregating by the sum of the revenue

5:31:05this time we've actually brought through

5:31:07a few different columns from that Branch

5:31:10Dimension table so let's just run it all

5:31:12and see what we get here okay so this

5:31:15query has returned the top three and

5:31:17what we've done is we've filtered this

5:31:20results for only country names that are

5:31:22equal to Spain now we've had to add in

5:31:25some new things into this group by

5:31:27selection because we've also got country

5:31:29name and Branch name mentioned here so

5:31:32we've also brought brought through the

5:31:33branch name so that we can just you know

5:31:35we've got more than the branch ID maybe

5:31:37you don't not familiar with the branch

5:31:38IDs you want the actual Branch names in

5:31:41this example and we can see that that

5:31:43has brought through the top three

5:31:46branches that are in Spain by the total

5:31:49revenue so that has sorted that problem

5:31:52but what if we wanted to filter after

5:31:54the aggregate so for example give me all

5:31:56of the branches that have a revenue of

5:32:00greater than some amount so here we're

5:32:03not going to be doing the wear statement

5:32:05because we want to be using having so

5:32:07you want to be doing filtering after the

5:32:09aggregate or on that aggregate value

5:32:12then we're going to be wanting to use

5:32:13having and that's going to come after

5:32:15our group by statement so previously the

5:32:17wear statement here because we were kind

5:32:19of pre-filtering on this country name

5:32:21now we're looking for the results of an

5:32:24aggregation that are greater than in

5:32:26this example 5 million or 50 million so

5:32:28we've got a very similar setup we've

5:32:30removed the wear statement because now

5:32:32we're not just interested in Spain we

5:32:34want all the branches and we're going to

5:32:36use this having some of the Reven Vue

5:32:38greater than 50 million so here you can

5:32:40see it's returned only two branches

5:32:43which makes sense what if we change this

5:32:46having statement we removed one of the

5:32:48zeros yeah so here we're just looking at

5:32:495 million so when we've got 50 million

5:32:51there's only actually two branches that

5:32:53have more than 50 million Revenue over

5:32:56this time period or over the full data

5:32:58set in that fact table now in this

5:33:00example we've been using the inner join

5:33:02because that's what we wanted for this

5:33:04specific use case but for the exam

5:33:07you're going to be need to be familiar

5:33:09with left join right join inner join F

5:33:12outer join all of these different join

5:33:14types and when they are useful for

5:33:16different scenarios now we're not going

5:33:18to be going through all of these here

5:33:20because there's quite a lot to go

5:33:21through but I definitely recommend you

5:33:22know if you're not familiar with these

5:33:23things learning the differences between

5:33:25these and when you might want to use one

5:33:28or the other so another thing that you

5:33:30might need to be familiar with for the

5:33:32exam is commentable Expressions now

5:33:35these are really heavily used in the

5:33:38world of SQL when you're creating views

5:33:40or things like that so you need to be

5:33:41familiar with how you construct a Common

5:33:45Table expression what the the key words

5:33:47are what the general structure is and

5:33:49that kind of thing here in general the

5:33:51Comon table expression allows us to

5:33:53Define kind of like variables that we

5:33:56can use later on in our script so here

5:33:59we've got these three lines are the

5:34:02query we're selecting the top five

5:34:04Branch IDs and summing the revenue we're

5:34:06grouping by that branch ID so it's

5:34:08similar to what we saw before and so we

5:34:10can actually just select these three

5:34:11rows and have a look at what that

5:34:13returns us and then what we're doing is

5:34:15we're kind of saving that in this

5:34:16variable or this commentable expression

5:34:19called top five rev and the syntax here

5:34:21is always going to start with with which

5:34:24is the keyword for a commentable

5:34:25expression you're going to give it a

5:34:26name and then as Open brackets put in

5:34:30your expression within those brackets

5:34:32and then we can use this top five rev

5:34:34later on in our query now one of the

5:34:36other benefits of table Expressions is

5:34:38that we can create more than one of

5:34:40these so for any subsequent Expressions

5:34:43that we declare firstly we're going to

5:34:45need a comma there and then we can call

5:34:47a second one branches so for subsequent

5:34:50ones we don't need the with keyword we

5:34:52only need that on the first one and so

5:34:53here I've defined branches as select

5:34:56Branch ID and Branch name from dbo dim

5:35:00Branch so this is what this one looks

5:35:02like so we're just getting the branch ID

5:35:03and the branch name and we're storing it

5:35:05as this branches and then the final

5:35:08kind of section of a CTE is the select

5:35:11statement so you're always going to need

5:35:12to return something from this comment

5:35:14table expression and we can do that with

5:35:16a select statement so as we can see here

5:35:18we're selecting the branch ID the branch

5:35:22name and the total revenue and these are

5:35:24coming from top five rev so that's our

5:35:27keyword for our first expression that we

5:35:30defined up here and we're joining it

5:35:32with our branches which is our second

5:35:33one here we're giving it these aliases

5:35:35and to run this we have to select all

5:35:37all of the different sections and then

5:35:38press execute and we get this result

5:35:41here so we've got the top five branches

5:35:43by revenue and then we've joined that on

5:35:47the last section of our CTE back to the

5:35:49branch data set now obviously this could

5:35:52also be written in a different way you

5:35:53probably don't especially need a CTE to

5:35:56return that result but I just want to

5:35:58show you the structure of a commentable

5:36:00expression because you might get asked

5:36:02about that in the exam or you might get

5:36:04shown some code that is in this format

5:36:06and you need to understand and how it

5:36:08works another thing that you might need

5:36:09to be aware of for the exam are the lag

5:36:11and the lead function so this is a way

5:36:15of creating new columns that reference

5:36:19other columns but with some specified

5:36:22offset so probably worthwhile just

5:36:24taking a look at a bit of an example

5:36:25here so what I've done is I've defined

5:36:27another Common Table expression here so

5:36:30in the first part we getting the top

5:36:32five branches and then we're dividing

5:36:34this Revenue by 1 million just to give

5:36:38us you know some a bit easier values to

5:36:40comprehend and we're getting the floor

5:36:43of that division so it's going to round

5:36:44down to single digits in this case and

5:36:47we're defining that column as Revenue in

5:36:49millions just to make it easier to

5:36:52understand what's going on with this lag

5:36:54and Lead function that we're going to

5:36:55introduce shortly so this is what our

5:36:57data set looks like we've got five

5:36:59branches and we've got rev which are all

5:37:01singled digit whole numbers now so we've

5:37:03stored that as rev T and next we're

5:37:05going to introduce a lag function over

5:37:09this data set so we got select Branch ID

5:37:11and rev M which is looks like this and

5:37:14we're going to add on another column

5:37:16into this using this lag function so

5:37:19what the lag function is going to do if

5:37:21I just run this let's just start with an

5:37:23offset of one and we rerun this so the

5:37:25lag function is going to look at the

5:37:29column which you pass in which is

5:37:31revenue M and it's going to look at the

5:37:34value of the row number minus the offset

5:37:37so for row number one here it's going to

5:37:40look at row zero which doesn't exist so

5:37:43that will return null but for row two

5:37:45it's going to look for the value in row

5:37:48one of the column that we pass into the

5:37:51function now another thing that we've

5:37:52done here is we've ordered it by this

5:37:54rev M so the order here has changed it's

5:37:57now in ascending order from 1 2 2 3 6

5:38:00like so and we can change this lag

5:38:02function to anything we want so here

5:38:05we're doing it with a lag of two so the

5:38:08offset parameter here is going to be two

5:38:10so it's going to introduce another null

5:38:11because now we're offsetting by two and

5:38:13it looks like this now hand in hand with

5:38:15the lag function we also have the lead

5:38:18function which looks in the opposite

5:38:21direction really so rather than looking

5:38:23at the lag so Looking Back In Time the

5:38:27lead function is going to look forward

5:38:28in time or at least forward in your row

5:38:31index so it's actually going to start on

5:38:33this row here now let's just change this

5:38:36back to a offset of one so we have and

5:38:39let's just change this to lead so when

5:38:41we run the lead function here on the

5:38:44same column Revenue M with an offset of

5:38:47one this is going to be the result here

5:38:49so now our first value in this lead call

5:38:53column is actually going to be the value

5:38:54here similarly this value is going to be

5:38:56here this value is three six comes from

5:38:59here and now we have an old value at the

5:39:01end of our data set because we don't

5:39:02have you know a sixth Row from which to

5:39:05pull that data from now one thing to

5:39:07bear in mind is that we can't do

5:39:09something like this so we can't do a

5:39:11lead or a lag with an offset of minus

5:39:13one the offset parameter cannot be a

5:39:15negative value so bear that in mind now

5:39:16another function to be aware of is the

5:39:19row number and partitioning so row

5:39:22number as a name suggest basically runs

5:39:24through your data set and sequentially

5:39:26numbers your rows okay so let's just run

5:39:30this first section of code just to have

5:39:32a look at what this is doing here so

5:39:34here we've defined our row number and

5:39:36we've given at this name row num and

5:39:39we've also passed in a partition value

5:39:42so what it's going to do is it's going

5:39:43to group all of our dealer IDs so here

5:39:47you can see in the output here all of

5:39:49these dealer IDs so dealer ID DLR 001

5:39:53this first one all of these values are

5:39:55within that partition okay they're all

5:39:58that same dealer ID and the row number

5:40:00is going to assign a row number based on

5:40:03the order so the revenue so these are

5:40:05all ordered in Revenue sending order and

5:40:08then we're creating our row number 1 to8

5:40:11within this partition then you'll notice

5:40:13for the next dealer ID DLR 002 now we've

5:40:17got 10 values within this partition

5:40:19ordered in the same way and we've got

5:40:21row numbers defined here from 1 through

5:40:2410 so that's this section within the CTE

5:40:27what I've done is I've just selected

5:40:28star from Parts which is the name of our

5:40:31CTE and I've just done a bit of a wear

5:40:33statement just to only get two of the

5:40:36dealer IDs back so here you can see this

5:40:39example here we've got DLR 001 which is

5:40:42this one and we got DLR 001 7 which is

5:40:46this one so make sure you understand row

5:40:48numbering and partitioning and ordering

5:40:51by for the exam because this is

5:40:53something that could easily come up now

5:40:55the last section I wanted to mention or

5:40:57the last topic that I wanted to cover is

5:40:59subqueries so for all of the other

5:41:01examples within this tutorial or this

5:41:04video so far we've been using select

5:41:06star from table right so if we go back

5:41:10up here we' got select all of this stuff

5:41:13from dbo fact Revenue now a subquery

5:41:16basically doesn't have that structure of

5:41:18Select star from dbo doable instead we

5:41:22can actually open up a brackets here and

5:41:26rather than doing a whole table or

5:41:27getting the data from the whole table we

5:41:29can create a subquery that's going to

5:41:31basically prefilter or do some sort of

5:41:34tsql query return the results of that

5:41:36sub query back up to the top level here

5:41:39and we've got to give it this Alias of

5:41:42sub doesn't have to be sub but that's

5:41:43just what I've called it here and we'll

5:41:45see the results there so what this is

5:41:47doing is first going to calculate this

5:41:49so whatever is in our brackets we're

5:41:51going to get the revenue figures just

5:41:53for these two dealer IDs and then we're

5:41:55going to get select star from that

5:41:57result so again this is just a bit of a

5:41:59toy example just to show you subqueries

5:42:01in a bit of action at least the

5:42:02structure of a subquery just so that if

5:42:04you come across it in the exam you know

5:42:06what that is okay so here we are in a

5:42:09data warehouse and it's the same data

5:42:11warehouse that we were using for the

5:42:13first part of this video the DW gold so

5:42:15we got our fact Revenue table our dim

5:42:18date and dim branch in here now

5:42:20specifically I want to show you the

5:42:22visual query editor so we can access it

5:42:24by clicking on a new visual query and

5:42:26immediately we're going to get this

5:42:28introduction here to build a visual

5:42:30query need to drag some tables onto our

5:42:33canvas here so let's just drag on the

5:42:35fact revenue and we'll see what we get

5:42:37so the visual query engine basically

5:42:39tries to make it as simple as possible

5:42:41to transform your data that is in your

5:42:45data warehouse or also in the Lakehouse

5:42:47tsql endpoint as well so we've got our

5:42:50fact Revenue table so what can we

5:42:52actually do with it here well you can

5:42:54see along the top here we've got some

5:42:55kind of quick options for choosing

5:42:58specific columns removing columns that

5:43:00kind of thing filtering so removing

5:43:03certain rows sorting rows transforming

5:43:06we can do group by

5:43:07we can also do merging and appending

5:43:09different data sets together so if we

5:43:11want to drag another one of these onto

5:43:13the canvas as well then we can kind of

5:43:15combine those using either merge or

5:43:19append but we're not going to do that in

5:43:21this example now another thing that you

5:43:22can do within the visual query engine is

5:43:24to click on this plus button here and we

5:43:27get access to a few more commands so

5:43:29we've got all of the ones that are

5:43:31available in that top menu plus we've

5:43:33also got some transformation so we can

5:43:36do some text Transformations on a

5:43:38specific column we can do length finding

5:43:40the First characters adding columns

5:43:43using for example a conditional column

5:43:45or column from examples so if you're

5:43:47used to using the data flow power query

5:43:50engine you might be familiar with

5:43:52creating a new column from examples and

5:43:54we can also do some adding column from

5:43:56text using these methods here so say for

5:43:58example maybe I want to do a group Buy

5:44:01on this fact Revenue table and I want to

5:44:05get the median value just going to

5:44:08change this for the median and I want it

5:44:10on the revenue column and I want to

5:44:12group by our branches so we also have

5:44:15this fuzzy grouping as well so if this

5:44:18column is not particularly good quality

5:44:20you might want to add in some fuzzy

5:44:22grouping which is basically going to

5:44:24look at likeness or similarity between

5:44:27different values in that column and if

5:44:29it's above a certain threshhold so

5:44:31they're very similar so maybe there's

5:44:33just one character that's different it's

5:44:35going to add it into the same group so

5:44:37that's what fuzzy grouping would be we

5:44:38don't want to enable that there cuz

5:44:40we've already managed our data quality

5:44:42so if we do an okay here you can see

5:44:43it's add in this step here so now along

5:44:47the bottom you can see the result has

5:44:49actually updated so we've got this group

5:44:51by the branch ID and we've got our new

5:44:53column which is the median value within

5:44:55that aggregate basically so say I wanted

5:44:58to I just remove this one because we

5:44:59don't don't want to be doing anything

5:45:00with our dim date we got our fat Revenue

5:45:02now and what we can do is either save as

5:45:06table so this is going to save the

5:45:09results as a new table or we can save it

5:45:12as a view so you can see that this is

5:45:15actually gr out here cuz it's you can't

5:45:17save this as a view because the query

5:45:19fact revenue is not supported as a

5:45:21warehouse view since it cannot be fully

5:45:23translated to SQL and the reason for

5:45:25that is because we've used this median

5:45:28and median is not actually a function in

5:45:30tsql so what we can instead do is rather

5:45:32than grouping by and aggregating on that

5:45:35median value we change this to sum I

5:45:37expect that yeah so now that error

5:45:39message or that warning message has now

5:45:41gone away we might want to also change

5:45:44this to some of the revenue and then we

5:45:46can save it as a SQL view so maybe we

5:45:50want to do sum of the revenue

5:45:52aggregation give our viewer name it also

5:45:54gives you the SQL statement for that

5:45:56view which we can copy to the clipboard

5:45:58if we want and we just save it as a view

5:46:00which you can query from a powerbi

5:46:03semantic model obviously acknowledging

5:46:05the fact that with a view it's going to

5:46:07fall back to direct query mode if you're

5:46:09using direct Lake mode to access this

5:46:12data so that was just a tour of the

5:46:15visual query engine in the data

5:46:17warehouse think for the exam just

5:46:19understand what it's capable of have a

5:46:21look at the different functionality here

5:46:23understand what you can do here because

5:46:25you might get a question or two about

5:46:27the visual query engine okay so let's

5:46:29just round up this video by looking at

5:46:31some practice questions that you could

5:46:33expect for this section of the exam

5:46:35question one the SQL scripts creates the

5:46:38results shown below what is function in

5:46:42this example so we're doing select

5:46:44Branch ID column one and then your

5:46:46function which you have to work out what

5:46:48that function is from sales and

5:46:50returning that table below so take a

5:46:52look at the answers on the right hand

5:46:55side have a little think about what that

5:46:57function would be to produce the results

5:47:00in the table at the bottom of the page

5:47:02pause the video here have a think and

5:47:04then I'll reveal the answer to you

5:47:05shortly so the answer here is a we want

5:47:08to be using a lag function because we

5:47:12can see that what we're trying to do is

5:47:14make a transformation from column 1 to

5:47:17column 2 and we can see the pattern

5:47:19there that on Row three in column 2

5:47:22there's obviously an offset going on

5:47:24there and we can see the offset is two

5:47:27because Row three in column two maps to

5:47:31row one in column one and we know there

5:47:32going to be a lag function because

5:47:35column 3 is actually two rows behind

5:47:38what's on column one and you can see

5:47:40that replicates below with rows four and

5:47:43five as well and another clue here is

5:47:45that we' got two null values at the Top

5:47:47If You Got null values at the top it's

5:47:49going to be a lag function because it

5:47:50doesn't have a value for that Row one in

5:47:53column 2 and row two in column two so

5:47:56that one's always going to be null so we

5:47:57know it's got to be a lag function and

5:47:58we know that the offset is two so it

5:48:00can't be B which is a lead function the

5:48:03lead function is going to look ahead of

5:48:05time rather than looking back in time

5:48:07and the bottom two there have an offset

5:48:09of minus two so the offset in a lag a

5:48:13lead is always positive so those two

5:48:15would not be correct either question two

5:48:19you have the following query analyzing

5:48:21sales data for various products your

5:48:24goal is to analyze the sales data by

5:48:26product name and year but only for the

5:48:28products that have a yearly sales amount

5:48:31of more than $50,000 how would you

5:48:33complete this query so take a look at

5:48:36the different options there have a think

5:48:38about how you would come to that

5:48:40conclusion that this question is asking

5:48:42for and I'll reveal the answer to you

5:48:44shortly so the answer here that we were

5:48:46looking for is B so the question here is

5:48:49asking you to analyze sales data by

5:48:53product name and by year so that is the

5:48:56first clue here in our group by we need

5:48:59to be grouping by the product name and

5:49:02the year not the product key and the

5:49:05date key cuz that wouldn't be the right

5:49:07aggregation for what the question is

5:49:09asking here now the second part of the

5:49:11question is we're asking for products

5:49:14that have a yearly sales amount of more

5:49:16than 50,000 so the yearly sales amount

5:49:18of more than 50,000 that's going to come

5:49:21from our sum of the sales amount and the

5:49:23sum of the sales amount is going to be

5:49:24calculated during that aggregation so we

5:49:26need to be using having here what we're

5:49:28trying to do here is filter out the

5:49:31result of the aggregate we're not

5:49:32filtering out before the aggregate cuz

5:49:34that would just filter out individual

5:49:36rows in this table we want to be looking

5:49:39at the result of an aggregate function

5:49:42where the sum of the sales amount not

5:49:44just individual sales amounts the sum of

5:49:46the sales amount is more than $50,000 so

5:49:49the answer here is B we need the group

5:49:50buy product name and year and then

5:49:53having we need to use the having

5:49:54statement to only return the product

5:49:57names with a yearly sales of more than

5:49:5950,000 question three you're trying to

5:50:01inspect a join between two tables to

5:50:04spot referential integ violations which

5:50:07of the following tsql join types would

5:50:10be easiest to identify keys on both

5:50:13sides of the join that do not have a

5:50:15match on the other side of the join it's

5:50:17a left join B right join C inner join D

5:50:21full outer join e cross join so the

5:50:24answer here is going to be the full

5:50:25alter join so again there's two parts to

5:50:28this question really firstly we need to

5:50:30understand what a referential Integrity

5:50:32violation is so referential Integrity is

5:50:35when you're joining two tables together

5:50:38obviously you're going to be joining on

5:50:39a specific key now a referential

5:50:42Integrity violation would occur when one

5:50:45of the keys in the left hand table is

5:50:48not in the right hand data set and vice

5:50:52versa as well so to be able to identify

5:50:54that using a SQL join we're going to be

5:50:57needing to use the full outer join

5:51:00because the full out to join is going to

5:51:02bring back firstly the instances where

5:51:05those keys do match and then secondly

5:51:08all of the other results that don't

5:51:10match from both tables so when we get

5:51:13this back we're going to bring back some

5:51:15null values where there isn't an

5:51:17appropriate join or is an appropriate

5:51:19match on that join for the table on the

5:51:21left and equally the same on the table

5:51:24on the right so that's going to help us

5:51:25identify referential Integrity

5:51:27violations by investigating where there

5:51:30are null values in that output the left

5:51:32joint and the right joint they're not

5:51:34going to help us here CU we need to

5:51:35identify VI Rel ation on both sides of

5:51:38that join the inner join is only going

5:51:40to return the the matching keys from

5:51:42both sides and the cross join is just

5:51:44going to return us all the different

5:51:45combinations of the keys that exist in

5:51:47these tables so that's not going to help

5:51:49us really identify referential Integrity

5:51:52violations question four you have two

5:51:54warehouses Warehouse 1 and Warehouse 2

5:51:57you want to create a SQL view in

5:51:58Warehouse 2 that combines data from both

5:52:01warehouses Your solution should minimize

5:52:03dwell effort which solution do you

5:52:05recommend is it a a user data pipeline

5:52:07copy activity to copy the table in

5:52:10Warehouse 1 to Warehouse 2 B create a

5:52:13shortcut from Warehouse 1 to Warehouse 2

5:52:16to perform the query C use cross

5:52:18database querying between Warehouse 1

5:52:21and Warehouse 2 or D use a data flow Gen

5:52:242 to read the table in Warehouse 1 and

5:52:27set Warehouse 2 as the destination so

5:52:29the answer here is C use cross database

5:52:32querying between Warehouse 1 and

5:52:34Warehouse 2 now within a warehouse we

5:52:36can query any other Warehouse within the

5:52:39same workspace so this would be the most

5:52:42economical or require the least amount

5:52:44of effort because we can just query

5:52:46directly the data the table Warehouse 1

5:52:48from Warehouse 2 B would be incorrect

5:52:51create a shortcut because we can't

5:52:52actually shortcut from a warehouse one

5:52:54to Warehouse 2 functionality doesn't

5:52:56exist at the moment and the data

5:52:58pipeline copy activity and the data flow

5:53:00Gen 2 wouldn't minimize the development

5:53:04effort so that would technically meet

5:53:06the requirements apart from the

5:53:08requirement that says your solution

5:53:10should minimize development effort and

5:53:12this is something that you might see on

5:53:14quite a few questions within the exam so

5:53:16when you see that it's a bit of a flag

5:53:18that you should always look for the most

5:53:20efficient way of doing things or

5:53:22normally there's a method that requires

5:53:24little to no effort against some that

5:53:26require more development effort so the

5:53:28answer here is C using cross database

5:53:31query question five you're using the

5:53:33tcal query editor and your goal is to

5:53:36add a new column to your data set called

5:53:39salary bins and this column is going to

5:53:41bin a continuous salary variable into

5:53:45three different bins less than 30,000

5:53:48between 30,000 and 60,000 and more than

5:53:5160,000 which functionality should you

5:53:53use to add this new column A add

5:53:56additional column B add column from

5:53:59examples from all columns C add column

5:54:02from examples from selection D duplicate

5:54:04the column e add column from text so the

5:54:07answer here is a add a conditional

5:54:09column so when we're in the TC cor

5:54:12visual query editor there's a few

5:54:14different options there to help us add

5:54:16columns and for this specific use case

5:54:18creating a new column which is going to

5:54:20bin the values in a specific column

5:54:23we're going to be wanting to use the

5:54:24conditional column because that's where

5:54:26we can add in the logic for less than

5:54:29less than or equal to and more than and

5:54:31we can add more than one conditions on

5:54:32that column which would help us achieve

5:54:35our goal of creating the bins in this

5:54:37salary bins column now add columns from

5:54:39examples so these two are functionality

5:54:42within the tsql visual editor but it's

5:54:44going to be very difficult to implement

5:54:47that logic from a columns from examples

5:54:51it's not really a good use case for that

5:54:53type of adding column functionality and

5:54:55D duplicating the column well that's not

5:54:57going to do much it's just going to

5:54:58duplicate the column and similarly e

5:55:00adding a column from text that's also

5:55:03not going to achieve what we're looking

5:55:04for so the answer here is a adding a

5:55:07conditional column we did it

5:55:08congratulations we made it to the end of

5:55:10the study guide this is video 12 out of

5:55:1412 and we've covered an awful lot of

5:55:15ground over the last 6 weeks so well

5:55:18done for sticking with it it's a very

5:55:19tough exam we've covered a lot of

5:55:21different things this exam covers a very

5:55:24wide range of topic from tsql py spark

5:55:28Dax all of the different planning and

5:55:31all that sort of things so to get this

5:55:32far is a really good job now as a bonus

5:55:35I'll be recording one more video in this

5:55:38series basically to help you prepare to

5:55:41take the exam so things you need to know

5:55:44before the exam when you're booking your

5:55:46exam and whil you're in the exam to try

5:55:49and help you get as good a score as

5:55:51possible so thank you very much for

5:55:53joining us in this series I'll see you

5:55:56in the next video which will be the

5:55:57final video I look forward to seeing you

5:55:59all there hey everyone you thought the

TOP TIPS for the exam

5:56:02series was over but I've just got one

5:56:04more bonus video in this dp600 exam

5:56:08preparation course I really wanted to

5:56:10bring you some top tips for the exam so

5:56:13we're not going to be covering any of

5:56:14the technical content but I just have

5:56:16some words of advice or some tips that

5:56:19might help you when you're actually

5:56:21doing the exam itself so this is the

5:56:23bonus round of our course plan if you're

5:56:25just joining us here in this video I

5:56:27recommend you go back through all of the

5:56:28last 12 videos CU that's where we cover

5:56:30most of the content and I wanted to talk

5:56:31to you about booking your exam preparing

5:56:35for the exam some some advice for during

5:56:37the exam and then also what you should

5:56:39do after the exam as well so when you're

5:56:41booking the exam if you've used the AI

5:56:44skills challenge voucher then make sure

5:56:47you schedule your exam before the

5:56:49expiration date and that means you have

5:56:51to schedule the exam to be conducted

5:56:54before that expiration date now I think

5:56:56that was the 24th of June 2024 but

5:56:59double check that in your own emails

5:57:01that you have that date there now if you

5:57:03don't have that AI skills voucher that

5:57:06was given away maybe 2 months ago now

5:57:08then you can get a 50% off voucher for

5:57:12the exam by completing the cloud skills

5:57:15challenge I'll leave a link to that in

5:57:16the description so if you want 50% of

5:57:18the dp600 exam and you still haven't

5:57:21booked your exam yet or you still

5:57:22haven't got a voucher for it you can use

5:57:24that there and I've said this before but

5:57:26I always recommend if possible to take

5:57:28the exam in person now you might have

5:57:31heard there's a lot of horror stories

5:57:33for people that take it online you know

5:57:34have an online proor experience and

5:57:37they're very very busy and there's

5:57:38always problems with technology and you

5:57:41know you have to clear your desk and

5:57:43clear your room and there's all these

5:57:45kind of hurdles that if you can avoid by

5:57:48going to an inperson test center I would

5:57:50very much recommend doing that

5:57:52appreciate not everyone lives near a

5:57:54test center but if you do or you can

5:57:56commute to one just for the exam I would

5:57:58definitely recommend that reduces a lot

5:58:00of stress you just walk in take the exam

5:58:03and walk out so preparing for the exam

5:58:05well obviously you can go back through

5:58:07these videos and look in the school

5:58:10Community as well I've got lots of notes

5:58:12in there the reason I kind of created

5:58:14those was to help people who are you

5:58:15know just about to take the exam they

5:58:17can go through the key points and just

5:58:19make sure they remember everything as

5:58:21they go through each of the different

5:58:23sections in that school community and if

5:58:25you're still not sure about anything in

5:58:27any section of the exam you can dig into

5:58:29the further learning resources that I've

5:58:31linked in every section there now one of

5:58:33the key resources that I would

5:58:34definitely use when you're preparing is

5:58:36the official practice assessment for the

5:58:38dp600 exam and again I'll leave a link

5:58:40to that in the description now you can

5:58:42go through this multiple times I think

5:58:44it's a 50 question assessment but you

5:58:46can refresh the page and you get fresh

5:58:48questions when you refresh I don't know

5:58:50exactly how many different questions

5:58:51there are in that exam set but there's

5:58:54definitely more than 50 so go through

5:58:56that practice assessment numerous times

5:58:58to get a really good idea of the types

5:59:00of questions that you're going to be

5:59:01asked and to highlight any gaps in your

5:59:04knowledge another really good resource

5:59:06is the Microsoft learn together series

5:59:08that's kind of been running side by side

5:59:10with this series I think we started

5:59:12around the same time I think they

5:59:13finished maybe a week or two ago now

5:59:15this is a really good series delivered

5:59:17by Microsoft MVPs so there's lots and

5:59:20lots of content that you can go through

5:59:22there to dig in a bit deeper about any

5:59:24of the aspects in the dp600 study guide

5:59:27again I'll leave a link to that in the

5:59:28description as well as the technical

5:59:30content for Microsoft MVPs they've also

5:59:33got a few lectures or a few videos that

5:59:36describe to you what the exam entails so

5:59:39if you've never taken a Microsoft exam

5:59:42before I would definitely recommend

5:59:44watching some of those videos I think

5:59:45they're at the end of that Series where

5:59:47they walk through types of questions you

5:59:49might get how to prepare for it how much

5:59:51time you have all of that kind of stuff

5:59:53just so that you can enter that exam as

5:59:55prepared as possible now as well as that

5:59:57there's lots of other great content from

5:59:59other people in the community definitely

6:00:02recommend data Mozart Nicola's blog

6:00:05there and also you you Channel as well

6:00:07now the data guy newsletter on LinkedIn

6:00:10something managed by Abu Baka he

6:00:12basically collated a lot of learning

6:00:14resources for each section of the study

6:00:16guide definitely recommend that

6:00:18similarly there's a post on serverless

6:00:20squl by Andy Cutler which is basically a

6:00:22guide to the exam as well and in there

6:00:25he highlights a lot of resources for the

6:00:27exam as well so during the exam I would

6:00:30say don't spend too long on one question

6:00:33if you get stuck then you can flag it

6:00:35for VI and you can come back to it at

6:00:37the end if you have time now a lot of

6:00:39people are saying that there's not much

6:00:42time in this exam you have 100 minutes

6:00:45to actually do the exam and there's

6:00:47normally around 55 to 60 Questions and

6:00:50some of the questions are very wordy so

6:00:52it might take you 30 seconds or even a

6:00:54minute just to try and understand what

6:00:56they're asking you so if you get stuck

6:00:58and you don't know the answer don't

6:01:00waste too long on one question just flag

6:01:02it for review and come back to it at the

6:01:04end now you will receive a at least one

6:01:06case study question where you'll be

6:01:08given a lot of context about a business

6:01:11and you'll be asked a series of

6:01:13questions afterwards now I definitely

6:01:15recommend that you use a pen and paper

6:01:17to draw the case study architecture as

6:01:20you're reading that question because

6:01:22normally these case studies they're very

6:01:24complex and they'll describe a scenario

6:01:27to you for example oh company has

6:01:30capacity one and in capacity one there's

6:01:32workspace one and workspace 2 and in

6:01:34workspace one there's Lake housee one

6:01:36and Lake housee when you're reading this

6:01:38you don't really take any of it in so I

6:01:40definitely recommend drawing it as

6:01:42you're reading it this can really help

6:01:44you when you go forward to answering the

6:01:46questions because you don't have to keep

6:01:48on going back to the the case study the

6:01:50context when you're answering the

6:01:52questions you can just refer to your

6:01:54your diagram if you don't know an answer

6:01:56guess you don't lose marks for an

6:01:58incorrect guess so you might as well and

6:02:01bear in mind that the questions were

6:02:03created many months ago so I don't think

6:02:06they've updated the questions since it

6:02:08was first came out so bear that in mind

6:02:11so a lot of the more recent features

6:02:14they are not going to be the correct

6:02:15answer because they weren't in general

6:02:17availability when the questions were

6:02:19created so bear that in mind now a lot

6:02:21of the questions I would argue are

6:02:23ambiguous or at least the answers so

6:02:25what you have to bear in mind or try and

6:02:27keep asking yourself is what do I think

6:02:29that the Microsoft examiner is expecting

6:02:32to see here so don't try and be too

6:02:34clever always think about what AIC

6:02:35Microsoft trying to teach us about

6:02:37Microsoft fabric how do they want us to

6:02:39use the platform and always answer your

6:02:41questions with that in mind so after the

6:02:43exam so you will know if you pass or

6:02:46fail directly after the exam you get the

6:02:48results straight after now if you fail I

6:02:51would say don't be too hard on yourself

6:02:52honestly it's a very very tough exam it

6:02:55expects you to be familiar with a very

6:02:57wide range of topics right for starters

6:03:01SQL py spark M Dax plus all of the

6:03:04planning stuff and all the lake houses

6:03:07data pipelines data flows all of this

6:03:10stuff is a very wide ranging exam that

6:03:12you need to be at least familiar with so

6:03:14if you do fail don't be too hard on

6:03:16yourself at all and if you pass then

6:03:18well congratulations you can be very

6:03:20very proud of your achievement so that

6:03:22is all I have thank you so much for

6:03:25joining me in this series and the best

6:03:27of luck for all of you that are taking

6:03:29the exam in the future thank you so much

6:03:31for joining me in this series I'm going

6:03:34to be taking a short break now and be

6:03:36back on the Channel with more videos

6:03:38teaching Fabric in the future thank you

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.