Free YouTube Transcribe

Video transcript

Master Data Analysis with ChatGPT (in just 12 minutes)

Jeff Su · 2,269 words · 11 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

ChatGPT for Data Analysis

0:00Let's get straight to the point. All of

0:01us work with data regardless of our

0:03role. But very few of us were taught how

0:06to analyze data in a structured way. So

0:09in this video, I'll bridge that gap by

0:11sharing a simple three-step framework

0:13that basically turns chatbt into our

0:15personal data analyst with zero

0:17technical skills required. Let's get

0:19started. First, a bit of context. In all

0:21my past roles, management consultant,

0:22account manager, product marketing

0:24manager, I've had to work with data and

0:26present findings. Like most of you

0:28though, I've never received any formal

0:29training in data analysis. But I also

0:31knew AI could help. So after taking the

0:33top AI for data analysis course on

0:35Corsera, I learned that the key was

0:37giving Chachi PT a proven framework to

0:39follow and have the AI perform all the

0:42hard analytical tasks on our behalf.

The DIG Data Analysis Framework

0:45Diving right in the framework I learned

0:46is called DIG description,

0:48introspection, and goal setting. And in

0:50a nutshell, by using chatbt to apply the

0:52dig framework to any data set, we're

0:54able to number one, understand data

0:55we've never seen before in a matter of

0:57minutes instead of hours. And number

0:59two, extract insights that we as

1:01non-data analysts would have missed.

1:03Here's a simple visualization. When you

1:05get handed a spreadsheet with no

1:07context, you're at 0% understanding. But

1:10with every dig prompt you input into

1:12chatbt, your understanding of the data

1:15increases. And by the end of the dig

1:17framework, you've uncovered insights

1:19that would have taken hours to find

1:21manually, if you found them at all. Two

1:23quick things before diving to a real

1:24case study. First, I'm using a free

1:26Apple TV Plus data set that you can

1:28download and follow along. And it's

1:30actually pretty cool to analyze real

1:31data from popular TV shows and movies

1:33like Avatar the Last Airbender, The

1:35Godfather, and Sherlock. Second, the

1:37industry standard framework is actually

1:39called EDA, exploratory data analysis.

1:42But I'm using dig in this video because

1:43one, the principles are the same. And

1:45second, the professor on Corsera,

1:47probably used dig because it's easier to

Step 1: Description

1:49remember. Step one, description. Picture

1:52this hypothetical scenario. Your

1:53colleague, let's call him Tim Cookie,

1:55just rage quit for absolutely no reason

1:57and left you with a spreadsheet with

1:59zero context. At this point, we need

2:01ChachiBt to explain or describe what's

2:04in the file as quickly and effectively

2:06as possible. So let's open up ChachiBT,

2:08upload the data set, select the latest

2:11reasoning model, and start with a first

2:13description prompt. List all the columns

2:16in the attached spreadsheet and show me

2:18a sample of data from each column. The

2:20reason we start off with this prompt is

2:22because it forces chache to actually

2:24look at every single column in our data

2:26set and more importantly gives us a

2:28quick overview of the data we're working

2:30with. Looking at this output, I want to

2:32point out two things. First, having

2:33ChachiBT return just one sample output

2:36is much easier for us, the human to

2:38digest versus having to make sense of an

2:41entire spreadsheet we've never seen

2:43before. Second, the sample is selected

2:44as Forest Gump, a classic. Nice. And it

2:47returned all eight columns from the

2:48original spreadsheet. Great. But the

2:50release year is 994.0 and there are two

2:53genres separated by a comma. So, this

2:56might represent issues for Chacht down

2:58the road. So, we want to make a note of

2:59that. And I'm actually not really sure

3:01what IMDb ID means. Does every show in T

3:04does every show in movie have a unique

3:06ID? I'm not sure. So I might want to

3:08follow up and confirm. Next description

3:10prompt number two. Oh, I'll link to all

3:11these prompts down below. By the way, uh

3:13take five more random samples of the

3:15data for each column to make sure you

3:17understand the format and type of

3:18information in each column. So why are

3:21we asking for more samples? Because that

3:23one sample we received might be an

3:24outlier and therefore misleading.

3:27Multiple samples help us spot

3:28inconsistencies. Looking at this output,

3:30we see there are TV and movies under

3:33type, but we knew that already. This TV

3:36show has three genres. Okay, this these

3:38movies have one. All right, and uh some

3:42of this is available in one country,

3:44others are available in multiple

3:45countries. Okay, so our understanding of

3:47the data set is increasing. Moving on to

3:50description prompt number three. Run a

3:52data quality check on each column.

3:54specifically look for missing or empty

3:56values, unexpected formats or data

3:59types, outliers or suspicious values.

4:02This is pretty self-explanatory. We want

4:03Chhatabt to explicitly tell us if

4:06there's anything weird about the data we

4:07should know about before proceeding with

4:09our analysis. Okay, there's several

4:11tables here. This first one tells us how

4:13many values are missing from each

4:14column. So, for the title column, we're

4:16missing 589 values uh representing 3.1%.

4:20Going down, 10% is pretty high. Whoa.

4:22Okay. 99.7. This number tells us that

4:26for the available countries column,

4:28we're missing 99.7%

4:31of the values. I can double check really

4:33quickly by going into the raw data set

4:35and doing a quick filter or sort rather.

4:38Going to sort this and I scroll down.

4:40Yeah, most of these rows are completely

4:44empty for the available countries

4:45column. This means we should not perform

4:48any geographical analysis with this data

4:51set because we're missing that

4:52information. At this point, it should be

4:54pretty clear that first, although Chachi

4:56PT is not doing 100% of work for us,

4:58it's making our job as a human analyst

5:01much easier. Second, remember how the

5:03goal of the description step is for us

5:05to understand the data set as

5:08effectively as possible. I'm not going

5:10to waste your time here, but in real

5:12life, I would have at this point asked

5:14follow-up questions like, "Hey, what

5:16does this TT number mean here?" And

5:19chatbt would have confirmed it is the

5:22unique IMDb number. By the way, if you

5:24use Google Workspace tools at work, you

5:26might want to join my newsletter to

5:28receive an insanely actionable tip every

5:30week. Link down below. Next up, we have

Step 2: Introspection

5:32introspection. And the purpose of this

5:34step is to have Chat JBT brainstorm

5:36questions it could answer with our data.

5:38This shows whether ChachiPT truly gets

5:41our data and often services insights we

5:44hadn't considered. Prompt one, tell me

5:4610 interesting questions we could answer

5:48with this data set and explain why each

5:50would be valuable. And long story short,

5:53good questions mean CatchPT understands

5:55our data. Bad questions equal there's a

5:57misunderstanding that needs fixing

5:59before we proceed. These first three

6:02questions are solid. How has Apple TV's

6:04yearly output grown since launch? If

6:06we're putting out more TV shows and

6:08movies yearon year, it might mean we're

6:09capturing more market share. Uh, what

6:12share of releases are movies versus

6:13series each year? This might tell us

6:15about viewer behavior. Are we trending

6:17more towards TV shows or movies? I love

6:19this one. Which genres dominate the

6:21catalog and how have they shifted over

6:22time? Imagine you were on the Apple

6:24content team, right? You might want to

6:26invest more in the most popular genre

6:28next year. Or maybe additional analysis

6:31tells you the genre is oversaturated, so

6:33you want to pull back. Prompt two. For

6:35the first three questions, tell me

6:37exactly which columns you need to use

6:39and whether the current data is

6:40sufficient to answer it. This basically

6:43forces Chachbt to show its work and

6:45tells us whether or not we can perform

6:47these analyses. We see that for question

6:49one, yes, we just need to fix 0.3% of

6:53non-numeric entries. We can ignore that.

6:55For question two, yes, we need to do

6:57some light data cleanup. That's fine. We

6:59can tell chatbt to do it. For number

7:01three, yes, we have all the information

7:02we need to perform the analysis.

7:04Awesome. Prompt three is my personal

7:06favorite for the introspection step.

7:07What questions do you think someone

7:08would want to ask about this data but we

7:11can't answer due to missing information.

7:13This basically surfaces gaps in our data

7:16set and helps us manage our boss's

7:18expectations about what insights we can

7:21uncover. Here we see questions like

7:22what's the most watched genre? We can't

7:24answer that because we don't have the

7:25viewing metrics or from an ROI

7:28perspective which genres deliver the

7:30best cost per hour of content. We don't

7:32have the production budget, revenue, or

7:34cost fields. But here's where it gets

7:36interesting. What if I had access to

7:38some of that data? Let's say my friends

7:39at Apple invited me into Apple Part and

7:41I hacked their servers. That was

7:43obviously a joke. Very unrealistic

7:45scenario. I don't really have friends.

7:48So, I created a fake second data set.

7:51And to be very clear, this is made up.

7:52Don't report me to Apple. uh with the

7:54IMDb ID in column A, total viewership in

7:57column B, and the total cost of

7:59producing that show or movie in column

8:01C. I can now upload this onto the same

8:03ChachiPT thread and say, I just received

8:06this data set from a colleague. Your

8:07task is to explore and explain the

8:09relationships between this new data set

8:11with the original one and how they might

8:14be used to join the data together. After

8:16running for a bit, ChachiP confirms we

8:18can use the IMDb ID field to join the

8:21two data sets together and even gives us

8:24suggestions on how we can use this newly

8:26merged table. For example, we can

8:28calculate the cost per viewer ROI by

8:32genre. For example, after instructing

8:34Chacht to merge the data sets using the

8:36IMDb ID field, it gives us a sample

8:40output of this newly merged spreadsheet.

8:43Right? We see that for Forest Gump, if I

8:45scroll all the way to the right, yes, it

8:47now has the total viewership and total

8:49cost data in its row along with

8:53everything else. And I can even click

8:55here to download this merge CSV file.

8:59That's pretty awesome. Quick note, I'm

9:00focusing on the core dig prompts in this

9:02video to not waste your time. In real

9:05life, I would have branched out much

9:07sooner. For example, earlier when

9:09Chachib mentioned genre popularity, I'd

9:12immediately ask for that analysis and

9:14keep digging based on what it find. The

Step 3: Goal Setting

9:16third step, goal setting. This is

9:18extremely important to get right because

9:20imagine if our manager asked us to

9:22analyze the sales data and after working

9:24hard to create 20 beautiful slides, our

9:26manager says, "Wait, I just wanted to

9:28know if we should discontinue product

9:30X." This is what happens when our

9:32manager is an idiot. I mean uh when we

9:34analyze data without setting clear

9:35goals, we have something that's

9:37technically correct but ultimately

9:39useless. Obviously, the prompts we use

9:41in this step depend on the specific

9:42goal. So, I'll just share one example.

9:44Here's the prompt. My goal is to

9:46understand and you specify your goal

9:47here. I want to understand what content

9:49Apple TV should invest in next. Given

9:52this goal, which aspects of the data

9:54should we focus on? This is basically

9:56like giving CatchBT a mission briefing.

9:58It helps the AI prioritize what's

10:00important and ignore what's not. Okay,

10:03this is very useful. So, Traded D first

10:05breaks down our options for us. For

10:07example, if you're in the Apple content

10:09team, you might care about viewership,

10:11audience demand, and content supply,

10:12right? So, you would want to do this. If

10:14you're in the finance team, you might

10:16want to learn more about unit economics

10:18and do that. Scrolling down, we even see

10:20a step-by-step road map. So, first we

10:22might want to clean our data. Okay.

10:24Then, we build a genre scorecard. That

10:26could be very interesting. Then we rank

10:28our opportunities, layer in trend

10:30velocity. I would have never thought to

10:32do that. And finally, stress test with

10:34outliers. Makes sense. And here's the

10:36type of insight this process would

10:38surface. Uh, true crime series deliver

10:40three times the medium views of all

10:42series. They cost 18% less per finished

10:44hour and have climbed from 4% to 9%

10:47share of total watch time in the last 3

10:49years. Okay. Wow, that's that's really

10:51impressive. Pro tip. A final question I

10:53always like to ask Chad PT before any

Bonus Prompt

10:55presentation is what are the key

10:57questions someone reading my analysis

10:59would ask and how should we proactively

11:02address them. This prompt

11:03single-handedly saved my ass multiple

11:05times by anticipating but Jeff what

11:08about this questions from managers and

11:11overly ambitious peers trying to put me

11:12down. Just kidding. Everyone loves me.

11:14How could they not? Two things I'd like

11:15to leave you with. First, the dig

11:17framework plus chajbt levels the playing

11:20field for regular untrained people like

11:22us. It's a simple repeatable process we

11:25can all use immediately. Second,

11:27although I cover the essentials today,

11:29the full Corsera course touches on other

11:31important concepts like how to mitigate

11:33hallucinations and debug weird data

11:35errors. So, if you want to level up your

11:37data skills, sign up for Corsera using

11:38the link in the description to take

11:40advantage of my special offer of 40% off

11:42for 3 months of Corsera Plus. If you

11:44enjoyed this, check out my comprehensive

11:46Chachi PT pro tips video next. See you

11:49all there. And in the meantime, have a

11:51great one.

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com: free, unlimited, no sign-up.