Full transcript
ChatGPT for Data Analysis
0:00Let's get straight to the point. All of
0:01us work with data regardless of our
0:03role. But very few of us were taught how
0:06to analyze data in a structured way. So
0:09in this video, I'll bridge that gap by
0:11sharing a simple three-step framework
0:13that basically turns chatbt into our
0:15personal data analyst with zero
0:17technical skills required. Let's get
0:19started. First, a bit of context. In all
0:21my past roles, management consultant,
0:22account manager, product marketing
0:24manager, I've had to work with data and
0:26present findings. Like most of you
0:28though, I've never received any formal
0:29training in data analysis. But I also
0:31knew AI could help. So after taking the
0:33top AI for data analysis course on
0:35Corsera, I learned that the key was
0:37giving Chachi PT a proven framework to
0:39follow and have the AI perform all the
0:42hard analytical tasks on our behalf.
The DIG Data Analysis Framework
0:45Diving right in the framework I learned
0:46is called DIG description,
0:48introspection, and goal setting. And in
0:50a nutshell, by using chatbt to apply the
0:52dig framework to any data set, we're
0:54able to number one, understand data
0:55we've never seen before in a matter of
0:57minutes instead of hours. And number
0:59two, extract insights that we as
1:01non-data analysts would have missed.
1:03Here's a simple visualization. When you
1:05get handed a spreadsheet with no
1:07context, you're at 0% understanding. But
1:10with every dig prompt you input into
1:12chatbt, your understanding of the data
1:15increases. And by the end of the dig
1:17framework, you've uncovered insights
1:19that would have taken hours to find
1:21manually, if you found them at all. Two
1:23quick things before diving to a real
1:24case study. First, I'm using a free
1:26Apple TV Plus data set that you can
1:28download and follow along. And it's
1:30actually pretty cool to analyze real
1:31data from popular TV shows and movies
1:33like Avatar the Last Airbender, The
1:35Godfather, and Sherlock. Second, the
1:37industry standard framework is actually
1:39called EDA, exploratory data analysis.
1:42But I'm using dig in this video because
1:43one, the principles are the same. And
1:45second, the professor on Corsera,
1:47probably used dig because it's easier to
Step 1: Description
1:49remember. Step one, description. Picture
1:52this hypothetical scenario. Your
1:53colleague, let's call him Tim Cookie,
1:55just rage quit for absolutely no reason
1:57and left you with a spreadsheet with
1:59zero context. At this point, we need
2:01ChachiBt to explain or describe what's
2:04in the file as quickly and effectively
2:06as possible. So let's open up ChachiBT,
2:08upload the data set, select the latest
2:11reasoning model, and start with a first
2:13description prompt. List all the columns
2:16in the attached spreadsheet and show me
2:18a sample of data from each column. The
2:20reason we start off with this prompt is
2:22because it forces chache to actually
2:24look at every single column in our data
2:26set and more importantly gives us a
2:28quick overview of the data we're working
2:30with. Looking at this output, I want to
2:32point out two things. First, having
2:33ChachiBT return just one sample output
2:36is much easier for us, the human to
2:38digest versus having to make sense of an
2:41entire spreadsheet we've never seen
2:43before. Second, the sample is selected
2:44as Forest Gump, a classic. Nice. And it
2:47returned all eight columns from the
2:48original spreadsheet. Great. But the
2:50release year is 994.0 and there are two
2:53genres separated by a comma. So, this
2:56might represent issues for Chacht down
2:58the road. So, we want to make a note of
2:59that. And I'm actually not really sure
3:01what IMDb ID means. Does every show in T
3:04does every show in movie have a unique
3:06ID? I'm not sure. So I might want to
3:08follow up and confirm. Next description
3:10prompt number two. Oh, I'll link to all
3:11these prompts down below. By the way, uh
3:13take five more random samples of the
3:15data for each column to make sure you
3:17understand the format and type of
3:18information in each column. So why are
3:21we asking for more samples? Because that
3:23one sample we received might be an
3:24outlier and therefore misleading.
3:27Multiple samples help us spot
3:28inconsistencies. Looking at this output,
3:30we see there are TV and movies under
3:33type, but we knew that already. This TV
3:36show has three genres. Okay, this these
3:38movies have one. All right, and uh some
3:42of this is available in one country,
3:44others are available in multiple
3:45countries. Okay, so our understanding of
3:47the data set is increasing. Moving on to
3:50description prompt number three. Run a
3:52data quality check on each column.
3:54specifically look for missing or empty
3:56values, unexpected formats or data
3:59types, outliers or suspicious values.
4:02This is pretty self-explanatory. We want
4:03Chhatabt to explicitly tell us if
4:06there's anything weird about the data we
4:07should know about before proceeding with
4:09our analysis. Okay, there's several
4:11tables here. This first one tells us how
4:13many values are missing from each
4:14column. So, for the title column, we're
4:16missing 589 values uh representing 3.1%.
4:20Going down, 10% is pretty high. Whoa.
4:22Okay. 99.7. This number tells us that
4:26for the available countries column,
4:28we're missing 99.7%
4:31of the values. I can double check really
4:33quickly by going into the raw data set
4:35and doing a quick filter or sort rather.
4:38Going to sort this and I scroll down.
4:40Yeah, most of these rows are completely
4:44empty for the available countries
4:45column. This means we should not perform
4:48any geographical analysis with this data
4:51set because we're missing that
4:52information. At this point, it should be
4:54pretty clear that first, although Chachi
4:56PT is not doing 100% of work for us,
4:58it's making our job as a human analyst
5:01much easier. Second, remember how the
5:03goal of the description step is for us
5:05to understand the data set as
5:08effectively as possible. I'm not going
5:10to waste your time here, but in real
5:12life, I would have at this point asked
5:14follow-up questions like, "Hey, what
5:16does this TT number mean here?" And
5:19chatbt would have confirmed it is the
5:22unique IMDb number. By the way, if you
5:24use Google Workspace tools at work, you
5:26might want to join my newsletter to
5:28receive an insanely actionable tip every
5:30week. Link down below. Next up, we have
Step 2: Introspection
5:32introspection. And the purpose of this
5:34step is to have Chat JBT brainstorm
5:36questions it could answer with our data.
5:38This shows whether ChachiPT truly gets
5:41our data and often services insights we
5:44hadn't considered. Prompt one, tell me
5:4610 interesting questions we could answer
5:48with this data set and explain why each
5:50would be valuable. And long story short,
5:53good questions mean CatchPT understands
5:55our data. Bad questions equal there's a
5:57misunderstanding that needs fixing
5:59before we proceed. These first three
6:02questions are solid. How has Apple TV's
6:04yearly output grown since launch? If
6:06we're putting out more TV shows and
6:08movies yearon year, it might mean we're
6:09capturing more market share. Uh, what
6:12share of releases are movies versus
6:13series each year? This might tell us
6:15about viewer behavior. Are we trending
6:17more towards TV shows or movies? I love
6:19this one. Which genres dominate the
6:21catalog and how have they shifted over
6:22time? Imagine you were on the Apple
6:24content team, right? You might want to
6:26invest more in the most popular genre
6:28next year. Or maybe additional analysis
6:31tells you the genre is oversaturated, so
6:33you want to pull back. Prompt two. For
6:35the first three questions, tell me
6:37exactly which columns you need to use
6:39and whether the current data is
6:40sufficient to answer it. This basically
6:43forces Chachbt to show its work and
6:45tells us whether or not we can perform
6:47these analyses. We see that for question
6:49one, yes, we just need to fix 0.3% of
6:53non-numeric entries. We can ignore that.
6:55For question two, yes, we need to do
6:57some light data cleanup. That's fine. We
6:59can tell chatbt to do it. For number
7:01three, yes, we have all the information
7:02we need to perform the analysis.
7:04Awesome. Prompt three is my personal
7:06favorite for the introspection step.
7:07What questions do you think someone
7:08would want to ask about this data but we
7:11can't answer due to missing information.
7:13This basically surfaces gaps in our data
7:16set and helps us manage our boss's
7:18expectations about what insights we can
7:21uncover. Here we see questions like
7:22what's the most watched genre? We can't
7:24answer that because we don't have the
7:25viewing metrics or from an ROI
7:28perspective which genres deliver the
7:30best cost per hour of content. We don't
7:32have the production budget, revenue, or
7:34cost fields. But here's where it gets
7:36interesting. What if I had access to
7:38some of that data? Let's say my friends
7:39at Apple invited me into Apple Part and
7:41I hacked their servers. That was
7:43obviously a joke. Very unrealistic
7:45scenario. I don't really have friends.
7:48So, I created a fake second data set.
7:51And to be very clear, this is made up.
7:52Don't report me to Apple. uh with the
7:54IMDb ID in column A, total viewership in
7:57column B, and the total cost of
7:59producing that show or movie in column
8:01C. I can now upload this onto the same
8:03ChachiPT thread and say, I just received
8:06this data set from a colleague. Your
8:07task is to explore and explain the
8:09relationships between this new data set
8:11with the original one and how they might
8:14be used to join the data together. After
8:16running for a bit, ChachiP confirms we
8:18can use the IMDb ID field to join the
8:21two data sets together and even gives us
8:24suggestions on how we can use this newly
8:26merged table. For example, we can
8:28calculate the cost per viewer ROI by
8:32genre. For example, after instructing
8:34Chacht to merge the data sets using the
8:36IMDb ID field, it gives us a sample
8:40output of this newly merged spreadsheet.
8:43Right? We see that for Forest Gump, if I
8:45scroll all the way to the right, yes, it
8:47now has the total viewership and total
8:49cost data in its row along with
8:53everything else. And I can even click
8:55here to download this merge CSV file.
8:59That's pretty awesome. Quick note, I'm
9:00focusing on the core dig prompts in this
9:02video to not waste your time. In real
9:05life, I would have branched out much
9:07sooner. For example, earlier when
9:09Chachib mentioned genre popularity, I'd
9:12immediately ask for that analysis and
9:14keep digging based on what it find. The
Step 3: Goal Setting
9:16third step, goal setting. This is
9:18extremely important to get right because
9:20imagine if our manager asked us to
9:22analyze the sales data and after working
9:24hard to create 20 beautiful slides, our
9:26manager says, "Wait, I just wanted to
9:28know if we should discontinue product
9:30X." This is what happens when our
9:32manager is an idiot. I mean uh when we
9:34analyze data without setting clear
9:35goals, we have something that's
9:37technically correct but ultimately
9:39useless. Obviously, the prompts we use
9:41in this step depend on the specific
9:42goal. So, I'll just share one example.
9:44Here's the prompt. My goal is to
9:46understand and you specify your goal
9:47here. I want to understand what content
9:49Apple TV should invest in next. Given
9:52this goal, which aspects of the data
9:54should we focus on? This is basically
9:56like giving CatchBT a mission briefing.
9:58It helps the AI prioritize what's
10:00important and ignore what's not. Okay,
10:03this is very useful. So, Traded D first
10:05breaks down our options for us. For
10:07example, if you're in the Apple content
10:09team, you might care about viewership,
10:11audience demand, and content supply,
10:12right? So, you would want to do this. If
10:14you're in the finance team, you might
10:16want to learn more about unit economics
10:18and do that. Scrolling down, we even see
10:20a step-by-step road map. So, first we
10:22might want to clean our data. Okay.
10:24Then, we build a genre scorecard. That
10:26could be very interesting. Then we rank
10:28our opportunities, layer in trend
10:30velocity. I would have never thought to
10:32do that. And finally, stress test with
10:34outliers. Makes sense. And here's the
10:36type of insight this process would
10:38surface. Uh, true crime series deliver
10:40three times the medium views of all
10:42series. They cost 18% less per finished
10:44hour and have climbed from 4% to 9%
10:47share of total watch time in the last 3
10:49years. Okay. Wow, that's that's really
10:51impressive. Pro tip. A final question I
10:53always like to ask Chad PT before any
Bonus Prompt
10:55presentation is what are the key
10:57questions someone reading my analysis
10:59would ask and how should we proactively
11:02address them. This prompt
11:03single-handedly saved my ass multiple
11:05times by anticipating but Jeff what
11:08about this questions from managers and
11:11overly ambitious peers trying to put me
11:12down. Just kidding. Everyone loves me.
11:14How could they not? Two things I'd like
11:15to leave you with. First, the dig
11:17framework plus chajbt levels the playing
11:20field for regular untrained people like
11:22us. It's a simple repeatable process we
11:25can all use immediately. Second,
11:27although I cover the essentials today,
11:29the full Corsera course touches on other
11:31important concepts like how to mitigate
11:33hallucinations and debug weird data
11:35errors. So, if you want to level up your
11:37data skills, sign up for Corsera using
11:38the link in the description to take
11:40advantage of my special offer of 40% off
11:42for 3 months of Corsera Plus. If you
11:44enjoyed this, check out my comprehensive
11:46Chachi PT pro tips video next. See you
11:49all there. And in the meantime, have a
11:51great one.