Free YouTube Transcribe

Video transcript

Building a Chess Coach — Anant Dole and Asbjorn Steinskog, Take Take Take

AI Engineer · 3,238 words · 15 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:14Afternoon, everyone. So, our next talk

0:17will be something a little bit

0:19different. We're going to dive into the

0:21world of chess. Quick show of hands, who

0:24has heard of Magnus Carlsen?

0:27Okay, fantastic. No introduction needed,

0:30but widely considered the best chess

0:32player in the world. He also founded a

0:34company called Play Magnus.

0:37Uh this is where myself, Ananth, and my

0:39colleague Asbjørn currently work at. And

0:42we're going to talk to you today about

0:43how we built our AI chess coach that now

0:46you can use and is in production.

0:50So, first up, quick agenda. We'll

0:52quickly discuss a bit more about Play

0:53Magnus,

0:54what it is we actually built, what it is

0:56we actually launched. Uh Asbjørn will

0:58then go into a quick history of chess

1:00and AI. A lot of links there. We'll

1:02briefly touch on why LLMs are actually

1:04bad at chess and how we managed to solve

1:07this problem.

1:08We're then going to deep dive into

1:09actually understanding our game review

1:11and sort of closing the loop with our

1:12autonomous agent. And you'll get a demo.

1:15And then finally, some latency versus

1:17quality trade-offs, as this is a

1:19consumer-focused AI application. And

1:22then lastly, some learnings.

1:25So, first up, what is Play Magnus? In

1:27its simplest form today, it's currently

1:29an iOS and Android application. You can

1:31go on and play your friends, and you can

1:34post about your games.

1:36What's relevant for our particular talk

1:38is that after you play a game, you get

1:40presented with our game review. And this

1:43is powered by our AI pipeline.

1:46So, for example, just showing you how it

1:48works. In this particular position, it's

1:51leading to a checkmate. The last move

1:53that white has played has moved the

1:55knight from this yellow square over here

1:57on F3, captured the pawn on E5.

2:00It is a brilliant move, so automatically

2:02gets the brilliant sort of notation. And

2:04the commentary below is actually

2:06generated by our system. And we're using

2:09an LLM and the pipeline we'll get into

2:11in a second. But what's quite

2:13interesting about it is we're able to

2:15give you the nuance of why it is a

2:17tactic, what detectors from a positional

2:20and tactical sense have fired, what are

2:22the threats you're trying to do, and

2:24actually explain sort of the why behind

2:26the move. So, that's the system we're

2:28going to be talking about today.

2:31Finally, on the last little step of our

2:34of our application, we've started

2:36revealing insights about your play. And

2:38this could be things like how accurate

2:40you played in a particular game phase,

2:42maybe your current rating, or your

2:44current depth in a particular opening.

2:46And these insights form the next layer

2:48of analysis that we present to the

2:50coach, who then gives them back to you

2:52as opportunities for learning and

2:53improving. We hope by using this, you'll

2:56be able to improve and become better at

2:58the game.

3:00All right. So, first a brief history of

3:02chess and AI since they've been

3:03intertwined for so long, just to give

3:05you a little bit of backstory. 1949,

3:08Claude Shannon, the OG Claude, wrote the

3:11the paper programming a computer to play

3:13chess. And here he envisioned that or he

3:16proposed that there there are two types

3:18of of chess engines, type A and type B.

3:22Uh type A were these brute force engines

3:25that search through all possible

3:28possible moves and figure out the best

3:30move. While type B

3:32were those who we know from from 2017

3:35and and onward that can selectively

3:38pick out the the best moves.

3:40Back then, he assumed that we would need

3:43type B computers to to play chess

3:45because

3:46computers were so weak back then, you

3:48couldn't search through the whole whole

3:50tree of of moves. But computers quickly

3:52became better, and people just started

3:55scaling these type A computers.

3:57Uh they got better and better until they

3:59in 1997,

4:01Deep Blue versus Kasparov,

4:03uh the first time a chess engine beat

4:06the best chess player at the time.

4:09Uh so, people didn't really bother about

4:11these type B computers for a while,

4:13these uh intuitive

4:15uh engines, until

4:18DeepMind, shout out to DeepMind, uh

4:21released first AlphaGo, because Go is a

4:24much more complex game than chess. So,

4:26you you can't solve this with these type

4:28A computers. You would need this

4:29intuitive approach, neural network

4:31approach that actually selectively uh

4:33figure out which lines to to calculate.

4:35Uh but after that, they released

4:37AlphaZero, who could play not only Go,

4:39but also chess and and shogi.

4:41Uh

4:42and

4:44some some years later,

4:45uh LLMs came, and people started playing

4:48chess against the LLMs, and quickly

4:49turned out that they can't really play

4:51chess.

4:52Uh sometimes they they make some right

4:54moves, and they they can to an extent uh

4:57play play nice opening, but they quickly

4:59start to hallucinate.

5:02So,

5:05let's see if we can show the

5:07Yeah, there's a video of Grok went to

5:10Yeah, we see Poison Pawn line with Queen

5:12B6 early on and lost pretty badly, not

5:16necessarily because of the opening, but

5:18because it doesn't really know how to

5:20play chess.

5:22That that was Magnus Carlsen commenting

5:25a LLM chess tournament from our office

5:28in in Oslo.

5:30Um that was a tournament organized by

5:32Kaggle

5:34when they launched their game arena,

5:36which was a

5:37benchmark for benchmarking LLMs

5:40when playing different types of games.

5:42One of them was chess, and now they've

5:43started to to add more games. Also added

5:46werewolf recently, where you can watch

5:49LLMs try to deceive each other in in

5:51social deduction games, which is I can

5:53recommend watching.

5:55But yeah, we see that LLMs often

5:59uh

6:00hallucinate

6:01because obviously they're trained on

6:02language, they're not they can't

6:04calculate.

6:05Uh they can't like high reasoning models

6:07can to an extent calculate through the

6:10reasoning steps where they can actually

6:11play out moves, but they quickly uh fall

6:15apart. Um but there's nothing inherently

6:18wrong about the architecture of the like

6:19the transformer architectures to play

6:21chess. DeepMind has trained a

6:24transformer to instead of predicting the

6:27next token, they predict the evaluation

6:30based on a chess position, where they've

6:32trained it on millions of

6:34chess positions to uh Stockfish

6:36evaluations pair. And that has actually

6:38led the transformer to to play at a

6:42grandmaster level strength. But these

6:44aren't trained on language, so these

6:45can't explain chess. So, how do we

6:48uh

6:49bridge the gap between these old chess

6:52computers that can understand and play

6:54really good chess

6:56between the LLMs that can explain chess?

7:00So, we're going to go through our

7:01pipeline of how our game review explains

7:04chess in our app. When you play a game,

7:07first thing we do is we run Stockfish

7:10through the whole game. Stockfish is the

7:12leading chess engine now

7:15that's like a classical chess engine

7:17that that uh calculates the best move.

7:20So, it's what Stockfish says is

7:21considered to be the solution in a in a

7:23chess position.

7:26We then extract a lot of

7:29uh context in the position cuz we want

7:31to explain not only the best move, we

7:33want to explain the threats, the plans,

7:36um the tactics that could arise in the

7:39position, what you should have played, a

7:41lot of these nuances that is useful when

7:44if you want to learn how to become

7:46better at chess.

7:47Um so, we have a lot of detectors that

7:50tries to figure out all all of this, the

7:52forks, pins, skewers,

7:55uh

7:56positional structural themes.

7:58Doubled pawns, for example, is a

8:00disadvantage, so we need to be aware of

8:02all of those kind of things.

8:04And there's also a new novel chess

8:07engine called Maya,

8:09uh which is behind a research project by

8:12the University of Toronto, uh where

8:15instead of

8:16building a chess engine that is trained

8:18to play the best, they have trained a

8:21chess engine, uh it's a neural network

8:23that predicts the moves that humans

8:25would play in certain positions. So,

8:28given a rating, for example, an online

8:29rating of 1,500, it outputs the

8:32probability distribution over all the

8:34moves in the position.

8:35And by doing this, we we could actually

8:37say that a move is a move is really it's

8:40it's the best move. We know that because

8:42of Stockfish, but you also know it's

8:44really hard to find that move because

8:45the probability of playing it at certain

8:47levels are

8:48are so low.

8:49And all of this information, we feed

8:51that to the LLM, and that the LLMs

8:55for now, the LLM's job is only to

8:57translate this information

8:59uh into English, because we really don't

9:03want it to try to figure out too much on

9:05its own, because it quickly leads to

9:06hallucination. It still does, but we

9:08want everything to be grounded in the

9:10information that we uh give it.

9:12And that could result into a comment

9:14like this. If you play chess or know

9:17about chess, this is a game that I

9:18played. My opponent played F5 here, uh

9:21which is a bad move. So, by using

9:23Stockfish, you could see that you get

9:24like a bad move indicator.

9:26Um

9:27but that's not so useful to just know

9:28it's a bad move. So, we are running our

9:32detectors to figure out that, "Okay, F5

9:34is threatening to trap my queen." You

9:36can also see it draws a a line with

9:38Bishop G5. But it can also say, "While

9:40it threatens to to

9:42to trap the queen,

9:44I can just capture the pawn in the

9:46middle, because that's defense this

9:49square so that my queen can get out of

9:51the situation."

9:53Mm.

9:54So, that's how we get to that situation.

9:58Now I'm going to

9:59explain a bit on how we improve our

10:04our game review using agents. We have a

10:08We have closed the loop from user

10:10feedback to

10:12the public request essentially with

10:14humans in the in the loop, but what

10:17happens when users download the

10:19commentary in your app because you can

10:20download it if you think it's it's bad.

10:23It posts it to Slack, but it also sends

10:25it to Cloud Code channel. Channel is a

10:28new feature in the research preview that

10:30is essentially an MCP server that can

10:33inject events into a running Cloud Code

10:35session. So kind of like Open Claw if

10:37you use that. So you have this

10:39continuously running channel and you can

10:41inject events to it. So and then

10:44Cloud Code

10:46starts working or on the on the

10:48commentary. It gets all the information.

10:50It runs a commentary triage skill that

10:53we created that outlines its process how

10:56it should go about to investigate what's

10:58going what's wrong in the position. It

11:00has some scripts to actually run the

11:03generation so it can modify for example

11:06the prompt. It could change some of the

11:08detectors, create some new detectors and

11:10then it can generate the

11:12commentary again given this new

11:14information and verify its own work.

11:17And then it will also ask questions back

11:19to Slack so that I could be on the bus

11:21and I could get a message from Cloud who

11:22is working on this problem where it will

11:24ask me this seems right and I can guide

11:26it.

11:27And if it if it looks right, I'll just

11:29tell it to submit the PR and I open

11:31GitHub on my mobile.

11:33It works works fine and and I merge it.

11:36I'm going to show how this works

11:40by So we have a running Cloud Code

11:45Cloud channel here. Here is the check

11:47the the Slack channel where the

11:50commentary appears. This is just me

11:52having tested a bunch of time. I'm going

11:54to open up the app on my phone.

11:57Go to

11:58a

11:59comment

12:01and report it as as bad.

12:06Now we see it posts a

12:09a comment. We can see the position. The

12:12commentary that was generated is there.

12:14Now I haven't really looked at the

12:15commentary so it could be it's it's it's

12:17probably it's probably good. But we can

12:20also see that it injects it to the Cloud

12:24channel who invokes the commentary

12:25triage skill and starts working. Now

12:28this is now running on high effort so

12:30this could take a while so I'm thinking

12:31we should just go

12:33to the next slide and then we could get

12:34back to it to see if it is

12:36something is happening.

12:38Fantastic. So we'll come back to that in

12:41a few seconds. So as we built this for

12:44you know end users, we had to really

12:46kind of consider this trade-off between

12:47latency versus quality. So typically

12:49when you finish a chess game, you want

12:51to get the analysis and the results

12:52pretty quick. You want to cycle through

12:54the moves kind of one by one. So we

12:56really couldn't show you like a coach is

12:57thinking screen you know indefinitely

13:00while reasoning tokens are kind of

13:01running in the background as an example.

13:03So we had to get this done which felt

13:05almost instant. In AI world, that's a

13:08few seconds at at best. So we we're

13:10aiming for sub 3 seconds to generate our

13:13coach sort of feedback.

13:15How do we do this? We use Gemini 3

13:17flash. Time to first token is typically

13:19being about a second. End to end latency

13:21on average is about 3 seconds which kind

13:23of meets our criteria. We have

13:25experimented with other reasoning models

13:28and we'll get into that on the on the

13:29next slide. The analysis is not

13:31incorrect so the quality is is

13:33definitely good, but the the challenge

13:35is it's unpredictable as to how long

13:37it's going to take to finish. So we have

13:39a new set of features kind of planned

13:41for a more you know chat with your coach

13:42type experience where we can kind of

13:45expect the user to be more patient and

13:46and wait for a response rather than in

13:49the sort of phase where it needs to be

13:50more instantaneous.

13:52The the last thing about quality is

13:54Osborne and I are both a good chess

13:56players. So we ultimately kind of have

13:58the final say is when we look at the

13:59position, it's actually use how we would

14:02calculate and how we would play and

14:03compare it to the LLM's response. This

14:05allows us to actually evaluate whether

14:07it's doing the right thing or not.

14:10So if we talk about evals in more in

14:12more detail, like I said Gemini flash is

14:14kind of our our benchmark, but we have

14:16multiple chess scenarios. Currently we

14:18have 16 different scenarios that we

14:19created. These are around themes like

14:21tactical patterns, blunders, and sort of

14:24limiting hallucination. So as an

14:26example, you know there might be a

14:27knight fork on the particular chess

14:29position and we're trying to assert that

14:31can be LLM actually understand and

14:33mention this when we run it through with

14:35our sort of context engine. And how do

14:38we do this? We extract scenarios from

14:39real games. We use LLM as a judge, very

14:42powerful sort of technique that to test.

14:44We then run the model

14:46in Open Router. Open Router has come in

14:48handy because new models are being

14:49released you know so fast so frequently,

14:52we just want to be able to quickly swap

14:53in and swap out maybe new version of the

14:56Gemini. We want to check check out the

14:58latest GPT-5 model or one of the Cloud

15:00models. So

15:02we'll then compare and and see sort of

15:03the quality ultimately relying on our

15:06own skill to detect whether this is good

15:08or or bad. And as a sort of final point

15:11on this, we we ran all three models

15:13Gemini, Cloud, and and GPT-5 and you

15:16know typically Gemini flash is about

15:1875%. It still doesn't pass all the the

15:20scenarios we've set it so we're always

15:22kind of seeing if a new model will

15:23actually exceed some of the the tricky

15:25cases we've set up.

15:27Cloud on more thinking gets us to about

15:29just under 60% but the latency is much

15:31longer. GPT-5 mini giving us a smaller

15:34set of model, lower latency or other

15:36slower latency but lower accuracy as

15:38well. So we kind of continuously run

15:40through these to to update.

15:43Last thing on sort of our our learnings

15:44and how this can sort of apply to to

15:46your world sitting in front of us.

15:48Number one, you really important to

15:49separate that sort of data pipeline from

15:52the language generation. LLMs can do a

15:54lot of different things but if you need

15:55you know high latency or quick latency,

15:57it's a good sort of technique. Really

15:59try to close the loop with autonomous

16:01agents. Kind of the flow that Osborne

16:02showed is now very common and very

16:04powerful and really allows you to

16:06iterate quickly. Always try to build a

16:08very a clear sort of context extraction

16:10model. This unfortunately in the

16:12beginning is a very slow sort of painful

16:14process. It's a large you know

16:15ultimately JSON file that you keep sort

16:17of starting big and you start to prune

16:18step by step and you see how quality

16:21improves over time. Automated evals

16:23really do help and I I'm hoping in your

16:25domains you also have a set of you know

16:27SMEs that you can help rely on to

16:28evaluate the output and sometimes that's

16:30not necessarily the person actually

16:32building it. Could be someone else who

16:33is a domain expert. So remember to sort

16:35of partner if needed.

16:38The last thing just on the on a fun sort

16:40of note before we go back to the the

16:41output of the

16:43the sort of coding agent,

16:44we do have some chess sets on the the

16:46third floor at the entrance you may have

16:48seen. We're going to host a chess simul

16:51today in the afternoon around 3:45 p.m.

16:53A chess simul for those who are

16:55unfamiliar is when one person in this

16:57case me or Osborne plays multiple people

17:00at the same time. So we have four chess

17:02sets. We will play four people at the

17:03same time. We'll have a slightly more

17:06time for us cuz we have to walk around

17:07to play multiple boards. If you happen

17:09to play and you happen to win, you will

17:12get one of the wooden chess boards at

17:13the end of the event. They're very nice

17:15high quality chess sets. If no one wins,

17:17we will still determine who the two best

17:19players are and we will still give you a

17:21set. And if everyone wins, we need to

17:23buy more boards. That's that's

17:25Hopefully not everyone wins. There's a

17:26QR code if you want to sign up or you

17:28just stop by 3:45. You're welcome to to

17:31do that. And then yeah. Yeah, close it

17:34off. Let's go back to see if

17:36what has been happening here. Oh, we see

17:38it's still thinking. It is It has

17:40actually added a comment looking into

17:41this now investing in the position.

17:43Quick question. What specifically feels

17:45wrong about the commentary? Yeah, that's

17:46it got me there. It's there's nothing

17:48wrong.

17:49You're absolutely right. Nothing wrong.

17:54So yeah, it's

17:56Now it's going to close this off cuz

17:58yeah, it it worked well.

18:01Fantastic. Well, thank you so much and

18:02happy to take any questions.

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com: free, unlimited, no sign-up.