Full transcript
0:14Afternoon, everyone. So, our next talk
0:17will be something a little bit
0:19different. We're going to dive into the
0:21world of chess. Quick show of hands, who
0:24has heard of Magnus Carlsen?
0:27Okay, fantastic. No introduction needed,
0:30but widely considered the best chess
0:32player in the world. He also founded a
0:34company called Play Magnus.
0:37Uh this is where myself, Ananth, and my
0:39colleague Asbjørn currently work at. And
0:42we're going to talk to you today about
0:43how we built our AI chess coach that now
0:46you can use and is in production.
0:50So, first up, quick agenda. We'll
0:52quickly discuss a bit more about Play
0:53Magnus,
0:54what it is we actually built, what it is
0:56we actually launched. Uh Asbjørn will
0:58then go into a quick history of chess
1:00and AI. A lot of links there. We'll
1:02briefly touch on why LLMs are actually
1:04bad at chess and how we managed to solve
1:07this problem.
1:08We're then going to deep dive into
1:09actually understanding our game review
1:11and sort of closing the loop with our
1:12autonomous agent. And you'll get a demo.
1:15And then finally, some latency versus
1:17quality trade-offs, as this is a
1:19consumer-focused AI application. And
1:22then lastly, some learnings.
1:25So, first up, what is Play Magnus? In
1:27its simplest form today, it's currently
1:29an iOS and Android application. You can
1:31go on and play your friends, and you can
1:34post about your games.
1:36What's relevant for our particular talk
1:38is that after you play a game, you get
1:40presented with our game review. And this
1:43is powered by our AI pipeline.
1:46So, for example, just showing you how it
1:48works. In this particular position, it's
1:51leading to a checkmate. The last move
1:53that white has played has moved the
1:55knight from this yellow square over here
1:57on F3, captured the pawn on E5.
2:00It is a brilliant move, so automatically
2:02gets the brilliant sort of notation. And
2:04the commentary below is actually
2:06generated by our system. And we're using
2:09an LLM and the pipeline we'll get into
2:11in a second. But what's quite
2:13interesting about it is we're able to
2:15give you the nuance of why it is a
2:17tactic, what detectors from a positional
2:20and tactical sense have fired, what are
2:22the threats you're trying to do, and
2:24actually explain sort of the why behind
2:26the move. So, that's the system we're
2:28going to be talking about today.
2:31Finally, on the last little step of our
2:34of our application, we've started
2:36revealing insights about your play. And
2:38this could be things like how accurate
2:40you played in a particular game phase,
2:42maybe your current rating, or your
2:44current depth in a particular opening.
2:46And these insights form the next layer
2:48of analysis that we present to the
2:50coach, who then gives them back to you
2:52as opportunities for learning and
2:53improving. We hope by using this, you'll
2:56be able to improve and become better at
2:58the game.
3:00All right. So, first a brief history of
3:02chess and AI since they've been
3:03intertwined for so long, just to give
3:05you a little bit of backstory. 1949,
3:08Claude Shannon, the OG Claude, wrote the
3:11the paper programming a computer to play
3:13chess. And here he envisioned that or he
3:16proposed that there there are two types
3:18of of chess engines, type A and type B.
3:22Uh type A were these brute force engines
3:25that search through all possible
3:28possible moves and figure out the best
3:30move. While type B
3:32were those who we know from from 2017
3:35and and onward that can selectively
3:38pick out the the best moves.
3:40Back then, he assumed that we would need
3:43type B computers to to play chess
3:45because
3:46computers were so weak back then, you
3:48couldn't search through the whole whole
3:50tree of of moves. But computers quickly
3:52became better, and people just started
3:55scaling these type A computers.
3:57Uh they got better and better until they
3:59in 1997,
4:01Deep Blue versus Kasparov,
4:03uh the first time a chess engine beat
4:06the best chess player at the time.
4:09Uh so, people didn't really bother about
4:11these type B computers for a while,
4:13these uh intuitive
4:15uh engines, until
4:18DeepMind, shout out to DeepMind, uh
4:21released first AlphaGo, because Go is a
4:24much more complex game than chess. So,
4:26you you can't solve this with these type
4:28A computers. You would need this
4:29intuitive approach, neural network
4:31approach that actually selectively uh
4:33figure out which lines to to calculate.
4:35Uh but after that, they released
4:37AlphaZero, who could play not only Go,
4:39but also chess and and shogi.
4:41Uh
4:42and
4:44some some years later,
4:45uh LLMs came, and people started playing
4:48chess against the LLMs, and quickly
4:49turned out that they can't really play
4:51chess.
4:52Uh sometimes they they make some right
4:54moves, and they they can to an extent uh
4:57play play nice opening, but they quickly
4:59start to hallucinate.
5:02So,
5:05let's see if we can show the
5:07Yeah, there's a video of Grok went to
5:10Yeah, we see Poison Pawn line with Queen
5:12B6 early on and lost pretty badly, not
5:16necessarily because of the opening, but
5:18because it doesn't really know how to
5:20play chess.
5:22That that was Magnus Carlsen commenting
5:25a LLM chess tournament from our office
5:28in in Oslo.
5:30Um that was a tournament organized by
5:32Kaggle
5:34when they launched their game arena,
5:36which was a
5:37benchmark for benchmarking LLMs
5:40when playing different types of games.
5:42One of them was chess, and now they've
5:43started to to add more games. Also added
5:46werewolf recently, where you can watch
5:49LLMs try to deceive each other in in
5:51social deduction games, which is I can
5:53recommend watching.
5:55But yeah, we see that LLMs often
5:59uh
6:00hallucinate
6:01because obviously they're trained on
6:02language, they're not they can't
6:04calculate.
6:05Uh they can't like high reasoning models
6:07can to an extent calculate through the
6:10reasoning steps where they can actually
6:11play out moves, but they quickly uh fall
6:15apart. Um but there's nothing inherently
6:18wrong about the architecture of the like
6:19the transformer architectures to play
6:21chess. DeepMind has trained a
6:24transformer to instead of predicting the
6:27next token, they predict the evaluation
6:30based on a chess position, where they've
6:32trained it on millions of
6:34chess positions to uh Stockfish
6:36evaluations pair. And that has actually
6:38led the transformer to to play at a
6:42grandmaster level strength. But these
6:44aren't trained on language, so these
6:45can't explain chess. So, how do we
6:48uh
6:49bridge the gap between these old chess
6:52computers that can understand and play
6:54really good chess
6:56between the LLMs that can explain chess?
7:00So, we're going to go through our
7:01pipeline of how our game review explains
7:04chess in our app. When you play a game,
7:07first thing we do is we run Stockfish
7:10through the whole game. Stockfish is the
7:12leading chess engine now
7:15that's like a classical chess engine
7:17that that uh calculates the best move.
7:20So, it's what Stockfish says is
7:21considered to be the solution in a in a
7:23chess position.
7:26We then extract a lot of
7:29uh context in the position cuz we want
7:31to explain not only the best move, we
7:33want to explain the threats, the plans,
7:36um the tactics that could arise in the
7:39position, what you should have played, a
7:41lot of these nuances that is useful when
7:44if you want to learn how to become
7:46better at chess.
7:47Um so, we have a lot of detectors that
7:50tries to figure out all all of this, the
7:52forks, pins, skewers,
7:55uh
7:56positional structural themes.
7:58Doubled pawns, for example, is a
8:00disadvantage, so we need to be aware of
8:02all of those kind of things.
8:04And there's also a new novel chess
8:07engine called Maya,
8:09uh which is behind a research project by
8:12the University of Toronto, uh where
8:15instead of
8:16building a chess engine that is trained
8:18to play the best, they have trained a
8:21chess engine, uh it's a neural network
8:23that predicts the moves that humans
8:25would play in certain positions. So,
8:28given a rating, for example, an online
8:29rating of 1,500, it outputs the
8:32probability distribution over all the
8:34moves in the position.
8:35And by doing this, we we could actually
8:37say that a move is a move is really it's
8:40it's the best move. We know that because
8:42of Stockfish, but you also know it's
8:44really hard to find that move because
8:45the probability of playing it at certain
8:47levels are
8:48are so low.
8:49And all of this information, we feed
8:51that to the LLM, and that the LLMs
8:55for now, the LLM's job is only to
8:57translate this information
8:59uh into English, because we really don't
9:03want it to try to figure out too much on
9:05its own, because it quickly leads to
9:06hallucination. It still does, but we
9:08want everything to be grounded in the
9:10information that we uh give it.
9:12And that could result into a comment
9:14like this. If you play chess or know
9:17about chess, this is a game that I
9:18played. My opponent played F5 here, uh
9:21which is a bad move. So, by using
9:23Stockfish, you could see that you get
9:24like a bad move indicator.
9:26Um
9:27but that's not so useful to just know
9:28it's a bad move. So, we are running our
9:32detectors to figure out that, "Okay, F5
9:34is threatening to trap my queen." You
9:36can also see it draws a a line with
9:38Bishop G5. But it can also say, "While
9:40it threatens to to
9:42to trap the queen,
9:44I can just capture the pawn in the
9:46middle, because that's defense this
9:49square so that my queen can get out of
9:51the situation."
9:53Mm.
9:54So, that's how we get to that situation.
9:58Now I'm going to
9:59explain a bit on how we improve our
10:04our game review using agents. We have a
10:08We have closed the loop from user
10:10feedback to
10:12the public request essentially with
10:14humans in the in the loop, but what
10:17happens when users download the
10:19commentary in your app because you can
10:20download it if you think it's it's bad.
10:23It posts it to Slack, but it also sends
10:25it to Cloud Code channel. Channel is a
10:28new feature in the research preview that
10:30is essentially an MCP server that can
10:33inject events into a running Cloud Code
10:35session. So kind of like Open Claw if
10:37you use that. So you have this
10:39continuously running channel and you can
10:41inject events to it. So and then
10:44Cloud Code
10:46starts working or on the on the
10:48commentary. It gets all the information.
10:50It runs a commentary triage skill that
10:53we created that outlines its process how
10:56it should go about to investigate what's
10:58going what's wrong in the position. It
11:00has some scripts to actually run the
11:03generation so it can modify for example
11:06the prompt. It could change some of the
11:08detectors, create some new detectors and
11:10then it can generate the
11:12commentary again given this new
11:14information and verify its own work.
11:17And then it will also ask questions back
11:19to Slack so that I could be on the bus
11:21and I could get a message from Cloud who
11:22is working on this problem where it will
11:24ask me this seems right and I can guide
11:26it.
11:27And if it if it looks right, I'll just
11:29tell it to submit the PR and I open
11:31GitHub on my mobile.
11:33It works works fine and and I merge it.
11:36I'm going to show how this works
11:40by So we have a running Cloud Code
11:45Cloud channel here. Here is the check
11:47the the Slack channel where the
11:50commentary appears. This is just me
11:52having tested a bunch of time. I'm going
11:54to open up the app on my phone.
11:57Go to
11:58a
11:59comment
12:01and report it as as bad.
12:06Now we see it posts a
12:09a comment. We can see the position. The
12:12commentary that was generated is there.
12:14Now I haven't really looked at the
12:15commentary so it could be it's it's it's
12:17probably it's probably good. But we can
12:20also see that it injects it to the Cloud
12:24channel who invokes the commentary
12:25triage skill and starts working. Now
12:28this is now running on high effort so
12:30this could take a while so I'm thinking
12:31we should just go
12:33to the next slide and then we could get
12:34back to it to see if it is
12:36something is happening.
12:38Fantastic. So we'll come back to that in
12:41a few seconds. So as we built this for
12:44you know end users, we had to really
12:46kind of consider this trade-off between
12:47latency versus quality. So typically
12:49when you finish a chess game, you want
12:51to get the analysis and the results
12:52pretty quick. You want to cycle through
12:54the moves kind of one by one. So we
12:56really couldn't show you like a coach is
12:57thinking screen you know indefinitely
13:00while reasoning tokens are kind of
13:01running in the background as an example.
13:03So we had to get this done which felt
13:05almost instant. In AI world, that's a
13:08few seconds at at best. So we we're
13:10aiming for sub 3 seconds to generate our
13:13coach sort of feedback.
13:15How do we do this? We use Gemini 3
13:17flash. Time to first token is typically
13:19being about a second. End to end latency
13:21on average is about 3 seconds which kind
13:23of meets our criteria. We have
13:25experimented with other reasoning models
13:28and we'll get into that on the on the
13:29next slide. The analysis is not
13:31incorrect so the quality is is
13:33definitely good, but the the challenge
13:35is it's unpredictable as to how long
13:37it's going to take to finish. So we have
13:39a new set of features kind of planned
13:41for a more you know chat with your coach
13:42type experience where we can kind of
13:45expect the user to be more patient and
13:46and wait for a response rather than in
13:49the sort of phase where it needs to be
13:50more instantaneous.
13:52The the last thing about quality is
13:54Osborne and I are both a good chess
13:56players. So we ultimately kind of have
13:58the final say is when we look at the
13:59position, it's actually use how we would
14:02calculate and how we would play and
14:03compare it to the LLM's response. This
14:05allows us to actually evaluate whether
14:07it's doing the right thing or not.
14:10So if we talk about evals in more in
14:12more detail, like I said Gemini flash is
14:14kind of our our benchmark, but we have
14:16multiple chess scenarios. Currently we
14:18have 16 different scenarios that we
14:19created. These are around themes like
14:21tactical patterns, blunders, and sort of
14:24limiting hallucination. So as an
14:26example, you know there might be a
14:27knight fork on the particular chess
14:29position and we're trying to assert that
14:31can be LLM actually understand and
14:33mention this when we run it through with
14:35our sort of context engine. And how do
14:38we do this? We extract scenarios from
14:39real games. We use LLM as a judge, very
14:42powerful sort of technique that to test.
14:44We then run the model
14:46in Open Router. Open Router has come in
14:48handy because new models are being
14:49released you know so fast so frequently,
14:52we just want to be able to quickly swap
14:53in and swap out maybe new version of the
14:56Gemini. We want to check check out the
14:58latest GPT-5 model or one of the Cloud
15:00models. So
15:02we'll then compare and and see sort of
15:03the quality ultimately relying on our
15:06own skill to detect whether this is good
15:08or or bad. And as a sort of final point
15:11on this, we we ran all three models
15:13Gemini, Cloud, and and GPT-5 and you
15:16know typically Gemini flash is about
15:1875%. It still doesn't pass all the the
15:20scenarios we've set it so we're always
15:22kind of seeing if a new model will
15:23actually exceed some of the the tricky
15:25cases we've set up.
15:27Cloud on more thinking gets us to about
15:29just under 60% but the latency is much
15:31longer. GPT-5 mini giving us a smaller
15:34set of model, lower latency or other
15:36slower latency but lower accuracy as
15:38well. So we kind of continuously run
15:40through these to to update.
15:43Last thing on sort of our our learnings
15:44and how this can sort of apply to to
15:46your world sitting in front of us.
15:48Number one, you really important to
15:49separate that sort of data pipeline from
15:52the language generation. LLMs can do a
15:54lot of different things but if you need
15:55you know high latency or quick latency,
15:57it's a good sort of technique. Really
15:59try to close the loop with autonomous
16:01agents. Kind of the flow that Osborne
16:02showed is now very common and very
16:04powerful and really allows you to
16:06iterate quickly. Always try to build a
16:08very a clear sort of context extraction
16:10model. This unfortunately in the
16:12beginning is a very slow sort of painful
16:14process. It's a large you know
16:15ultimately JSON file that you keep sort
16:17of starting big and you start to prune
16:18step by step and you see how quality
16:21improves over time. Automated evals
16:23really do help and I I'm hoping in your
16:25domains you also have a set of you know
16:27SMEs that you can help rely on to
16:28evaluate the output and sometimes that's
16:30not necessarily the person actually
16:32building it. Could be someone else who
16:33is a domain expert. So remember to sort
16:35of partner if needed.
16:38The last thing just on the on a fun sort
16:40of note before we go back to the the
16:41output of the
16:43the sort of coding agent,
16:44we do have some chess sets on the the
16:46third floor at the entrance you may have
16:48seen. We're going to host a chess simul
16:51today in the afternoon around 3:45 p.m.
16:53A chess simul for those who are
16:55unfamiliar is when one person in this
16:57case me or Osborne plays multiple people
17:00at the same time. So we have four chess
17:02sets. We will play four people at the
17:03same time. We'll have a slightly more
17:06time for us cuz we have to walk around
17:07to play multiple boards. If you happen
17:09to play and you happen to win, you will
17:12get one of the wooden chess boards at
17:13the end of the event. They're very nice
17:15high quality chess sets. If no one wins,
17:17we will still determine who the two best
17:19players are and we will still give you a
17:21set. And if everyone wins, we need to
17:23buy more boards. That's that's
17:25Hopefully not everyone wins. There's a
17:26QR code if you want to sign up or you
17:28just stop by 3:45. You're welcome to to
17:31do that. And then yeah. Yeah, close it
17:34off. Let's go back to see if
17:36what has been happening here. Oh, we see
17:38it's still thinking. It is It has
17:40actually added a comment looking into
17:41this now investing in the position.
17:43Quick question. What specifically feels
17:45wrong about the commentary? Yeah, that's
17:46it got me there. It's there's nothing
17:48wrong.
17:49You're absolutely right. Nothing wrong.
17:54So yeah, it's
17:56Now it's going to close this off cuz
17:58yeah, it it worked well.
18:01Fantastic. Well, thank you so much and
18:02happy to take any questions.