Full transcript
Spotting subtle bugs in AI generated code
0:00Let's play a game. Have a look at this
0:02Convex code and see if you can spot the
0:03bug.
0:04So, it's a query that takes a task ID,
0:07it grabs the currently authenticated
0:08user, then it grabs the task,
0:11then the project for that task, and the
0:14assignee's name before returning all of
0:16it at the end.
0:17So, have you spotted where the bug is
0:19here?
0:21No? Well, not to worry, neither did I.
0:24See, the bug is quite subtle and is one
0:26of those that can easily slip by when
0:28you're in the vibe coding zone.
0:31You see, I, probably like most of you
0:33guys as well, have been letting the
0:35agents absolutely rip my projects
0:37lately, and I've been quite happily just
0:39merging stuff without doing a deep
0:41evaluation of the code for potential
0:44issues
0:45and stuff was definitely falling through
0:47the cracks.
0:48So, I decided enough was enough. I was
0:50going to try and tackle this problem by
0:52running this this particular code issue
0:55that I just showed you and nine others
0:57through a battery of code review bots to
0:59see which would catch the issue and
1:01which would not.
1:02And I have to say, the results are very
1:05interesting.
1:06I won't spoil it for you just yet, but I
1:08will just say that one tool got almost
1:11everything right, and one famous tool
1:14basically face-planted the [music] whole
1:15thing.
1:16So, if you want to find out what the
1:17answer is to that problem I just showed
1:19you, and you're also like me and you
1:22want to inject a little bit more
1:23confidence into your vibe coding
1:24sessions with Convex code, [music]
1:26then this is the right video for you.
1:29We're going to take a look at some of
1:30the most popular code review bots on the
1:32market and see how they stack up against
1:34each other.
1:35And by the way, if your favorite tool is
1:37missing from this list, stick around as
1:39I'm going to discuss that a little bit
1:40later.
1:41>> [music]
1:41>> So, make sure you grab yourself a lovely
1:43cup of tea, drop me a like and sub, and
1:46let's get into it. All right. So, just
Setting up the React and Convex evaluation project
1:48quickly before we dive into testing the
1:50review bots, I need to quickly explain
1:52how I went about evaluating these bots.
1:55So, I started by creating this little
1:57project here. It's a realistic project
2:00that I hope I can use as a baseline.
2:02It's basically a Trello clone that uses
2:04React, V and Convex Auth for
2:06authentication.
2:07So, after you log in, you can create a
2:09project. Each project can have any
2:11number of members assigned to them. Then
2:13the members create tasks, all which
2:15appear on this Kanban style board.
2:18Each task has a few properties and each
2:20task can also have comments.
2:22There's also an activity log which we
2:24don't display, [music]
2:25but basically it's keeping an audit
2:26trail of any changes behind the scenes.
2:29So, it's quite basic, but it has a lot
2:32of the essentials that you'll find in
2:33any modern SaaS, such as the nested
2:35related data and authentication and
2:38authorization. So, I feel like this is a
2:40good starting point to evaluate from.
2:43So, now I've got my baseline. I came up
Testing bot knowledge of Convex indexes
2:46with 10 different scenarios I wanted to
2:48test the code review bots on and these
2:50basically form our evals.
2:53So, let's just take a quick look at one
2:54of these to get a sense of what I'm
2:55going for.
2:56So, starting with what I thought would
2:57be an easy one [music] to test to see if
2:59the review bot can detect whether we are
3:01falsely using Convex's dot filter when
3:04we should be using dot with index
3:06instead.
3:07Now, this recommendation to use indices
3:10instead of filter is basically plastered
3:12all over the Convex docs, not to mention
3:15it being actually included inside this
3:16project in the dot cursor {slash} rules
3:19file directory.
3:21So, the bot should really have been
3:23>> [music]
3:23>> smart enough to pick this up.
3:25So, I thought this would be an easy test
3:27to check whether the review bot's basic
3:29understanding of Convex and to see
3:32whether they would actually go out and
3:33read the cursor rules file or not.
3:36>> [music]
3:36>> Now, I do realize that it might be a bit
3:38of a a stretch expecting the review bots
3:41to check every single IDE or tool
3:43special folder like this dot cursor one.
3:47And that's why I'm also working on this
3:49PR to make a change to how we inject
3:52these AI files into your Convex project.
3:55So, I'll be doing a video on that once
3:58this PR lands. So, make sure you get
3:59subscribed as I'm kind [music] of
4:01excited for this one.
4:02So, with that said, I was still hoping
4:04that the review bots would be smart
4:06enough to pick up like from one of the
4:08most important AI tools out there,
4:10Cursor.
4:12But, spoiler,
4:13they didn't.
Overview of the AI code review bots tested
4:15Actually, let's just take a quick aside
4:17to have a talk about the different
4:19review bots I chose for this series of
4:21tests.
4:22So, we have Codium, GitHub Copilot,
4:25Cubic, Code Rabbit, Reptile, Code Ant
4:27AI, Sourcery, Macropus, Graphite AI. So,
4:31I signed up for each of these in turn.
4:33Some required a credit card, but most
4:35gave me some sort of full trial, which
4:38is good as this little experiment could
4:39have gone quite expensive without that.
4:42Some of them had really smooth
4:44onboarding processes and a really clear
4:46and easy to understand dashboard, and
4:48others not so much.
4:50Um, we will talk more about that a bit
4:52later when we discuss the winners.
4:54>> [music]
4:54>> Now, just a note here that I
4:56intentionally did not mess with any of
4:58the settings on any of the review bots.
5:00I know some of you are probably going to
5:02complain at me for that, but I wanted a
5:04fresh out-of-the-box comparison, and I'm
5:07not an expert in any of these tools. So,
5:10configuring them all to be exactly right
5:12for this project would have been a bit
5:13of a fiddly affair and likely
5:15error-prone.
5:16And while we're on the topic of
5:18addressing your complaints before they
5:19arise, I know some of you are going to
5:21complain that I'm missing some tool or
5:23other in this list.
5:24>> [music]
5:25>> And
5:26I posted that I was going to be working
5:28on this video the other day, and I got a
5:29bunch of suggestions for other review
5:32bots to try out.
5:33Now, they are really good suggestions,
5:35and I want to look into those, but I
5:37didn't want to crowd out this video with
5:39every bot out there. But, people did
5:42mention some important ones such as
5:43BugBot and some of the new offerings
5:46from Open AI and Anthropic. So, I am
5:49going to do a follow-up video in the
5:51near future to test out those guys as
5:52well, and also to explore what I think
5:55might be the ultimate solution. So, make
5:58sure you get subscribed on here and on X
6:01so you get notified when that one drops.
Evaluating the first pull request
6:04Okay, so now with that out of the way,
6:05let's return back to the PR we've got
6:07set up for these bots to test the dot
6:10filter versus the dot with index. So,
6:13you can see here and generally
6:15throughout all these tests, I tried to
6:17keep the set of changes quite minimal so
6:20that we are usually only testing one
6:22thing.
6:23So, although I try and do this,
6:24sometimes it's quite difficult to have
6:26the bit of work be isolated and yet the
6:29PR still be realistic.
6:31One little trick I employed to get
6:33around that is to have
6:34um the new query here where we are
6:37testing for filter versus index, I
6:38marked it as an internal query and then
6:41added this comment on top which says
6:43that we're going to execute this query
6:45via the Convex dashboard, which is a
6:47common thing that Convex people do. So,
6:50don't worry Mr. Review Bot, where this
6:52isn't actually used anywhere in the code
6:54base, but we're still going to be used
6:56from the Convex dashboard.
6:58I should also just mention that I was
6:59very careful in this entire project and
7:02in the PR descriptions and everything
7:04not to add any files or anything that
7:06would give the game away that we are
7:08evaluating the review bots.
7:11I don't want a Volkswagen
7:13>> [laughter]
7:13>> incident on my hands here.
7:15All right, so finally Michael, stop
Reviewing bot spam and PR summaries
7:17talking. How did the bots do?
7:19All right, so if we just take a look at
7:20the review on GitHub, you'll notice that
7:23there's a lot of spam at the start here.
7:25So, up here in the user description,
7:27this part here is the bit that I added
7:30and all the rest of this stuff is just
7:31fluff that the review bots added.
7:34I personally don't think that this like
7:36PR summary stuff is useful
7:39um
7:39and I'm sure you can probably turn off
7:41in the settings.
7:42Sorcery in particular found to be quite
7:45verbose out of the box. And that is it
7:48adds all these massive sequence diagrams
7:50in every PR, which I personally don't
7:53think is that useful, but whatever.
7:55So, as we continue to scroll down here,
7:57one thing you might be thinking because
7:59I was too is that each of these bots,
8:02are they able to see each other's
8:05comments and reviews?
8:07I was actually really concerned about
8:08this because I didn't want to have to
8:10create an entirely separate repo and
8:13project for each review bot so that they
8:16could be isolated from each other. So,
8:19the approach that I decided to take was
8:20to to keep a very close eye on these and
8:23see if any of them mentioned any other
8:25reviews review bots or previous reviews.
8:29And I'm thankful to report that at no
8:32instance did I ever see any review bot
8:34mention any other other bots' reviews.
8:37I think this might be because I've got
8:39them all set up to start their review
8:41process is immediately as soon as the PR
8:44drops. So, basically they all get
8:46triggered at the same time, grab their
8:47context, and it's at that point that
8:50they then start to do the reviews. So,
8:51they don't have a chance to look at each
8:53other's reviews.
8:54Either way, as we continue to scroll
8:55down here, you can see that there's
8:57heaps and heaps of stuff here. Some have
9:00put stuff in some uh summary comments
9:03and some have done inline code comments.
9:05Some have done reviews. Some have done
9:07checks. Some have not.
9:09And I was finding it kind of annoying
9:12and error prone to go through this
9:13manually. So, I would do what any
Using an AI agent to grade the review bots
9:16engineer does in 2026. I had an agent do
9:19it for me.
9:21So, for each of these tasks I'm setting
9:23the review bot, I first had an agent
9:26write one of these documents where we
9:28detail what it is exactly that we're
9:30going to be testing, what we're
9:32expecting the model to say, what we want
9:34them to and things that we don't want
9:36them to say. [music] So, basically this
9:38is our grading criteria.
9:40And once again, I intentionally do not
9:42commit these files into the repo to make
9:45sure that the review bots can't see this
9:48and understand [music] that we are, you
9:50know, testing them. And I also make sure
9:52that we have a good cover story for each
9:54PR.
9:56Hi, Mike from the future here. Yes,
9:58weather's lovely. I just want to note
10:00that I have actually now committed these
10:03PR um expectation documents
10:06to the main branch. So, if you want to
10:11go ahead and try and run this yourself,
10:13then you will be able to see them on the
10:14main branch or see the detailed scores
10:16or see whatever else.
10:18But these were not included in any of
10:21the branches so that the review bots
10:24didn't see any of this.
10:26Now, back to Mike from the past.
10:28So, for this particular PR, our cover
10:30story is
10:32new convex/reporting.ts module with a
10:34single internal query for use with the
10:36convex dashboard. Get activity for user
10:39returns recent activity for a specific
10:42user within a project. So, what we want
10:44the bots to do is catch the dot filter
10:46and we want them to recommend that we
10:48should use dot with index instead.
10:50What we don't want the bots to say is
10:52anything to do with auth as this is an
10:55internal query and we know that internal
10:58queries don't need auth because they're
11:00going to get executed by the dashboard.
11:02All right. So, finally again, after all
11:04of this, let's take a look at the
11:05results.
Results of the index versus filter test
11:07So, Code Rabbit caught it. Nice. Reptile
11:10caught it. Macroscopy didn't, which is a
11:12shame. It's only reported uh code check.
11:16Cubic caught it. Graphite didn't and
11:19we'll see, spoiler, will continue to
11:21perform badly.
11:23Quodo caught it. Nice. Codon didn't.
11:26They vaguely mentioned indices, which
11:30might have gotten them a pass, but they
11:31also incorrectly flagged auth on the
11:34internal query. Whoops.
11:36Sourcery did a kind of half-assed job by
11:38mentioning that we might need an index
11:41if not already indexed. It should have
11:43checked the schema to see whether we
11:45have an index. It shouldn't say if we
11:47already have an index.
11:49Copilot caught it and added a nice
11:51helpful inline code comment.
11:54And so, I guess the takeaway is that six
11:57out of nine of the bots caught the
11:59issue, which is pretty good.
12:01So, I then also had the agent, um,
12:05Claude Opus, by the way,
12:06uh, do a detailed tool review. Uh, as
12:09some of the bots went the extra mile and
12:11caught some other issues or flagged
12:13other things which earned them bonus
12:15points, and I'll talk about the scoring
12:18process a little bit later.
12:19Yeah, and then this sort of detailed
12:22evaluation here, you might notice lines
12:23about table, uh, validation agrees.
12:27So, just to give myself a little bit
12:29more confidence that this review agent
12:32is doing a good job, I, um, had another
12:36agent run through the entire thing
12:37again, this time using Convex 5. Uh,
12:40sorry, Codex 5.3 in a clean con-
12:43context. So, that's what this is. It's
12:45basically saying, "Yes, I agree with the
12:47initial review."
12:49And Codex mostly agreed. There I think
12:51there was only like one or two places
12:52where it caught a couple of issues that
12:54Opus didn't catch.
12:57I also did do a few smoke tests on a few
12:59of these just to make sure, 100% sure,
13:01give myself more confidence that both of
13:03them are messing up somewhere.
13:06So, the final thing is in each of these
13:08PR expectations documents, we have a
13:10scores table at the bottom here. So, I
The scoring system explained
13:13guess we might as well talk about
13:14scoring now.
13:15So, basically, I gave the bot three
13:18points if it passed the primary
13:20objective and zero points if they
13:22failed. And if they got a mixed result,
13:25like Sourcery did here, they just get
13:27one point.
13:28Then I gave a bonus point if the review
13:31bot picked up extra good stuff.
13:34And this kind of extra good stuff is
13:36more of a judgment kind of call. So, if
13:39they did a generally good job, they
13:41would get one point. However, if they
13:43had lots of false positives, then I kind
13:46of punish them by giving them minus one
13:48because
13:49in my opinion, false positives in review
13:51bots are just really annoying to deal
13:53with and I don't want to see them.
13:55So, at the end of all that, Code Rabbit,
13:58Reptile, Cubic, and Quoddo and Co-pilot
14:01all tied with three points. Codant
14:03failed hard on this one by not catching
14:06the primary issue and also falsely
14:08flagging the auth issue,
14:10which is not great CommaX knowledge at
14:12all there.
14:13But will this trend continue? What do
14:16you think?
Testing for false positives with database queries
14:19Okay, so we talked a little bit about
14:20false positives just now. So, I want to
14:22take a look at this
14:24PR PR01, which is the first one that I
14:27actually tackled, which deals with false
14:28positives.
14:30So, I wanted to set the bots task to see
14:32if I could catch them out. I wanted to
14:35see whether they would falsely flag
14:37something as an issue even though it
14:40isn't an issue just because of the way
14:42that CommaX works.
14:44So, let's take a look at the code. So,
14:45here we have this internal query get
14:47project overview.
14:48I'm doing that internal query trick
14:51again here so I can isolate this bit of
14:52work.
14:53And this query is supposed to take in a
14:56project and return some data that can be
14:59used to give you an overview of that
15:01project.
15:03So, the first thing it does is grab a
15:04bunch of tasks.
15:06Now, I'm not using dot collect here
15:08because I don't want the bots to get
15:10confused and flag unbounded reads. I
15:13have a separate test for that coming up.
15:15So, I only select 50 tasks here.
15:19And
15:20when I make decisions like this, I
15:23explicitly put a comment on the top here
15:25so that the review bot knows I'm making
15:28this explicit decision.
15:30Again, I just don't I want to try and
15:31minimize the number of potential issues,
15:34side issues that get raised in here. I
15:36want the review bot to to focus on the
15:38one thing that I'm mainly testing for.
15:41But anyway, so for each task we then
15:43also grab the assignee and all the
15:46labels before finally returning all of
15:49the data at the end.
15:50Now, if you are a review bot, which is,
15:55you know, an AI model, and you've been
15:57trained on a bunch of traditional
15:59serverless code, you might look at this
16:01and go, "Oh, there's a bunch of database
16:05round trips here. Uh this is not good.
16:07This is what's known as an N+1 problem
16:09and is a pretty standard performance
16:11issue."
16:12But we all know, as Convex developers,
16:15that this is not true in Convex because
16:17in Convex the database and the compute
16:21are collocated together, so effectively
16:23this entire query is one big
16:25transaction, so there is no N+1 issue.
16:29Okay, so now in the markdown file we
16:31basically in this what I'll be testing
16:32section we have where the tools
16:34understand that nested uh context.db
16:37calls inside of a Convex query function
16:40are not an N+1 problem.
16:42So, how did they do?
16:44Well, Code Rabbit, whoops. Reptile,
16:48nope. Macoscope decided not to flag
16:50anything, but again, it didn't flag
16:53anything either before, so I've given it
16:55a pass here, but I have my suspicions.
16:58Uh Cubic saw nothing wrong uh too, which
17:01is good.
17:02Graphite, as with Macoscope, um only did
17:06code checks here, which got a pass, but
17:08again, I'm very suspicious. Uh Quoda
17:11falsely flagged this one, so I
17:14uh same with Coda and same with
17:16Sourcery, and even Co-pilot failed as
17:18well.
17:19>> [music]
17:19>> Wow, so okay, three out of the nine
17:22tools passed.
17:24And as we'll see, I think Macoscope and
17:26Graphite are kind of a little bit suss
17:28here, so I would say that Cubic is
17:30probably the only one that honestly did
17:32a good job here.
17:35Now, I'm going to skip over the details
17:36part, and let's just take a look at the
17:38scores table.
17:39So, Macoscope, Cubic, and Graphite all
17:41got three points.
17:43And a few of the others picked up bonus
17:45points, and Code Ant and Code Rabbit
17:49What's it with these code animal things?
17:51Anyway, they get negative points for
17:54false additional positives.
17:56So, super interesting results. Let's
17:58continue.
17:59All right. Now, I think you'll all agree
18:01that this is pretty interesting.
18:03But, I don't want to bore you guys by
18:05just going through list of all these in
18:07painful details. So,
18:09instead, what I'm going to do is do the
18:10rest of these
18:12tests that I gave the the bots in a few
18:15little mini arcs.
Checking for missing authentication and authorization
18:17All right. So, in the third one, I
18:18wanted to test for missing auth.
18:21So, you can see that this request export
18:23function here doesn't check for
18:25authentication before doing its thing.
18:27So, the results are as follows.
18:32So, in the fourth one, I set up a cron
18:35job to run this mutation clean up old
18:38activity.
18:39And [music]
18:40while there is no issue in the handler
18:42itself, the thing that I was testing the
18:44bot for to see whether they would pick
18:46up that this should be an internal
18:48mutation, not a public mutation.
18:51And here are the results.
18:54And staying on the auth authentication
18:56trip, this is the last one I did, but
18:59it's also the one that I I showed you
19:00right in the intro. So, did you spot the
19:03issue?
19:04I'll give you a few more seconds while
19:06you have a look.
19:09Okay. So, if you said it was an
19:11authorization issue, you would be
19:13correct. So, basically, even though that
19:16we check for authentication here, we
19:19grab the user's ID,
19:21we don't actually at any point in here
19:24check to see whether the past task
19:26actually belongs to this user.
19:29And thus, a malicious user could use
19:32this to read tasks from any other user,
19:35which is obviously not ideal and should
19:38be caught by the bots.
19:40And most did indeed catch it, which is
Performance tests on unbounded queries
19:43nice.
19:44All right, so let's take a look at some
19:45performance issues now. So, in the fifth
19:48one, I wanted to test for one of the big
19:51performance gotchas for agents plus
19:53Convex right now, and that is unbounded
19:56collects.
19:58So, here we can see that even though we
20:00are using an index, we are also using a
20:03dot collect, and that could be a
20:05potential issue, particularly when the
20:07table can grow in an unbounded way.
20:11>> [music]
20:11>> What I mean by unbounded is that um the
20:13number of rows here can be expected to
20:16grow large and larger as time goes on,
20:19which is the case for tasks in a
20:22project. And so, this is how the bots
20:25did.
20:27The sixth task is similar to the
20:29previous one and is there because I see
20:32agents doing this one a lot wrong at the
20:34moment as well.
20:36So, I've seen them using dot collect as
20:38a way to count the number of documents.
20:41Basically, I'm hoping that the bot will
20:43catch this and instead suggest a
20:47denormalized count process or use the
20:51aggregate component instead.
20:53And here's the results.
Schema design and array size limit tests
20:55The seventh test is again another
20:57performance one
20:58and is the first one where I made a
21:01schema change as well.
21:03So, I added this checklist array to the
21:05tasks table.
21:07Then, I added a couple of functions for
21:10adding and toggling a checklist item.
21:14So, the test here is to see whether the
21:15bots would catch the fact that this
21:18checklist array should probably be its
21:20own table.
21:22That's because this checklist array
21:24could grow in unbounded way, which would
21:27cause us to hit the row size limit for
21:30an individual row in a Convex table.
21:33And here's how the bots did.
21:36Okay, the last one in our little mini
21:38arc here is I wanted to test for
21:40correctness, basically. So, in the
Component correctness and state synchronization
21:42eighth test, I wanted to test
21:45the com- the bots' knowledge of Convex
21:47components.
21:48So, as part of the baseline, I added
21:50this aggregate that counts the number of
21:53tasks um in a given project, which is is
21:56all well and good. It's all fine. It's
21:58how you're supposed to do counts.
21:59But, this PR we just added to test the
22:01bot
22:02uh is a new mutation that does something
22:05that will change the number of tasks,
22:08but we intentionally forget to make the
22:11call to update the aggregate as well.
22:14So, hopefully the bot will be smart
22:16enough to catch that.
22:18And let's have a look at the results.
22:19And it seems that uh some weren't
22:21unfortunately smart enough to catch it.
Triggering Optimistic Concurrency Control errors
22:23Now, the ninth test has the most number
22:26of changes in it, and what I was
22:28thinking was going to be the most
22:29complex [music] one for bots to get.
22:32So, optimistic concurrency control or
22:35OCC
22:36errors in Convex are a fairly complex
22:38topic.
22:39They basically happen when you try and
22:42read and write to the same value in lots
22:44of different places all at the same
22:46time. They kind of basically conflict
22:47with each other. [music]
22:49So, in this PR, I created a new table,
22:51platform stats, which has a field total
22:53mutations,
22:55that I write to from every mutation in
22:58the app for every single user.
23:01So, this is obviously a very bad idea.
23:03It's going to cause OCC errors.
23:05>> [music]
23:05>> Now, originally I actually had this
23:07where
23:08the I I
23:10it wasn't for every single user. I
23:11actually had this platform stats for a
23:13single project, and none of the bots
23:16caught it as an OCC error. And I was
23:18like, "Wow, okay. I think maybe maybe I
23:21can make this more obvious because
23:24reflecting back in I realized that it
23:25could be argued that maybe there
23:27wouldn't be that many OCC errors
23:29depending on on how big the project is.
23:31So, I decided to change this this one to
23:33make it more explicit by making it every
23:36single mutation for every single user
23:38hits the same field.
23:39And now the results, yeah, they seemed
23:42much more realistic of what I was
23:44expecting.
Final scores and tier list for all bots
23:45All right, so at last, the moment you've
23:48been waiting for, the final score for
23:50all the models. Here we go.
23:52So, at the top, we would have what I
23:54would class as gold tier. We have Codo,
23:57closely followed by Co-pilot. Then in
24:00silver tier, we have Cubic, Code Rabbit,
24:03and Greptile. Then in bronze tier, we
24:05have Codant. And then down the bottom
24:08here, we have Sourcery, Macroscope, and
24:10Graphite, which I would class in the not
24:13worth your time tier.
24:15Codo being right at number one was
24:18a total surprise to me because I have to
24:21say its dashboard experience was
24:24absolutely the worst out of all of these
24:25tools.
24:27I mean, just look at this. This is all
24:29you get.
24:30It's and like it's just bugs all over
24:32the place. I mean, it says I need to
24:33connect repositories, but I've already
24:36done that several times. Look.
24:39It just refuses to show the fact that
24:40I'm already connected. I was also so
24:43frustrated and confused by this that I
24:45actually tried to exclude Codo from my
24:47testing earlier on. I deleted my entire
24:50Codo account and everything, but to my
24:53great surprise, the tool continued to
24:55submit PRs reviews for me, even though
24:58they kind of been totally deleted. Like,
25:00what?
25:03But, I guess if it's going to continue
25:05to catch issues like this like a total
25:06champ and cost nothing, then
25:11Now, Graphite, I was sure I must be
25:13doing something wrong here because
25:15zero's straight across the board other
25:18than, you know, the false positive one
25:20at the start is just incredible result
25:23in the wrong way.
25:25But, no, I logged into the dashboard,
25:28checked, double-checked, triple-checked
25:29all of this.
25:31And just just to make sure that Graphite
25:33is indeed connected and it indeed finds
25:36no issues.
25:37I just find this really hard to imagine
25:40that
25:41a tool this well-known is also making
25:45this many mistakes.
25:47I guess this leads us into the next
25:48topic.
UX complaints and onboarding friction
25:52So, I know that at least some of the
25:54owners behind some of these bots will be
25:56watching this. So, I just wanted to
25:58point out a couple of places that I
26:00think could be improved. So, firstly,
26:02onboarding experience varied widely
26:04between the tools. Some required that I
26:06hand over the credit card before I could
26:08even try the tool.
26:10Please don't do that. It's just super
26:12frustrating and puts me off right away
26:15when I want when I just want to try out
26:16something.
26:17The absolute worst offender for this was
26:19Akido, which you will notice is missing
26:22from the video as I was originally going
26:24to be the 10th review bot until I
26:27realized it was going to cost me $350
26:30per month just for this review process.
26:32So, yeah, thanks, but no, thanks.
26:35And while we're talking about
26:35onboarding, I really like simple wizard
26:38walkthroughs right at the start where it
26:40basically asked me what it is I'm
26:42looking to do, PR summaries, code
26:44reviews, things like that. It kind of
26:46sets up the basics for you, which eases
26:48me in and lets me know that those other
26:50features are also available.
26:52And while we're talking about
26:53dashboards, man,
26:54just confusing UIs. I'm just Again, I'm
26:57sorry to pick on you, Graphite, but
26:59icons for your main nav bar and no text
27:02is just
27:03Like, how am I supposed to know what
27:04this little down arrow thing is supposed
27:07to represent?
27:08Then, where on earth are the settings
27:10that allow me to configure the review
27:11bot? Like, is it here under settings to
27:15enable code reviews in the setting? No,
27:17no, not there.
27:18Is it this little settings icon here in
27:20the inbox that shows you reviews? No,
27:22no, not there. Ah, okay, it's this
27:25little icon here and then the settings
27:27here. Okay, yeah, got it.
27:29I think.
27:31Now, I know I picked on Graphite here,
27:32but a number of the other bots also
27:34suffer from this kind of like dashboard
27:36nonsense.
27:38And outside of like code review and
27:41dashboard issues, I would say that
27:43I was expecting more of the bots to do
27:46to look at my project more. They should
27:48look at the cursor the rules files. They
27:51may well be looking at the agent.md and
27:53and Claude.md files. I didn't include
27:55those with this project, but I did add
27:57cursor rules, so I was hoping that they
27:59would check that. And if they did, I
28:01would have thought that they would have
28:03done a lot better as a result because
28:04those cursor rules files would have
28:06included much of the knowledge that many
28:09of these review bots seem to have failed
28:11on.
28:12They could also have been a little bit
28:14more intelligent, like if they detected
28:15that this is a Convex project, maybe
28:18they could have
28:19you know,
28:20proactively gone out and find like our
28:22ESLint rules or something else and then
28:25stored that inside its knowledge for
28:28this particular project and used that as
28:31part of its review process. That would
28:32have been a very intelligent way to
28:35have reviews for this project.
Final recommendations and future testing
28:39All right, so that's about all I have
28:41for you for today. I hope you found this
28:43both useful and insightful.
28:45So, I'm am sure about you, but I found
28:47those results quite surprising and so I
28:50think personally for me I will probably
28:53just start to use Copilot for my code
28:56review needs. It's built into GitHub
28:59automatically so I don't need to mess
29:01around with anything else. The billing
29:02is already set up and it just does seem
29:04to do a really good job on Convex code
29:07at least from what I can see.
29:09Now as mentioned earlier I am going to
29:11have to do a follow up as a number of
29:13you mentioned that you would like to see
29:15other tools such as Cursor's BugBot
29:17compared here. So, please do let me know
29:20in the comments down below if there are
29:22other tools you would like me to
29:23evaluate.
29:25I am also very keen to check out
29:27GitHub's agentic workflows. I think this
29:29might be
29:31very well the ideal way to create your
29:33own project tailored review bot.
29:36But for now I think this video is
29:38probably long enough to this. So,
29:40>> [music]
29:41>> until next time thanks for watching.
29:44Cheerio.