Free YouTube Transcribe

Video transcript

Lauren Tan @poteto is an engineer at Cursor. Previously, she worked on the React Compiler at Meta...

UninformedInvestors · 10,840 words · 50 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Everyone, I'm Lauren, Lauren Tan. I

0:03guess not many people know my last name.

0:05Uh I am potato on Twitter.

0:08Uh

0:09potato with spelled with an E.

0:11Um and I have been at Cursor for about 5

0:16months.

0:17Uh previously I was at Meta where I

0:19worked on the React team, uh

0:22specifically working on the React

0:23compiler, uh which was a whole lot of

0:25fun. Uh I'm still on the on the core

0:28team and and uh contributing to open

0:30source here and there. Uh

0:32so that that's really nice that they

0:34still let me do that. Uh and before Meta

0:36I was at Netflix, uh where I was uh both

0:40a tech lead and uh I transitioned to be

0:44be an engineering manager

0:45um

0:46for about 2 years.

0:49So I've had a I've had a lot of

0:51experience going between engineering

0:54management and being an individual

0:55contributor. Uh

0:57and I think something I've noticed

0:59actually, which is quite interesting, is

1:01that there are so many parallels with

1:03you know, management skills and how to

1:05like manage agents.

1:07Uh and that's actually a a big part

1:09about what I wanted to chat with you and

1:11everybody else about today.

1:13Um but yeah, that's that's me. Uh I do

1:16have some like very light slides, but uh

1:20it's not going to be um just rambling.

1:22So let me just share my screen.

1:25And hope that I don't leak anything.

1:28Uh

1:33Oh no, I need to allow permissions.

1:36>> No worries. All good, take your time.

1:38>> There's always tech tech tech trouble.

1:40Uh give me 1 second to rejoin.

1:42>> Yeah, go for it.

1:46I see many of you already know Lauren

1:48from the looks of the chat here. Um so

1:51yeah, it's exciting to to get a chance

1:53to chat with her uh and and go through

1:55some of her her recent work. As you guys

1:57heard, you know, a lot of recent

1:58experience from from Netflix to to Meta

2:02and then now over at Cursor.

2:04Uh we're going to chat a little bit

2:05about Grok Bot as well. So, uh that'll

2:08be exciting. I don't know if you guys

2:09saw that was a recent release. I think

2:10literally maybe yesterday or the day

2:12before

2:13uh from the Cursor team, which is kind

2:14of like um let's call it like agents for

2:17everyone. You can go check it out if you

2:18want and learn a little bit more about

2:20the product. But um but yeah, we'll

2:22we'll explore that a little bit today as

2:23well.

2:26All righty, welcome back.

2:40And Lauren, you're just on mute there if

2:42uh you want to hop off and mute and

2:43we'll start chatting.

2:44>> Yeah, sorry.

2:45>> No worries.

2:47>> It's 2026 and I still don't know how to

2:48use Zoom.

2:49>> [laughter]

2:50>> That's all good.

2:51>> Uh okay, so I assume you can see my

2:53screen.

2:54>> Yes, yeah.

2:56>> So yeah, today

2:57yeah, I think I think the big theme for

2:59me as I've been using agents to write

3:01code, and I'm sure a lot of you have had

3:03the same experience as well, is

3:06how do you trust it? You know,

3:07especially if you are an engineer that's

3:09been

3:10writing code for a very long time, you

3:12have

3:13a lot of opinions and lessons that

3:16you've learned about doing good

3:17engineering.

3:19And when you see agents just, you know,

3:21winging it and, you know, guessing,

3:23hallucinating, uh you know, confidently

3:26stating that they found the smoking gun

3:29uh for the hundredth time, uh but it's

3:31actually not the real problem, you lose

3:33a lot of trust. And when you lose when

3:36you don't have much trust in your

3:37agents,

3:39I feel like you you really can't get the

3:40most out of them.

3:42And for me, the parallel is like with

3:44management.

3:45Uh so, if I'm an a manager, an

3:47engineering manager of a team,

3:49and I have a bunch of, you know, I have

3:51a team of engineers

3:53uh on my team and I don't trust them,

3:56then the mode of operation I'm going to

3:58be in is going to be like

4:00micromanagement, right? I'll have to

4:02spend a lot of time looking over my

4:04reports' shoulders and checking that

4:07they're doing their work well, you know,

4:08that they're not shipping bugs to

4:10production.

4:13And so, I drew this chart because

4:15uh it's not it's not a very scientific

4:17chart, but

4:18like this is how I imagine

4:21myself and my journey through using

4:23agents.

4:24So, you know, like fast forward or back

4:27forward

4:28uh

4:29or fast back

4:31uh fast backwards like a year or so

4:34when, you know, nobody was or not many

4:36people were using agents to code. Uh

4:38I think you

4:40uh you get into this mode where you are

4:45in very heavily in the loop with one or

4:47several like a handful of agents.

4:50And you find yourself just constantly

4:53uh you know, trying to understand what

4:54your agents are doing

4:56uh and you're very very in the loop.

4:58You're watching every single output. You

5:00are sitting there prompting

5:02um and you really can't parallelize

5:05beyond that because you don't again you

5:07don't have that trust, right? You can't

5:09go to 100 agents

5:11uh like spawn 100 agents when you don't

5:13even trust the output of one agent.

5:17So, over the past 5 months I feel like

5:19I've really been able to uh like ascend

5:23this

5:24trust curve and now I'm at the point

5:27where

5:28uh I actually have This sounds kind of

5:30scary to say this and it it makes me

5:31sound like a slot artist, but I I

5:33promise I'm not.

5:35But I actually have my agents now um

5:37auto merging PRs for me.

5:39Uh which is like a wild thing to say,

5:42but um like I woke up today and there

5:44were like 20 PRs landed and I just

5:46reviewed them on main like they were

5:49already landed.

5:50And they were good.

5:51Uh so how did I get to that point? It's

5:54basically what I wanted to talk about

5:56today.

5:58Uh and again like yeah, feel free to

6:00jump in if you have questions, Colin. Um

6:03but uh

6:05Oh yeah, of course I got to show this

6:07this chart.

6:08Uh where uh

6:12Oh do not trust to read Someone

6:15requested to control my computer. Uh

6:18probably won't do that. Uh but yeah, so

6:21this chart I think

6:22I I I'm sharing this chart not to kind

6:24of like flex but to kind of show like

6:27the journey. Like so you can see like

6:29the curve like it sort of like inversely

6:31matches the contributions I've been able

6:34to land at cursor.

6:36So I joined 5 months ago. And 5 months

6:38ago like I you know, my first month I

6:40was like not very productive cuz I was

6:43you know, I was learning the code base.

6:45Didn't know what the heck was going on.

6:47And as I got more confident in in my

6:50agents uh I've really been able to kind

6:52of ramp up my productivity.

6:55Uh and again like yeah, like last month

6:57I shipped 1,000 PRs which is ridiculous.

7:01Uh and then this month we're only on the

7:0412th, I'm already at like almost 800 PRs

7:07landed.

7:08Uh so the velocity is definitely high

7:12and you you I'm I'm sure a lot of you

7:14will definitely be questioning like how

7:16how much of this code is actually good.

7:18Um and I think yeah, like that's

7:20definitely fair to question.

7:23Um but uh yeah, I think

7:26I think if you

7:27set up your agents well, you can

7:30definitely get to a very similar level.

7:34Um and so I'm going to talk about how we

7:35do that.

7:38Uh so for me I think I'm curious like I

7:42guess calling your experience as well,

7:44but uh for me

7:46I think the most important skill that

7:49you should have in your toolbox when you

7:52work with agents is verification.

7:55Uh and by verification I mean the

7:57ability for an agent to actually

8:00run the code

8:01uh or take CPU traces or heap snapshots

8:06or

8:07uh

8:07you know, open an iOS simulator,

8:10whatever you know, however your

8:11application is exposed to your users

8:15it can do the same thing

8:17and uh run it for real and actually test

8:20and verify it'll work.

8:22Because that's the thing that really

8:24closes the loop. Uh it doesn't guarantee

8:26your agent writes good code,

8:29uh but it allows them to at least write

8:31correct code. Uh which is a big a really

8:34big step forward for being able to trust

8:37your agent.

8:38Um

8:40I will I can share one example that we

8:43have uh within

8:45cursor.

8:47Uh

8:48oops

8:50Where let me open this.

8:59Let me just make this bigger.

9:01Uh let me full screen.

9:03There you go.

9:04Uh

9:05so for the for cursors agent window, uh

9:09so this is actually an interesting

9:11story, but uh when I joined

9:14cursor 5 months ago

9:15uh they're actually uh

9:18well, I was supposed to join a different

9:19team. I was supposed to join like the

9:21cloud agents team. Uh but then since I

9:23have a lot of experience working on

9:25React and agents window is a React

9:28application uh I was

9:31uh

9:32I was asked to basically help out with

9:34uh uh, the

9:36agent window work.

9:37Uh, but,

9:39um,

9:40there wasn't really a lot of like skills

9:43to help me. So, I just found myself

9:45like, okay, uh, agent window is going to

9:46launch in like a week. All right, we

9:48have a really tight deadline.

9:50And,

9:51um,

9:52there was, uh, you know, I was just

9:54sitting there like, okay, I'm going to

9:55open up the performance the the Chrome

9:57DevTools and just like take a trace,

9:59look at it myself, and try to make sense

10:01of this flame graph. And keep in mind, I

10:03was just like in my first week. So, I

10:05had no idea what I was looking at. No

10:07idea what, you know, I mean, I had some

10:09idea, but, you know, the the code base

10:11was completely fresh to me.

10:13Uh, and I realized like my agent had no

10:16idea either, you know, like I would take

10:18a screenshot of the trial download

10:19trace, I'll send it to him and it'd be

10:20like, yeah, it kind of looks like this,

10:23you know,

10:24uh, and it would like confidently state

10:25like it's this thing, and then I try to

10:27fix that, and it turns out that's not

10:29the actual thing.

10:31So, this was very very slow process. And

10:34if you've ever done any like performance

10:36work yourself, or, you know, just even

10:38development with an agent where you

10:39don't have a verification skill,

10:42you are the verifier. All right, you

10:44you're the bottleneck. You you you tell

10:46your agent to do something, and then it

10:48goes off and writes some code, then you

10:50open up your, you know, local dev build,

10:52and then you start to say, oh, you know,

10:53that doesn't work. Then you got to copy

10:55paste screen, uh, you know, screenshots

10:57or console errors or whatever,

11:00uh, and then your agent like slowly kind

11:02of like uh, you know, works with that,

11:05and then tries to understand it, and,

11:07um,

11:08uh, fix the thing. But then you're

11:10constantly just in the loop and and

11:12being a bottleneck. So, there's really

11:13no way to parallelize.

11:15So, the control glass goes like one of

11:17the first skills I built, uh, for

11:19Cursor.

11:20Uh, and glass, by the way, is the code

11:22name for agent window that we use

11:24internally, but it's just

11:27Cursor, I guess.

11:28Um, and so Well, skill, uh, is I guess

11:32that that the code itself is not super

11:34interesting. Your agent can very easily

11:36make one for you. Uh where if if you're

11:39building an Electron app or a web app or

11:41even iOS uh applications, uh you can

11:45teach your agent how to use like the

11:47Chrome DevTools protocol or through

11:51uh Apple has some utilities as well for

11:54running the simulator and taking traces

11:56and controlling programmatic control

11:58as well.

12:00Uh so, that's really useful.

12:03Uh but one thing I want actually want to

12:04talk about is

12:06uh the

12:08this thing.

12:10Uh

12:11where is the read me?

12:13Uh so, this skill comes with this very

12:15unique feature called or not feature yet

12:18uh unique file called a feature map.

12:21And so, the story then is like I built

12:24this skill and so now the agent was able

12:26to

12:27uh actually run the agent window and

12:30take traces and whatnot.

12:32Uh

12:33but it had no idea what

12:35what the agent's window was. So, um you

12:38know, like someone say like oh, the the

12:41left side bar is like laggy or something

12:44like that or you know, the right side

12:46the the PR tab is not working. And the

12:49agent would just be like kind of

12:50flailing around. It would spend a lot of

12:51time trying to like look up the code and

12:53you know, where is this feature? How do

12:55I actually get to it on the UI? Which

12:57made it basically completely useless. Uh

13:00you know, like we would I would run

13:03this skill locally

13:04and you know, it would spawn a dev

13:06build.

13:07Uh but then it just be churning. Like it

13:10would just try to click here. It

13:11wouldn't know how to get to things.

13:14Um and it was just an awful experience.

13:16Uh so, who was putting arrows on my

13:19screen?

13:20Um

13:21so,

13:22uh

13:23yeah, this this feature map has has

13:25really useful

13:26uh because it teaches the agent how to

13:28get to all of the features that you

13:30have.

13:31Um and in P-SAC, the plugin that I I've

13:34made,

13:35uh if you

13:36search for P-SAC cursor on Google,

13:39you'll find it.

13:40Uh but there is a creative verification

13:43skill in that plugin where it actually

13:45helps you set up something like this for

13:47yourself.

13:48Um including the feature map, so it will

13:50actually explore the code and build up

13:52this initial feature map that tells your

13:54agent how to get to all of the different

13:57features that you have.

13:58Uh and this is extremely powerful

14:00because now that you have these user

14:02reports that come in,

14:04uh you you can actually map even like a

14:07vague report or even a screenshot.

14:10So we have this

14:11uh internally at Cursor where uh we have

14:14a Slack channel with, you know, lots of

14:17people giving us feedback on the agent's

14:19window and rock bar and whatnot.

14:21Uh and oftentimes the report is very

14:24bad. Like very low quality, like someone

14:27just put

14:28very often we get like a screenshot like

14:30and then someone just says question mark

14:31question mark question mark, like what

14:32is this?

14:33And you know, like the without this,

14:35your agent says, "I have no clue."

14:38Right? But with a feature map like this,

14:40it has a lot more context and

14:43understanding of how to actually

14:45navigate, how to get to all of the

14:47different features. Uh so like, you

14:49know, example like, I guess like the

14:51sidebar, like what is the sidebar?

14:53Uh you know, like all the different sub

14:55features that are present in it. Um like

14:58from the user point of view, here's

15:00where to how to get to it, all the

15:02different keyboard shortcuts.

15:05Uh even like the

15:06the what do you call it? The DOM

15:08elements or yeah, like the attributes

15:11that you use for selecting things

15:13through the CDP

15:15uh are all there. So uh again, you know,

15:18this is like really really powerful

15:20uh for for agents. And uh and a piece

15:24that ships uh that create verification

15:26skill but also a maintain verification

15:28skill.

15:29Uh so you can keep this up to date.

15:32>> Cool. Yeah, I was just going to ask how

15:34you created that. So do you mind sharing

15:35a little bit more about um

15:37that that process in the context of P

15:38stack and maybe just what P stack is for

15:40the folks who aren't familiar?

15:42>> Yeah, so P stack is pretty interesting

15:44because uh

15:46well, first of all, the name is kind of

15:47goofy. Like the P the P in P stack is

15:51like potato potato sack.

15:53Uh because I um

15:55So uh the uh there is a pretty

15:59uh famous person Garry Tan who is the

16:02CEO of Y Combinator and he's come up

16:04with this plugin called G stack. Uh

16:07Garry stack and uh

16:10funnily enough, we share the last name.

16:12We have no relation.

16:13Uh but I thought it'd be funny to kind

16:16of, you know, poke fun at Garry and make

16:18P stack my version of of of

16:21>> [laughter]

16:21>> of his plugin.

16:22Uh but kind of just tailor it to my own

16:25set of prac engineering practices.

16:29Uh but I honestly actually never set out

16:30to build P stack. Uh it just started

16:33with a bunch of skills. All right, like

16:34I started with that control glass skill.

16:37And then I started with another skill

16:38like called howl which I also noticed

16:41through like observing agents.

16:44Um so like, you know, in the early days

16:46of me, you know, trying to climb this

16:48ladder, I was like super in the loop.

16:50And I was basically nitpicking my agents

16:52to an extreme degree. I was uh like I

16:55would tell it um

16:57you know, this feature has stopped

16:59working. Here's a bug report. Like why

17:01isn't it working?

17:04And very often the agent would just like

17:06confidently state like, "Oh, it has to

17:08be this, right? It has to be this

17:10thing."

17:11And I noticed like when I looked at the

17:13actual tool calls, I noticed it wasn't

17:15actually reading the code

17:17that I thought should be affected.

17:20And that made me just extremely

17:21suspicious. And at that point I was just

17:23like, I'm not going to I can't trust any

17:24this agent anymore because it's just

17:26it's just completely hallucinating.

17:29And I think

17:31I think it's very easy to just, you

17:33know, like build up that distrust and

17:35not and kind of feel

17:37helpless. Like, you know, you don't know

17:39how to help your agents succeed.

17:41But like again, I think the the the

17:43management analogies super helpful

17:45because like imagine if you were a

17:46manager of an engineering team and you

17:49had an engineer in your team who was a

17:51really good coder, no business context

17:54whatsoever, you know, they they just you

17:56just hired them and they they onboarded,

17:58you know, like uh 5 seconds ago.

18:00Uh and so how do you actually teach that

18:03person to be effective?

18:05So how you do that is through skill. Uh

18:08skill being just, you know, it's just

18:09markdown, right? But you know, it it

18:11codes a lot of information,

18:13instructions, a lot of

18:15uh you can really draw out a lot of

18:19intelligence from an agent by well, some

18:22people on Twitter call it like, you

18:23know, pull the agent to a different

18:25latent space. Which is kind of like a

18:27fancy way of just saying, like since uh

18:29you know, LLMs are sort of like they

18:31predict the next token.

18:33Uh when you give it some high-quality

18:36tokens uh to begin with, then, you know,

18:38it it can kind of pattern match on like

18:40a higher space that's, you know,

18:42smarter.

18:44Um so that's like a very interesting

18:48model there. But yeah, I built Pstack

18:50very, very incrementally. Uh so uh

18:53started with just really observing how

18:55agents, you know, all the failed

18:57different failure modes of of that

18:59agents were having. And every time I saw

19:01that, I just, okay, I'm just going to

19:02make that a skill. All right, like stop

19:04hallucinating, actually go and search up

19:07look up the code, use a lot of

19:09subagents, uh and yeah, stop guessing.

19:16>> Yeah, that makes sense. One one kind of

19:18follow-up question here, both for myself

19:19and for a bunch of people in the chat.

19:20So,

19:21I guess it's two two parts. So, one is

19:23like, how do you maintain these skills?

19:25So, like the product changes over time.

19:27Obviously, there's a lot of people who

19:27are shipping against the code base.

19:29So, how do these skills get maintained?

19:32And then second to that is like, how do

19:33you know when your verification is is

19:35good enough? Like in you know, you can

19:38trust that the very verification loops

19:40that you've built are going to

19:41I guess you trust that the outputs when

19:43they're done.

19:47>> Uh, yeah, maybe I'll talk about um,

19:50I think I have something really down.

19:51Maybe I'll start with this one first.

19:53So, like how do I

19:55maintain these skills?

19:58So, um,

20:00if you're not familiar with this

20:01concept, an eval [clears throat] is

20:02essentially like a way to

20:05Uh, well, I think the mental model I

20:07have is like it's like a unit test for

20:08an agent.

20:09Um, and

20:11you can actually make your own evals.

20:14You don't need like a special framework

20:15for them. You can build you can you can

20:18build one depending on like, you know,

20:20how scientific and how rigorous you want

20:22to be.

20:23Uh, my screen is red.

20:26>> Yeah, there's a little button. Um,

20:28sorry.

20:29>> Are you going to like disabling the

20:30drawing or something? Like I can't see

20:32my screen.

20:33>> Yeah, sorry. If you guys could not draw

20:35on the screen, that'd be great. But, um,

20:36there's a little button in the

20:38>> troll.

20:39>> Yeah, the the little drop down.

20:41>> Um,

20:42how do I clear?

20:44>> Yeah.

20:45>> Okay, yeah.

20:46>> You got it. Perfect.

20:47>> Um,

20:47>> Continue.

20:49>> Yeah, so evals are a way to unit test

20:52your skills, basically.

20:55And actually in P stack, we ship

20:57under potato mode, there's a playbook.

20:59If you search for it, called eval

21:01playbook.

21:02Um, and it's

21:03uh,

21:05Uh, it's like not It's actually pretty

21:07pretty rigorous the way it's done. Uh,

21:10But essentially what I do is I spawn a

21:13lot of different sub agents. I have like

21:15my main coordinator agent

21:18come up with a rubric for

21:21what I want the skill to do. Um

21:25and then it spawns all these sub agents

21:26and it it creates individual directories

21:29for them

21:31which are cleverly named to not let the

21:34sub agent know that it's being evaluated

21:37because agents can actually tell and

21:40when they do they change their behavior.

21:42Uh but it does a bunch of stuff like

21:44that to

21:46um

21:46essentially yeah like test whether or

21:49not the skill I'm making or changing is

21:52actually doing what I think it does. Um

21:55and one of the really nice things about

21:56cursor is that we are we have we support

21:59so many different models. So you can

22:01actually eval your skill across all

22:03sorts of different models. Um and you

22:06know get a sense of how well it performs

22:08across that different matrix. Um

22:12especially for the models that you use.

22:15Uh so I do this a lot. Every time I I

22:17modify a skill I will run one of these

22:20like the eval playbook

22:22and make sure that you know it's

22:24actually leading to a result I want. Uh

22:27but I will say like

22:29maintaining skills is actually pretty

22:31hard. It requires I think a lot of

22:34taste and observation. So you kind of

22:37need to be very good at being a backseat

22:40driver. You know what I mean? Like if

22:42you do pair if you've ever done pair

22:44programming for example

22:46and you watch a coworker code and you

22:48just like you could probably do this

22:49better or you know you could do you know

22:51do you like why did you not do this?

22:52Right? You you ask a lot of questions to

22:54your coworker and it's kind of a similar

22:56thing here. You like you don't want to

22:57just be a passive observer agent. You

23:00want to be very

23:01in the driver seat in the initial stages

23:03when you're building up your own set of

23:05skills.

23:06Uh you know, obviously you can use

23:07something like P set, but if you're

23:09building your own set of skills, it's

23:11very I think you know, opening up the

23:13all the tool calls and like reading the

23:15code and

23:17reading all the uh

23:18the agent behavior and their thinking

23:20blocks is uh a really great way to see

23:23where they they fail, right? Like what

23:26what

23:27you know, where are they being done? And

23:29then you can go and build the skill for

23:30that.

23:31And then with verification, how you

23:33trust it is it's I think it's also a

23:35very similar iteration loop uh where you

23:38know, like I actually did the same

23:40process for verifying the verification

23:42skill where I actually get um

23:47So, one thing that's interesting about

23:48evals is that you can sort of hill climb

23:50them, meaning that uh

23:53your eval can produce a score, right? Uh

23:55a score that you can get your

23:57coordinator to produce, uh but also uh

24:00you can have a judge agent of a

24:02different model to uh kind of

24:05cross-reference and make sure that the

24:08first model is not being biased, right?

24:10The model that's judging all of the sub

24:12agents that are running the thing.

24:14Uh but you can also like hill climb. So,

24:16meaning that you can you can use like

24:18{slash} loop in cursor

24:20and you can say, "Okay, keep looping on

24:22this eval, right? Until everything is 10

24:25out of 10." As an example. Uh and I did

24:28the same the basically the same approach

24:30with the control skill. And so, I kind

24:31of it was very it was very hands-off

24:33actually.

24:34Uh so, you know, I uh I kind of built I

24:37built that skill that way, like the CLI

24:40in that skill. Um and over time

24:43it's gotten really good. Uh but yeah, it

24:45was definitely not super smooth

24:48at the beginning. It required a lot of

24:50iteration. And I think there's an

24:53analogy here for me, which is um

24:56Well, I make this analogy later in a

24:57different slide on my drawing here,

25:00uh, but I think of it like

25:03uh, you know, as a as a

25:05engineer now, you're sort of more like

25:08you

25:09like maybe a manager or the analogy I

25:11like is like you're like a a chef in a

25:14restaurant. Uh, you you're the head

25:16chef. Uh, you're not cooking all the

25:18food yourself anymore. You have a team

25:20of cooks, right? You have a line cooks,

25:22you have a sous chef, you have you know,

25:23all these different stations.

25:26Um, and it's your job to really design

25:28the environment. You know, you you

25:30you're in charge of setting up the

25:31kitchen. You're in charge of, you know,

25:34like giving tasks to different people.

25:37So,

25:39um, yeah, it's a very interesting way of

25:42working.

25:43Uh, but yeah, that's that's how I

25:45basically built uh, these verification

25:47skills.

25:48>> Yeah, just just one follow up there and

25:50like to go try to go one layer deeper.

25:52So, are you, let's say we wanted to

25:54build um,

25:56uh, an an eval or a skill for for

25:58something

25:59and we wanted to kind of get better on

26:01its own, which is is what what I think

26:03you're suggesting. Uh, are you doing

26:04that in like a work tree, a kind of

26:07isolated with like the sub agents and

26:09and then the reviewer agents and and all

26:11that? Is it happening like in some type

26:12of cloud hosted environment? Like what's

26:15the the more the practical steps? If I

26:16wanted to go do this, uh, and like set

26:18up a verification system for something,

26:20what would I what would I do or where

26:21would I start?

26:24Uh, I think that uh, the best place to

26:27start is local because you can observe.

26:31You can definitely observe what your

26:32agents are doing. So,

26:33uh, if you're building a verification

26:35skill for yourself,

26:37uh, I would definitely start local and

26:39just have your agent bring up the

26:40application, whether it's like a CLI or

26:44uh, desktop app or whatever.

26:45And so, you can actually observe, right?

26:47You can see how the agent is interacting

26:49with the

26:51the application. You can see it, you

26:53know, how it calls like the different

26:56APIs that that allow it to interact with

26:58the

26:59uh the application.

27:02Um but uh for me personally, uh I have

27:06basically been kind of all in mostly all

27:08in on cloud agents because they're

27:10extremely powerful.

27:12Uh and the really powerful thing about

27:14Cursor is the the cloud agents actually.

27:17Where if you spend a little bit of time

27:18setting up your environment,

27:21these control skills, these verification

27:23skills pay a huge amount of dividend

27:26because it's not just something that

27:28makes you as a single engineer better,

27:31it actually levels up your whole team.

27:33Uh and even your whole company because

27:36uh you can actually start thinking about

27:38cloud agents being that thing about

27:39automations that automatically do things

27:42like

27:44uh I'll buy I get I I cannot talk about

27:46this a bit later, but I'll just kind of

27:49get into it. Uh where where, you know,

27:51for example, like I talk a lot about

27:53this agent we have called Benny, right?

27:55Who

27:56who uh you know, takes all of the bug

27:59reports that we get and it automatically

28:02goes off in the cloud, opens up a cloud

28:04uh it's, you know, it's desktop. It runs

28:07Cursor in its own computer

28:09and it uses the same control skills to

28:12interact with the application and try to

28:14reproduce the bug

28:15uh or the user report. All right, and

28:17this is so so powerful because at once I

28:20can immediately I I get so much

28:22information from this automatically.

28:24Like here in this example, you can see

28:25that uh the Benny actually reproduced

28:28the bug,

28:29uh but it's already fixed on main.

28:33So, it actually confirms that we fixed

28:35this problem already. And all I need to

28:37do is just release another build of of

28:39Cursor.

28:40Uh so, that's like huge information

28:42there that I didn't have to go off and

28:44sit with an agent, you know, and spend

28:46an hour trying to figure out, like, is

28:47this fixed? Is this not fixed?

28:49So, you you you gain back so much time,

28:52uh but, you know, everybody on my team

28:54benefits from this. Everybody in the

28:55company benefits from this.

28:57Uh so,

28:59definitely think that, uh you know,

29:01keeping these uh using cloud agents is

29:04super powerful.

29:05Uh but, yeah, it's like a journey. You

29:07have to trust it first, right? Before

29:09you you get to this point. And that's it

29:12goes back to what I was saying here,

29:13where, you know, it's very hard it's

29:15it's almost impossible, and I would

29:17definitely encourage you not to try to

29:19jump from,

29:20you know, like, if you're still in this

29:22zone, you don't want to jump to, like,

29:25I'm going to spawn a hundred a thousand

29:27or a thousand of cloud agents right now,

29:29because you're just going to waste a lot

29:31of tokens,

29:32um and it's going to be extremely

29:33expensive.

29:35>> Yeah, so just to kind of recap so far,

29:37basically the if we wanted to go on the

29:39journey that you've kind of gone on, it

29:40would be just start with verification,

29:43uh building some some skills and some

29:45some ways of determining that the agents

29:47are producing

29:48at least like correct code, whether like

29:50you said, whether it's good code or not

29:51is maybe a separate question, but like

29:52it's it's technically solving the

29:54problem by looking at, you know, stack

29:56traces, looking at, you know, the the

29:58actual behavior in the app, and so on.

30:00Um and then once we trust it locally,

30:02then we can start to think about scaling

30:03into the cloud and running more agents

30:06that are picking up signals, I guess, on

30:08their own, right? So, whether that's

30:09like a bug report that comes in or

30:10something, they can go and pick it up

30:11and solve the problem and and give us

30:14back a PR. And then maybe the last step

30:16is like auto-merging the PRs, which uh

30:18is where you're at.

30:20>> Yeah, yeah.

30:21>> Uh and then reviewing them on main, but

30:23um

30:24is that is that about right?

30:26>> Yeah, exactly. I think yeah, that's why

30:27I drew this this uh this this curve,

30:29right? Because that this this basically

30:31describes my journey of, you know, when

30:33I started

30:35barely could use a couple agents, and I

30:36was just observing every single thing.

30:39I think there's really no shortcut for

30:40going from here to there because this is

30:43really about your personal level of

30:45trust in agents. Right?

30:47Obviously, you know, as a as engineer

30:49you don't want to just slop code into

30:51production.

30:52So, how do you actually build up that

30:53trust? Takes

30:55um a lot of

30:57uh I guess taste and judgment. Um but

31:00uh you know, like I think plugins like

31:02Pstack definitely kind of help you

31:05get up to speed much quicker.

31:07Uh and so I guess it's It's like if you

31:10trust me and you trust Pstack, then in

31:13by extension you can maybe trust your

31:15agents. But if you don't trust me and I

31:17I definitely would not encourage people

31:19to blindly trust me.

31:22Uh

31:23uh you know, if you build up your own

31:24set of skills that you can obviously,

31:26you know, take a look at Pstack and kind

31:28of fork it, make it your own, improve

31:31the skills. Definitely encourage that.

31:33Um but for me it's really all about it

31:36just keeps coming back to trust. You

31:38know, every one of us here in this chat

31:40have a different standard for

31:41engineering. Uh and there are different

31:43things that are important for us in our

31:45code base. And uh when you are able to

31:49encode all of that into skills and you

31:51can verify that your agents actually

31:53doing them,

31:54that allows you to really kind of ascend

31:56this curve. And

31:59uh you know, start automating things.

32:02Uh there's another piece I wanted to

32:03talk about. Um if there's

32:06>> Yeah, go for it. I'll I'll think of more

32:07questions as I go, but yeah.

32:08>> Yeah, I think there's a third to this

32:10which I haven't talked about yet, which

32:12is

32:13kind of an interesting one, which is

32:14like refactoring and rewriting. Like one

32:17of the uh I guess most controversial one

32:20of the most controversial topics in the

32:22industry, I think, is like should you

32:25rewrite your app or not?

32:27Um because I think engineers are very

32:30prone to this where especially when you

32:32join a company, you come in and you see

32:34like the code base and you're like,

32:35"Man, this is

32:37Like who wrote this code? You know, it's

32:39it's terrible. I want to rewrite the

32:40whole thing. There is a very common

32:43inclination and I think a lot of, you

32:45know, before agents, um and I guess

32:48arguably even now, people will

32:49definitely discourage you from re-

32:51rewriting stuff.

32:53But I'm actually here to make a case for

32:54why you might want to consider it.

32:57Um because

33:00I think it really depends. Uh you know,

33:03brownfield applications I think are

33:05actually in a pretty good spot,

33:07especially if they're set up well

33:09already.

33:10Uh and like recently I've been talking

33:12to some people, but uh you know, I I I

33:14was just observing. I I just noticed

33:17this

33:18parallel, which is that a lot of big

33:21tech company problems are now

33:23everybody's problems.

33:25Um and the big tech company problem, you

33:26know, like when I was working at Meta,

33:28like we had this giant mono repo. We had

33:31like, I don't know, tens of thousands of

33:33engineers just, you know, like banging

33:35on their keyboards and and shipping

33:37code.

33:38And

33:40a lot of really great engineers at Meta,

33:42uh but uh I'll say like, you know,

33:45you'll be surprised that the code

33:46quality is actually not that good.

33:48>> [laughter]

33:48>> Um and so

33:50I often joke that like, you know, before

33:51AI slop, we had human slop.

33:54Um and so uh you know, I think a lot of

33:56big tech infra, like uh like what Meta

33:59has or Google, you know, you know,

34:01really big tech companies,

34:03are actually designed for that, where

34:05you you're sort of like, you're catering

34:07to the the you know, like

34:09uh this sounds so bad to say, but like

34:11the the least capable engineer on your

34:13team, right? You build you build

34:15frameworks, you build conventions, you

34:17build guardrails, you know, you restrict

34:20credentials so that, you know, your

34:21intern doesn't wipe your production

34:23database.

34:24Um

34:26there's uh you know, if you have that

34:28level of infra already, I think your

34:31agents can actually already do a very

34:32solid job. Right? Because they have the

34:35the guardrails are already in place for

34:38agents to not cause havoc or not cause

34:41too much havoc in your codebase. Um and

34:45you can always add more, you know,

34:46guardrails.

34:48Uh but I think like greenfield

34:49applications especially are you know,

34:50like the brand new applications are like

34:53the biggest risk in my opinion. Uh and

34:56also the greatest opportunity.

34:58Because, you know, if you vibe code a

35:00project uh a prototype

35:03um like we did for Grokbot, you know,

35:04Grokbot was spun up very very very

35:06quickly. Um and if you if you haven't

35:09heard of of Grokbot, it's like our a new

35:11application we just launched yesterday.

35:13Uh it's it's really cool.

35:15Uh lets you orchestrate your create like

35:18individual agents that have their own

35:20identity and you can kind of orchestrate

35:22them. It's super cool. Definitely check

35:23it out.

35:25Um but yeah, that was it's like a very

35:26it was a very greenfield application

35:28like most prototypes are.

35:30So, it was like vibe coded very quickly.

35:32Humans were not reading the code at all.

35:35And

35:36uh I had this tweet recently

35:39uh where I said something about organic

35:42architecture.

35:43Um

35:45Let me I'll find it.

35:46Uh but the idea is that

35:50uh when you have a completely vibe coded

35:52application, you essentially have no

35:54guardrails whatsoever. So,

35:56uh your agents

35:59when you give them a task, they will

36:01just solve it in whatever method is the

36:03most convenient.

36:05And over time, you get into this

36:07uh

36:08situation where you have a codebase that

36:10is spiraling out of control because you

36:12don't understand it. Uh your agents

36:15understand it, I guess, in a way, but

36:17like they've built something that is,

36:19you know, optimized for short for

36:21shortcuts.

36:22Uh and uh you know, it will you will

36:25suffer you have a lot of of issues with

36:27that application.

36:30Uh so, I think starting your code base

36:33with uh like very strong strains is

36:37very much needed.

36:39Uh because like when you have a code

36:41base that you can trust, right? When you

36:43have guardrails that actually help you

36:45uh

36:47uh help your agents write good code, you

36:49can get into the you know like into this

36:51part of the curve where I I where I like

36:54I I said, you know, I woke up today and

36:56I had like 20 PRs merged uh by my

36:59agents. And that's because I invested a

37:02lot lot a lot of time

37:03uh over 600 PRs I I I calculated

37:06yesterday uh

37:09when I refactored all of GrokBot to this

37:11new architecture that I've been

37:12building.

37:13Um

37:15and yeah, I've I've gotten to a point

37:17where I

37:19I don't really look I really don't look

37:20at the code anymore. And um I say that

37:23not just, you know, to sell you tokens,

37:25but

37:26because I you know, it it it took a lot

37:29of work to get to that point. I spent a

37:30lot of tokens to get the code base to

37:33this point where I no longer have to

37:34look at it.

37:36Uh but I'm very excited because

37:38you know, of the potential where, you

37:40know, it's not just it this doesn't just

37:42benefit me. It benefits everyone

37:44contributing to GrokBot.

37:46And it also empowers, you know,

37:48designers and product managers and, you

37:50know,

37:51even GTM people to add features to

37:54GrokBot. And I don't have to worry, you

37:56know, I don't have to to wake up at

37:58night in in the middle of the night and

37:59worry like, "Oh, Someone's just

38:01merged a perf regression." Right? I have

38:03a ton of

38:04constraints and CIs like it's actually

38:07very annoying to write code in in

38:09GrokBot, but like agents absorb all of

38:11that annoyance.

38:13Um but yeah, I'm happy to talk about

38:15what exactly that is. Um

38:18>> Yeah, I think one question

38:20um

38:20>> Yeah.

38:21>> before we get into the this part here is

38:23just around that element of like what

38:25your your your CI looks like or maybe

38:27some of the constraints and then also

38:29like the average PR size. I saw a

38:30question about that earlier. Just to

38:32give people a you know, kind of a a

38:34glance. It doesn't have to be like

38:35mathematically average, but just a you

38:37know, like what generally the size of

38:39the PR is

38:41um if it's only a couple lines of code

38:42or you know, um yeah.

38:45>> Uh

38:45um

38:47I think it depends. Uh let me

38:51I'm trying to do this in a way where I'm

38:53not going to like

38:53>> Yeah, yeah, you don't have to share the

38:55actual number. If you like the actual

38:56number just

38:57>> this is is fine.

38:58>> benchmark

38:59>> But like we have So okay.

39:01This is not

39:02that interesting, but uh well, fun fact

39:05is that virtualization in Grokbot and in

39:09uh Cursor is actually powered by uh

39:12Pretext,

39:13uh which is a sort of new library that

39:16someone's built. Um that's really

39:19interesting. You should You should check

39:20it out.

39:21But that's not really that important.

39:23Uh I think the average PR size I

39:25actually don't know I I don't know if I

39:27want to click on these.

39:29Uh I probably can, but I would say like

39:32they can range anywhere from a few

39:34hundred lines or 50 lines to like a

39:36thousand depending on what the thing is

39:39doing.

39:40Uh so like here I'm actually like

39:41deleting a bunch of files, so I expect

39:43that it's just this like mostly

39:45deletion.

39:46Uh but yeah, it it kind of varies.

39:49>> There's no like

39:50>> There's no like hard cap or hard limit.

39:52They're all like 50 line PRs.

39:54>> there's definitely no hard cap, but I I

39:55do encourage my agents to split up their

39:57work into multiple PRs.

39:59Uh I do that mostly because uh

40:03I like I like the idea of but I guess

40:06maybe this is much harder to do now as

40:08in the world of agents and you have like

40:10so many commits.

40:12But I like the idea that you know, the

40:13Git history is a very rich source of

40:15context. Uh and I like the I like each

40:19PR to sort of atomically describe what

40:22that

40:23small piece of thing is doing,

40:25which also makes it easier for me to

40:26revert changes and like figure out, you

40:28know, oh, I shipped a bug and it's this

40:30it's here, right? It's not in this

40:3240,000 line PR where you who knows what

40:35landed in there.

40:38Uh but I I don't have a hard cap on PR

40:41size.

40:42>> Cool. And then um yeah, also quick

40:45question on like CI. So again, you don't

40:46have to go into like uh the screen share

40:48of like your your CI does, but just

40:50generally, would you describe what the

40:52CI kind of looks like uh or how strict

40:54it is?

40:56>> Uh yeah, so

40:58uh well, specifically for Grokbot, so

41:01Dune is the is the sort of cheeky code

41:04code name for the architecture that

41:06we've built for

41:07Grokbot. Um the CI looks pretty annoying

41:12because there's checks for everything.

41:14So, like literally, I have um

41:18uh well, if you've written any React for

41:19example, you know, you know that one of

41:21the biggest foot guns in React is use

41:23effect. Uh so, in

41:26uh Dune and in Grokbot, we've banned use

41:29effect. So, Dune is just you can the the

41:32mental model of what Dune is,

41:34uh you can kind of think of it as like

41:35Next.js for

41:37uh Electron apps and it's designed for

41:40agents to write uh and it's like custom

41:42for, you know, our agent for the

41:45applications.

41:46Um so, the CI checks are very like

41:48specific to that, like, you know, don't

41:50use use effect. It's it's it's banned,

41:52like, CI will fail

41:54uh and yell at you. We have like some of

41:57the more interesting ones that people

41:58might raise eyebrows is like I actually

42:00banned code comments as well,

42:03uh which is very interesting. Um but I

42:06noticed that 99% of the time agents just

42:10write code comments that kind of

42:12describe some historical thing that is

42:14actually totally irrelevant to the code.

42:17Um,

42:18like it will often say like, you know,

42:19"Oh, Lauren said you should never do

42:21this." and it's now in the in the code

42:22comment. I'm like, "What? Like, why what

42:24that that was I didn't say that as like

42:26a durable, you know, global rule. I just

42:28meant like your this PR sucks and you

42:31should change that part."

42:33Uh, agents

42:35don't really understand us that well,

42:36surprisingly.

42:38Uh, and or they kind of assume too much

42:40and they kind of do things in like very

42:42stupid ways.

42:43So, like yeah, we just ban everything

42:46everything you can imagine like the

42:47agents are bad at, we ban.

42:50Uh, so one example that we actually

42:53suffer a lot in the agents window is

42:56we have, uh, you know, you know, if

42:58you've used agents window, you've

42:59definitely seen performance issues and,

43:01you know, we're constantly trying to fix

43:02them.

43:03Uh, but it's like a

43:05it's a never-ending struggle because

43:07there's so many pull requests that get

43:09merged.

43:10Every any one of them could just regress

43:12performance or stability or reliability.

43:15Uh, you know, the agents window doesn't

43:16have this architecture yet. I plan to do

43:18bring this learning back there and kind

43:21of refactor everything there.

43:23Uh, but

43:25uh, it just regresses super often

43:27uh, because uh, there's there's just one

43:30example is like we have very poor

43:33um, isolation between processes. So,

43:35like on you know, on Electron, you have

43:37a renderer thread that renders your UI.

43:40But you also have like a main thread

43:41that you can run other code that, you

43:43know, doesn't need to block the

43:44renderer.

43:46Um, but we do a poor job of separating

43:48those things. And so, often times you

43:51just accidentally have code that gets

43:53pulled into running on the renderer

43:55thread. All right, and then all of a

43:56sudden you're competing with the the

43:58renderer that you know, that has a very

44:01If you want like 60 FPS, you have to

44:03every frame that gets drawn has to be

44:05done in 16 milliseconds. It's a very,

44:08very small, you know, deadline per

44:09frame.

44:10Uh if you want, you know, a very smooth

44:13product. Uh and when you start building

44:15bringing in accidentally bringing in,

44:17you know, things that are like very

44:19computationally heavy or they have a lot

44:21of IO, uh

44:23then you just get into like a lot of

44:24jank. Uh your your FPS really drops. You

44:27start uh you know, losing frames. You

44:30get long tasks that

44:32take more than 16 milliseconds and you

44:33just get this really choppy experience.

44:36So, all of those patterns that we've

44:38learned basically building Electron

44:40apps, we've encoded into this framework

44:42and it becomes like a hard failure. So,

44:45I literally in in Grokbot, we literally

44:47have a directory called Electron main,

44:49Electron renderer, and we have uh

44:52import

44:54uh CI, I guess, where we actually check

44:57the dependency graph to make sure you're

44:59not accidentally importing code from one

45:02directory to another.

45:03Uh so, that's enforced by CI,

45:07um as well as Bugbot, uh which is our

45:10which cursors

45:11um like code review tool that runs on

45:14CI.

45:15Uh you know, in our agents MD, it's

45:17everywhere. Like so, I I I I um

45:20I have this thing here where I talk

45:22about like

45:24um you know, like there are multiple

45:25layers, I think, for building a good

45:28code base.

45:29Uh obviously, the code base is one where

45:31uh if you have an architecture like

45:33this, where it's extremely strict, uh

45:36you know, the the the way to build

45:38features is very conventional. That's

45:40like the strongest strongest level of

45:42enforcement because agents just love to

45:45copy existing patterns.

45:47So,

45:48uh one example of this in Grokbot is

45:50like we have this these concepts called

45:52like a feature and we have entry points

45:54and transcript cards. Like all of you

45:56know, the cards that you see in the

45:57chat.

45:58These are all like

46:00like nouns, I guess, in in the

46:02framework. And so, there's a very

46:04conventional way of creating them. And

46:08so, like a feature is all in in a single

46:09directory, as an example. And so, all of

46:12the code that contributes to that

46:13feature lives in one directory. So, it's

46:16all co-located in one place. Makes it

46:18super easy, you know, agents don't have

46:19to like uh grep around and try to figure

46:22out like where all the things are.

46:24It just looks at the feature and like,

46:25"Oh, okay, I'm working on the onboarding

46:28feature in Grokbot. Uh I'm just going to

46:31work in this directory." And for 80% of

46:33the work, it's mostly just very

46:35encapsulated there.

46:37But uh like it's like

46:40very It's like designed again for, you

46:42know, like the dumbest agent. Like, you

46:44don't have to think. All right, the the

46:46the One of the key principles I have for

46:49this framework is like the shortest the

46:51shortest path is the best path.

46:54So, uh because that plays exactly to how

46:57agents love to write code. Like, they

46:59like to take shortcuts, really. You

47:01know, they'll they'll find the quickest

47:02way to solve the problem. So, why not

47:06make that the best way to solve the

47:08problem?

47:09Uh so, I I I I probably won't get into

47:11all the specific details. Um and uh the

47:15the

47:16this framework is really more of a

47:17collection of ideas and principles,

47:19rather than something that we'll open

47:20source.

47:21Uh you can you can, you know, screenshot

47:23this, I guess, if you want and uh tell

47:25your agent to

47:27uh do some build build something like

47:29this for you to

47:31um Yeah, but it's really all about the

47:33layers. Uh you know, like the the code

47:36base is one part with features uh and

47:38directories and you know, import

47:41uh blocking import dependencies uh that

47:44shouldn't be imported. Uh but and and it

47:47all enforces that and static analysis.

47:49So, like uh there's CI checks. We have a

47:52lot of lints for bad patterns that we

47:55observe. Uh compiler diagnostics. Uh,

47:59there's also rules and bug bot which are

48:01um, I think like three, four, five are

48:04more soft, right? These two actually

48:07make

48:08make CI red, right? So, that you know,

48:11there's a hard constraint where the

48:13agent can just write crappy code.

48:17For rules and skills and bug bot,

48:20your agents can still forget, right? You

48:22can still

48:23or it may not always consistently apply

48:26them. So, I like to layer them, but I

48:30don't I don't like to rely on them as

48:32the only source

48:34of enforcement because it's very, very

48:35soft, right? And if you if you only have

48:37rules and bug bot and skills and a style

48:40guide for your code,

48:41you will it's only a matter of time

48:43before your code base looks like

48:45complete trash. Uh, I'm sorry to say

48:47that, but uh, I definitely recommend you

48:50all like, you know, investing

48:51in, you know, things that can be hard

48:53enforced. Right? And this is why,

48:56you know, maybe uh, the choice of tech

48:58stack that you use is also very

48:59important.

49:01Um,

49:02like I think for example, Rust is sort

49:04of making, you know, it's like getting

49:06super popular again, uh, because the

49:09compiler is so strict, right? The

49:11compiler enforces so many different

49:13things, you know, there's a borrow

49:14checker that you have to appease.

49:16And if as long as you make sure your

49:18agents don't write unsafe code blocks,

49:20uh, you can more or less feel somewhat

49:23confident that if the code compiles, it

49:25probably works and it's good.

49:27Uh, but you know, you see, it gives you

49:28that level of trust and confidence that

49:32you as a human engineer no longer need

49:35to go and check it yourself. You you

49:37know, you you rely on

49:39code and static analysis to actually

49:43make that uh, a lot smoother.

49:46Um, and I I guess the worst part the

49:49worst place to be in is if you are stuck

49:52in code review land where you actually

49:54enforce all of the constraints, the

49:56invariants in your code base

49:58by literally the human person saying,

50:02you know, reading the code and be like,

50:03"Okay, you should not do this."

50:04Right?

50:06Every time you have to do that, you

50:07should consider that as a code smell,

50:08like a uh anti-pattern. And you should

50:11say, "Okay, instead of me commenting on

50:13the PR,

50:14how do I turn this into a hard rule?

50:17Right? How do I turn this into a lint

50:18rule? How do I turn this into a CI

50:21failure? Or how do I even categorically

50:23eliminate this problem

50:25uh entirely?

50:27Uh

50:28I I can talk about another migration

50:30I've done, but I'll probably pause here.

50:32>> Sure.

50:33Yeah, I feel like that's that's where I

50:34am, to be honest, is is what you're

50:36describing right now, which is that like

50:38I don't have all of these rules. Uh so,

50:40I have some things to go do after this

50:41session

50:42in terms of being able to scale my

50:43agents. I'm I'm definitely on like the

50:46uh you know, maybe a couple of parallel

50:48ones locally stage, so like two to three

50:50locally. And I'm sure most people here

50:52are on the same, so uh

50:54yeah, I know we only have a couple

50:55minutes left. Uh Lauren, was there

50:56anything else that you wanted to to

50:57highlight? I obviously there's lots of

50:59questions, so I can grab more, but I

51:01want to give you a few minutes uh if

51:02there's anything else you want to talk

51:03about.

51:03>> I think I've been yapping for quite a

51:05lot, so I'm Maybe let's just do

51:06questions.

51:08>> Okay, cool. Uh one question I had uh a

51:10couple of uh came up a couple of times

51:12was just around like token usage.

51:14So, the the question is like is is what

51:17you're describing a realistic thing for

51:19people who are on, you know, uh

51:22a a normal set of token usage, they

51:24don't have, you know, basically

51:25unlimited tokens uh to work with.

51:29>> Well, I think that's a really good

51:30point. I mean, like obviously, you know,

51:31I work at a AI lab where we have

51:34unlimited tokens, so uh

51:36I definitely cannot

51:39say that, you know, this is something

51:40everyone

51:42should do in the exact same way that I

51:43did it. I think it's possible to get to

51:45this point without, you know, breaking

51:47the bank.

51:49But, you know, if you're like an

51:50engineering leader or, you know, you're

51:52you're you have a startup that you lead,

51:54um I think that may be a question of

51:56ROI.

51:58Um and it's like

52:00uh

52:00yes, you spend a lot of money on tokens

52:04in the upfront stage. You know, you like

52:06refactoring your code base is going to

52:07take a lot of tokens. Uh adding all

52:09these things uh is going to take a bunch

52:11of tokens.

52:13But, if we're heading to a world where

52:15agents are writing all the code,

52:17and you know, you want to be very lean,

52:20right? You don't want to have to hire

52:22You don't want to be You don't want to

52:23become like meta, right? Like I mean

52:25like in terms of You don't want to

52:26become a 10,000-person engineering org

52:29because

52:30I mean, that's a cool problem to have,

52:32but also you know, you you have so much

52:34overhead. There's like planning, you

52:36know, like you it it's it's uh

52:39personally I I wouldn't uh

52:41it it's not super fun. But, um

52:44I think you want to stay very nimble,

52:45right? And you want to you want to be

52:47like agents are all about allowing you

52:49to do things that you couldn't do

52:51before. That's really to me like the

52:53value of agents. You know, it's not just

52:56throwing tokens on every single little

52:58thing.

52:59But, um to me like the thing I couldn't

53:01do before is like

53:02enforce this level of constraints in a

53:05code base by myself.

53:07Right? Like I'm just a single person,

53:09you know,

53:10uh it would have taken me years to build

53:13this framework uh and do it all the

53:15refactoring and

53:18test everything myself and verify, you

53:20know, like run Imagine if there was just

53:22me right now in in pre-agent era, just

53:25like running, you know, by my It would

53:27take me so long, right? And my salary is

53:30pretty high, right? Like So, you know,

53:33the the question I think an engineering

53:34leader might have is just then,

53:36you know, like what is there's a

53:38tradeoff of do you hire someone to do

53:41this,

53:42or do you spend the tokens to set up a

53:44code base so that even

53:46the the most naive, right, the dumbest

53:49agents can do a good job.

53:51And when you actually get to this point,

53:53like even agents that are not, you know,

53:55Fable size do an excellent job of

53:58writing code.

53:59And this pays a lot of dividends as well

54:01for me and personally where I've

54:03empowered not just myself, but

54:06again, like PMs, designers, engineers

54:09who are not familiar with Grok but to

54:11just contribute in a way that

54:14is sustainable.

54:16So, I think yeah, it's definitely like a

54:18trade-off for sure. You know, like

54:19nothing's like free for sure.

54:22Uh and tokens are pretty expensive.

54:24Uh but oh, actually,

54:26I don't I I I don't know how many of you

54:28have seen this, but we actually

54:29announced Grok 4.6 today. So, very

54:32exciting. Finally out. Um so, yeah, Grok

54:354.6 would be like a great It's very very

54:37smart. Uh it's really good on on the

54:40benchmarks. Uh and it's the same the

54:43tokens Uh well, uh I hopefully I'm not

54:45saying this thing correctly, but uh I

54:47believe the cost per token is the same

54:50as 4.5.

54:51So, you're actually getting more

54:53intelligence for the same cost.

54:56Uh I think this is an area that Cursor

54:58tries to Cursor and SpaceX AI

55:01try to really optimize for like that

55:04Pareto frontier of, you know, cost

55:06versus intelligence. Uh you know, we

55:08don't necessarily want to build the

55:10biggest model ever because that is

55:12extremely expensive to run.

55:14It's really about like how do you find

55:16that sweet spot, right? You don't You

55:17don't need a giant model, but it's just

55:19super smart, right? And it's not very

55:21expensive for inference.

55:24Uh but um yeah, I think to kind of round

55:27it up, um

55:29I think it's like a it's it's there's a

55:31If you do your own analysis, I feel like

55:33it's pretty positive.

55:35It It'll be pretty positive that the ROI

55:38you get from investing in stuff like

55:40this uh just empowers not just yourself,

55:43but your whole team to be so much more

55:45productive. Right? Like imagine if you

55:47have an army of engineers like me who

55:50are shipping so much improvements and

55:52and bug fixes uh you know, every day.

55:56Right? Like

55:57that is pretty exciting.

55:59>> Cool.

56:00One last question before we wrap up.

56:02This one is for the people in product on

56:04the on the call. So, let's say we do

56:06have an army of engineers who are

56:08shipping like Lauren. I'm just curious

56:10like how is the product team or other

56:12functions of your company keeping up

56:13given that like if you're shipping so

56:16quickly, have have are they using AI

56:18more to do their jobs? Like as much as

56:20you can speak to that. I obviously you

56:21don't have like you're not in that role,

56:23but just curious about how that works.

56:25>> Um I think this is where Grokbot has

56:28been actually exceedingly powerful. Uh

56:30where

56:32so before Grokbot like you know,

56:34obviously Cursor only had Cursor. Like

56:36we only had agent window. We had a CLI.

56:38We had an IDE.

56:40And these are really like

56:42power user tools, right? Like the

56:44they're designed for developers. So,

56:46it's very very developer-centric. You

56:48can do knowledge work in them, but it

56:50like the UI is not really optimized for

56:52that.

56:53So, we actually didn't really have uh

56:56well, I think like a lot of people like

56:58you know, GTM, product, like they might

57:00have used Cursor uh to do their work,

57:03but it definitely wasn't like a the like

57:05full experience with them.

57:07Um I think now with Grokbot

57:10uh it's become Grokbot is basically like

57:13the Cursor moment for people who are not

57:16in tech in my opinion. Like it's like

57:18it's like a very very accessible way to

57:21use agents in a very comfortable, very

57:23familiar interface. It looks like

57:25iMessage. Um and it's very fun too, you

57:29know, you can give your agent to be

57:30name. Uh you can have you can kind of do

57:33orchestration with in a very like

57:35natural way where you can sort of

57:37you know, each agent is like a person

57:39and now you're on a team of agents like

57:40working on you have one one agent per

57:42account that you manage as an example.

57:44Or if you're a PM, you have you know,

57:46you can have an agent that summarizes

57:47all the work that Lauren did last night

57:49and then now you know what I did, right?

57:52So, I think our PMs are leveraging that

57:54a lot. And they're shipping code, too.

57:57Uh, so you know, I often times they will

57:59just say, "Oh, here's a bug. I fix it.

58:01Can you look at it?" And then I'll go

58:02review it and actually it's just

58:04perfect. I'm like, "Okay, stamp." Uh, so

58:07uh, I think that shows that you know,

58:08the the Dune architecture is holding up,

58:11right? The all the the really strict

58:13constraints allow

58:15people who are not experts in

58:17engineering to contribute at a high

58:19level.

58:20Uh, so I'm I feel like I'm already

58:21seeing that pay off a lot where uh, you

58:24know, designers and PMs are just able to

58:27to to ship features directly.

58:30Um, and that just makes

58:32the Grock bot team super fast, right?

58:35Where we can ship so quickly.

58:37Um,

58:39and we have a lot planned, so I'm very

58:41excited uh, to you know, uh,

58:44to to ship more uh, ship more stuff.

58:46>> Yeah, that's awesome. Uh, we are at

58:48time, so I I guess Lauren, if if folks

58:51want to support you, maybe go try Grock

58:53bot, try out uh, 46 and uh, you know,

58:57just get to get provide some feedback.

58:58But yeah, this was awesome. Really

58:59appreciate you taking the time. Uh,

59:01thanks everyone for all the messages in

59:02the chat. Uh, lots of good questions. I

59:04know we didn't get through everything,

59:05but uh, as I kind of said at the top,

59:06way more questions than than we could

59:07get through. But uh, yeah, really really

59:10thanks thanks for for joining. Thanks

59:11everyone for joining and uh, hopefully

59:13you you enjoyed the session.

59:15>> Yep.

59:16>> All right.

59:16>> Yeah, I see see you thanks for having me

59:18and uh, if you have any more questions,

59:19yeah, just DM me on Twitter. Uh, I'll

59:21I'll open them up, I guess. I'll let the

59:23I'll let the floodgates

59:24>> to get you're going to get a lot of DMs.

59:25>> Yeah, I'll open the floodgates. So,

59:27yeah, DM me. Maybe I'll do like a

59:29Twitter Space at some point as well for

59:31more questions, but

59:33really appreciate everyone for showing

59:34up, uh you know, taking an hour out of

59:36your day.

59:37>> Yeah. All right. Thanks, all. I'll see

59:38you in the next one.

59:39>> Okay, thanks, everyone. Bye.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.