Free YouTube Transcribe

Video transcript

Lauren Tan workshop XAi Grokbot

AI Agents for Business · 10,753 words · 49 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Everyone, I am Lauren Laurent Tan. I

0:03guess not many people know my last name.

0:05Uh, I am Potato on Twitter. Uh, potato

0:10with spelled with an E. Um, and I have

0:13been at Cursor for about 5 months. Uh,

0:18previously I was at Meta where I worked

0:20on the React team. Uh, specifically

0:22working on the React compiler, uh, which

0:25was a whole lot of fun. Uh I'm still on

0:27the on the core team and and uh

0:29contributing to open source here and

0:30there. Uh so that that's really nice

0:33that they still let me do that. Uh and

0:36before Meta, I was at Netflix uh where I

0:39was uh both a tech lead and uh I

0:43transitioned to be be an engineering

0:45manager um for about two years. So I've

0:50had a I've have had a lot of experience

0:52going between engineering management and

0:55being an individual contributor. Uh and

0:58I think something I've noticed actually

0:59which is quite interesting is that there

1:02are so many parallels with you know

1:04management skills and how to like manage

1:06agents. Uh and that's actually a big

1:09part about what I wanted to chat with

1:10you and everybody else about today. Um

1:14but yeah that's that's me. Uh I do have

1:17some like very light slides but uh it's

1:20not going to be uh just rambling. So let

1:23me just share my screen

1:25and hope that I don't leak anything. Uh

1:33oh no I need to allow permissions.

1:36>> No worries. Take your time. There's

1:38always tech tech trouble. Uh give me one

1:41second to rejoin.

1:42>> Yeah, go for it.

1:46I see many of you already know Lauren

1:48from the looks of the chat here. Um, so

1:51yeah, it's exciting to to get a chance

1:53to chat with her uh and and go through

1:55some of her her recent work. As you guys

1:57heard, you know, a lot of recent

1:59experience from from Netflix to to Meta

2:02and then now over at Cursor. Uh, we're

2:04going to chat a little bit about

2:05Grockbot as well. So, uh, that'll be

2:08exciting. I don't know if you guys saw

2:09that was a recent release. I think

2:10literally maybe yesterday or the day

2:12before uh from the cursor team which is

2:14kind of like um let's call it like

2:16agents for everyone. You can go check it

2:18out if you want and learn a little bit

2:20more about the product but um but yeah

2:22we'll we'll explore that a little bit

2:23today as well.

2:26All righty. Welcome back.

2:40>> And Lauren, you're just on mute there if

2:42uh you want to hop off of mute if you're

2:43chatting.

2:44>> Yeah, sorry.

2:45>> No worries.

2:47>> It's 2026 and I still don't know how to

2:49use Zoom. [laughter]

2:50>> That's all good.

2:51>> Uh okay. So, I assume you can see my

2:53screen.

2:54>> Yes. Yeah, we're good. Um, so yeah,

2:56today, yeah, I think I think the big

2:58theme for me as I've been using agents

3:01to write code, and I'm sure a lot of you

3:03have had the same experience as well, is

3:06how do you trust it? You know,

3:07especially if you are an engineer that's

3:10been writing code for a very long time,

3:12you have a lot of opinions and lessons

3:15that you've learned about doing good

3:18engineering. And when you see agents

3:20just, you know, winging it and, you

3:22know, guessing, hallucinating,

3:24uh, you know, confidently stating that

3:27they found the smoking gun, uh, for the

3:29hundth time, uh, but it's actually not

3:32the real problem. You lose a lot of

3:34trust. And when you lose when you don't

3:36have much trust in your agents,

3:39I feel like you you really can't get the

3:41most out of them. And for me, the

3:43parallel is like with management. Uh so

3:46if I'm an a manager an engineering

3:48manager of a team and I have a bunch of

3:51you know I have a team of engineers uh

3:54on my team and I don't trust them then

3:57the mode of operation I'm going to be in

3:59is going to be like micromanagement

4:00right I'll have to spend a lot of time

4:03looking over my reports shoulders and

4:06checking that they're doing their work

4:08well you know that they're not shipping

4:09bugs to production

4:13and so I drew this chart cuz uh it's not

4:16it's not a very scientific chart but

4:18like this is how I imagine

4:21myself and my journey through using

4:23agents. So you know like fast forward or

4:27back forward uh or fast back uh fast

4:32backwards like a year or so when you

4:34know nobody was or not many people were

4:36using agents to code. Uh I think you

4:41uh you know get into this mode where you

4:43are

4:45in very heavily in the loop with one or

4:48several like a handful of agents and you

4:51find yourself just constantly figure uh

4:53you know trying to understand what your

4:55agents are doing uh and you're very very

4:57in loop. You're watching every single

4:59output. you are sitting there prompting

5:03um and you really can't parallelize

5:05beyond that because you don't again you

5:08don't have that trust right you can't go

5:09to a 100 agents uh like spawn 100 agents

5:13when you don't even trust the output of

5:15one agent

5:17so over the past 5 months I feel like

5:20I've really been able to uh like ascend

5:23this trust curve and now I'm at the

5:26point where uh I actually have this

5:29Sounds kind of scary to say this and it

5:31makes me sound like a slop artist, but I

5:33I promise I'm not, but I actually have

5:36my agents now um automerging PRs for me.

5:39Uh which is like a wild thing to say,

5:42but um like I woke up today and there

5:44were like 20 PRs landed and I just

5:47reviewed them on Maine like they were

5:49already landed and they were good. Uh so

5:52how did I get to that point is basically

5:55what I wanted to talk about today.

5:58Uh and again like yeah feel free to jump

6:00in if you have questions Colin. Um but

6:04uh oh yeah of course I got to show this

6:07this chart. Uh where uh

6:12oh do do not trust to someone requested

6:16to control my computer. Uh probably

6:18won't do that. Uh but yeah, so this

6:21chart I think I I'm I'm sharing this

6:24chart not to kind of like flex but to

6:26kind of show like the journey like so

6:29you can see like the curve like it sort

6:30of like inversely matches the

6:33contributions I've been able to land at

6:35cursor. So I joined five months ago and

6:38five months ago like I you know my first

6:40month I was like not very productive

6:42because I was obvious you know I was

6:43learning the codebase didn't know what

6:45the heck was going on and as I got more

6:48confident in in my agents uh I've really

6:52been able to kind of ramp up my

6:54productivity. Uh and again like yeah

6:56like last month I shipped a thousand PRs

7:00which is ridiculous. Uh, and then this

7:02month we're only on the 12th. I'm

7:05already at like almost 800 PRs landed.

7:09Uh, so the velocity is definitely high

7:12and you you I'm sure a lot of you will

7:14definitely be questioning like how how

7:16much of this code is actually good. Um,

7:18and I think yeah, like that's definitely

7:21fair to question.

7:23Um, but uh, yeah, I think I think if you

7:28set up your agents well, you can

7:30definitely get to a very similar level.

7:34Um, and so I'm going to talk about how

7:36we do that.

7:39Uh so for me I think I'm curious like I

7:42guess call in your experience as well

7:44but uh for me I think the most important

7:48skill that you should have in your

7:51toolbox when you work with agents is

7:53verification.

7:55Uh and by verification I mean the

7:57ability for an agent to actually run the

8:00code uh or take CPU traces or heap

8:05snapshots or uh you know open an iOS

8:09simulator whatever you know however your

8:12application is exposed to your users it

8:15can do the same thing and uh run it for

8:19real and actually test and verify it

8:21don't work because that's the thing that

8:24really closes the loop. Uh it doesn't

8:26guarantee your agent writes good code.

8:29Uh but it allows them to at least write

8:31correct code. Uh which is a big a really

8:34big step forward for being able to trust

8:37your agent. Um

8:40I will I can share one example that we

8:43have uh within cursor. Uh

8:49oops

8:50where let me open this.

9:00Let's make this uh make me full screen.

9:03There you go. Uh so for the for cursors

9:08agent window uh so this is actually an

9:11interesting story but uh when I joined

9:14cursor 5 months ago uh they're actually

9:18uh well I was supposed to join a

9:19different team. I was supposed to join

9:20like the cloud agents team. Uh but then

9:23since I have a lot of experience working

9:25on react and agents window is a react

9:28application. Uh I was

9:31uh I was asked to basically help out

9:34with uh the agent window work. Uh but um

9:41there wasn't really a lot of like skills

9:43to help me. So I just found myself like

9:45okay uh agents which is going to launch

9:47in like a week right we have a really

9:49tight deadline and um there was uh you

9:53know I was just sitting there like okay

9:55I'm going to open up the performant the

9:57the chrome dev tools and just like take

9:59a trace look at it myself and try to

10:01make sense of this flame graph and keep

10:03in mind I was just like in my first week

10:05so I had no idea what I was looking at

10:07no idea where you know I mean I had some

10:09idea but you know the the code base was

10:11completely fresh to me Uh, and I

10:14realized like my agent had no idea

10:17either, you know, like I would take a

10:18screenshot of the trailer, I would

10:19download trades, I would send it to it,

10:20and it be like, "Yeah, it kind of looks

10:22like this, you know, uh, and it would

10:24like confidently state like it's this

10:26thing." And then I try to fix that and

10:28turns out that's not the actual thing.

10:32So, this was very very slow process. And

10:34if you've ever done any like performance

10:36work yourself or you know just even

10:38development with an agent where you

10:40don't have a verification skill, you are

10:43the verifier, right? You you're the

10:45bottleneck. You you you tell your agent

10:47to do something and then it goes off and

10:49write some code. Then you open up your

10:51you know local dev build and then you

10:53start to say oh you know doesn't work.

10:54Then you got to copy paste screen uh you

10:56know screenshots or console errors or

10:59whatever. uh and then your agent like

11:01slowly kind of like uh you know works

11:04with that and then tries to understand

11:06it and um fix the thing but then you're

11:10constantly just in the loop and and

11:12being a bottleneck. So there's really no

11:13way to parallelize. So the control glass

11:16skill is like one of the first skills I

11:18built uh for cursor. Uh and glass by the

11:21way is the code name for agents window

11:24that we use internally but it's just

11:27cursor I guess. Um, and so this skill uh

11:31is I guess the the the code itself is

11:33not super interesting. Your agent can

11:35very easily make one for you. Uh where

11:38if if you're building an electron app or

11:40a web app or even iOS uh applications,

11:45uh you can teach your agent how to use

11:47like the ChromeDev Tools protocol or

11:50through uh Apple has some utilities as

11:53well for running the simulator and

11:55taking traces and controlling

11:56programmatic control.

11:58as well. Uh so that's really useful. Uh

12:03but one thing I want to talk about is uh

12:06the

12:08this thing

12:10uh where is the read me? Uh so this

12:14skill comes with this very unique

12:16feature called or not feature uh unique

12:19file called a feature map. And so the

12:23story then is like I built this skill

12:25and so now the agent was able to uh

12:28actually run the agent window and take

12:30traces and whatnot. Uh but it had no

12:34idea what what the agent's window was.

12:37So um you know like someone would say

12:40like oh the the left sidebar is like

12:43laggy or something like that or you know

12:45the right side the the PR tab is not

12:48working and the agent would just be like

12:50kind of flailing around. it would spend

12:51a lot of time trying to like look up the

12:53code and you know where is this feature?

12:54How do I actually get to it on the UI

12:57which made it basically completely

12:59useless. Uh you know like we would I

13:01would run this skill locally and you

13:05know it would spawn a dev build uh but

13:08then it just be turnurning like it just

13:10try to click here. It it wouldn't know

13:12how to get to things um and it was just

13:15an awful experience. Uh so who's putting

13:19arrows on my screen? Um so uh yeah this

13:24this feature map has been really useful

13:26uh because it teaches the agent how to

13:28get to all of the features that you

13:30have. Um and in PAC the plugin that I

13:34I've made uh if you search for PAC

13:38cursor on Google you you'll find it. Uh

13:41but there is a create verification skill

13:43in that plug-in where it actually helps

13:45you set up something like this for

13:47yourself. Um including the feature map.

13:50So it will actually explore the code and

13:52build up this initial feature map that

13:54tells your agent how to get to all of

13:56the different features that you have. Uh

13:59and this is extremely powerful because

14:01now like you have these user reports

14:03that come in uh you you can actually map

14:06even like a vague report or even a

14:08screenshot. So we have this uh

14:11internally at cursor where uh we have a

14:14slack channel where you know lots of

14:17people giving us feedback on the agents

14:19window and rockbot and whatnot. Uh, and

14:22often times the report is very bad, like

14:26very low quality, like someone just put

14:28very often we get like a screenshot like

14:30and then someone just says question mark

14:31question mark question mark like what is

14:33this [laughter] and you know like the

14:35without this your agent is like I have

14:37no clue, right? But with a feature map

14:40like this, it has a lot more context and

14:43understanding of how to actually

14:45navigate, how to get to all the

14:47different features. Uh so like you know

14:49example like I guess like the sidebar

14:51like what is the sidebar uh you know

14:54like all the different sub features that

14:56are present in it um like from the user

14:59point of view here's where how do I get

15:01to it all the different keyboard

15:03shortcuts

15:05uh even like the the what do you call it

15:07the DOM elements or yeah like the

15:10attributes that you use for selecting

15:12things through the CDP uh are all there.

15:17So uh again yeah this is like really

15:19really powerful uh for for agents

15:23>> uh and pstack ships uh that create

15:26verification skill but also a maintain

15:28verification skill uh so you can keep

15:30this up to date.

15:32>> Cool. Yeah, I was just going to ask how

15:34you created that. So do you mind sharing

15:35a little bit more about um that that

15:37process in the context of Pstack and

15:39maybe just what Pstack is for the folks

15:40who aren't familiar?

15:42Yeah. So, Pstack is pretty interesting

15:44because uh well, first of all, the name

15:47is kind of goofy. like the P the P in P

15:50sack is like potato potato stack because

15:54I um so uh uh there is a pretty uh

16:00famous person Gary Tan who is the CEO of

16:03Y Combinator and he's come up with this

16:05plugin called GStack uh Gary Stack and

16:10uh funnily enough we share the last name

16:12we have no relations uh but I thought it

16:15would be funny to kind of you know poke

16:16fun at Gary and make Pack back my

16:19version of of of of his plug-in. Uh but

16:23kind of tailor it to my own set of pract

16:26engineering practices.

16:29Uh but I honestly actually never set out

16:31to build PAC. Uh it just started with a

16:33bunch of skills, right? Like I started

16:35with that control glass skill and then I

16:37started with another skill like called

16:39how which I also noticed through like

16:42observing agents. Um, so like you know

16:46in the early days of me, you know,

16:47trying to climb this ladder, I was like

16:49super in the loop and I was basically

16:52nitpicking my agents to an extreme

16:53degree. I was uh I would tell it um you

16:58know this feature has stopped working.

16:59Here's a bug report like why isn't it

17:02working?

17:04And very often the agent would just like

17:06confidently state like oh it has to be

17:09this right it has to be this thing. And

17:11I noticed like when I looked at the

17:13actual tool calls, I noticed it wasn't

17:15actually reading the code that I thought

17:18should be affected. And that made me

17:20just extremely suspicious. And at that

17:22point, I was like, I'm not going to I

17:24can't trust any this agent anymore cuz

17:26it's just it's just completely

17:27hallucinating. And I think

17:31I think it's very easy to just you know

17:33like build up that distrust and not and

17:36kind of feel helpless like you know you

17:39don't know how to help your agents

17:40succeed. But like again I think the the

17:43the management analogy is super helpful

17:45because like imagine if you were a

17:47manager of an engineering team and you

17:49had an engineer on your team who was a

17:52really good coder. No business context

17:54whatsoever. you know they they just you

17:56just hired them and they they onboarded

17:58you know like 5 seconds ago. Uh and so

18:02how do you actually teach that person to

18:04be effective? So how you do that is

18:06through a skill. uh skill being just you

18:09know it's just markdown right but you

18:11know it encodes a lot of information

18:13instructions a lot of uh you can really

18:17draw out a lot of intelligence from an

18:20agent by well some people on Twitter

18:23call it like you know pull the agent to

18:25a different latent space which is kind

18:27of like a fancy way of just saying like

18:29since uh you LLMs are sort of like they

18:31predict the next token uh when you give

18:35it some high quality tokens uh to begin

18:37with then you know it it can kind of

18:39pattern match on like a higher space

18:42that's you know smarter

18:45um so that's like a very interesting

18:48model there but yeah I built Pstack very

18:50very incrementally uh so uh started with

18:54just really observing how agents you

18:56know all the fail different failure

18:58modes of of that agents were having and

19:01every time I saw that I just okay I'm

19:02just going to make that a skill right

19:04like stop hallucinating actually go and

19:07search up, look up the code, use a lot

19:09of sub agents, uh, and yeah, stop

19:12guessing.

19:16>> Yeah, that makes sense. One, one kind of

19:18followup question here, uh, both from

19:19myself and from a bunch of people in the

19:20chat. So, I guess it's two two parts.

19:23So, one is like how do you maintain

19:24these skills? So, like the product

19:26changes over time. Obviously, there's a

19:27lot of people shipping against the

19:28codebase. So, how do these skills get

19:31maintained? Uh and then second to that

19:33is like how do you know when your

19:34verification is is good enough? Uh like

19:37and you know you can trust that the ver

19:39verification loops that you've built are

19:41going to I guess you trust that the

19:43outputs uh when they're done.

19:47>> Uh yeah maybe I'll talk about um I think

19:50there are some related maybe I'll start

19:52with this one first. So like how do I

19:56maintain these skills?

19:59So, um, if you're not familiar with this

20:01concept, an eval [clears throat] is

20:03essentially like a way to, uh, well, I

20:06the mental model I have is like it's

20:07like a unit test for an agent. Um, and,

20:11uh, you can actually make your own eval.

20:14You don't need like a special framework

20:16for them. You can build you can you can

20:18build one depending on like, you know,

20:20how scientific and how rigorous you want

20:23to be. Uh, my screen is red.

20:26>> Yeah, there's a little button. Um,

20:28sorry. Do

20:29>> you mind like disabling the drawing or

20:31something? I I can't see my screen.

20:33>> Yeah, sorry. If you guys could not draw

20:35on the screen, that'd be great. But, um,

20:36there's a little button in the

20:38>> Is that a troll?

20:39>> Yeah, the little drop down.

20:42>> How do I clear?

20:44>> Yeah, you got it. Perfect.

20:49>> Yeah. So, eval

20:52test your skills basically. And actually

20:56in Pstack we ship uh under potato mode

20:59there's a playbook if you search for it

21:00called eval playbook. Um and it's uh

21:05uh it's like not it's actually pretty

21:07pretty rigorous the way it's done. Uh

21:10but um essentially what I do is I spawn

21:13a lot of different sub agents. I have

21:15like my main coordinator agent uh come

21:18up with a rubric for uh what I want the

21:22skill to do. Um, and then it spawns all

21:26these sub aents and it it creates

21:28individual directories for them uh which

21:31are cleverly named to not let the sub

21:35agent know that it's being evaluated

21:37because uh agents can actually tell and

21:40when they do they change their behavior.

21:42Uh, but it does a bunch of stuff like

21:44that to um essentially yeah like test

21:49whether or not the skill I'm making or

21:51changing is actually doing what I think

21:53it does. Um, and one of the really nice

21:56things about cursor is that we are we we

21:58support so many different models. So you

22:01can actually eval your skill across all

22:03sorts of different models um and you

22:06know get a sense of how well it performs

22:09across that different matrix. um

22:12especially for the models that you use.

22:15Uh so I do this a lot. Every time I I

22:17modify a skill, I will run one of these

22:20uh like the Eval playbook uh and make

22:22sure that you know it's actually leading

22:25to a result I want. Uh but I will say

22:28like

22:30maintaining skills is actually pretty

22:32hard. Uh it requires I think a lot of

22:34taste and observation. So, you kind of

22:37need to be very good at being a backseat

22:40driver, you know what I mean? Like, if

22:42you do pair, if you ever done pair

22:44programming, for example, uh, and you

22:46watch a co-orker code and you're just

22:48like, you you could probably do this

22:50better, you know, you could do, you

22:51know, like, why did you not do this,

22:52right? You you ask a lot of questions to

22:54your coworker. And it's kind of a

22:56similar thing here. You like you don't

22:57want to just be a passive observer of

22:59your agent. You want to be very in the

23:02driver seat in the initial stages when

23:04you're building up your own set of

23:05skills. uh you know obviously you can

23:07use something like PAC but if you're

23:09building your own set of skills it's

23:11very I think you know opening up the all

23:14the tool calls and like reading the code

23:16and reading all the uh the agent

23:19behavior and their thinking blocks is a

23:22really great way to see where they they

23:24fail right like what what you know where

23:28are they being done and then you can go

23:29and build a skill for that and then with

23:32verification how you trust it is it's I

23:35think it's also a very similar iteration

23:37loop uh where you know like I actually

23:39did the same process for verifying the

23:42verification skill where I actually get

23:46um so one thing that's interesting about

23:48eval is that you can sort of hill climb

23:51them meaning that uh your eval can

23:54produce a score right uh a score that

23:56you can get your coordinator to produce

23:59uh but also comp uh you can have a judge

24:01agent of a different model to uh kind

24:05cross reference and make sure that the

24:08first model is not being biased, right?

24:10The model that's judging all of the sub

24:12aents that are running the thing. Uh but

24:14you can also like hill climb. So meaning

24:16that you can you can use like /loop in

24:19cursor and you can say okay keep looping

24:22on this eval right until everything is

24:2510 out of 10 as an example. Uh and I did

24:28the same the basically the same approach

24:30with the control skill. And so I kind of

24:32it was very it was very hands-off

24:33actually. Uh so you know I uh I kind of

24:37built I built that skill that way like

24:39the CLI and that skill. Um and over time

24:43it's gotten really good. Uh but yeah it

24:46was definitely not super smooth at the

24:48beginning. It required a lot of

24:50iteration and I think there's an analogy

24:53here for me which is um well I make this

24:57analogy later in a different slide on my

24:59drawing here. Uh but I think of it like

25:03uh you know as a as a engineer now

25:06you're sort of more like you uh like

25:09maybe a manager or the analogy I like is

25:12like you're like a a chef in a

25:14restaurant. Uh you you're the head chef.

25:17Uh you're not cooking all the food

25:18yourself anymore. You have a team of

25:20cooks, right? You have line cooks, you

25:22have a sue chef, you have, you know, all

25:24these different stations.

25:26Um and it's your job to really design

25:28the environment. you know, you you're in

25:31charge of setting up the kitchen. You're

25:32in charge of, you know, like giving

25:35tasks to different people. So,

25:39um yeah, it's a very interesting way of

25:42working. Uh but yeah, that's that's how

25:45I've basically built uh these

25:47verification skills.

25:48>> Yeah, just just one thought there on

25:50like to go try to go one layer deeper.

25:52So, are you let's say we wanted to build

25:56um an eval or a skill for for something

25:59and we wanted to kind of get better on

26:01its own, which is is what I think you're

26:03suggesting. Uh are you doing that in

26:05like a work tree kind of isolated with

26:08like the sub agents and and then the

26:09reviewer agent and and all that? Is it

26:11happening like in some type of cloud

26:13hosted environment? Like what's the the

26:15more the practical steps? If I wanted to

26:17go do this uh and like set up a

26:19verification system for something, what

26:20would I what would I do or where would I

26:21start?

26:24Um I think that uh the best place to

26:27start is local because you can observe

26:31you can definitely observe what your

26:32agents are doing. So, uh, if you're

26:34building a verification skill for

26:36yourself, uh, I would definitely start

26:38local and just have your agent bring up

26:40the application, whether it's like a CLI

26:43or, uh, desktop app or whatever. And so,

26:46you can actually observe, right? You can

26:47see how the agent is interacting with

26:50the the application. You can see it, you

26:53know, how it calls like the different

26:56APIs that that allow it to interact with

26:58the uh the application.

27:02Um but uh for me personally uh I have

27:06basically been kind of all in mostly all

27:08in on cloud agents because they're

27:10extremely powerful. Uh and the really

27:13powerful thing about cursor is the the

27:15cloud agents actually where if you spend

27:18a little bit of time setting up your

27:19environment

27:21these control skills these verification

27:23skills pay a huge amount of dividends

27:26because it's not just something that

27:28makes you as a single engineer better.

27:31It actually levels up your whole team uh

27:33and even your whole company because uh

27:36you can actually start thinking about

27:38cloud agents. change that thing about

27:39automations that automatically do things

27:43like uh I I get I I kind of talk about

27:46this a bit later but I'll just kind of

27:49get into it. uh where where you know for

27:51example like I talk a lot about this

27:54agent we have called Benny right who who

27:57uh you know takes all of the bug reports

28:00that we get and it automatically goes

28:02off in the cloud opens up a cloud uh

28:04it's you know its desktop it runs cursor

28:08in its own computer and it uses the same

28:11control skills to interact with the

28:13application and try to reproduce the bug

28:15uh or the user report right and this is

28:18so so powerful because at once I can

28:20immediately I I get so much information

28:23from this automatically like here in

28:25this example you can see that uh the

28:27benny actually reproduced the bug uh but

28:30it's already fixed on main so it

28:33actually confirms that we fixed this

28:35problem already and all I need to do is

28:37just release another build of of cursor

28:40uh so that's like huge information there

28:43that I didn't have to go off and sit

28:44with an agent you know and spend an hour

28:46trying to figure like is this fixed is

28:48this not fixed

28:49So you you you gain back so much time.

28:52Uh but you know everybody on my team

28:54benefits from this. Everybody at the

28:55company benefits from this. Uh so

28:59definitely think that uh you know

29:01keeping these uh using cloud agents is

29:04super powerful. Uh but yeah it's like a

29:06journey. You have to trust it first

29:09right before you you get to this point.

29:11And that's it goes back to what I was

29:13saying here where you know it's very

29:15hard. It's almost impossible. And I

29:17would definitely encourage you not to

29:18try to jump from, you know, like if

29:21you're still in this zone, you don't

29:24want to jump to like I'm going to spawn

29:26hundred of thousand or thousands of

29:28cloud agents right now because you're

29:30just going to waste a lot of tokens. Um,

29:32and it's going to be extremely

29:34expensive.

29:35>> Yeah. So, just to kind of recap so far,

29:37basically the if we wanted to go on the

29:39journey that you've kind of gone on, it

29:40would be just start with verification.

29:43um building some some skills and some

29:45some ways of determining that the agents

29:47are producing at least like correct code

29:50whether like you said whether it's good

29:51code or not is maybe a separate question

29:52but like it's it's technically solving

29:54the problem by looking at you know stack

29:56traces looking at you know the the

29:58actual behavior in the app and so on um

30:00and then once we trust it locally then

30:02we can start to think about scaling into

30:04the cloud and running more agents that

30:06are picking up signals I guess on their

30:08own right so whether that's like a bug

30:09report that comes in or something they

30:11can go and pick it up and solve the

30:13problem and and give us back a PR and

30:15then maybe the last step is like

30:17automerging the PRs which uh is where

30:19you're at and maybe not where I want to

30:20go.

30:21>> Um and then reviewing them on main but

30:24um is that is that about right?

30:26>> Yeah, exactly. I think yeah that's why I

30:27drew this this uh this curve, right?

30:29Because that this this basically

30:31describes my journey of you know when I

30:34started barely could use a couple agents

30:36and I was just observing every single

30:38thing. I think there's really no

30:40shortcut for going from here to there

30:43because this is really about your

30:44personal level of trust in agents,

30:47right? Um obviously, you know, as a as

30:49engineer, you don't want to just slop

30:50code into production. So, how do you

30:53actually build up that trust takes um a

30:56lot of uh I guess taste and judgment. Um

30:59but uh you know, like I think plugins

31:02like Pstack definitely kind of help you

31:05uh get up to speed much quicker. Uh, and

31:08so I guess it's like if you trust me and

31:11you trust Pstack, then in by extension

31:14you can maybe trust your agents. But if

31:16you don't trust me and I I definitely

31:18would not encourage people to blindly

31:20trust me. Uh, uh, you know, if you build

31:24up your own set of skills that you can

31:26obviously, you know, take a look at PA

31:28and kind of fork it, make it your own,

31:30improve the skills. Definitely encourage

31:32that. Uh but for me it's really all

31:35about it just keeps coming back to

31:37trust. You know every one of us here in

31:39this chat have a different standard for

31:41engineering. Uh and there are different

31:43things that are important for us in our

31:46codebase. And uh when you are able to

31:50encode all of that into skills and you

31:51can verify that your agent is actually

31:53doing them that allows you to really

31:55kind of ascend this curve and um uh you

32:00know start automating things. Uh there's

32:02another piece I wanted to talk about. Um

32:05if there's more

32:06>> Yeah, go for it. I'll I'll pick up more

32:08questions as I go. But

32:09>> yeah, I think there's a third part to

32:10this which I haven't talked about yet

32:12which is kind of an interesting one

32:14which is like refactoring and rewriting

32:17like one of the uh I guess most

32:19controversial one of the most

32:21controversial topics in the industry I

32:23think is like should you rewrite your

32:26app or not? Um because I think engineers

32:30are very prone to this where especially

32:32when you join a company you come in and

32:34you see like the codebase and you're

32:35like oh man this is like who wrote

32:38this code you know it's terrible I want

32:40to rewrite the whole thing there is a

32:42very common inclination and I think a

32:44lot of you know before agents um and I

32:48guess arguably even now people will

32:50definitely discourage you from re

32:51rewriting stuff but I'm actually here to

32:54make a case for why you might want to

32:56consider it.

32:58Um because

33:00I think it really depends. Uh you know,

33:03uh brownfield applications I think are

33:05actually in a pretty good spot,

33:07especially if they're set up well

33:09already. Uh and like recently I've been

33:12talking to some people, but uh you know,

33:14I I was just observing I I just noticed

33:17this parallel, which is that a lot of

33:21big tech company problems are now

33:23everybody's problems. Um, and the big

33:26tech company problem, you know, like

33:27when I was working at Meta, like we had

33:29this giant monor repo, we had like, I

33:31don't know, tens of thousands of

33:33engineers just, you know, like banging

33:35on their keyboards and and shipping code

33:38and

33:40a lot of really great engineers at Meta.

33:42Uh, but, uh, I'll say like, you know,

33:45you'll be surprised that the code

33:46quality is actually not that good.

33:48[laughter] Um, and so I often joke that

33:51like, you know, before AI sloth, we had

33:53human sloth. Um and so uh you know I

33:56think a lot of big tech infra like uh

33:59like what Meta has or Google you know

34:01you know really big tech companies are

34:03actually designed for that where you

34:05you're sort of like you're catering to

34:07the the you know like uh this sounds so

34:10bad to say but like the the least

34:12capable engineer on your team right you

34:14build you build frameworks you build

34:16conventions you build guard rails you

34:19know you restrict credentials so that

34:21you know your intern doesn't wipe your

34:22production database

34:24Um

34:26there's uh you know if you have that

34:28level of infra already I think your

34:31agents can actually already do a very

34:33solid job right because they have the

34:35the guard rails are already in place for

34:38agents to not cause havoc or not cause

34:41too much havoc uh in your codebase um

34:44and you can always add more you know

34:46guardrails. Uh but I think like green

34:49field applications especially are you

34:50know like the brand new applications are

34:53like the biggest risk in my opinion. Uh

34:56and also the greatest opportunity

34:58because you know if you vibe code a

35:00project uh a prototype um like we did

35:04for Grockbot you know Grockbot was spun

35:05up very very very quickly. Um and if you

35:08if you haven't heard of of Grockbot it's

35:10like our a new application we just

35:12launched yesterday. Uh it's it's really

35:14cool. uh lets you orchestrate your

35:18create like individual agents that have

35:20their own identity and you can kind of

35:21orchestrate them. It's super cool.

35:23Definitely check it out. Um but yeah,

35:25that was it's like a very it was a very

35:27green field application like most

35:29prototypes are so like vibe coded very

35:32quickly. Humans were not reading the

35:34code at all. And uh I had this tweet

35:38recently uh where I said something about

35:41organic architecture. Um

35:45maybe I'll find it. Uh but the idea is

35:49that

35:50uh when you have a completely vibe coded

35:52application, you essentially have no

35:54guard rails whatsoever. So uh your

35:57agents

35:59when you give them a task, they will

36:01just solve it in whatever method is the

36:03most convenient. And over time you get

36:06into this uh situation where you have a

36:09code base that is spiraling out of

36:11control because you don't understand it.

36:14Uh your agents understand it I guess in

36:17a way but like they've built something

36:18that is you know optimized for short for

36:21shortcuts. Uh and uh you know it will

36:24you will suffer you'll have a lot of of

36:27issues with that application.

36:30Uh so I think starting your codebase

36:33with uh like very strong constraints is

36:37very much needed uh because like when

36:41you have a codebase that you can trust,

36:43right? when you have guardrails that

36:44actually help you uh uh help your agents

36:48write good code, you can get into the

36:50you know like into this part of the

36:52curve where I I where like I I said you

36:55know I woke up today and I had like 20

36:57PRs merged u by my agents and that's

37:00because I invested a lot a lot of time

37:04uh over 600 PRs I I I calculated

37:07yesterday uh when I refactored all of

37:10Grockbot to this new architecture that

37:12I've been

37:14Um,

37:15and yeah, I've gotten to a point where I

37:19I don't really look I really don't look

37:20at the code anymore. And um, I say that

37:24not just, you know, to sell you tokens,

37:25but because I, you know, it it it took a

37:29lot of work to get to that point. I

37:30spent a lot of tokens to get the

37:32codebase to this point where I no longer

37:34have to look at it. Uh but I'm very

37:37excited because you know of the

37:39potential where you know it's not just

37:41this doesn't just benefit me it benefits

37:43everyone contributing to Grothbot and it

37:46also empowers you know designers and

37:49product managers and you know pe uh even

37:52GTM people to add features to grabbot

37:55and I don't have to worry you know I

37:56don't have to to wake up at night in in

37:58the middle of the night and worry like

38:00oh someone's just merged a perf

38:01regression right I have a ton of

38:04constraints and CI is like it's actually

38:07very annoying to write code in in grabb

38:10but like agents absorb all of that

38:11annoyance.

38:13Um but yeah I'm happy to talk about what

38:16exactly that is. Um

38:18>> yeah I think one question um

38:21>> before we get into the this part here is

38:24just around that element of like what

38:26your your your CI looks like or maybe

38:28some of the constraints and then also

38:29like the average PR size. I saw a

38:31question about that earlier just to give

38:32people you know kind of a a glance. So

38:35it doesn't have to be like

38:35mathematically average but just uh you

38:37know like what generally the size of the

38:39a PR is. Um if it's only a couple lines

38:42of code or you know um yeah

38:45>> um

38:47I think it depends. Uh let me

38:51I'm trying to do this in a way where I'm

38:53not going to like

38:54>> you yeah you don't have to share the

38:55actual number like an actual average.

38:57This this is fine

38:58>> benchmark

38:59>> but like we have so okay this is not

39:03that interesting but uh well fun fact is

39:05that virtualization in Grockbot and in

39:09uh cursor is actually powered by uh

39:12pretext uh which is a sort of new

39:15library that someone's built um that's

39:19really interesting you should you should

39:20check it out but that's not really that

39:22important uh I think the average PR size

39:25I actually don't No, I pro I don't know

39:27if I want to click on these. Uh, I

39:29probably can, but I would say

39:31[clears throat] like they can range

39:33anywhere from a few hundred lines or 50

39:35lines to like a thousand depending on

39:38what the thing is doing. Uh, so like

39:41here I'm actually like deleting a bunch

39:42of files. So I expect that it's just

39:44like mostly deletion. Uh, but yeah, it

39:48kind of varies.

39:49>> There's no like Yeah,

39:51>> there's no like hard cap or hard limit.

39:52Are they're all like 50 line VR?

39:53>> There's no hard cap. Yeah, there's

39:54definitely no hard cap. But I I do

39:56encourage my agents to split up their

39:57work into multiple PRs. Uh I do that

40:01mostly because uh I like I like the idea

40:05of the I guess maybe this is much harder

40:07to do now as in the world of agents and

40:10you have like so many commits, but I

40:12like the idea that you know the git

40:13history is a very rich source of

40:15context. Uh, and I like the I like each

40:19PR to sort of atomically describe what

40:22that small piece of thing is doing,

40:25which also makes it easier for me to

40:26revert changes and like figure out, you

40:28know, oh, I shipped a bug and it's just

40:30it's here, right? It's not in this

40:3240,000 line PR where you who knows what

40:36landed in there.

40:38Uh, but I don't have a hard cap on PR

40:41size.

40:42>> Cool. And then um yeah, also quick

40:45question on like CI. So again, you don't

40:47have to go into like uh the screen share

40:48of like your CI does, but just generally

40:51would you describe what the CI kind of

40:53looks like uh or how strict it is?

40:56>> Uh yeah. So

40:58uh well specifically for Grockbot. So

41:01Dune is the is the sort of cheeky code

41:04code name for the architecture that

41:06we've built for Grockbot. Um the CI

41:10looks pretty annoying because there's

41:13checks for everything. So like literally

41:16I have um uh well if you've written any

41:19React for example you know you know that

41:21one of the biggest foot guns in React is

41:23use effect. Uh so in

41:26uh Dune and in graphbot we've banned use

41:29effect. So Dune is just you can the the

41:32mental model of what Dune is uh you can

41:34kind of think of it as like Nex.js JS

41:36for uh electron apps and it's designed

41:39for agents to write uh and it's like

41:42custom for you know our agent powered

41:45applications. Um so the CI checks are

41:48very like specific to that like you know

41:50don't use use effect. It's it's it's

41:52banned like CI will fail uh and yell at

41:55you. We have like some of the more

41:57interesting ones that people might raise

41:59eyebrows is like I actually ban code

42:01comments as well uh which is very

42:04interesting. Uh, but I've noticed that

42:0899% of the time agents just write code

42:11comments that kind of describe some

42:13historical thing that is actually

42:15totally irrelevant to the code. Um, like

42:18it will often say like, you know, oh,

42:19Lauren said you should never do this and

42:21it's now in in a code comment. I'm like,

42:23what? Like why what that was? I didn't

42:26say that as like a durable, you know,

42:27global rule. I just meant like your this

42:30PR sucks and you should change that

42:32part.

42:34uh agents don't really understand us

42:36that well surprisingly uh and or they

42:39kind of assume too much and they kind of

42:41do things in like very stupid ways. So

42:44like yeah we just ban everything

42:46everything you can imagine like the

42:48agents are bad at we ban. Uh so one

42:52example that we actually suffer a lot in

42:54the agents window is we have uh you know

42:57you know if you've used the agents

42:58window you've definitely seen

42:59performance issues and you know we're

43:01constantly trying to fix them. Uh but

43:04it's like a

43:05it's a never- ending struggle because

43:07there's so many pull requests that get

43:09merged. Every any one of them could just

43:11regress performance or stability or

43:13reliability. Uh you know the agents

43:16window doesn't have this architecture

43:17yet. I plan to do bring this learning

43:20back there and kind of refactor

43:22everything there. Uh but uh it just

43:25regresses super often. uh because uh

43:29there's just one example is like we have

43:31very poor um isolation between

43:34processes. So like on you know on on

43:36electron you have a renderer thread that

43:38renders your UI but you also have like a

43:40main thread that you can run other code

43:42that you know doesn't need to block the

43:44renderer.

43:46Um, but we do a poor job of separating

43:49those things and so often times you just

43:51accidentally have code that gets pulled

43:53into running on the renderer thread and

43:56then all of a sudden you're competing

43:57with the the renderer that you know that

44:00has a very if you want like 60 fps you

44:03have to every frame that gets drawn has

44:05to be done in 16 milliseconds. So very

44:08very small you know deadline per frame

44:11uh if you want you know a very smooth

44:13product. Uh and when you start building

44:15bringing in accidentally bringing in you

44:17know things that are like very

44:19computationally heavy or they have a lot

44:21of IO uh then you just get into like a

44:24lot of jank right your your FPS really

44:27drops you start uh you know losing

44:29frames you get long tasks that take more

44:32than 16 milliseconds and you just get

44:34this really choppy experience.

44:37So all of those patterns that we've

44:38learned basically building electron apps

44:40we've encoded into this framework and it

44:43becomes like a hard failure. So I

44:45literally in in grabbot we literally

44:47have a directory called electron main

44:50electron renderer and we have uh import

44:54uh CI guess where we actually check the

44:58dependency graph to make sure you're not

44:59accidentally importing code from one

45:02directory to another. Uh so that's

45:05enforced by CI um as well as bug bots uh

45:09which is our which cursors um like code

45:13review tool that runs on CI uh you know

45:15in our agents MD it's everywhere like so

45:18I I I I um I have this thing here where

45:22I I talk about like um you know like

45:25there are multiple layers I think for

45:27building a good codebase. Uh obviously

45:30the codebase is one where uh if you have

45:32an architecture like this where it's

45:35extremely strict uh you know the the the

45:38way to build features is very

45:39conventional that's like the strongest

45:42strongest level of enforcement because

45:44agents just love to copy existing

45:46patterns. So uh one example of this in

45:49rockbot is like we have this these

45:51concepts called like a feature and we

45:54have entry points and transcript cards

45:56like oh you know the cards that you see

45:57in the chat these are all like like

46:01nouns I guess in in the framework and so

46:03there's a very conventional way of

46:05creating them and so like a feature is

46:09all in in a single directory as an

46:10example and so all of the code that

46:13contributes to that feature lives in one

46:15directory so it's all coll-located in

46:17one place makes it super easy. You know,

46:19agents don't have to like uh grap around

46:21and try to figure out like where all the

46:23things are. It just looks at the feature

46:25and like, oh, okay, I'm working on the

46:27onboarding feature in Grockbot. Uh I'm

46:31just going to work in this directory.

46:32And for 80% of the work, it's mostly

46:35just very encapsulated there.

46:37[clears throat] But, uh like it's like

46:40very it's like designed again for you

46:42know like the dumbest agent like you

46:44don't have to think, right? the the the

46:47one of the key principles I have for

46:49this framework is like the shortest the

46:51shortest path is the best path.

46:55So uh because that plays exactly to how

46:57agents love to write code is like they

46:59like to take shortcuts really you know

47:01they they'll find the quickest way to

47:04solve the problem. So why not make that

47:06the best way to solve the problem? Uh so

47:10I I probably won't get into all the

47:12specific details. Um and uh the the this

47:16framework is really more of a collection

47:17of ideas and principles rather than

47:19something that will open source. Uh you

47:22can you can you know screenshot this I

47:23guess if you want and uh tell your agent

47:26to uh do some build build something like

47:29this for you too.

47:31Um yeah, but it's really all about the

47:34layers uh you know like the the codebase

47:36is one part with features uh and

47:39directories and you know import or

47:41blocking import dependencies uh that

47:44shouldn't be imported uh but and and it

47:47all enforces that and static analysis.

47:49So like uh there's CI checks, we have a

47:52lot of lints for bad patterns that we

47:55observe. Uh compiler diagnostics,

47:58uh there's also rules and bugbot which

48:01are um I think like three, four, five

48:04are more soft, right? These two actually

48:07make make CI red, right? So that you

48:11know there's a hard constraint where the

48:13agent can't just write crappy code

48:17for rules and skills in Bogbot. Your

48:20agents can still forget, right? You can

48:22still or it may not always consistently

48:25apply them. So I like to layer them, but

48:30I don't I don't like to rely on them as

48:32the only source of enforcement because

48:35it's very very soft, right? And if you

48:37if you only have rules and bug bot and

48:39skills and a style guide for your code,

48:41you will it's only a matter of time

48:43before your codebase looks like complete

48:46trash. I'm sorry to say that but uh I

48:49definitely recommend yeah like you know

48:50investing in you know things that can be

48:53hard and forced right and this is why

48:56you know maybe the choice of tech stack

48:59that you use is also very important. Um

49:02like I think for example Rust is sort of

49:04making you know it's like getting super

49:06popular again uh because the compiler is

49:10so strict right the compiler enforces so

49:13many different things you know there's a

49:14borrow checker that you have to appease

49:16and if as long as you make sure your

49:18agents don't write unsafe code blocks uh

49:21you can more or less feel somewhat

49:23confident that if the code compiles it

49:25probably works and it's good. Uh but you

49:28see it gives you that level of trust and

49:30confidence that you as a human engineer

49:34no longer need to go and check it

49:36yourself. You know you you rely on code

49:40and static analysis to actually make

49:43that uh a lot smoother. Um, and I I

49:48guess the worst part, the worst place to

49:50be in is if you are stuck in code review

49:53land where you actually enforce all of

49:55the constraints, the invariance in your

49:57codebase by literally the human person

50:01saying, you know, reading the code and

50:03like, okay, you should not do this,

50:04right?

50:06Every time you have to do that, you

50:07should consider that as a code smell,

50:08like a anti- pattern. And you should

50:11say, okay, instead of me commenting on

50:13the PR, how do I turn this into a hard

50:17rule, right? How do I turn this into a

50:18lint rule? How do I turn this into a CI

50:21failure? Or how do I even categorically

50:23eliminate this problem uh entirely? Uh I

50:28I can talk about another migration I've

50:30done, but I'll probably pause here.

50:32>> Sure.

50:33>> Yeah. I feel like that's that's where I

50:34am to be honest is is what you're

50:36describing right now which is that like

50:38I don't have all of these rules. So I

50:40have some things to go do after this

50:41session in terms of being able to scale

50:44my agents. I'm I'm definitely on like

50:45the uh you know maybe a couple of

50:47parallel ones locally stage. So like two

50:50to three locally and I'm sure most

50:51people here are on the same. So uh yeah

50:54I know we only couple minutes left.

50:56Lauren, was there anything else that you

50:57wanted to to highlight? I obviously

50:59there's lots of questions so I can grab

51:00more but I want to give you a few

51:02minutes if there's anything else you

51:03want to talk about.

51:03>> I think I've been yapping for quite a

51:05lot so I maybe let's just do questions.

51:08>> Okay, cool. Uh one question that had uh

51:10a couple of uh came up a couple times

51:12was just around like token usage.

51:15>> So the the question is like is is what

51:17you're describing a realistic thing for

51:19people who are on you know uh a normal

51:23set of token usage. They don't have you

51:25know basically unlimited tokens uh to

51:27work with.

51:29I think that's a really good point. I

51:30mean like obviously, you know, I work at

51:32a AI lab where we have unlimited tokens.

51:35So, uh I definitely cannot

51:39say that, you know, this is something

51:41everyone should do in the exact same way

51:43that I did it. I think it's possible to

51:45get to this point without, you know,

51:47breaking the bank.

51:49But you know if you're like an

51:50engineering leader or you know you're

51:52you you have a startup that you lead um

51:55I think to me it's a question of ROI um

51:58and it's like uh yes you spend a lot of

52:03money on tokens in the upfront stage you

52:06know like refactoring your code base is

52:07going to take a lot of tokens uh adding

52:09all these things uh is going to take a

52:11bunch of tokens but if we're heading to

52:14a world where agents are writing all the

52:16code and you know You want to be very

52:20lean, right? You don't want to have to

52:22hire, you don't want to be, you don't

52:23want to become like meta, right? Like I

52:25mean like in terms of you don't want to

52:26become a 10,000 person engineering org

52:29because I mean that's a cool problem to

52:32have, but also you you have so much

52:34overhead. There's like planning, you

52:37know, like you it's it's a personally I

52:39I wouldn't uh it it's not super fun, but

52:43um I think you want to stay very nimble,

52:46right? And you want to you want to be

52:47like agents are all about allowing you

52:50to do things that you couldn't do

52:51before. That's really to me like the

52:53value of agents, you know, it's not just

52:56storing tokens on every single little

52:58thing, but um to me like the thing I

53:00couldn't do before is like enforce this

53:03level of constraints in a codebase by

53:06myself, right? Like I'm just a single

53:09person, you know? Uh it would have taken

53:11me years to build this framework uh and

53:15do all the refactoring and test

53:18everything myself and verify you know

53:20like run imagine if there it was just me

53:22right no in in pre- agent era just like

53:25running you know by it would take me so

53:28long right and my salary is pretty high

53:30right like so you know the the question

53:34I think an engineering leader might have

53:35is just then you know like what is

53:38there's a trade-off of do to hire

53:40someone to do this or do you spend the

53:43tokens to set up a code base so that

53:45even the the most naive, right, the

53:48dumbest agents can do a good job. And

53:51when you actually get to this point,

53:53like even agents that are not, you know,

53:55fable size do an excellent job of

53:58writing code. And this pays a lot of

54:00dividends as well for me personally

54:02where I've empowered not just myself but

54:06again like PMs, designers, engineers who

54:10are not familiar with Grockbot to just

54:12contribute in a way that is sustainable.

54:16So I think yeah it's definitely like a

54:18trade-off for sure. You know like

54:20nothing is like free for sure. Uh and

54:22tokens are pretty expensive. Uh but oh

54:25actually uh I I I don't know how many of

54:27you have seen this but we actually

54:29announced Grock 4.6 today. So very

54:32exciting finally out. Um so yeah Gro 4.6

54:35would be like a great it was very very

54:37smart. Uh it's really good on the on the

54:40benchmarks. Uh and it's the same the

54:43tokens uh well uh I hopefully I'm not

54:45saying this incorrectly but uh I believe

54:48the cost per token is the same as 4.5.

54:51So you're actually getting more

54:53intelligence for the same cost. Uh I

54:57think this is an area that cursor tries

54:59to cursor and SpaceX AI try to really

55:02optimize for like that heredto frontier

55:05of you know cost versus intelligence. Uh

55:08you know we don't necessarily want to

55:10build the biggest model ever because

55:12that is extremely expensive to run. It's

55:14really about like how do you find that

55:16sweet spot right? you don't you don't

55:18need a giant model, but it's just super

55:19smart, right? And it's not very

55:21expensive for inference.

55:24Uh but um yeah, I think to kind of round

55:27it up, um I think it's like a it's it's

55:30there's a if you do your own analysis, I

55:33feel like it's pretty positive. It it'll

55:36be pretty positive that the ROI you get

55:38from investing in stuff like this uh

55:41just empowers not just yourself, but

55:44your whole team to be so much more

55:46productive, right? Right? Like imagine

55:47if you have an army of engineers like me

55:49who are shipping so much improvements

55:52and and bug fixes uh you know every day,

55:56right? Like that is pretty exciting.

56:00>> Cool. Uh one last question before we

56:01wrap up. This one is for the people in

56:03product on the on the call.

56:05>> So let's say we do have an army of

56:07engineers who are shipping like Lauren.

56:09I'm just curious like how is the product

56:11team or other functions of your company

56:13keeping up given that like if you're

56:15shipping so quickly have are they using

56:18AI more to do their jobs like as much as

56:20you can speak to that obviously you

56:22don't have like you're not in that role

56:23but just curious about how that works.

56:26Um I think this is where grabbot has

56:28been actually exceedingly powerful. Uh

56:31where so before grabbot like you know uh

56:34obviously cursor only had cursor like we

56:37only had agents window we had a CLI we

56:39had an IDE and these are really like

56:42power user tools right like de they're

56:44designed for developers so it's very

56:46very developerentric you can do

56:48knowledge work in them but it like the

56:50UI is not really optimized for that. So

56:54we actually didn't really have uh well I

56:57think like a lot of people like you know

56:58GTM product like they might have used

57:01cursor uh to do their work but it

57:04definitely wasn't like a delightful

57:05experience for them. Um I think now with

57:09Grogbot

57:10uh it's become Grogbot is basically like

57:13the Kusher moment for people who are not

57:16in tech in my opinion like it's like

57:18it's like a very very accessible way to

57:21use agents in a very comfortable very

57:24familiar interface. It looks like

57:25iMessage

57:27um and it's very fun to you know you can

57:29give your agent a fun name. uh you can

57:32have you can kind of do orchestration

57:34with in a very like natural way where

57:36you can sort of you know each agent is

57:38like a person right and now you got a

57:39team of agents like working on you have

57:41one one agent per account that you

57:43manage as an example or if you're a PM

57:45you have you know you can have an agent

57:47that summarizes all the work that Lauren

57:49did last night and then now you know

57:50what I did right so I think our PMs are

57:53leveraging that a lot and they're

57:56shipping code too uh so you know like

57:58often times they will just say oh here's

58:00a bug I fixed can you look at it and

58:02then I'll go review it and actually it's

58:04just perfect. I'm like okay stamp. Uh so

58:07uh that I think that shows that you know

58:09the the Dune architecture is holding up

58:11right the all the the really strict

58:14constraints allow people who are not

58:16experts in engineering to contribute at

58:19a high level. Uh so I'm I feel like I'm

58:21already seeing that pay off a lot where

58:24uh you know designers and PMs are just

58:26able to to to ship features directly.

58:30Um, and that just makes the Grockbot

58:33team super fast, right? Where we can

58:36ship so quickly. Um, and we have a lot

58:40planned, so I'm very excited uh to, you

58:43know, uh to to ship more ship more

58:46stuff.

58:47>> Yeah, that's awesome. We are at time, so

58:50uh I guess Lauren, if if folks want to

58:52support you, maybe go try out Grockbot,

58:54try out uh 46 and uh you know, get get

58:57provide some feedback. But yeah, this

58:59was awesome. really appreciate you

59:00taking the time. Uh thanks everyone for

59:02all the messages in the chat. Lots of

59:03good questions. I know we didn't get

59:04through everything, but as I kind of

59:06said at the top, way more questions than

59:07than we could get through, but uh yeah,

59:09really really thanks thanks for for

59:11joining. Thanks everyone for joining and

59:13hopefully you enjoyed the session.

59:15>> Yep.

59:16>> All right.

59:16>> Yeah, I see see thanks for having me and

59:18uh if you have any more questions yet,

59:20just DM me on Twitter. I'll I'll open

59:22them up. I guess I'll let the

59:24>> You're gonna get a lot of DMs.

59:26>> I'll open the floodgates. So yeah, DM

59:28me. Maybe I'll do like a Twitter space

59:30at some point as well for more

59:32questions. But really appreciate

59:34everyone for showing up. Uh, you know,

59:35taking an hour out of your day.

59:37>> Yeah. All right. Thanks all. I'll see

59:38you the next one.

59:39>> Okay. Thanks everyone. Bye.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.