Full transcript
0:00Everyone, I am Lauren Laurent Tan. I
0:03guess not many people know my last name.
0:05Uh, I am Potato on Twitter. Uh, potato
0:10with spelled with an E. Um, and I have
0:13been at Cursor for about 5 months. Uh,
0:18previously I was at Meta where I worked
0:20on the React team. Uh, specifically
0:22working on the React compiler, uh, which
0:25was a whole lot of fun. Uh I'm still on
0:27the on the core team and and uh
0:29contributing to open source here and
0:30there. Uh so that that's really nice
0:33that they still let me do that. Uh and
0:36before Meta, I was at Netflix uh where I
0:39was uh both a tech lead and uh I
0:43transitioned to be be an engineering
0:45manager um for about two years. So I've
0:50had a I've have had a lot of experience
0:52going between engineering management and
0:55being an individual contributor. Uh and
0:58I think something I've noticed actually
0:59which is quite interesting is that there
1:02are so many parallels with you know
1:04management skills and how to like manage
1:06agents. Uh and that's actually a big
1:09part about what I wanted to chat with
1:10you and everybody else about today. Um
1:14but yeah that's that's me. Uh I do have
1:17some like very light slides but uh it's
1:20not going to be uh just rambling. So let
1:23me just share my screen
1:25and hope that I don't leak anything. Uh
1:33oh no I need to allow permissions.
1:36>> No worries. Take your time. There's
1:38always tech tech trouble. Uh give me one
1:41second to rejoin.
1:42>> Yeah, go for it.
1:46I see many of you already know Lauren
1:48from the looks of the chat here. Um, so
1:51yeah, it's exciting to to get a chance
1:53to chat with her uh and and go through
1:55some of her her recent work. As you guys
1:57heard, you know, a lot of recent
1:59experience from from Netflix to to Meta
2:02and then now over at Cursor. Uh, we're
2:04going to chat a little bit about
2:05Grockbot as well. So, uh, that'll be
2:08exciting. I don't know if you guys saw
2:09that was a recent release. I think
2:10literally maybe yesterday or the day
2:12before uh from the cursor team which is
2:14kind of like um let's call it like
2:16agents for everyone. You can go check it
2:18out if you want and learn a little bit
2:20more about the product but um but yeah
2:22we'll we'll explore that a little bit
2:23today as well.
2:26All righty. Welcome back.
2:40>> And Lauren, you're just on mute there if
2:42uh you want to hop off of mute if you're
2:43chatting.
2:44>> Yeah, sorry.
2:45>> No worries.
2:47>> It's 2026 and I still don't know how to
2:49use Zoom. [laughter]
2:50>> That's all good.
2:51>> Uh okay. So, I assume you can see my
2:53screen.
2:54>> Yes. Yeah, we're good. Um, so yeah,
2:56today, yeah, I think I think the big
2:58theme for me as I've been using agents
3:01to write code, and I'm sure a lot of you
3:03have had the same experience as well, is
3:06how do you trust it? You know,
3:07especially if you are an engineer that's
3:10been writing code for a very long time,
3:12you have a lot of opinions and lessons
3:15that you've learned about doing good
3:18engineering. And when you see agents
3:20just, you know, winging it and, you
3:22know, guessing, hallucinating,
3:24uh, you know, confidently stating that
3:27they found the smoking gun, uh, for the
3:29hundth time, uh, but it's actually not
3:32the real problem. You lose a lot of
3:34trust. And when you lose when you don't
3:36have much trust in your agents,
3:39I feel like you you really can't get the
3:41most out of them. And for me, the
3:43parallel is like with management. Uh so
3:46if I'm an a manager an engineering
3:48manager of a team and I have a bunch of
3:51you know I have a team of engineers uh
3:54on my team and I don't trust them then
3:57the mode of operation I'm going to be in
3:59is going to be like micromanagement
4:00right I'll have to spend a lot of time
4:03looking over my reports shoulders and
4:06checking that they're doing their work
4:08well you know that they're not shipping
4:09bugs to production
4:13and so I drew this chart cuz uh it's not
4:16it's not a very scientific chart but
4:18like this is how I imagine
4:21myself and my journey through using
4:23agents. So you know like fast forward or
4:27back forward uh or fast back uh fast
4:32backwards like a year or so when you
4:34know nobody was or not many people were
4:36using agents to code. Uh I think you
4:41uh you know get into this mode where you
4:43are
4:45in very heavily in the loop with one or
4:48several like a handful of agents and you
4:51find yourself just constantly figure uh
4:53you know trying to understand what your
4:55agents are doing uh and you're very very
4:57in loop. You're watching every single
4:59output. you are sitting there prompting
5:03um and you really can't parallelize
5:05beyond that because you don't again you
5:08don't have that trust right you can't go
5:09to a 100 agents uh like spawn 100 agents
5:13when you don't even trust the output of
5:15one agent
5:17so over the past 5 months I feel like
5:20I've really been able to uh like ascend
5:23this trust curve and now I'm at the
5:26point where uh I actually have this
5:29Sounds kind of scary to say this and it
5:31makes me sound like a slop artist, but I
5:33I promise I'm not, but I actually have
5:36my agents now um automerging PRs for me.
5:39Uh which is like a wild thing to say,
5:42but um like I woke up today and there
5:44were like 20 PRs landed and I just
5:47reviewed them on Maine like they were
5:49already landed and they were good. Uh so
5:52how did I get to that point is basically
5:55what I wanted to talk about today.
5:58Uh and again like yeah feel free to jump
6:00in if you have questions Colin. Um but
6:04uh oh yeah of course I got to show this
6:07this chart. Uh where uh
6:12oh do do not trust to someone requested
6:16to control my computer. Uh probably
6:18won't do that. Uh but yeah, so this
6:21chart I think I I'm I'm sharing this
6:24chart not to kind of like flex but to
6:26kind of show like the journey like so
6:29you can see like the curve like it sort
6:30of like inversely matches the
6:33contributions I've been able to land at
6:35cursor. So I joined five months ago and
6:38five months ago like I you know my first
6:40month I was like not very productive
6:42because I was obvious you know I was
6:43learning the codebase didn't know what
6:45the heck was going on and as I got more
6:48confident in in my agents uh I've really
6:52been able to kind of ramp up my
6:54productivity. Uh and again like yeah
6:56like last month I shipped a thousand PRs
7:00which is ridiculous. Uh, and then this
7:02month we're only on the 12th. I'm
7:05already at like almost 800 PRs landed.
7:09Uh, so the velocity is definitely high
7:12and you you I'm sure a lot of you will
7:14definitely be questioning like how how
7:16much of this code is actually good. Um,
7:18and I think yeah, like that's definitely
7:21fair to question.
7:23Um, but uh, yeah, I think I think if you
7:28set up your agents well, you can
7:30definitely get to a very similar level.
7:34Um, and so I'm going to talk about how
7:36we do that.
7:39Uh so for me I think I'm curious like I
7:42guess call in your experience as well
7:44but uh for me I think the most important
7:48skill that you should have in your
7:51toolbox when you work with agents is
7:53verification.
7:55Uh and by verification I mean the
7:57ability for an agent to actually run the
8:00code uh or take CPU traces or heap
8:05snapshots or uh you know open an iOS
8:09simulator whatever you know however your
8:12application is exposed to your users it
8:15can do the same thing and uh run it for
8:19real and actually test and verify it
8:21don't work because that's the thing that
8:24really closes the loop. Uh it doesn't
8:26guarantee your agent writes good code.
8:29Uh but it allows them to at least write
8:31correct code. Uh which is a big a really
8:34big step forward for being able to trust
8:37your agent. Um
8:40I will I can share one example that we
8:43have uh within cursor. Uh
8:49oops
8:50where let me open this.
9:00Let's make this uh make me full screen.
9:03There you go. Uh so for the for cursors
9:08agent window uh so this is actually an
9:11interesting story but uh when I joined
9:14cursor 5 months ago uh they're actually
9:18uh well I was supposed to join a
9:19different team. I was supposed to join
9:20like the cloud agents team. Uh but then
9:23since I have a lot of experience working
9:25on react and agents window is a react
9:28application. Uh I was
9:31uh I was asked to basically help out
9:34with uh the agent window work. Uh but um
9:41there wasn't really a lot of like skills
9:43to help me. So I just found myself like
9:45okay uh agents which is going to launch
9:47in like a week right we have a really
9:49tight deadline and um there was uh you
9:53know I was just sitting there like okay
9:55I'm going to open up the performant the
9:57the chrome dev tools and just like take
9:59a trace look at it myself and try to
10:01make sense of this flame graph and keep
10:03in mind I was just like in my first week
10:05so I had no idea what I was looking at
10:07no idea where you know I mean I had some
10:09idea but you know the the code base was
10:11completely fresh to me Uh, and I
10:14realized like my agent had no idea
10:17either, you know, like I would take a
10:18screenshot of the trailer, I would
10:19download trades, I would send it to it,
10:20and it be like, "Yeah, it kind of looks
10:22like this, you know, uh, and it would
10:24like confidently state like it's this
10:26thing." And then I try to fix that and
10:28turns out that's not the actual thing.
10:32So, this was very very slow process. And
10:34if you've ever done any like performance
10:36work yourself or you know just even
10:38development with an agent where you
10:40don't have a verification skill, you are
10:43the verifier, right? You you're the
10:45bottleneck. You you you tell your agent
10:47to do something and then it goes off and
10:49write some code. Then you open up your
10:51you know local dev build and then you
10:53start to say oh you know doesn't work.
10:54Then you got to copy paste screen uh you
10:56know screenshots or console errors or
10:59whatever. uh and then your agent like
11:01slowly kind of like uh you know works
11:04with that and then tries to understand
11:06it and um fix the thing but then you're
11:10constantly just in the loop and and
11:12being a bottleneck. So there's really no
11:13way to parallelize. So the control glass
11:16skill is like one of the first skills I
11:18built uh for cursor. Uh and glass by the
11:21way is the code name for agents window
11:24that we use internally but it's just
11:27cursor I guess. Um, and so this skill uh
11:31is I guess the the the code itself is
11:33not super interesting. Your agent can
11:35very easily make one for you. Uh where
11:38if if you're building an electron app or
11:40a web app or even iOS uh applications,
11:45uh you can teach your agent how to use
11:47like the ChromeDev Tools protocol or
11:50through uh Apple has some utilities as
11:53well for running the simulator and
11:55taking traces and controlling
11:56programmatic control.
11:58as well. Uh so that's really useful. Uh
12:03but one thing I want to talk about is uh
12:06the
12:08this thing
12:10uh where is the read me? Uh so this
12:14skill comes with this very unique
12:16feature called or not feature uh unique
12:19file called a feature map. And so the
12:23story then is like I built this skill
12:25and so now the agent was able to uh
12:28actually run the agent window and take
12:30traces and whatnot. Uh but it had no
12:34idea what what the agent's window was.
12:37So um you know like someone would say
12:40like oh the the left sidebar is like
12:43laggy or something like that or you know
12:45the right side the the PR tab is not
12:48working and the agent would just be like
12:50kind of flailing around. it would spend
12:51a lot of time trying to like look up the
12:53code and you know where is this feature?
12:54How do I actually get to it on the UI
12:57which made it basically completely
12:59useless. Uh you know like we would I
13:01would run this skill locally and you
13:05know it would spawn a dev build uh but
13:08then it just be turnurning like it just
13:10try to click here. It it wouldn't know
13:12how to get to things um and it was just
13:15an awful experience. Uh so who's putting
13:19arrows on my screen? Um so uh yeah this
13:24this feature map has been really useful
13:26uh because it teaches the agent how to
13:28get to all of the features that you
13:30have. Um and in PAC the plugin that I
13:34I've made uh if you search for PAC
13:38cursor on Google you you'll find it. Uh
13:41but there is a create verification skill
13:43in that plug-in where it actually helps
13:45you set up something like this for
13:47yourself. Um including the feature map.
13:50So it will actually explore the code and
13:52build up this initial feature map that
13:54tells your agent how to get to all of
13:56the different features that you have. Uh
13:59and this is extremely powerful because
14:01now like you have these user reports
14:03that come in uh you you can actually map
14:06even like a vague report or even a
14:08screenshot. So we have this uh
14:11internally at cursor where uh we have a
14:14slack channel where you know lots of
14:17people giving us feedback on the agents
14:19window and rockbot and whatnot. Uh, and
14:22often times the report is very bad, like
14:26very low quality, like someone just put
14:28very often we get like a screenshot like
14:30and then someone just says question mark
14:31question mark question mark like what is
14:33this [laughter] and you know like the
14:35without this your agent is like I have
14:37no clue, right? But with a feature map
14:40like this, it has a lot more context and
14:43understanding of how to actually
14:45navigate, how to get to all the
14:47different features. Uh so like you know
14:49example like I guess like the sidebar
14:51like what is the sidebar uh you know
14:54like all the different sub features that
14:56are present in it um like from the user
14:59point of view here's where how do I get
15:01to it all the different keyboard
15:03shortcuts
15:05uh even like the the what do you call it
15:07the DOM elements or yeah like the
15:10attributes that you use for selecting
15:12things through the CDP uh are all there.
15:17So uh again yeah this is like really
15:19really powerful uh for for agents
15:23>> uh and pstack ships uh that create
15:26verification skill but also a maintain
15:28verification skill uh so you can keep
15:30this up to date.
15:32>> Cool. Yeah, I was just going to ask how
15:34you created that. So do you mind sharing
15:35a little bit more about um that that
15:37process in the context of Pstack and
15:39maybe just what Pstack is for the folks
15:40who aren't familiar?
15:42Yeah. So, Pstack is pretty interesting
15:44because uh well, first of all, the name
15:47is kind of goofy. like the P the P in P
15:50sack is like potato potato stack because
15:54I um so uh uh there is a pretty uh
16:00famous person Gary Tan who is the CEO of
16:03Y Combinator and he's come up with this
16:05plugin called GStack uh Gary Stack and
16:10uh funnily enough we share the last name
16:12we have no relations uh but I thought it
16:15would be funny to kind of you know poke
16:16fun at Gary and make Pack back my
16:19version of of of of his plug-in. Uh but
16:23kind of tailor it to my own set of pract
16:26engineering practices.
16:29Uh but I honestly actually never set out
16:31to build PAC. Uh it just started with a
16:33bunch of skills, right? Like I started
16:35with that control glass skill and then I
16:37started with another skill like called
16:39how which I also noticed through like
16:42observing agents. Um, so like you know
16:46in the early days of me, you know,
16:47trying to climb this ladder, I was like
16:49super in the loop and I was basically
16:52nitpicking my agents to an extreme
16:53degree. I was uh I would tell it um you
16:58know this feature has stopped working.
16:59Here's a bug report like why isn't it
17:02working?
17:04And very often the agent would just like
17:06confidently state like oh it has to be
17:09this right it has to be this thing. And
17:11I noticed like when I looked at the
17:13actual tool calls, I noticed it wasn't
17:15actually reading the code that I thought
17:18should be affected. And that made me
17:20just extremely suspicious. And at that
17:22point, I was like, I'm not going to I
17:24can't trust any this agent anymore cuz
17:26it's just it's just completely
17:27hallucinating. And I think
17:31I think it's very easy to just you know
17:33like build up that distrust and not and
17:36kind of feel helpless like you know you
17:39don't know how to help your agents
17:40succeed. But like again I think the the
17:43the management analogy is super helpful
17:45because like imagine if you were a
17:47manager of an engineering team and you
17:49had an engineer on your team who was a
17:52really good coder. No business context
17:54whatsoever. you know they they just you
17:56just hired them and they they onboarded
17:58you know like 5 seconds ago. Uh and so
18:02how do you actually teach that person to
18:04be effective? So how you do that is
18:06through a skill. uh skill being just you
18:09know it's just markdown right but you
18:11know it encodes a lot of information
18:13instructions a lot of uh you can really
18:17draw out a lot of intelligence from an
18:20agent by well some people on Twitter
18:23call it like you know pull the agent to
18:25a different latent space which is kind
18:27of like a fancy way of just saying like
18:29since uh you LLMs are sort of like they
18:31predict the next token uh when you give
18:35it some high quality tokens uh to begin
18:37with then you know it it can kind of
18:39pattern match on like a higher space
18:42that's you know smarter
18:45um so that's like a very interesting
18:48model there but yeah I built Pstack very
18:50very incrementally uh so uh started with
18:54just really observing how agents you
18:56know all the fail different failure
18:58modes of of that agents were having and
19:01every time I saw that I just okay I'm
19:02just going to make that a skill right
19:04like stop hallucinating actually go and
19:07search up, look up the code, use a lot
19:09of sub agents, uh, and yeah, stop
19:12guessing.
19:16>> Yeah, that makes sense. One, one kind of
19:18followup question here, uh, both from
19:19myself and from a bunch of people in the
19:20chat. So, I guess it's two two parts.
19:23So, one is like how do you maintain
19:24these skills? So, like the product
19:26changes over time. Obviously, there's a
19:27lot of people shipping against the
19:28codebase. So, how do these skills get
19:31maintained? Uh and then second to that
19:33is like how do you know when your
19:34verification is is good enough? Uh like
19:37and you know you can trust that the ver
19:39verification loops that you've built are
19:41going to I guess you trust that the
19:43outputs uh when they're done.
19:47>> Uh yeah maybe I'll talk about um I think
19:50there are some related maybe I'll start
19:52with this one first. So like how do I
19:56maintain these skills?
19:59So, um, if you're not familiar with this
20:01concept, an eval [clears throat] is
20:03essentially like a way to, uh, well, I
20:06the mental model I have is like it's
20:07like a unit test for an agent. Um, and,
20:11uh, you can actually make your own eval.
20:14You don't need like a special framework
20:16for them. You can build you can you can
20:18build one depending on like, you know,
20:20how scientific and how rigorous you want
20:23to be. Uh, my screen is red.
20:26>> Yeah, there's a little button. Um,
20:28sorry. Do
20:29>> you mind like disabling the drawing or
20:31something? I I can't see my screen.
20:33>> Yeah, sorry. If you guys could not draw
20:35on the screen, that'd be great. But, um,
20:36there's a little button in the
20:38>> Is that a troll?
20:39>> Yeah, the little drop down.
20:42>> How do I clear?
20:44>> Yeah, you got it. Perfect.
20:49>> Yeah. So, eval
20:52test your skills basically. And actually
20:56in Pstack we ship uh under potato mode
20:59there's a playbook if you search for it
21:00called eval playbook. Um and it's uh
21:05uh it's like not it's actually pretty
21:07pretty rigorous the way it's done. Uh
21:10but um essentially what I do is I spawn
21:13a lot of different sub agents. I have
21:15like my main coordinator agent uh come
21:18up with a rubric for uh what I want the
21:22skill to do. Um, and then it spawns all
21:26these sub aents and it it creates
21:28individual directories for them uh which
21:31are cleverly named to not let the sub
21:35agent know that it's being evaluated
21:37because uh agents can actually tell and
21:40when they do they change their behavior.
21:42Uh, but it does a bunch of stuff like
21:44that to um essentially yeah like test
21:49whether or not the skill I'm making or
21:51changing is actually doing what I think
21:53it does. Um, and one of the really nice
21:56things about cursor is that we are we we
21:58support so many different models. So you
22:01can actually eval your skill across all
22:03sorts of different models um and you
22:06know get a sense of how well it performs
22:09across that different matrix. um
22:12especially for the models that you use.
22:15Uh so I do this a lot. Every time I I
22:17modify a skill, I will run one of these
22:20uh like the Eval playbook uh and make
22:22sure that you know it's actually leading
22:25to a result I want. Uh but I will say
22:28like
22:30maintaining skills is actually pretty
22:32hard. Uh it requires I think a lot of
22:34taste and observation. So, you kind of
22:37need to be very good at being a backseat
22:40driver, you know what I mean? Like, if
22:42you do pair, if you ever done pair
22:44programming, for example, uh, and you
22:46watch a co-orker code and you're just
22:48like, you you could probably do this
22:50better, you know, you could do, you
22:51know, like, why did you not do this,
22:52right? You you ask a lot of questions to
22:54your coworker. And it's kind of a
22:56similar thing here. You like you don't
22:57want to just be a passive observer of
22:59your agent. You want to be very in the
23:02driver seat in the initial stages when
23:04you're building up your own set of
23:05skills. uh you know obviously you can
23:07use something like PAC but if you're
23:09building your own set of skills it's
23:11very I think you know opening up the all
23:14the tool calls and like reading the code
23:16and reading all the uh the agent
23:19behavior and their thinking blocks is a
23:22really great way to see where they they
23:24fail right like what what you know where
23:28are they being done and then you can go
23:29and build a skill for that and then with
23:32verification how you trust it is it's I
23:35think it's also a very similar iteration
23:37loop uh where you know like I actually
23:39did the same process for verifying the
23:42verification skill where I actually get
23:46um so one thing that's interesting about
23:48eval is that you can sort of hill climb
23:51them meaning that uh your eval can
23:54produce a score right uh a score that
23:56you can get your coordinator to produce
23:59uh but also comp uh you can have a judge
24:01agent of a different model to uh kind
24:05cross reference and make sure that the
24:08first model is not being biased, right?
24:10The model that's judging all of the sub
24:12aents that are running the thing. Uh but
24:14you can also like hill climb. So meaning
24:16that you can you can use like /loop in
24:19cursor and you can say okay keep looping
24:22on this eval right until everything is
24:2510 out of 10 as an example. Uh and I did
24:28the same the basically the same approach
24:30with the control skill. And so I kind of
24:32it was very it was very hands-off
24:33actually. Uh so you know I uh I kind of
24:37built I built that skill that way like
24:39the CLI and that skill. Um and over time
24:43it's gotten really good. Uh but yeah it
24:46was definitely not super smooth at the
24:48beginning. It required a lot of
24:50iteration and I think there's an analogy
24:53here for me which is um well I make this
24:57analogy later in a different slide on my
24:59drawing here. Uh but I think of it like
25:03uh you know as a as a engineer now
25:06you're sort of more like you uh like
25:09maybe a manager or the analogy I like is
25:12like you're like a a chef in a
25:14restaurant. Uh you you're the head chef.
25:17Uh you're not cooking all the food
25:18yourself anymore. You have a team of
25:20cooks, right? You have line cooks, you
25:22have a sue chef, you have, you know, all
25:24these different stations.
25:26Um and it's your job to really design
25:28the environment. you know, you you're in
25:31charge of setting up the kitchen. You're
25:32in charge of, you know, like giving
25:35tasks to different people. So,
25:39um yeah, it's a very interesting way of
25:42working. Uh but yeah, that's that's how
25:45I've basically built uh these
25:47verification skills.
25:48>> Yeah, just just one thought there on
25:50like to go try to go one layer deeper.
25:52So, are you let's say we wanted to build
25:56um an eval or a skill for for something
25:59and we wanted to kind of get better on
26:01its own, which is is what I think you're
26:03suggesting. Uh are you doing that in
26:05like a work tree kind of isolated with
26:08like the sub agents and and then the
26:09reviewer agent and and all that? Is it
26:11happening like in some type of cloud
26:13hosted environment? Like what's the the
26:15more the practical steps? If I wanted to
26:17go do this uh and like set up a
26:19verification system for something, what
26:20would I what would I do or where would I
26:21start?
26:24Um I think that uh the best place to
26:27start is local because you can observe
26:31you can definitely observe what your
26:32agents are doing. So, uh, if you're
26:34building a verification skill for
26:36yourself, uh, I would definitely start
26:38local and just have your agent bring up
26:40the application, whether it's like a CLI
26:43or, uh, desktop app or whatever. And so,
26:46you can actually observe, right? You can
26:47see how the agent is interacting with
26:50the the application. You can see it, you
26:53know, how it calls like the different
26:56APIs that that allow it to interact with
26:58the uh the application.
27:02Um but uh for me personally uh I have
27:06basically been kind of all in mostly all
27:08in on cloud agents because they're
27:10extremely powerful. Uh and the really
27:13powerful thing about cursor is the the
27:15cloud agents actually where if you spend
27:18a little bit of time setting up your
27:19environment
27:21these control skills these verification
27:23skills pay a huge amount of dividends
27:26because it's not just something that
27:28makes you as a single engineer better.
27:31It actually levels up your whole team uh
27:33and even your whole company because uh
27:36you can actually start thinking about
27:38cloud agents. change that thing about
27:39automations that automatically do things
27:43like uh I I get I I kind of talk about
27:46this a bit later but I'll just kind of
27:49get into it. uh where where you know for
27:51example like I talk a lot about this
27:54agent we have called Benny right who who
27:57uh you know takes all of the bug reports
28:00that we get and it automatically goes
28:02off in the cloud opens up a cloud uh
28:04it's you know its desktop it runs cursor
28:08in its own computer and it uses the same
28:11control skills to interact with the
28:13application and try to reproduce the bug
28:15uh or the user report right and this is
28:18so so powerful because at once I can
28:20immediately I I get so much information
28:23from this automatically like here in
28:25this example you can see that uh the
28:27benny actually reproduced the bug uh but
28:30it's already fixed on main so it
28:33actually confirms that we fixed this
28:35problem already and all I need to do is
28:37just release another build of of cursor
28:40uh so that's like huge information there
28:43that I didn't have to go off and sit
28:44with an agent you know and spend an hour
28:46trying to figure like is this fixed is
28:48this not fixed
28:49So you you you gain back so much time.
28:52Uh but you know everybody on my team
28:54benefits from this. Everybody at the
28:55company benefits from this. Uh so
28:59definitely think that uh you know
29:01keeping these uh using cloud agents is
29:04super powerful. Uh but yeah it's like a
29:06journey. You have to trust it first
29:09right before you you get to this point.
29:11And that's it goes back to what I was
29:13saying here where you know it's very
29:15hard. It's almost impossible. And I
29:17would definitely encourage you not to
29:18try to jump from, you know, like if
29:21you're still in this zone, you don't
29:24want to jump to like I'm going to spawn
29:26hundred of thousand or thousands of
29:28cloud agents right now because you're
29:30just going to waste a lot of tokens. Um,
29:32and it's going to be extremely
29:34expensive.
29:35>> Yeah. So, just to kind of recap so far,
29:37basically the if we wanted to go on the
29:39journey that you've kind of gone on, it
29:40would be just start with verification.
29:43um building some some skills and some
29:45some ways of determining that the agents
29:47are producing at least like correct code
29:50whether like you said whether it's good
29:51code or not is maybe a separate question
29:52but like it's it's technically solving
29:54the problem by looking at you know stack
29:56traces looking at you know the the
29:58actual behavior in the app and so on um
30:00and then once we trust it locally then
30:02we can start to think about scaling into
30:04the cloud and running more agents that
30:06are picking up signals I guess on their
30:08own right so whether that's like a bug
30:09report that comes in or something they
30:11can go and pick it up and solve the
30:13problem and and give us back a PR and
30:15then maybe the last step is like
30:17automerging the PRs which uh is where
30:19you're at and maybe not where I want to
30:20go.
30:21>> Um and then reviewing them on main but
30:24um is that is that about right?
30:26>> Yeah, exactly. I think yeah that's why I
30:27drew this this uh this curve, right?
30:29Because that this this basically
30:31describes my journey of you know when I
30:34started barely could use a couple agents
30:36and I was just observing every single
30:38thing. I think there's really no
30:40shortcut for going from here to there
30:43because this is really about your
30:44personal level of trust in agents,
30:47right? Um obviously, you know, as a as
30:49engineer, you don't want to just slop
30:50code into production. So, how do you
30:53actually build up that trust takes um a
30:56lot of uh I guess taste and judgment. Um
30:59but uh you know, like I think plugins
31:02like Pstack definitely kind of help you
31:05uh get up to speed much quicker. Uh, and
31:08so I guess it's like if you trust me and
31:11you trust Pstack, then in by extension
31:14you can maybe trust your agents. But if
31:16you don't trust me and I I definitely
31:18would not encourage people to blindly
31:20trust me. Uh, uh, you know, if you build
31:24up your own set of skills that you can
31:26obviously, you know, take a look at PA
31:28and kind of fork it, make it your own,
31:30improve the skills. Definitely encourage
31:32that. Uh but for me it's really all
31:35about it just keeps coming back to
31:37trust. You know every one of us here in
31:39this chat have a different standard for
31:41engineering. Uh and there are different
31:43things that are important for us in our
31:46codebase. And uh when you are able to
31:50encode all of that into skills and you
31:51can verify that your agent is actually
31:53doing them that allows you to really
31:55kind of ascend this curve and um uh you
32:00know start automating things. Uh there's
32:02another piece I wanted to talk about. Um
32:05if there's more
32:06>> Yeah, go for it. I'll I'll pick up more
32:08questions as I go. But
32:09>> yeah, I think there's a third part to
32:10this which I haven't talked about yet
32:12which is kind of an interesting one
32:14which is like refactoring and rewriting
32:17like one of the uh I guess most
32:19controversial one of the most
32:21controversial topics in the industry I
32:23think is like should you rewrite your
32:26app or not? Um because I think engineers
32:30are very prone to this where especially
32:32when you join a company you come in and
32:34you see like the codebase and you're
32:35like oh man this is like who wrote
32:38this code you know it's terrible I want
32:40to rewrite the whole thing there is a
32:42very common inclination and I think a
32:44lot of you know before agents um and I
32:48guess arguably even now people will
32:50definitely discourage you from re
32:51rewriting stuff but I'm actually here to
32:54make a case for why you might want to
32:56consider it.
32:58Um because
33:00I think it really depends. Uh you know,
33:03uh brownfield applications I think are
33:05actually in a pretty good spot,
33:07especially if they're set up well
33:09already. Uh and like recently I've been
33:12talking to some people, but uh you know,
33:14I I was just observing I I just noticed
33:17this parallel, which is that a lot of
33:21big tech company problems are now
33:23everybody's problems. Um, and the big
33:26tech company problem, you know, like
33:27when I was working at Meta, like we had
33:29this giant monor repo, we had like, I
33:31don't know, tens of thousands of
33:33engineers just, you know, like banging
33:35on their keyboards and and shipping code
33:38and
33:40a lot of really great engineers at Meta.
33:42Uh, but, uh, I'll say like, you know,
33:45you'll be surprised that the code
33:46quality is actually not that good.
33:48[laughter] Um, and so I often joke that
33:51like, you know, before AI sloth, we had
33:53human sloth. Um and so uh you know I
33:56think a lot of big tech infra like uh
33:59like what Meta has or Google you know
34:01you know really big tech companies are
34:03actually designed for that where you
34:05you're sort of like you're catering to
34:07the the you know like uh this sounds so
34:10bad to say but like the the least
34:12capable engineer on your team right you
34:14build you build frameworks you build
34:16conventions you build guard rails you
34:19know you restrict credentials so that
34:21you know your intern doesn't wipe your
34:22production database
34:24Um
34:26there's uh you know if you have that
34:28level of infra already I think your
34:31agents can actually already do a very
34:33solid job right because they have the
34:35the guard rails are already in place for
34:38agents to not cause havoc or not cause
34:41too much havoc uh in your codebase um
34:44and you can always add more you know
34:46guardrails. Uh but I think like green
34:49field applications especially are you
34:50know like the brand new applications are
34:53like the biggest risk in my opinion. Uh
34:56and also the greatest opportunity
34:58because you know if you vibe code a
35:00project uh a prototype um like we did
35:04for Grockbot you know Grockbot was spun
35:05up very very very quickly. Um and if you
35:08if you haven't heard of of Grockbot it's
35:10like our a new application we just
35:12launched yesterday. Uh it's it's really
35:14cool. uh lets you orchestrate your
35:18create like individual agents that have
35:20their own identity and you can kind of
35:21orchestrate them. It's super cool.
35:23Definitely check it out. Um but yeah,
35:25that was it's like a very it was a very
35:27green field application like most
35:29prototypes are so like vibe coded very
35:32quickly. Humans were not reading the
35:34code at all. And uh I had this tweet
35:38recently uh where I said something about
35:41organic architecture. Um
35:45maybe I'll find it. Uh but the idea is
35:49that
35:50uh when you have a completely vibe coded
35:52application, you essentially have no
35:54guard rails whatsoever. So uh your
35:57agents
35:59when you give them a task, they will
36:01just solve it in whatever method is the
36:03most convenient. And over time you get
36:06into this uh situation where you have a
36:09code base that is spiraling out of
36:11control because you don't understand it.
36:14Uh your agents understand it I guess in
36:17a way but like they've built something
36:18that is you know optimized for short for
36:21shortcuts. Uh and uh you know it will
36:24you will suffer you'll have a lot of of
36:27issues with that application.
36:30Uh so I think starting your codebase
36:33with uh like very strong constraints is
36:37very much needed uh because like when
36:41you have a codebase that you can trust,
36:43right? when you have guardrails that
36:44actually help you uh uh help your agents
36:48write good code, you can get into the
36:50you know like into this part of the
36:52curve where I I where like I I said you
36:55know I woke up today and I had like 20
36:57PRs merged u by my agents and that's
37:00because I invested a lot a lot of time
37:04uh over 600 PRs I I I calculated
37:07yesterday uh when I refactored all of
37:10Grockbot to this new architecture that
37:12I've been
37:14Um,
37:15and yeah, I've gotten to a point where I
37:19I don't really look I really don't look
37:20at the code anymore. And um, I say that
37:24not just, you know, to sell you tokens,
37:25but because I, you know, it it it took a
37:29lot of work to get to that point. I
37:30spent a lot of tokens to get the
37:32codebase to this point where I no longer
37:34have to look at it. Uh but I'm very
37:37excited because you know of the
37:39potential where you know it's not just
37:41this doesn't just benefit me it benefits
37:43everyone contributing to Grothbot and it
37:46also empowers you know designers and
37:49product managers and you know pe uh even
37:52GTM people to add features to grabbot
37:55and I don't have to worry you know I
37:56don't have to to wake up at night in in
37:58the middle of the night and worry like
38:00oh someone's just merged a perf
38:01regression right I have a ton of
38:04constraints and CI is like it's actually
38:07very annoying to write code in in grabb
38:10but like agents absorb all of that
38:11annoyance.
38:13Um but yeah I'm happy to talk about what
38:16exactly that is. Um
38:18>> yeah I think one question um
38:21>> before we get into the this part here is
38:24just around that element of like what
38:26your your your CI looks like or maybe
38:28some of the constraints and then also
38:29like the average PR size. I saw a
38:31question about that earlier just to give
38:32people you know kind of a a glance. So
38:35it doesn't have to be like
38:35mathematically average but just uh you
38:37know like what generally the size of the
38:39a PR is. Um if it's only a couple lines
38:42of code or you know um yeah
38:45>> um
38:47I think it depends. Uh let me
38:51I'm trying to do this in a way where I'm
38:53not going to like
38:54>> you yeah you don't have to share the
38:55actual number like an actual average.
38:57This this is fine
38:58>> benchmark
38:59>> but like we have so okay this is not
39:03that interesting but uh well fun fact is
39:05that virtualization in Grockbot and in
39:09uh cursor is actually powered by uh
39:12pretext uh which is a sort of new
39:15library that someone's built um that's
39:19really interesting you should you should
39:20check it out but that's not really that
39:22important uh I think the average PR size
39:25I actually don't No, I pro I don't know
39:27if I want to click on these. Uh, I
39:29probably can, but I would say
39:31[clears throat] like they can range
39:33anywhere from a few hundred lines or 50
39:35lines to like a thousand depending on
39:38what the thing is doing. Uh, so like
39:41here I'm actually like deleting a bunch
39:42of files. So I expect that it's just
39:44like mostly deletion. Uh, but yeah, it
39:48kind of varies.
39:49>> There's no like Yeah,
39:51>> there's no like hard cap or hard limit.
39:52Are they're all like 50 line VR?
39:53>> There's no hard cap. Yeah, there's
39:54definitely no hard cap. But I I do
39:56encourage my agents to split up their
39:57work into multiple PRs. Uh I do that
40:01mostly because uh I like I like the idea
40:05of the I guess maybe this is much harder
40:07to do now as in the world of agents and
40:10you have like so many commits, but I
40:12like the idea that you know the git
40:13history is a very rich source of
40:15context. Uh, and I like the I like each
40:19PR to sort of atomically describe what
40:22that small piece of thing is doing,
40:25which also makes it easier for me to
40:26revert changes and like figure out, you
40:28know, oh, I shipped a bug and it's just
40:30it's here, right? It's not in this
40:3240,000 line PR where you who knows what
40:36landed in there.
40:38Uh, but I don't have a hard cap on PR
40:41size.
40:42>> Cool. And then um yeah, also quick
40:45question on like CI. So again, you don't
40:47have to go into like uh the screen share
40:48of like your CI does, but just generally
40:51would you describe what the CI kind of
40:53looks like uh or how strict it is?
40:56>> Uh yeah. So
40:58uh well specifically for Grockbot. So
41:01Dune is the is the sort of cheeky code
41:04code name for the architecture that
41:06we've built for Grockbot. Um the CI
41:10looks pretty annoying because there's
41:13checks for everything. So like literally
41:16I have um uh well if you've written any
41:19React for example you know you know that
41:21one of the biggest foot guns in React is
41:23use effect. Uh so in
41:26uh Dune and in graphbot we've banned use
41:29effect. So Dune is just you can the the
41:32mental model of what Dune is uh you can
41:34kind of think of it as like Nex.js JS
41:36for uh electron apps and it's designed
41:39for agents to write uh and it's like
41:42custom for you know our agent powered
41:45applications. Um so the CI checks are
41:48very like specific to that like you know
41:50don't use use effect. It's it's it's
41:52banned like CI will fail uh and yell at
41:55you. We have like some of the more
41:57interesting ones that people might raise
41:59eyebrows is like I actually ban code
42:01comments as well uh which is very
42:04interesting. Uh, but I've noticed that
42:0899% of the time agents just write code
42:11comments that kind of describe some
42:13historical thing that is actually
42:15totally irrelevant to the code. Um, like
42:18it will often say like, you know, oh,
42:19Lauren said you should never do this and
42:21it's now in in a code comment. I'm like,
42:23what? Like why what that was? I didn't
42:26say that as like a durable, you know,
42:27global rule. I just meant like your this
42:30PR sucks and you should change that
42:32part.
42:34uh agents don't really understand us
42:36that well surprisingly uh and or they
42:39kind of assume too much and they kind of
42:41do things in like very stupid ways. So
42:44like yeah we just ban everything
42:46everything you can imagine like the
42:48agents are bad at we ban. Uh so one
42:52example that we actually suffer a lot in
42:54the agents window is we have uh you know
42:57you know if you've used the agents
42:58window you've definitely seen
42:59performance issues and you know we're
43:01constantly trying to fix them. Uh but
43:04it's like a
43:05it's a never- ending struggle because
43:07there's so many pull requests that get
43:09merged. Every any one of them could just
43:11regress performance or stability or
43:13reliability. Uh you know the agents
43:16window doesn't have this architecture
43:17yet. I plan to do bring this learning
43:20back there and kind of refactor
43:22everything there. Uh but uh it just
43:25regresses super often. uh because uh
43:29there's just one example is like we have
43:31very poor um isolation between
43:34processes. So like on you know on on
43:36electron you have a renderer thread that
43:38renders your UI but you also have like a
43:40main thread that you can run other code
43:42that you know doesn't need to block the
43:44renderer.
43:46Um, but we do a poor job of separating
43:49those things and so often times you just
43:51accidentally have code that gets pulled
43:53into running on the renderer thread and
43:56then all of a sudden you're competing
43:57with the the renderer that you know that
44:00has a very if you want like 60 fps you
44:03have to every frame that gets drawn has
44:05to be done in 16 milliseconds. So very
44:08very small you know deadline per frame
44:11uh if you want you know a very smooth
44:13product. Uh and when you start building
44:15bringing in accidentally bringing in you
44:17know things that are like very
44:19computationally heavy or they have a lot
44:21of IO uh then you just get into like a
44:24lot of jank right your your FPS really
44:27drops you start uh you know losing
44:29frames you get long tasks that take more
44:32than 16 milliseconds and you just get
44:34this really choppy experience.
44:37So all of those patterns that we've
44:38learned basically building electron apps
44:40we've encoded into this framework and it
44:43becomes like a hard failure. So I
44:45literally in in grabbot we literally
44:47have a directory called electron main
44:50electron renderer and we have uh import
44:54uh CI guess where we actually check the
44:58dependency graph to make sure you're not
44:59accidentally importing code from one
45:02directory to another. Uh so that's
45:05enforced by CI um as well as bug bots uh
45:09which is our which cursors um like code
45:13review tool that runs on CI uh you know
45:15in our agents MD it's everywhere like so
45:18I I I I um I have this thing here where
45:22I I talk about like um you know like
45:25there are multiple layers I think for
45:27building a good codebase. Uh obviously
45:30the codebase is one where uh if you have
45:32an architecture like this where it's
45:35extremely strict uh you know the the the
45:38way to build features is very
45:39conventional that's like the strongest
45:42strongest level of enforcement because
45:44agents just love to copy existing
45:46patterns. So uh one example of this in
45:49rockbot is like we have this these
45:51concepts called like a feature and we
45:54have entry points and transcript cards
45:56like oh you know the cards that you see
45:57in the chat these are all like like
46:01nouns I guess in in the framework and so
46:03there's a very conventional way of
46:05creating them and so like a feature is
46:09all in in a single directory as an
46:10example and so all of the code that
46:13contributes to that feature lives in one
46:15directory so it's all coll-located in
46:17one place makes it super easy. You know,
46:19agents don't have to like uh grap around
46:21and try to figure out like where all the
46:23things are. It just looks at the feature
46:25and like, oh, okay, I'm working on the
46:27onboarding feature in Grockbot. Uh I'm
46:31just going to work in this directory.
46:32And for 80% of the work, it's mostly
46:35just very encapsulated there.
46:37[clears throat] But, uh like it's like
46:40very it's like designed again for you
46:42know like the dumbest agent like you
46:44don't have to think, right? the the the
46:47one of the key principles I have for
46:49this framework is like the shortest the
46:51shortest path is the best path.
46:55So uh because that plays exactly to how
46:57agents love to write code is like they
46:59like to take shortcuts really you know
47:01they they'll find the quickest way to
47:04solve the problem. So why not make that
47:06the best way to solve the problem? Uh so
47:10I I probably won't get into all the
47:12specific details. Um and uh the the this
47:16framework is really more of a collection
47:17of ideas and principles rather than
47:19something that will open source. Uh you
47:22can you can you know screenshot this I
47:23guess if you want and uh tell your agent
47:26to uh do some build build something like
47:29this for you too.
47:31Um yeah, but it's really all about the
47:34layers uh you know like the the codebase
47:36is one part with features uh and
47:39directories and you know import or
47:41blocking import dependencies uh that
47:44shouldn't be imported uh but and and it
47:47all enforces that and static analysis.
47:49So like uh there's CI checks, we have a
47:52lot of lints for bad patterns that we
47:55observe. Uh compiler diagnostics,
47:58uh there's also rules and bugbot which
48:01are um I think like three, four, five
48:04are more soft, right? These two actually
48:07make make CI red, right? So that you
48:11know there's a hard constraint where the
48:13agent can't just write crappy code
48:17for rules and skills in Bogbot. Your
48:20agents can still forget, right? You can
48:22still or it may not always consistently
48:25apply them. So I like to layer them, but
48:30I don't I don't like to rely on them as
48:32the only source of enforcement because
48:35it's very very soft, right? And if you
48:37if you only have rules and bug bot and
48:39skills and a style guide for your code,
48:41you will it's only a matter of time
48:43before your codebase looks like complete
48:46trash. I'm sorry to say that but uh I
48:49definitely recommend yeah like you know
48:50investing in you know things that can be
48:53hard and forced right and this is why
48:56you know maybe the choice of tech stack
48:59that you use is also very important. Um
49:02like I think for example Rust is sort of
49:04making you know it's like getting super
49:06popular again uh because the compiler is
49:10so strict right the compiler enforces so
49:13many different things you know there's a
49:14borrow checker that you have to appease
49:16and if as long as you make sure your
49:18agents don't write unsafe code blocks uh
49:21you can more or less feel somewhat
49:23confident that if the code compiles it
49:25probably works and it's good. Uh but you
49:28see it gives you that level of trust and
49:30confidence that you as a human engineer
49:34no longer need to go and check it
49:36yourself. You know you you rely on code
49:40and static analysis to actually make
49:43that uh a lot smoother. Um, and I I
49:48guess the worst part, the worst place to
49:50be in is if you are stuck in code review
49:53land where you actually enforce all of
49:55the constraints, the invariance in your
49:57codebase by literally the human person
50:01saying, you know, reading the code and
50:03like, okay, you should not do this,
50:04right?
50:06Every time you have to do that, you
50:07should consider that as a code smell,
50:08like a anti- pattern. And you should
50:11say, okay, instead of me commenting on
50:13the PR, how do I turn this into a hard
50:17rule, right? How do I turn this into a
50:18lint rule? How do I turn this into a CI
50:21failure? Or how do I even categorically
50:23eliminate this problem uh entirely? Uh I
50:28I can talk about another migration I've
50:30done, but I'll probably pause here.
50:32>> Sure.
50:33>> Yeah. I feel like that's that's where I
50:34am to be honest is is what you're
50:36describing right now which is that like
50:38I don't have all of these rules. So I
50:40have some things to go do after this
50:41session in terms of being able to scale
50:44my agents. I'm I'm definitely on like
50:45the uh you know maybe a couple of
50:47parallel ones locally stage. So like two
50:50to three locally and I'm sure most
50:51people here are on the same. So uh yeah
50:54I know we only couple minutes left.
50:56Lauren, was there anything else that you
50:57wanted to to highlight? I obviously
50:59there's lots of questions so I can grab
51:00more but I want to give you a few
51:02minutes if there's anything else you
51:03want to talk about.
51:03>> I think I've been yapping for quite a
51:05lot so I maybe let's just do questions.
51:08>> Okay, cool. Uh one question that had uh
51:10a couple of uh came up a couple times
51:12was just around like token usage.
51:15>> So the the question is like is is what
51:17you're describing a realistic thing for
51:19people who are on you know uh a normal
51:23set of token usage. They don't have you
51:25know basically unlimited tokens uh to
51:27work with.
51:29I think that's a really good point. I
51:30mean like obviously, you know, I work at
51:32a AI lab where we have unlimited tokens.
51:35So, uh I definitely cannot
51:39say that, you know, this is something
51:41everyone should do in the exact same way
51:43that I did it. I think it's possible to
51:45get to this point without, you know,
51:47breaking the bank.
51:49But you know if you're like an
51:50engineering leader or you know you're
51:52you you have a startup that you lead um
51:55I think to me it's a question of ROI um
51:58and it's like uh yes you spend a lot of
52:03money on tokens in the upfront stage you
52:06know like refactoring your code base is
52:07going to take a lot of tokens uh adding
52:09all these things uh is going to take a
52:11bunch of tokens but if we're heading to
52:14a world where agents are writing all the
52:16code and you know You want to be very
52:20lean, right? You don't want to have to
52:22hire, you don't want to be, you don't
52:23want to become like meta, right? Like I
52:25mean like in terms of you don't want to
52:26become a 10,000 person engineering org
52:29because I mean that's a cool problem to
52:32have, but also you you have so much
52:34overhead. There's like planning, you
52:37know, like you it's it's a personally I
52:39I wouldn't uh it it's not super fun, but
52:43um I think you want to stay very nimble,
52:46right? And you want to you want to be
52:47like agents are all about allowing you
52:50to do things that you couldn't do
52:51before. That's really to me like the
52:53value of agents, you know, it's not just
52:56storing tokens on every single little
52:58thing, but um to me like the thing I
53:00couldn't do before is like enforce this
53:03level of constraints in a codebase by
53:06myself, right? Like I'm just a single
53:09person, you know? Uh it would have taken
53:11me years to build this framework uh and
53:15do all the refactoring and test
53:18everything myself and verify you know
53:20like run imagine if there it was just me
53:22right no in in pre- agent era just like
53:25running you know by it would take me so
53:28long right and my salary is pretty high
53:30right like so you know the the question
53:34I think an engineering leader might have
53:35is just then you know like what is
53:38there's a trade-off of do to hire
53:40someone to do this or do you spend the
53:43tokens to set up a code base so that
53:45even the the most naive, right, the
53:48dumbest agents can do a good job. And
53:51when you actually get to this point,
53:53like even agents that are not, you know,
53:55fable size do an excellent job of
53:58writing code. And this pays a lot of
54:00dividends as well for me personally
54:02where I've empowered not just myself but
54:06again like PMs, designers, engineers who
54:10are not familiar with Grockbot to just
54:12contribute in a way that is sustainable.
54:16So I think yeah it's definitely like a
54:18trade-off for sure. You know like
54:20nothing is like free for sure. Uh and
54:22tokens are pretty expensive. Uh but oh
54:25actually uh I I I don't know how many of
54:27you have seen this but we actually
54:29announced Grock 4.6 today. So very
54:32exciting finally out. Um so yeah Gro 4.6
54:35would be like a great it was very very
54:37smart. Uh it's really good on the on the
54:40benchmarks. Uh and it's the same the
54:43tokens uh well uh I hopefully I'm not
54:45saying this incorrectly but uh I believe
54:48the cost per token is the same as 4.5.
54:51So you're actually getting more
54:53intelligence for the same cost. Uh I
54:57think this is an area that cursor tries
54:59to cursor and SpaceX AI try to really
55:02optimize for like that heredto frontier
55:05of you know cost versus intelligence. Uh
55:08you know we don't necessarily want to
55:10build the biggest model ever because
55:12that is extremely expensive to run. It's
55:14really about like how do you find that
55:16sweet spot right? you don't you don't
55:18need a giant model, but it's just super
55:19smart, right? And it's not very
55:21expensive for inference.
55:24Uh but um yeah, I think to kind of round
55:27it up, um I think it's like a it's it's
55:30there's a if you do your own analysis, I
55:33feel like it's pretty positive. It it'll
55:36be pretty positive that the ROI you get
55:38from investing in stuff like this uh
55:41just empowers not just yourself, but
55:44your whole team to be so much more
55:46productive, right? Right? Like imagine
55:47if you have an army of engineers like me
55:49who are shipping so much improvements
55:52and and bug fixes uh you know every day,
55:56right? Like that is pretty exciting.
56:00>> Cool. Uh one last question before we
56:01wrap up. This one is for the people in
56:03product on the on the call.
56:05>> So let's say we do have an army of
56:07engineers who are shipping like Lauren.
56:09I'm just curious like how is the product
56:11team or other functions of your company
56:13keeping up given that like if you're
56:15shipping so quickly have are they using
56:18AI more to do their jobs like as much as
56:20you can speak to that obviously you
56:22don't have like you're not in that role
56:23but just curious about how that works.
56:26Um I think this is where grabbot has
56:28been actually exceedingly powerful. Uh
56:31where so before grabbot like you know uh
56:34obviously cursor only had cursor like we
56:37only had agents window we had a CLI we
56:39had an IDE and these are really like
56:42power user tools right like de they're
56:44designed for developers so it's very
56:46very developerentric you can do
56:48knowledge work in them but it like the
56:50UI is not really optimized for that. So
56:54we actually didn't really have uh well I
56:57think like a lot of people like you know
56:58GTM product like they might have used
57:01cursor uh to do their work but it
57:04definitely wasn't like a delightful
57:05experience for them. Um I think now with
57:09Grogbot
57:10uh it's become Grogbot is basically like
57:13the Kusher moment for people who are not
57:16in tech in my opinion like it's like
57:18it's like a very very accessible way to
57:21use agents in a very comfortable very
57:24familiar interface. It looks like
57:25iMessage
57:27um and it's very fun to you know you can
57:29give your agent a fun name. uh you can
57:32have you can kind of do orchestration
57:34with in a very like natural way where
57:36you can sort of you know each agent is
57:38like a person right and now you got a
57:39team of agents like working on you have
57:41one one agent per account that you
57:43manage as an example or if you're a PM
57:45you have you know you can have an agent
57:47that summarizes all the work that Lauren
57:49did last night and then now you know
57:50what I did right so I think our PMs are
57:53leveraging that a lot and they're
57:56shipping code too uh so you know like
57:58often times they will just say oh here's
58:00a bug I fixed can you look at it and
58:02then I'll go review it and actually it's
58:04just perfect. I'm like okay stamp. Uh so
58:07uh that I think that shows that you know
58:09the the Dune architecture is holding up
58:11right the all the the really strict
58:14constraints allow people who are not
58:16experts in engineering to contribute at
58:19a high level. Uh so I'm I feel like I'm
58:21already seeing that pay off a lot where
58:24uh you know designers and PMs are just
58:26able to to to ship features directly.
58:30Um, and that just makes the Grockbot
58:33team super fast, right? Where we can
58:36ship so quickly. Um, and we have a lot
58:40planned, so I'm very excited uh to, you
58:43know, uh to to ship more ship more
58:46stuff.
58:47>> Yeah, that's awesome. We are at time, so
58:50uh I guess Lauren, if if folks want to
58:52support you, maybe go try out Grockbot,
58:54try out uh 46 and uh you know, get get
58:57provide some feedback. But yeah, this
58:59was awesome. really appreciate you
59:00taking the time. Uh thanks everyone for
59:02all the messages in the chat. Lots of
59:03good questions. I know we didn't get
59:04through everything, but as I kind of
59:06said at the top, way more questions than
59:07than we could get through, but uh yeah,
59:09really really thanks thanks for for
59:11joining. Thanks everyone for joining and
59:13hopefully you enjoyed the session.
59:15>> Yep.
59:16>> All right.
59:16>> Yeah, I see see thanks for having me and
59:18uh if you have any more questions yet,
59:20just DM me on Twitter. I'll I'll open
59:22them up. I guess I'll let the
59:24>> You're gonna get a lot of DMs.
59:26>> I'll open the floodgates. So yeah, DM
59:28me. Maybe I'll do like a Twitter space
59:30at some point as well for more
59:32questions. But really appreciate
59:34everyone for showing up. Uh, you know,
59:35taking an hour out of your day.
59:37>> Yeah. All right. Thanks all. I'll see
59:38you the next one.
59:39>> Okay. Thanks everyone. Bye.