Full transcript
0:00Everyone, I'm Lauren, Lauren Tan. I
0:03guess not many people know my last name.
0:05Uh I am potato on Twitter.
0:08Uh
0:09potato with spelled with an E.
0:11Um and I have been at Cursor for about 5
0:16months.
0:17Uh previously I was at Meta where I
0:19worked on the React team, uh
0:22specifically working on the React
0:23compiler, uh which was a whole lot of
0:25fun. Uh I'm still on the on the core
0:28team and and uh contributing to open
0:30source here and there. Uh
0:32so that that's really nice that they
0:34still let me do that. Uh and before Meta
0:36I was at Netflix, uh where I was uh both
0:40a tech lead and uh I transitioned to be
0:44be an engineering manager
0:45um
0:46for about 2 years.
0:49So I've had a I've had a lot of
0:51experience going between engineering
0:54management and being an individual
0:55contributor. Uh
0:57and I think something I've noticed
0:59actually, which is quite interesting, is
1:01that there are so many parallels with
1:03you know, management skills and how to
1:05like manage agents.
1:07Uh and that's actually a a big part
1:09about what I wanted to chat with you and
1:11everybody else about today.
1:13Um but yeah, that's that's me. Uh I do
1:16have some like very light slides, but uh
1:20it's not going to be um just rambling.
1:22So let me just share my screen.
1:25And hope that I don't leak anything.
1:28Uh
1:33Oh no, I need to allow permissions.
1:36>> No worries. All good, take your time.
1:38>> There's always tech tech tech trouble.
1:40Uh give me 1 second to rejoin.
1:42>> Yeah, go for it.
1:46I see many of you already know Lauren
1:48from the looks of the chat here. Um so
1:51yeah, it's exciting to to get a chance
1:53to chat with her uh and and go through
1:55some of her her recent work. As you guys
1:57heard, you know, a lot of recent
1:58experience from from Netflix to to Meta
2:02and then now over at Cursor.
2:04Uh we're going to chat a little bit
2:05about Grok Bot as well. So, uh that'll
2:08be exciting. I don't know if you guys
2:09saw that was a recent release. I think
2:10literally maybe yesterday or the day
2:12before
2:13uh from the Cursor team, which is kind
2:14of like um let's call it like agents for
2:17everyone. You can go check it out if you
2:18want and learn a little bit more about
2:20the product. But um but yeah, we'll
2:22we'll explore that a little bit today as
2:23well.
2:26All righty, welcome back.
2:40And Lauren, you're just on mute there if
2:42uh you want to hop off and mute and
2:43we'll start chatting.
2:44>> Yeah, sorry.
2:45>> No worries.
2:47>> It's 2026 and I still don't know how to
2:48use Zoom.
2:49>> [laughter]
2:50>> That's all good.
2:51>> Uh okay, so I assume you can see my
2:53screen.
2:54>> Yes, yeah.
2:56>> So yeah, today
2:57yeah, I think I think the big theme for
2:59me as I've been using agents to write
3:01code, and I'm sure a lot of you have had
3:03the same experience as well, is
3:06how do you trust it? You know,
3:07especially if you are an engineer that's
3:09been
3:10writing code for a very long time, you
3:12have
3:13a lot of opinions and lessons that
3:16you've learned about doing good
3:17engineering.
3:19And when you see agents just, you know,
3:21winging it and, you know, guessing,
3:23hallucinating, uh you know, confidently
3:26stating that they found the smoking gun
3:29uh for the hundredth time, uh but it's
3:31actually not the real problem, you lose
3:33a lot of trust. And when you lose when
3:36you don't have much trust in your
3:37agents,
3:39I feel like you you really can't get the
3:40most out of them.
3:42And for me, the parallel is like with
3:44management.
3:45Uh so, if I'm an a manager, an
3:47engineering manager of a team,
3:49and I have a bunch of, you know, I have
3:51a team of engineers
3:53uh on my team and I don't trust them,
3:56then the mode of operation I'm going to
3:58be in is going to be like
4:00micromanagement, right? I'll have to
4:02spend a lot of time looking over my
4:04reports' shoulders and checking that
4:07they're doing their work well, you know,
4:08that they're not shipping bugs to
4:10production.
4:13And so, I drew this chart because
4:15uh it's not it's not a very scientific
4:17chart, but
4:18like this is how I imagine
4:21myself and my journey through using
4:23agents.
4:24So, you know, like fast forward or back
4:27forward
4:28uh
4:29or fast back
4:31uh fast backwards like a year or so
4:34when, you know, nobody was or not many
4:36people were using agents to code. Uh
4:38I think you
4:40uh you get into this mode where you are
4:45in very heavily in the loop with one or
4:47several like a handful of agents.
4:50And you find yourself just constantly
4:53uh you know, trying to understand what
4:54your agents are doing
4:56uh and you're very very in the loop.
4:58You're watching every single output. You
5:00are sitting there prompting
5:02um and you really can't parallelize
5:05beyond that because you don't again you
5:07don't have that trust, right? You can't
5:09go to 100 agents
5:11uh like spawn 100 agents when you don't
5:13even trust the output of one agent.
5:17So, over the past 5 months I feel like
5:19I've really been able to uh like ascend
5:23this
5:24trust curve and now I'm at the point
5:27where
5:28uh I actually have This sounds kind of
5:30scary to say this and it it makes me
5:31sound like a slot artist, but I I
5:33promise I'm not.
5:35But I actually have my agents now um
5:37auto merging PRs for me.
5:39Uh which is like a wild thing to say,
5:42but um like I woke up today and there
5:44were like 20 PRs landed and I just
5:46reviewed them on main like they were
5:49already landed.
5:50And they were good.
5:51Uh so how did I get to that point? It's
5:54basically what I wanted to talk about
5:56today.
5:58Uh and again like yeah, feel free to
6:00jump in if you have questions, Colin. Um
6:03but uh
6:05Oh yeah, of course I got to show this
6:07this chart.
6:08Uh where uh
6:12Oh do not trust to read Someone
6:15requested to control my computer. Uh
6:18probably won't do that. Uh but yeah, so
6:21this chart I think
6:22I I I'm sharing this chart not to kind
6:24of like flex but to kind of show like
6:27the journey. Like so you can see like
6:29the curve like it sort of like inversely
6:31matches the contributions I've been able
6:34to land at cursor.
6:36So I joined 5 months ago. And 5 months
6:38ago like I you know, my first month I
6:40was like not very productive cuz I was
6:43you know, I was learning the code base.
6:45Didn't know what the heck was going on.
6:47And as I got more confident in in my
6:50agents uh I've really been able to kind
6:52of ramp up my productivity.
6:55Uh and again like yeah, like last month
6:57I shipped 1,000 PRs which is ridiculous.
7:01Uh and then this month we're only on the
7:0412th, I'm already at like almost 800 PRs
7:07landed.
7:08Uh so the velocity is definitely high
7:12and you you I'm I'm sure a lot of you
7:14will definitely be questioning like how
7:16how much of this code is actually good.
7:18Um and I think yeah, like that's
7:20definitely fair to question.
7:23Um but uh yeah, I think
7:26I think if you
7:27set up your agents well, you can
7:30definitely get to a very similar level.
7:34Um and so I'm going to talk about how we
7:35do that.
7:38Uh so for me I think I'm curious like I
7:42guess calling your experience as well,
7:44but uh for me
7:46I think the most important skill that
7:49you should have in your toolbox when you
7:52work with agents is verification.
7:55Uh and by verification I mean the
7:57ability for an agent to actually
8:00run the code
8:01uh or take CPU traces or heap snapshots
8:06or
8:07uh
8:07you know, open an iOS simulator,
8:10whatever you know, however your
8:11application is exposed to your users
8:15it can do the same thing
8:17and uh run it for real and actually test
8:20and verify it'll work.
8:22Because that's the thing that really
8:24closes the loop. Uh it doesn't guarantee
8:26your agent writes good code,
8:29uh but it allows them to at least write
8:31correct code. Uh which is a big a really
8:34big step forward for being able to trust
8:37your agent.
8:38Um
8:40I will I can share one example that we
8:43have uh within
8:45cursor.
8:47Uh
8:48oops
8:50Where let me open this.
8:59Let me just make this bigger.
9:01Uh let me full screen.
9:03There you go.
9:04Uh
9:05so for the for cursors agent window, uh
9:09so this is actually an interesting
9:11story, but uh when I joined
9:14cursor 5 months ago
9:15uh they're actually uh
9:18well, I was supposed to join a different
9:19team. I was supposed to join like the
9:21cloud agents team. Uh but then since I
9:23have a lot of experience working on
9:25React and agents window is a React
9:28application uh I was
9:31uh
9:32I was asked to basically help out with
9:34uh uh, the
9:36agent window work.
9:37Uh, but,
9:39um,
9:40there wasn't really a lot of like skills
9:43to help me. So, I just found myself
9:45like, okay, uh, agent window is going to
9:46launch in like a week. All right, we
9:48have a really tight deadline.
9:50And,
9:51um,
9:52there was, uh, you know, I was just
9:54sitting there like, okay, I'm going to
9:55open up the performance the the Chrome
9:57DevTools and just like take a trace,
9:59look at it myself, and try to make sense
10:01of this flame graph. And keep in mind, I
10:03was just like in my first week. So, I
10:05had no idea what I was looking at. No
10:07idea what, you know, I mean, I had some
10:09idea, but, you know, the the code base
10:11was completely fresh to me.
10:13Uh, and I realized like my agent had no
10:16idea either, you know, like I would take
10:18a screenshot of the trial download
10:19trace, I'll send it to him and it'd be
10:20like, yeah, it kind of looks like this,
10:23you know,
10:24uh, and it would like confidently state
10:25like it's this thing, and then I try to
10:27fix that, and it turns out that's not
10:29the actual thing.
10:31So, this was very very slow process. And
10:34if you've ever done any like performance
10:36work yourself, or, you know, just even
10:38development with an agent where you
10:39don't have a verification skill,
10:42you are the verifier. All right, you
10:44you're the bottleneck. You you you tell
10:46your agent to do something, and then it
10:48goes off and writes some code, then you
10:50open up your, you know, local dev build,
10:52and then you start to say, oh, you know,
10:53that doesn't work. Then you got to copy
10:55paste screen, uh, you know, screenshots
10:57or console errors or whatever,
11:00uh, and then your agent like slowly kind
11:02of like uh, you know, works with that,
11:05and then tries to understand it, and,
11:07um,
11:08uh, fix the thing. But then you're
11:10constantly just in the loop and and
11:12being a bottleneck. So, there's really
11:13no way to parallelize.
11:15So, the control glass goes like one of
11:17the first skills I built, uh, for
11:19Cursor.
11:20Uh, and glass, by the way, is the code
11:22name for agent window that we use
11:24internally, but it's just
11:27Cursor, I guess.
11:28Um, and so Well, skill, uh, is I guess
11:32that that the code itself is not super
11:34interesting. Your agent can very easily
11:36make one for you. Uh where if if you're
11:39building an Electron app or a web app or
11:41even iOS uh applications, uh you can
11:45teach your agent how to use like the
11:47Chrome DevTools protocol or through
11:51uh Apple has some utilities as well for
11:54running the simulator and taking traces
11:56and controlling programmatic control
11:58as well.
12:00Uh so, that's really useful.
12:03Uh but one thing I want actually want to
12:04talk about is
12:06uh the
12:08this thing.
12:10Uh
12:11where is the read me?
12:13Uh so, this skill comes with this very
12:15unique feature called or not feature yet
12:18uh unique file called a feature map.
12:21And so, the story then is like I built
12:24this skill and so now the agent was able
12:26to
12:27uh actually run the agent window and
12:30take traces and whatnot.
12:32Uh
12:33but it had no idea what
12:35what the agent's window was. So, um you
12:38know, like someone say like oh, the the
12:41left side bar is like laggy or something
12:44like that or you know, the right side
12:46the the PR tab is not working. And the
12:49agent would just be like kind of
12:50flailing around. It would spend a lot of
12:51time trying to like look up the code and
12:53you know, where is this feature? How do
12:55I actually get to it on the UI? Which
12:57made it basically completely useless. Uh
13:00you know, like we would I would run
13:03this skill locally
13:04and you know, it would spawn a dev
13:06build.
13:07Uh but then it just be churning. Like it
13:10would just try to click here. It
13:11wouldn't know how to get to things.
13:14Um and it was just an awful experience.
13:16Uh so, who was putting arrows on my
13:19screen?
13:20Um
13:21so,
13:22uh
13:23yeah, this this feature map has has
13:25really useful
13:26uh because it teaches the agent how to
13:28get to all of the features that you
13:30have.
13:31Um and in P-SAC, the plugin that I I've
13:34made,
13:35uh if you
13:36search for P-SAC cursor on Google,
13:39you'll find it.
13:40Uh but there is a creative verification
13:43skill in that plugin where it actually
13:45helps you set up something like this for
13:47yourself.
13:48Um including the feature map, so it will
13:50actually explore the code and build up
13:52this initial feature map that tells your
13:54agent how to get to all of the different
13:57features that you have.
13:58Uh and this is extremely powerful
14:00because now that you have these user
14:02reports that come in,
14:04uh you you can actually map even like a
14:07vague report or even a screenshot.
14:10So we have this
14:11uh internally at Cursor where uh we have
14:14a Slack channel with, you know, lots of
14:17people giving us feedback on the agent's
14:19window and rock bar and whatnot.
14:21Uh and oftentimes the report is very
14:24bad. Like very low quality, like someone
14:27just put
14:28very often we get like a screenshot like
14:30and then someone just says question mark
14:31question mark question mark, like what
14:32is this?
14:33And you know, like the without this,
14:35your agent says, "I have no clue."
14:38Right? But with a feature map like this,
14:40it has a lot more context and
14:43understanding of how to actually
14:45navigate, how to get to all of the
14:47different features. Uh so like, you
14:49know, example like, I guess like the
14:51sidebar, like what is the sidebar?
14:53Uh you know, like all the different sub
14:55features that are present in it. Um like
14:58from the user point of view, here's
15:00where to how to get to it, all the
15:02different keyboard shortcuts.
15:05Uh even like the
15:06the what do you call it? The DOM
15:08elements or yeah, like the attributes
15:11that you use for selecting things
15:13through the CDP
15:15uh are all there. So uh again, you know,
15:18this is like really really powerful
15:20uh for for agents. And uh and a piece
15:24that ships uh that create verification
15:26skill but also a maintain verification
15:28skill.
15:29Uh so you can keep this up to date.
15:32>> Cool. Yeah, I was just going to ask how
15:34you created that. So do you mind sharing
15:35a little bit more about um
15:37that that process in the context of P
15:38stack and maybe just what P stack is for
15:40the folks who aren't familiar?
15:42>> Yeah, so P stack is pretty interesting
15:44because uh
15:46well, first of all, the name is kind of
15:47goofy. Like the P the P in P stack is
15:51like potato potato sack.
15:53Uh because I um
15:55So uh the uh there is a pretty
15:59uh famous person Garry Tan who is the
16:02CEO of Y Combinator and he's come up
16:04with this plugin called G stack. Uh
16:07Garry stack and uh
16:10funnily enough, we share the last name.
16:12We have no relation.
16:13Uh but I thought it'd be funny to kind
16:16of, you know, poke fun at Garry and make
16:18P stack my version of of of
16:21>> [laughter]
16:21>> of his plugin.
16:22Uh but kind of just tailor it to my own
16:25set of prac engineering practices.
16:29Uh but I honestly actually never set out
16:30to build P stack. Uh it just started
16:33with a bunch of skills. All right, like
16:34I started with that control glass skill.
16:37And then I started with another skill
16:38like called howl which I also noticed
16:41through like observing agents.
16:44Um so like, you know, in the early days
16:46of me, you know, trying to climb this
16:48ladder, I was like super in the loop.
16:50And I was basically nitpicking my agents
16:52to an extreme degree. I was uh like I
16:55would tell it um
16:57you know, this feature has stopped
16:59working. Here's a bug report. Like why
17:01isn't it working?
17:04And very often the agent would just like
17:06confidently state like, "Oh, it has to
17:08be this, right? It has to be this
17:10thing."
17:11And I noticed like when I looked at the
17:13actual tool calls, I noticed it wasn't
17:15actually reading the code
17:17that I thought should be affected.
17:20And that made me just extremely
17:21suspicious. And at that point I was just
17:23like, I'm not going to I can't trust any
17:24this agent anymore because it's just
17:26it's just completely hallucinating.
17:29And I think
17:31I think it's very easy to just, you
17:33know, like build up that distrust and
17:35not and kind of feel
17:37helpless. Like, you know, you don't know
17:39how to help your agents succeed.
17:41But like again, I think the the the
17:43management analogies super helpful
17:45because like imagine if you were a
17:46manager of an engineering team and you
17:49had an engineer in your team who was a
17:51really good coder, no business context
17:54whatsoever, you know, they they just you
17:56just hired them and they they onboarded,
17:58you know, like uh 5 seconds ago.
18:00Uh and so how do you actually teach that
18:03person to be effective?
18:05So how you do that is through skill. Uh
18:08skill being just, you know, it's just
18:09markdown, right? But you know, it it
18:11codes a lot of information,
18:13instructions, a lot of
18:15uh you can really draw out a lot of
18:19intelligence from an agent by well, some
18:22people on Twitter call it like, you
18:23know, pull the agent to a different
18:25latent space. Which is kind of like a
18:27fancy way of just saying, like since uh
18:29you know, LLMs are sort of like they
18:31predict the next token.
18:33Uh when you give it some high-quality
18:36tokens uh to begin with, then, you know,
18:38it it can kind of pattern match on like
18:40a higher space that's, you know,
18:42smarter.
18:44Um so that's like a very interesting
18:48model there. But yeah, I built Pstack
18:50very, very incrementally. Uh so uh
18:53started with just really observing how
18:55agents, you know, all the failed
18:57different failure modes of of that
18:59agents were having. And every time I saw
19:01that, I just, okay, I'm just going to
19:02make that a skill. All right, like stop
19:04hallucinating, actually go and search up
19:07look up the code, use a lot of
19:09subagents, uh and yeah, stop guessing.
19:16>> Yeah, that makes sense. One one kind of
19:18follow-up question here, both for myself
19:19and for a bunch of people in the chat.
19:20So,
19:21I guess it's two two parts. So, one is
19:23like, how do you maintain these skills?
19:25So, like the product changes over time.
19:27Obviously, there's a lot of people who
19:27are shipping against the code base.
19:29So, how do these skills get maintained?
19:32And then second to that is like, how do
19:33you know when your verification is is
19:35good enough? Like in you know, you can
19:38trust that the very verification loops
19:40that you've built are going to
19:41I guess you trust that the outputs when
19:43they're done.
19:47>> Uh, yeah, maybe I'll talk about um,
19:50I think I have something really down.
19:51Maybe I'll start with this one first.
19:53So, like how do I
19:55maintain these skills?
19:58So, um,
20:00if you're not familiar with this
20:01concept, an eval [clears throat] is
20:02essentially like a way to
20:05Uh, well, I think the mental model I
20:07have is like it's like a unit test for
20:08an agent.
20:09Um, and
20:11you can actually make your own evals.
20:14You don't need like a special framework
20:15for them. You can build you can you can
20:18build one depending on like, you know,
20:20how scientific and how rigorous you want
20:22to be.
20:23Uh, my screen is red.
20:26>> Yeah, there's a little button. Um,
20:28sorry.
20:29>> Are you going to like disabling the
20:30drawing or something? Like I can't see
20:32my screen.
20:33>> Yeah, sorry. If you guys could not draw
20:35on the screen, that'd be great. But, um,
20:36there's a little button in the
20:38>> troll.
20:39>> Yeah, the the little drop down.
20:41>> Um,
20:42how do I clear?
20:44>> Yeah.
20:45>> Okay, yeah.
20:46>> You got it. Perfect.
20:47>> Um,
20:47>> Continue.
20:49>> Yeah, so evals are a way to unit test
20:52your skills, basically.
20:55And actually in P stack, we ship
20:57under potato mode, there's a playbook.
20:59If you search for it, called eval
21:01playbook.
21:02Um, and it's
21:03uh,
21:05Uh, it's like not It's actually pretty
21:07pretty rigorous the way it's done. Uh,
21:10But essentially what I do is I spawn a
21:13lot of different sub agents. I have like
21:15my main coordinator agent
21:18come up with a rubric for
21:21what I want the skill to do. Um
21:25and then it spawns all these sub agents
21:26and it it creates individual directories
21:29for them
21:31which are cleverly named to not let the
21:34sub agent know that it's being evaluated
21:37because agents can actually tell and
21:40when they do they change their behavior.
21:42Uh but it does a bunch of stuff like
21:44that to
21:46um
21:46essentially yeah like test whether or
21:49not the skill I'm making or changing is
21:52actually doing what I think it does. Um
21:55and one of the really nice things about
21:56cursor is that we are we have we support
21:59so many different models. So you can
22:01actually eval your skill across all
22:03sorts of different models. Um and you
22:06know get a sense of how well it performs
22:08across that different matrix. Um
22:12especially for the models that you use.
22:15Uh so I do this a lot. Every time I I
22:17modify a skill I will run one of these
22:20like the eval playbook
22:22and make sure that you know it's
22:24actually leading to a result I want. Uh
22:27but I will say like
22:29maintaining skills is actually pretty
22:31hard. It requires I think a lot of
22:34taste and observation. So you kind of
22:37need to be very good at being a backseat
22:40driver. You know what I mean? Like if
22:42you do pair if you've ever done pair
22:44programming for example
22:46and you watch a coworker code and you
22:48just like you could probably do this
22:49better or you know you could do you know
22:51do you like why did you not do this?
22:52Right? You you ask a lot of questions to
22:54your coworker and it's kind of a similar
22:56thing here. You like you don't want to
22:57just be a passive observer agent. You
23:00want to be very
23:01in the driver seat in the initial stages
23:03when you're building up your own set of
23:05skills.
23:06Uh you know, obviously you can use
23:07something like P set, but if you're
23:09building your own set of skills, it's
23:11very I think you know, opening up the
23:13all the tool calls and like reading the
23:15code and
23:17reading all the uh
23:18the agent behavior and their thinking
23:20blocks is uh a really great way to see
23:23where they they fail, right? Like what
23:26what
23:27you know, where are they being done? And
23:29then you can go and build the skill for
23:30that.
23:31And then with verification, how you
23:33trust it is it's I think it's also a
23:35very similar iteration loop uh where you
23:38know, like I actually did the same
23:40process for verifying the verification
23:42skill where I actually get um
23:47So, one thing that's interesting about
23:48evals is that you can sort of hill climb
23:50them, meaning that uh
23:53your eval can produce a score, right? Uh
23:55a score that you can get your
23:57coordinator to produce, uh but also uh
24:00you can have a judge agent of a
24:02different model to uh kind of
24:05cross-reference and make sure that the
24:08first model is not being biased, right?
24:10The model that's judging all of the sub
24:12agents that are running the thing.
24:14Uh but you can also like hill climb. So,
24:16meaning that you can you can use like
24:18{slash} loop in cursor
24:20and you can say, "Okay, keep looping on
24:22this eval, right? Until everything is 10
24:25out of 10." As an example. Uh and I did
24:28the same the basically the same approach
24:30with the control skill. And so, I kind
24:31of it was very it was very hands-off
24:33actually.
24:34Uh so, you know, I uh I kind of built I
24:37built that skill that way, like the CLI
24:40in that skill. Um and over time
24:43it's gotten really good. Uh but yeah, it
24:45was definitely not super smooth
24:48at the beginning. It required a lot of
24:50iteration. And I think there's an
24:53analogy here for me, which is um
24:56Well, I make this analogy later in a
24:57different slide on my drawing here,
25:00uh, but I think of it like
25:03uh, you know, as a as a
25:05engineer now, you're sort of more like
25:08you
25:09like maybe a manager or the analogy I
25:11like is like you're like a a chef in a
25:14restaurant. Uh, you you're the head
25:16chef. Uh, you're not cooking all the
25:18food yourself anymore. You have a team
25:20of cooks, right? You have a line cooks,
25:22you have a sous chef, you have you know,
25:23all these different stations.
25:26Um, and it's your job to really design
25:28the environment. You know, you you
25:30you're in charge of setting up the
25:31kitchen. You're in charge of, you know,
25:34like giving tasks to different people.
25:37So,
25:39um, yeah, it's a very interesting way of
25:42working.
25:43Uh, but yeah, that's that's how I
25:45basically built uh, these verification
25:47skills.
25:48>> Yeah, just just one follow up there and
25:50like to go try to go one layer deeper.
25:52So, are you, let's say we wanted to
25:54build um,
25:56uh, an an eval or a skill for for
25:58something
25:59and we wanted to kind of get better on
26:01its own, which is is what what I think
26:03you're suggesting. Uh, are you doing
26:04that in like a work tree, a kind of
26:07isolated with like the sub agents and
26:09and then the reviewer agents and and all
26:11that? Is it happening like in some type
26:12of cloud hosted environment? Like what's
26:15the the more the practical steps? If I
26:16wanted to go do this, uh, and like set
26:18up a verification system for something,
26:20what would I what would I do or where
26:21would I start?
26:24Uh, I think that uh, the best place to
26:27start is local because you can observe.
26:31You can definitely observe what your
26:32agents are doing. So,
26:33uh, if you're building a verification
26:35skill for yourself,
26:37uh, I would definitely start local and
26:39just have your agent bring up the
26:40application, whether it's like a CLI or
26:44uh, desktop app or whatever.
26:45And so, you can actually observe, right?
26:47You can see how the agent is interacting
26:49with the
26:51the application. You can see it, you
26:53know, how it calls like the different
26:56APIs that that allow it to interact with
26:58the
26:59uh the application.
27:02Um but uh for me personally, uh I have
27:06basically been kind of all in mostly all
27:08in on cloud agents because they're
27:10extremely powerful.
27:12Uh and the really powerful thing about
27:14Cursor is the the cloud agents actually.
27:17Where if you spend a little bit of time
27:18setting up your environment,
27:21these control skills, these verification
27:23skills pay a huge amount of dividend
27:26because it's not just something that
27:28makes you as a single engineer better,
27:31it actually levels up your whole team.
27:33Uh and even your whole company because
27:36uh you can actually start thinking about
27:38cloud agents being that thing about
27:39automations that automatically do things
27:42like
27:44uh I'll buy I get I I cannot talk about
27:46this a bit later, but I'll just kind of
27:49get into it. Uh where where, you know,
27:51for example, like I talk a lot about
27:53this agent we have called Benny, right?
27:55Who
27:56who uh you know, takes all of the bug
27:59reports that we get and it automatically
28:02goes off in the cloud, opens up a cloud
28:04uh it's, you know, it's desktop. It runs
28:07Cursor in its own computer
28:09and it uses the same control skills to
28:12interact with the application and try to
28:14reproduce the bug
28:15uh or the user report. All right, and
28:17this is so so powerful because at once I
28:20can immediately I I get so much
28:22information from this automatically.
28:24Like here in this example, you can see
28:25that uh the Benny actually reproduced
28:28the bug,
28:29uh but it's already fixed on main.
28:33So, it actually confirms that we fixed
28:35this problem already. And all I need to
28:37do is just release another build of of
28:39Cursor.
28:40Uh so, that's like huge information
28:42there that I didn't have to go off and
28:44sit with an agent, you know, and spend
28:46an hour trying to figure out, like, is
28:47this fixed? Is this not fixed?
28:49So, you you you gain back so much time,
28:52uh but, you know, everybody on my team
28:54benefits from this. Everybody in the
28:55company benefits from this.
28:57Uh so,
28:59definitely think that, uh you know,
29:01keeping these uh using cloud agents is
29:04super powerful.
29:05Uh but, yeah, it's like a journey. You
29:07have to trust it first, right? Before
29:09you you get to this point. And that's it
29:12goes back to what I was saying here,
29:13where, you know, it's very hard it's
29:15it's almost impossible, and I would
29:17definitely encourage you not to try to
29:19jump from,
29:20you know, like, if you're still in this
29:22zone, you don't want to jump to, like,
29:25I'm going to spawn a hundred a thousand
29:27or a thousand of cloud agents right now,
29:29because you're just going to waste a lot
29:31of tokens,
29:32um and it's going to be extremely
29:33expensive.
29:35>> Yeah, so just to kind of recap so far,
29:37basically the if we wanted to go on the
29:39journey that you've kind of gone on, it
29:40would be just start with verification,
29:43uh building some some skills and some
29:45some ways of determining that the agents
29:47are producing
29:48at least like correct code, whether like
29:50you said, whether it's good code or not
29:51is maybe a separate question, but like
29:52it's it's technically solving the
29:54problem by looking at, you know, stack
29:56traces, looking at, you know, the the
29:58actual behavior in the app, and so on.
30:00Um and then once we trust it locally,
30:02then we can start to think about scaling
30:03into the cloud and running more agents
30:06that are picking up signals, I guess, on
30:08their own, right? So, whether that's
30:09like a bug report that comes in or
30:10something, they can go and pick it up
30:11and solve the problem and and give us
30:14back a PR. And then maybe the last step
30:16is like auto-merging the PRs, which uh
30:18is where you're at.
30:20>> Yeah, yeah.
30:21>> Uh and then reviewing them on main, but
30:23um
30:24is that is that about right?
30:26>> Yeah, exactly. I think yeah, that's why
30:27I drew this this uh this this curve,
30:29right? Because that this this basically
30:31describes my journey of, you know, when
30:33I started
30:35barely could use a couple agents, and I
30:36was just observing every single thing.
30:39I think there's really no shortcut for
30:40going from here to there because this is
30:43really about your personal level of
30:45trust in agents. Right?
30:47Obviously, you know, as a as engineer
30:49you don't want to just slop code into
30:51production.
30:52So, how do you actually build up that
30:53trust? Takes
30:55um a lot of
30:57uh I guess taste and judgment. Um but
31:00uh you know, like I think plugins like
31:02Pstack definitely kind of help you
31:05get up to speed much quicker.
31:07Uh and so I guess it's It's like if you
31:10trust me and you trust Pstack, then in
31:13by extension you can maybe trust your
31:15agents. But if you don't trust me and I
31:17I definitely would not encourage people
31:19to blindly trust me.
31:22Uh
31:23uh you know, if you build up your own
31:24set of skills that you can obviously,
31:26you know, take a look at Pstack and kind
31:28of fork it, make it your own, improve
31:31the skills. Definitely encourage that.
31:33Um but for me it's really all about it
31:36just keeps coming back to trust. You
31:38know, every one of us here in this chat
31:40have a different standard for
31:41engineering. Uh and there are different
31:43things that are important for us in our
31:45code base. And uh when you are able to
31:49encode all of that into skills and you
31:51can verify that your agents actually
31:53doing them,
31:54that allows you to really kind of ascend
31:56this curve. And
31:59uh you know, start automating things.
32:02Uh there's another piece I wanted to
32:03talk about. Um if there's
32:06>> Yeah, go for it. I'll I'll think of more
32:07questions as I go, but yeah.
32:08>> Yeah, I think there's a third to this
32:10which I haven't talked about yet, which
32:12is
32:13kind of an interesting one, which is
32:14like refactoring and rewriting. Like one
32:17of the uh I guess most controversial one
32:20of the most controversial topics in the
32:22industry, I think, is like should you
32:25rewrite your app or not?
32:27Um because I think engineers are very
32:30prone to this where especially when you
32:32join a company, you come in and you see
32:34like the code base and you're like,
32:35"Man, this is
32:37Like who wrote this code? You know, it's
32:39it's terrible. I want to rewrite the
32:40whole thing. There is a very common
32:43inclination and I think a lot of, you
32:45know, before agents, um and I guess
32:48arguably even now, people will
32:49definitely discourage you from re-
32:51rewriting stuff.
32:53But I'm actually here to make a case for
32:54why you might want to consider it.
32:57Um because
33:00I think it really depends. Uh you know,
33:03brownfield applications I think are
33:05actually in a pretty good spot,
33:07especially if they're set up well
33:09already.
33:10Uh and like recently I've been talking
33:12to some people, but uh you know, I I I
33:14was just observing. I I just noticed
33:17this
33:18parallel, which is that a lot of big
33:21tech company problems are now
33:23everybody's problems.
33:25Um and the big tech company problem, you
33:26know, like when I was working at Meta,
33:28like we had this giant mono repo. We had
33:31like, I don't know, tens of thousands of
33:33engineers just, you know, like banging
33:35on their keyboards and and shipping
33:37code.
33:38And
33:40a lot of really great engineers at Meta,
33:42uh but uh I'll say like, you know,
33:45you'll be surprised that the code
33:46quality is actually not that good.
33:48>> [laughter]
33:48>> Um and so
33:50I often joke that like, you know, before
33:51AI slop, we had human slop.
33:54Um and so uh you know, I think a lot of
33:56big tech infra, like uh like what Meta
33:59has or Google, you know, you know,
34:01really big tech companies,
34:03are actually designed for that, where
34:05you you're sort of like, you're catering
34:07to the the you know, like
34:09uh this sounds so bad to say, but like
34:11the the least capable engineer on your
34:13team, right? You build you build
34:15frameworks, you build conventions, you
34:17build guardrails, you know, you restrict
34:20credentials so that, you know, your
34:21intern doesn't wipe your production
34:23database.
34:24Um
34:26there's uh you know, if you have that
34:28level of infra already, I think your
34:31agents can actually already do a very
34:32solid job. Right? Because they have the
34:35the guardrails are already in place for
34:38agents to not cause havoc or not cause
34:41too much havoc in your codebase. Um and
34:45you can always add more, you know,
34:46guardrails.
34:48Uh but I think like greenfield
34:49applications especially are you know,
34:50like the brand new applications are like
34:53the biggest risk in my opinion. Uh and
34:56also the greatest opportunity.
34:58Because, you know, if you vibe code a
35:00project uh a prototype
35:03um like we did for Grokbot, you know,
35:04Grokbot was spun up very very very
35:06quickly. Um and if you if you haven't
35:09heard of of Grokbot, it's like our a new
35:11application we just launched yesterday.
35:13Uh it's it's really cool.
35:15Uh lets you orchestrate your create like
35:18individual agents that have their own
35:20identity and you can kind of orchestrate
35:22them. It's super cool. Definitely check
35:23it out.
35:25Um but yeah, that was it's like a very
35:26it was a very greenfield application
35:28like most prototypes are.
35:30So, it was like vibe coded very quickly.
35:32Humans were not reading the code at all.
35:35And
35:36uh I had this tweet recently
35:39uh where I said something about organic
35:42architecture.
35:43Um
35:45Let me I'll find it.
35:46Uh but the idea is that
35:50uh when you have a completely vibe coded
35:52application, you essentially have no
35:54guardrails whatsoever. So,
35:56uh your agents
35:59when you give them a task, they will
36:01just solve it in whatever method is the
36:03most convenient.
36:05And over time, you get into this
36:07uh
36:08situation where you have a codebase that
36:10is spiraling out of control because you
36:12don't understand it. Uh your agents
36:15understand it, I guess, in a way, but
36:17like they've built something that is,
36:19you know, optimized for short for
36:21shortcuts.
36:22Uh and uh you know, it will you will
36:25suffer you have a lot of of issues with
36:27that application.
36:30Uh so, I think starting your code base
36:33with uh like very strong strains is
36:37very much needed.
36:39Uh because like when you have a code
36:41base that you can trust, right? When you
36:43have guardrails that actually help you
36:45uh
36:47uh help your agents write good code, you
36:49can get into the you know like into this
36:51part of the curve where I I where I like
36:54I I said, you know, I woke up today and
36:56I had like 20 PRs merged uh by my
36:59agents. And that's because I invested a
37:02lot lot a lot of time
37:03uh over 600 PRs I I I calculated
37:06yesterday uh
37:09when I refactored all of GrokBot to this
37:11new architecture that I've been
37:12building.
37:13Um
37:15and yeah, I've I've gotten to a point
37:17where I
37:19I don't really look I really don't look
37:20at the code anymore. And um I say that
37:23not just, you know, to sell you tokens,
37:25but
37:26because I you know, it it it took a lot
37:29of work to get to that point. I spent a
37:30lot of tokens to get the code base to
37:33this point where I no longer have to
37:34look at it.
37:36Uh but I'm very excited because
37:38you know, of the potential where, you
37:40know, it's not just it this doesn't just
37:42benefit me. It benefits everyone
37:44contributing to GrokBot.
37:46And it also empowers, you know,
37:48designers and product managers and, you
37:50know,
37:51even GTM people to add features to
37:54GrokBot. And I don't have to worry, you
37:56know, I don't have to to wake up at
37:58night in in the middle of the night and
37:59worry like, "Oh, Someone's just
38:01merged a perf regression." Right? I have
38:03a ton of
38:04constraints and CIs like it's actually
38:07very annoying to write code in in
38:09GrokBot, but like agents absorb all of
38:11that annoyance.
38:13Um but yeah, I'm happy to talk about
38:15what exactly that is. Um
38:18>> Yeah, I think one question
38:20um
38:20>> Yeah.
38:21>> before we get into the this part here is
38:23just around that element of like what
38:25your your your CI looks like or maybe
38:27some of the constraints and then also
38:29like the average PR size. I saw a
38:30question about that earlier. Just to
38:32give people a you know, kind of a a
38:34glance. It doesn't have to be like
38:35mathematically average, but just a you
38:37know, like what generally the size of
38:39the PR is
38:41um if it's only a couple lines of code
38:42or you know, um yeah.
38:45>> Uh
38:45um
38:47I think it depends. Uh let me
38:51I'm trying to do this in a way where I'm
38:53not going to like
38:53>> Yeah, yeah, you don't have to share the
38:55actual number. If you like the actual
38:56number just
38:57>> this is is fine.
38:58>> benchmark
38:59>> But like we have So okay.
39:01This is not
39:02that interesting, but uh well, fun fact
39:05is that virtualization in Grokbot and in
39:09uh Cursor is actually powered by uh
39:12Pretext,
39:13uh which is a sort of new library that
39:16someone's built. Um that's really
39:19interesting. You should You should check
39:20it out.
39:21But that's not really that important.
39:23Uh I think the average PR size I
39:25actually don't know I I don't know if I
39:27want to click on these.
39:29Uh I probably can, but I would say like
39:32they can range anywhere from a few
39:34hundred lines or 50 lines to like a
39:36thousand depending on what the thing is
39:39doing.
39:40Uh so like here I'm actually like
39:41deleting a bunch of files, so I expect
39:43that it's just this like mostly
39:45deletion.
39:46Uh but yeah, it it kind of varies.
39:49>> There's no like
39:50>> There's no like hard cap or hard limit.
39:52They're all like 50 line PRs.
39:54>> there's definitely no hard cap, but I I
39:55do encourage my agents to split up their
39:57work into multiple PRs.
39:59Uh I do that mostly because uh
40:03I like I like the idea of but I guess
40:06maybe this is much harder to do now as
40:08in the world of agents and you have like
40:10so many commits.
40:12But I like the idea that you know, the
40:13Git history is a very rich source of
40:15context. Uh and I like the I like each
40:19PR to sort of atomically describe what
40:22that
40:23small piece of thing is doing,
40:25which also makes it easier for me to
40:26revert changes and like figure out, you
40:28know, oh, I shipped a bug and it's this
40:30it's here, right? It's not in this
40:3240,000 line PR where you who knows what
40:35landed in there.
40:38Uh but I I don't have a hard cap on PR
40:41size.
40:42>> Cool. And then um yeah, also quick
40:45question on like CI. So again, you don't
40:46have to go into like uh the screen share
40:48of like your your CI does, but just
40:50generally, would you describe what the
40:52CI kind of looks like uh or how strict
40:54it is?
40:56>> Uh yeah, so
40:58uh well, specifically for Grokbot, so
41:01Dune is the is the sort of cheeky code
41:04code name for the architecture that
41:06we've built for
41:07Grokbot. Um the CI looks pretty annoying
41:12because there's checks for everything.
41:14So, like literally, I have um
41:18uh well, if you've written any React for
41:19example, you know, you know that one of
41:21the biggest foot guns in React is use
41:23effect. Uh so, in
41:26uh Dune and in Grokbot, we've banned use
41:29effect. So, Dune is just you can the the
41:32mental model of what Dune is,
41:34uh you can kind of think of it as like
41:35Next.js for
41:37uh Electron apps and it's designed for
41:40agents to write uh and it's like custom
41:42for, you know, our agent for the
41:45applications.
41:46Um so, the CI checks are very like
41:48specific to that, like, you know, don't
41:50use use effect. It's it's it's banned,
41:52like, CI will fail
41:54uh and yell at you. We have like some of
41:57the more interesting ones that people
41:58might raise eyebrows is like I actually
42:00banned code comments as well,
42:03uh which is very interesting. Um but I
42:06noticed that 99% of the time agents just
42:10write code comments that kind of
42:12describe some historical thing that is
42:14actually totally irrelevant to the code.
42:17Um,
42:18like it will often say like, you know,
42:19"Oh, Lauren said you should never do
42:21this." and it's now in the in the code
42:22comment. I'm like, "What? Like, why what
42:24that that was I didn't say that as like
42:26a durable, you know, global rule. I just
42:28meant like your this PR sucks and you
42:31should change that part."
42:33Uh, agents
42:35don't really understand us that well,
42:36surprisingly.
42:38Uh, and or they kind of assume too much
42:40and they kind of do things in like very
42:42stupid ways.
42:43So, like yeah, we just ban everything
42:46everything you can imagine like the
42:47agents are bad at, we ban.
42:50Uh, so one example that we actually
42:53suffer a lot in the agents window is
42:56we have, uh, you know, you know, if
42:58you've used agents window, you've
42:59definitely seen performance issues and,
43:01you know, we're constantly trying to fix
43:02them.
43:03Uh, but it's like a
43:05it's a never-ending struggle because
43:07there's so many pull requests that get
43:09merged.
43:10Every any one of them could just regress
43:12performance or stability or reliability.
43:15Uh, you know, the agents window doesn't
43:16have this architecture yet. I plan to do
43:18bring this learning back there and kind
43:21of refactor everything there.
43:23Uh, but
43:25uh, it just regresses super often
43:27uh, because uh, there's there's just one
43:30example is like we have very poor
43:33um, isolation between processes. So,
43:35like on you know, on Electron, you have
43:37a renderer thread that renders your UI.
43:40But you also have like a main thread
43:41that you can run other code that, you
43:43know, doesn't need to block the
43:44renderer.
43:46Um, but we do a poor job of separating
43:48those things. And so, often times you
43:51just accidentally have code that gets
43:53pulled into running on the renderer
43:55thread. All right, and then all of a
43:56sudden you're competing with the the
43:58renderer that you know, that has a very
44:01If you want like 60 FPS, you have to
44:03every frame that gets drawn has to be
44:05done in 16 milliseconds. It's a very,
44:08very small, you know, deadline per
44:09frame.
44:10Uh if you want, you know, a very smooth
44:13product. Uh and when you start building
44:15bringing in accidentally bringing in,
44:17you know, things that are like very
44:19computationally heavy or they have a lot
44:21of IO, uh
44:23then you just get into like a lot of
44:24jank. Uh your your FPS really drops. You
44:27start uh you know, losing frames. You
44:30get long tasks that
44:32take more than 16 milliseconds and you
44:33just get this really choppy experience.
44:36So, all of those patterns that we've
44:38learned basically building Electron
44:40apps, we've encoded into this framework
44:42and it becomes like a hard failure. So,
44:45I literally in in Grokbot, we literally
44:47have a directory called Electron main,
44:49Electron renderer, and we have uh
44:52import
44:54uh CI, I guess, where we actually check
44:57the dependency graph to make sure you're
44:59not accidentally importing code from one
45:02directory to another.
45:03Uh so, that's enforced by CI,
45:07um as well as Bugbot, uh which is our
45:10which cursors
45:11um like code review tool that runs on
45:14CI.
45:15Uh you know, in our agents MD, it's
45:17everywhere. Like so, I I I I um
45:20I have this thing here where I talk
45:22about like
45:24um you know, like there are multiple
45:25layers, I think, for building a good
45:28code base.
45:29Uh obviously, the code base is one where
45:31uh if you have an architecture like
45:33this, where it's extremely strict, uh
45:36you know, the the the way to build
45:38features is very conventional. That's
45:40like the strongest strongest level of
45:42enforcement because agents just love to
45:45copy existing patterns.
45:47So,
45:48uh one example of this in Grokbot is
45:50like we have this these concepts called
45:52like a feature and we have entry points
45:54and transcript cards. Like all of you
45:56know, the cards that you see in the
45:57chat.
45:58These are all like
46:00like nouns, I guess, in in the
46:02framework. And so, there's a very
46:04conventional way of creating them. And
46:08so, like a feature is all in in a single
46:09directory, as an example. And so, all of
46:12the code that contributes to that
46:13feature lives in one directory. So, it's
46:16all co-located in one place. Makes it
46:18super easy, you know, agents don't have
46:19to like uh grep around and try to figure
46:22out like where all the things are.
46:24It just looks at the feature and like,
46:25"Oh, okay, I'm working on the onboarding
46:28feature in Grokbot. Uh I'm just going to
46:31work in this directory." And for 80% of
46:33the work, it's mostly just very
46:35encapsulated there.
46:37But uh like it's like
46:40very It's like designed again for, you
46:42know, like the dumbest agent. Like, you
46:44don't have to think. All right, the the
46:46the One of the key principles I have for
46:49this framework is like the shortest the
46:51shortest path is the best path.
46:54So, uh because that plays exactly to how
46:57agents love to write code. Like, they
46:59like to take shortcuts, really. You
47:01know, they'll they'll find the quickest
47:02way to solve the problem. So, why not
47:06make that the best way to solve the
47:08problem?
47:09Uh so, I I I I probably won't get into
47:11all the specific details. Um and uh the
47:15the
47:16this framework is really more of a
47:17collection of ideas and principles,
47:19rather than something that we'll open
47:20source.
47:21Uh you can you can, you know, screenshot
47:23this, I guess, if you want and uh tell
47:25your agent to
47:27uh do some build build something like
47:29this for you to
47:31um Yeah, but it's really all about the
47:33layers. Uh you know, like the the code
47:36base is one part with features uh and
47:38directories and you know, import
47:41uh blocking import dependencies uh that
47:44shouldn't be imported. Uh but and and it
47:47all enforces that and static analysis.
47:49So, like uh there's CI checks. We have a
47:52lot of lints for bad patterns that we
47:55observe. Uh compiler diagnostics. Uh,
47:59there's also rules and bug bot which are
48:01um, I think like three, four, five are
48:04more soft, right? These two actually
48:07make
48:08make CI red, right? So, that you know,
48:11there's a hard constraint where the
48:13agent can just write crappy code.
48:17For rules and skills and bug bot,
48:20your agents can still forget, right? You
48:22can still
48:23or it may not always consistently apply
48:26them. So, I like to layer them, but I
48:30don't I don't like to rely on them as
48:32the only source
48:34of enforcement because it's very, very
48:35soft, right? And if you if you only have
48:37rules and bug bot and skills and a style
48:40guide for your code,
48:41you will it's only a matter of time
48:43before your code base looks like
48:45complete trash. Uh, I'm sorry to say
48:47that, but uh, I definitely recommend you
48:50all like, you know, investing
48:51in, you know, things that can be hard
48:53enforced. Right? And this is why,
48:56you know, maybe uh, the choice of tech
48:58stack that you use is also very
48:59important.
49:01Um,
49:02like I think for example, Rust is sort
49:04of making, you know, it's like getting
49:06super popular again, uh, because the
49:09compiler is so strict, right? The
49:11compiler enforces so many different
49:13things, you know, there's a borrow
49:14checker that you have to appease.
49:16And if as long as you make sure your
49:18agents don't write unsafe code blocks,
49:20uh, you can more or less feel somewhat
49:23confident that if the code compiles, it
49:25probably works and it's good.
49:27Uh, but you know, you see, it gives you
49:28that level of trust and confidence that
49:32you as a human engineer no longer need
49:35to go and check it yourself. You you
49:37know, you you rely on
49:39code and static analysis to actually
49:43make that uh, a lot smoother.
49:46Um, and I I guess the worst part the
49:49worst place to be in is if you are stuck
49:52in code review land where you actually
49:54enforce all of the constraints, the
49:56invariants in your code base
49:58by literally the human person saying,
50:02you know, reading the code and be like,
50:03"Okay, you should not do this."
50:04Right?
50:06Every time you have to do that, you
50:07should consider that as a code smell,
50:08like a uh anti-pattern. And you should
50:11say, "Okay, instead of me commenting on
50:13the PR,
50:14how do I turn this into a hard rule?
50:17Right? How do I turn this into a lint
50:18rule? How do I turn this into a CI
50:21failure? Or how do I even categorically
50:23eliminate this problem
50:25uh entirely?
50:27Uh
50:28I I can talk about another migration
50:30I've done, but I'll probably pause here.
50:32>> Sure.
50:33Yeah, I feel like that's that's where I
50:34am, to be honest, is is what you're
50:36describing right now, which is that like
50:38I don't have all of these rules. Uh so,
50:40I have some things to go do after this
50:41session
50:42in terms of being able to scale my
50:43agents. I'm I'm definitely on like the
50:46uh you know, maybe a couple of parallel
50:48ones locally stage, so like two to three
50:50locally. And I'm sure most people here
50:52are on the same, so uh
50:54yeah, I know we only have a couple
50:55minutes left. Uh Lauren, was there
50:56anything else that you wanted to to
50:57highlight? I obviously there's lots of
50:59questions, so I can grab more, but I
51:01want to give you a few minutes uh if
51:02there's anything else you want to talk
51:03about.
51:03>> I think I've been yapping for quite a
51:05lot, so I'm Maybe let's just do
51:06questions.
51:08>> Okay, cool. Uh one question I had uh a
51:10couple of uh came up a couple of times
51:12was just around like token usage.
51:14So, the the question is like is is what
51:17you're describing a realistic thing for
51:19people who are on, you know, uh
51:22a a normal set of token usage, they
51:24don't have, you know, basically
51:25unlimited tokens uh to work with.
51:29>> Well, I think that's a really good
51:30point. I mean, like obviously, you know,
51:31I work at a AI lab where we have
51:34unlimited tokens, so uh
51:36I definitely cannot
51:39say that, you know, this is something
51:40everyone
51:42should do in the exact same way that I
51:43did it. I think it's possible to get to
51:45this point without, you know, breaking
51:47the bank.
51:49But, you know, if you're like an
51:50engineering leader or, you know, you're
51:52you're you have a startup that you lead,
51:54um I think that may be a question of
51:56ROI.
51:58Um and it's like
52:00uh
52:00yes, you spend a lot of money on tokens
52:04in the upfront stage. You know, you like
52:06refactoring your code base is going to
52:07take a lot of tokens. Uh adding all
52:09these things uh is going to take a bunch
52:11of tokens.
52:13But, if we're heading to a world where
52:15agents are writing all the code,
52:17and you know, you want to be very lean,
52:20right? You don't want to have to hire
52:22You don't want to be You don't want to
52:23become like meta, right? Like I mean
52:25like in terms of You don't want to
52:26become a 10,000-person engineering org
52:29because
52:30I mean, that's a cool problem to have,
52:32but also you know, you you have so much
52:34overhead. There's like planning, you
52:36know, like you it it's it's uh
52:39personally I I wouldn't uh
52:41it it's not super fun. But, um
52:44I think you want to stay very nimble,
52:45right? And you want to you want to be
52:47like agents are all about allowing you
52:49to do things that you couldn't do
52:51before. That's really to me like the
52:53value of agents. You know, it's not just
52:56throwing tokens on every single little
52:58thing.
52:59But, um to me like the thing I couldn't
53:01do before is like
53:02enforce this level of constraints in a
53:05code base by myself.
53:07Right? Like I'm just a single person,
53:09you know,
53:10uh it would have taken me years to build
53:13this framework uh and do it all the
53:15refactoring and
53:18test everything myself and verify, you
53:20know, like run Imagine if there was just
53:22me right now in in pre-agent era, just
53:25like running, you know, by my It would
53:27take me so long, right? And my salary is
53:30pretty high, right? Like So, you know,
53:33the the question I think an engineering
53:34leader might have is just then,
53:36you know, like what is there's a
53:38tradeoff of do you hire someone to do
53:41this,
53:42or do you spend the tokens to set up a
53:44code base so that even
53:46the the most naive, right, the dumbest
53:49agents can do a good job.
53:51And when you actually get to this point,
53:53like even agents that are not, you know,
53:55Fable size do an excellent job of
53:58writing code.
53:59And this pays a lot of dividends as well
54:01for me and personally where I've
54:03empowered not just myself, but
54:06again, like PMs, designers, engineers
54:09who are not familiar with Grok but to
54:11just contribute in a way that
54:14is sustainable.
54:16So, I think yeah, it's definitely like a
54:18trade-off for sure. You know, like
54:19nothing's like free for sure.
54:22Uh and tokens are pretty expensive.
54:24Uh but oh, actually,
54:26I don't I I I don't know how many of you
54:28have seen this, but we actually
54:29announced Grok 4.6 today. So, very
54:32exciting. Finally out. Um so, yeah, Grok
54:354.6 would be like a great It's very very
54:37smart. Uh it's really good on on the
54:40benchmarks. Uh and it's the same the
54:43tokens Uh well, uh I hopefully I'm not
54:45saying this thing correctly, but uh I
54:47believe the cost per token is the same
54:50as 4.5.
54:51So, you're actually getting more
54:53intelligence for the same cost.
54:56Uh I think this is an area that Cursor
54:58tries to Cursor and SpaceX AI
55:01try to really optimize for like that
55:04Pareto frontier of, you know, cost
55:06versus intelligence. Uh you know, we
55:08don't necessarily want to build the
55:10biggest model ever because that is
55:12extremely expensive to run.
55:14It's really about like how do you find
55:16that sweet spot, right? You don't You
55:17don't need a giant model, but it's just
55:19super smart, right? And it's not very
55:21expensive for inference.
55:24Uh but um yeah, I think to kind of round
55:27it up, um
55:29I think it's like a it's it's there's a
55:31If you do your own analysis, I feel like
55:33it's pretty positive.
55:35It It'll be pretty positive that the ROI
55:38you get from investing in stuff like
55:40this uh just empowers not just yourself,
55:43but your whole team to be so much more
55:45productive. Right? Like imagine if you
55:47have an army of engineers like me who
55:50are shipping so much improvements and
55:52and bug fixes uh you know, every day.
55:56Right? Like
55:57that is pretty exciting.
55:59>> Cool.
56:00One last question before we wrap up.
56:02This one is for the people in product on
56:04the on the call. So, let's say we do
56:06have an army of engineers who are
56:08shipping like Lauren. I'm just curious
56:10like how is the product team or other
56:12functions of your company keeping up
56:13given that like if you're shipping so
56:16quickly, have have are they using AI
56:18more to do their jobs? Like as much as
56:20you can speak to that. I obviously you
56:21don't have like you're not in that role,
56:23but just curious about how that works.
56:25>> Um I think this is where Grokbot has
56:28been actually exceedingly powerful. Uh
56:30where
56:32so before Grokbot like you know,
56:34obviously Cursor only had Cursor. Like
56:36we only had agent window. We had a CLI.
56:38We had an IDE.
56:40And these are really like
56:42power user tools, right? Like the
56:44they're designed for developers. So,
56:46it's very very developer-centric. You
56:48can do knowledge work in them, but it
56:50like the UI is not really optimized for
56:52that.
56:53So, we actually didn't really have uh
56:56well, I think like a lot of people like
56:58you know, GTM, product, like they might
57:00have used Cursor uh to do their work,
57:03but it definitely wasn't like a the like
57:05full experience with them.
57:07Um I think now with Grokbot
57:10uh it's become Grokbot is basically like
57:13the Cursor moment for people who are not
57:16in tech in my opinion. Like it's like
57:18it's like a very very accessible way to
57:21use agents in a very comfortable, very
57:23familiar interface. It looks like
57:25iMessage. Um and it's very fun too, you
57:29know, you can give your agent to be
57:30name. Uh you can have you can kind of do
57:33orchestration with in a very like
57:35natural way where you can sort of
57:37you know, each agent is like a person
57:39and now you're on a team of agents like
57:40working on you have one one agent per
57:42account that you manage as an example.
57:44Or if you're a PM, you have you know,
57:46you can have an agent that summarizes
57:47all the work that Lauren did last night
57:49and then now you know what I did, right?
57:52So, I think our PMs are leveraging that
57:54a lot. And they're shipping code, too.
57:57Uh, so you know, I often times they will
57:59just say, "Oh, here's a bug. I fix it.
58:01Can you look at it?" And then I'll go
58:02review it and actually it's just
58:04perfect. I'm like, "Okay, stamp." Uh, so
58:07uh, I think that shows that you know,
58:08the the Dune architecture is holding up,
58:11right? The all the the really strict
58:13constraints allow
58:15people who are not experts in
58:17engineering to contribute at a high
58:19level.
58:20Uh, so I'm I feel like I'm already
58:21seeing that pay off a lot where uh, you
58:24know, designers and PMs are just able to
58:27to to ship features directly.
58:30Um, and that just makes
58:32the Grock bot team super fast, right?
58:35Where we can ship so quickly.
58:37Um,
58:39and we have a lot planned, so I'm very
58:41excited uh, to you know, uh,
58:44to to ship more uh, ship more stuff.
58:46>> Yeah, that's awesome. Uh, we are at
58:48time, so I I guess Lauren, if if folks
58:51want to support you, maybe go try Grock
58:53bot, try out uh, 46 and uh, you know,
58:57just get to get provide some feedback.
58:58But yeah, this was awesome. Really
58:59appreciate you taking the time. Uh,
59:01thanks everyone for all the messages in
59:02the chat. Uh, lots of good questions. I
59:04know we didn't get through everything,
59:05but uh, as I kind of said at the top,
59:06way more questions than than we could
59:07get through. But uh, yeah, really really
59:10thanks thanks for for joining. Thanks
59:11everyone for joining and uh, hopefully
59:13you you enjoyed the session.
59:15>> Yep.
59:16>> All right.
59:16>> Yeah, I see see you thanks for having me
59:18and uh, if you have any more questions,
59:19yeah, just DM me on Twitter. Uh, I'll
59:21I'll open them up, I guess. I'll let the
59:23I'll let the floodgates
59:24>> to get you're going to get a lot of DMs.
59:25>> Yeah, I'll open the floodgates. So,
59:27yeah, DM me. Maybe I'll do like a
59:29Twitter Space at some point as well for
59:31more questions, but
59:33really appreciate everyone for showing
59:34up, uh you know, taking an hour out of
59:36your day.
59:37>> Yeah. All right. Thanks, all. I'll see
59:38you in the next one.
59:39>> Okay, thanks, everyone. Bye.