3 October 2026
Poteto (creator of pstack) & Matt Pocock · watch on YouTube ↗ · click any timestamp to jump the video
Machine-generated by an AI from the transcript and the top comments. Not my writing, and it may contain errors.
The first ten minutes are mutual appreciation and warm-up. The substance runs from 15:55 to about 56:40, and the single best stretch is the environment argument from 25:18 through the "dark factory" admission at 51:53. Captions are auto-generated, so the >> marks a change of speaker.
The first skill she built at Cursor gave the agent hands and eyes — run the app, drive it, take traces and heap snapshots. Her framing of why it mattered: before that she was "the meat proxy" between the agent and Chrome DevTools. If an agent cannot observe the result of its own work it cannot iterate, so the word loop is doing no work; verification is the part that closes it.
Before the CLI existed, every agent rebuilt the verification scaffolding from scratch, differently each time, then threw it away — wasting context, and more importantly wall-clock time. Her rule of thumb is to treat agent work as a gradient from judgement to mechanics, push the mechanical end into scripts the skill ships with, and leave the model only what actually needs thought.
The discipline is to watch how agents fail and then ask how to make that failure impossible, rather than correcting the agent in front of you. Early Grokbot was eight god files of 10,000+ lines; it became per-feature directories plus a registry and restrictive lint rules, so there is essentially one way to do anything. She links it back to TypeScript type narrowing — constrain the space until only the correct value is representable.
The inner loop is agents working toward a snapshot of what you wanted. New information — bug reports, feature requests, an infra limitation someone mentioned — lands in Slack, Linear or X, outside that loop. If nothing pulls it back in, you are the proxy ferrying context by hand. Connecting the outer loop to the inner one is what lets agents fetch their own context instead of asking you.
One of her routines scans constantly for bad React patterns — and is explicitly told not to fix them, only to append to a document. She reads the buffer every few days and discovers that twenty findings are one finding. Pure execution mode misses the pattern; a deliberate buffer forces the zoom-out that produces the real fix.
At volume you sample like a factory quality supervisor rather than inspecting every item. A single agent taking a shortcut is a one-off and may need nothing. The same shortcut appearing across several agents is the signal — and the fix goes into the skills, lints and types, never into a correction aimed at one chat.
Asked whether the lights are on, she cheerfully concedes they aren't: she has the equivalent of ten-plus "chiefs of staff", each owning an area, merging their own pull requests while she sleeps. Review happens after landing, by reading the morning's commits and then reverting or adding a lint rule. Her verification mode spawns a swarm of verifier agents per PR that fuzz the running app and fix what they find — token-hungry, and tunable down to one.
Pushed on domains where a bad merge causes data loss, she doesn't reach for a workaround — she says it comes down to how verifiable the domain is, that software happens to be unusually verifiable, and that she doesn't have the answer. The forward-looking bit is her bet on agent-oriented languages that marry code with proofs (she names Bend), so that "it compiles" can carry the weight that review carries today.
Her claim is that once the model stops being the bottleneck, the bottleneck becomes your ability to express intent clearly enough to be carried out — which advantages the doctor or lawyer with deep non-engineering expertise and just enough technical curiosity, not less.
If you haven't invested in skills and tooling, the only way to cope with low trust is to micromanage — which consumes exactly the time you would need to build the thing that would raise the trust. Her analogy: dull knives, a looming deadline, and no slack to sharpen them.
Spawning an agent per report duplicates work and, worse, loses the thread between them. Several slightly different reports are often how you discover the real defect sits a level above where any single report pointed.
Rather than inventing a skill, mine past chats for the moments you had to intervene and correct the agent — that is the actual process rather than an abstraction of it. She wrapped this into a pstack skill called recall after repeatedly wanting last session's hard-won context in a new chat.
Last year's skills encoded exact commands and implementation detail. With current models you can delete most of that and keep only the workflow — a series of steps. She expects skills to keep getting smaller and more compact.
The most-liked reply by a distance, and the obvious retort: if all this validation machinery works, measure recurring bugs over time rather than merged PRs. It goes unaddressed in the interview.
A commenter points out these can't be 2,500 features, so they are mostly fixes — which raises the question of why a codebase with this many validation layers generates that much repair work, and whether automating your own exit from the bug-fixing loop is the win it sounds like.
A sharp one: a flow that ingests public bug reports and merges its own PRs with little human supervision invites deliberately crafted reports designed to walk exploitable code into production — worth auditing in exactly the kind of organisation that attracts that attention.
A nice paradox from the comments: every deterministic gate built to constrain the agents — lints, types, verification, CI — is just good engineering practice, and makes the codebase better for the humans whether or not an agent ever touches it.
Auto-generated captions. Click any line or timestamp to seek; the current line highlights as the video plays.
0:00So, hello folks. I've got another treat for you today.
0:03Last time on this kind of podcasty thing, I suppose, we had Uncle Bob and we talked about software quality.
0:09We talked about agents. We talked about lots of cool stuff.
0:12Now, we have uh an incredible guest, someone who I'm delighted to welcome on, who's been exploding on Twitter recently about software
0:21factories, um about increasing the quality of your work, about increasing your velocity and climbing the trust ladder with agents so that
0:30you can ship more and more and more.
0:32And it is potato. Welcome. Thank you so much for joining.
0:36>> Thanks for having me. Yeah, very excited to be here.
0:39Yeah, big fan of yours >> and a huge fan of yours.
0:42I think people have been talking about this like it's like the meeting of the skill minds, the skill Mount Olympus or
0:49something because both of us have very popular skill libraries.
0:51Um I've not, as I was saying before we started, I've not used a ton of yours and like I want to
0:57get all of the juice out of your brain so that I can go and use it properly and use it better.
1:02And I think where I want to start with this is you gave a talk um pretty recently like um about 10
1:08days ago and posted on X which went absolutely nuts as about how I shipped 2,500 PRs last month to production got
1:16about 3 million views or something on X and I watched it and I loved it and I recommended it and I
1:23kind of want to run this as almost like a Q&A of that talk basically of giving you because it just I
1:29just had tons of questions about it and I wanted to dive into it.
1:33And I think where I want to start is you talk about a trust ladder with agents where you as you trust
1:40agents more, you can get them to do better and better things and or scale them to up to use more and
1:47more agents. So what is your story of how you climbed the trust ladder and how did that work when like you
1:54got SpaceX >> started climbing more and more?
1:58So I think this the the journey sort of began even before I joined cursor uh which is now SpaceX AI.
2:05Uh so the story is um after Meta so I I used to work at Meta on the React team.
2:13Uh I took a month off uh because I was feeling kind of burnt out and of course when what what do
2:19you do when you're burnt out?
2:20You go and start a new side project.
2:22Um and so I started a side project.
2:24you know, I was uh of course using AI to to write code.
2:28Uh but then I started to realize uh you know, I was spending like so many hours just micromanaging one agent, right?
2:35And you know, at the time, this was back in February, maybe February, early February or January, you know, people were really
2:43obsessed with this idea of like orchestration.
2:45This was like, you know, before, you know, things like cursor, you know, like the agents window was had become popular.
2:51So people were still in like like 2 land you know in their terminal and they were all talking about okay here
2:57you know I built a custom orchestrator right and so of course I had I was a bit nerd sniped by that
3:03and you know as I was building my toy project uh I got nerd sniped by oh how do I make my
3:09AI coding setup more efficient and so you know I I kind of started the journey there where I just you know
3:17took a step back and realized you know I was spending all this time micromanaging a single agent you know I was
3:23creating skills and I was like finding it quite difficult to measure the output or the the result the impact of the
3:30skill as well so I was kind of flying blind but I was you know iterating really fast um and um so
3:39that project eventually sort of became the basis of PAC even though I didn't know it at the time um and a
3:47lot of some tricks I had learned like building that early set of skills.
3:51Actually, it's still open source if you want to if anybody wants to take a look.
3:55It's on my GitHub like potato noodle n o d l e.
4:01Um, and in there you will see some skills and a brain directory.
4:05And so I was really interested in this idea of how do I, you know, extract my own ability, if that makes
4:12sense, and give it to the agent, right?
4:14cuz I was I I realized that you know all I was trying to do was trying to teach the agent to
4:19write code more like me you know do do you do do workflows more like me.
4:24So you know the skills were like an entry point to doing Um and then you know after I joined cursor uh
4:32I was starting to work on the agents window and uh it had a lot of performance issues.
4:37Uh it was it was it was pretty laggy.
4:40Uh and so since I had experience working in React, I was asked like, "Hey, do you want to come and help
4:45out uh with the agents window?" Um and so the the this beginning of the cursor journey was very manual.
4:54Uh I was deep in like looking at like flame graphs and heap snapshots and trying to see like why exactly is
5:02the app so slow. Uh but then coming back to the same realization like you know I was sort of the bottleneck.
5:08I was doing everything manually. I was sort of the meat proxy in a way, right?
5:12I was the meat proxy between my agent and Chrome DevTools.
5:16Uh and I was like really annoyed by that.
5:18>> And what month of the year is that?
5:20Let's say where are we in the timeline?
5:22>> Uh so I joined Cursor in March.
5:25So this was like early early April probably early April is when you know uh I joined and I didn't have any
5:32skills, right? I had I I sort of abandoned my personal skills because I didn't think they'd be relevant anymore.
5:38Uh but then working on the agents window uh and now working on grockbot uh I sort of realized that a lot
5:45of the lessons I had learned from those skill time building the the initial set of skills were very relevant especially around
5:53things like uh you know being very rigorous in your work um because I think from my experience even the the frontier
6:04ones tend to tend to take shortcuts.
6:09Uh they tend to do the easy thing.
6:12Uh so uh a lot of the skills that I've built have been around how do I make the easy thing the
6:19right thing? You know, how do I make that the best thing?
6:22>> The idea of sort of distilling your expertise and turning what you do every day into processes, that's something that feels
6:31super familiar to me. That's exactly what I've been doing with the skills.
6:34And I suppose there's something in that which is a lot of people think domain expertise is getting less useful now as
6:42people uh start to rely more on AI where what do you think about that just as a sort of vibe check
6:48before we start talking >> I actually feel like domain expertise is is more important than ever you know uh I think
6:57I wrote this on my ex at some point but you know at times I sometimes think of you know AI as
7:03is like especially as the models get smarter and smarter and more capable and the frontier models are just getting so good
7:10like I love Opus 5.5 by the way um uh you know as the models get really really really good it almost
7:18becomes like the bottleneck is no longer the agent right it becomes your ability to express your intent and your goals in
7:27a clear way that the agent can understand and actually carry out and That's why I think you know like people with
7:35a lot of domain expertise are extremely have a have a huge advantage in my opinion especially if you're a little bit
7:42like you know tech technoc curious you know so I I think of people like you know like uh like a doctor
7:49or a lawyer or you know someone who who has a deep expertise in a particular non-engineering domain and if they're actually
7:56just a little bit techsavvy and they can figure out how to use agents they can actually build really really great products,
8:04right? If they if they have a clear enough vision in their head and they can articulate it in a way that
8:09the agent can build it, you know, I think that that is really the the bottleneck these days is is like the
8:18transfer of your intent, right, and your vision to the agent.
8:23>> Yeah. I've been obsessed with language basically since agents um dropped.
8:27are just obsessed 100% and thinking constantly about the the composition of words, how I can make things sharper, what um what
8:36might be hidden in the phrases that I'm using.
8:39And it's and finding what I love is when you find a word that the agent then hooks on to and then
8:45goes, "Okay, I'm going to reinforce that word.
8:48I'm going to reuse that in my thinking traces." You know, I found that with um TDD was an early example of
8:53that. a lot of chat about TDD recently of like, you know, people say, should you use TDD with agents?
8:58Doesn't matter. What you're doing is you're getting the agent to think about TDD, getting it to write tests, getting it to
9:04prioritize things in a different way than it did before.
9:07And that's why sort of grilling, I think, works effectively.
9:10Grilling is >> Yeah. Yeah. It it draws those words out of you, right?
9:14or or at least it helps the agent understand your thinking so that they can propose those words to you and you
9:21can pick up and say yes exactly >> Uh I've actually copied some of the the tips that you've shared as well
9:26where you know one of my favorite ones that you've shared recently or or not or like maybe in the past couple
9:31weeks is about uh reducing or eliminating tautological tests.
9:37Like one of my pet peeves of agents is like all of the useless tests that they write.
9:41And so, you know, that was one thing where, you know, the word tutology, right, is is is I guess, you know,
9:47not many people necessarily know that if if especially if English isn't your first language, but there's a lot of meaning to
9:53that word. And it's like it's almost like compressed, right?
9:57Like you compress a lot of intent and meaning into words.
10:02And so I I I totally agree with you.
10:04I think language I've always been interested in language actually uh like programming languages natural human languages and how they came to
10:12be and it's so interesting that now with agents it's sort of like this meeting of natural language with programming language but
10:20it's all it's all language out of the hood it's all communication >> totally I did a drama degree right so you
10:25know I've been thinking about language and Shakespeare and stuff for a long time and so this all feels very >> um
10:32so okay there's sort before we get into like because I think the thing I want from you is like software factory
10:39stuff, right? Software factory is the big buzzword.
10:42Software factory is the thing that I'm thinking about too.
10:44I'm sort of releasing a course in that direction too.
10:47>> And it's this sort of scaling yourself up to un unrealistic numbers of PRs basically or PR numbers that sound ridiculous
10:56to people who don't understand how this works.
10:59So where I want to get to is sort of from people who are doing kind of like one to five agents
11:04today up to, you know, hundreds of agents running at once and how that sort of functions.
11:09And so I'd love to hear about your metaphor of the Michelin Kitchen instead of the software factory because I think that
11:15says a bit about the way you think about this stuff.
11:20Yeah. I I I've I've never really liked the term software factory.
11:24Not because you know it's not accurate but I think I think the a lot of people when they think factory right
11:30they don't necessarily equate that with quality or craft right things which are very important to me and a lot of people
11:38and technologists who work you know building products we care about the user experience we care about the things we're building.
11:47So while so while I do think software factory is an apt term, it also I guess maybe conjures up negative, you
11:54know, maybe sometimes negative connotations. So Michelin Kitchen is the thing that I've sort of landed on where it's much more I
12:02feel like it's much more aspirational and uh I like the metaphor a lot cuz you know I like food.
12:07I'm called potato of course and I like cooking and I see a lot of parallels right like with food right when
12:15you're cooking a meal for yourself for example it's both utilitarian like you're trying to just feed yourself right and and survive
12:22uh but it can actually be transformed into art right and that's what what a Michelin starred chef or even just a
12:29chef or a cook can do with food is take something very ordinary and turn it into a delicious meal that you
12:36know takes you back to your childhood days or something like that.
12:40Um and so it almost like mirrors that trust letter that I talk about where uh you can sort of imagine your
12:48own journey as a home cook, right?
12:50Uh as a home cook, you are doing all of the food, the cooking yourself.
12:54You cut all the vegetables, you do all the prep work, you do all the cleanup, you know, you are the one
13:01man or one woman show really.
13:04Um, and it's an interesting thought experiment like, okay, if you were to cook a meal and then you add people, right,
13:12your your your partner trying to your brother, your sister, and now suddenly you have your whole family in the kitchen.
13:18I think most people would get very stressed by that, right?
13:21The thought of, oh, so many people are just mocking around in my kitchen.
13:24They have no no idea where all the utensils >> I have a max capacity of one person in the kitchen.
13:30Yeah, absolutely. So, I feel like that that's really apt because when you ask yourself that question of how do I go
13:35from being a solo cook, right, to having an army or even not not even an army but a few sue chefs,
13:44right, that that are helping me in the kitchen.
13:45How do I think about dividing the work in a way that makes sense?
13:50You know, I'm not dividing work just for the sake of it, but in a way that actually makes the sum the
13:55to the better than, you know, the total of its parts.
13:59And so the Michelin kitchen metaphor to me like works really well in that regard because you know as a chef you're
14:07you know if you become a chef you're in a position where you're not necessarily cooking all the food yourself anymore but
14:14you are thinking you're almost like the tech lead right for the kitchen where uh you know chefs have to think about
14:21you know not just cooking but they have to basically organize the whole kitchen and they're like the CEO of the kitchen
14:27they have to think about when do you order ingredients, how do you store them, how do you prepare them, when do
14:32they have to be prepared, you know, it's a whole it's a whole job, right?
14:35That's not just cooking. Um, and I think that again it mirrors so much of how engineers write code today where you
14:44are not writing the code yourself anymore.
14:46You have agents, right? But you as the human are still responsible for the final outcome, right?
14:51your name still is associated with the work that you do, your reputation and you know so how you set up your
14:59kitchen right and how you set up your skills your environment your codebase I think are ultimately the new ingredients that go
15:08into um building >> yeah I think what I love about your approach is the amount of focus that you put into
15:16the environment that the agent operates in right because I think a lot of people they think, right, the agent is good.
15:23I'm probably not going to be able to make it better.
15:26Let's just trust what these magic model people have put into the harness and the model combination.
15:32Uh, there's nothing I can really do, like I can't mess about with claw codes internals or something or whatever you're using.
15:38Um, but what I love about your approach, and it's something I advocate for too, is that you can change the environment
15:45the agent operates in, right? you can make changes in the codebase and also give it tools for verification as well and
15:53allow it to verify its own work.
15:55So the thing I I loved about watching that talk is the amount of focus you put in verification and like that
16:02is the lever that you can start to generate trust.
16:05Can you talk about that and what that concretely looks like?
16:07Let's start like looking at practical ways that people can improve their own processes, their own kitchens.
16:14Yeah, I've I've I've said this a lot actually that you know even if you don't use PAC or you know your
16:20skills I think that the single most important skill that should be in your toolkit is verification because without verification and for
16:30for by the way for those watching who don't know what that means it's this idea that you can give you can
16:36sort of give your agent uh hands and eyes in a way that's the the analogy where the agent is able to
16:45run the code, right? And actually uh interact with it like a normal human user would and also do things like you
16:53know debug it, you know, take traces and snapshots.
16:57Um and uh funnily enough like that was actually the first skill I built when I joined Cursor.
17:02uh that gave me a lot of that was that was the thing that actually started to let me ascend the trust
17:07ladder a little bit in a way that some of the other skills I had looked at or built had not really
17:13let me do because no matter how good you know some of the other skills were like the how skill, the why
17:20skill, the unsop skill were, I was still relying on me right as the proxy between my agent and the output.
17:29So that you know if the agent can't actually see the result of its work there's no way it can actually iterate
17:35right and so this is where people start to talk about loops this idea of a loop and really I think the
17:40term loop you know seems kind of uh almost abstract like people like what what what is a loop what is an
17:47agent loop but really to me like the most important part of a loop that allows it to be a loop is
17:54the verification part because the agent is able to to verify by its own work and uh you know that takes you
18:02out of the equation where now I can actually do something like so the very one of the very first use cases
18:07I had for verification was you know like the performance work that I was doing on cursors agent window and I want
18:13I wanted to get to a point where I could do something called hill climbing uh which is a term that I
18:19I think the labs uh talk about a lot which is this idea that you know you have some kind of rubric
18:25or a way to judge or score something And now because you have a loop, you can have an agent continually try
18:32to make improvements to that. Uh I think Carpathy, Andre Carpathy also famously released uh something called auto research that has a
18:41lot of these ideas. Um but yeah, verification I would say is probably the most important skill in PAC uh and many
18:50other you know tool sets. Uh and I think it's the most important thing to focus on.
18:57So a lot of the a lot of I spent a lot of time actually you know tuning the verification the creative
19:04verification skill um and also internally the the we have so many verification skills now like every app that cursor has or
19:12spaceexai has has a uh verification skill that is automaintained as well >> uh and it's become critical infrastructure for our team
19:22because everybody uses it >> and you went pretty far with that too right like you had a um in your talk
19:28I saw that you actually built a custom CLI for that too.
19:31So what does that CLI do?
19:32Like how does it execute things and why did you I mean that's proof of how deep you're going right of how
19:37much you're pushing that. >> Yeah.
19:40So this is actually a tip I learned early on where um I guess you know back in January or or late
19:50last year the thing that people were concerned about was context window, right?
19:54That was the big the big topic at the time was how do I you know manage the context window because you
20:00know compaction summarization wasn't really that good yet and people were always people had this there was almost this meme in the
20:08community that you know once your agent summarized or compacted once it would become sort of stupid right for the rest of
20:16your session. So there was a lot of thinking around like you know being very efficient with your context usage and so
20:24that was actually the inspiration for some of the uh the CLI work inside of the verification skills.
20:32I guess now it's less so about context because uh you know agents are much better or harnesses have gotten a lot
20:40better with summarization. Um, I still think there's some benefits to, you know, uh, having a clean context window.
20:50Uh, so the CLI is really just more of a way for me to take the deterministic parts of what the skill
20:56does and encode that into a script or CLI to reduce to kind of take away the judgment that would otherwise unnecessarily
21:06be used because with judgment so I also think of you know agents and skills in sort of like it's like a
21:13gradient you have some parts of the work that are entirely judge measurement based right you know something that requires thought you
21:22know putting together multiple pieces of context thinking um and then you have the more deterministic parts like I don't know if
21:29you wanted to uh refactor some code right from one pattern to another that's very mechanical right you don't you don't need
21:37an agent to think about it and come up with it in a novel way each time right and so that that
21:44was really the inspiration for the CLI and you'll see this in a lot of the other skills that I built is
21:49like I try to extract out the deterministic parts and turn that into code and just leave only the parts that actually
21:57require judgment to the agent. So in a way I think of the seal as kind of like your wrapper, right?
22:02It's a wrapper with some light instructions around how to use these custom tools that are inside of the skill.
22:09Um but yeah, I don't think the CL is really that interesting in its own really.
22:13It's not like a novel piece of software.
22:15It's just something that interacts with like Playright and the Chrome DevTools protocol and calls a bunch of APIs.
22:22It's like it's just a bunch of glue.
22:24>> No, it's fascinating because it's a way of hiding information from the skill, right?
22:27It's a way of conserving the skill, keeping the skill quite small, I imagine, and then you're able to delegate more of
22:33the complicated deterministic stuff into a script within the skill.
22:37So it's almost you're compressing information and making the agent do more consistent things more >> that's fascinating.
22:47And it it also helps I guess if you care about context window it it does help because now the agent doesn't
22:53need to uh you know re reinvent uh things cuz uh one thing I had noticed early on when we didn't have
23:00a CLI was that uh well the the agent would try to verify it work but it would basically rebuild the world
23:08each time and then every agent did it differently and I was starting to notice like that's very inefficient right I was
23:13wasting it it was actually not just about context usage but also speed, right?
23:18Like because now an agent had to actually go off and write the scripts or the CLI and test it and you
23:23know and it doesn't work and the last agent did it and it worked but it discarded it.
23:27So it was just very obvious at that point like I should just turn this into a CLI and put that inside
23:32of the skill uh so that every agent that uses it now benefits from that same piece.
23:39Um but I I also think like you know it's a good push for people to think about is how much of
23:46your skills and rules could actually be Um that's like another core thing or or one of my core principles that I
23:55like to think about is yeah how do I uh make very efficient use of determinism and and you know let Asians
24:06shine at the non-deterministic parts right because that's what they're trained to do.
24:11Um and the other parts which are much more mechanical or you know straightforward can be just pure determinism.
24:18Um, and you'll see this as well for things like doing migrations.
24:22Um, which is another big thing that I've I've talked about is, you know, going from one technology to another, especially one
24:31that is better for agents, right?
24:33And a lot of how you can do that migration is, I think, through things like scripts and CLIs, like the deterministic
24:40parts like code mods, you know, like crawling the abstract syntax tree and transforming code literally mechanically, right?
24:48like a script does it for you instead of the >> Totally makes sense.
24:53I I mean I think what there's another thing there which is you're taking stuff away from the agent and you're kind
25:02of putting it in the environment too a little bit which is let's say you have a a thing that you notice
25:08the agent always gets wrong. You want to make that um just impossible within the environment.
25:14And that sort of comes down to code quality as well.
25:18I mean, I talk about a lot like having a what a good codebase means, right?
25:24What is a good codebase? And there's a definition I like which is a a good codebase is a codebase that's easy
25:30to make changes in, right? Easy to um change stuff without things screwing up.
25:36And that means that you have a lot of guard rails that you have a lot of um the agent or the
25:41human is constrained to very narrow paths.
25:44And that's again something you talk about in your talk.
25:46>> And you talk about this not only on the kind of sort of automated checks side of things.
25:52So linting and type checking blah blah blah but also in the way you design abstractions.
25:56And you guys even I think built a framework uh for your agent to work in too.
26:01>> I think what I'd love to hear is you obviously think of that as very important, right?
26:07And that's how important is that compared to other things you could be doing like building features or shipping work.
26:14Yeah, I think that's a um I almost feel like the new job of the engineer is really to to spend time
26:20on the environment. Um I almost actually wrote a tweet about this yesterday, but I but I didn't.
26:26But I think that I think that you you know if you if you haven't really spent time, you know, building trust
26:33in your agents and building skills and tools, you can get stuck in this mode where you're very low on that trust
26:40ladder, right? you don't have a lot of trust in your agents work.
26:43And so the only way to cope in that when you're in that situation is just to kind of lock in and
26:50micromanage your agents. And that's very time consuming.
26:53And when you're stuck in that mode, you don't really have the luxury to think about, you know, uh higher level things
27:01like like making yourself more productive.
27:05In the same way that uh I guess analogy would be like if you've never taken the time to learn like your
27:12tools right as a developer when you were writing code yourself and you know you've never heard of VS Code, you've never
27:18heard of Vim, you only knew about Notepad uh and you had hadn't even heard about Git.
27:24That's sort of the analogy. It's like you you haven't spent the time sharpening your own knives, right?
27:29And so, of course, if you have a dull knife, then everything's going to take a long time.
27:34Um, and you're going to be you're just going to be and and especially if you know deadlines are looming, then you
27:41don't have the now you're stuck in this rut, right?
27:43Where where you you you don't have sharp knives, you don't have good tools, but you're under all this pressure to ship,
27:50right? And so, all you can do is just focus on that.
27:53But I do think that, you know, if you can find yourself the time to actually spend time thinking about your setup,
27:59it's again going back to the cooking, you know, it's like uh, you know, if you, for example, if if cutting cutting
28:09cutting the garlic is like super slow, right?
28:11There are garlic mashers, right? You can you can buy and you put it in the thing and you like squeeze it
28:16out, right? It's super fast. Uh, machines and tools were invented for a reason, right?
28:22And so if you're operating a Michelin kitchen and your your your cooks had no tools, then of course everything's going to
28:28be extremely inefficient, very very, you know, every every every cook is going to make something up of their own.
28:35So I think the tools and the determinism to me are you know taking that part away and and just like you
28:42said about constraints as well. It's the constraints are are to me as well like uh actually a slight tangent on that
28:50is uh I think we should talk about TypeScript cuz like we we actually both share like a background in Typescript where
28:57you know you obviously have done a lot of work with TypeScript and total TypeScript and you know you're a leader in
29:02that space and I uh had adopted TypeScript pretty early and I had given like a talk or two at Typescript conf
29:10uh many years ago and so one of the the talk that I did actually was about type systems and constraining the
29:18constraining types. Like one of my most favorite things about Typescript is actually type narrowing, right?
29:23This idea that you go from a very broad type, right?
29:26That could be anything and then you through type guards and you know type narrowing and you know runtime checks you can
29:34actually narrow the space and say like oh this isn't just a string this is a very special type of string.
29:40It's a constant, right? like I but I I determine that through the type system and in a way it's like uh
29:46there's a lot of parallels I think to that with constraints in your codebase where is it you're you're constraining the space
29:55right if you if you think about category theory as well you know you're constraining the the number of possible types right
30:02that can can exist and you're saying there's only one type right and for for us like that framework that I'm called
30:11Dune. Uh it's not an open source framework.
30:14It's the the way I describe it to people.
30:17It's it's kind of like a internal Nex.js for our Electron apps.
30:22Uh but it comes with a lot of really really restrictive lit rules and the codebase is designed in a way that
30:29there's really only one way to do something.
30:31So we make use a we make use of a lot of conventional patterns.
30:36So like features all go into a specific directory.
30:40Well, every feature has its own directory.
30:42As an example, you know, there's like a a thing that discovers features like through a registry and like crawling the codebase
30:48and stuff like that. But this conventional pattern and the lint rules make for an environment where it's actually very hard to
30:58write bad code. And that sort of frees up the it both frees up your own mental uh you know capacity as
31:09well as the agent sort of doesn't have to think about that anymore where it's just like oh there's only there's I
31:16should just if I want to add a new feature it just goes in the feature the new feature directory and all
31:20the code goes in there and I'm not going to append to a god file right that was really actually the inspiration
31:26for those feature directories is the very first couple of versions of Grockbot were composed of like eight god files which were
31:36like at least 10,000 lines long if not longer and so I kind of had to break it up into smaller pieces.
31:43Uh but it was just observing you know actually that's another important part is observing how agents fail and then every time
31:50you see a mistake every time you see something that could be done better you think you step back and think how
31:57do I turn this into a lint rule?
31:58How do I make it so that the code base makes this impossible?
32:02Right? And it comes back to me for my you know my background learning Typescript and types uh type systems is how
32:10do I constrain the space so that you know I know precisely what I'm working with and I think yeah there's a
32:16lot of parallels there. >> Totally makes sense.
32:19And don't I mean it's funny that you mentioned TypeScript and Goth files in the same sentence because Typescript famously has a
32:2525,000line type uh file. Although I don't know if they've rewritten that and go as they probably have, haven't they?
32:33Um, okay. So, environment is important.
32:37You should watch your agent like a hawk to make sure that any mistakes it makes.
32:42You turn them into things in the environment.
32:44And the benefit of the environment is you're not overloading your agent, right, in terms of rules, in terms of things it
32:50has to remember. It's just in the environment.
32:52And so it stumbles into the rules and exactly um you know bounces off them and hits them at the right moment.
32:59>> So okay, we still haven't talked about the 2,500 PRs.
33:02Where do those come from? Like how do you you've built your trust ladder, you've worked on your environment, and you understand,
33:09okay, um I now want to scale up.
33:12So what are the mechanics of that scaling?
33:15Are you um initiating 2500 like chats per month?
33:21That can't be right. So there must be are there any kind of automated triggers that trigger stuff in your repo?
33:27Like how do you get the software factory kind of triggering work by itself?
33:34Um I'll definitely say that the prerequisite to you know something like a very high volume of of pull requests um is
33:44the environment. you know, the the stuff we just talked about where I definitely would not have been able to do this
33:49if I had not spent the time, you know, thinking about the kitchen, right, and the knives and the tools for my
33:55Asians. And so, in a way, I I think of this as I've spent the time building one kitchen and one restaurant.
34:04And now I I'm in a position where I don't actually have to be there anymore because the environment, you know, that
34:11the same analogy, right? It works really well.
34:14Yeah. You open chain of restaurants, right?
34:16That's >> Yeah. Exactly. Yeah. Exactly.
34:18It's like you're Gordon Ramsay and you know, you you've taught your executive chef like all the tricks of the of coming
34:25up with great menu. Uh and like the kitchen is set up really well.
34:29Everything's just perfect and you're now in a position where you can open your second your third restaurant.
34:35And I guess I I sort of see each project that I work on, like each big chat is sort of like
34:41a restaurant, right? and and I'm I'm I have multiple of them operating at the same time and I'm sort of like
34:48helicoptering between them sometimes some more than others depending on how in the loop I am but yeah definitely I think there's
34:56there there are external triggers and context that those projects don't have that for a long time I was the proxy for
35:07that so uh the best example I have is like you know you have a project that's working on a feature uh
35:13or you're trying to fix a bug and you're getting bug reports, but the bug reports are going to things like Slack
35:19or linear or X, right? And these are external systems that aren't connected to your inner loop.
35:27So, I like to talk about this outer loop and the inner Uh I don't know if I'm using the definition correctly
35:34but to me my inner loop is like basically my engineers my agent engineers working on the code to an building towards
35:43an intent or snapshot of my intent right and the thing about that is that the snapshot can go stale right new
35:51information comes to light that I then have to be the proxy of and you know transfer that context to my agent
35:58so you know if you if you don't have these triggers pulling information back into the interloop, then you sort of have
36:05to play that role where you're you're off, you know, in Slack or X or or whatever and you're gathering context, right?
36:14You're getting context about bug reports, about feature requests, about, you know, something someone said about, you know, our backend infrastructure has
36:21some limitation, you know, all that information, you have to f that across to your agent.
36:28So that's where I think like tools like Grogbot are really good because they help you automate the outer loop as well.
36:36And when you connect those two loops, it's very very powerful because now all of a sudden your agents have the ability
36:42to get context for this for themselves, right?
36:46If for example uh you know either through just as a simple example like maybe you have the Slack MCP, right?
36:54Or you have uh your own harness, right, that you've built a Slack subscription into for a particular Slack channel.
37:02Now all of a sudden you can tell your agents, okay, subscribe to the Slack channel.
37:06Every time there's a uh, you know, bug report about something, go off and triage that thing, right?
37:14Go reproduce the issue, right? Using the verification skills that we've already spent time building and all of those other skills that
37:21we've set up so that I have a lot of trust, right?
37:24I have a lot of trust that these agents can actually go off and understand the bug, you know, uh verify that
37:31the bug actually still exists on main and it something about you maybe the users setup or their data or maybe I
37:40don't know they didn't install a dependency or something like that like uh basically I think uh creating that yeah creating those
37:49two loops and connecting them is really a very important part of the job these days.
37:54Um, especially if you are thinking about how to scale yourself.
37:58So, a big theme here is really just like always thinking about like what where am I the bottleneck in this process?
38:06Why do my agents need me, you know, to answer this question?
38:10I I always like to think about that.
38:11And so I try to think about how do I actually get the agent to answer its own question, right?
38:16But not by hallucinating, not by guessing, but actually real data.
38:21And you know, a lot of people talk about this idea of a company brain, right, or a context graph.
38:26I feel like those terms are complex uh or even abstract.
38:33To me, it's just about um how do I take information that my agent needs that I would otherwise have to go
38:40and pass it myself and just teach it how to do it, right?
38:44And that removes me from the equation.
38:47And so how I arrive at 2,000 or however many PRs is the fact that I have all these loops set up,
38:54right? And so uh it allows me to open chain restaurants, right?
39:00I can I can really parallels myself.
39:03So yeah, I'm not sitting there creating 2,500 chats, right?
39:06Of course, it's really like these projects are um actually cursor has a new feature called projects which are these like coordinator
39:14agents. Um and so the coordination co coordinator agents are really good at sort of delegating and not doing work of their
39:22own but they manage and supervise like almost a list of tasks and they spawn sub agents to go and do them.
39:28And so I'm just constantly feeding context or teaching the agents how to get their own context and then they're going off
39:34and doing the work for me.
39:36Uh and really the big the last thing I'll say to this is like the big unlock for me for getting to
39:412,000 PRs is starting from the question and working backwards of how do I get to the point where my agent can
39:49merge its own code? the obvious thing people ask me when when they when I tell them, "Oh, I shipped 2,000 and
39:582,500 pull requests last month." They'll be like, "How did you review that?" Right?
40:02That that's a lot of PRs to review.
40:04Like your team must hate you.
40:05>> Do do you mind if we go there in a second?
40:07Because a good question about >> Yeah.
40:09Yeah. Yeah. >> I want to like this analogy is great.
40:12I want to like deepen it a bit which is before if you're like manually initiating all those chats it's like you're
40:18bringing the orders to your chefs manually right whereas if you've got an agent sort of like doing the expo then you're
40:25able to sort of run it yourself itself what is what does that concretely look like then you've got these sort of
40:31grock bots that are um subscribing to channels pulling in Slack messages and you it sounds like have a couple of coordinator
40:38agents or like chief of staff agents that like monitor that or something like when you look at your computer to manage
40:45your agents, what does it look like?
40:49>> Yeah. So, so uh this is I guess somewhat confusing but we're working on you know simplifying and unifying but so
40:57uh there's graphbot uh which or you know you can use other tools of course as well but I I largely think
41:03of these tools as like your outer loop.
41:05These are tools like you know Grabbot that have connectors right these are connectors I guess they a lot of people call
41:11them personal agents um but they're connectors to things like your email your calendar slack uh plaid I don't know like all
41:21these different services and they are a great source of pulling context in to your work so the same way that a
41:30human like you know if I were if I was a manager and I was leading a team of engineers years.
41:36Um, you know, like when I used to work in Netflix, one of the biggest things that managers would talk about was
41:41this idea of context not control, which funnily enough, you know, has so much uh has so much uh carry over to
41:50the agents world. Uh, of you know, you you know, you you of course can drive to an outcome you want by
41:57control, right? Like by micromanaging, but what you want is to provide context instead, right?
42:02like teach the agent, teach your engineers how to be self-sufficient and then you don't have to micromanage them.
42:09>> Um, and so I see a lot of parallels there.
42:12Uh, but yeah, graphbot. So, concretely, I have some graph bots that look at my Slack channels, look at my X, uh,
42:19or my emails, uh, or linear, and they're just constantly they have routines that subscribe.
42:26So they're constantly watching and I have I I'll tell them things like you know uh I'll watch for issues with uh
42:33bugs in the graphbot desktop app as an example.
42:36Uh and whenever you find that send it to my cursor project.
42:41So one of the really cool things about grabbot is it connects to cursor.
42:45So cursor has uh like I I just mentioned this new feature called projects.
42:50And a project is really a uh again like a you get a coordinator agent that's in the cloud.
42:56It has its own computer and all it really does is like it's a manager of agents.
43:00It's like your executive chef, right?
43:02Your your chief of staff. It doesn't do the work itself.
43:06It delegates and orchestrates and manages the work of other sub agents to you know that report to your chief your chief
43:16uh of staff. And it basically is responsible for driving the work forward and managing things and uh passing context to them.
43:27>> So if you get a sudden burst of issues, let's say you get 30 issues at once in one payload or
43:31something or very quickly the coordinator agent can figure it out and delegate.
43:34>> Yeah, exactly. It gets like uh you know 30 the 30 or so payloads and spawns a sub agent or a
43:40single coordinator agent. It can actually do a bunch of different topologies of agents and it will sort of figure out the
43:46best way to uh you know efficiently distribute the tasks to your team of agents.
43:54Um so I use uh cursor projects a lot um and I also use grapot a lot and cursor projects are my
44:03inner loop and grabbot is my outer loop.
44:05Grabbot takes all the context, external context, gives it to the projects because it can actually just send messages to those projects,
44:14right? You don't even have to open cursor.
44:16You can just tell your grabbot, okay, create a project, right, for these series of tasks.
44:21They're all related, right? Maybe as an example, you know, you've had a uh a big burst of issues that are all
44:28about performance, right? Your app is slow uh and they're all connected, right?
44:34Maybe some of them even have a similar fix, right?
44:37But and you can certainly go off and just spawn one agent per task, but then you've lost that sort of thread
44:43between them, right? And and you may duplicate work or you may not really think about the higher level problem.
44:49You know, sometimes when you you you solve bugs, you know, it helps to have multiple bug reports that are are slightly
44:55different because it helps you really, you know, zoom out and see actually, you know, the problem when I looked at this
45:00one report, I thought the bug was here, but actually when when I see the other multitude of bugs is actually up
45:06here, >> Yeah. Got you. So that that's why you have so many agents in that loop then, right?
45:11Because it's not just you have um like you have a bug report comes in, you spawn a single agent to look
45:16at that bug report. that a that single agent will be duplicating work with other um other agents, right?
45:22Because if there are multiple bug reports coming in through the same thing, that can be duplicated >> That's really fascinating.
45:29Okay. And so this just this endless series of triggers um coming from real users reporting real reports um builds up this
45:38sort of and accelerates the factory sort of adds more orders in.
45:42Other than bug reports, are there any other sources that you use for like um accelerating for pushing these PRs?
45:51>> Uh well, funnily enough, it's some of it comes from uh reading the code, too.
45:56So, I guess I have sort of uh well, so to clarify that, you know, the 2,500 PRs, they're not obviously like
46:042,500 features, right? they are a lot of the work actually is spent on gardening like another term that I really love.
46:14Uh so I guess this is more important when you have a big team of engineers human engineers that you work with
46:23where and also this goes back a little bit to what I was talking about with the environment.
46:27You know, setting up a really good environment that doesn't just help you and your agents, but everybody on your team, right?
46:32Think of a new hire who doesn't have a lot of context on all of your engineering practices joining your your team.
46:39And if you have a really good environment, they can be productive from day one, right?
46:43They can they don't have to like, you know, make open a bunch of lowquality PRs.
46:47They can start, you know, they can start just turning out really good code.
46:54and uh, yeah, I think I I sort of lost my train of thought.
46:58>> I've got a I've got a followup, which is what's what are the mechanics of like >> like how when when
47:05do you trigger a a to go and look at the code, right?
47:10Because some people might say, "Oh, let's just do that every hour or something or like on a chron job or >>
47:16Oh, yeah. Yeah. Yeah. Yeah. Uh I saw some of your recent tweets as well about like you know the some of
47:21the tweets you've been doing which are great for setting up your routines.
47:26Uh I have some routines like that as well.
47:29Um so uh one of them is uh like looking through just another simple example is you know React has a lot
47:41of foot guns. Um, so, uh, as as I'm sure you're aware.
47:45And so I have an agent that's just constantly looking for band patterns.
47:49And the interesting thing about that one is that I don't actually tell it to fix the issue first.
47:54I tell it to append it to a document.
47:56And then every couple of days I look at it and I see actually these are all the same thing, you know,
48:02and so that gives me, you know, you almost want like a buffer, a queue.
48:06Sometimes that's actually more effective than just spawning off a couple of like a lot of sub agents to fix every single
48:12thing because when you are in kind of pure execution mode and just trying to like you know f uh you know
48:20execute on the orders that are coming in very fast you sometimes miss the big picture.
48:25So sometimes having a buffer forces you to think about the big picture because you you you have these artifacts and things
48:32that you can look at as a human um and sort of use your own human judgment to or I guess you
48:39can use an agent to do that as well.
48:41But you give the agent and yourself a way to identify patterns, right, that you might otherwise miss if you're just only
48:49solving each bug at a time.
48:52And that's also really the benefit of having something like a chief of staff agent is uh it can see the forest
48:59right uh in addition to actually doing the >> Fascinating.
49:05That's I mean my brain is exploding a bit there with the sort of chief of staff at the software factory.
49:10I might have to change some of the course that I'm filming next week.
49:15>> uh all right. Let's talk about let's talk about review, right?
49:18because this is the reply that you get, you know, is >> did you read did you taste all 2500 of those
49:24dishes as they swept past you?
49:27>> And I assume the answer is a variety is a version of no.
49:34>> Yeah, I think you you you don't want to be in a position where you're not tasting your food ever again.
49:39Uh but you also, you know, for scale, you cannot be tasting every single dish that comes out of your kitchen, especially
49:45if you have multiple restaurants. So it becomes more about sampling right and thinking about the processes in the same way that
49:53you know if I guess maybe this is where the the the factory analogy is a bit more apt is you know
50:00as a quality supervisor on a factory you you can't look at every single item you sample right you take you you
50:07you go in there every day and you look at the quality of the pull requests you look at the code that
50:12the agents are writing and you scrutinize it very rig rigorously and you think about all the inefficiencies, the bad patterns that
50:22the agents are doing and then you think about how to course correct the environment, right?
50:27Not not that single agent. Uh because if maybe if it if it was a one-off incident, it's fine.
50:35You know that maybe there's nothing to fix there.
50:38But if you actually notice that multiple agents are are having the same issue, right?
50:42They're taking the same shortcut. they're they're propagating the same workaround everywhere.
50:47Uh that's a sign that you should go off and think about how to uh amend your kitchen or your factory, right?
50:55Like thinking about your skills, your constraints, your lints, your type systems um and setting or adjusting it so that that problem
51:05doesn't happen again. And when you do that enough times, then you get to a place where the codebase is again like
51:10the environment is so constrained and so it guides you so well that you can just you can just step away, right?
51:19That's the dream. And I'll I'll definitely say it um it's very hard to get to this point.
51:25I don't want to sell this as like, you know, something that you can just do easily by using PAC.
51:30Like it takes a lot of time and effort to think about your code and where you see your agents failing and
51:37thinking very thoughtfully, intentionally and setting up guard rails and constraints so that they do the right thing by >> And you're
51:47not like if to go back to the software factory analogy, this isn't a dark factory, right?
51:53This the lights are on, right?
51:55>> It kind of is. Yeah, actually.
51:56>> Is it? >> Yeah. Well, it's dark in the sense that so um it's dark in the sense that well I
52:02think if my agents are merging their own pull requests it's sort of become dark where I go to sleep my agents
52:09now work uh I have I have the equivalent of like more than 10 chiefs of staff right each working on a
52:17different area like for example I have one that's working on performance of the Grockbot desktop app I have one that's working
52:24on uh fixing bugs that users report I have one that's exploring rewriting it in a different language just for fun, you
52:32know, like what if what if, you know, just reimagining what what it would be if it was like a native app.
52:37It's just a toy. Um, but the idea is like yeah, I uh I when you spend the time setting up your
52:43environment, I've gotten to a point where I review the pull request after it's landed, right?
52:48I I tell my agents full autopilot is is something that you can do in in PAC and that will trigger off
52:56this very intense rigorous verification loop where it will spawn a bunch of verifier agents for every pull request and it will
53:04fuzz right fuzzing meaning that it will actually run the application.
53:08It's going to click around and try to use it like a real human.
53:11look for regressions, look for bugs in your implementation and um it will try to find issues with the thing and then
53:19it will fix it itself. It'll do that again and eventually get the PR to a state where it can land.
53:26Uh so it does it does it is quite token intensive.
53:30You can tune this of course.
53:32Uh so you know instead of like 10 verifier agents you might do like one, right?
53:37Or you just tell the agent to verify it's done work.
53:40But yeah, the key thing is the verification part is really the key piece that gives me a lot of that I
53:48guess verification plus the environment, right?
53:50It's these the combination of these two things that allow me to step away and say agents go off and merge your
53:55thing. I'll review it in the morning by looking at my commit >> and if I see problems, I go and course
54:02>> right? And I'll go and revert or modify, add new link rules and whatever.
54:09Um, so it does it does take time to get to that point, but once you get it, oh, it's so it
54:14feels so magical. Uh, I I I tell people like I'm sleeping so much better now because, you know, it took it
54:22the very first day I turned on the sort of dark factory was very scary because I was like, "Ooh, what if
54:27I call the SE, right? What if I break something >> Uh, and it took a lot of it took a lot
54:34of uh bravery, I think, to do that, but >> somehow I did it.
54:37And yeah, now I'm in a place where my Asians are are merging their own code while I sleep.
54:43>> It sounds like >> I think it's dark in that sense.
54:46>> Yes, it's dark sometimes, right?
54:47You do >> That's true. That's true.
54:49>> Because I think of a dark factory is like almost like if you take the original definition of Kapathy's vibe coding,
54:56right, which is the code almost doesn't exist.
54:58You forget that code might be a thing.
55:00I think your approach is totally different from that, which is that code and the environment is essential.
55:06And if the code in the environment are bad, then you will get bad outputs.
55:10Garbage in, garbage out. So I I I think this is a this is a different thing.
55:15It's like, you know, the I don't know, maybe there's a dimmer switch or something, right?
55:21Like, you know, some parts of dark, some parts were light.
55:23This is why the maybe the the kitchen is a better analogy >> the restaurant because you know even as a as
55:30a as a as a restaurant restaurant you still might go to your restaurants every now and then to take take a
55:38peek in taste the food right >> uh I think that's >> the idea of sampling instead of blocking I think is
55:43really important >> I think what would you say to people who are in I guess you're obviously in a pretty security
55:49conscious environment where you're working very security >> mh Um maybe there are folks working in like um medical applications or law
55:59or finance or something. I think of the like some PRs are kind of like two-way doors which is you can merge
56:07it and then revert it, right?
56:08It's cheap back through. But there are some PRs that are one-way doors, right?
56:12That will cause data loss of some kind that will >> do something that can't be easily walked back.
56:19How do you deal with situations where most of your PRs, let's say, are one-way doors?
56:24Like, is this something you just wouldn't recommend or like what do you think?
56:28>> Yeah, I think that's a really good question.
56:30I think that it all comes back to me to the quality of the verification that you're able to um get out
56:39of your agent. And I think for domains where the work is this is easier, right?
56:47and the the oneway doors become two-way doors in a But I guess I don't know if you're working on something that
56:55is like is very hard to verify programmatically then I think yeah you're definitely in a position where it's very hard to
57:04get to that point. Um so I do think like yeah verifiability of the domain is an important aspect to be able
57:12to do this. Um and software engineering is just one of those things where it's quite verifiable in in a lot of
57:18cases maybe not totally um you know like other domains like mathematics I think are another example of not all of it
57:26of course but some aspects of mathematics can be verifiable if you write a proof for example um and so yeah I
57:37think it's a great question that I don't really have the answer to and I think that this is something the industry
57:42and us as engineers will have to figure out is you know my sort of uh hope and prediction for the future
57:49is that we'll see more and more interesting new agentoriented programming languages and one of the most fascinating ones that I've seen
57:58so far is this one called bend bend d bend um and that language is one where it kind of marries programming
58:08with proofs right there used to be a time, you know, where you actually had to write your proofs in a different
58:16language. And proofs, by the way, for those uh who who aren't familiar is this idea of uh that you can sort
58:23of formally verify that some code is correct mathematically, right?
58:28Especially if you've written your code in a very functional programming way.
58:34uh but for the longest time you had to do that in a separate language like lean or tla+ or uh I'm
58:41blanking on some of the other other examples but uh like languages like that where you would construct the mathematical proof and
58:48then use a solver essentially to det that that you've covered all the cases you don't have like a race condition or
58:56or whatever. So yeah, I think trying to sum up the question, I think yeah, if you are in a position where
59:05you can figure out how your agents can truly verify the work in a way that gives you confidence, you can actually,
59:13you know, uh have the PRs merge cuz if it compiles, right, if it if the proofs show you that it's correct,
59:20then why wouldn't you just merge it?
59:22Um but of course, yeah, not all the means are verifiable.
59:27Yeah, it's a tough one. Um, okay.
59:32I think we've got to think about wrapping up because we are nearly on the hour.
59:36Have you Have you got something after this?
59:37I mean, I've got something before I give my son dinner, but >> I I can go a bit longer after you.
59:42>> Okay, let's let's go five minutes longer then.
59:44Um I think I just want to have one more question which is I think I want to ask how you see
59:53Pstack and how you see skills in general like in terms of we talked about this before we went on air which
1:00:01is like people think of as like my skills versus your skills and how do you combine frameworks together?
1:00:09How do you use Pstack with my stuff?
1:00:12like what should you take from each one and because I think I see skills as sort of just derived from process
1:00:20basically like they're just processes turned into words and I would love to know how you recommend people take Pstack and take
1:00:29my stuff as well and turn it into their own processes.
1:00:35I think you shared a tip actually today that I thought was actually very relevant, which is this idea that you go
1:00:40off and look at your previous transcripts, right?
1:00:43And you sort of mine for information of, you know, your own look through your own your own prompts, right, to the
1:00:50agents where you correct them where you have to constantly intervene and uh you know take that higher level learning and turn
1:00:58that into a a reusable skill, right?
1:01:00So that agents stop repeating that mistake.
1:01:04I think that uh PAC and your skills are very complimementaryary.
1:01:09I I totally agree with you that they're like a skill is really much just process.
1:01:13I mean it's just at the end of the day a skill is just English or or language.
1:01:17It's just >> Um and I think you can you can definitely weave them, combine them in a way that makes sense
1:01:25to you. But I do think that uh everyone should have their own set of knives, right?
1:01:33I I keep going back to the the the cooking analogy, but it's so apt because like, you know, every chef when
1:01:39they go to a different job, right, when they go to a different restaurant, they carry they bring their knives with them.
1:01:44The tools go with them, right?
1:01:46And so trust to me is really about trust in your own tools.
1:01:49And when you spend the time sharpening them and understanding them really, really well, you can do great things.
1:01:54And everybody's skills and tool set is going to look different.
1:01:58you know, someone might find a lot of success combining, you know, like your grill me with docs, uh, or wayfinder skill
1:02:06with some of the execution skills in Pstack as an example.
1:02:09Some people might use more of your skills, some people might use more of my skills.
1:02:13I think at the end end of the day, it really just comes back to how much do you trust, you know,
1:02:18me and Matt, right? Like if you if you trust us both, of course, use our skills, but I also encourage you
1:02:24to, you know, look at your own transcripts.
1:02:26Um um tell the agent to look through, you know, some of the all of the patterns that you've used, the the
1:02:34times you've had to intervene, you know, suggest turning them into lint rules or new skills, right?
1:02:41The the past chats I I often say is like a a treasure trove of context because that, you know, it's it's
1:02:48like the process materialized, right? Like it's the real process.
1:02:52It's not an abstract idea in your head.
1:02:55it's the actual thing right and you can actually see how it happened in practice and extract so much information from that
1:03:01and there's so much so that I actually turn I have a skill in pac called recall which is exactly that um
1:03:07where this was a pattern where you know I was working in I was working on a similar problem so specifically I
1:03:14was working on virtualization for the cursor application and there were a lot of bugs and so you know every time I
1:03:21started a new chat I was like ah this there's so much good context from the last one So, you know, I
1:03:25want to bring it over to the new chat.
1:03:27How do I do that? And that's where the transcript came, uh, you know, looking at the past transcript came about.
1:03:32And then recall was just a way for me to collapse collapse and compress that workflow into a skill.
1:03:38So that I didn't have to just say I didn't have to write a long essay every time.
1:03:42Go look at all these chats, right?
1:03:43And, you know, blah blah blah.
1:03:45So, I I largely think of skills, especially as agents get more capable as really encoding workflows.
1:03:52you know skills from last year were really more about like almost like implementation details like here are the exact script commands
1:04:00you know you should use right I think with the latest models you can just delete those parts and just really focus
1:04:06on the workflow right until it it's more the skill becomes more like a series of steps a series of your process
1:04:15uh and I think over time we'll see that skills get smaller and smaller you know more compact Um, and yeah, they're
1:04:24very compatible. Or you can, you know, if you want, why not read our skills, right, and com and combine them in
1:04:31your of your own, right? Combine Wfinder with potato mode and make your own custom mode, right?
1:04:36Like like skills are the the thing I love about skills that is are that they're so malleable.
1:04:42You can do anything you want.
1:04:43It's just language. >> Absolutely. There's nothing magical in them, right?
1:04:46They're just words. And >> Exactly.
1:04:48If if there is any magic in them, it's just the words chosen and the phrases used and the thinking that's been
1:04:56done to turn those like take abstract process and turn them into language.
1:05:02And once that thinking has been done, then it's just there.
1:05:05It's available. It's on the surface and you just nick it.
1:05:08Um Lauren, thank you so much.
1:05:10This has been glorious. >> Yeah, this has been super fun.
1:05:12I really enjoyed talking to you.
1:05:15Hope we can do it again.
1:05:17I'd love to do it again.
1:05:18I'd love to do it again.
1:05:19Absolutely. Um yeah, we'll check in in uh >> I don't know.
1:05:23Yeah, Monday. Let's do it. >> Yeah, let's do it.
1:05:27Part two. >> Well, thank you so much.
1:05:29I'm going to close the stream here.
1:05:30Laura and I will uh chat a little bit and stay here.
1:05:33But thank you guys so much for watching.
1:05:34The