← All videos

Building pi in a World of Slop

16 April 2026

Mario Zechner · AI Engineer · watch on YouTube ↗ · click any timestamp to jump the video

Most interesting ideas

Machine-generated by an AI from the transcript and the top comments. Not my writing, and it may contain errors.

Eighteen minutes, three acts, almost no padding. Act 1 (to 10:57) is why he abandoned the mainstream harnesses and wrote his own; act 2 is a short detour on bot-generated noise in open source; act 3, from 12:02, is the argument people quote him for and the best part of the talk.

The genuinely fresh ideas

▶ 1:55Your context is not your context

The bugs and the flicker are symptoms; the real objection is that the harness owns your context and edits it behind you. System prompts and tool definitions change with each release, tools get removed or modified, and reminders are injected mid-context telling the model the information may or may not be relevant to what you're doing — which he argues confuses the model and quietly breaks workflows you had working.

▶ 4:42The most minimal harness is near the top of the leaderboard

Terminal-Bench gives the model one capability: send keystrokes to a tmux session and read what comes back. No file tools, no sub-agents, none of it. That harness outscores much richer ones, and often beats a model's own native harness — the strongest empirical point in the talk, and an uncomfortable one for anyone adding features.

▶ 6:58The model already knows it is a coding agent

Models are post-trained inside coding-agent harnesses, so the harness is the thing they were trained in rather than something that needs explaining. Spending ten thousand tokens establishing the role is therefore waste. pi ships a system prompt of a few lines and four tools — read, write, edit, bash — and he puts the tool definitions on screen specifically so you can see how short they are.

▶ 8:10Extensions are TypeScript files, distributed on npm

An extension is a module on disk that hooks events, defines tools and slash commands, replaces compaction or providers, and hot-reloads inside the running session — a habit he traces to game development, where iteration speed is everything. His aside on distribution is pointed: package managers already exist, so there is no reason to invent another marketplace silo.

▶ 11:31A filter that works precisely because bots don't re-read

Automated pull requests are closed with a comment asking for an issue written in the sender's own voice, under a screen of text. A human reads it and gets added to an allowlist for next time; a bot never goes back to look at the comment it received. The filter works not by detecting bots but by requiring the one behaviour they don't have.

▶ 14:16A sufficiently detailed spec is a program

The neatest line in the talk, aimed at the "but my spec was detailed" defence. Whatever you leave blank gets filled in, and it gets filled in from the average of public code — so the gaps in your specification are precisely where the mediocrity enters.

▶ 14:46Humans are valuable because they are bottlenecks that feel pain

An inversion worth sitting with: people are fallible, but there is a hard limit on how much damage one can do per day, and they suffer when a codebase becomes unpleasant. That suffering is the feedback signal that eventually triggers a refactor. Agents have no such limit and no such signal, so nothing stops the accumulation.

▶ 15:53Patches land locally and break globally — and the tests are suspect too

Once the codebase outgrows what any agent can hold in context, every decision it makes is a local one, and fixes start breaking things elsewhere. His sharpest consequence: you cannot fall back on the test suite to catch it, because the agent wrote the tests as well.

Quieter but sharp points

▶ 7:41A confirmation dialog is not a security model

pi runs unsandboxed by default, on the reasoning that a prompt asking you to approve each bash call trains you to click through it. He would rather hand you enough rope to build the isolation your situation actually warrants than ship a gesture at safety.

▶ 12:41Errors compound, and the pain arrives later

His framing of the arithmetic: one human, one agent and ten agents produce very different error rates against a review capacity that doesn't move. A review agent only catches some of it — pointing an agent at agent output is a snake eating its own tail.

▶ 13:45Agents learned complexity from our old code

The taste for layered abstraction, duplication and defensive scaffolding came from training on public code, most of which is nobody's best work. Hence "enterprise-grade complexity within two weeks with two humans and ten agents".

▶ 16:25What a good agent task actually looks like

Concrete and practical: scope it so everything the agent needs is findable (which means modularising the codebase), give it a function that scores its own result if you can, and hand it the non-mission-critical work — boring chores, reproduction cases for partial bug reports, or just rubber-ducking.

▶ 17:28"How do you know what's critical? You read the code."

The closing move, and the one that resists being automated away. Write important things by hand, use the agent to help rather than to decide, and accept the friction — because that friction is what builds the model of the system in your head, and where you learn anything new.

From the comments

"The arch linux of coding agent harnesses"

The most quotable reaction, and a fair summary of the trade: you assemble what you want, and nothing arrives that you did not ask for.

The minimal system prompt is the actual selling point

One commenter argues every other harness is too opinionated about how the model should be prompted — if you want yours fiddling with test suites and commit hooks after a one-line change that's your business, but it shouldn't be the default you cannot remove.

This used to just be the Unix philosophy

Fewer features, done properly, composed together — a viewer points out the talk's prescription is the old argument for small sharp tools, arriving back by a different route.

The dissent: not enough technical depth

Worth noting against the near-unanimous praise — at least one viewer wanted more substance on how pi actually works and less on the state of the industry. The talk is a position piece more than an architecture walkthrough.

Auto-generated captions. Click any line or timestamp to seek; the current line highlights as the video plays.

Intro and motivation for building pi 0:00

0:14>> Hey there. I'm Mario. I built pie in a world of slop and this is a tragedy tragedy in three acts.

0:21Just to talk about this real quick.

0:22Bunch of people on the internet gave me money for ad space on my torso and all of that goes to a

0:27charity. So, yeah, thanks guys. So, act one, building pie.

Act 1: Building pi and the frustration with existing agent harnesses 0:29

0:31In the beginning there was cloud code and it was good, right?

0:34We all got basically catnipped by that thing and stopped bunch of stuff before that, but cloud cloud code was the one

0:43thing that kind of clicked with me the most.

0:45And to preface all of this, I love the cloud cloud team.

0:48They're brilliant people, talented, super high velocity.

0:51So, uh they also created the entire game.

0:54Major props to them. So, this is not a roast.

0:56This is just me, an old man, telling you why I stopped using cloud code and built my own thing.

1:01Um in 2025 I started using cloud code in about April, I think, thanks to Peter uh because he told us the

1:09agents are working now. And back then it was simple and predictable and fit my workflow, but the token madness got hold

1:18of them, I think, and the team got bigger and they started uh dog fooding that stuff and built a lot of

1:23features. A lot of features I don't need, which is fine.

1:26I can just ignore them. But with velocity and more features come more bugs and that's bad because I used to work

1:33at construction sites and if my hammer breaks every day, I'm getting really mad.

1:36And if my development tools break every day, I'm also getting mad.

1:40So, there was this. It's just a running gag and here's Tariq telling us that cloud code is now a game engine.

1:45And here's Mitchell from Ghosty telling us, "No, it's not." And eventually they fixed the flicker, but then other stuff broke and

1:51I think they're now in the third iteration of a tool renderer.

1:55Yeah, but that's just a symptom.

Why current context management in tools like Cloud Code and Open Code fails 1:56

1:57The real problem is that my context wasn't my context.

2:00Cloud code is the thing that controls my context and behind my back cloud code does things uh to the context.

2:07So, you have the system prompt which changes on every release, including the tool definitions.

2:11They would remove tools, modify tools.

2:14It's not good. They would insert system reminders in the most inopportuned place in your context telling the model, "Here's some information.

2:22It may or may not be relevant to what you're doing." That actually says it may or may not be relevant what

2:28you're doing. And that kind of confused the model and that kind of broke my workflows.

2:34On top of all that, there's zero observability because that's how the tool is constructed and I like knowing what my agents

2:39are doing. There's zero model choice which is obvious.

2:42It's the native Anthropic harness, so it makes sense for them to want you to use Claude, right?

2:47And there's almost zero extensibility and some of you might have written some hooks for Claude code, but I'm telling you the

2:53number of hooks and the depth of those hooks is very shallow.

2:56Um and every time a hook triggers, what actually happens is a new process gets spawned.

3:01Basically, the command you specified for that hook to be executed and I don't find that specifically efficient.

3:06So, I uh took a step back and looked around for alternatives and I'd like to especially call out Amp and the

3:13Porsche and Lamborghini of coding agent harnesses.

3:16So, if you can afford them, please use them.

3:18They're at the frontier. They're really good and the teams are fantastic.

3:21And there's a bunch of other options and I have history in OSS, so naturally I kind of gravitated towards open code.

3:27And again, brilliant team, super high execution velocity and they don't sell you hype.

3:33They sell you tools that work for the most part.

3:36I started looking under the hood of open code uh with respect to context handling as well because that's the most important

3:41part for me and I found a bunch of things like given some conditions, open code code would just uh prune tool

3:49outputs after a specific minimum amount of tokens.

3:53And that basically lobotomizes the model.

3:56Uh there's also LSP server support, which means every time your model is calling the edit tool, open code goes to the

4:02LSP server that's connected, asks, [snorts] are there any errors?

4:06And if so, injects that as part of the edit tool uh result.

4:10Which is bad, because think about how you are editing code.

4:13You're not writing a line of code, checking the errors, writing the next line, checking the errors.

4:17You don't do that. You finish your work and then you check the errors.

4:21This confuses the model. There's a bunch of other things like storing individual messages of a session in a JSON file.

4:27Each message message is a JSON file on Uh there was this, and this happens to all of us, no no blame

4:33there, but it's not great if by default a server spins up, course headers are set in such a way that any

4:39website you open in your browser can now access your open code server.

4:42That's And uh entirely unrelated to all of this, I started looking into benchmarks for coding agent harnesses and found uh Terminal

The importance of minimal harnesses and the "Terminal" benchmark 4:44

4:50Bench, um which is a pretty good benchmark, all things considered.

4:54And the funny part about it is that it's the most minimal kind of thing you can think of.

4:58All it gives the model is a tool to send keystrokes to to a tmux session and read the output of that

5:04tmux There's no file tools, no sub-agents, none of that stuff.

5:10And it's one of the best performing harnesses in the leaderboard.

5:13Here's the leaderboard from December Irrespective of model family, Terminus scores higher, mostly high even higher than the native harness of that

5:23model. So, what does that tell us?

5:25The form two thesis is we are in the around and find out phase of coding agents, and their current form is

5:31not their final form, second thesis is we need better ways to around.

Introducing pi: A self-modifying, extensible agent core 5:35

5:37And for me, that means self-modifying malleable agents.

5:41Things that the agent itself can modify, and I can modify, depending on my So, I stripped away all the things, built

5:48a minimal core, but made it super and made it so that the agent can modify With some creature comforts, it's not

5:56entirely bare-bones. Uh so, that's Pie.

5:59It's an agent that adapts to your workflow instead of the other way It comes with four packages, uh an AI package,

6:05which is basically just an abstraction across providers and context handoff between providers, an agent core, uh which is just a while

6:12loop and the tool calling, a bespoke tweezer frame work.

6:15I come out of game development, so I built a thing that actually doesn't flicker too much, and the coding agent itself.

6:21Here's Pie's system prompt. >> That's it.

6:25Eventually, the industry created a new standard called skills, which is basically just markdown files.

6:30So, we added that as well, and that needs to go in the system prompt.

6:32So, begrudgingly, we had to add a couple more lines.

6:36And finally, here's the magic that makes Pie able to modify itself.

6:40We ship the documentation, which was handcrafted by me and an agent, um and code examples of extensions.

6:48And [clears throat] all we need to do for the agent to modify itself is tell it, "Here's the documentation.

6:53Here's some code that shows you how to modify yourself by writing extensions." It comes with four tools.

6:58That's all it has, read, write, edit, bash.

7:00Here's the tool definitions. Don't read the the text, just look at the size.

7:05That's it. Here's what happens when you start a new session in one of these tools.

7:11So, the thing is, the models are actually reinforcement trained up to a zoo.

7:15So, they didn't know what a coding agent is, because the coding agent harness is basically what they're being trained when they

7:20are post-trained. You don't need 10,000 tokens to tell them, "You're a coding agent." They know, because they are coding agents Pie

The "YOLO" security philosophy and extensibility through TypeScript 7:27

7:27is also yellow by default, because my security needs are different than yours, and I don't think a little dialogue that pops

7:33up every now every time you call bash, asking you to approve, is a smart security uh mechanism.

7:41So, instead, I give you so much rope that you can build anything that's fit for your specific security There's also stuff

7:49that's not built in. I'm a heathen.

7:53Because this is how I do it.

7:55But if you don't like that, then you just ask Pi to build you sub-agent support on plan mode or MCP support,

8:00whatever you need. Extensibility comes with a bunch of table stakes and then with the extensions itself.

8:07And extensions in Pi are just TypeScript modules.

8:10In the simplest case, a TypeScript file on disk.

8:12You point Pi at that. Here's an extension, load that as part of the harness.

8:16And with that, you get a basically an extension API that lets you hook into everything and define stuff for the harness

8:23to expose to the to the model.

8:25And that includes tools, slash command shortcuts.

8:28You can listen in on any kind of event and react and then save state in the session optionally provided to the

8:37agent as well or stored there for tools that analyze sessions as part of your organizational workflows.

8:43You can do custom compaction, custom providers, and you have full control over the tools.

8:46So, you can modify everything in Pi.

8:49And you can then bundle all of that up and put it on NPM or on GitHub because I think we don't

8:54need to reinvent another bunch of silos called marketplaces.

8:59We already have package managers managers.

9:02And all of that hot reloads.

Examples of pi extensions (chat rooms, NES, Doom) 9:03

9:04So, if you develop an extension for Pi, you do so in the session and you hot reloads changes and see the

9:11the the effects of that immediately, which is very great and that's also game development thing is in game development, you want

9:17high very low iteration uh speeds and that's great.

9:22So, a couple of examples. Cloud or Anthropic ships the slash, by the way, which lets you talk to the agent while

9:28goes on its main quest. I posted this little prompt on Twitter jokingly and somebody built it in 5 minutes with more

9:34features. And they didn't have to fork or clone it just let the agent write the extension based on the prompt.

9:42Here's Nico as one of the most prolific uh extension writers.

9:45I don't know what the is going on here.

9:46It's a chat room for all of his Pi agents and they talk with each other.

9:49I would never use this, but all of this is custom including the UI.

9:53Or you can play NES games.

9:55Or you can play Doom. And there's a bunch of other examples I'm not going to talk about.

10:01So, how do you build a Pi extension?

10:02You don't. You tell Pi to build it for you based on your specifications and then you just iterate with it on

10:07that and hot reload during the session.

10:10Going to skip that example as well.

10:11And if you don't like building things yourself and I hope you do like building things yourself, but if you don't you

10:17can look on NPM or our little search uh interface on top of NPM to find packages for sub agents, MCP, and

10:23so on. So, does it actually work?

10:25Well, here's the terminal bench leaderboard from October before Pi had compaction.

10:29I added that for Peter's claw thingy.

10:32It scored sixth place. Uh but none of this is actually about Pi.

10:37If you want to read I basically want you to retake control of your tools and workflows.

10:41So, build your own. Um and if you want to know more about Pi and Open Claw, go to this talk, please.

Act 2: OSS in the age of "clankers" and how to fight them 10:46

10:46Yeah, and then eventually Peter happened.

10:48He put Pi inside of Open Claw as it's a gent core, which meant my open source project became the target of

10:53a lot of Open Claw instances unbeknownst to their users.

10:57So, this is act two, OSS in the age of clankers.

11:00Clankers are destroying OSS. Here's Till Draw.

11:02They closed down the issue and pull request tracker.

11:04Here's Open Claw's uh trackers. Here's mine.

11:08Half of that is Open Claw instances who post garbage.

11:11So, I started to rage against the Um if you send a pull request, it gets auto closed with a comment that

11:18asks you to please write a nice issue in your human voice no longer than a screen worth of text.

11:23And if I see that, I write looks good to me and your account name gets put in a file in the

11:28repository and the next time you send a pull request, it's let through.

11:31Clankers don't read that comment. They don't go back once they posted a pull request.

11:35So, that's a perfect filter. Uh Mitchell eventually turned that into vouch.

11:39Here's a clanker. Uh I also labeled them.

11:42If you had interactions with open claw, your issues get deprioritized.

11:46I also built tools where I embed uh issues and pull request texts into 3D space, so I see clusters of issues.

11:53Uh I also invented OS certification.

11:54I just close the tracker whenever I want, so I have my life back.

11:58So, does this work? Yes, sort of.

12:02>> Which leads me to act three, slow the Everything's broken.

Act 3: A plea to slow down and stop the "slop" in software development 12:03

12:08And then there's people that say, "Our product's been 100% built by agents." Yes, we know it sucks now.

12:22>> And I'm hearing this from my peers, and this is entirely unhealthy.

12:26Um so, here's how we should not work with agents and why, at least in my opinion.

12:31I wrote this on my blog a while ago, but the basic gist is we're having army of agents in your using

12:35beats on and you don't know that it's basically uninstallable malware, and Entropic built a C compiler.

12:41It kind of works, but actually doesn't, and we're hoping the next generation of molds will fix it.

12:45And here is Kerbal building a browser, and that's also super broken.

12:48Uh but the next generation will fix it.

12:50And SaaS is dead software is often 6 months, and my grandma just built herself a Spotify with her open Come on,

12:57people. So, agents are actually compounding booboos, which is my word for errors, with zero learning and no bottlenecks and uh delayed

13:05pain. The delayed pain is for you.

13:07Here's your code base on a human, on one agent, and 10 agents.

13:12How much of the agent code can you review?

13:14Here's the same code base, but expressed in number of booboos per day.

13:19How much of those booboos do you think you'll find?

13:22Then you say, "Oh, I have a review agent." Let me introduce you to the wonderful world of the ouroboros.

13:28Doesn't work. It catches some issues.

13:30Um the problem is that agents and merchants have learned complexity.

13:33Where did they learn that complexity from?

13:35From the internet. What's on the internet?

13:37All our old garbage code. There are some pearls on the internet, really well-designed systems, but 90% of code on the internet

13:43is our old garbage. And that's what the models learn from.

13:47And every decision of an agent is local, especially if the code base is so big that it doesn't fit into its

13:52context. And if you let it go wild and add abstractions everywhere that are and Um so that leads to a lot

How agents create "enterprise-grade complexity" and why humans are still the bottleneck 13:58

13:59of abstractions and duplication and backwards compatibility.

14:03Who has seen that in the output of their agents?

14:05It's annoying. Or defense in depth.

14:09So yeah, you get enterprise-grade complexity within 2 weeks with just two humans and 10 agents.

14:16And then you say, "But my detailed spec." Yes, sure.

14:20You know what we call a sufficiently detailed spec?

14:23It's a program. So if you leave blanks in your spec, what do you think happens?

14:29How does the model fill in the blanks?

14:31And with what does it fill that in?

14:33It fills it in with the garbage that it learned on the internet from our old code, which is garbage from mediocre.

14:39And then you say, "But humans also." Yes, humans are horrible, failed fallible beings, but they can learn.

14:45And they are bottlenecks. There's only so many booboos they can add to your code base on a daily basis.

14:51And humans feel pain. Which is a very interesting property because humans hate pain.

14:56And once there's too much pain, the human has a bunch of options.

14:59It can quit their job. It can uh blame somebody else and make them fix it.

15:05Or everybody bands together and starts refactoring the out out of the garbage code base, right?

15:11will happily keep into your code base.

15:16And now your agents and their super complex memory systems will not save you.

15:20Agents don't learn the way we Those are my most most beloved people.

15:26I don't even read the code anymore.

15:28Congratulations. Something is broken and your users are screaming.

15:32So, who you going to call?

15:33Not yourself, because you haven't read the code.

15:36So, you're relying on your agents, but they are now also overwhelmed because the code base is so humongous that there's absolutely

15:42zero chance they can get all the context they need to fix the issues.

15:46And long context windows are a hack, as most of you will find those this year as everybody's switching to 1 million

15:52tokens context windows. And agentic search is also failing.

15:56So, the agent patches locally and up globally.

16:00If you see this in your code base, So, you cannot trust your code base anymore and also not your tests because

16:09your agent wrote your tests. So, good So, here's how I think we should work.

Practical advice: How to effectively integrate agents into your workflow 16:12

16:13Um there's a bunch of properties for good agent tasks.

16:16That means scope. If you can scope it in such a way that the agent is guaranteed to find all the things

16:22it needs to find to do a good job, you're done.

16:25That means modularize your code base.

16:27If you can give it a function to evaluate how well it did the job, even better.

16:31Hill climbing, auto Uh anything non-mission critical, let it wipe.

16:36Boring stuff, let it wipe. Reproduction cases for user issues, which are usually only partial in information, perfect.

16:42I don't spend any mornings anymore doing that.

16:44Or if you don't have a human near you, rubber duck.

16:47So, lots of tasks you can use them for and save time.

16:50At the you evaluate. You take what's reasonable, most of it isn't, and then My final slide, more or less.

16:58Slow the >> Think about what you're building and why, and don't just build because your agent can do it now.

17:03That's Uh learn to say no.

17:07This is your most valuable capability at the moment.

17:11Fewer features, but the ones that matter, and then use your agents to polish the out of that.

17:15Enlighten your users, not your uh token maxing desires.

17:21Cap the amount of generated code uh that you need to review.

17:25And non-critical code, sure, five slope ahead.

17:28Critical code, read every See the keynote after me for more info on that.

17:34So, how do you know what's critical?

17:36Any guesses? you read the code.

17:41>> Uh if you do anything important, write it by hand.

17:43You can use a clanker to help you with that, but don't make let it make the decisions for you because we've

17:48learned all the decisions it makes are learned from the internet.

17:52And that friction is the thing that builds the understanding of the system in your head, which is important.

17:58And it's also where you learn new And all of this requires discipline and And all of this still requires humans.

18:07Thank you.