I often find it reductive when people say "just tell Pi to build you an extension". Having used it as my one and only harness for a few months now, it's easy to get an extension, but hard to get a good one, that actually works well and helps.
My advice: focus on getting work done and slowly adapt Pi with small augmentations as you go. You can start getting work done on vanilla setup. When the right idea comes along, try it. Be ready to refine it, and most importantly, rollback the addition. I've rolled back a bunch.
Many "batteries" that are "included" come from speculative and half-baked ideas, from people who were excited about something at some point in their journey. In practice, those ideas may not bring the desired results, and their creator may've moved on already. So it's better to either learn very well established tools, or mold your own slowly.
For example, many automatic memory systems are not helpful. I built a small extension that asked me whether it should remember something (and write it down to a properly scoped SKILL or AGENTS file). Turned out I accepted less than 5% of suggestions. Most were useless one-offs that would pollute the context. Can't imagine how much crap would accumulate if I wasn't in the loop.
I have the same opinion as your first paragraph, but I don't want to spend weeks or months vibe-coding basic features which come built into almost every other agent.
Yeah maybe Claude/OpenCode/KiloCode/Hermes/whatever are not as minimal as Pi but they also work right now.
We probably differ a lot in what we consider "basic". Are subagents basic? I found them only useful in very few situations. Is LSP support basic? There are mixed results on whether it helps or hurts. Are multiple choice asking tools basic? I found that they add extra unnecessary ceremony, eat extra context, and I almost never answer with one of the choices exactly.
And if you try vanilla Pi, you will also find out that it works right now.
> Are subagents basic? I found them only useful in very few situations.
I've found them to be extraordinarily helpful, because they allow me to much more carefully control context and reduce token spend by using a smart model for the parent agent and cheap models for the subagents. Do you just have a big token budget?
I started out using OpenCode with subagents. Then, after switching to Pi, went completely subagentless, usually on the current frontier GPT model.
I didn't notice any significant change in context usage, and tasks were completed faster. That surprised me, I'm still not sure (not an expert on this), but maybe the handoff boundary was the problem. When the main model gives an isolated task to the subagent, the latter goes wild producing a comprehensive report, trying to satisfy every possibility. Without the handoff, the main model does the job much more precisely and conservatively, checks only specific/narrow things, and stops sooner.
Recently I decided to reintroduce 2 subagents to see how it goes. First was to have a cheaper model drive my real Safari browser instead of using agent-browser and the like. Second, to see if having a cheaper model navigate/search my file system helps in any way.
I think there's some benefit to having a cheap model drive Safari, because there's so much unavoidable garbage produced in that interaction. The filesystem one I don't think I see any benefit, just a lot of unnecessary work that (albeit cheap) wastes more time.
Of course I'm eyeballing this, not benchmarking formally, but I see so many people just onboard these mindlessly. Are you sure that you saw a real improvement in the produced outcomes/timing, or was it based on seeing subagents do a lot of stuff and assuming that the main model would've been doing the same at higher cost?
I admit that subagents may have great benefits, but I wouldn't treat it as just out-of-the-box basic feature that always improves your outcomes.
I found the new deepseek flash is so good I rarely have to pull out gpt sol anymore. if anything for coding tasks, gpt sol ends to go out of scope and I have to reign it in. Trying to keep my tasks managably small for human review is a endless battle with the frontier models.
Subagents are good when the harness (And agent?) understand how smart they are.
My root level CLAUDE.md has pretty much just "use a lower tier agent when relevant".
Then I daily-drive Opus, it automatically offloads simpler stuff to Sonnet or even Haiku based its own reasoning because it "knows" their capabilities.
It's so much more cost/token efficient to do it like this. Opus writes the exact implementation plan for Sonnet and then waits for it to complete. After that it checks the work and fixes any issues itself.
In Codex, for example, this doesn't work because the whole system doesn't know about agent tiers and barely can use subagents. So I'm just running Sol all the time.
You don't actually need those features. Initially I began using Pi thinking I would customize the hell out of it. I've installed one extension for guardrails and that's about it. I've been able to do everything I did before just with vanilla pi.
What extension do you use ? With all the supply chain attacks I am warry of adding extensions, so I wonder if there's like a go-to one that everyone uses
I think for indie hackers and people that build their own stack is great, but real scenario and people with money Enterprise likes the idea of batteries included.
I found this one a little bit better and they do support Extensions like Pi. But comes with all features like codex, claude code and it's open-source.
My advice: focus on getting work done and slowly adapt Pi with small augmentations as you go. You can start getting work done on vanilla setup. When the right idea comes along, try it. Be ready to refine it, and most importantly, rollback the addition. I've rolled back a bunch.
Many "batteries" that are "included" come from speculative and half-baked ideas, from people who were excited about something at some point in their journey. In practice, those ideas may not bring the desired results, and their creator may've moved on already. So it's better to either learn very well established tools, or mold your own slowly.
For example, many automatic memory systems are not helpful. I built a small extension that asked me whether it should remember something (and write it down to a properly scoped SKILL or AGENTS file). Turned out I accepted less than 5% of suggestions. Most were useless one-offs that would pollute the context. Can't imagine how much crap would accumulate if I wasn't in the loop.