I get that this is more of a rant, and I agree with a lot of it, but some of these are Claude-specific or have plausible reasons.
For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.
Likewise, when it comes to model/task suitability, self-reflection on model suitability is non-existent in general, yes, but there are passable alternatives.
In Claude, I instruct it to use models appropriate for the subtask, and it kinda works, but if you use Oh-my-pi, it allows you to specify models for different roles (slow, smol, review, basic tasks, etc), and will use them consistently.
Oh-my-pi also directly supports advisor agents, where a smarter model watches another subagent, and interjects as needed, which is a hybrid solution.
For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.
I think this is a plausible explanation, but I still find it unintuitive. Wouldn't you expect the agent eat up more tokens to do 10 tasks in the same session/context than to have one supervisor delegate to 10 tightly-scoped subagents?
But even token efficiency aside, I imagine there are lots of users like me who don't run up against quota limits often and would happily trade tokens for faster execution (in wall time).
Your wish list for coding agents is essentially 70 years worth of computer science, design and product discipline, coupled with super human intelligence and human level restraint. If we had all of that we'd have solved pretty much every problem in CS. And you want it to be open source.
We're closer to having humanoid robots that can reproduce themselves than that.
One thing I'll often do is ask my agent to update the project docs first (or I'll make initial edits). I want to see what this will look like from a user's perspective. Once I'm happy with that, I'll move to plan, and then code.
I agree, I rarely read every last plan detail, but they are probably still a useful exercise for the agent to go through the planning step for itself (like a human).
It's both model and agent. The model has a context limit, so until someone figures out a way to efficiently flush and make it long lived, you have to treat every LLM interaction as if with a baby freshly spawned but with a vast compilation of knowledge but that knows nothing about the specific environment except AGENTS.md or whatever shitty crutches we have.
I suspect frontier labs see harnesses as just an LLM-to-bash adapter. You won't see smart harness features from them, because everything a harness could help with, they hope to RL-train into their next model.
a lot of this cognitive dissonance goes away if you just substitute "tool" or "code generator" for "agent" throughout. anecdotally you will also have a much better time using them, in terms of satisfaction vs frustration.
kingmob | 6 hours ago
I get that this is more of a rant, and I agree with a lot of it, but some of these are Claude-specific or have plausible reasons.
For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.
Likewise, when it comes to model/task suitability, self-reflection on model suitability is non-existent in general, yes, but there are passable alternatives.
In Claude, I instruct it to use models appropriate for the subtask, and it kinda works, but if you use Oh-my-pi, it allows you to specify models for different roles (slow, smol, review, basic tasks, etc), and will use them consistently.
Oh-my-pi also directly supports advisor agents, where a smarter model watches another subagent, and interjects as needed, which is a hybrid solution.
[OP] mtlynch | 6 hours ago
Thanks for reading!
I think this is a plausible explanation, but I still find it unintuitive. Wouldn't you expect the agent eat up more tokens to do 10 tasks in the same session/context than to have one supervisor delegate to 10 tightly-scoped subagents?
But even token efficiency aside, I imagine there are lots of users like me who don't run up against quota limits often and would happily trade tokens for faster execution (in wall time).
janxdevil | 7 hours ago
I think I found the root of the issue here.
dlisboa | 5 hours ago
Your wish list for coding agents is essentially 70 years worth of computer science, design and product discipline, coupled with super human intelligence and human level restraint. If we had all of that we'd have solved pretty much every problem in CS. And you want it to be open source.
We're closer to having humanoid robots that can reproduce themselves than that.
[OP] mtlynch | 5 hours ago
Thanks for reading!
Which wish list items do you mean? The models are capable of these things but the agents don't take advantage.
cbrake | 4 hours ago
This is a great rant :-)
One thing I'll often do is ask my agent to update the project docs first (or I'll make initial edits). I want to see what this will look like from a user's perspective. Once I'm happy with that, I'll move to plan, and then code.
I agree, I rarely read every last plan detail, but they are probably still a useful exercise for the agent to go through the planning step for itself (like a human).
benjajaja | 5 hours ago
It's both model and agent. The model has a context limit, so until someone figures out a way to efficiently flush and make it long lived, you have to treat every LLM interaction as if with a baby freshly spawned but with a vast compilation of knowledge but that knows nothing about the specific environment except AGENTS.md or whatever shitty crutches we have.
kornel | 3 hours ago
I suspect frontier labs see harnesses as just an LLM-to-bash adapter. You won't see smart harness features from them, because everything a harness could help with, they hope to RL-train into their next model.
zem | an hour ago
a lot of this cognitive dissonance goes away if you just substitute "tool" or "code generator" for "agent" throughout. anecdotally you will also have a much better time using them, in terms of satisfaction vs frustration.
codekobold | an hour ago
I think the part about OS level sandboxing is factually wrong? Claude supports that via bubblewrap: https://code.claude.com/docs/en/sandboxing