Why Are Coding Agents So Dumb?

24 points by mtlynch 7 hours ago on lobsters | 10 comments

kingmob | 6 hours ago

I get that this is more of a rant, and I agree with a lot of it, but some of these are Claude-specific or have plausible reasons.

For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.

Likewise, when it comes to model/task suitability, self-reflection on model suitability is non-existent in general, yes, but there are passable alternatives.

In Claude, I instruct it to use models appropriate for the subtask, and it kinda works, but if you use Oh-my-pi, it allows you to specify models for different roles (slow, smol, review, basic tasks, etc), and will use them consistently.

Oh-my-pi also directly supports advisor agents, where a smarter model watches another subagent, and interjects as needed, which is a hybrid solution.

[OP] mtlynch | 6 hours ago

Thanks for reading!

For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.

I think this is a plausible explanation, but I still find it unintuitive. Wouldn't you expect the agent eat up more tokens to do 10 tasks in the same session/context than to have one supervisor delegate to 10 tightly-scoped subagents?

But even token efficiency aside, I imagine there are lots of users like me who don't run up against quota limits often and would happily trade tokens for faster execution (in wall time).

janxdevil | 7 hours ago

"If I had a human employee tell me they sat idle their whole shift because they wanted my input on some superficial detail, I’d quickly fire them."

I think I found the root of the issue here.

dlisboa | 5 hours ago

Your wish list for coding agents is essentially 70 years worth of computer science, design and product discipline, coupled with super human intelligence and human level restraint. If we had all of that we'd have solved pretty much every problem in CS. And you want it to be open source.

We're closer to having humanoid robots that can reproduce themselves than that.

[OP] mtlynch | 5 hours ago

Thanks for reading!

Which wish list items do you mean? The models are capable of these things but the agents don't take advantage.

cbrake | 4 hours ago

This is a great rant :-)

One thing I'll often do is ask my agent to update the project docs first (or I'll make initial edits). I want to see what this will look like from a user's perspective. Once I'm happy with that, I'll move to plan, and then code.

I agree, I rarely read every last plan detail, but they are probably still a useful exercise for the agent to go through the planning step for itself (like a human).

benjajaja | 5 hours ago

It's both model and agent. The model has a context limit, so until someone figures out a way to efficiently flush and make it long lived, you have to treat every LLM interaction as if with a baby freshly spawned but with a vast compilation of knowledge but that knows nothing about the specific environment except AGENTS.md or whatever shitty crutches we have.

kornel | 3 hours ago

I suspect frontier labs see harnesses as just an LLM-to-bash adapter. You won't see smart harness features from them, because everything a harness could help with, they hope to RL-train into their next model.

a lot of this cognitive dissonance goes away if you just substitute "tool" or "code generator" for "agent" throughout. anecdotally you will also have a much better time using them, in terms of satisfaction vs frustration.

codekobold | an hour ago

I think the part about OS level sandboxing is factually wrong? Claude supports that via bubblewrap: https://code.claude.com/docs/en/sandboxing