The Four Horsemen of Agentic Coding

55 points by guillego a day ago on lobsters | 21 comments

Even as someone who is more tolerant of agentic coding than many folks here, I thought this was brilliant.

(My personal trend has been to limit the agent to a virtual pairing partner, watch it closely, and write half the code myself. It really is a nice rubber duck, for better or worse. But if you let it work unsupervised, it will get into trouble.)

arialdo | 10 hours ago

I would love to read a blog post about your setup and workflow.

Sure, I should definitely write this up properly at some point!

The basic idea behind my workflow is that writing code isn't the real bottleneck. The real bottleneck is understanding the problem and the code. So anything that generates code faster than I can understand it is a problem. So agent swarms are mostly useless, because they speed up code generation without improving understanding. Asking an agent to look at my code and see if I've overlooked anything, on the other hand, increases my understanding. Especially if I do the fixes myself.

Now, it's also important that not all code requires equal understanding. Just like on human projects, there are parts I don't care too much about, and parts that I watch like a hawk. Knowing which part is which is where experience helps. The core Rust traits in a system are very important. The 9th implementation of some trait may not be.

Similarly, I choose my models carefully. My model needs to be smart enough to debug a broken library, or to help me figure out mathematical invariants for a prop test. But it shouldn't be so smart that I'm tempted to let it do a lot of work unsupervised. Somewhere between Sonnet 4.5 and Opus 4.6 is ideal. Fable is too smart. In practice, this means running open models like Qwen3.8 27B, Qwen3.8 Flash Next, or DeepSeek V4 Flash 0731. These are all smart enough to walk me through the interaction of two monads (say, flatMap and propFlatMap) if I ask. But they're not smart enough to encourage true vibecoding and "not even reading the code." Probably the largest I would consider for my own work is GLM 5.3 Flash, but you have to be really rich to run it locally.

My favorite agent harness is Pi. One session, no background agents, no TODO list, no plan mode, no fancy workflows. Maybe 1,500 tokens in my AGENTS.md, plus 2 or 3 tiny skills that tell the agent how I like to break down problems. (Which I do not share publicly, because I don't want future agents to learn how I break down problems.)

For version control, I use jj. Using it really well isn't any simpler than learning git. But it makes it easy to break up changes into multiple tiny pieces, and then interactively edit any of those pieces. I have some skills telling the agent to "break changes up Linux kernel style," and how to write a good commit message. This allows me to review tiny, standalone pieces right in Zed and edit patch 3 of 4 all that I want.

I still invest time in my editor. I'm trying Zed these days, with Zeta2 as a code completion model. Zeta2 is nice, because it only completes obvious things, and it rarely completes more than 3 lines. If it has any doubt, it moves my cursor to the part I need to fill in. I work in Rust whenever I can, because I like Rust. And I very regularly drop into my editor and rough out traits, refactor things, and think with my fingers.

Overall, this process is good because it maximizes my understanding. And once a project passes 3,000 lines, a deeper understanding helps me move faster. Like, I started programming on the Apple IIe, and I like to imagine I've gotten pretty decent at it. Plus, it's fun. Letting an agent write all the code would be like letting an agent go watch a fun movie for me.

But the north star of this process is that I'm optimizing for understanding, not speed. And from there, everything else is just making deliberate choices. My process has actually gotten more minimalist with time. And as agents have gotten smarter, I've become increasingly alienated from frontier models. I want a model that will punish me swiftly for trying to delegate too much understanding. It's sort of like my kayak: It's quite maneuverable, but it punishes sloppy technique by making me swim.

This is not the only process I use. Sometimes I just turn the agents off entirely, or limit them to research. But at its best, this understanding-focused process is like having a smart pair programmer or a faculty advisor. I can ask it for tips on what I'm overlooking, or I can say, "I don't really feel like converting from the old Rust error handling library to the new one, so please do that for me." But at each step, I focus on maintaining understanding. And if I do lose understanding, I also lose productivity.

ocharles | 2 hours ago

Lovely write up, and really hits home for me. I haven't considered simpler models but might try - I use Opus 5.5 and Fable and often find it way oversteps and drowns me in code.

One thing I would add - it may or may not be in your workflow - is the use of agents to understand ideas. Not only do I use agents to iterate on ideas and explore the design space, I will also have them walk me through ideas. I find agents really nice to play with a concept and break it apart. I do a lot of work in ML and computer vision, and I've found agents super useful to deep dive into concepts. They can generate examples, edge cases, intermediate representations - all sorts of good stuff. This is equally true when reviewing my colleagues work - I can go beyond line my line reading and really stress test things. But it comes back to understanding. The moment the agent is ahead of me I have to pump the breaks. Simply continuing with more code always ends in disaster and ultimately burns time.

I haven't considered simpler models but might try - I use Opus 5.5 and Fable and often find it way oversteps and drowns me in code.

Using simpler models is partly about keeping me honest. Fable actually has pretty good taste about certain domains. Which means that by the time it does mess up, I have 2,000 lines of code to deal with.

For the "not getting drowned in code" part, I'm relying more on the other parts:

  • Using jj, which gives me patch-stack superpowers (at a price, though jjui and stakk help).
  • Asking the LLM for "Linux style patches" (one logical change per patch). This forces the LLM to think in small chunks.
  • Providing a custom skill about how to write a clear commit message.

Taken together, this can really simplify deep review.

One thing I would add - it may or may not be in your workflow - is the use of agents to understand ideas.

Yes. Agents are excellent at this. I often say, "So here's how I'd like to handle this. Can you see a reason why this won't work?" Then the agent points out 5 problems. Three of them are silly, but one of them means I need to to do a major rethink.

I also have a specialized skill which tells the model that the human is an experienced software designer with the magical ability to change the requirements. So if the model gets stuck on a messy design challenge, it gets maximum credit for flagging that to me. This actually gets many overenthusiastic models to stop more often and to explain why my plan was far too simple. That's one of the points where I tend to start refactoring and typing.

lthms | 10 hours ago

Personally I have written a small util sidekick that allows me to write comments for a LLM and have it write back directly in the editor buffers.

My mental health has improved so much since I did that 😅

I need to clean up and write a blogpost about it

kaspar | 10 hours ago

Could you expand on your set up and workflow for that? I have been trying out an "output style" plugin for the pi harness and setting it to "learning" (mimicking claude-code functionality of the same name) but that ends up making me implement way less than half the code.

antoinewdg | 16 hours ago

[The agent] will never report you to HR

Somehow, I doubt that.

k749gtnc9l3w | 11 hours ago

If you use a hosted model, it will hallucinate a reason to go straight to police, and if you sandbox a local one at all, maybe it won't succeed in reaching HR?

antoinewdg | 3 hours ago

To be honest I did not have the local models in mind ! My reasoning was: I think my employer has access to my claude history (not actually sure, but I would be very surprised otherwise). From then the agent does not even need to report me to.

k749gtnc9l3w | 2 hours ago

Technically they also have access to the entire team chat history, but it doesn't mean they practically react to things without being asked to (when they do, it is often a part of the company being doomed already…)

antoinewdg | 2 hours ago

Agreed it's pretty much the same as chat history. And I guess I'm a bit less optimistic about companies' practices than you are. I would not recommend using a LLM paid by one's employer to look up union rights for example. I would also not be surprised to learn that a lot of companies are monitoring this specific word on regular chat history.

k749gtnc9l3w | an hour ago

I mostly do not expect specifically HR to be the ones doing union-busting crimes.

(Of course, personally I work at a European university where large enough unions literally get small offices on campus, so here union-related conflicts of interest play out differently)

gasche | 7 hours ago

I read that the University of Sidney is moving to a "dual lane" approach where some assignments are meant to be done using AI tools, and others are done is strictly controlled conditions where AI usage is forbidden -- this would be a way for students to both practice for themselves, yet also become fluent in AI tooling usage.

Maybe a similar approach could work in software companies: mandate that a certain share of tasks be done without AI tools. For example the company could have a control where company-funded AI usage is available from Monday to Wednesday, and shut off Thursday and Friday. Would such a hybrid approach avoid some of the downsides? Of course this relies on companies being willing to accept a short-term productivity hit (assuming that AI-assisted days are indeed noticeably more productive), for the benefit of long-term skills and work dynamics. I don't expect most companies to be able to make this choice...

k749gtnc9l3w | 2 hours ago

Tuesday and Thursday are «autopilot» days, you can queue background agentic work for these days in advance with specified maximum token budgets, but you are not allowed to look at the results or course-correct during the auto-pilot day. This is how the company stays sharp and ready for long-horizon agentic future!

Any benefits from people actually working on the code on Tuesdays and Thursdays are purely accidental.

peter-leonov | 3 hours ago

I mostly disagree.

Risky take:

  1. Slop: heavily depends on established coding standards and models
  2. Alienation: the memory of a code decays faster true, but the product vision is still king
  3. Deskilling: AI massively levels the skills at mid-senior level for virtually no effort, prototyping new ideas is cheaper than ever, code reviews are finally politics free
  4. Team Fallout: the opposite, actually; now we have way more time to talk architecture and product future; coding doesn't distract too much anymore

Like, really, hold it by the handle already, it's a tool after all 😅

zem | a day ago

heh, I'm probably the guy on the team who annoys people because I insist on checking md files in, but ironically I'm also one of the few people in the team who doesn't believe in prompt engineering at all.

what I want to check in are the bots' planning notes - I will work with the LLM to plan some feature out, iterate on it a bit, and then tell it to refine the plan into a series of self contained steps and then implement those as a stack of commits. I'll admit, I barely read the md files, what I do is read every commit thoroughly, rework it, have claude or whatever keep refining individual commits and restacking, until everything looks good and I send the code out for review, at which point it goes through more refining, sometimes extensive.

through all that, my theory is that the plan md files help act as the bot's working memory, and having them checked in and modified alongside the code keeps it on track.

as an aside I'm super happy to be working in a team where LLM generated code is held to the same "read every line and make sure commits are good" standard that we used to hold human code to.

carlana | 19 hours ago

Lately I’ve had better luck with having the bot write an MD plan that is detailed enough that a bot or human can implement it and then implementing it myself because the final result needs less polish to be accepted when I do it.

gbalduzzi | 8 hours ago

The real problem of this horsemen is that there is no incentive to fix them by the big players, because it is not something companies are interested in

agnishom | 7 hours ago

Not the author's main point, but I still want to nitpick on this.

For plain prose, humanity seems to be converging on the opinion that AI prose is bland at best and an insult at worst.

Is it? It could be, but I have two concerns before I can accept this conclusion.

  1. Usually the problem is that the human using the LLM output has poor understanding of the sociolinguistic layer.

For example, imagine my friend is mad at me, and I ask the clanker to write an apology for him, and paste it and send it to him. Of course, the prose is going to be off. Presumably, I didn't give the LLM enough context to figure out what tone or phrasing to use in the apology message. The LLM will use a cookie cutter response, but that doesn't mean it was bad prose.

  1. Em dashes, parallel constructions, rule of threes, etc are not bad prose. They are just the default kind of prose LLMs output unless you prompt them to do so otherwise.

sunshowers | 2 hours ago

Em dashes, parallel constructions, rule of threes, etc are not bad prose.

The first by itself, I agree — I've seen plenty of effective prose that uses em dashes well (LLMs are not great at using them, though, seemingly because they don't really have the notion of catching a breath). The latter two are powerful and meant to be used sparingly; anything more and they certainly become bad prose.