I’m running little offices of 5-10 persistent agents with defined roles and it’s crazy productive and code quality has never been better. I’ve been working on management of physical systems as well, and so far it’s been pretty solid.
Is there any great documentation you have followed to build your workflow or has it been experimentation? I've been experimenting a lot and gotten some cool things done but every time I read up on what everyone else is doing I realize I don't have much imagination and so I miss out.
Do you basically know that quality of code hasn't declined based on the number of incidents that are occurring? How do you even know what is happening?
I'd instead say that for both it's really at 100%, and not any number in between.
Problem is that we can't define precisely what is good enough or when software is "finished": I mean, we are trying to do that with human languages, so it should not be a surprise.
Yes, LLMs are now similarly aware of the average context a human would be aware of, but for anything specific to the situation a human has better chances of resolving the ambiguity.
It's great. I remember bemoaning how little things had really changed between 2006 and 2016. Given that in theory software is infinitely malleable I found it depressing how it felt like we were just doing the same things. It's exciting we've finally unblocked progress and are finding amazing new ways to work, although I never would have foreseen this being the way! 2026 sure feels nothing like 2016.
Well, we've unlocked something. How much of it is progress remains to be seen. Definitely interesting times.
A lot of the AI hype these days reminds me of the infamous pets.com; but, of course, the general ideas of the dot-com boom did come to pass. (And if all ideas people had tried back then had turned out to be sound and successful, that would have been a surefire sign that they weren't trying out enough crazy stuff.)
I would argue really it's less than 9 months. At the start of the year then models/AI really weren't capable of shipping production code etc without hand holding...
Todo's left in code, mock implementations instead of fully working, people doing Ralph loops etc
Take a stepback and understand why standup was invented in the first place. The reason was: information flow bottleneck - with agents information flow should not be a bottleneck because at any given time the agent can grab as much context as they need, including talking to other agents (althought that nit necessary imo since agents should work visibly through issues and PRs)
It could, that's why there are teams that don't have a standup and work very well (I've been with both and I much prefer the ones that dont have a standup OR they meet on a standup to chit-chat and keep up the connections rather than discuss the actual work - simply because they have autonomy and can do it async)
However, a "standup" was invented to solve a particular problem: people, in isolation, do not always understand that they are stuck, and talking it through lets fellow humans identify their challenge.
The stand up part is mostly to encourage keeping it concise and short, so you only focus on what is actionable between all participants (waiting on a review, I am looking at this for the 3rd day and I think I am ok, anyone wants to brainstorm this...).
With agents, they get similarly stuck (completely or in a loop), and with each keeping a different context at any given moment, they might still be able to help each other.
Obviously, their memories are not like humans' (there is no "I touched that two months ago, let's discuss after"), but there might be a different sweet spot where having them sync regularly (every 1M tokens instead of daily?) might be useful as they are also parallelizing work.
One useful part of standup is to elicit connections where one wasn't expected. For example, someone might say "I've been working on X" and another person might say "Oh, I worked on X a couple of days ago and you should know ..."
It is an opportunity for things people didn't realize they needed to communicate to get stated publicly in a way that can create connections neither side would have initiated.
Been using the Pullboard workflow to build software that MUST be correct, and recently open sourced it. There was another HN link today involving agents role-playing and I can't take that seriously, no offense. (Love fun projects like this, seriously, it's cool.)
From my perspective, agents are great, they code better than most of us. But like all of us, they also make many mistakes or can't see how their changes impact the codebase. But that is a test-driven development mitigation along with some sort of scrutinizing process, "prove it". I'm actively working on this, and by actively working on it, I mean building serious apps that hit correctness walls, figuring out why Pullboard didn't catch it. To me, this is all about the process.
Tests are inadequate to catch bugs. Tests are like putting thin net in some little sections of a window, but property based testing is like something more like a decent net but not a solid barrier. And agents will lie, misunderstand or hallucinate things when asked to "prove it." And none of this fights off the bloat, performance issues, and decrease in maintainability. People keep thinking the process will save them from actually having to think through and put things together correctly. It's too bad that people are missing the satisfaction from personally building solid, correct things. And users are getting increasingly worse software because of this trend.
Entirely agree, also things can be logically correct and well tested but not the behaviour the user intended.
Even if that's down to a bad initial prompt, or lack of data for the agent to notice the edge case I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.
You are correct that it can all be correct in one sense, and wrong in the product (“what the user wants”) sense. Of course you need to monitor the output, that’s a given. You need to question the entire system. That is the frontier for engineers.
> I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.
I also totally agree with your response. But my point was that you can't engineer a better agentic solution. I'm arguing to never use agentic development because it is fundamentally flawed and inferior.
For a long time I've avoided anthropomorphizing agents like this, but it really is hard to ignore the effectiveness of using them with human processes and systems. Are the processes (standup, sprint, etc.) effective because they are using our languages? Or are the processes just universally good?
I hope when the bubble pops people will reference this and go “yeah it got so stupid that people started making teams of cartoons you could work with” - our industry is a complete clown car right now
I don’t think it will disappear as the tools can be useful (finding bugs, searching large stores of info etc), but obviously bogus things like this meant to impress the gullible will disappear yes.
Sadly the internet has already been ruined by bots pretending to be human, and there’s no way back from that.
> Made by a bored dev and his robots.
This dev was replaced by his own agents.
Maybe if you hadn't outsourced your job to Claude, you wouldn't be bored. And if you're forced to do so by manglement, why would you also outsource this toy project?
The dependency tree and "who is blocked" view feels more useful than the call itself. When a few agents run in parallel, the failure mode I keep hitting is not missing status updates, it is silent waits on a shared file or API that nobody surfaces. Curious whether the standup format actually changed how you unblocked them, or whether the value was mostly having that tree visible.
Would love for my agent to chime in 'nothing new on my end'
on a serious note, the classic enterprise stack of Gemini and Jira really has put me in a position of questioning the standup. Kept mostly for vibes and routines I would say.
holdupagain | 17 hours ago
K0balt | 16 hours ago
I’m running little offices of 5-10 persistent agents with defined roles and it’s crazy productive and code quality has never been better. I’ve been working on management of physical systems as well, and so far it’s been pretty solid.
technocratius | 16 hours ago
saturn8601 | 16 hours ago
seer | 16 hours ago
my-next-account | 16 hours ago
duttish | 16 hours ago
Claude has let me build really cool things but so far I'm not letting it run on it's own.
Olscore | 16 hours ago
eru | 15 hours ago
necovek | 15 hours ago
Problem is that we can't define precisely what is good enough or when software is "finished": I mean, we are trying to do that with human languages, so it should not be a surprise.
Yes, LLMs are now similarly aware of the average context a human would be aware of, but for anything specific to the situation a human has better chances of resolving the ambiguity.
GolfPopper | 15 hours ago
coubri | 16 hours ago
Just imagine seeing this in 2021
pmg101 | 15 hours ago
eru | 15 hours ago
A lot of the AI hype these days reminds me of the infamous pets.com; but, of course, the general ideas of the dot-com boom did come to pass. (And if all ideas people had tried back then had turned out to be sound and successful, that would have been a surefire sign that they weren't trying out enough crazy stuff.)
CurleighBraces | 14 hours ago
Todo's left in code, mock implementations instead of fully working, people doing Ralph loops etc
It's an entirely different situation now
lelanthran | 16 hours ago
You don't need the self-policing enforcement of micromanagement that humans use standups for, after all.
I mean, why not just write the harness so that a management agent constantly reads those sub-agent's files and course-correct the subs?
gwt4life | 16 hours ago
Gigachad | 16 hours ago
piterrro | 16 hours ago
spacebanana7 | 15 hours ago
newsicanuse | 15 hours ago
darkwater | 15 hours ago
eru | 15 hours ago
ranguna | 12 hours ago
piterrro | 11 hours ago
faithlv | 10 hours ago
necovek | 15 hours ago
However, a "standup" was invented to solve a particular problem: people, in isolation, do not always understand that they are stuck, and talking it through lets fellow humans identify their challenge.
The stand up part is mostly to encourage keeping it concise and short, so you only focus on what is actionable between all participants (waiting on a review, I am looking at this for the 3rd day and I think I am ok, anyone wants to brainstorm this...).
With agents, they get similarly stuck (completely or in a loop), and with each keeping a different context at any given moment, they might still be able to help each other.
Obviously, their memories are not like humans' (there is no "I touched that two months ago, let's discuss after"), but there might be a different sweet spot where having them sync regularly (every 1M tokens instead of daily?) might be useful as they are also parallelizing work.
stillpointlab | 15 hours ago
It is an opportunity for things people didn't realize they needed to communicate to get stated publicly in a way that can create connections neither side would have initiated.
Olscore | 16 hours ago
https://github.com/pullboard-dev/pullboard
Been using the Pullboard workflow to build software that MUST be correct, and recently open sourced it. There was another HN link today involving agents role-playing and I can't take that seriously, no offense. (Love fun projects like this, seriously, it's cool.)
From my perspective, agents are great, they code better than most of us. But like all of us, they also make many mistakes or can't see how their changes impact the codebase. But that is a test-driven development mitigation along with some sort of scrutinizing process, "prove it". I'm actively working on this, and by actively working on it, I mean building serious apps that hit correctness walls, figuring out why Pullboard didn't catch it. To me, this is all about the process.
adamddev1 | 14 hours ago
CurleighBraces | 14 hours ago
Even if that's down to a bad initial prompt, or lack of data for the agent to notice the edge case I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.
Olscore | 14 hours ago
“How is the agent lying to me?” Etc.
adamddev1 | 14 hours ago
I also totally agree with your response. But my point was that you can't engineer a better agentic solution. I'm arguing to never use agentic development because it is fundamentally flawed and inferior.
cyberrock | 16 hours ago
woggy | 16 hours ago
bazza451 | 15 hours ago
hlynurd | 15 hours ago
blamestross | 15 hours ago
grey-area | 15 hours ago
Sadly the internet has already been ruined by bots pretending to be human, and there’s no way back from that.
bazza451 | 14 hours ago
Bring this into a professional workplace and watch yourself get laughed out of the building
lelanthran | 10 hours ago
The internet didn't disappear after dot-bomb, but the cue-cat did.
teaearlgraycold | 15 hours ago
Kwpolska | 14 hours ago
Maybe if you hadn't outsourced your job to Claude, you wouldn't be bored. And if you're forced to do so by manglement, why would you also outsource this toy project?
N_Lens | 14 hours ago
apt-apt-apt-apt | 14 hours ago
danielsorok | 13 hours ago
donkey_brains | 8 hours ago
Oscalemor | 3 hours ago
on a serious note, the classic enterprise stack of Gemini and Jira really has put me in a position of questioning the standup. Kept mostly for vibes and routines I would say.