I just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
curious what moved it the most for you? I have been measuring claude PR reviews in CI, biggest difference was giving it the diff with the surrounding code upfront instead of letting it go read the repo by itself, roughly half the cost
We saw similar (slightly smaller effect, ~30% less instead of the 50% less you saw) improvement from trimming the number of turns. I'm not at all confident we've found the best way to do this here. We do need the thinking tokens to be spent in order to get the less obvious findings, but we couldn't get them from just increasing thinking effort. We went with forcing the model to use a tool that parallelizes the calls, and has an output floor so there's no turns with very little content read. Again, not at all confident this is the best way to do this, just what we've thrown together that works.
The actual biggest improvement is reducing the number of review rounds a PR has to go through. We were having too many rounds of drip-feeding new findings, because the AI reviewer is a subjective grader. Even if it found the same finding before, the next time it runs it might score it wildly differently. And it tends to want to score its findings across the whole grading curve, because it thinks that's more correct-looking. Our answer is to use previous rounds scores as anchors for the next round. Since we can't just use the same context and still have good performance, we needed to condense the score report into a small json we send forward.
I've been doing this for a long time now. I have the agent report its "user story" from how it went about building said thing and the tweaks/hacks it had to make in the process. This is the "output" of the test I read and then use to formulate the next set of changes.
I'm just actually using the thing I'm building while building it, so I feel the sharp edges and the project quickly evolves based on actual need. Nothing new really.
We use NX in our monorepo, and it is great at determining which testing/linting tasks need to be run based on which libraries in the repo were "affected". We have a bunch of e2e tests, and I've been having a lot of success getting Claude with Opus to run only the relevant e2e tests when appropriate during development.
Seems like one more arrow in the toolkit. But the best testing looks at the sources and explores the cracks between the strata with edge cases, looks at limits, and where one method changes to another. (And hats off to Murphy, for waiting until after you ship...)
> In effect, instead of building the core while trying to anticipate what might be needed at the other layers, you just simulate the other layers by actually building them.
Eh. I think you're just going to end up with slop, or sloppy recommendations?
My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.
Unless I'm missing something, this doesn't help with preventing regressions. In the end, as the author already puts it, it's an integration test in the end, why not just write the integration tests directly?
The word "testing" is a bit misleading. It's not about test coverage, it's about experimenting to find architecture decisions that are suitable for further development. So, not generating automated tests, but instead building throw-away features on top of the new thing and seeing if they turn out okay. (Though the post remains a bit vague on how to judge "okay".) Maybe it should be called something like "ephemeral build-out" or "future usage trial" or so.
I've been using this pattern quite a bit recently for API design, and I really like it.
The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?
With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
That's a thing with AI-generated code. The lying machine happily reports that "boss, everything is clean, tests pass, code gate green", but once we need to build on top it's always "preexisting flaky tests, not related to this section" and trying to do commit --no-veriry before it's slapped on it's little robot hands twice.
There is a solution for green test suite - mutation tests. Testing your tests gives you solid ground to treat green tests as reliable. It takes time but also brings quality.
In terms of an agent reporting green tests never executed, a proper step in deployment workflow may be a good gate.
i'm testing to see if doing new feature roll-outs can help me eval whether a refactor was good or not - very similar to approach here. my intuition is that good refactors should reduce tokens used by downstream coding agents. haven't seen a big difference yet but it might just be that i need to do more rollouts (lots of variance in tokens used per run). in my experience you have to intentionally 'mow the lawn' or things get out of hand so i'm always looking for slop signals.
I had good results with a similar technique this summer.
I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
Not unthinkable at all. Been doing this for several years as an opt-in on pr tags usually but now with agents writing code I have it on by default. I've got several skills files about accessing with a read only account and gitops done through the PR. It's an excellent dev environment for the agents.
This is an interesting idea. It's effectively what a good software engineer already does in their head, except it's doing it with a real compiler.
Richard Gabriel wrote something that has really stuck with me:
> Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.
This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.
A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!
Another trick I have found useful is to do integration testing with coverage enabled. One agent creates tasks for subagents that run the application with coverage enabled, it merges the reports, and based on that it comes up with new tasks for subagents, repeat until coverage no longer moves.
instead of? IMO You ARE writing the proper test. If your API doesn't have state, your throwaway UI will have imparted the same test coverage as a repeatedly running unit test where it matters. If your API has state, well ...
IMO, It's the same idea with REPL-driven development in a functional language, why write unit tests if an interface is impossible to behave differently over time without changes?
Typically someone then argues, well, what if you change the implementation. To that person I will point out that the change will be developed in a REPL just like the original version, and thus be tested when it hits. Oh well...
To some, QA means a lot of compute and green outputs, to others, QA means spending quality hammock time before hitting the REPL :)
> A library with hidden state, surprising defaults, or incomplete docs produces a pile of patches and failures.
Sigh.
Write a C or C++ API that works with pointers. Make it handle null pointers and errors elegantly so that the API user can safely chain calls and only check the final result. Claude decides it's better to be safe than sorry and peppers its code with intermediate nullptr and return value checks anyway.
This is an application of the rule of three: you need at least three clients to write a good library, protocol, or other underlying abstraction. Except it proposes using a large language model to write the three clients and then throwing them away.
Unsteered coding agents like to write unit tests on the smallest possible functions (based on their training data) - but that was important when humans were writing code. AIs rarely make mistakes on simple functions, what is more important is testing the most outer layers, the business requirements rather than implementation.
One idea I've had for a while which feels similar:
- We haven't really solved software engineering yet, it's hard to tell upfront whether a design adapts well to future needs. The community has a collection of heuristics which are sometimes at odds with each other.
- Therefore: write all designs, each simulated against a random tree of future needs.
- Pick whichever minimises the average future diff.
In other words, monte carlo simulation for software design. This scheme obviously assumes the price of code production goes even closer to zero.
I can tell there aren't a lot of other QA people here because this is just usability/user testing [0] with a throwaway prototype add-on (which is valuable but also kind of an obvious thing in the age of AI)
that gut check of 'does this feel good to use' is funny to hear as QA because most of the feedback I give that boils down to 'this sucks to use' nearly universally gets met with 'but the AC! SLO! SLA! it's designed this way on purpose! it's a feature not a bug' etc.
I mean, I'm glad SWEs are finally discovering that using the product of your own work is actually a net good but it's a little like seeing a bullet coffee evangelist author a blog post about how he's found the real secret to preventing congestive heart failure and it's to not drink butter regularly
CBLT | 21 hours ago
raviinits | 8 hours ago
CBLT | 4 hours ago
The actual biggest improvement is reducing the number of review rounds a PR has to go through. We were having too many rounds of drip-feeding new findings, because the AI reviewer is a subjective grader. Even if it found the same finding before, the next time it runs it might score it wildly differently. And it tends to want to score its findings across the whole grading curve, because it thinks that's more correct-looking. Our answer is to use previous rounds scores as anchors for the next round. Since we can't just use the same context and still have good performance, we needed to condense the score report into a small json we send forward.
aliasxneo | 21 hours ago
hugs | 21 hours ago
skeledrew | 20 hours ago
exac | 20 hours ago
SA9G | 20 hours ago
rgoulter | 20 hours ago
Eh. I think you're just going to end up with slop, or sloppy recommendations?
My experience is that you can make different trade-offs for different reasons. I think even asking for the best answer to "improve the code, make better trade-offs".. even if you got a perfect response, there's no reason to think that it's the same set of trade-offs your actual use cases would benefit from.
folkrav | 20 hours ago
zavec | 14 hours ago
kqr | 14 hours ago
jaynetics | 13 hours ago
simonw | 20 hours ago
The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?
With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.
Muromec | 20 hours ago
Gets annoying pretty quickly
beybol | 13 hours ago
theodorewiles | 20 hours ago
bunderbunder | 20 hours ago
I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.
bakies | 19 hours ago
kqr | 14 hours ago
Richard Gabriel wrote something that has really stuck with me:
> Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.
This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.
A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!
stabbles | 14 hours ago
olalonde | 11 hours ago
ephaeton | 7 hours ago
IMO, It's the same idea with REPL-driven development in a functional language, why write unit tests if an interface is impossible to behave differently over time without changes?
Typically someone then argues, well, what if you change the implementation. To that person I will point out that the change will be developed in a REPL just like the original version, and thus be tested when it hits. Oh well...
To some, QA means a lot of compute and green outputs, to others, QA means spending quality hammock time before hitting the REPL :)
anilakar | 9 hours ago
Sigh.
Write a C or C++ API that works with pointers. Make it handle null pointers and errors elegantly so that the API user can safely chain calls and only check the final result. Claude decides it's better to be safe than sorry and peppers its code with intermediate nullptr and return value checks anyway.
glownagger | 9 hours ago
faremint | 9 hours ago
tartakovsky | 9 hours ago
tshaw96 | an hour ago
- We haven't really solved software engineering yet, it's hard to tell upfront whether a design adapts well to future needs. The community has a collection of heuristics which are sometimes at odds with each other.
- Therefore: write all designs, each simulated against a random tree of future needs.
- Pick whichever minimises the average future diff.
In other words, monte carlo simulation for software design. This scheme obviously assumes the price of code production goes even closer to zero.
paimapi | 45 minutes ago
that gut check of 'does this feel good to use' is funny to hear as QA because most of the feedback I give that boils down to 'this sucks to use' nearly universally gets met with 'but the AC! SLO! SLA! it's designed this way on purpose! it's a feature not a bug' etc.
I mean, I'm glad SWEs are finally discovering that using the product of your own work is actually a net good but it's a little like seeing a bullet coffee evangelist author a blog post about how he's found the real secret to preventing congestive heart failure and it's to not drink butter regularly
[0] https://en.wikipedia.org/wiki/Usability_testing