So they used Prime Agent and GLM 5.3 to swarm and rewrite the code in Rust. This also shows Prime Agent doing what it preaches by rebuilding itself. Since Prime Agent is Pi under the hood, will they push a Rust rewrite to Pi? Pi extensions use Typescript so I wonder if they will work.
I don’t see many people talk about Prime Agent, I always wondered if it could just be a Pi extension cause it seems to be a subagent orchestration agent.
I feel like I talk about it often enough people might think I have an agenda. For a while it felt like a superpower compared to other harnesses, although they've caught up with persistent/backgrounded agents.
Very happy user. Loved it with deepseek-v4-flash, although I've moved on because newer models are so compelling.
I just get good results from it; I think it's the python requirement.
It has been very good with sub-agents, and also finding old context, for a long time. And had backgrounded agents that you can run from one instance without herdr. Herdr is kinda redundant (better UI than PA though).
I see nobody else talk about it. It's so little talk that when I mention it on X, the devs comment on my posts sometimes.
Kinda wish people break down token usage into input (cache hit), input (cache miss) and output when talking about it. Giving a total 200B tokens number doesn’t help gauge costs.
I was checking openrouter a few days ago to see if they provided that information yet. Would be a really nice breakdown to have under total task price. Doesn't map to every task, but averages would still be nice.
For Anthropic's reported benchmark numbers for Sonnet 5.5, when running terminal-bench 4.0, they had a comparison between $0.10 and $0.20 cache read for total task cost. That 50% cost reduction resulted in a 20.9-26.2% total cost reduction. Pretty specific though... single benchmark, model, harness, and provider. Cache reads make up ~42-52% of the total cost at the standard $0.20 price.
I'm interested in their Planner -> Implementer -> Reviewer -> Verifier process they used for this transition to Rust. I've see similar but curious how they actually implemented this.
Curious how this could be applied to greenfield coding rather than just making a copy in a new language or performance optimizing.
I know people tend to hate on rust rewrites. But having something that compiles to a single binary that you can just copy over and get started has its advantages especially when working with sandboxes etc.
> Overall, our Rust rewrite and following performance hillclimbing has made Prime Agent significantly faster and more resource efficient. With time to input roughly 14x faster than TypeScript and using over 80% less memory after startup
That sounds like a nightmare for comprehensibility, incremental development, and maintenance.
Combine that with the fact that, unlike other code generators, LLM output can't as a rule be reliably reproduced from the original input in the future, sounds like a recipe for unpredictable long-term costs and regular regressions.
> Eventually AIs will generate Assembly code directly anyway, no need for intermediate 3 GL languages.
For what its worth, Someone created an text editor application in pure assembly using AI (which is genuinely really cool!)[0]
To test this hypothesis, I tried to add scratchpad functionality to it and I was able to do so with an open source model almost autonomously by just having a single prompt which can have auto-saving, easy way to generate new files and its honestly really minimalist and that's the point while having multiple things as well and I personally really love it :-D[1]
I daily drive my scratchpad application and use it quite a lot
I was actually quite surprised by the fact that it actually worked and the fact that rhun itself can exist.
So it is possible that if AI agentic capabilities grow even further then assembly generation can be possible but I am unsure about preferred about some small use cases requiring pure minimalism.
But nonetheless, I think the way LLM's are trained, programming languages are good enough for them and they can express what they want to accomplish easily and without repeating themselves easily through it. I have thought about creating some more assembly applications but I just think that it would be a more token-intensive process, that's all.
Though pure speculation at this point but if some architecture optimizes for assembly (maybe speculative decoding extremely fine tuned to assembly?) or maybe some other architecture itself (world model?) who can express their ideas in better formats than tokens than perhaps its possible as code written in programming languages is much more token-efficient than assembly
theturtletalks | 21 hours ago
I don’t see many people talk about Prime Agent, I always wondered if it could just be a Pi extension cause it seems to be a subagent orchestration agent.
sejje | 17 hours ago
Very happy user. Loved it with deepseek-v4-flash, although I've moved on because newer models are so compelling.
I just get good results from it; I think it's the python requirement.
It has been very good with sub-agents, and also finding old context, for a long time. And had backgrounded agents that you can run from one instance without herdr. Herdr is kinda redundant (better UI than PA though).
I see nobody else talk about it. It's so little talk that when I mention it on X, the devs comment on my posts sometimes.
theturtletalks | 5 hours ago
As far as the dev comments, it seems Prime Agent doesn't have the hype as other harnesses but doesn't mean its not worth trying.
oefrha | 21 hours ago
ram1500natrluvr | 16 hours ago
For Anthropic's reported benchmark numbers for Sonnet 5.5, when running terminal-bench 4.0, they had a comparison between $0.10 and $0.20 cache read for total task cost. That 50% cost reduction resulted in a 20.9-26.2% total cost reduction. Pretty specific though... single benchmark, model, harness, and provider. Cache reads make up ~42-52% of the total cost at the standard $0.20 price.
mgreg | 21 hours ago
Curious how this could be applied to greenfield coding rather than just making a copy in a new language or performance optimizing.
rpunkfu | 15 hours ago
https://ctx.company/blog/introducing-ctx-traits/
helsinki | 21 hours ago
esafak | 19 hours ago
Tsarp | 20 hours ago
binary132 | 18 hours ago
AbuAssar | 19 hours ago
very nice outcome of this port
esafak | 19 hours ago
rvz | 18 hours ago
Now there is no excuses to not use Rust and it just shows in raw performance alone.
pjmlp | 13 hours ago
Eventually AIs will generate Assembly code directly anyway, no need for intermediate 3 GL languages.
jasomill | 13 hours ago
Combine that with the fact that, unlike other code generators, LLM output can't as a rule be reliably reproduced from the original input in the future, sounds like a recipe for unpredictable long-term costs and regular regressions.
pjmlp | 12 hours ago
JIT compilers, GC workflows, PGO, and ML optimising compilers passes are also non deterministic, and yet work gets done.
If you are curious, there is already enough work out there into this direction.
Imustaskforhelp | 8 hours ago
For what its worth, Someone created an text editor application in pure assembly using AI (which is genuinely really cool!)[0]
To test this hypothesis, I tried to add scratchpad functionality to it and I was able to do so with an open source model almost autonomously by just having a single prompt which can have auto-saving, easy way to generate new files and its honestly really minimalist and that's the point while having multiple things as well and I personally really love it :-D[1]
I daily drive my scratchpad application and use it quite a lot
I was actually quite surprised by the fact that it actually worked and the fact that rhun itself can exist.
So it is possible that if AI agentic capabilities grow even further then assembly generation can be possible but I am unsure about preferred about some small use cases requiring pure minimalism.
But nonetheless, I think the way LLM's are trained, programming languages are good enough for them and they can express what they want to accomplish easily and without repeating themselves easily through it. I have thought about creating some more assembly applications but I just think that it would be a more token-intensive process, that's all.
Though pure speculation at this point but if some architecture optimizes for assembly (maybe speculative decoding extremely fine tuned to assembly?) or maybe some other architecture itself (world model?) who can express their ideas in better formats than tokens than perhaps its possible as code written in programming languages is much more token-efficient than assembly
[0]: https://rhun.app/
[1]: https://github.com/serJaimeLannister/rhunpad
pjmlp | 13 hours ago
Naturally then there isn't source material for "we rewrote yet another slow scripting project into Go/Rust/Zig/C/C++/..." blog posts.