Rewriting Prime Agent in Rust

63 points by piotrgrabowski 23 hours ago on hackernews | 19 comments

theturtletalks | 21 hours ago

So they used Prime Agent and GLM 5.3 to swarm and rewrite the code in Rust. This also shows Prime Agent doing what it preaches by rebuilding itself. Since Prime Agent is Pi under the hood, will they push a Rust rewrite to Pi? Pi extensions use Typescript so I wonder if they will work.

I don’t see many people talk about Prime Agent, I always wondered if it could just be a Pi extension cause it seems to be a subagent orchestration agent.

sejje | 17 hours ago

I feel like I talk about it often enough people might think I have an agenda. For a while it felt like a superpower compared to other harnesses, although they've caught up with persistent/backgrounded agents.

Very happy user. Loved it with deepseek-v4-flash, although I've moved on because newer models are so compelling.

I just get good results from it; I think it's the python requirement.

It has been very good with sub-agents, and also finding old context, for a long time. And had backgrounded agents that you can run from one instance without herdr. Herdr is kinda redundant (better UI than PA though).

I see nobody else talk about it. It's so little talk that when I mention it on X, the devs comment on my posts sometimes.

theturtletalks | 5 hours ago

Thanks for your feedback, I'll give it a shot!

As far as the dev comments, it seems Prime Agent doesn't have the hype as other harnesses but doesn't mean its not worth trying.

oefrha | 21 hours ago

Kinda wish people break down token usage into input (cache hit), input (cache miss) and output when talking about it. Giving a total 200B tokens number doesn’t help gauge costs.

ram1500natrluvr | 16 hours ago

I was checking openrouter a few days ago to see if they provided that information yet. Would be a really nice breakdown to have under total task price. Doesn't map to every task, but averages would still be nice.

For Anthropic's reported benchmark numbers for Sonnet 5.5, when running terminal-bench 4.0, they had a comparison between $0.10 and $0.20 cache read for total task cost. That 50% cost reduction resulted in a 20.9-26.2% total cost reduction. Pretty specific though... single benchmark, model, harness, and provider. Cache reads make up ~42-52% of the total cost at the standard $0.20 price.

mgreg | 21 hours ago

I'm interested in their Planner -> Implementer -> Reviewer -> Verifier process they used for this transition to Rust. I've see similar but curious how they actually implemented this.

Curious how this could be applied to greenfield coding rather than just making a copy in a new language or performance optimizing.

rpunkfu | 15 hours ago

I’ll have to update the blog post, since API and capabilities of ctx grew a lot but one of examples here explain for we’re doing it.

https://ctx.company/blog/introducing-ctx-traits/

helsinki | 21 hours ago

If anyone wants a 100% parity Rust port of v1.0 DeepSeek Harness, I have one here: https://github.com/trevorprater/SeekDeep-Harness

esafak | 19 hours ago

What benefits have you seen from using the DeepSeek Harness? What were you using before?

Tsarp | 20 hours ago

I know people tend to hate on rust rewrites. But having something that compiles to a single binary that you can just copy over and get started has its advantages especially when working with sandboxes etc.

binary132 | 18 hours ago

Convenient yes, but also not just a rust thing!

AbuAssar | 19 hours ago

> Overall, our Rust rewrite and following performance hillclimbing has made Prime Agent significantly faster and more resource efficient. With time to input roughly 14x faster than TypeScript and using over 80% less memory after startup

very nice outcome of this port

esafak | 19 hours ago

Does anyone have experience to share about their Prime Agent harness? Does it do anything that regular harnesses paired with a memory plugin can't do?
We already knew TypeScript was the wrong language for many use cases from the start as an excuse to not learn Rust.

Now there is no excuses to not use Rust and it just shows in raw performance alone.

pjmlp | 13 hours ago

Any scripting language is the wrong language, when the job is performance only delivered by compiled languages.

Eventually AIs will generate Assembly code directly anyway, no need for intermediate 3 GL languages.

jasomill | 13 hours ago

That sounds like a nightmare for comprehensibility, incremental development, and maintenance.

Combine that with the fact that, unlike other code generators, LLM output can't as a rule be reliably reproduced from the original input in the future, sounds like a recipe for unpredictable long-term costs and regular regressions.

pjmlp | 12 hours ago

Spoken like an Assembly programmer when optimising compilers arrived into town.

JIT compilers, GC workflows, PGO, and ML optimising compilers passes are also non deterministic, and yet work gets done.

If you are curious, there is already enough work out there into this direction.

Imustaskforhelp | 8 hours ago

> Eventually AIs will generate Assembly code directly anyway, no need for intermediate 3 GL languages.

For what its worth, Someone created an text editor application in pure assembly using AI (which is genuinely really cool!)[0]

To test this hypothesis, I tried to add scratchpad functionality to it and I was able to do so with an open source model almost autonomously by just having a single prompt which can have auto-saving, easy way to generate new files and its honestly really minimalist and that's the point while having multiple things as well and I personally really love it :-D[1]

I daily drive my scratchpad application and use it quite a lot

I was actually quite surprised by the fact that it actually worked and the fact that rhun itself can exist.

So it is possible that if AI agentic capabilities grow even further then assembly generation can be possible but I am unsure about preferred about some small use cases requiring pure minimalism.

But nonetheless, I think the way LLM's are trained, programming languages are good enough for them and they can express what they want to accomplish easily and without repeating themselves easily through it. I have thought about creating some more assembly applications but I just think that it would be a more token-intensive process, that's all.

Though pure speculation at this point but if some architecture optimizes for assembly (maybe speculative decoding extremely fine tuned to assembly?) or maybe some other architecture itself (world model?) who can express their ideas in better formats than tokens than perhaps its possible as code written in programming languages is much more token-efficient than assembly

[0]: https://rhun.app/

[1]: https://github.com/serJaimeLannister/rhunpad

pjmlp | 13 hours ago

There are enough compiled languages to chose from, start there.

Naturally then there isn't source material for "we rewrote yet another slow scripting project into Go/Rust/Zig/C/C++/..." blog posts.