Incremental seems to keep winning the easy wins, so the interesting part is whether the remaining compile-time still lives in the same places as last year.
The EverInitializedPlaces example really stands out. Going from ~1.5M to ~90K apply_effects_in_block calls by changing the CFG traversal is a reminder that the biggest compiler optimizations often come from changing the algorithm, not optimizing the hot loop itself.
It also seems like the new Polonius/trait-solver work is pushing compiler performance toward a more interesting problem: doing expensive analysis only when it is actually needed.
4.57% mean wall-time reduction across 629 benchmarks in two months is pretty remarkable. Great progress.
> reminder that the biggest compiler optimizations often come from changing the algorithm, not optimizing the hot loop itself.
Isn't this true for most optimizations, not just in compilers? My usual goto process for optimizing is "Find stuff we're doing that we don't have to do, re-evaluate what data structures we use and then re-evaluate what algorithms we use" basically, with minor changes depending on the results. Served me well so far, and haven't (intentionally) written any compilers.
Its similar to C++ compilers, there is no one reason. It is a language that tries to optimize a lot, its a big language, it does safety checks, it uses llvm which is a bit slow, its a language that makes use of generics which generate extra code etc etc.
C++ compilers have it better, despite the fame, because on the C++ world, people like their binary libraries (not compiling the world from scratch on each git clone), there better support for incremental compilation, and incremental linking, there are interpreters and hot code reload tooling available.
Some C++ compilers and linkers have better support for incremental compilation and linking, and do it more fine grained than Rust, e.g. VC++.
Some tooling ideas go all the way back to when C++ vendors started adopting ideas from Smalltalk and Lisp, e.g. Energize C++ or Visual Age for C++ v4.
Or build systems like ClearMake from ClearCase, where the object files and binary libraries are shared across everyone on the cluster with the same views (ClearMake speak for what files/branches are selected).
This is not actually the main reason, most of the time.
Generics/monomorphization and how iterators work results in a lot of compiler bytecode that has to be churned through. More bytecode = longer compilation. It increases the size of the (debug) binaries, the debuginfo in general, causes performance issues with debug binaries in some situations unless you bump the optimization level, causes more IO, etc.
It's unlikely that many people are using Rust without using Option<T> or Result<T, U> a fair bit. Idiomatic Rust fundamentally uses a lot of generics. And a basic for loop expands into quite a lot of intermediate representation due to Iterator.
Other languages support generics, monomorphization, and iterators (e.g. Zig or D), but they're not as slow as Rust to compile, what's the reason for that?
Trait solving isn't a bottleneck for Rust compilation unless you're doing some extremely cursed things like trying to implement Doom in the type system. At the end of the day it's still mostly that Rust just generates a ton of IR for LLVM to chew on (because of monomorphization), then LLVM takes a while to process all of it (because LLVM is designed primarily to produce high-performance code, often resulting in reasonable tradeoffs against compilation speed), and then the linker takes a while to connect it all together. With an alternative backend to LLVM (e.g. Cranelift) you could choose to design it with a greater emphasis on compilation speed (likely trading off generated code quality in the process). And with a more radical vertically-integrated model where Rust controls the linker you could do in-place linking and nearly completely skip the final step (at least for debug builds), but that's a big change compared to the classic C-style build model.
It's likely easier to answer for specific languages
All three of those words can mean "just like Rust" but equally "Not at all like Rust" for different languages.
Two examples to contrast: In C++ the iterators are basically a pointer analog (in some cases they're just literally pointers) and that's a very difference "feature" but it's still definitely iterators. In Ginger Bill's Odin, the iterators are a function, possibly generic, which returns a pair, the next item and a boolean telling you whether the iterator was exhausted.
First of all, the fact that the article talks about "speeding up" the rust compiler doesn't automatically mean that the compiler is "slow"[0].
Now, is rustc slower than e.g. clang? by how much? why?
Those are different (and complicated) questions. It really depends on what you're compiling, but I'd say rustc can be 1-5x slower (maybe more at times?).
The reasons are many and varied, but in general rust compilation is slower because the compiler is doing way more things compared to C (monomorphization, complex trait resolution + type inference, borrow checker..)
[0]: Also I'd argue that "slow" without a concrete point of reference is a meaningless term in this context.
The go compiler is considered fast, yet Google spends effort speeding up the go compiler. Not to the same extent, but Google has a lot of go code, so it does make a difference.
Doing stuff isn't free. For instance, Go compiles relatively quickly for a modern language, but the biggest reason for that is that it does less stuff than most compilers... less optimization, less checking, and some stuff built into the language to avoid some of the problems with having to read lots of headers just to compile a file and other ways of doing less stuff, but mostly the key is it does less stuff, in both the good and bad senses of that.
If you want something like Rust that offers guarantees and checks and cross-checks by the boatload, it adds up. Macros, monomorphization, implicit code generation with traits and all those other things add up too. And you can't always get O(n) or O(n log n) code to implement those checks. Maybe it can be sped up and maybe there's tricks here or there, but at the Pareto frontier, a language that has more checks will be slower to compile than one that has fewer.
And that's not a bad thing or a deficit in Rust, it's just the nature of the beast.
And since you bring up Go and contrast the compile time with Rust, it's been experienced at Google (and also Volvo and other places) that Rust and Go teams are as productive whereas C++ is less than half as productive: https://www.youtube.com/watch?t=27012&v=6mZRWFQRvmw&feature=...
So the focus on Rust compile time is misplaced. It's not a big deal in terms of overall productivity.
I don't think that follows - just that compile times aren't the only thing that matters. Maybe rust could be twice as productive as go if compile times went to zero for instance (or maybe not - just disputing that the evidence proves the claim here).
But it can't go to zero, that was the point -- at some level you have to pay for what you get. I'm not saying the compile times couldn't be better, but Rust can't be Go in terms of compile times if it wants to offer the guarantees and options it does. Meanwhile if Go wants to offer more options and stronger guarantees, it won't be able to preserve the short compile times everyone points to.
The Cranelift backend is extremely fast. Most of the slowness is all the crazy optimizations that LLVM does + other things (generating debug info etc).
For symbol heavy projects, linking is a surprising bottleneck.
Some ideas to speed up compilation by not evaluating items that are not used might pan out significantly for big crates in your dep tree (that behavior might never be stable because that would allow items with compile errors in a crate that would still let your application compile, which is against the Rust approach). The same work to do that would also allow overlapping of crate evaluation between different rustc instances called by cargo, as it would require partial evaluation of crates (to do name res only and gather the symbols needed from its deps).
Another thing is that stable rust doesn't treat macros as idempotent (because that wasn't a requirement from the start, there are crates that do dynamic IO to generate types), but if they are then incr comp can be faster by not evaluating them unnecessarily.
I know people are working on a bunch of different strategies to improve both first and incremental compile times, and I'm looking forward to the fruit of their labor.
Compile times are rarely the bottleneck for me, but that doesn't mean I won't welcome any improvements on that front.
> For symbol heavy projects, linking is a surprising bottleneck.
Is that still true with modern linkers like `wild`? Seems like it can link Chromium in 1-2s.
> Another thing is that stable rust doesn't treat macros as idempotent
This seems absolutely insane to me given how pervasive macros are in Rust (I wonder how much of incremental compilation time is just repeated serde derives?). Obviously we can't just blindly treat all macros as idempotent, but an opt-in attribute on a macros that pinky-promises that it is seems like it ought to be pretty easy to implement?
Maybe somebody's looked into it and it doesn't help much? But I've heard (unverified) rumours of the opposite.
> Is that still true with modern linkers like `wild`? Seems like it can link Chromium in 1-2s.
Wild improves things significantly, but it really depends on the kind of project. For some a third of the time can be linking. And not everyone configures their environment to use a different linker. Another thing that can easily improve performance is changing the global allocator, but I've seen teams that measured 10% perf improvements in their own metrics decide against going with anything other than the "default".
I fully expect that if wild delivers an effective incremental linking architecture, then rustc will be able to produce the patches directly (instead of wild having to produce patches from two versions of the object files), which would mean both that rustc is producing less LLVM bytecode and that wild gets to do what it (will) do best and make linking as fast as mechanically possible.
> This seems absolutely insane to me given how pervasive macros are in Rust
Part of the problem I see is that crates will have to opt-in (and we might be able to change the default over an edition boundary) to get the perf benefit, and it might require a granularity lower than "a crate" (which would complicate implementation). I haven't seen additional discussions around these design considerations (needed to stabilize), which will need to happen before people can see progress.
Sadly, Rust is, like many open source projects, a show-up-o-cracy: you need a motivated (group of?) individual(s) to deliver a feature to fruition, and if the person driving a feature is fine with using nightly for their purposes, and only needs a subset of a feature, they will drive the feature to that state (lets say 80% completion), and then the feature will linger (until someone else with enough motivation to see the other 80% through shows up). This is exacerbated because the project is unwilling to have 80% solutions on stable unless the state is very clearly not going to preclude other future work, or the path to completion is 100% visible (but if that were the case, then it would have been completed already). So we end up with situations like the Allocator APIs.
(I did a quick test, and I'm seeing between 15-20% for https://github.com/servo/stylo which is a macro-heavy crate. I guess it only matter for crates you're actually editing though, macro will already be cached as part of "entire crate is unchanged")
> Another thing that can easily improve performance is changing the global allocator, but I've seen teams that measured 10% perf improvements in their own metrics decide against going with anything other than the "default".
Yeah, although I think there's more of a trade-off there. Alternative allocators can add significant amounts of compile time. Whereas, modulo maturity, I think a faster linker is more of a pure win. I would imagine we will make wild the default linker at some point if development continues on the trajectory it seems to be on.
I’m not sure if Cranelift support will ever become stable. It’s really complex, it has tradeoffs like you mentioned, and it requires its own backend to be maintained. TPDE looks like something that could more realistically be stabilized as it just slots into the existing LLVM support.
Rust wants badly to have its cake and eat it too. I get it, but it's not the sensibilities I have. The rustc compiler, if it really can achieve everything all at once, will be a beast like none other, making C++ compilers look simple by comparison.
I'd love something halfway. Go is maybe a bit radical in some regards, but also, with the news of the new SIMD package for Go, it has occured to me just how little I missed having things like, say, autovectorization.
(I know also that some people have tried halfway, but the big thing is figuring out how to keep a relatively simple type system that can still support a borrow checker. Even if there is some middleground, is it truly worth it? As nice as it sounds, I've been more skeptical. Go seems to exist in a very narrow space where its simplifications barely can be made to work.)
Instead of Go, you could have reached out to complex languages with fast compilation times like D, OCaml, Haskell, Ada, Delphi, C++.
All of them have alternative implementations with fast compilation times.
D, use dmd for fast development workflows, gdc or ldc for the ultimate performance at the expense of compilation times.
OCaml, use the REPL or bytecode interpreter for fast development times, the full blow compiler for ultimate performance.
Haskell, use the REPL, GHCi for the fast development cycles, GHC for the release build.
Ada and Delphi, have had fast implementations since forever, although Ada/SPARK is indeed somehow expensive.
C++, yes it isn't a mistake. Use Live++, VS hot reload, coupled with binary libraries, or a REPL like CINT (nee ROOT), binary libraries for dependencies, incremental compilation and incremental linking for the development workflow.
The problem with Rust isn't the language itself, rather the ecosystem currently lacking such kind of options being available.
including c++ with all of those qualifiers would be like me saying rust is fast with sccache, subsecond, cranelift, and making every single module a separate crate.
The difference being the adoption culture across the ecosystem.
C++ game engines also don't need Bevy like tutorials, because most studios aren't compiling them from scratch, and tools like Live++ or VC++ hot reload are relatively easy to use.
unreal engine is compiled from scratch (depending on your definition of "from scratch") and has notoriously long compile times. heck you have to recompile the editor to make certain gameplay code changes. and it's by far the most popular AAA game engine on the planet. if it was so easy to fix with your suggestions they would have done it already.
but hey, just add subsecond and switch to cranelift. problem solved :)
Unreal Engine has an installer, and they have done my suggestions, as they are the main customers of Live++ [0], advocates of tooling like Blueprints and now Verse for doing full games.
Many studios have delivered games with little to no changes to the underlying C++ code.
so... c++ compile times are fast, because you can simply not compile c++ and instead use blueprints? seems odd since we're talking about c++/rust compile times. turns out you can make rust compile times infinitely faster by not compiling it too...
but when you do have to compile c++, it can be real slow. maybe not as slow as rust, but it's not in the same league as eg. go, c#, zig, etc.
ue5's hot reload is by no means perfect either - many gameplay code changes require recompiling the editor. exposing c++ properties on a blueprint? recompile. modify a constructor? recompile. changing parameters, return types, or adding/removing UFUNCTION or UPROPERTY macros? recompile. you can simply search the web to read the experiences of thousands of ue devs complaining about slow compile times and workflows. it's the same with unity btw, search "reloading domain unity".
also verse is not in unreal engine. no studio is using verse for ue games. not sure why you threw that in.
i get you love c++ and dislike rust, but you aren't really making a good case for c++ when you group it in with other languages that have fast compile times ootb, make arguments that involve not using c++ (like blueprints, lol), and rely on brittle third-party hot-reloading solutions.
Zero-cost abstractions aren't zero cost in compilation time. High-level abstractions translate to a lot of boilerplate that the compiler has to optimize out.
In unoptimized builds often the linker is the bottleneck. Rust/Cargo can parallelize most of the build, generating tons of code and debug info, but then the poor linker has to consume all of it at once. The object/exe formats were designed in ancient times, so they're hard to build incrementally or in parallel (some linkers are trying).
Is there any experimentation with new ABIs? I mean certain linker flags like --f-lto literally hijack the linker protocol to dump an AST into the backend.
At least for the fully static binary part of rust, there should be some optimizations there w.r.t. compilation. Sure you're not going to interface with shared libraries well but maybe a small experimental feature for fully owned projects? Idk.
I once heard (and don't know if it's still true) that the biggest compile-time sinks are macros and codegen.
Both are kind of outside the Rust compiler's influence. Macros can be almost arbitrarily complex: you pay for what you order. Codegen is LLVM, and that's a fixed choice. You can use Cranelift to get around it, but then you pay elsewhere.
Also, generics and monomorphization regularly come up in these discussion, while common wisdom seems to be that cost for the additional static analysis over other languages like C++ is no a major contributor.
Regardless, it's nice to see performance improvements in the compiler, even if you have to cooperate to benefit from them (e.g. by keeping your macros light and use less generics).
Ultimately it's the fact that compile time was not a first-class consideration during the design of Rust's important features. There's only so much you can do to mitigate the consequences.
Compilation times were a consideration, but they were a subservient consideration to the prime considerations of 1) being as memory-safe as GC'd languages, 2) exploring the limits of statically-guaranteed thread safety, and 3) being as fast and memory-efficient as C and C++. If Rust had been willing to compromise on those goals and instead had just, for example, used a virtual machine with pervasive garbage collection and been designed for dynamically-dispatched generics then you'd get a lot of compilation time reductions for free, but the world didn't need another Java, it needed a more secure systems language to stand up against C++ where every previous challenger had failed.
No, Rust is not at the paretto frontier for your mentioned 1, 2, 3 and also 4 compilation speed. It could have all of those things and also just have faster compilation. Rust devs have talked about how they have some regrets about not optimizing more for compilation, but instead another attribute got optimized in terms of paretto efficiency - the language's development time. Arguably I think that is the single worst trait you can optimize for in language development, because by saving a little time developing the language you cost humanity, let's say hundreds of millions of hours of development time downstream from you, with millions of users being less productive.
I'm glad to see the donations from big companies to open source maintainers are making measurable difference to the Rust experience. Telling these companies that their employees spend 5% less time waiting for compilation might motivate future investment in people like Nick and the others mentioned.
LLMs are useful in engineering even if the PRs don't include LLM-written code.
>I’m still writing all my own code and text, because (a) that’s paramount, and (b) the project policy requires it, but I had useful LLM analysis assistance on several of the PRs mentioned in this post.
This is one of the most important objectives in software today.
Rust is the best language to serialize agentic LLM output to. It's native, well constructed, low-defect due to design. It's also easy for humans to read and debug if necessary.
The biggest problem with Rust is the compile times. The cycle has to get faster. And we need to start thinking about making the artifact cache non-blocking so multiple agents can work simultaneously - that'll be a big task, but essential if we want to speed up work on one machine rather than spinning up clusters of agent sandboxes (the alternative, perhaps superior solution).
OpenAI is a Rust Foundation platinum level donor. They gave straight cash which can go toward paying maintainers, which is better IMO. This actually made a lot of people angry because they saw it as a risk of OpenAI gaining influence or something. Can't make everyone happy.
OpenAI also hands out free subscriptions to open source maintainers: https://developers.openai.com/community/codex-for-oss They're time limited for now, but they've been pretty generous about who gets them. Anyone who can show Rust contributions would have been able to get one before.
I tend to disregard anyone who complains about foundational open source being funded like this. As long as the donor doesn't have dictatorial control over the product, why obsess over this?
Because not having dictatorial control doesn't protect anyone from anything; politicians take 'donations' (bribes) and then argue or enact laws for what the donor wants, no dictating involved.
Or in software terms, BigCorp sponsors you to develop an open source tool fulltime. If you make the tool more useful for them while making it more complex or worse or ignoring things for others, they might fund you next year as well. If you don't develop things that help them, they might stop and send that money elsewhere and you will need to get a normal job instead and drop your tool work to part-time. Even though they aren't dictating anything, what are you incentivised to do?
There's definitely a cost in terms of features BigCo doesn't need that get deprioritized. But that work still gets done if other people need it, it's just slower, and I don't think you could ever pull off making things worse on purpose in a big community run project like Rust.
The only times this actually happened are C++ and CORBA as far as I remember. Both of them were bureaucratic design by committee efforts from the start, the people doing it weren't users but implementors looking to stop competition, and it all happened before the open source era.
Can't confirm this "pretty generous", many people with lots of contributions to popular projects got any from OpenAI, not even a response to their request.
Yes this seems like the main way to go. With dependency injection using "dyn" traits on top of that, it seems that this mostly solves the compilation issues with Rust. It might be better for tests and agents to split code in small chunks too.
I am experimenting with this in my project right now. I haven't reached a verdict yet, but in theory it sounds good. I think you're still paying the linking cost though?
Splitting out a project into multiple crates reduces what needs to be recompiled during development. I tend to break out areas of my application that aren't going to change regularly so I'm only having to rebuild the main application.
In a production scenario where you're probably building your application in some CI environment, unless you cache the release build artifacts to re-use when applicable, the entire project will be rebuilt every run.
I've moved from Rust to Go for most things because in the era of agents, being able to iterate quickly on a project is a huge advantage and Rust is way, way slower than Go for compilation. There are times when Rust is more appropriate, but for the vast majority of things Go is perfectly fine.
Rust doesn't let you write an impl a trait that wasn't declared in your crate for a type that wasn't declared in your crate. At least one of the two needs to be locally declared. This makes splitting projects into multiple crates harder than it could be, but on the other hand it makes the crates ecosystem less brittle than it otherwise would be.
Some people propose relaxing this rule for "workspaces" (local projects with multiple crates that are not published individually). I haven't researched whether that would be technically feasible.
If you built two impls, your compiler wouldn't know which to pick. Or to phrase it differently, if you wanted to be able to choose the right one, it wouldn't be 'orphan impls', it would be 'orphan impls plus some selection mechanism'. E.g. Scala implicits. And once you have Scala implicits, non-orphan impls probably start looking like a pretty sweet alternative!
btw. That could actually make it slower to compile. The compiler uses the orphan rule to short circuit some checks IIRC.
Nevertheless, people in the project are trying to figuring out a way to relax the rule while still maintaining coherence, so this might happen some day.
Very cool projects! I agree, desktop apps are one of the areas that Rust beats Go. It's too bad though, since I feel anything with a heavy GUI like this can benefit from a rapid iteration cycle which is painful with Rust.
I finally gave in to the rust hype and built a new project in it (or my agent fleet did). Almost instant regret. My poor machine with just 2TB HDD and 24 cores was almost immediately crippled by agents each working in their own sandbox, each one building and compiling
Most projects my fleet works on are in other languages (TS, Go, Elixir) and can comfortably handle 10+ agents working in parallel, but not rust. I had to cap the fleet to 5 workers and build a dedicated resource monitor to step in and tidy up every time the disk almost filled up
That and general progress on building is far slower, with far more time spent building and testing than any other language I use.
Ended up rebuilding in Go, the perf gains weren't worth it
You should have set up https://github.com/mozilla/sccache at minimum. Setting up compiler caching is routine for people coming from older compiled language backgrounds but it's often missed by people coming from languages like JS/TS who have never been exposed to these ideas.
There's a new project called kache https://github.com/kunobi-ninja/kache that takes this even further and handles more edge cases with worktrees, but it's newer.
If you were pulling the entire project fresh and rebuilding everything in every worktree all the time without any caching then it would be frustrating. You could have asked your agent to set up basic caching and solved most of your problem in minutes.
Cargo could definitely improve with some inter-project sharing by default, but simultaneously working with a bunch of workspaces of the same project is a very new phenomena...
You really shouldn't build software with rust if your workflow involves 10+ agents. Why aren't you using Python or Typescript for things like that?
Rust is what you use when you want to maximize on the runtime performance. It will build software that requires significantly less memory and maximize the CPU that will save time and energy. It's a system language for writing software meant to access the hardware.
Funny to see this now. I’ve got a private branch I am in the process of shaping up this weekend to show the compiler team. For deep nested projects like rust analyzer, if you emit the meta data about function types earlier for downstream slots to use, before successful type checking, you can start other crates earlier and use all the slots you have instead of sitting around waiting for the full type checking of the bodies (which other crates largely don’t care about). Something like 40% wall time speed up, maybe 10% to 15% if you have the parallel frontend on.
Sounds lovely, give it a try on the codex-rs which seems to be getting close to using 2K crates in their repository any time now. Takes like 30 minutes to compile on a my beefy machine and it's a goddamn TUI, not sure what's going on, and not too interested in diving into that beast either.
I will email you when I have something up for sure! First cargo check has it at about 28% reduced on codex rs, from 383 seconds down to 277 seconds on a 16 core machine. No -Zthreads yet though, and that typically has a similar effect and reduces the speed up the early meta data gives.
Nice, excited to see that. This is basically deeper pipelining, IIUC. Nick (the author of the blog) implemented the current version of this "emit metadata sooner," optimization and already experimented with pushing it even sooner a while back.
We have discussed this briefly on last All Hands, the problem will probably be dealing with compiler errors, when they happen in builds that already emitted metadata. I believe this was the reason why it wasn't merged in the first place.
Yes I think I read a post of his from like 2019 were something was tried like this before. My guess is that this will be an interesting and hopefully informative PR for you all to see, but will likely not be merged as is. I had to add the ability for rustc to pause after early meta data emission and that was hacked together I feel, not well designed per se.
I view this as speculative execution/compilation and just throw out any errors from a crate that depended on another crate that ultimately errors out. It has the same outcome (same errors are printed in both cases) you just have a chance of having totally wasted some cpu time (that was otherwise just sitting around though).
I recall a talk about makepad.dev, I think it was by Rik Arends, that explained how they achieved outstanding compilation times in rust for makepad.
I can't find that talk again, but it was quite interesting: the rust compiler is fast, but often times it has to perform a lot of unnecessary checks because crates contain more stuff than needed. By stripping unnecessary work, the makepad team made building pretty fast.
I think there are a few more. Especially cranelift, but I would not use it for production stuff and it is missing a lot of llvm intrinsics that will just trap if you are trying to use them. But for most it seems to be just fine for regular dev builds with a HUGE speedup.
torutofu | 8 hours ago
Citrusoff | 8 hours ago
It also seems like the new Polonius/trait-solver work is pushing compiler performance toward a more interesting problem: doing expensive analysis only when it is actually needed.
4.57% mean wall-time reduction across 629 benchmarks in two months is pretty remarkable. Great progress.
embedding-shape | 8 hours ago
Isn't this true for most optimizations, not just in compilers? My usual goto process for optimizing is "Find stuff we're doing that we don't have to do, re-evaluate what data structures we use and then re-evaluate what algorithms we use" basically, with minor changes depending on the results. Served me well so far, and haven't (intentionally) written any compilers.
dev_hugepages | 5 hours ago
Surac | 8 hours ago
What is the performance killer?
JMKH42 | 8 hours ago
pjmlp | 6 hours ago
panstromek | 6 hours ago
What do you mean by this?
pjmlp | 4 hours ago
Some tooling ideas go all the way back to when C++ vendors started adopting ideas from Smalltalk and Lisp, e.g. Energize C++ or Visual Age for C++ v4.
Or build systems like ClearMake from ClearCase, where the object files and binary libraries are shared across everyone on the cluster with the same views (ClearMake speak for what files/branches are selected).
ModernMech | 8 hours ago
dralley | 8 hours ago
Generics/monomorphization and how iterators work results in a lot of compiler bytecode that has to be churned through. More bytecode = longer compilation. It increases the size of the (debug) binaries, the debuginfo in general, causes performance issues with debug binaries in some situations unless you bump the optimization level, causes more IO, etc.
ModernMech | 8 hours ago
dralley | 8 hours ago
ModernMech | 7 hours ago
moritzruth | 7 hours ago
kibwen | 6 hours ago
tialaramex | 7 hours ago
All three of those words can mean "just like Rust" but equally "Not at all like Rust" for different languages.
Two examples to contrast: In C++ the iterators are basically a pointer analog (in some cases they're just literally pointers) and that's a very difference "feature" but it's still definitely iterators. In Ginger Bill's Odin, the iterators are a function, possibly generic, which returns a pair, the next item and a boolean telling you whether the iterator was exhausted.
gf000 | 6 hours ago
On top the borrow checker and other features are non-existent in these languages.
thevinter | 8 hours ago
Now, is rustc slower than e.g. clang? by how much? why?
Those are different (and complicated) questions. It really depends on what you're compiling, but I'd say rustc can be 1-5x slower (maybe more at times?).
The reasons are many and varied, but in general rust compilation is slower because the compiler is doing way more things compared to C (monomorphization, complex trait resolution + type inference, borrow checker..)
[0]: Also I'd argue that "slow" without a concrete point of reference is a meaningless term in this context.
Joker_vD | 7 hours ago
Well, if it weren't "slow" for some definition of "slow", nobody would bother speeding it up, would they?
bryanlarsen | 7 hours ago
jerf | 8 hours ago
If you want something like Rust that offers guarantees and checks and cross-checks by the boatload, it adds up. Macros, monomorphization, implicit code generation with traits and all those other things add up too. And you can't always get O(n) or O(n log n) code to implement those checks. Maybe it can be sped up and maybe there's tricks here or there, but at the Pareto frontier, a language that has more checks will be slower to compile than one that has fewer.
And that's not a bad thing or a deficit in Rust, it's just the nature of the beast.
ch4s3 | 8 hours ago
ModernMech | 8 hours ago
So the focus on Rust compile time is misplaced. It's not a big deal in terms of overall productivity.
gpm | 7 hours ago
ModernMech | 7 hours ago
vlovich123 | 6 hours ago
estebank | 6 hours ago
Some ideas to speed up compilation by not evaluating items that are not used might pan out significantly for big crates in your dep tree (that behavior might never be stable because that would allow items with compile errors in a crate that would still let your application compile, which is against the Rust approach). The same work to do that would also allow overlapping of crate evaluation between different rustc instances called by cargo, as it would require partial evaluation of crates (to do name res only and gather the symbols needed from its deps).
Another thing is that stable rust doesn't treat macros as idempotent (because that wasn't a requirement from the start, there are crates that do dynamic IO to generate types), but if they are then incr comp can be faster by not evaluating them unnecessarily.
I know people are working on a bunch of different strategies to improve both first and incremental compile times, and I'm looking forward to the fruit of their labor.
Compile times are rarely the bottleneck for me, but that doesn't mean I won't welcome any improvements on that front.
nicoburns | 6 hours ago
Is that still true with modern linkers like `wild`? Seems like it can link Chromium in 1-2s.
> Another thing is that stable rust doesn't treat macros as idempotent
This seems absolutely insane to me given how pervasive macros are in Rust (I wonder how much of incremental compilation time is just repeated serde derives?). Obviously we can't just blindly treat all macros as idempotent, but an opt-in attribute on a macros that pinky-promises that it is seems like it ought to be pretty easy to implement?
Maybe somebody's looked into it and it doesn't help much? But I've heard (unverified) rumours of the opposite.
estebank | 6 hours ago
Wild improves things significantly, but it really depends on the kind of project. For some a third of the time can be linking. And not everyone configures their environment to use a different linker. Another thing that can easily improve performance is changing the global allocator, but I've seen teams that measured 10% perf improvements in their own metrics decide against going with anything other than the "default".
I fully expect that if wild delivers an effective incremental linking architecture, then rustc will be able to produce the patches directly (instead of wild having to produce patches from two versions of the object files), which would mean both that rustc is producing less LLVM bytecode and that wild gets to do what it (will) do best and make linking as fast as mechanically possible.
> This seems absolutely insane to me given how pervasive macros are in Rust
There was some work done on this front, but I haven't kept up to date on the current status of that. There was a PR showing promise https://github.com/rust-lang/rust/pull/129102 (later landed as https://github.com/rust-lang/rust/pull/145354, 10% on a specific serde-heavy crate, reports of 32% improvements in the original PR). The tracking issue doesn't have any updates https://github.com/rust-lang/rust/issues/151364, but you can try out nightly with -Zcache-proc-macros to see what the effect could be on your projects.
Part of the problem I see is that crates will have to opt-in (and we might be able to change the default over an edition boundary) to get the perf benefit, and it might require a granularity lower than "a crate" (which would complicate implementation). I haven't seen additional discussions around these design considerations (needed to stabilize), which will need to happen before people can see progress.
Sadly, Rust is, like many open source projects, a show-up-o-cracy: you need a motivated (group of?) individual(s) to deliver a feature to fruition, and if the person driving a feature is fine with using nightly for their purposes, and only needs a subset of a feature, they will drive the feature to that state (lets say 80% completion), and then the feature will linger (until someone else with enough motivation to see the other 80% through shows up). This is exacerbated because the project is unwilling to have 80% solutions on stable unless the state is very clearly not going to preclude other future work, or the path to completion is 100% visible (but if that were the case, then it would have been completed already). So we end up with situations like the Allocator APIs.
nicoburns | 4 hours ago
> Another thing that can easily improve performance is changing the global allocator, but I've seen teams that measured 10% perf improvements in their own metrics decide against going with anything other than the "default".
Yeah, although I think there's more of a trade-off there. Alternative allocators can add significant amounts of compile time. Whereas, modulo maturity, I think a faster linker is more of a pure win. I would imagine we will make wild the default linker at some point if development continues on the trajectory it seems to be on.
dicytea | 6 hours ago
https://github.com/rust-lang/rustc_codegen_cranelift/issues/...
dabinat | an hour ago
jchw | 6 hours ago
I'd love something halfway. Go is maybe a bit radical in some regards, but also, with the news of the new SIMD package for Go, it has occured to me just how little I missed having things like, say, autovectorization.
(I know also that some people have tried halfway, but the big thing is figuring out how to keep a relatively simple type system that can still support a borrow checker. Even if there is some middleground, is it truly worth it? As nice as it sounds, I've been more skeptical. Go seems to exist in a very narrow space where its simplifications barely can be made to work.)
oscillonoscope | 5 hours ago
pjmlp | 6 hours ago
Instead of Go, you could have reached out to complex languages with fast compilation times like D, OCaml, Haskell, Ada, Delphi, C++.
All of them have alternative implementations with fast compilation times.
D, use dmd for fast development workflows, gdc or ldc for the ultimate performance at the expense of compilation times.
OCaml, use the REPL or bytecode interpreter for fast development times, the full blow compiler for ultimate performance.
Haskell, use the REPL, GHCi for the fast development cycles, GHC for the release build.
Ada and Delphi, have had fast implementations since forever, although Ada/SPARK is indeed somehow expensive.
C++, yes it isn't a mistake. Use Live++, VS hot reload, coupled with binary libraries, or a REPL like CINT (nee ROOT), binary libraries for dependencies, incremental compilation and incremental linking for the development workflow.
The problem with Rust isn't the language itself, rather the ecosystem currently lacking such kind of options being available.
Shorel | 6 hours ago
slopinthebag | 4 hours ago
pjmlp | 4 hours ago
C++ game engines also don't need Bevy like tutorials, because most studios aren't compiling them from scratch, and tools like Live++ or VC++ hot reload are relatively easy to use.
slopinthebag | 4 hours ago
but hey, just add subsecond and switch to cranelift. problem solved :)
pjmlp | 3 hours ago
Many studios have delivered games with little to no changes to the underlying C++ code.
[0] - https://dev.epicgames.com/documentation/unreal-engine/using-...
slopinthebag | an hour ago
but when you do have to compile c++, it can be real slow. maybe not as slow as rust, but it's not in the same league as eg. go, c#, zig, etc.
ue5's hot reload is by no means perfect either - many gameplay code changes require recompiling the editor. exposing c++ properties on a blueprint? recompile. modify a constructor? recompile. changing parameters, return types, or adding/removing UFUNCTION or UPROPERTY macros? recompile. you can simply search the web to read the experiences of thousands of ue devs complaining about slow compile times and workflows. it's the same with unity btw, search "reloading domain unity".
also verse is not in unreal engine. no studio is using verse for ue games. not sure why you threw that in.
i get you love c++ and dislike rust, but you aren't really making a good case for c++ when you group it in with other languages that have fast compile times ootb, make arguments that involve not using c++ (like blueprints, lol), and rely on brittle third-party hot-reloading solutions.
pornel | 8 hours ago
In unoptimized builds often the linker is the bottleneck. Rust/Cargo can parallelize most of the build, generating tons of code and debug info, but then the poor linker has to consume all of it at once. The object/exe formats were designed in ancient times, so they're hard to build incrementally or in parallel (some linkers are trying).
sigbottle | 6 hours ago
At least for the fully static binary part of rust, there should be some optimizations there w.r.t. compilation. Sure you're not going to interface with shared libraries well but maybe a small experimental feature for fully owned projects? Idk.
kreco | 7 hours ago
weinzierl | 7 hours ago
Both are kind of outside the Rust compiler's influence. Macros can be almost arbitrarily complex: you pay for what you order. Codegen is LLVM, and that's a fixed choice. You can use Cranelift to get around it, but then you pay elsewhere.
Also, generics and monomorphization regularly come up in these discussion, while common wisdom seems to be that cost for the additional static analysis over other languages like C++ is no a major contributor.
Regardless, it's nice to see performance improvements in the compiler, even if you have to cooperate to benefit from them (e.g. by keeping your macros light and use less generics).
senderista | 6 hours ago
kibwen | 5 hours ago
applfanboysbgon | 2 hours ago
jiehong | 2 hours ago
Zig and Go are similar on that front.
adamch | 8 hours ago
bryanlarsen | 8 hours ago
Sometimes we really can have our cake and eat it too.
maherbeg | 7 hours ago
OG_BME | 7 hours ago
https://forge.rust-lang.org/policies/llm-usage.html
poly2it | 7 hours ago
https://forge.rust-lang.org/policies/llm-usage.html#experime...
sigmar | 6 hours ago
>I’m still writing all my own code and text, because (a) that’s paramount, and (b) the project policy requires it, but I had useful LLM analysis assistance on several of the PRs mentioned in this post.
from the article
echelon | 6 hours ago
Rust is the best language to serialize agentic LLM output to. It's native, well constructed, low-defect due to design. It's also easy for humans to read and debug if necessary.
The biggest problem with Rust is the compile times. The cycle has to get faster. And we need to start thinking about making the artifact cache non-blocking so multiple agents can work simultaneously - that'll be a big task, but essential if we want to speed up work on one machine rather than spinning up clusters of agent sandboxes (the alternative, perhaps superior solution).
Aurornis | 6 hours ago
OpenAI also hands out free subscriptions to open source maintainers: https://developers.openai.com/community/codex-for-oss They're time limited for now, but they've been pretty generous about who gets them. Anyone who can show Rust contributions would have been able to get one before.
rirze | 6 hours ago
jodrellblank | 5 hours ago
Or in software terms, BigCorp sponsors you to develop an open source tool fulltime. If you make the tool more useful for them while making it more complex or worse or ignoring things for others, they might fund you next year as well. If you don't develop things that help them, they might stop and send that money elsewhere and you will need to get a normal job instead and drop your tool work to part-time. Even though they aren't dictating anything, what are you incentivised to do?
tancop | 3 hours ago
The only times this actually happened are C++ and CORBA as far as I remember. Both of them were bureaucratic design by committee efforts from the start, the people doing it weren't users but implementors looking to stop competition, and it all happened before the open source era.
maherbeg | 5 hours ago
silverwind | 5 hours ago
nicoburns | 6 hours ago
godwinson__4-8 | 6 hours ago
What you suggest may thus be more or less already happening.
hnp9j9qtda | 7 hours ago
clarus | 6 hours ago
rapind | 6 hours ago
geauxvirtual | 6 hours ago
Splitting out a project into multiple crates reduces what needs to be recompiled during development. I tend to break out areas of my application that aren't going to change regularly so I'm only having to rebuild the main application.
In a production scenario where you're probably building your application in some CI environment, unless you cache the release build artifacts to re-use when applicable, the entire project will be rebuilt every run.
ddalcino | 7 hours ago
slowin | 6 hours ago
echelon | 6 hours ago
I've stopped using Tauri and gone with 100% egui. It's cross platform and excellent, and if you give it design constraints it will look beautiful.
Check out my 100% adobe clean room reimplementations:
https://github.com/storytold/filmcraft
https://github.com/storytold/photocraft
https://github.com/storytold/drawcraft (going to rename this vectorcraft)
The #1 thing for the Rust project to do is make Rust faster to compile.
Rust is the agentic AI language. It just needs to lean in and go faster.
tcfhgj | 6 hours ago
#0 get rid of the orphan rule
Quitschquat | 6 hours ago
estebank | 6 hours ago
Some people propose relaxing this rule for "workspaces" (local projects with multiple crates that are not published individually). I haven't researched whether that would be technically feasible.
echelon | 4 hours ago
Workspaces are a great way to architect larger projects and monorepos, and they'd at least be internally consistent.
mrkeen | 2 hours ago
panstromek | 6 hours ago
Nevertheless, people in the project are trying to figuring out a way to relax the rule while still maintaining coherence, so this might happen some day.
the_sleaze_ | 6 hours ago
slowin | 6 hours ago
xutopia | 6 hours ago
jlahijani | 5 hours ago
mixmastamyk | 4 hours ago
Also it says egui is immediate mode, does that use a lot of {C,G}PU or have any other issues?
s08148692 | 5 hours ago
Most projects my fleet works on are in other languages (TS, Go, Elixir) and can comfortably handle 10+ agents working in parallel, but not rust. I had to cap the fleet to 5 workers and build a dedicated resource monitor to step in and tidy up every time the disk almost filled up
That and general progress on building is far slower, with far more time spent building and testing than any other language I use.
Ended up rebuilding in Go, the perf gains weren't worth it
slopinthebag | 5 hours ago
paholg | 5 hours ago
https://github.com/mozilla/sccache
pjmlp | 3 hours ago
Aurornis | 5 hours ago
There's a new project called kache https://github.com/kunobi-ninja/kache that takes this even further and handles more edge cases with worktrees, but it's newer.
If you were pulling the entire project fresh and rebuilding everything in every worktree all the time without any caching then it would be frustrating. You could have asked your agent to set up basic caching and solved most of your problem in minutes.
slopinthebag | 4 hours ago
jiehong | 3 hours ago
gpm | 2 hours ago
jarjoura | 2 hours ago
Rust is what you use when you want to maximize on the runtime performance. It will build software that requires significantly less memory and maximize the CPU that will save time and energy. It's a system language for writing software meant to access the hardware.
ameliaquining | an hour ago
virtualritz | an hour ago
Cargo run/test/nextest/run only, check/fmt etc should be excluded.
This avoids contention/oversubscription of the CPU which makes builds up to 300% slower from what I measured.
muthuh | 5 hours ago
knuckleheads | 5 hours ago
embedding-shape | 5 hours ago
knuckleheads | 3 hours ago
embedding-shape | 2 hours ago
knuckleheads | 2 hours ago
yearolinuxdsktp | 34 minutes ago
Incremental compilation is not cleaned up. Older deps pile up and don’t get cleaned.
Run out of disk space? Oh it’s just the 200+ GB codex-rs target folder.
Forget about doing worktrees.
Rust has a lot of work to do.
panstromek | 5 hours ago
We have discussed this briefly on last All Hands, the problem will probably be dealing with compiler errors, when they happen in builds that already emitted metadata. I believe this was the reason why it wasn't merged in the first place.
knuckleheads | 3 hours ago
I view this as speculative execution/compilation and just throw out any errors from a crate that depended on another crate that ultimately errors out. It has the same outcome (same errors are printed in both cases) you just have a chance of having totally wasted some cpu time (that was otherwise just sitting around though).
randypewick | 5 hours ago
I can't find that talk again, but it was quite interesting: the rust compiler is fast, but often times it has to perform a lot of unnecessary checks because crates contain more stuff than needed. By stripping unnecessary work, the makepad team made building pretty fast.
dabinat | 5 hours ago
But the real speedup will happen when TPDE is enabled: https://goals.rust-lang.org/2026/tpde.html
1vuio0pswjnm7 | 3 hours ago
Is it slower
sharktheone | an hour ago