I wish this covered more features specific to rust that make it more runtime safer not just compile time safer. I guess by the time it's IR it's the same but moving it back up to C would be nice to see. Like bounds checking.
I think those are covered. Any checks that Rust adds will be present in the IR (either coded in the original source or implied through rust semantics). Those checks would then get put in the compiled C output. The ideal output of Eurydice would be to have exactly the same semantics and safety as provided by Rust. Since their use case is to use Rust to code on platform without support for a rust compiler, it is a rather important goal to have too.
I don't see what's wrong with e.g. LLVM IR for that purpose.
C seems to me a really poor choice because it's barely more expressive than one of the machine code architectures and yet unlike them it's not even a real machine we can know about - some of the edge cases are just a shrug emoji.
The LLVM IR is buggy of course, and under-specified, but we could imagine improving those things more easily than fixing C for this purpose.
Both you and I think C is more readable than LLVM IR but I'm not at all convinced that's just universally true -- I think it's because we're both experienced C programmers and so it's just our bias.
In particular I think LLVM's poison and undef values are valuable for understanding what some C++ means and why. Seeing that some C++ ends up as a select to choose between certain values, some of which are poison, makes it apparent why an optimisation choice is sound or not, where the likely C would leave this unsaid.
LLVM IR is very low-level and verbose and I think that makes it generally less readable. It is much like assembler. I have not seen people using it to write useful programs, and I think this would be painful.
I agree that undef / poison etc. are perhaps useful for optimization, but I was not proposing C as an intermediate language for optimization but as the result of C++ lowering. There are still some weird C++ semantics and specific features that would need extensions.
In fact, looking at several languages since Fortran, some kind of minimal bytecode was a common approach.
That is also how Pascal P-Code started, UCSD reused and extend it for UCSD Pascal, however Niklaus Wirth designed it originally to port Pascal compilers.
Then he did it again with Modula-2 (M-Code), and Oberon (slim binaries, although here the idea came from Michael Franz).
Naturally there were many others, and best of all no UB surprises with backend optimisations.
The C++ subset target is of theoretical interest and because some compilers cannot support a C++ compiler of the latest standard. OpenWatcom comes to mind.
Of course the question is which C++ standard in and what C standard out!
It depends entirely on your trust model. If you just want to be able to compile a rust compiler on a new platform that only has a C compiler, you wouldn't care. If you worry about a Thompson-style trojan horse, then you want to be able to audit all sources in the bootstrap chain.
Some idealistic projects like live-bootstrap go a step further and don't even allow generated sources in the full source bootstrap chain, even if they are readable in principle, because the xz debacle showed that a configure script, which is also readable in principle, can contain a backdoor. In those projects you could probably still bootstrap Eurydice from sources, and then do the rust -> c translation of the rust compiler as an ordinary build step, before compiling it with a bootstrapped C compiler.
> If you worry about a Thompson-style trojan horse, then you want to be able to audit all sources in the bootstrap chain.
Wouldn't you want to audit the final outputs? Except nobody is going to bother analyzing my hand all the assembly produced by the final tool.
In other words, I am not seeing the value of pursuing the discovery of Thompson-style Trojans, no matter what the original source language was. It is one of these risks that one has to live with.
Stabbles gets it! (Also, thank you for your other excellent posts on HN -- you are clearly not a bot like 50% of the other accounts are!)
Now, trust/security is one aspect of things, that's true...
But another important aspect of all of this is building understandable, maintainable, controllable software and AI's from the ground up, for future generations...
Let's suppose that in the future, there is no longer source code for compilers, they are only black-box compilers that future AI (which is also black-box by then) uses... At that point, humanity has lost control! No more source can be compiled without a black-box compiler, without an AI, and any source code so compiled will be controlled by whatever unknown things have been implanted in the original compiler!
Phrased another way, when humanity requires AI to generate new AI's (i.e., no scientists, computer scientists, programmers or other people know how to generate an AI from the ground up, from transistors on up, and control every layer of abstraction on top of that from assemblers to compilers to neural networks)... then humanity is lost, because at that point, humanity (not knowing how to build AI again from scratch, and control it at every step) quite literally has lost control of AI!
Future humanity cannot have that.
Future humanity requires, absolutely requires, transparency at every abstraction level... no black boxes should exist anywhere!
Future humanity also requires the requisite education, to be able to understand, manage and control each level of hardware/software abstraction, but that is outside of the scope of this HN post... :-)
I highly recommend checking out the other projects from AeneasVerif
(and Jonathan Protzenko).
Scylla is kind of Eurydice's dual. But also Charon and the eponymous Aeneas.
Also,
"There is ongoing work to integrate Eurydice-generated code for both Microsoft and Google’s respective crypto libraries." [1]
which I find incredibly cool. This link is also a brilliant intro to Eurydice.
”This can result in several different implementations of a function that differ only by type — often, the more idiomatic C approach would be to use macros or void * arguments.”
”Idiomatic” - is a ridiculous tautology I wish people discussing languages would stop using.
The only non-idiomatic code written seriosly is that which the compiler/interpreter does not understand.
Claiming universal superiority of style is shallow and thoughtless.
Sure - when discussing C of all languages this expression felt shallower than usual though.
If you are really enthusiastic you can write _obsfuscated_ C. And people can kind of agree what that is when they see it.
But _idiomatic_ ? Good luck defining what that is.
The term has sort of stuck in Python and there it's purpose is a bit more clear. There is this _object_ way of doing things, there is this _expression_ way of doing things etc and the interaction patterns work really nicely.
But you can't really use that expression in other languages which lack the same level of institutional design without sounding very much out of touch.
Neywiny | 21 hours ago
Lvl999Noob | 21 hours ago
Neywiny | 10 hours ago
chias | 17 hours ago
"You're da C"
nfw2 | 16 hours ago
Only know the pronunciation due to hadestown
nfw2 | 16 hours ago
egeozcan | 15 hours ago
ELI5, anyone?
Twisol | 15 hours ago
scns | 11 hours ago
True, it is "you-ri-di-cee"
1718627440 | 10 hours ago
fithisux | 17 hours ago
pjmlp | 15 hours ago
C++ started as a preprocessor that translated into C, nowadays all relevant C compilers are written in C++.
D has enough backeds already available.
uecker | 14 hours ago
I also think it would be helpful to how have a useful intermediate result to look at. So I agree, having C++ to C compiler would be useful.
pjmlp | 13 hours ago
Those people should move on, there isn't a single relevant platform without at least C++98 compiler, and there are other ways to bootstrap compilers.
tialaramex | 11 hours ago
C seems to me a really poor choice because it's barely more expressive than one of the machine code architectures and yet unlike them it's not even a real machine we can know about - some of the edge cases are just a shrug emoji.
The LLVM IR is buggy of course, and under-specified, but we could imagine improving those things more easily than fixing C for this purpose.
uecker | 11 hours ago
tialaramex | 9 hours ago
In particular I think LLVM's poison and undef values are valuable for understanding what some C++ means and why. Seeing that some C++ ends up as a select to choose between certain values, some of which are poison, makes it apparent why an optimisation choice is sound or not, where the likely C would leave this unsaid.
uecker | 8 hours ago
I agree that undef / poison etc. are perhaps useful for optimization, but I was not proposing C as an intermediate language for optimization but as the result of C++ lowering. There are still some weird C++ semantics and specific features that would need extensions.
pjmlp | 10 hours ago
That is also how Pascal P-Code started, UCSD reused and extend it for UCSD Pascal, however Niklaus Wirth designed it originally to port Pascal compilers.
Then he did it again with Modula-2 (M-Code), and Oberon (slim binaries, although here the idea came from Michael Franz).
Naturally there were many others, and best of all no UB surprises with backend optimisations.
uecker | 10 hours ago
It is easy to avoid UB surprises in C when generating code, so I do not think there is any good argument against C at this point.
fithisux | 12 hours ago
The C++ subset target is of theoretical interest and because some compilers cannot support a C++ compiler of the latest standard. OpenWatcom comes to mind.
Of course the question is which C++ standard in and what C standard out!
pjmlp | 12 hours ago
As for not having a compiler, do it like in the old days, cross compilation.
See Go and Zig for good examples of old practices brought into modern times.
IshKebab | 10 hours ago
fnord77 | 16 hours ago
kevincox | 10 hours ago
freecodeio | 3 hours ago
stabbles | 15 hours ago
knxsnsn | 13 hours ago
stabbles | 12 hours ago
Some idealistic projects like live-bootstrap go a step further and don't even allow generated sources in the full source bootstrap chain, even if they are readable in principle, because the xz debacle showed that a configure script, which is also readable in principle, can contain a backdoor. In those projects you could probably still bootstrap Eurydice from sources, and then do the rust -> c translation of the rust compiler as an ordinary build step, before compiling it with a bootstrapped C compiler.
david-gpu | 11 hours ago
Wouldn't you want to audit the final outputs? Except nobody is going to bother analyzing my hand all the assembly produced by the final tool.
In other words, I am not seeing the value of pursuing the discovery of Thompson-style Trojans, no matter what the original source language was. It is one of these risks that one has to live with.
vlovich123 | 7 hours ago
* compile “real” compiler with a different trusted compiler, producing “independent real compiler”
* compile the “real” compiler again using itself and the “independent compiler” - if the outputs don’t match, then you have detected the issue.
Source level transformation is a brilliant way to generate the “independent compiler”.
https://en.wikipedia.org/wiki/Backdoor_(computing)#Counterme...
genxy | 3 hours ago
david-gpu | 50 minutes ago
[OP] peter_d_sherman | 5 hours ago
Now, trust/security is one aspect of things, that's true...
But another important aspect of all of this is building understandable, maintainable, controllable software and AI's from the ground up, for future generations...
Let's suppose that in the future, there is no longer source code for compilers, they are only black-box compilers that future AI (which is also black-box by then) uses... At that point, humanity has lost control! No more source can be compiled without a black-box compiler, without an AI, and any source code so compiled will be controlled by whatever unknown things have been implanted in the original compiler!
Phrased another way, when humanity requires AI to generate new AI's (i.e., no scientists, computer scientists, programmers or other people know how to generate an AI from the ground up, from transistors on up, and control every layer of abstraction on top of that from assemblers to compilers to neural networks)... then humanity is lost, because at that point, humanity (not knowing how to build AI again from scratch, and control it at every step) quite literally has lost control of AI!
Future humanity cannot have that.
Future humanity requires, absolutely requires, transparency at every abstraction level... no black boxes should exist anywhere!
Future humanity also requires the requisite education, to be able to understand, manage and control each level of hardware/software abstraction, but that is outside of the scope of this HN post... :-)
Anyway, Stabbles, you get it!
nlehuen | 9 hours ago
If readability is the selling point, why not go the extra millimeter and add a heuristic so that this variable is named return_value for instance?
weinzierl | 9 hours ago
Scylla is kind of Eurydice's dual. But also Charon and the eponymous Aeneas.
Also, "There is ongoing work to integrate Eurydice-generated code for both Microsoft and Google’s respective crypto libraries." [1] which I find incredibly cool. This link is also a brilliant intro to Eurydice.
[1] https://jonathan.protzenko.fr/2025/10/28/eurydice.html
fsloth | 8 hours ago
”Idiomatic” - is a ridiculous tautology I wish people discussing languages would stop using.
The only non-idiomatic code written seriosly is that which the compiler/interpreter does not understand.
Claiming universal superiority of style is shallow and thoughtless.
Intentional obfuscation is a different thing.
genxy | 3 hours ago
When computer people name things, it is with a tenuous grasp of meaning of the words they are using. You kinda have to roll with it.
fsloth | 2 hours ago
If you are really enthusiastic you can write _obsfuscated_ C. And people can kind of agree what that is when they see it.
But _idiomatic_ ? Good luck defining what that is.
The term has sort of stuck in Python and there it's purpose is a bit more clear. There is this _object_ way of doing things, there is this _expression_ way of doing things etc and the interaction patterns work really nicely.
But you can't really use that expression in other languages which lack the same level of institutional design without sounding very much out of touch.