Incidentally, Zig 0.17.0 just shipped full support for x32, which means no-libc w/ direct syscalls, or w/ glibc, or musl. We also test it natively in CI; Debian still supports syscall.x32=y on the kernel command line. (We also shipped support for N32 which is the MIPS equivalent, but somehow I suspect people aren't going to be as interested in that one.)
I did it mostly as a fun exercise and to help catch bad pointer size assumptions in CI, but it would be nice if it accidentally turned out to be actually useful.
The RAM savings look very high, but I guess a language with that header structure and a lot of small objects would have these properties. In some of the analysis that we did for CHERI, we typically found that under 20% of pages contained one or more pointers. A lot of memory ends up being big buffers, images, and other data that doesn't contain pointers.
This is also a big part of the reason that x32 largely went away. One process using it may be smaller, but once you factor in the extra disk space and RAM from 32-bit and 64-bit shared libraries, you end up with a much less clear win. If everything used x32 then it would be a win probably but it turns out that JavaScript and some similar scripting languages like NaN boxing, so 32-bit isn't a win. Similarly, lightweight SFI things (e.g. WebAssembly) work nicely if you can carve out a 4 GiB chunk of the address space and just truncate pointers to that size (optionally with a base offset). So you find that things like Node.js and web browsers want a 64-bit system.
The same thing happens with JVMs, where it's nice to be able to map the heap in at least two places for generational GC. You also want to be able to have a large address range to put your Java objects in that doesn't overlap anything else, so a 2 GiB userspace address space may limit you to something on the order of a 300 MiB heap. And the pointer size increase doesn't matter for JVMs because 64-bit JVMs typically use a 32-bit integer to represent objects, which is an offset from the base of the heap right shifted by 3 or 4, so you can have a 64 GiB heap. And you can often put NIO buffers and large arrays in a separate heap region with only a small object header in the 64 GiB heap, which lets you actually scale to heaps that are a large fraction of a TiB, without any real pointer-size overhead.
I believe v8 (and consequently Node and browsers) doesn't use NaN boxing, and instead puts doubles on the heap: https://v8.dev/blog/pointer-compression . Ironically it's exactly to get the smaller pointer cost we're chasing with x32!
Memory use has grown to the point that it is challenging to shoehorn even a basic "hello world" app into the 2-3GB of virtual address space that 32-bit pointers afford.
~ Corbet, on LWN on (yet another) proposed removal for x32
Maybe one of the positive externalities of all these chip shortages will a second wind for memory frugality. Just because memory sizes had been doubling doesn't mean our software needs to use all of it, even if "unused RAM is wasted RAM".
Very nice article, I only learned about -mx32 after the Arch run failed, trying to figure out what the compiled support was supposed to be. It's perhaps a bit late, but I'd like to increase its use, determine when it offers the bigger performance gains @hailey found with mastadon etc.
So I went back and added a call to action:
Linux distros should include -mx32 support and daemons and utilities should be compiled down for more efficient idling and day to day computing! It only costs a flag.
every time i used x32 over the years it was really, really good performance gains, like 10-25%. I wonder if with the new AVX10 CPUs there could be an x32+ that'd get even more registers?
I searched a bit and found this thread: "I was shocked to see that a basic [GTK4] UI -- before anything actually intensive is done in the application -- already takes up 200MB of reserved memory", ha. (In practice measuring real memory use is kind of difficult, so take claims like that skeptically.)
I remember the discussions back in the day when amd64 showed up and how there were concerns about extra memory usage from the larger pointers. It's interesting that those were disregarded, probably as unimportant, because resources were ever-growing. Funny we've hit a limit (in terms of what one can buy), and agree that maybe it'd be nice to revisit those choices.
Why ... do you still have that cover image there if you know it makes people assume your blog post is slop? Do you like pre-emptively defending your cover image every time you discuss it?
This is historical content. There was a brief period of time when I did use AI to generate cover images and going back and changing it doesn't feel right. It was useful/OK back then to do so even if today it has become "harmful".
It's kind of my knee-jerk reaction to the takes from all people around here that assume "oh no, AI was in the vicinity of this so it's SlOp!". If you can't look past the use of AI and value content/code/whatever by what it is, well... too bad.
But actually, I have considered changing the images, but it's a non-trivial amount of work. And no, removing is not really an option; I need replacements.
If you just want to use 32-bit pointers for a subset of the process's heap allocations, like in a VM, you can do it without such an intrusive change to the system ABI. In fact, mmap on Linux makes it easy with MAP_32BIT and on other operating systems without that feature you can emulate it by trying MAP_FIXED with different base addresses and allocation sizes, ideally right after program startup, to carve out chunks from the 0..2^32 address range which are then sub-allocated.
With this approach you won't use the standard pointer types for the small pointer-at-rest representation, just like you wouldn't for 32-bit relative offsets or arena indices. You lose some provenance fidelity in the transition between the at-rest and in-flight forms, but that's pretty much exactly where provenance isn't as useful as a compiler optimization enabler, and your in-flight form (which is what lives in registers and on the stack and travels across function boundaries, including to FFI endpoints in something like a VM) can and should be the regular pointer type.
And, of course, if you control allocation yourself somehow, nothing stops you from storing pointers as 32-bit (or maybe even 16-bit!) indexes into a buffer. Translate it into a pointer for memory operations by adding it to the base of your buffer (maybe left-shifting first if you only need aligned offsets). It requires up-front thought about how you design your application, but if you've identified it as desirable, it's not difficult.
In Ninja, we used Node* as the unique identifier for files, so any given path was canonicalized into one of those and then we could use pointer comparisons instead of string comparisons for equailty. In my followup n2 experiment, I switched to using integer indexes within an array as ids for Rust ownership reasons, but also found that made it easy to then switch them to 32-bit integers.
This matters because these systems are basically large graph where each build step depends on a list of files, and with 32-bit file ids those lists can be arrays of 32-bit integers. At the time I found it wrote it made a few ms of difference on my builds and I didn't measure memory, but due to how the program was structured it was a trivial change so why not!
"Arena indices" in the post you're replying to was referring to exactly that; I assumed people would know what that was shorthand for and didn't elaborate. I was just pointing out that you can also easily do x32-style 32-bit absolute pointers within the standard 64-bit host ABI.
alexrp | 17 hours ago
Incidentally, Zig 0.17.0 just shipped full support for x32, which means no-libc w/ direct syscalls, or w/ glibc, or musl. We also test it natively in CI; Debian still supports
syscall.x32=yon the kernel command line. (We also shipped support for N32 which is the MIPS equivalent, but somehow I suspect people aren't going to be as interested in that one.)I did it mostly as a fun exercise and to help catch bad pointer size assumptions in CI, but it would be nice if it accidentally turned out to be actually useful.
[OP] veqq | 16 hours ago
Awesome! I'm currently testing crosscomping via
zig cc -target x86_64-linux-muslx32which would greatly ease the distribution process!david_chisnall | 13 hours ago
The RAM savings look very high, but I guess a language with that header structure and a lot of small objects would have these properties. In some of the analysis that we did for CHERI, we typically found that under 20% of pages contained one or more pointers. A lot of memory ends up being big buffers, images, and other data that doesn't contain pointers.
This is also a big part of the reason that x32 largely went away. One process using it may be smaller, but once you factor in the extra disk space and RAM from 32-bit and 64-bit shared libraries, you end up with a much less clear win. If everything used x32 then it would be a win probably but it turns out that JavaScript and some similar scripting languages like NaN boxing, so 32-bit isn't a win. Similarly, lightweight SFI things (e.g. WebAssembly) work nicely if you can carve out a 4 GiB chunk of the address space and just truncate pointers to that size (optionally with a base offset). So you find that things like Node.js and web browsers want a 64-bit system.
The same thing happens with JVMs, where it's nice to be able to map the heap in at least two places for generational GC. You also want to be able to have a large address range to put your Java objects in that doesn't overlap anything else, so a 2 GiB userspace address space may limit you to something on the order of a 300 MiB heap. And the pointer size increase doesn't matter for JVMs because 64-bit JVMs typically use a 32-bit integer to represent objects, which is an offset from the base of the heap right shifted by 3 or 4, so you can have a 64 GiB heap. And you can often put NIO buffers and large arrays in a separate heap region with only a small object header in the 64 GiB heap, which lets you actually scale to heaps that are a large fraction of a TiB, without any real pointer-size overhead.
evmar | 8 hours ago
I believe v8 (and consequently Node and browsers) doesn't use NaN boxing, and instead puts doubles on the heap: https://v8.dev/blog/pointer-compression . Ironically it's exactly to get the smaller pointer cost we're chasing with x32!
vpr | 20 hours ago
~ Corbet, on LWN on (yet another) proposed removal for x32
Maybe one of the positive externalities of all these chip shortages will a second wind for memory frugality. Just because memory sizes had been doubling doesn't mean our software needs to use all of it, even if "unused RAM is wasted RAM".
[OP] veqq | 19 hours ago
Very nice article, I only learned about -
mx32after the Arch run failed, trying to figure out what the compiled support was supposed to be. It's perhaps a bit late, but I'd like to increase its use, determine when it offers the bigger performance gains @hailey found with mastadon etc.So I went back and added a call to action:
jcelerier | 18 hours ago
every time i used x32 over the years it was really, really good performance gains, like 10-25%. I wonder if with the new AVX10 CPUs there could be an x32+ that'd get even more registers?
mort | 14 hours ago
jaredkrinke | 18 hours ago
Is the "hello world apps not fitting in a 32-bit address space" comment a joke? I assumed it was, but I'm actually afraid to ask.
vpr | 5 hours ago
A Java hello world process gives 4,049 MB virtual (VSZ) and 38MB resident, -Xmx64m gets us down to 2,589 MB.
Go 1.24, 1,197MB for its heap arena...
GHC's two-step allocator (since 8.0) reserves 1TB contiguous on 64-bit systems.
evmar | 8 hours ago
I searched a bit and found this thread: "I was shocked to see that a basic [GTK4] UI -- before anything actually intensive is done in the application -- already takes up 200MB of reserved memory", ha. (In practice measuring real memory use is kind of difficult, so take claims like that skeptically.)
jmmv | 18 hours ago
I wrote about this in some more detail two years ago here: https://blogsystem5.substack.com/p/x86-64-programming-models . (Yes, yes, "generated AI cover image that sucks, oh no!", but the content was all hand-written and researched.)
I remember the discussions back in the day when amd64 showed up and how there were concerns about extra memory usage from the larger pointers. It's interesting that those were disregarded, probably as unimportant, because resources were ever-growing. Funny we've hit a limit (in terms of what one can buy), and agree that maybe it'd be nice to revisit those choices.
mort | 14 hours ago
Why ... do you still have that cover image there if you know it makes people assume your blog post is slop? Do you like pre-emptively defending your cover image every time you discuss it?
jmmv | 5 hours ago
Because:
This is historical content. There was a brief period of time when I did use AI to generate cover images and going back and changing it doesn't feel right. It was useful/OK back then to do so even if today it has become "harmful".
It's kind of my knee-jerk reaction to the takes from all people around here that assume "oh no, AI was in the vicinity of this so it's SlOp!". If you can't look past the use of AI and value content/code/whatever by what it is, well... too bad.
But actually, I have considered changing the images, but it's a non-trivial amount of work. And no, removing is not really an option; I need replacements.
pervognsen | 12 hours ago
If you just want to use 32-bit pointers for a subset of the process's heap allocations, like in a VM, you can do it without such an intrusive change to the system ABI. In fact, mmap on Linux makes it easy with MAP_32BIT and on other operating systems without that feature you can emulate it by trying MAP_FIXED with different base addresses and allocation sizes, ideally right after program startup, to carve out chunks from the 0..2^32 address range which are then sub-allocated.
With this approach you won't use the standard pointer types for the small pointer-at-rest representation, just like you wouldn't for 32-bit relative offsets or arena indices. You lose some provenance fidelity in the transition between the at-rest and in-flight forms, but that's pretty much exactly where provenance isn't as useful as a compiler optimization enabler, and your in-flight form (which is what lives in registers and on the stack and travels across function boundaries, including to FFI endpoints in something like a VM) can and should be the regular pointer type.
mort | 7 hours ago
And, of course, if you control allocation yourself somehow, nothing stops you from storing pointers as 32-bit (or maybe even 16-bit!) indexes into a buffer. Translate it into a pointer for memory operations by adding it to the base of your buffer (maybe left-shifting first if you only need aligned offsets). It requires up-front thought about how you design your application, but if you've identified it as desirable, it's not difficult.
evmar | 7 hours ago
In Ninja, we used
Node*as the unique identifier for files, so any given path was canonicalized into one of those and then we could use pointer comparisons instead of string comparisons for equailty. In my followupn2experiment, I switched to using integer indexes within an array as ids for Rust ownership reasons, but also found that made it easy to then switch them to 32-bit integers.This matters because these systems are basically large graph where each build step depends on a list of files, and with 32-bit file ids those lists can be arrays of 32-bit integers. At the time I found it wrote it made a few ms of difference on my builds and I didn't measure memory, but due to how the program was structured it was a trivial change so why not!
pervognsen | 6 hours ago
"Arena indices" in the post you're replying to was referring to exactly that; I assumed people would know what that was shorthand for and didn't elaborate. I was just pointing out that you can also easily do x32-style 32-bit absolute pointers within the standard 64-bit host ABI.
mort | 6 hours ago
Ah, no that's on me for not reading thoroughly.
easrng | 9 hours ago
Is it not possible to compile for 32bit pointers and 64bit syscalls?