You know, I've always liked Futurama but I always kind of thought it was silly that literally everything has an AI and a personality.
But, you know, I actually think that there might be a logic to it. Economies of scale might mean that almost-literally every computer you buy in the year 3000 has some kind of AI-assistance chip in there, and sure maybe it will have full AI with a personality spitting out one-liners.
I'm curious how this would handle grammar checking on a basic word processor. Or maybe generate worlds for small text based games. I have no idea what the capabilities are of a cluster like this.
This is something I've been fantasizing about for long.
Let's say we took Rust, a language that makes parallelization easier than others (as it helps you avoid some common footguns). How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips? Let's say we picked many little Risc-V's. Surely this would be an interesting experiment (though I'm not sure whether it'd make economic sense or not...)
It would certainly not make any economic sense, and I guess that's also why noone is seriously looking into stuff like volunteer/enthusiast clusters of home computers to do inference in the same way that e.g. LHC@home works. The main bottleneck for LLMs is still memory bandwidth. Any memory bus not directly soldered on your GPU is terribly slow. That's why one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each. And it's also not like you can just solder more memory onto a chip. At modern speeds, the speed of light is a hard limit. For current GDDR7, signals may only travel like 10mm per cycle.
If you spread such a system out over dozens or hundreds of tiny chips, you'll be wasting most of its resources and lose hard to anyone who built a single chip setup.
Look at Xmos, founded by a transputers-dad. Small uCs, that can be connected together into a massive cluster, while being first-citizen of their xC language.
Mojo might be what you want, especially with "MAX" which is their AI modeling framework, where in other languages you need NVidia's libraries, or AMDs, etc the Max libraries just let you talk directly to the GPU / CPU / ASIC with Mojo. I think Mojo is very underrated in this space right now, but assuming they don't mess it up, it could be a major contender in AI. In theory, if someone releases a board, and Modular (company that 'owns' Mojo) adds it to Max, you're basically in the green to experiment as much as you want.
Seriously though, what are the low cost chips that can usefully run LLMs? Is a Mac Mini the lowest we can go? Are there iGPUs on mini-itx that can do it, or are there dedicated AI chips that one could turn into a pi HAT?
Depends on what you consider to "usefully run LLMs".
Earlier this year, I bought a mini pc from Aliexpress, specs are roughly Ryzen H255, 24GB LPDDR5, 1TB SSD. This was around 350€ including VAT, customs, shipping etc. I would personally consider this somewhat of a lowest class of useful LLM box. It can run 8B models well, up to somewhere around 24B. I currently run Gemma 4 26B A4B Q5 on it, with MTP, and it is quite slow, but smaller models would run okay on it.
An Orange Pi 5 Max does this job for real. It's an RK3588 board — $75 for 4GB, $95 for 8GB on AliExpress. A community test got Qwen2.5-0.5B at about 12 tok/s on the CPU via llama.cpp. The chip also has a 6 INT8 TOPS NPU if you'd rather go the RKNN route. Won't beat a Mac Mini, but it's an actual computer for under a hundred bucks.
I regret to inform you that prices have long departed the lower atmosphere. Even on AliExpress, you're looking at a few hundred bucks for those. Still less than a Mac Mini, but less less.
cameron_b | 21 hours ago
sjakati98 | 20 hours ago
nkozyra | 17 hours ago
pantalaimon | 12 hours ago
[OP] nkko | 11 hours ago
tdhz77 | 19 hours ago
tombert | 17 hours ago
But, you know, I actually think that there might be a logic to it. Economies of scale might mean that almost-literally every computer you buy in the year 3000 has some kind of AI-assistance chip in there, and sure maybe it will have full AI with a personality spitting out one-liners.
abroadwin | 16 hours ago
KeplerBoy | 15 hours ago
oneZergArmy | 15 hours ago
Tade0 | 15 hours ago
_joel | 10 hours ago
jagged-chisel | 9 hours ago
matthewfcarlson | 18 hours ago
NDlurker | 18 hours ago
nonasking_ | 17 hours ago
librasteve | 15 hours ago
don’t get too excited until we get the TinyGo backend built though ;-)
ladyanita22 | 15 hours ago
Let's say we took Rust, a language that makes parallelization easier than others (as it helps you avoid some common footguns). How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips? Let's say we picked many little Risc-V's. Surely this would be an interesting experiment (though I'm not sure whether it'd make economic sense or not...)
sigmoid10 | 15 hours ago
If you spread such a system out over dozens or hundreds of tiny chips, you'll be wasting most of its resources and lose hard to anyone who built a single chip setup.
ur-whale | 13 hours ago
It's scaling the communication that becomes hard.
In this project they daisy-chain SPI. I don't believe that would scale very far.
akavel | 12 hours ago
alfanick | 8 hours ago
tyingq | 8 hours ago
https://milkv.io/cluster-08
godojo | 8 hours ago
giancarlostoro | 5 hours ago
https://max.modular.com/
sneak | 14 hours ago
Risse | 13 hours ago
Earlier this year, I bought a mini pc from Aliexpress, specs are roughly Ryzen H255, 24GB LPDDR5, 1TB SSD. This was around 350€ including VAT, customs, shipping etc. I would personally consider this somewhat of a lowest class of useful LLM box. It can run 8B models well, up to somewhere around 24B. I currently run Gemma 4 26B A4B Q5 on it, with MTP, and it is quite slow, but smaller models would run okay on it.
bahmboo | 10 hours ago
chorylee | 9 hours ago
cameron_b | 6 hours ago
AmazingTurtle | 9 hours ago