23/08/2026
The big players seem to be allowed to hack away at whatever they like — publicly, apparently. 😆
Meanwhile, I think we’ve been sold a slightly misleading idea about what it takes to run genuinely useful AI.
Decent-quality LLMs do not have to live in the cloud.
You don’t necessarily need a £10k GPU either. With modern MoE models, sensible model placement, enough RAM, fast NVMe storage and a decent last-gen CPU, things get considerably more interesting.
You can even start looking at huge 200B+ parameter MoE models, because you aren't necessarily running every parameter for every token. Add a relatively cheap GPU and some architectures can use it for additional acceleration rather than requiring the entire model to fit in VRAM.
It won't always be datacentre-fast. That's rather missing the point.
You own the machine.
You own the data.
You choose the models.
You choose what gets connected to the internet.
So I'm going down the rabbit hole properly:
Multiple local LLMs, MoE model placement, fast NVMe, lots of RAM, and old enterprise servers that everyone else decided were obsolete.
Because apparently a rack full of old Xeons needed another purpose. 😆
Join me on the quest.
The future of AI doesn't have to be in the cloud.