Posts / ai

The 162GB Model That's Making Everyone Nervous


I check the LocalLLaMA subreddit the way some people check the footy scores. Rationed, but regular. This week it delivered a proper spectacle: DeepSeek quietly updated their Flash model, benchmarks came out, and half the thread devolved into disbelief. A 162GB model apparently going toe to toe with GLM 5.2, which is nearly ten times its size. Someone posted a chart. Then someone else posted a bigger chart. By the third “holy shit” comment I stopped counting.

I’ve been half paying attention to the open-weight AI race for a couple of years now, mostly because it’s the one bit of the AI story that doesn’t make me want to lie down in a dark room. The proprietary labs, OpenAI, Anthropic, the lot, are locked in a spend war that looks increasingly like a bubble propping itself up with vibes and IPO ambitions. Meanwhile DeepSeek keeps turning up with something smaller, cheaper, and stubbornly competitive, and dropping the weights on Hugging Face like it’s no big deal.

What struck me in the thread wasn’t just the benchmark numbers, it was the reaction to them. There’s a real “here we go again” energy whenever DeepSeek ships. OpenAI cut Luna’s pricing by 80 percent literally the same week, and more than one commenter joined the dots: they saw this coming. Whether that’s true or just pattern-matching on a Tuesday, I don’t know. But the price war is real, and for once it’s a price war where the customer (or at least the bloke running a couple of old 3090s in his garage) actually wins.

That’s the part I find genuinely interesting, not the frontier chase, but the efficiency chase. One user in the thread made the point better than I could: DeepSeek’s compute per dollar is starting to look like the actual innovation, more than raw scale. A model that fits a million tokens of context into 6GB of VRAM is a different kind of achievement to “we spent another half a billion dollars on training.” I’ve spent enough of my career watching companies solve problems by throwing more hardware at them to have a soft spot for anyone solving it with better engineering instead.

None of this erases my unease about the bigger picture. Every new model, however efficient per-token, still sits on top of an industry burning through electricity and water at a rate nobody’s being fully honest about. Somebody in the thread cracked a joke about Elon building datacentres in space to dodge the pollution question, which is funny until you remember it’s not entirely a joke. I can be impressed by a 162GB model punching above its weight and still think the whole AI build-out needs a lot more scrutiny than it’s getting. Both things sit in my head at once and I’ve stopped trying to force them into agreement.

There’s also a running joke through the comments about DeepSeek models occasionally lapsing into Chinese mid-response, no matter how firmly you ask for English. Small thing, but it’s a decent reminder that “better than Opus on this benchmark” and “reliable tool you’d trust with real work” are not the same claim. Benchmarks are useful. They are also, as more than one person pointed out, easy to game and easy to misread. DeepSWE apparently measures something closer to “expands vague prompts well” than “is technically capable,” which is not nothing, but it’s not everything either.

I don’t do a lot of local hosting myself, my home setup runs Home Assistant and not much else that would strain a graphics card, but I like knowing the option exists. I like that a curious teenager, or a small business, or someone in regional Victoria without an enterprise contract, might soon run something genuinely capable on hardware they can actually buy. That’s a different world to the one we were in two years ago, when frontier capability meant a server farm and a contact at Nvidia.

Pro is apparently coming soon, and if the pattern holds it’ll be smaller than its rivals and cheaper again. I’ve got no idea if that’s DeepSeek playing a long game or just good engineering culture, and I’m not going to pretend I do. But I’ll be watching, with a coffee, mildly delighted that for once the underdog story in AI isn’t the same three American companies restating their own greatness.