The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them. Reddit
128gb of unified memory is somewhere north of four grand right now. A subscription that covers what I actually do is twenty a month. That's a payback period measured in decades and the machine will be a paperweight in five years.
I keep doing this arithmetic, keep getting the same answer, and keep wanting the answer to be different. So here's why I still think about it and someone can tell me which part is cope.
Privacy is real but it isn't four thousand dollars real for what I do. Offline is nice, I have wifi. The one that genuinely holds up is that a rented model can be changed underneath you. Quantised down quietly, rate limited when demand spikes, or retired outright. A local one can't. I've had a workflow break twice this year because something upstream moved and nobody told me.
The other thing keeping the idea alive is that the models are drifting toward a shape that suits this hardware. Low active param counts run at a speed that makes a big-memory slow-compute machine viable in a way a dense model of the same total size never would. ling-3.0-flash is the shape I mean, 124b with about 5b live per token, though it's api only so far (free until Aug 3.) so it isn't actually an argument for buying anything yet.
So: has anyone here bought the expensive box and been glad a year later? Not "it was fun". Glad. I want to hear from someone past the honeymoon.
