124B total but only ~5B active-this is exactly the shape I want for my box. Are low-active MoEs just the local sweet spot now?

I keep getting more excited about the active param count than the total lately. Something like Ling-3.0-flash is 124B but only ~5.1B active per token, and that's the shape that actually runs decently on the weird bandwidth-limited hardware a lot of us have-unified-memory Macs, Strix Halo, DGX Spark, big-RAM CPU boxes.

(It's API-only for now so this is me daydreaming about if/when weights show up, but still.) Someone in the thread just said "nice one for strix halo" and, yeah, that.

For people running low-active MoEs today (Qwen's a3b ones etc.)-where's the sweet spot for you on active vs total? When does "tiny active params" start feeling too thin next to a dense 30B, and does prompt processing become the real bottleneck instead of generation? Trying to work out if this is the direction or just nice on paper.