I’ve been running open-weight models locally for a while. GLM-5.2 on a 512GB Mac Studio was capable, but you always felt the ceiling: 2-bit quant, 5-9 tok/s, tool calls that got flaky in long agent sessions. Local was the thing you tolerated to keep data on your network. DeepSeek V4 Flash was the first real step forward — finally a local model with a genuine balance of performance and speed, one you could actually leave running under agent traffic without babysitting it. It proved the “fast MoE with real agentic chops” formula worked. 5.3-Flash on two DGX Sparks takes that further. 320B params but only 18B active, so it’s running at good speeds. 1M context. Thinking is always on but tunable (low/high/max), and at max it holds up in real agent work through Hermes: multi-step coding, sustained tool calls, no hand-holding. I’m still reaching for cloud models. But for the first time I can see a point coming where you won’t have to. The gap is closing faster than I expected — if you have the hardware, this is the one to try.