The flash model has really good vibes. Feels like the best model in the 300B to 550B class. Better than DeepSeek v4.1 and GLM 5.3 Flash. Autocompacts really well, with perfect performance for a session with >3M tokens. Around 250 tok/s decode 2.6 ktok/s prefill.