
DeepSeek shipped V4.1-Flash on September 10, 2026, and the pitch is blunt: frontier-class agentic performance, native image understanding, and a price table that makes flagship models look like a luxury good. The model is live on the DeepSeek API as deepseek-flash, the open-source weights and a tech report are already on Hugging Face, and it is displacing the company's own flagship — literally. Starting September 14, every DeepSeek-V4-Pro request gets silently routed to V4.1-Flash.
V4.1-Flash is not another checkpoint of the V4 line. It is the smallest member of an entirely new architecture family: a 552B-parameter Mixture-of-Experts built on what DeepSeek calls a Causal Encoder–Decoder design, with an unusual asymmetric split — about 8B active parameters for input processing and 16B for output generation.
Two things make this release more than a spec bump:

Image: DeepSeek's official agentic benchmark results, September 10, 2026.
On DeepSeek's published table, V4.1-Flash doesn't just beat its Pro sibling — it beats several frontier flagships at the tasks agents actually perform:
| Benchmark | V4.1-Flash | V4-Pro (0813) | Claude Opus 5 | GPT5.6-Sol |
|---|---|---|---|---|
| DeepSWE v1.1 | 74.2 | 62.7 | 74.0 | 73.0 |
| CyberGym | 88.1 | 83.3 | – | 84.5 |
| Automation-Bench | 54.8 | 43.2 | 50.3 | 45.8 |
| Terminal-Bench 3.0 | 30.0 | 11.8 | 43.3 | 34.4 |
| Codeforces (Rating) | 3471 | 3348 | – | – |
The honest caveat: pure knowledge benchmarks still belong to the incumbents. GPQA Diamond sits at 90.9 (vs. GPT5.6-Sol's 94.1) and HLE at 36.8 against Opus 5's 56.3. V4.1-Flash competes where code, terminals, and multi-step agent work live — and mostly wins there.

Image: DeepSeek's full benchmark comparison table.

Image: DeepSeek-V4.1-Flash API pricing, effective 04:00 UTC, September 10, 2026.
Off-peak, you pay $0.15 per million input tokens and $0.60 per million output tokens — and "off-peak" covers everything except 01:00–04:00 and 06:00–10:00 UTC on weekdays. Weekends and public holidays are entirely off-peak. Cache hits collapse input costs to $0.003 per million tokens, which is the number that matters for agent workloads that re-read the same context all day.
The community has already run its own math: a popular r/singularity thread reports V4.1-Flash reaching 98% of GPT-6 "Astra"'s score on OpenDesign Arena at 1.4% of the cost.

Image: DeepSeek's KV cache per token across four model generations.
Serving cost is the quiet moat here. DeepSeek has compressed per-token KV cache from 389,120 bytes (V1, 2023) to 890 bytes, and V4.1-Flash's architecture extends that to 1/4 the HBM and 1/8 the SSD of its predecessor. Cheap cache is what makes a $0.003 cache-hit rate economically possible rather than promotional.
The boldest part of the announcement barely got a paragraph: from 04:00 UTC on September 14, all V4-Pro requests route to V4.1-Flash at V4.1-Flash rates until a V4.1-Pro launches. The old V4-Flash and V4-Flash-Vision-Exp model names already redirect. DeepSeek is not keeping a premium tier alive out of nostalgia — the flash tier is now the strategy.
The Hacker News thread on the V4-Flash line captures both sides. Fans point out that "DSV4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects." Skeptics counter that cheap tokens aren't the same as efficient ones — "you can blast DeepSeek for 1 hour solving a hard task, or Opus for 5 minutes" — and long-context stress tests of the V4 line's 1M window show quality degradation on very large repositories. Both things can be true: V4.1-Flash is spectacular value per token, and your agent's token budget will decide whether that value is real.
The model is served as deepseek-flash on the DeepSeek API (old model names route to it during the transition). Open-source weights and the tech report are on Hugging Face, and partners CodeBuddy (via WorkBuddy) and OpenCode shipped day-one support.
Carefully selected AI tools to improve your work, study, and live efficiency.
A major breakthrough has been achieved in the core architecture of large-scale models! The release of Kimi Linear marks the first time that linear attention technology has comprehensively surpassed and significantly outperformed the traditional Transformer full-attention model in both performance and efficiency. This "win-win" achievement is expected to significantly reduce the computational barriers and costs for long text processing, complex reasoning, and AI agent applications, potentially changing the competitive landscape of underlying technologies for large-scale models.
Over the past week, the AI community's attention has been drawn to a mysterious model that quietly emerged on the OpenRouter platform—Polaris Alpha. As a direct continuation of yesterday's discussion of the GPT-5.1 leak, this suddenly appearing model brings more technical details and strategic signals worthy of in-depth exploration.
A new paradigm in knowledge acquisition has arrived, this time powered by AI.
Standing at this moment in 2025, when we look back at the development journey of artificial intelligence, we witness how this revolutionary technology has reshaped every aspect of human society. From initial theoretical concepts to today's practical applications, each step forward in AI technology has changed the way we live. Let's revisit this fascinating journey together.
Sponsored by Lipsync Studio