Back to Blog List

DeepSeek V4.1-Flash Launches: Agentic Benchmarks at Flash Prices

9/10/2026
Author: Q Yang
Category: AI
DeepSeek V4.1-Flash Launches: Agentic Benchmarks at Flash Prices

DeepSeek shipped V4.1-Flash on September 10, 2026, and the pitch is blunt: frontier-class agentic performance, native image understanding, and a price table that makes flagship models look like a luxury good. The model is live on the DeepSeek API as deepseek-flash, the open-source weights and a tech report are already on Hugging Face, and it is displacing the company's own flagship — literally. Starting September 14, every DeepSeek-V4-Pro request gets silently routed to V4.1-Flash.

A new architecture, not a refresh

V4.1-Flash is not another checkpoint of the V4 line. It is the smallest member of an entirely new architecture family: a 552B-parameter Mixture-of-Experts built on what DeepSeek calls a Causal Encoder–Decoder design, with an unusual asymmetric split — about 8B active parameters for input processing and 16B for output generation.

Two things make this release more than a spec bump:

  • Native visual understanding. The interim V4-Flash-Vision-Exp had been testing image input since August 21; V4.1-Flash makes multimodality a first-class capability rather than a bolt-on.
  • A radical KV cache diet. The model stores just 890 bytes of KV cache per token — 3.9x less than V4-Flash and roughly 437x less than the original DeepSeek-V1 — cutting HBM usage to 1/4 and SSD storage to 1/8 of the previous generation.

DeepSeek-V4.1-Flash leads three of four agentic benchmarks against Kimi-K3, GLM-5.3, Claude Opus 5 and GPT5.6-Sol

Image: DeepSeek's official agentic benchmark results, September 10, 2026.

Benchmarks: it leads where agents live

On DeepSeek's published table, V4.1-Flash doesn't just beat its Pro sibling — it beats several frontier flagships at the tasks agents actually perform:

BenchmarkV4.1-FlashV4-Pro (0813)Claude Opus 5GPT5.6-Sol
DeepSWE v1.174.262.774.073.0
CyberGym88.183.384.5
Automation-Bench54.843.250.345.8
Terminal-Bench 3.030.011.843.334.4
Codeforces (Rating)34713348

The honest caveat: pure knowledge benchmarks still belong to the incumbents. GPQA Diamond sits at 90.9 (vs. GPT5.6-Sol's 94.1) and HLE at 36.8 against Opus 5's 56.3. V4.1-Flash competes where code, terminals, and multi-step agent work live — and mostly wins there.

Full benchmark table comparing DeepSeek-V4.1-Flash with V4-Pro, V4-Flash, GLM-5.3, Kimi K3, GPT5.6-Sol and Claude Opus 5

Image: DeepSeek's full benchmark comparison table.

The pricing table is the real headline

Official DeepSeek-V4.1-Flash API pricing table showing off-peak and peak rates per million tokens

Image: DeepSeek-V4.1-Flash API pricing, effective 04:00 UTC, September 10, 2026.

Off-peak, you pay $0.15 per million input tokens and $0.60 per million output tokens — and "off-peak" covers everything except 01:00–04:00 and 06:00–10:00 UTC on weekdays. Weekends and public holidays are entirely off-peak. Cache hits collapse input costs to $0.003 per million tokens, which is the number that matters for agent workloads that re-read the same context all day.

The community has already run its own math: a popular r/singularity thread reports V4.1-Flash reaching 98% of GPT-6 "Astra"'s score on OpenDesign Arena at 1.4% of the cost.

Why it can afford this: a four-generation cache diet

Global KV cache per token across DeepSeek generations, from 389,120 bytes in V1 to 890 bytes in V4.1-Flash

Image: DeepSeek's KV cache per token across four model generations.

Serving cost is the quiet moat here. DeepSeek has compressed per-token KV cache from 389,120 bytes (V1, 2023) to 890 bytes, and V4.1-Flash's architecture extends that to 1/4 the HBM and 1/8 the SSD of its predecessor. Cheap cache is what makes a $0.003 cache-hit rate economically possible rather than promotional.

V4-Pro is being retired, quietly

The boldest part of the announcement barely got a paragraph: from 04:00 UTC on September 14, all V4-Pro requests route to V4.1-Flash at V4.1-Flash rates until a V4.1-Pro launches. The old V4-Flash and V4-Flash-Vision-Exp model names already redirect. DeepSeek is not keeping a premium tier alive out of nostalgia — the flash tier is now the strategy.

Community reaction: impressed, with caveats

The Hacker News thread on the V4-Flash line captures both sides. Fans point out that "DSV4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects." Skeptics counter that cheap tokens aren't the same as efficient ones — "you can blast DeepSeek for 1 hour solving a hard task, or Opus for 5 minutes" — and long-context stress tests of the V4 line's 1M window show quality degradation on very large repositories. Both things can be true: V4.1-Flash is spectacular value per token, and your agent's token budget will decide whether that value is real.

How to try it

The model is served as deepseek-flash on the DeepSeek API (old model names route to it during the transition). Open-source weights and the tech report are on Hugging Face, and partners CodeBuddy (via WorkBuddy) and OpenCode shipped day-one support.

Resources

  1. Official announcement: https://api-docs.deepseek.com/news/news260910
  2. Weights & tech report: https://huggingface.co/deepseek-ai
  3. r/LocalLLaMA discussion: https://www.reddit.com/r/LocalLLaMA/comments/1wan3nl/deepseek_flash_41_is_already_being_tested_via_api/
  4. OpenDesign Arena cost-efficiency thread: https://www.reddit.com/r/singularity/comments/1wbmy11/deepseek_v41_flash_reaches_98_of_astras_score_at/
Share this article

Leave your comment

  • No comments yet.
Ad
Ad not loaded or not displayed

Recommended AI Tools

Carefully selected AI tools to improve your work, study, and live efficiency.

 Lipsync Studio

Transform your videos with advanced lip sync technology.

61.2K
SPONSORED
Image to Image AI

AI-powered image transformation for professional creative workflows.

SPONSORED
SAM TTS

Experience the nostalgic Microsoft SAM voice from Windows XP in your browser.

23.2K
SPONSORED
OpenArt

OpenArt is a versatile AI image and video generator.

SPONSORED
Grayscale Image

Grayscale Image is a free online tool for converting color photos to black and white with professional controls.

SPONSORED
Virtual Try On

AI-powered virtual try-on for clothes, hairstyles, and accessories.

SPONSORED
Circle Crop Image

Circle Crop Image is a free online tool for creating round images.

SPONSORED

Related Articles

Kimi Linear emerges: revolutionizing the attention architecture of Transformer, boosting long text processing efficiency by 6 times.
News
10/31/2025
Kimi Linear emerges: revolutionizing the attention architecture of Transformer, boosting long text processing efficiency by 6 times.
Author: Kimi Lv

A major breakthrough has been achieved in the core architecture of large-scale models! The release of Kimi Linear marks the first time that linear attention technology has comprehensively surpassed and significantly outperformed the traditional Transformer full-attention model in both performance and efficiency. This "win-win" achievement is expected to significantly reduce the computational barriers and costs for long text processing, complex reasoning, and AI agent applications, potentially changing the competitive landscape of underlying technologies for large-scale models.

In-depth analysis of OpenAI Polaris Alpha technology: A key sequel to the GPT-5.1 leak incident
News
11/12/2025
In-depth analysis of OpenAI Polaris Alpha technology: A key sequel to the GPT-5.1 leak incident
Author: Lydia

Over the past week, the AI ​​community's attention has been drawn to a mysterious model that quietly emerged on the OpenRouter platform—Polaris Alpha. As a direct continuation of yesterday's discussion of the GPT-5.1 leak, this suddenly appearing model brings more technical details and strategic signals worthy of in-depth exploration.

Grokipedia - xAI Launches New AI Knowledge Platform to Challenge Traditional Encyclopedias with AI Revolution
AI
10/28/2025
Grokipedia - xAI Launches New AI Knowledge Platform to Challenge Traditional Encyclopedias with AI Revolution
Author: Lucas

A new paradigm in knowledge acquisition has arrived, this time powered by AI.

2025, looking at the evolution of artificial intelligence
AI
4/24/2025
2025, looking at the evolution of artificial intelligence
Author: Q Yang

Standing at this moment in 2025, when we look back at the development journey of artificial intelligence, we witness how this revolutionary technology has reshaped every aspect of human society. From initial theoretical concepts to today's practical applications, each step forward in AI technology has changed the way we live. Let's revisit this fascinating journey together.

Most Popular AI Tools

Base44

Base44 is an AI-powered platform for building fully-functional apps with no code required.

105.8K
Magic Patterns

Magic Patterns is an AI design tool for product teams.

Midjourney API by PiAPI
5% offCode:AIWITHME

Transform text into stunning images with Midjourney API.

FLUX API - PiAPI
5% offCode:AIWITHME

FLUX API by PiAPI offers advanced image generation capabilities.

Typeless

Speak naturally, and Typeless will turn your words into polished messages, emails, and documents that read like you carefully typed them.

627.7K
LogoAi
30% offCode:aiwithme

Create a stunning logo effortlessly with LogoAi.

Pollo AI

Pollo AI is a versatile AI image and video generator.

Klap
30% offCode:AIWITHME

Klap transforms long videos into engaging shorts effortlessly.

458.4K