Suggested searches

GPT Image Introduction GPT Image Free Guides GPT Image Developer Docs Midjourney Introduction Midjourney Free Guides Midjourney Developer Docs Google Nano Banana Introduction Google Nano Banana Free Guides Google Nano Banana Developer Docs Adobe Firefly Image Introduction Adobe Firefly Image Free Guides Adobe Firefly Image Developer Docs FLUX Introduction FLUX Free Guides FLUX Developer Docs Ideogram Introduction Ideogram Free Guides Ideogram Developer Docs Recraft Introduction Recraft Free Guides Recraft Developer Docs Stable Diffusion Introduction ByteDance Seedream Introduction Grok Imagine Image Introduction Google Veo Introduction Google Veo Free Guides Google Veo Developer Docs Runway Introduction Kling AI Introduction ByteDance Seedance Introduction ByteDance Seedance Free Guides ByteDance Seedance Developer Docs Luma AI Introduction Adobe Firefly Video Introduction Adobe Firefly Video Free Guides Adobe Firefly Video Developer Docs Hailuo AI Introduction PixVerse Introduction PixVerse Free Guides PixVerse Developer Docs Pika Introduction Pika Free Guides Pika Developer Docs Alibaba Wan Introduction LTX Video Introduction Grok Imagine Video Introduction Google Gemini Introduction Google Gemini Free Guides Google Gemini Developer Docs OpenAI GPT Introduction Claude Fable Introduction Claude Fable Free Guides Claude Fable Developer Docs DeepSeek Introduction Qwen Introduction Llama Introduction Codex Introduction Codex Free Guides Codex Developer Docs Cursor Introduction Cursor Free Guides Cursor Developer Docs OpenClaw Introduction OpenClaw Free Guides OpenClaw Developer Docs Perplexity Introduction ElevenLabs Introduction Suno Introduction Manus Introduction Claude vs ChatGPT: Compare Them Without Borrowed Numbers Claude vs Gemini: Comparing Two Assistants Honestly Is Adobe Firefly Free? The Free Membership, the Credits and the First Year Adobe Firefly Pricing: Plans, Generative Credits and the Cost of One Image Adobe Firefly Video Cost: Credits per Second and per Clip Is Adobe Firefly Video Free? Where the Free Plan Stops FLUX.2 Prompt Guide: The Structure Black Forest Labs Documents FLUX.2 vs Nano Banana Pro: What Each Vendor Actually Publishes Ideogram 4.0 Prompt Guide: The Documented Structure and How Text Renders Ideogram Pricing: Plans, Credits and the Free Limits Midjourney Prompt Guide: Structure, Elements and Edit Instructions Midjourney Pricing: Four Plans, GPU Hours and What a Job Costs Nano Banana Pro Pricing: What Google Publishes and What It Does Not Nano Banana Pro vs Midjourney: Which One Fits Your Workflow Pika Credits: What One Pika 2.5 Video Costs Is Pika Free? What the $0 Plan Actually Includes What PixVerse Credits Cost, and Which Ledger You Are Paying Is PixVerse Free? Two Free Tiers, and What Each One Costs You Recraft V4.1: The Eight Variants and Which to Pick Recraft Pricing and the Free Plan: What the Vendor Publishes Seedance Official Website: Which Entrances Are First-Party Seedance 2.5 vs 2.0: The Parameters ByteDance Publishes Veo 3.1 vs Sora 2: What Each Vendor Still Confirms Veo 3.1 vs Kling 3.0: What Each Vendor Publishes Is Claude AI Free? What the $0 Plan Actually Gives You Claude Pricing Plans: Pro vs Max 5x vs Max 20x How to Use Adobe Firefly (Adobe Firefly Image 5) Adobe Firefly API: Credentials, Endpoints and the Image5 Schema How to Use Adobe Firefly Video Adobe Firefly Video API: /v3/videos/generate Explained Designing UI Mockups and App Screens with GPT Image 2.5 Sticker Packs and Transparent Emoji with GPT Image 2.5 Nano Banana Pro Prompt Guide: The Official Frameworks Is Nano Banana Pro Free? What Google Actually Confirms How to Use Pika 2.5 Pika API: the official developer site, billing and endpoints How to Use PixVerse PixVerse API: Platform, Endpoints and Credits How to Use ByteDance Seedance 2.5 Seedance API: The Real Endpoint, Model ID and Fields Veo 3.1 Prompt Guide: The Seven Elements Google Names Veo 3.1 Price and Free Access: The Official Numbers How to Use Claude AI Claude API: Getting Started How to Use FLUX AI: A FLUX.2 Getting-Started Guide FLUX.2 API: Getting Started with Black Forest Labs GPT Image 2.5 vs DALL·E 3 Making Infographics and Diagrams with GPT Image 2.5 How to Use Ideogram Ideogram API: Access, Endpoints and Text Rendering How to Use Midjourney V8.2 Midjourney API: What Exists and What Does Not How to Use Nano Banana Pro Nano Banana Pro API: Getting Started How to Use Recraft Recraft API: Access, Endpoints and Style Consistency How to Use Veo 3.1 Veo 3.1 API Pricing and Vertex AI Access GPT Image 2.5 Prompt Sharing GPT Image 2.5 vs Nano Banana 2 GPT Image 2.5 vs Seedream 5.0 Pro GPT Image 2.5 vs FLUX 2 GPT Image 2.5 vs Ideogram Is GPT Image 2.5 Free GPT Image 2.5 Sketch GPT Image 2.5 Templates GPT Image 2.5 Comment Editing GPT Image 2.5 Character Consistency GPT Image 2.5 Combine Images GPT Image 2.5 Text Rendering What Is GPT Image 2.5 GPT Image 2.5 API Overview Codex vs. ChatGPT: Which Should You Use? Codex app, CLI, IDE, or cloud: how to choose the right surface Your First Low-Risk Coding Task with Codex: A Safe Walkthrough What Is Cursor? Its AI Coding Workflow Explained Cursor vs. VS Code: Which Editor Fits Your Workflow? Cursor Features Explained: Agent, Tab, Context, and More What Is Gemini? Apps, Models, AI Studio, and API Explained Gemini Apps vs. Gemini API: How to Choose the Right Tool for the Job What Can Gemini Do? A Practical Capability Guide OpenClaw Foundation Explained: Governance and Independence OpenClaw Skill Workshop Guide: Review Reusable Workflows OpenClaw Skill Cards: Read ClawHub Security Scans OpenClaw 2.0 Guide: New Features and Upgrade Checks OpenClaw LTS Guide: Choosing extended-stable or stable Install OpenClaw: Desktop, Script, npm, and Source Options OpenClaw Node.js Setup: Versions, Installation, and PATH How to Write Better Codex Prompts: A Practical Framework How to Review Codex Code Changes Before You Commit Cursor Beginner Tutorial: From Install to First Reviewed Edit Cursor Rules Tutorial: Project Rules, User Rules, and AGENTS.md Install Cursor on Windows and Configure a Chinese Interface Cursor MCP Tutorial: Configure, Verify, and Secure MCP Servers Gemini Prompt Guide: Better Instructions and Templates Gemini API Quickstart: Key, Python SDK and First Call Gemini Web App Guide: Login, Files, Chats and Privacy Gemini API Key Security: Storage, Restrictions and Rotation What Is Codex? Capabilities, Limits, and Ways to Use It Codex Beginner Tutorial: Complete Your First Safe Task Install Codex CLI: Sign In and Run Your First Safe Task Codex AGENTS.md Guide: Layered Rules and Validation Codex CLI Commands: Sessions, Review, and Automation How to Use Gemini: Web, Android & iPhone Setup Gemini Features Guide: Chat, Files, Images & Live How to Chat with Gemini: Prompts, Follow-Ups & Live Gemini AI Image Generator Guide: Prompts & Editing Gemini vs GPT-4: Features, Limits & Which to Use Gemini AI Assistant Guide: Mobile, Apps & Privacy Gemini Prompt Engineering Guide: Patterns & Examples Gemini Chat API Guide: Multi-Turn Prompts in Python Gemini System Instructions: API Guide & Examples Gemini Context Caching Guide: Cost, Latency & API
AI Tool Blog DeepSeek V4.1 Flash learning hub

Read the KV cache numbers behind DeepSeek V4.1 Flash.

DeepSeek V4.1 Flash is DeepSeek’s current model, and its technical report is built around a single number: the global KV cache it needs per token. At 890 bytes that is roughly a quarter of V4-Flash and 437 times smaller than the first DeepSeek model, which is what makes million-token agent workloads affordable to serve.

Specifications, benchmarks and architecture notes on this page come from DeepSeek’s own model card and technical report for V4.1 Flash. The weights are MIT licensed, but DeepSeek publishes no per-token rate on the card.

Official figures

The two charts DeepSeek publishes for V4.1 Flash

The model card carries two figures: an agentic benchmark comparison and a chart of how far the KV cache has fallen across four model generations. Everywhere else DeepSeek draws its numbers in JavaScript, so these are the only static figures the vendor publishes.

DeepSeek’s published V4.1 Flash agentic benchmark chart against Kimi-K3, GLM-5.3, Opus 5 and GPT-5.6 Sol
Agentic benchmarks: DeepSeek V4.1 Flash against Kimi-K3, GLM-5.3, Opus 5 and GPT-5.6 Sol. V4.1 Flash leads DeepSWE v1.1 at 74.2, CyberGym at 88.1 and Automation-Bench at 54.8, and trails on Terminal-Bench 3.0, where it scores 30.0 against Opus 5’s 43.3.
DeepSeek’s published chart of global KV cache bytes per token, falling from V1 to V4.1 Flash
Global KV cache per token, in bytes, plotted across four generations: DeepSeek-V1 at 389,120 in November 2023, V3.2 at 48,068 in December 2025, V4-Flash at 3,514 in April 2026 and V4.1 Flash at 890 in September 2026 — an 8.1×, then a 13.7×, then a 3.9× reduction, each measured against the generation before it.

Both figures belong to DeepSeek and are reproduced with credit: DeepSeek V4.1 Flash model card

What it is

How DeepSeek describes V4.1 Flash

DeepSeek titles its technical report “Pushing the Limits of KV Cache Compression”, and the numbers back the title. This is a model sold on the cost of holding a long context, not on winning every benchmark.

01

The headline is memory, not raw capability

890 bytes of global KV cache per token, about a quarter of V4-Flash’s 3,514 and 437 times less than V1’s 389,120. For input-heavy agentic workloads served at long context, that is the difference between one node and a rack.

02

An encoder-decoder that shares one cache

V4.1 Flash uses a 40-layer Causal Encoder-Decoder: 20 encoder layers followed by 20 decoder layers, where the decoder’s global KV cache is projected from the final encoder hidden states instead of being derived layer by layer. DeepSeek says that is why only 8B parameters are activated per token during prefill, and 16B during decode.

03

Sparse attention with three static modes

Compressed Sparse Attention 2 gives each attention layer one of three modes — Full, Reindex or Reuse — so main KV and indexer K are shared across layers and Top-K indices are reused. A hierarchical sparse indexer then bounds deeper indexer cost independently of context length, and FP4 main KV caching stores one E4M3 scale per 16 channels.

04

Native multimodality from pre-training

A DeepSeek-ViT encoder trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle feeds a two-layer MLP projector, and images are processed jointly with text from the start of language-model pre-training. The corpus is 45T tokens; sparse attention was trained at 64K and the context extended to 1M at the 34T mark.

05

Reasoning effort is a dial, not three settings

V4.1 Flash takes a reasoning_effort value from 1 to 100 — continuous, rather than the usual low, medium and high. DeepSeek built the model around that trade: its post-training changes are almost entirely data-pipeline work, automating the synthesis of agent tasks and environments, rather than algorithmic ones.

Official benchmarks

What the model card reports

The vendor’s own table at maximum reasoning effort. The last column names the best rival score in the same table, so the rows where V4.1 Flash loses are visible beside the ones where it wins.

Benchmark DeepSeek V4.1 Flash Best rival in the same table
Terminal-Bench 2.1 90.6 Opus-5.0 — 89.1
Terminal-Bench 3.0 30.0 Opus-5.0 — 43.3
DeepSWE v1.1 74.2 Opus-5.0 — 74.0
NL2Repo-Bench 64.0 Opus-5.0 — 75.3
CyberGym 88.1 GPT-5.6 Sol — 84.5
HLE with tools 63.9 Opus-5.0 — 63.6
Automation-Bench 54.8 Opus-5.0 — 50.3
GPQA Diamond 90.9 GPT-5.6 Sol — 94.1
HLE 36.8 Opus-5.0 — 56.3
Codeforces (rating) 3471 DeepSeek-V4-Pro — 3348

DeepSeek states that scores within 0.3 of each other are treated as equivalent in its own base-model evaluations. The card also notes that agentic results depend heavily on the harness: the same model scores 74.2 on DeepSWE v1.1 with mini-SWE and 65.6 with Codex.

Specifications

Documented specifications

Every row below is stated on DeepSeek’s model card or in its technical report.

Developer
DeepSeek
Model
DeepSeek V4.1 Flash
Architecture
Causal Encoder-Decoder, 20 + 20 layers
Backbone parameters
552B
Activated parameters
8B prefill · 16B decode
Experts
1 shared + 384 routed, 6 routed per token
Context
Up to 1,000,000 tokens
Pre-training corpus
45T multimodal tokens
Vision encoder
DeepSeek-ViT, trained from scratch
KV cache
890 bytes per token
Reasoning effort
Continuous, 1 to 100
Licence
MIT
How to use it

Documented access channels

DeepSeek ships this one as downloadable weights with a technical report attached, and leaves hosting to you or to a provider. That makes the licence the first thing to read, not the last.

Official model card

The architecture notes, both figures, the full evaluation tables and the recommended sampling parameters.

Open

Technical report

The PDF the model card links to, describing the compression methods in full.

Open

Weights and licence

Both the repository and the model weights are MIT licensed — no user threshold in the licence text, and no field-of-use restriction beyond applicable law.

deepseek-recipe

A set of Rust libraries with Python bindings that encode Messages, Chat Completions and Responses API requests into V4 and V4.1 prompts, and parse streamed output back.

Open
Common questions

DeepSeek V4.1 Flash questions

Is DeepSeek V4.1 Flash really multimodal?

Yes. It takes text and images and returns text. Images go through a DeepSeek-ViT encoder trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling and then a two-layer MLP projector, and the model card reports DocVQA at 95.6, RefCOCO-avg at 86.0 and MMMU-Pro at 56.5.

What does 890 bytes per token mean in practice?

It is the global KV cache the model must keep per token of context, measured in bytes. The same figure was 3,514 bytes for V4-Flash and 389,120 for the original DeepSeek-V1, so the serving arithmetic for a million-token request changes by orders of magnitude rather than by a few percent.

Can I run it locally?

The repository, its inference folder and its evaluation folder assume serious hardware: a 552B backbone with 8B parameters activated during prefill and 16B during decode. DeepSeek documents weight conversion and the recommended sampling parameters, but it publishes no minimum hardware figure, so this page does not invent one.

What licence do the weights carry?

MIT. Both the repository and the model weights are MIT licensed, which is unusual for a model of this class and means there is no revenue threshold or user-count clause in the licence itself.

How much does the API cost?

DeepSeek’s model card does not publish per-token pricing, and this page will not guess at it. Rate cards for the hosted service live on DeepSeek’s own platform pages. What the card does publish is the recommended sampling setup: temperature 1.0, top_p 0.95 or 1.0, a 1M-token context window and at least 256K output tokens.

What is reasoning_effort, and why 1 to 100?

It is a continuous dial rather than three named levels: any integer from 1 to 100 trades inference cost for accuracy. Every instruct score in the table above was measured at 100, the maximum setting.

Why extend the context to 1M rather than train at it?

On the vendor’s own account, sparse attention was trained at a 64K sequence length and the context was extended to 1M after 34T of the 45T training tokens. That schedule is also why the compression work matters: a long context is only useful if the cache it needs stays small.