Suggested searches

GPT Image Introduction GPT Image Free Guides GPT Image Developer Docs Midjourney Introduction Midjourney Free Guides Midjourney Developer Docs Google Nano Banana Introduction Google Nano Banana Free Guides Google Nano Banana Developer Docs Adobe Firefly Image Introduction Adobe Firefly Image Free Guides Adobe Firefly Image Developer Docs FLUX Introduction FLUX Free Guides FLUX Developer Docs Ideogram Introduction Ideogram Free Guides Ideogram Developer Docs Recraft Introduction Recraft Free Guides Recraft Developer Docs Stable Diffusion Introduction ByteDance Seedream Introduction Grok Imagine Image Introduction Google Veo Introduction Google Veo Free Guides Google Veo Developer Docs Runway Introduction Kling AI Introduction ByteDance Seedance Introduction ByteDance Seedance Free Guides ByteDance Seedance Developer Docs Luma AI Introduction Adobe Firefly Video Introduction Adobe Firefly Video Free Guides Adobe Firefly Video Developer Docs Hailuo AI Introduction PixVerse Introduction PixVerse Free Guides PixVerse Developer Docs Pika Introduction Pika Free Guides Pika Developer Docs Alibaba Wan Introduction LTX Video Introduction Grok Imagine Video Introduction Google Gemini Introduction Google Gemini Free Guides Google Gemini Developer Docs OpenAI GPT Introduction Claude Fable Introduction Claude Fable Free Guides Claude Fable Developer Docs DeepSeek Introduction Qwen Introduction Llama Introduction Codex Introduction Codex Free Guides Codex Developer Docs Cursor Introduction Cursor Free Guides Cursor Developer Docs OpenClaw Introduction OpenClaw Free Guides OpenClaw Developer Docs Perplexity Introduction ElevenLabs Introduction Suno Introduction Manus Introduction Claude vs ChatGPT: Compare Them Without Borrowed Numbers Claude vs Gemini: Comparing Two Assistants Honestly Is Adobe Firefly Free? The Free Membership, the Credits and the First Year Adobe Firefly Pricing: Plans, Generative Credits and the Cost of One Image Adobe Firefly Video Cost: Credits per Second and per Clip Is Adobe Firefly Video Free? Where the Free Plan Stops FLUX.2 Prompt Guide: The Structure Black Forest Labs Documents FLUX.2 vs Nano Banana Pro: What Each Vendor Actually Publishes Ideogram 4.0 Prompt Guide: The Documented Structure and How Text Renders Ideogram Pricing: Plans, Credits and the Free Limits Midjourney Prompt Guide: Structure, Elements and Edit Instructions Midjourney Pricing: Four Plans, GPU Hours and What a Job Costs Nano Banana Pro Pricing: What Google Publishes and What It Does Not Nano Banana Pro vs Midjourney: Which One Fits Your Workflow Pika Credits: What One Pika 2.5 Video Costs Is Pika Free? What the $0 Plan Actually Includes What PixVerse Credits Cost, and Which Ledger You Are Paying Is PixVerse Free? Two Free Tiers, and What Each One Costs You Recraft V4.1: The Eight Variants and Which to Pick Recraft Pricing and the Free Plan: What the Vendor Publishes Seedance Official Website: Which Entrances Are First-Party Seedance 2.5 vs 2.0: The Parameters ByteDance Publishes Veo 3.1 vs Sora 2: What Each Vendor Still Confirms Veo 3.1 vs Kling 3.0: What Each Vendor Publishes Is Claude AI Free? What the $0 Plan Actually Gives You Claude Pricing Plans: Pro vs Max 5x vs Max 20x How to Use Adobe Firefly (Adobe Firefly Image 5) Adobe Firefly API: Credentials, Endpoints and the Image5 Schema How to Use Adobe Firefly Video Adobe Firefly Video API: /v3/videos/generate Explained Designing UI Mockups and App Screens with GPT Image 2.5 Sticker Packs and Transparent Emoji with GPT Image 2.5 Nano Banana Pro Prompt Guide: The Official Frameworks Is Nano Banana Pro Free? What Google Actually Confirms How to Use Pika 2.5 Pika API: the official developer site, billing and endpoints How to Use PixVerse PixVerse API: Platform, Endpoints and Credits How to Use ByteDance Seedance 2.5 Seedance API: The Real Endpoint, Model ID and Fields Veo 3.1 Prompt Guide: The Seven Elements Google Names Veo 3.1 Price and Free Access: The Official Numbers How to Use Claude AI Claude API: Getting Started How to Use FLUX AI: A FLUX.2 Getting-Started Guide FLUX.2 API: Getting Started with Black Forest Labs GPT Image 2.5 vs DALL·E 3 Making Infographics and Diagrams with GPT Image 2.5 How to Use Ideogram Ideogram API: Access, Endpoints and Text Rendering How to Use Midjourney V8.2 Midjourney API: What Exists and What Does Not How to Use Nano Banana Pro Nano Banana Pro API: Getting Started How to Use Recraft Recraft API: Access, Endpoints and Style Consistency How to Use Veo 3.1 Veo 3.1 API Pricing and Vertex AI Access GPT Image 2.5 Prompt Sharing GPT Image 2.5 vs Nano Banana 2 GPT Image 2.5 vs Seedream 5.0 Pro GPT Image 2.5 vs FLUX 2 GPT Image 2.5 vs Ideogram Is GPT Image 2.5 Free GPT Image 2.5 Sketch GPT Image 2.5 Templates GPT Image 2.5 Comment Editing GPT Image 2.5 Character Consistency GPT Image 2.5 Combine Images GPT Image 2.5 Text Rendering What Is GPT Image 2.5 GPT Image 2.5 API Overview Codex vs. ChatGPT: Which Should You Use? Codex app, CLI, IDE, or cloud: how to choose the right surface Your First Low-Risk Coding Task with Codex: A Safe Walkthrough What Is Cursor? Its AI Coding Workflow Explained Cursor vs. VS Code: Which Editor Fits Your Workflow? Cursor Features Explained: Agent, Tab, Context, and More What Is Gemini? Apps, Models, AI Studio, and API Explained Gemini Apps vs. Gemini API: How to Choose the Right Tool for the Job What Can Gemini Do? A Practical Capability Guide OpenClaw Foundation Explained: Governance and Independence OpenClaw Skill Workshop Guide: Review Reusable Workflows OpenClaw Skill Cards: Read ClawHub Security Scans OpenClaw 2.0 Guide: New Features and Upgrade Checks OpenClaw LTS Guide: Choosing extended-stable or stable Install OpenClaw: Desktop, Script, npm, and Source Options OpenClaw Node.js Setup: Versions, Installation, and PATH How to Write Better Codex Prompts: A Practical Framework How to Review Codex Code Changes Before You Commit Cursor Beginner Tutorial: From Install to First Reviewed Edit Cursor Rules Tutorial: Project Rules, User Rules, and AGENTS.md Install Cursor on Windows and Configure a Chinese Interface Cursor MCP Tutorial: Configure, Verify, and Secure MCP Servers Gemini Prompt Guide: Better Instructions and Templates Gemini API Quickstart: Key, Python SDK and First Call Gemini Web App Guide: Login, Files, Chats and Privacy Gemini API Key Security: Storage, Restrictions and Rotation What Is Codex? Capabilities, Limits, and Ways to Use It Codex Beginner Tutorial: Complete Your First Safe Task Install Codex CLI: Sign In and Run Your First Safe Task Codex AGENTS.md Guide: Layered Rules and Validation Codex CLI Commands: Sessions, Review, and Automation How to Use Gemini: Web, Android & iPhone Setup Gemini Features Guide: Chat, Files, Images & Live How to Chat with Gemini: Prompts, Follow-Ups & Live Gemini AI Image Generator Guide: Prompts & Editing Gemini vs GPT-4: Features, Limits & Which to Use Gemini AI Assistant Guide: Mobile, Apps & Privacy Gemini Prompt Engineering Guide: Patterns & Examples Gemini Chat API Guide: Multi-Turn Prompts in Python Gemini System Instructions: API Guide & Examples Gemini Context Caching Guide: Cost, Latency & API
AI Tool Blog Qwen3.8 Max learning hub

Understand Qwen3.8 Max before you size the hardware.

Qwen3.8 Max is Alibaba’s flagship model, and Qwen3.8 is the first generation in which the Max class was released as open weights at all: the hosted Max builds on the published 2.4-trillion-parameter checkpoint. Its architecture is a 3-to-1 hybrid — three linear-attention layers for every sparse-attention layer — which is what makes a 262,144-token context affordable to serve.

Architecture, benchmarks and licence terms on this page come from Qwen’s own model cards and the Qwen3.8 release blog. Qwen publishes no per-token price on the model card; hosted rates live on Qwen Cloud.

Official figures

The architecture Qwen draws for the 3.8 generation

Qwen’s cards carry one static figure: the block diagram of the hybrid attention stack. The 2.4T card describes that layout in text, and the diagram comes from the sibling Qwen3.8-Flash-Next card, which draws it out.

Qwen’s official diagram of the hybrid Gated DeltaNet and sparse-attention layout used by the Qwen3.8 generation
Inside a hybrid block: three Gated DeltaNet layers — the linear-attention path — for every sparse-attention layer, each followed by a mixture-of-experts module. An N-gram embedding layer feeds the second layer, GR Read and GR Write move state between blocks, and multi-token-prediction modules sit on top. The 2.4T card describes the same three-to-one rhythm for Qwen3.8 Max, naming its sparse path Gated Attention.

The diagram belongs to Qwen and is reproduced with credit: Qwen3.8 Flash Next model card

What it is

How Qwen describes Qwen3.8 Max

Qwen calls Qwen3.8 the most capable generation in its open-model family, and the release blog is titled “A New Bar for Coding and Cowork”. The architecture underneath it is the more interesting part: a deliberate mix of linear and sparse attention.

01

The Max class went open

Qwen says this is the first time a Qwen-Max-class model has been released openly, and publishes it as Qwen3.8-2.4T-A95B. The hosted Qwen3.8 Max is that checkpoint plus serving features the weights do not carry.

02

Three linear layers to one sparse layer

The hidden layout is 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)) across 92 layers. Gated DeltaNet carries 128 linear-attention heads for values against 16 for queries and keys; the sparse path uses 64 query heads against 4 key-value heads with a 256-wide head dimension. The linear path is what makes the long context affordable.

03

A very large vocabulary and an N-gram embedding layer

The token embedding is 248,320 entries, padded. The 2.4T card also carries multi-token prediction trained over multiple steps, and the generation’s block diagram shows an N-gram embedding layer feeding layer two.

04

What Max adds over the raw weights

Qwen’s card is explicit: the hosted Max adds vision input, non-thinking mode, a 1M context window by default and official built-in tools. The published weights are text-only and require thinking mode — the card states that thinking cannot be turned off there.

05

Benchmarked on an unusually wide surface

The comparison table covers coding agents, general agent tasks, professional work in law, finance and health, and long-context retrieval. That breadth is the argument: Qwen is selling a general worker rather than a coding model, and the table shows where it still loses.

Official benchmarks

What the model card reports

Qwen’s own comparison table, with the best rival column added so the rows where Qwen3.8 Max loses stay visible. It is a vendor-run table, and the competitor figures are those vendors’ published scores.

Benchmark Qwen3.8 Max Best rival in the same table
Terminal Bench 2.1 86.6 GPT 5.6 Sol — 88.8
SWE-bench Pro 67.7 Fable 5 — 80.0
PaperBench 93.0 GPT 5.6 Sol — 90.5
CoWorkBench 74.8 Fable 5 — 75.9
WideSearch 81.9 Fable 5 — 81.2
IFBench 82.8 GPT 5.6 Sol — 72.7
GPQA Diamond 92.6 GPT 5.6 Sol — 94.1
HLE 43.6 Fable 5 — 53.3
HLE with tools 56.2 Fable 5 — 64.5
MRCR v2 256K 92.9 GPT 5.6 Sol — 93.8
HealthBench 60.2 GPT 5.6 Sol — 55.3

Qwen ran its own Terminal Bench 2.1 evaluation with the Claude Code harness at average@10, while taking other models’ best published scores from different harnesses — Artificial Analysis for the Claude models, Codex for GPT-5.6 Sol. The card flags that Fable 5 results may involve fallbacks, and that several of its rows are in-house benchmarks a reader cannot reproduce.

Specifications

Documented specifications

Every row below is stated on Qwen’s model card for the 2.4T checkpoint.

Developer
Alibaba Qwen Team
Model
Qwen3.8 Max
Open weights
Qwen3.8-2.4T-A95B
Parameters
2.4T total, 95B activated
Layers
92
Hidden layout
23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))
Experts
512 total, 10 routed + 1 shared per token
Vocabulary
248,320, padded
Context
262,144 native, extensible to 1,010,000
Max adds
Vision input, non-thinking mode, 1M context, built-in tools
Reasoning effort
xhigh (default), medium, low
Licence
Qwen3.8-Max licence
How to use it

Documented access channels

There are two ways in, and they are not the same model: the downloadable 2.4T checkpoint, or the hosted Max that adds vision, a non-thinking mode and built-in tools.

Official model card

Qwen’s own card for the 2.4T checkpoint: the architecture, the context length, the sampling parameters and the full comparison table.

Open

Qwen3.8 Max on Qwen Cloud

The hosted model page, which lists the features the checkpoint does not carry: vision input, non-thinking mode, 1M context by default and the official built-in tools.

Open

Qwen Studio

The consumer entry point, which can be pointed straight at Qwen3.8 Max for a hands-on check before you commit to an integration.

Open

Serving frameworks

Qwen links official recipes for SGLang, vLLM and TokenSpeed and recommends them for production or high-throughput work, warning that throughput varies significantly between frameworks and versions.

Common questions

Qwen3.8 Max questions

Is Qwen3.8 Max open weights?

Partly, and the split matters. Qwen released the underlying 2.4T checkpoint openly — the first Max-class model to get that treatment — but the hosted Qwen3.8 Max adds vision input, non-thinking mode, a 1M context window by default and official built-in tools that the checkpoint does not have. The licence is the custom qwen3.8-max licence, not an OSI licence.

Can I run it on my own hardware?

This is a 2.4T-parameter model with 95B activated per token, so it is data-centre territory rather than workstation territory. Qwen documents serving through SGLang, vLLM and TokenSpeed and recommends the latest framework versions; it publishes no minimum hardware figure.

How large is the context window?

262,144 tokens natively, extensible to 1,010,000. The hosted Max runs with a 1M context by default. On Qwen’s own MRCR v2 256K retrieval test the model scores 92.9, against 93.8 for GPT 5.6 Sol and 86.7 for its own predecessor, Qwen3.7-Max.

Why does the model card say thinking cannot be disabled?

Because the published checkpoint is text-only and requires thinking mode for every interaction: multimodal inputs are unsupported there, and every response begins with reasoning before the final answer. Non-thinking support is one of the features Qwen reserves for the hosted Max.

How do I control how much it thinks?

Through reasoning_effort, which takes xhigh (the default), medium or low. A second switch, preserve_thinking, is on by default and retains reasoning context from earlier messages. Qwen recommends temperature 1.0, top_p 0.95, top_k 20, min_p 0.0, presence_penalty 0.0 and repetition_penalty 1.0.

How much does Qwen3.8 Max cost?

Qwen’s model card publishes no per-token rate, so this page states none. The card points to Qwen Cloud for managed inference, and its model page is the place to check current pricing.

Is it actually better than the closed frontier models?

It depends on the row. On Qwen’s own table Qwen3.8 Max leads PaperBench at 93.0 against 90.5, WideSearch and IFBench, and trails on SWE-bench Pro at 67.7 against 80.0, HLE at 43.6 against 53.3 and Terminal Bench 2.1 at 86.6 against 88.8. Several rows are Qwen’s in-house benchmarks, and the card notes that different models were evaluated with different harnesses.