Suggested searches

GPT Image Introduction GPT Image Free Guides GPT Image Developer Docs Midjourney Introduction Midjourney Free Guides Midjourney Developer Docs Google Nano Banana Introduction Google Nano Banana Free Guides Google Nano Banana Developer Docs Adobe Firefly Image Introduction Adobe Firefly Image Free Guides Adobe Firefly Image Developer Docs FLUX Introduction FLUX Free Guides FLUX Developer Docs Ideogram Introduction Ideogram Free Guides Ideogram Developer Docs Recraft Introduction Recraft Free Guides Recraft Developer Docs Stable Diffusion Introduction ByteDance Seedream Introduction Grok Imagine Image Introduction Google Veo Introduction Google Veo Free Guides Google Veo Developer Docs Runway Introduction Kling AI Introduction ByteDance Seedance Introduction ByteDance Seedance Free Guides ByteDance Seedance Developer Docs Luma AI Introduction Adobe Firefly Video Introduction Adobe Firefly Video Free Guides Adobe Firefly Video Developer Docs Hailuo AI Introduction PixVerse Introduction PixVerse Free Guides PixVerse Developer Docs Pika Introduction Pika Free Guides Pika Developer Docs Alibaba Wan Introduction LTX Video Introduction Grok Imagine Video Introduction Google Gemini Introduction Google Gemini Free Guides Google Gemini Developer Docs OpenAI GPT Introduction Claude Fable Introduction Claude Fable Free Guides Claude Fable Developer Docs DeepSeek Introduction Qwen Introduction Llama Introduction Codex Introduction Codex Free Guides Codex Developer Docs Cursor Introduction Cursor Free Guides Cursor Developer Docs OpenClaw Introduction OpenClaw Free Guides OpenClaw Developer Docs Perplexity Introduction ElevenLabs Introduction Suno Introduction Manus Introduction Claude vs ChatGPT: Compare Them Without Borrowed Numbers Claude vs Gemini: Comparing Two Assistants Honestly Is Adobe Firefly Free? The Free Membership, the Credits and the First Year Adobe Firefly Pricing: Plans, Generative Credits and the Cost of One Image Adobe Firefly Video Cost: Credits per Second and per Clip Is Adobe Firefly Video Free? Where the Free Plan Stops FLUX.2 Prompt Guide: The Structure Black Forest Labs Documents FLUX.2 vs Nano Banana Pro: What Each Vendor Actually Publishes Ideogram 4.0 Prompt Guide: The Documented Structure and How Text Renders Ideogram Pricing: Plans, Credits and the Free Limits Midjourney Prompt Guide: Structure, Elements and Edit Instructions Midjourney Pricing: Four Plans, GPU Hours and What a Job Costs Nano Banana Pro Pricing: What Google Publishes and What It Does Not Nano Banana Pro vs Midjourney: Which One Fits Your Workflow Pika Credits: What One Pika 2.5 Video Costs Is Pika Free? What the $0 Plan Actually Includes What PixVerse Credits Cost, and Which Ledger You Are Paying Is PixVerse Free? Two Free Tiers, and What Each One Costs You Recraft V4.1: The Eight Variants and Which to Pick Recraft Pricing and the Free Plan: What the Vendor Publishes Seedance Official Website: Which Entrances Are First-Party Seedance 2.5 vs 2.0: The Parameters ByteDance Publishes Veo 3.1 vs Sora 2: What Each Vendor Still Confirms Veo 3.1 vs Kling 3.0: What Each Vendor Publishes Is Claude AI Free? What the $0 Plan Actually Gives You Claude Pricing Plans: Pro vs Max 5x vs Max 20x How to Use Adobe Firefly (Adobe Firefly Image 5) Adobe Firefly API: Credentials, Endpoints and the Image5 Schema How to Use Adobe Firefly Video Adobe Firefly Video API: /v3/videos/generate Explained Designing UI Mockups and App Screens with GPT Image 2.5 Sticker Packs and Transparent Emoji with GPT Image 2.5 Nano Banana Pro Prompt Guide: The Official Frameworks Is Nano Banana Pro Free? What Google Actually Confirms How to Use Pika 2.5 Pika API: the official developer site, billing and endpoints How to Use PixVerse PixVerse API: Platform, Endpoints and Credits How to Use ByteDance Seedance 2.5 Seedance API: The Real Endpoint, Model ID and Fields Veo 3.1 Prompt Guide: The Seven Elements Google Names Veo 3.1 Price and Free Access: The Official Numbers How to Use Claude AI Claude API: Getting Started How to Use FLUX AI: A FLUX.2 Getting-Started Guide FLUX.2 API: Getting Started with Black Forest Labs GPT Image 2.5 vs DALL·E 3 Making Infographics and Diagrams with GPT Image 2.5 How to Use Ideogram Ideogram API: Access, Endpoints and Text Rendering How to Use Midjourney V8.2 Midjourney API: What Exists and What Does Not How to Use Nano Banana Pro Nano Banana Pro API: Getting Started How to Use Recraft Recraft API: Access, Endpoints and Style Consistency How to Use Veo 3.1 Veo 3.1 API Pricing and Vertex AI Access GPT Image 2.5 Prompt Sharing GPT Image 2.5 vs Nano Banana 2 GPT Image 2.5 vs Seedream 5.0 Pro GPT Image 2.5 vs FLUX 2 GPT Image 2.5 vs Ideogram Is GPT Image 2.5 Free GPT Image 2.5 Sketch GPT Image 2.5 Templates GPT Image 2.5 Comment Editing GPT Image 2.5 Character Consistency GPT Image 2.5 Combine Images GPT Image 2.5 Text Rendering What Is GPT Image 2.5 GPT Image 2.5 API Overview Codex vs. ChatGPT: Which Should You Use? Codex app, CLI, IDE, or cloud: how to choose the right surface Your First Low-Risk Coding Task with Codex: A Safe Walkthrough What Is Cursor? Its AI Coding Workflow Explained Cursor vs. VS Code: Which Editor Fits Your Workflow? Cursor Features Explained: Agent, Tab, Context, and More What Is Gemini? Apps, Models, AI Studio, and API Explained Gemini Apps vs. Gemini API: How to Choose the Right Tool for the Job What Can Gemini Do? A Practical Capability Guide OpenClaw Foundation Explained: Governance and Independence OpenClaw Skill Workshop Guide: Review Reusable Workflows OpenClaw Skill Cards: Read ClawHub Security Scans OpenClaw 2.0 Guide: New Features and Upgrade Checks OpenClaw LTS Guide: Choosing extended-stable or stable Install OpenClaw: Desktop, Script, npm, and Source Options OpenClaw Node.js Setup: Versions, Installation, and PATH How to Write Better Codex Prompts: A Practical Framework How to Review Codex Code Changes Before You Commit Cursor Beginner Tutorial: From Install to First Reviewed Edit Cursor Rules Tutorial: Project Rules, User Rules, and AGENTS.md Install Cursor on Windows and Configure a Chinese Interface Cursor MCP Tutorial: Configure, Verify, and Secure MCP Servers Gemini Prompt Guide: Better Instructions and Templates Gemini API Quickstart: Key, Python SDK and First Call Gemini Web App Guide: Login, Files, Chats and Privacy Gemini API Key Security: Storage, Restrictions and Rotation What Is Codex? Capabilities, Limits, and Ways to Use It Codex Beginner Tutorial: Complete Your First Safe Task Install Codex CLI: Sign In and Run Your First Safe Task Codex AGENTS.md Guide: Layered Rules and Validation Codex CLI Commands: Sessions, Review, and Automation How to Use Gemini: Web, Android & iPhone Setup Gemini Features Guide: Chat, Files, Images & Live How to Chat with Gemini: Prompts, Follow-Ups & Live Gemini AI Image Generator Guide: Prompts & Editing Gemini vs GPT-4: Features, Limits & Which to Use Gemini AI Assistant Guide: Mobile, Apps & Privacy Gemini Prompt Engineering Guide: Patterns & Examples Gemini Chat API Guide: Multi-Turn Prompts in Python Gemini System Instructions: API Guide & Examples Gemini Context Caching Guide: Cost, Latency & API
AI Tool Blog Llama 4 Maverick learning hub

Understand Llama 4 Maverick before you download it.

Llama 4 Maverick is Meta’s open-weights flagship: a mixture of experts with 17 billion parameters active per token, 128 experts in total and a one-million-token context window. It is the model teams download when the licence, not just the benchmark score, is the deciding factor — and the licence is the part most summaries leave out.

Every figure on this page comes from Meta’s published Llama 4 model card, the Llama 4 Community Licence and the Llama Models repository. There is no gallery: Meta publishes no static official figure for the Llama 4 family, and a page with no images is better than one padded with images we have no right to reproduce. Meta’s product navigation now leads with its Muse models, and the Llama 4 model page it used to serve is no longer reachable; the model card, the weights and the licence terms are unchanged.

What it is

How Meta describes Llama 4

Meta released Llama 4 in two sizes and describes both as natively multimodal mixture-of-experts models. Which one you pick depends on whether you are buying context or capability.

01

Two models, one collection

Scout is the context play: 17B activated across 16 experts, 109B parameters in total, a 10M-token window and a ~40T-token training corpus. Maverick is the capability play: 17B activated across 128 experts, 400B parameters in total, a 1M window and ~22T training tokens.

02

Multimodal by early fusion

Both fuse image and text inputs early rather than bolting a vision adapter onto a text model. The card lists multilingual text and images in, multilingual text and code out, and 12 supported languages: Arabic, English, French, German, Hindi, Indonesian, Italian, Portuguese, Spanish, Tagalog, Thai and Vietnamese.

03

22T tokens and a cutoff you should check

The card puts Maverick’s corpus at roughly 22T tokens with a knowledge cutoff of August 2024, and describes the models as static, trained on an offline dataset. Meta says future tuned models may follow as community feedback accumulates, but nothing after that date is in the weights you download today.

04

The licence is not an open-source licence

Llama 4 ships under the Llama 4 Community Licence Agreement, a custom commercial licence. It grants worldwide, non-exclusive, royalty-free use and modification, and then adds conditions: past 700 million monthly active users you must request a licence from Meta; you must display “Built with Llama” prominently; derivative models must carry a “Llama 4” name prefix. The Acceptable Use Policy applies on top.

05

Where the training data came from

Meta describes the corpus as a mix of publicly available and licensed data plus information from its own products and services — explicitly including publicly shared Instagram and Facebook posts and people’s interactions with Meta AI, with its Privacy Center linked for the detail. That disclosure is in the model card, not an inference from it.

The family

How the Llama models compare

Both Llama 4 rows come from the model card Meta publishes; the two older models are included to show the size of the jump.

Model Architecture Context length
Llama 4 Maverick MoE — 17B activated of 400B, 128 experts 1M
Llama 4 Scout MoE — 17B activated of 109B, 16 experts 10M
Llama 3.1 405B Dense — 405B 128K
Llama 3.3 70B Dense — 70B 128K

Meta’s own benchmarks put Maverick ahead of Llama 3.1 405B on MMLU (85.5 against 85.2), MMLU-Pro (62.9 against 61.6), MATH (61.2 against 53.5) and MBPP (77.6 against 74.4) — at a similar total parameter count, but with only 17B parameters active per token.

Capabilities

What the model card documents

Meta publishes specifications and benchmark tables rather than example images, so this section carries the capability claims the card actually makes.

01

Native image reasoning

Maverick scores 73.4 on MMMU, 59.6 on MMMU Pro and 73.7 on MathVista in Meta’s instruction-tuned table, and 85.3 on ChartQA and 91.6 on DocVQA in the pre-trained table.

02

A million-token window

1M for Maverick, 10M for Scout, both stated in the model card. Meta publishes no retrieval benchmark to go with those numbers, so treat them as capacity rather than as a measured long-context capability.

03

Twelve supported languages

Arabic, English, French, German, Hindi, Indonesian, Italian, Portuguese, Spanish, Tagalog, Thai and Vietnamese. Pre-training spanned 200 languages, and the licence allows fine-tuning into languages beyond the twelve.

04

Mixture-of-experts efficiency

17B parameters are activated per token out of 400B in total. That is the point of the architecture: near-large-model quality at a fraction of the per-token compute, provided you have the memory to hold the experts.

05

Fine-tuning and distillation are allowed

The card names synthetic data generation, distillation and adapting the pretrained models among the supported use cases, and the licence grants derivative works — subject to the naming, attribution and acceptable-use conditions.

06

Training cost, disclosed

Meta publishes the training footprint: 2.38M H100-80GB GPU hours for Maverick and 5.0M for Scout, 7.38M in total, with 1,999 tons of location-based CO2 equivalent offset to zero on a market basis. Few vendors put this in a model card.

Specifications

Documented specifications

Every row below is stated on Meta’s published Llama 4 model card or in the Llama 4 Community Licence Agreement.

Developer
Meta
Model
Llama 4 Maverick (17B×128E)
Also in the collection
Llama 4 Scout (17B×16E)
Total parameters
400B (Maverick) · 109B (Scout)
Activated parameters
17B per token
Experts
128 (Maverick) · 16 (Scout)
Context length
1M (Maverick) · 10M (Scout)
Input modalities
Multilingual text and image
Output modalities
Multilingual text and code
Training corpus
~22T tokens (Maverick) · ~40T (Scout)
Knowledge cutoff
August 2024
Release date
5 April 2025
Licence
Llama 4 Community Licence
How to use it

Documented access channels

Everything Llama 4 gives you arrives as a download: the weights, the model card, the licence and the CLI that fetches them. There is no hosted Meta endpoint for this model.

Llama 4 model card

The specification table, both benchmark tables, the intended-use statement, the training footprint and the safety notes — the source for every figure on this page.

Open

Llama 4 Community Licence

The binding terms: the 700-million-user clause, the “Built with Llama” display requirement and the attribution notice you must ship in a copy of the materials.

Open

Llama Models repository

Meta’s own repository for the whole family, with the model table across versions and the CLI documentation.

Open

Weights on Hugging Face

Published under Meta’s organisation on Hugging Face. The repositories are gated: you request access, accept the licence, and receive a signed download URL that expires after 24 hours.

Open
Common questions

Llama 4 Maverick questions

Is Llama 4 open source?

No, and the distinction is worth keeping. The weights are downloadable and the licence is royalty-free for most users, but it is the Llama 4 Community Licence Agreement — a custom commercial licence rather than an OSI-approved one. It adds a 700-million-user threshold, a “Built with Llama” display rule and a naming rule for derivative models.

What exactly does the 700-million-user clause say?

If the monthly active users of the products or services made available by you or your affiliates exceeded 700 million in the previous calendar month as of the model’s release date, you must request a licence from Meta and may not exercise any rights under the agreement until Meta grants one.

How much context does it have?

1M tokens for Maverick and 10M for Scout, both from Meta’s model card. Meta publishes no retrieval benchmark alongside those figures, so they describe capacity rather than measured long-context accuracy.

When was it released, and how current is it?

Llama 4 was released on 5 April 2025 with a knowledge cutoff of August 2024, and Meta describes the models as static, trained on an offline dataset. Meta’s newer frontier work sits under the Muse brand and its product navigation lists those models first; Llama 4 remains published, downloadable and licensed under the same terms.

Can I fine-tune it?

Yes. The model card names synthetic data generation, distillation and adapting the pretrained checkpoints among the supported use cases, and the licence grants derivative works — with the naming, attribution and acceptable-use conditions attached.

Can I run it locally?

Maverick holds 400B parameters, and the memory cost follows the total rather than the 17B active per token. Meta publishes no minimum hardware figure; the card documents the training configuration, not a deployment one. This is not a laptop model.

Does it support images?

Yes, natively and from early in pre-training rather than through a bolt-on adapter. For Maverick the card reports MMMU at 73.4 and MMMU Pro at 59.6, against 69.4 and 52.2 for Scout.