Suggested searches

GPT Image Introduction GPT Image Free Guides GPT Image Developer Docs Midjourney Introduction Midjourney Free Guides Midjourney Developer Docs Google Nano Banana Introduction Google Nano Banana Free Guides Google Nano Banana Developer Docs Adobe Firefly Image Introduction Adobe Firefly Image Free Guides Adobe Firefly Image Developer Docs FLUX Introduction FLUX Free Guides FLUX Developer Docs Ideogram Introduction Ideogram Free Guides Ideogram Developer Docs Recraft Introduction Recraft Free Guides Recraft Developer Docs Stable Diffusion Introduction ByteDance Seedream Introduction Grok Imagine Image Introduction Google Veo Introduction Google Veo Free Guides Google Veo Developer Docs Runway Introduction Kling AI Introduction ByteDance Seedance Introduction ByteDance Seedance Free Guides ByteDance Seedance Developer Docs Luma AI Introduction Adobe Firefly Video Introduction Adobe Firefly Video Free Guides Adobe Firefly Video Developer Docs Hailuo AI Introduction PixVerse Introduction PixVerse Free Guides PixVerse Developer Docs Pika Introduction Pika Free Guides Pika Developer Docs Alibaba Wan Introduction LTX Video Introduction Grok Imagine Video Introduction Google Gemini Introduction Google Gemini Free Guides Google Gemini Developer Docs OpenAI GPT Introduction Claude Fable Introduction Claude Fable Free Guides Claude Fable Developer Docs DeepSeek Introduction Qwen Introduction Llama Introduction Codex Introduction Codex Free Guides Codex Developer Docs Cursor Introduction Cursor Free Guides Cursor Developer Docs OpenClaw Introduction OpenClaw Free Guides OpenClaw Developer Docs Perplexity Introduction ElevenLabs Introduction Suno Introduction Manus Introduction Claude vs ChatGPT: Compare Them Without Borrowed Numbers Claude vs Gemini: Comparing Two Assistants Honestly Is Adobe Firefly Free? The Free Membership, the Credits and the First Year Adobe Firefly Pricing: Plans, Generative Credits and the Cost of One Image Adobe Firefly Video Cost: Credits per Second and per Clip Is Adobe Firefly Video Free? Where the Free Plan Stops FLUX.2 Prompt Guide: The Structure Black Forest Labs Documents FLUX.2 vs Nano Banana Pro: What Each Vendor Actually Publishes Ideogram 4.0 Prompt Guide: The Documented Structure and How Text Renders Ideogram Pricing: Plans, Credits and the Free Limits Midjourney Prompt Guide: Structure, Elements and Edit Instructions Midjourney Pricing: Four Plans, GPU Hours and What a Job Costs Nano Banana Pro Pricing: What Google Publishes and What It Does Not Nano Banana Pro vs Midjourney: Which One Fits Your Workflow Pika Credits: What One Pika 2.5 Video Costs Is Pika Free? What the $0 Plan Actually Includes What PixVerse Credits Cost, and Which Ledger You Are Paying Is PixVerse Free? Two Free Tiers, and What Each One Costs You Recraft V4.1: The Eight Variants and Which to Pick Recraft Pricing and the Free Plan: What the Vendor Publishes Seedance Official Website: Which Entrances Are First-Party Seedance 2.5 vs 2.0: The Parameters ByteDance Publishes Veo 3.1 vs Sora 2: What Each Vendor Still Confirms Veo 3.1 vs Kling 3.0: What Each Vendor Publishes Is Claude AI Free? What the $0 Plan Actually Gives You Claude Pricing Plans: Pro vs Max 5x vs Max 20x How to Use Adobe Firefly (Adobe Firefly Image 5) Adobe Firefly API: Credentials, Endpoints and the Image5 Schema How to Use Adobe Firefly Video Adobe Firefly Video API: /v3/videos/generate Explained Designing UI Mockups and App Screens with GPT Image 2.5 Sticker Packs and Transparent Emoji with GPT Image 2.5 Nano Banana Pro Prompt Guide: The Official Frameworks Is Nano Banana Pro Free? What Google Actually Confirms How to Use Pika 2.5 Pika API: the official developer site, billing and endpoints How to Use PixVerse PixVerse API: Platform, Endpoints and Credits How to Use ByteDance Seedance 2.5 Seedance API: The Real Endpoint, Model ID and Fields Veo 3.1 Prompt Guide: The Seven Elements Google Names Veo 3.1 Price and Free Access: The Official Numbers How to Use Claude AI Claude API: Getting Started How to Use FLUX AI: A FLUX.2 Getting-Started Guide FLUX.2 API: Getting Started with Black Forest Labs GPT Image 2.5 vs DALL·E 3 Making Infographics and Diagrams with GPT Image 2.5 How to Use Ideogram Ideogram API: Access, Endpoints and Text Rendering How to Use Midjourney V8.2 Midjourney API: What Exists and What Does Not How to Use Nano Banana Pro Nano Banana Pro API: Getting Started How to Use Recraft Recraft API: Access, Endpoints and Style Consistency How to Use Veo 3.1 Veo 3.1 API Pricing and Vertex AI Access GPT Image 2.5 Prompt Sharing GPT Image 2.5 vs Nano Banana 2 GPT Image 2.5 vs Seedream 5.0 Pro GPT Image 2.5 vs FLUX 2 GPT Image 2.5 vs Ideogram Is GPT Image 2.5 Free GPT Image 2.5 Sketch GPT Image 2.5 Templates GPT Image 2.5 Comment Editing GPT Image 2.5 Character Consistency GPT Image 2.5 Combine Images GPT Image 2.5 Text Rendering What Is GPT Image 2.5 GPT Image 2.5 API Overview Codex vs. ChatGPT: Which Should You Use? Codex app, CLI, IDE, or cloud: how to choose the right surface Your First Low-Risk Coding Task with Codex: A Safe Walkthrough What Is Cursor? Its AI Coding Workflow Explained Cursor vs. VS Code: Which Editor Fits Your Workflow? Cursor Features Explained: Agent, Tab, Context, and More What Is Gemini? Apps, Models, AI Studio, and API Explained Gemini Apps vs. Gemini API: How to Choose the Right Tool for the Job What Can Gemini Do? A Practical Capability Guide OpenClaw Foundation Explained: Governance and Independence OpenClaw Skill Workshop Guide: Review Reusable Workflows OpenClaw Skill Cards: Read ClawHub Security Scans OpenClaw 2.0 Guide: New Features and Upgrade Checks OpenClaw LTS Guide: Choosing extended-stable or stable Install OpenClaw: Desktop, Script, npm, and Source Options OpenClaw Node.js Setup: Versions, Installation, and PATH How to Write Better Codex Prompts: A Practical Framework How to Review Codex Code Changes Before You Commit Cursor Beginner Tutorial: From Install to First Reviewed Edit Cursor Rules Tutorial: Project Rules, User Rules, and AGENTS.md Install Cursor on Windows and Configure a Chinese Interface Cursor MCP Tutorial: Configure, Verify, and Secure MCP Servers Gemini Prompt Guide: Better Instructions and Templates Gemini API Quickstart: Key, Python SDK and First Call Gemini Web App Guide: Login, Files, Chats and Privacy Gemini API Key Security: Storage, Restrictions and Rotation What Is Codex? Capabilities, Limits, and Ways to Use It Codex Beginner Tutorial: Complete Your First Safe Task Install Codex CLI: Sign In and Run Your First Safe Task Codex AGENTS.md Guide: Layered Rules and Validation Codex CLI Commands: Sessions, Review, and Automation How to Use Gemini: Web, Android & iPhone Setup Gemini Features Guide: Chat, Files, Images & Live How to Chat with Gemini: Prompts, Follow-Ups & Live Gemini AI Image Generator Guide: Prompts & Editing Gemini vs GPT-4: Features, Limits & Which to Use Gemini AI Assistant Guide: Mobile, Apps & Privacy Gemini Prompt Engineering Guide: Patterns & Examples Gemini Chat API Guide: Multi-Turn Prompts in Python Gemini System Instructions: API Guide & Examples Gemini Context Caching Guide: Cost, Latency & API
Gemini context caching

Reuse large Gemini context when repeated requests justify the cache.

Context caching lets applications reuse large, repeated input such as documents, media, code, or extensive system instructions. It can reduce repeated processing cost and latency, but only when cache reuse, lifetime, model support, and storage risk fit the workload.

On this page

Caching decisions at a glance

01

Implicit caching

Some supported models may automatically reuse matching prompt prefixes. Applications do not create or manage a cache resource.

02

Explicit caching

Create a reusable cached-content resource, reference it in later requests, and manage its lifetime directly.

03

Reuse economics

Caching is most useful when a large stable prefix is reused enough times before it expires to offset storage and setup cost.

04

Data lifecycle

Cached content has retention and access implications. Apply least privilege, expiration, deletion, regional, and privacy controls.

01

Choose a suitable workload

Good candidates have a large stable context followed by many smaller questions: a long manual, video, codebase, policy collection, or extensive system instruction. One-off prompts and frequently changing context usually gain less.

  • Estimate context size, expected cache hits, lifetime, and request frequency.
  • Keep the reusable prefix stable and place changing user input after it.
  • Compare total uncached cost with cache creation, storage, and cached-input charges.
02

Understand implicit and explicit caching

Implicit caching is automatic on supported models and may report cached tokens in usage metadata. Explicit caching creates a named resource that later generation requests reference, giving the application more predictable reuse and lifecycle control.

  • Do not assume every model, input type, or request is eligible for caching.
  • Inspect cached-content token metadata rather than inferring a hit from latency alone.
  • Use explicit caching when controlled reuse and a managed TTL matter to the design.
03

Create and reference cached content

With explicit caching, create the cache using a supported model and repeated content, then pass the returned cached-content name in later generation configuration. Keep dynamic questions outside the cached resource.

  • Use a descriptive display name and store the provider resource name with your application record.
  • Set a TTL aligned with the real reuse window and refresh it only when necessary.
  • Handle cache creation, lookup, expiration, deletion, and fallback to an uncached request.
04

Measure cost and latency

A cache is an optimization that should be proven with production-like measurements. Track creation latency, hit rate, cached and uncached input tokens, storage duration, end-to-end latency, and total cost per completed task.

  • Benchmark against an uncached baseline using the same model and content.
  • Separate time to first token from total response time for streaming interfaces.
  • Revisit the decision when pricing, token thresholds, model support, or traffic patterns change.
05

Protect cached data

Cached content can include documents, media, source code, and system instructions. Treat the cache as stored data: minimize content, control access, define deletion, and review provider retention and regional behavior.

  • Do not cache secrets or data that the application is not authorized to retain.
  • Keep cache identifiers server-side and authorize every request that references them.
  • Delete obsolete caches and document incident, privacy, and retention procedures.
Python explicit cache example

Python

This simplified flow creates cached content with a one-hour TTL and references it in a later request. Replace the placeholder with eligible repeated context and verify current SDK syntax.

                import os
from google import genai
from google.genai import types

client = genai.Client()
model = os.environ["GEMINI_MODEL"]
cached = client.caches.create(
    model=model,
    config=types.CreateCachedContentConfig(
        display_name="product-manual",
        contents=["<large repeated context>"],
        ttl="3600s",
    ),
)

response = client.models.generate_content(
    model=model,
    contents="Summarize the upgrade procedure.",
    config=types.GenerateContentConfig(cached_content=cached.name),
)
print(response.text)
              
Published
By
AI Tool Blog