Prompt Caching
Enable Claude cache_system_prompt for a large system message and print cache write and read token counts across two runs.
Use prompt caching with Anthropic agents to cache the system prompt passed to the model.
Anthropic retired the source's claude-sonnet-4-20250514 model on June 15, 2026. Replace it with claude-sonnet-4-6 before running. See Anthropic model deprecations.
"""
This cookbook shows how to use prompt caching with Agents using Anthropic models, to catch the system prompt passed to the model.
This can significantly reduce processing time and costs.
Use it when working with a static and large system prompt.
You can check more about prompt caching with Anthropic models here: https://docs.anthropic.com/en/docs/prompt-caching
"""
from pathlib import Path
from agno.agent import Agent
from agno.models.anthropic import Claude
from agno.utils.media import download_file
# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
# Load an example large system message from S3. A large prompt like this would benefit from caching.
txt_path = Path(__file__).parent.joinpath("system_prompt.txt")
download_file(
"https://agno-public.s3.amazonaws.com/prompts/system_promt.txt",
str(txt_path),
)
system_message = txt_path.read_text()
agent = Agent(
model=Claude(
id="claude-sonnet-4-20250514",
cache_system_prompt=True, # Activate prompt caching for Anthropic to cache the system prompt
),
system_message=system_message,
markdown=True,
)
# First run - this will create the cache
response = agent.run(
"Explain the difference between REST and GraphQL APIs with examples"
)
if response and response.metrics:
print(f"First run cache write tokens = {response.metrics.cache_write_tokens}")
# Second run - this will use the cached system prompt
response = agent.run(
"What are the key principles of clean code and how do I apply them in Python?"
)
if response and response.metrics:
print(f"Second run cache read tokens = {response.metrics.cache_read_tokens}")
# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
passRun the Example
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno anthropicExport your Anthropic API key
export ANTHROPIC_API_KEY="your_anthropic_api_key_here"Update the Claude model
Replace Claude(id="claude-sonnet-4-20250514") with Claude(id="claude-sonnet-4-6") in the saved file.
Run the example
Save the code above as prompt_caching.py, then run:
python prompt_caching.pyFull source: cookbook/90_models/anthropic/prompt_caching.py