Cloudflare AI Gateway (basic)

Run a Workers AI model through Cloudflare AI Gateway in sync, async, and streaming modes.

Default model is Workers AI (only Cloudflare token + account id). You can paste a catalog binding: Agent(model="cloudflare:@cf/meta/llama-3.3-70b-instruct-fp8-fast"), the full gateway id, or Cloudflare() defaults.

The current Agno Cloudflare adapter uses AI Gateway's /compat/chat/completions endpoint. Cloudflare deprecates this endpoint for single-model calls while keeping it working for existing integrations. Dynamic routes still require it. For new direct model integrations, follow Cloudflare's current REST API guide.

basic.py
"""
Cloudflare AI Gateway (basic)
=============================

Cookbook example for Cloudflare AI Gateway OpenAI-compatible unified API.

Requires:
- CLOUDFLARE_API_TOKEN
- CLOUDFLARE_ACCOUNT_ID

Optional:
- CLOUDFLARE_AI_GATEWAY_ID (defaults to the ``default`` gateway)

Default model is **Workers AI** (only Cloudflare token + account id). You can paste a catalog binding:
``Agent(model="cloudflare:@cf/meta/llama-3.3-70b-instruct-fp8-fast")``, the full gateway id, or ``Cloudflare()`` defaults.

Set ``CLOUDFLARE_API_TOKEN`` and ``CLOUDFLARE_ACCOUNT_ID`` in your shell before running (see README),
not in this file.

For switching models (OpenRouter-style ``id`` string), see ``switch_model.py`` in this folder.

For ``google/...`` or ``openai/...`` you need a supported gateway provider name, BYOK keys in the
dashboard, and a model id from the Cloudflare docs — otherwise you may see **Invalid provider** (HTTP 400)
or upstream **401** errors.
"""

import asyncio

from agno.agent import Agent
from agno.models.cloudflare import Cloudflare

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------

agent = Agent(
    model=Cloudflare("@cf/meta/llama-3.3-70b-instruct-fp8-fast"),
    markdown=True,
)

# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    # --- Sync ---
    agent.print_response("Share a 2 sentence horror story")

    # --- Sync + Streaming ---
    agent.print_response("Share a 2 sentence horror story", stream=True)

    # --- Async ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story"))

    # --- Async + Streaming ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story", stream=True))

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno openai

Set up Cloudflare access

Use a Cloudflare account and an API token with the permissions required for AI Gateway and Workers AI. The default gateway is created on the first request; set CLOUDFLARE_AI_GATEWAY_ID only to use another existing gateway.

export CLOUDFLARE_ACCOUNT_ID="your_cloudflare_account_id_here"
export CLOUDFLARE_API_TOKEN="your_cloudflare_api_token_here"

Workers AI examples use this Cloudflare token. Other providers can require stored provider keys or supported unified billing; configure the selected route in the gateway before switching to it.

Run the example

Save the code above as basic.py, then run:

python basic.py

Full source: cookbook/90_models/cloudflare/basic.py