Llama OpenAI Basic

Run Llama 4 Maverick through Meta's OpenAI-compatible API with sync, async, and streaming calls.

At the linked source revision, LlamaOpenAI has a message-formatter signature mismatch and fails before sending a request. Apply the compatible adapter instructions below before running this example.

basic.py
"""
Meta Basic
==========

Cookbook example for `meta/llama_openai/basic.py`.
"""

from agno.agent import Agent, RunOutput  # noqa
from agno.models.meta import LlamaOpenAI
import asyncio

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------

agent = Agent(
    model=LlamaOpenAI(id="Llama-4-Maverick-17B-128E-Instruct-FP8"),
    markdown=True,
)

# Get the response in a variable
# run: RunOutput = agent.run("Share a 2 sentence horror story")
# print(run.content)

# Print the response in the terminal

# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    # --- Sync ---
    agent.print_response("Share a 2 sentence horror story")

    # --- Sync + Streaming ---
    agent.print_response("Share a 2 sentence horror story", stream=True)

    # --- Async ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story"))

    # --- Async + Streaming ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story", stream=True))

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno llama-api-client openai

Export your Meta Llama API key

export LLAMA_API_KEY="your_llama_api_key_here"

Use the compatible API adapter

Replace the LlamaOpenAI import (or Llama in the byte-image source) with this helper. Then replace every LlamaOpenAI(...) or Llama(...) construction in the saved file with llama_model(...). Keep the existing id, temperature, and any retry options inside those calls.

Compatible model helper
from os import getenv

from agno.models.openai.like import OpenAILike


def llama_model(**kwargs):
    return OpenAILike(
        api_key=getenv("LLAMA_API_KEY"),
        base_url="https://api.llama.com/compat/v1/",
        supports_native_structured_outputs=False,
        supports_json_schema_outputs=True,
        **kwargs,
    )

This uses Meta's OpenAI-compatible endpoint. You need a Meta API account with access to the selected model; check your account's current model catalog before running.

Run the example

Save the code above as basic.py, then run:

python basic.py

Full source: cookbook/90_models/meta/llama_openai/basic.py