Llama Basic

Run Llama 4 Maverick through Meta's Llama API with sync, async, and streaming calls.

basic.py
"""
Meta Basic
==========

Cookbook example for `meta/llama/basic.py`.
"""

from agno.agent import Agent, RunOutput  # noqa
from agno.models.meta import Llama
import asyncio

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------

agent = Agent(
    model=Llama(id="Llama-4-Maverick-17B-128E-Instruct-FP8"),
    markdown=True,
)

# Get the response in a variable
# run: RunOutput = agent.run("Share a 2 sentence horror story")
# print(run.content)

# Print the response in the terminal

# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    # --- Sync ---
    agent.print_response("Share a 2 sentence horror story")

    # --- Sync + Streaming ---
    agent.print_response("Share a 2 sentence horror story", stream=True)

    # --- Async ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story"))

    # --- Async + Streaming ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story", stream=True))

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno llama-api-client

Export your Meta Llama API key

export LLAMA_API_KEY="your_llama_api_key_here"

Run the example

Save the code above as basic.py, then run:

python basic.py

Full source: cookbook/90_models/meta/llama/basic.py