Llama Cpp Basic

Run a local GGUF model with LlamaCpp and print sync and streamed responses.

basic.py
"""
Llama Cpp Basic
===============

Cookbook example for `llama_cpp/basic.py`.
"""

from agno.agent import Agent, RunOutput  # noqa
from agno.models.llama_cpp import LlamaCpp

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------

agent = Agent(model=LlamaCpp(id="ggml-org/gpt-oss-20b-GGUF"), markdown=True)

# Get the response in a variable
# run: RunOutput = agent.run("Share a 2 sentence horror story")
# print(run.content)

# Print the response in the terminal

# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    # --- Sync ---
    agent.print_response("Share a 2 sentence horror story")

    # --- Sync + Streaming ---
    agent.print_response("Share a 2 sentence horror story", stream=True)

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno openai

Install llama.cpp

Install the llama-server binary. This command supports macOS and Linux with Homebrew; see the llama.cpp installation guide for other platforms:

brew install llama.cpp

Start llama.cpp

Serve ggml-org/gpt-oss-20b-GGUF at http://127.0.0.1:8080/v1:

llama-server -hf ggml-org/gpt-oss-20b-GGUF --ctx-size 0 --jinja -ub 2048 -b 2048

Run the example

Save the code above as basic.py, then run:

python basic.py

Full source: cookbook/90_models/llama_cpp/basic.py