Ollama Basic
Run a local Llama 3.1 agent on Ollama in sync, async, and streaming modes.
"""
Ollama Basic
============
Cookbook example for `ollama/chat/basic.py`.
"""
from agno.agent import Agent, RunOutput # noqa
from agno.models.ollama import Ollama
import asyncio
# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
agent = Agent(model=Ollama(id="llama3.1:8b"), markdown=True)
# Get the response in a variable
# run: RunOutput = agent.run("Share a 2 sentence horror story")
# print(run.content)
# Print the response in the terminal
# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
# --- Sync ---
agent.print_response("Share a 2 sentence horror story")
# --- Sync + Streaming ---
agent.print_response("Share a 2 sentence horror story", stream=True)
# --- Async ---
asyncio.run(agent.aprint_response("Share a breakfast recipe.", markdown=True))
# --- Async + Streaming ---
asyncio.run(agent.aprint_response("Share a 2 sentence horror story", stream=True))Run the Example
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno ollamaSelect the local Ollama server
In the shell used for the pull commands and Python example, clear a previous cloud key and point native clients and embeddings at your local server:
unset OLLAMA_API_KEY
export OLLAMA_HOST="http://localhost:11434"Without this reset, OLLAMA_API_KEY makes Agno's default Ollama model route to https://ollama.com even when OLLAMA_HOST points locally. Keep a local Ollama server running for the following steps.
Prepare Ollama
Install and start Ollama, then pull the model used by this example:
ollama pull llama3.1:8bRun the example
Save the code above as basic.py, then run:
python basic.pyFull source: cookbook/90_models/ollama/chat/basic.py