Basic Accuracy Evaluation
Run AccuracyEval with run() and arun() over a multi-step calculator prompt, using 1 and 3 iterations.
Demonstrates synchronous and asynchronous accuracy evaluations.
Scores include only successfully evaluated iterations. A high average can therefore accompany missing results. Replace the score-only assertions with completeness checks before using this example as a quality gate.
"""
Basic Accuracy Evaluation
=========================
Demonstrates synchronous and asynchronous accuracy evaluations.
"""
import asyncio
from typing import Optional
from agno.agent import Agent
from agno.eval.accuracy import AccuracyEval, AccuracyResult
from agno.models.openai import OpenAIChat
from agno.tools.calculator import CalculatorTools
# ---------------------------------------------------------------------------
# Create Sync Evaluation
# ---------------------------------------------------------------------------
evaluation = AccuracyEval(
name="Calculator Evaluation",
model=OpenAIChat(id="o4-mini"),
agent=Agent(
model=OpenAIChat(id="gpt-5.6-luna"),
tools=[CalculatorTools()],
),
input="What is 10*5 then to the power of 2? do it step by step",
expected_output="2500",
additional_guidelines="Agent output should include the steps and the final answer.",
num_iterations=1,
)
# ---------------------------------------------------------------------------
# Create Async Evaluation
# ---------------------------------------------------------------------------
async_evaluation = AccuracyEval(
model=OpenAIChat(id="o4-mini"),
agent=Agent(
model=OpenAIChat(id="gpt-5.6-luna"),
tools=[CalculatorTools()],
),
input="What is 10*5 then to the power of 2? do it step by step",
expected_output="2500",
additional_guidelines="Agent output should include the steps and the final answer.",
num_iterations=3,
)
# ---------------------------------------------------------------------------
# Run Evaluation
# ---------------------------------------------------------------------------
if __name__ == "__main__":
result: Optional[AccuracyResult] = evaluation.run(print_results=True)
assert result is not None and result.avg_score >= 8
async_result: Optional[AccuracyResult] = asyncio.run(
async_evaluation.arun(print_results=True)
)
assert async_result is not None and async_result.avg_score >= 8Run the Example
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno openaiExport your OpenAI API key
export OPENAI_API_KEY="your_openai_api_key_here"Check every requested iteration
Replace the complete if __name__ == "__main__": block with:
if __name__ == "__main__":
result = evaluation.run(print_results=True)
assert result is not None, "No sync evaluation result"
assert len(result.results) == evaluation.num_iterations, "Incomplete sync evaluation"
assert all(item.score is not None for item in result.results)
assert result.avg_score is not None and result.avg_score >= 8
async_result = asyncio.run(async_evaluation.arun(print_results=True))
assert async_result is not None, "No async evaluation result"
assert len(async_result.results) == async_evaluation.num_iterations, "Incomplete async evaluation"
assert all(item.score is not None for item in async_result.results)
assert async_result.avg_score is not None and async_result.avg_score >= 8These assertions are development checks; run Python without -O so they remain enabled.
Run the example
Save the code above as accuracy_basic.py, then run:
python accuracy_basic.pyFull source: cookbook/09_evals/accuracy/accuracy_basic.py