Basic Accuracy Evaluation

Run AccuracyEval with run() and arun() over a multi-step calculator prompt, using 1 and 3 iterations.

Demonstrates synchronous and asynchronous accuracy evaluations.

Scores include only successfully evaluated iterations. A high average can therefore accompany missing results. Replace the score-only assertions with completeness checks before using this example as a quality gate.

accuracy_basic.py
"""
Basic Accuracy Evaluation
=========================

Demonstrates synchronous and asynchronous accuracy evaluations.
"""

import asyncio
from typing import Optional

from agno.agent import Agent
from agno.eval.accuracy import AccuracyEval, AccuracyResult
from agno.models.openai import OpenAIChat
from agno.tools.calculator import CalculatorTools

# ---------------------------------------------------------------------------
# Create Sync Evaluation
# ---------------------------------------------------------------------------
evaluation = AccuracyEval(
    name="Calculator Evaluation",
    model=OpenAIChat(id="o4-mini"),
    agent=Agent(
        model=OpenAIChat(id="gpt-5.6-luna"),
        tools=[CalculatorTools()],
    ),
    input="What is 10*5 then to the power of 2? do it step by step",
    expected_output="2500",
    additional_guidelines="Agent output should include the steps and the final answer.",
    num_iterations=1,
)

# ---------------------------------------------------------------------------
# Create Async Evaluation
# ---------------------------------------------------------------------------
async_evaluation = AccuracyEval(
    model=OpenAIChat(id="o4-mini"),
    agent=Agent(
        model=OpenAIChat(id="gpt-5.6-luna"),
        tools=[CalculatorTools()],
    ),
    input="What is 10*5 then to the power of 2? do it step by step",
    expected_output="2500",
    additional_guidelines="Agent output should include the steps and the final answer.",
    num_iterations=3,
)

# ---------------------------------------------------------------------------
# Run Evaluation
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    result: Optional[AccuracyResult] = evaluation.run(print_results=True)
    assert result is not None and result.avg_score >= 8

    async_result: Optional[AccuracyResult] = asyncio.run(
        async_evaluation.arun(print_results=True)
    )
    assert async_result is not None and async_result.avg_score >= 8

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno openai

Export your OpenAI API key

export OPENAI_API_KEY="your_openai_api_key_here"

Check every requested iteration

Replace the complete if __name__ == "__main__": block with:

accuracy completeness checks
if __name__ == "__main__":
    result = evaluation.run(print_results=True)
    assert result is not None, "No sync evaluation result"
    assert len(result.results) == evaluation.num_iterations, "Incomplete sync evaluation"
    assert all(item.score is not None for item in result.results)
    assert result.avg_score is not None and result.avg_score >= 8

    async_result = asyncio.run(async_evaluation.arun(print_results=True))
    assert async_result is not None, "No async evaluation result"
    assert len(async_result.results) == async_evaluation.num_iterations, "Incomplete async evaluation"
    assert all(item.score is not None for item in async_result.results)
    assert async_result.avg_score is not None and async_result.avg_score >= 8

These assertions are development checks; run Python without -O so they remain enabled.

Run the example

Save the code above as accuracy_basic.py, then run:

python accuracy_basic.py

Full source: cookbook/09_evals/accuracy/accuracy_basic.py