Evals

Evaluate agents and teams for accuracy, model-judged quality, performance, reliability, and reusable suites.

ExampleDescription
AccuracyAccuracy examples evaluate how well responses match expected outputs.
Agent As JudgeAgent-as-judge examples evaluate output quality with model-based scoring.
PerformancePerformance examples benchmark runtime and memory impact for agents and teams.
ReliabilityReliability examples validate whether expected expected tool executions are recorded.
SuiteDeclare reusable eval cases and run them together with the built-in suite CLI.