Retry
Review retry settings and why invalid model IDs cannot reliably exercise the retry path.
The source imports vLLM, but the current exported class is VLLM; the import fails before a request. This example also assumes an invalid model ID triggers the configured retries. Invalid-model responses commonly use terminal 400 or 404 statuses, which Agno does not retry. Do not run this source as a retry test.
"""Example demonstrating how to set up retries with vLLM."""
from agno.agent import Agent
from agno.models.vllm import vLLM
# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
# We will use a deliberately wrong model ID, to trigger retries.
wrong_model_id = "vllm-wrong-id"
agent = Agent(
model=vLLM(
id=wrong_model_id,
retries=3, # Number of times to retry the request.
delay_between_retries=1, # Delay between retries in seconds.
exponential_backoff=True, # If True, the delay between retries is doubled each time.
),
)
agent.print_response("What is the capital of France?")
# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
passCurrent Alternative
Complete the basic example's two-terminal setup. In the activated client terminal, save this program as vllm_retry.py and run python vllm_retry.py.
SDK retries are disabled. Agno excludes terminal 400/404 errors; test a controlled transient 429, connection failure or 5xx response to exercise retries. A successful call alone does not test the retry path.
from agno.agent import Agent
from agno.models.vllm import VLLM
agent = Agent(
model=VLLM(
id="Qwen/Qwen2.5-7B-Instruct",
client_params={"max_retries": 0},
retries=3,
delay_between_retries=1,
exponential_backoff=True,
)
)
agent.print_response("What is the capital of France?")Full source: cookbook/90_models/vllm/retry.py