Models

Provider examples for Agno models, with setup requirements and historical integrations identified.

ExampleDescription
AimlapiAIML API examples for basic runs, multimodal input, memory, retries, structured output, and tool use.
AnthropicClaude examples for multimodal input, context management, caching, knowledge, memory, thinking, structured output, server tools, and skills.
AWSRun Claude and Amazon Nova models on AWS Bedrock.
AzureRun Claude and open-source models on Azure AI Foundry and OpenAI endpoints.
CerebrasCerebras examples: basic runs, storage, knowledge, structured output, retries, and tool use.
Cerebras OpenAIRun Cerebras models through the OpenAI-compatible endpoint: streaming, tools, structured output, storage, and knowledge.
ClientsConfigure Agno's default sync httpx.Client with headers, logging, request IDs, timeouts, and error tracking.
CohereRun Cohere Command A and Aya Vision models with tools, knowledge, memory, retries, and structured output.
CometapiRun GPT, Claude, Gemini, DeepSeek, and Qwen models through CometAPI's OpenAI-compatible gateway.
DashscopeBrowse DashScope model examples with Qwen models, image analysis, knowledge tools, and retry patterns.
DeepInfraDeepInfra examples: basic agent runs, JSON output, tool use, and retries.
DeepSeekRun DeepSeek models with reasoning, thinking mode, structured output, retries, and tool use.
FireworksRun Fireworks models with streaming, structured output, web search, and retry configuration.
GoogleUse Gemini for audio, video, image, PDF, grounding, file search, and thinking-budget examples.
GroqGroq examples for agents and teams, multimodal input, knowledge, reasoning, research, transcription, translation, structured output, and tools.
Hugging FaceHugging Face examples for basic and streaming runs, essay generation, retries, and web-search tool use.
IBMIBM watsonx examples for model retries, storage, knowledge, structured output, and tools.
InternlmRun InternLM models with basic responses, tools, knowledge, storage, retries, and structured output.
LangDBRun LangDB models with basic responses, tools, retries, and structured output.
LiteLLMRun agents through the LiteLLM Python SDK with tools, knowledge, structured output, and audio, image, and PDF input.
LiteLLM OpenAIExamples for LiteLLM with OpenAI-compatible models.
Llama CppRun agents against a local llama.cpp server serving ggml-org/gpt-oss-20b-GGUF at http://127.0.0.1:8080/v1.
LmstudioLM Studio examples for local models, images, knowledge, memory, storage, retries, structured output, and tools.
MetaLlama and Llama OpenAI examples covering tool use, knowledge, memory, metrics, storage, and retries.
MistralRun Mistral models with image input, memory, structured output, retries, and tool use.
MoonshotKimi K3 and K2.6 examples for reasoning, files, schemas and web-search tools.
N1NN1N gateway examples: running OpenAI models via N1N with basic streaming and web-search tool calls.
NebiusNebius model examples: basic runs, Postgres sessions, PgVector knowledge, retries, structured output, and tool use.
NeosantaraNeosantara examples covering basic runs, structured output, and web-search tool use.
NexusNexus examples covering basic runs, retry configuration, and tool use.
NVIDIANVIDIA API examples: basic runs, retry configuration, and tool use.
OllamaOllama Chat and Responses API examples for local and cloud models, knowledge, memory, reasoning, structured output, and tools.
OpenAIOpenAI Chat and Responses API examples for multimodal input, tools, reasoning, structured output, storage, and streaming.
OpenRouterOpenRouter Chat and Responses API examples for model routing, retries, structured output, and tools.
PerplexityIndex of Perplexity sonar-pro agent examples: basic runs, knowledge, memory, retries, structured output, and web search.
PortkeyIndex of Agno examples routing agents through the Portkey AI gateway: basic runs, retries, structured output, and tool use.
RequestyRequesty AI is an LLM gateway with AI governance. See their website for more information.
SambanovaSambaNova model examples: basic sync/stream/async runs and retry configuration.
SiliconflowExamples for SiliconFlow model integration.
TogetherRun Together models with streaming, image input, reasoning, structured output, web search, and retry configuration.
VercelHistorical examples for the removed v0 Model API; current alternatives linked.
Vertex AIVertex AI examples for Claude models, retries, multimodal input, knowledge, memory, caching, structured output, and tools.
vLLMRun local vLLM models with explicit server setup, tools, storage, memory and schemas.
xAIxAI model examples for building agents with Grok, including vision, web search, and financial analysis.
CloudflareCloudflare AI Gateway model examples.
InceptionInception Labs Mercury model examples.
MiniMaxMiniMax M3 examples for basic runs, web search tools, and locally validated JSON.
Xiaomi MiMoXiaomi MiMo model examples.
Ramp RouterResponses model routing, fallback, schemas and tools.
SynthorAIBasic and tool-using agents through SynthorAI.
TokenLabBasic agents through the TokenLab endpoint.
TrustedRouterStreaming agents, schemas and tools through TrustedRouter.
Tuning EnginesRun an agent against a tenant-enabled model deployment.