Multimodal Agents
Build agents that process and generate images, audio, video, and files.
Agno agents support multimodal input and output using text, image, audio, video and files.
Guides
Image As Input
Analyze and describe images with agents.
Image As Output
Return generated images from agent responses.
Image to Text
Convert input image to text.
OpenAI Image Generation
Generate images with OpenAI tool.
Image Generation
Legacy DalleTools example for image generation.
Image Analysis in Same Run
Legacy DalleTools example for generation and same-run analysis.
Image Analysis in Multi-turn Runs
Legacy DalleTools example for generation and multi-turn analysis.
Image I/O with Fal API
Use input image and Fal API to generate new images.
Image to Structured Output
Convert input image to structured output using Pydantic models.
Generate Image with Intermediate Steps
Legacy DalleTools example with intermediate run events.
High Fidelity Image Analysis
Analyze images with high fidelity.
Image to Audio
Convert input image to audio.
Image input for Tools
Legacy example that passes uploaded and DALL-E-generated images to tools.
Audio As Input
Analyze and understand audio input with agents.
Audio As Output
Return audio responses from agents.
Audio I/O
Use audio as input and output in agents.
Generate Music
Generate classical music using agents.
Speech-to-Text
Transcribe audio conversations with Gemini.
OpenAI Speech-to-Text
Transcribe local audio with OpenAI tools.
Audio Generation
Generate speech and sound effects with ElevenLabs.
Multi-turn Audio
Multi-turn audio conversation with AI models.
Audio Streaming
Stream audio responses from agents.
Audio Sentiment Analysis
Analyze sentiment of audio using agents.
Convert Blog to Podcast
Convert blog to podcast using agents.
Video Input
Analyze and understand video input with agents.
Video Output
Generate video output using FAL.
Generate Video Captions
Use video as input to generate captions.
Generate Shorts
Generate Shorts from Video.
Generate Video with Replicate
Generate video using Replicate.
Generate Video with ModelsLab
Generate video using ModelsLab.