Image To Audio
Write a story about a local image, validate the result, then narrate it with gpt-audio and save a WAV.
Image To Audio.
"""
Image To Audio
=============================
Image To Audio.
"""
from pathlib import Path
from agno.agent import Agent, RunOutput
from agno.media import Image
from agno.models.openai import OpenAIChat
from agno.utils.audio import write_audio_to_file
from rich import print
from rich.text import Text
cwd = Path(__file__).parent.resolve()
# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
image_agent = Agent(model=OpenAIChat(id="gpt-5.6-luna"))
image_path = Path(__file__).parent.joinpath("sample.jpg")
# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
image_story: RunOutput = image_agent.run(
"Write a 3 sentence fiction story about the image",
images=[Image(filepath=image_path)],
)
formatted_text = Text.from_markup(
f":sparkles: [bold magenta]Story:[/bold magenta] {image_story.content} :sparkles:"
)
print(formatted_text)
audio_agent = Agent(
model=OpenAIChat(
id="gpt-audio",
modalities=["text", "audio"],
audio={"voice": "sage", "format": "wav"},
),
)
audio_story: RunOutput = audio_agent.run(
f"Narrate the story with flair: {image_story.content}"
)
if audio_story.response_audio is not None:
write_audio_to_file(
audio=audio_story.response_audio.content, filename="tmp/sample_story.wav"
)Stop the chain on failed output
The archived program continues to narration even when image analysis fails. Add from agno.run.base import RunStatus to its imports, then insert this guard immediately after image_agent.run(...), before printing or narrating:
if (
image_story.status != RunStatus.completed
or not isinstance(image_story.content, str)
or not image_story.content.strip()
):
raise RuntimeError("Image analysis did not produce a story")Replace the final if audio_story.response_audio is not None: block with:
if (
audio_story.status != RunStatus.completed
or audio_story.response_audio is None
or not audio_story.response_audio.content
):
raise RuntimeError("Narration did not produce audio")
write_audio_to_file(
audio=audio_story.response_audio.content,
filename="tmp/sample_story.wav",
)Keep both guards inside the existing if __name__ == "__main__": block. This prevents an error message from being narrated as a story.
OpenAI schedules gpt-audio for shutdown on January 20, 2027. It remains available until that announced date. See the provider lifecycle notice and verify audio formats and streaming support when selecting a replacement.
Run the Example
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno openaiExport your OpenAI API key
export OPENAI_API_KEY="your_openai_api_key_here"Add the sample image
Place a JPEG named sample.jpg in the same directory as image_to_audio.py.
Run the example
Save the code above as image_to_audio.py, then run:
python image_to_audio.pyFull source: cookbook/02_agents/12_multimodal/image_to_audio.py