Image To Audio

Write a story about a local image, validate the result, then narrate it with gpt-audio and save a WAV.

Image To Audio.

image_to_audio.py
"""
Image To Audio
=============================

Image To Audio.
"""

from pathlib import Path

from agno.agent import Agent, RunOutput
from agno.media import Image
from agno.models.openai import OpenAIChat
from agno.utils.audio import write_audio_to_file
from rich import print
from rich.text import Text

cwd = Path(__file__).parent.resolve()

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
image_agent = Agent(model=OpenAIChat(id="gpt-5.6-luna"))

image_path = Path(__file__).parent.joinpath("sample.jpg")

# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    image_story: RunOutput = image_agent.run(
        "Write a 3 sentence fiction story about the image",
        images=[Image(filepath=image_path)],
    )
    formatted_text = Text.from_markup(
        f":sparkles: [bold magenta]Story:[/bold magenta] {image_story.content} :sparkles:"
    )
    print(formatted_text)

    audio_agent = Agent(
        model=OpenAIChat(
            id="gpt-audio",
            modalities=["text", "audio"],
            audio={"voice": "sage", "format": "wav"},
        ),
    )

    audio_story: RunOutput = audio_agent.run(
        f"Narrate the story with flair: {image_story.content}"
    )
    if audio_story.response_audio is not None:
        write_audio_to_file(
            audio=audio_story.response_audio.content, filename="tmp/sample_story.wav"
        )

Stop the chain on failed output

The archived program continues to narration even when image analysis fails. Add from agno.run.base import RunStatus to its imports, then insert this guard immediately after image_agent.run(...), before printing or narrating:

if (
    image_story.status != RunStatus.completed
    or not isinstance(image_story.content, str)
    or not image_story.content.strip()
):
    raise RuntimeError("Image analysis did not produce a story")

Replace the final if audio_story.response_audio is not None: block with:

if (
    audio_story.status != RunStatus.completed
    or audio_story.response_audio is None
    or not audio_story.response_audio.content
):
    raise RuntimeError("Narration did not produce audio")
write_audio_to_file(
    audio=audio_story.response_audio.content,
    filename="tmp/sample_story.wav",
)

Keep both guards inside the existing if __name__ == "__main__": block. This prevents an error message from being narrated as a story.

OpenAI schedules gpt-audio for shutdown on January 20, 2027. It remains available until that announced date. See the provider lifecycle notice and verify audio formats and streaming support when selecting a replacement.

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno openai

Export your OpenAI API key

export OPENAI_API_KEY="your_openai_api_key_here"

Add the sample image

Place a JPEG named sample.jpg in the same directory as image_to_audio.py.

Run the example

Save the code above as image_to_audio.py, then run:

python image_to_audio.py

Full source: cookbook/02_agents/12_multimodal/image_to_audio.py