Audio to Text

Transcribe an MP3 conversation with Gemini, labeling each speaker in the output.

Download an MP3, transcribe it with Gemini, and label each speaker in the streamed response.

audio_to_text.py
"""
Audio To Text
=============================

Audio To Text.
"""

import requests
from agno.agent import Agent
from agno.media import Audio
from agno.models.google import Gemini

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
agent = Agent(
    model=Gemini(id="gemini-3.5-flash"),
    markdown=True,
)

url = "https://agno-public.s3.us-east-1.amazonaws.com/demo_data/QA-01.mp3"

response = requests.get(url)
audio_content = response.content

# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    # Give a transcript of this audio conversation. Use speaker A, speaker B to identify speakers.

    agent.print_response(
        "Give a transcript of this audio conversation. Use speaker A, speaker B to identify speakers.",
        audio=[Audio(content=audio_content)],
        stream=True,
    )

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno google-genai requests

Export your Google API key

export GOOGLE_API_KEY="your_google_api_key_here"

Run the example

Save the code above as audio_to_text.py, then run:

python audio_to_text.py

Full source: cookbook/02_agents/12_multimodal/audio_to_text.py