Audio to Text
Transcribe an MP3 conversation with Gemini, labeling each speaker in the output.
Download an MP3, transcribe it with Gemini, and label each speaker in the streamed response.
"""
Audio To Text
=============================
Audio To Text.
"""
import requests
from agno.agent import Agent
from agno.media import Audio
from agno.models.google import Gemini
# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
agent = Agent(
model=Gemini(id="gemini-3.5-flash"),
markdown=True,
)
url = "https://agno-public.s3.us-east-1.amazonaws.com/demo_data/QA-01.mp3"
response = requests.get(url)
audio_content = response.content
# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
# Give a transcript of this audio conversation. Use speaker A, speaker B to identify speakers.
agent.print_response(
"Give a transcript of this audio conversation. Use speaker A, speaker B to identify speakers.",
audio=[Audio(content=audio_content)],
stream=True,
)Run the Example
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno google-genai requestsExport your Google API key
export GOOGLE_API_KEY="your_google_api_key_here"Run the example
Save the code above as audio_to_text.py, then run:
python audio_to_text.pyFull source: cookbook/02_agents/12_multimodal/audio_to_text.py