Audio To Text
Demonstrates team-based audio transcription and follow-up content analysis.
"""
Audio To Text
=============================
Demonstrates team-based audio transcription and follow-up content analysis.
"""
import requests
from agno.agent import Agent
from agno.media import Audio
from agno.models.google import Gemini
from agno.team import Team
# ---------------------------------------------------------------------------
# Create Members
# ---------------------------------------------------------------------------
transcription_specialist = Agent(
name="Transcription Specialist",
role="Convert audio to accurate text transcriptions",
model=Gemini(id="gemini-3.5-flash"),
instructions=[
"Transcribe audio with high accuracy",
"Identify speakers clearly as Speaker A, Speaker B, etc.",
"Maintain conversation flow and context",
],
)
content_analyzer = Agent(
name="Content Analyzer",
role="Analyze transcribed content for insights",
model=Gemini(id="gemini-3.5-flash"),
instructions=[
"Analyze transcription for key themes and insights",
"Provide summaries and extract important information",
],
)
# ---------------------------------------------------------------------------
# Create Team
# ---------------------------------------------------------------------------
audio_team = Team(
name="Audio Analysis Team",
model=Gemini(id="gemini-3.5-flash"),
members=[transcription_specialist, content_analyzer],
instructions=[
"Work together to transcribe and analyze audio content.",
"Transcription Specialist: First convert audio to accurate text with speaker identification.",
"Content Analyzer: Analyze transcription for insights and key themes.",
],
markdown=True,
)
# ---------------------------------------------------------------------------
# Run Team
# ---------------------------------------------------------------------------
if __name__ == "__main__":
url = "https://agno-public.s3.us-east-1.amazonaws.com/demo_data/QA-01.mp3"
response = requests.get(url)
audio_content = response.content
audio_team.print_response(
"Give a transcript of this audio conversation. Use speaker A, speaker B to identify speakers.",
audio=[Audio(content=audio_content)],
stream=True,
)Before running
Add response.raise_for_status() after the download. The sample is MP3; Audio(content=audio_content, mime_type="audio/mp3") makes its format explicit. The leader chooses how to delegate transcription and analysis, and the prompt’s requested speaker labels are not guaranteed diarization output.
Run the Example
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno google-genai requestsExport your Google API key
export GOOGLE_API_KEY="your_google_api_key_here"Run the example
Save the code above as audio_to_text.py, then run:
python audio_to_text.pyFull source: cookbook/03_teams/19_multimodal/audio_to_text.py