NVIDIA Basic
Run Mistral Nemotron through the NVIDIA API in sync, async, and streaming modes.
NVIDIA marks the free endpoint for Llama 3.3 70B as deprecated. The current instructions use Mistral Nemotron, which lists an available free endpoint and function calling. You still need API account access. This does not describe partner endpoints or downloaded Llama weights.
"""
Nvidia Basic
============
Cookbook example for `nvidia/basic.py`.
"""
from agno.agent import Agent, RunOutput # noqa
from agno.models.nvidia import Nvidia
import asyncio
# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------
agent = Agent(model=Nvidia(id="meta/llama-3.3-70b-instruct"), markdown=True)
# Get the response in a variable
# run: RunOutput = agent.run("Share a 2 sentence horror story")
# print(run.content)
# Print the response in the terminal
# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
# --- Sync ---
agent.print_response("Share a 2 sentence horror story")
# --- Sync + Streaming ---
agent.print_response("Share a 2 sentence horror story", stream=True)
# --- Async ---
asyncio.run(agent.aprint_response("Share a 2 sentence horror story"))
# --- Async + Streaming ---
asyncio.run(agent.aprint_response("Share a 2 sentence horror story", stream=True))Run the Example
Set up your virtual environment
uv venv --python 3.12
source .venv/bin/activateInstall dependencies
uv pip install -U agno openaiExport your NVIDIA API key
export NVIDIA_API_KEY="your_nvidia_api_key_here"Use the current hosted model
Replace id="meta/llama-3.3-70b-instruct" with id="mistralai/mistral-nemotron" in the saved file. Keep the NVIDIA API key and existing run options.
Run the example
Save the code above as basic.py, then run:
python basic.pyFull source: cookbook/90_models/nvidia/basic.py