NVIDIA Basic

Run Mistral Nemotron through the NVIDIA API in sync, async, and streaming modes.

NVIDIA marks the free endpoint for Llama 3.3 70B as deprecated. The current instructions use Mistral Nemotron, which lists an available free endpoint and function calling. You still need API account access. This does not describe partner endpoints or downloaded Llama weights.

basic.py
"""
Nvidia Basic
============

Cookbook example for `nvidia/basic.py`.
"""

from agno.agent import Agent, RunOutput  # noqa
from agno.models.nvidia import Nvidia
import asyncio

# ---------------------------------------------------------------------------
# Create Agent
# ---------------------------------------------------------------------------

agent = Agent(model=Nvidia(id="meta/llama-3.3-70b-instruct"), markdown=True)

# Get the response in a variable
# run: RunOutput = agent.run("Share a 2 sentence horror story")
# print(run.content)

# Print the response in the terminal

# ---------------------------------------------------------------------------
# Run Agent
# ---------------------------------------------------------------------------
if __name__ == "__main__":
    # --- Sync ---
    agent.print_response("Share a 2 sentence horror story")

    # --- Sync + Streaming ---
    agent.print_response("Share a 2 sentence horror story", stream=True)

    # --- Async ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story"))

    # --- Async + Streaming ---
    asyncio.run(agent.aprint_response("Share a 2 sentence horror story", stream=True))

Run the Example

Set up your virtual environment

uv venv --python 3.12
source .venv/bin/activate

Install dependencies

uv pip install -U agno openai

Export your NVIDIA API key

export NVIDIA_API_KEY="your_nvidia_api_key_here"

Use the current hosted model

Replace id="meta/llama-3.3-70b-instruct" with id="mistralai/mistral-nemotron" in the saved file. Keep the NVIDIA API key and existing run options.

Run the example

Save the code above as basic.py, then run:

python basic.py

Full source: cookbook/90_models/nvidia/basic.py