ScrapeGraph
ScrapeGraphTools enable an Agent to extract structured data from webpages, convert content to markdown, and retrieve raw HTML content.
ScrapeGraphTools enable an Agent to extract structured data from webpages, convert content to markdown, and retrieve raw HTML content using the ScrapeGraphAI API.
The toolkit provides 5 core capabilities:
- smartscraper: Extract structured data using natural language prompts
- markdownify: Convert web pages to markdown format
- searchscraper: Search the web and extract information
- crawl: Crawl websites with structured data extraction
- scrape: Get raw HTML content from websites
The scrape method is particularly useful when you need:
- Complete HTML source code
- Raw content for further processing
- HTML structure analysis
- Content that needs to be parsed differently
All methods support heavy JavaScript rendering when needed.
Prerequisites
Install Agno with the ScrapeGraphAI and OpenAI clients:
uv pip install -U agno scrapegraph-py openaiGet API keys from ScrapeGraphAI and OpenAI, then set both environment variables:
export SGAI_API_KEY="your_scrapegraph_api_key_here"
export OPENAI_API_KEY="your_openai_api_key_here"Example
The following agent will extract structured data from a website using the smartscraper tool:
from agno.agent import Agent
from agno.models.openai import OpenAIResponses
from agno.tools.scrapegraph import ScrapeGraphTools
agent_model = OpenAIResponses(id="gpt-5.2")
scrapegraph_smartscraper = ScrapeGraphTools(enable_smartscraper=True)
agent = Agent(
tools=[scrapegraph_smartscraper], model=agent_model, markdown=True, stream=True
)
agent.print_response("""
Use smartscraper to extract the following from https://www.wired.com/category/science/:
- News articles
- Headlines
- Images
- Links
- Author
""")Raw HTML Scraping
Get complete HTML content from websites for custom processing:
# Enable scrape method for raw HTML content
scrapegraph_scrape = ScrapeGraphTools(enable_scrape=True, enable_smartscraper=False)
scrape_agent = Agent(
tools=[scrapegraph_scrape],
model=agent_model,
markdown=True,
stream=True,
)
scrape_agent.print_response(
"Use the scrape tool to get the complete raw HTML content from https://en.wikipedia.org/wiki/2025_FIFA_Club_World_Cup"
)All Functions with JavaScript Rendering
Enable all ScrapeGraph functions with heavy JavaScript support:
# Enable all ScrapeGraph functions
scrapegraph_all = Agent(
tools=[
ScrapeGraphTools(all=True, render_heavy_js=True)
], # render_heavy_js=True scrapes all JavaScript
model=agent_model,
markdown=True,
stream=True,
)
scrapegraph_all.print_response("""
Use any appropriate scraping method to extract comprehensive information from https://www.wired.com/category/science/:
- News articles and headlines
- Convert to markdown if needed
- Search for specific information
""")Toolkit Params
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key | Optional[str] | None | ScrapeGraph API key. If not provided, uses SGAI_API_KEY environment variable. |
enable_smartscraper | bool | True | Enable the smartscraper function for LLM-powered data extraction. |
enable_markdownify | bool | False | Enable the markdownify function for webpage to markdown conversion. |
enable_crawl | bool | False | Enable the crawl function for website crawling and data extraction. |
enable_searchscraper | bool | False | Enable the searchscraper function for web search and information extraction. |
enable_scrape | bool | False | Enable the scrape function for retrieving raw HTML content from websites. |
render_heavy_js | bool | False | Enable heavy JavaScript rendering for all scraping functions. Useful for SPAs and dynamic content. |
headers | Optional[Dict[str, str]] | None | Custom HTTP headers applied to every tool call. |
crawl_poll_interval | int | 3 | Seconds between crawl status polls. |
crawl_max_wait | int | 180 | Max seconds to wait for a crawl to complete. |
all | bool | False | Enable all available functions. When True, all enable flags are ignored. |
Toolkit Functions
| Function | Description |
|---|---|
smartscraper | Extract structured data from a webpage using LLM and natural language prompt. Parameters: url (str), prompt (str). |
markdownify | Convert a webpage to markdown format. Parameters: url (str). |
crawl | Crawl a website and extract structured data across multiple pages. Parameters: url (str), prompt (str), schema (dict), max_depth (int), max_pages (int). |
searchscraper | Search the web and extract information. Parameters: query (str). |
scrape | Get raw HTML content from a website. Useful for complete source code retrieval and custom processing. Parameters: url (str). |