ScrapeGraphAI
ScrapeGraphAI is an AI-powered web scraping API for developers and data teams: describe the data you want in a natural-language prompt and get clean structured JSON back, with no CSS selectors, no proxies and no scraper maintenance to carry.
What is ScrapeGraphAI?
ScrapeGraphAI is a web data API for developers. You pass a URL and a natural-language instruction, an LLM reads the page, and the fields you asked for come back as JSON. There are no CSS selectors to write, no XPath to maintain, no proxy pool to run and no headless browser to operate.
Since the V2 rewrite the product is organised around five verbs. Scrape returns page formats: markdown, HTML, summary, links, images, JSON, branding, screenshot. Extract returns structured JSON from a URL, an HTML string or a Markdown document, and accepts a JSON Schema — a Pydantic model, for instance — when the output shape must stay stable. Search runs a web query and extracts structured data from the results in one call. Crawl walks a site or a section as an asynchronous job with progress tracking. Monitor checks a page on a cron schedule, stores the diffs and fires a webhook on change.
A single request can ask for several formats at once instead of scraping the same page five times. JavaScript pages, single-page applications and dynamically loaded content are handled natively. A fetchConfig object exposes the levers when a target resists: auto, fast or js rendering mode, wait time, scrolls, headers, cookies, timeout, country routing and a stealth toggle. Most requests are announced to complete in two to ten seconds depending on complexity, and frequently requested URLs are cached.
What comes back is an API response, not a hosted dashboard export you download by hand. That framing is deliberate: the tool is built for backend jobs, notebooks, agents, ETL pipelines and internal tools. The integration surface is wide — Python and JavaScript SDKs, a just-scrape CLI, an official MCP server, plus LangChain, LangGraph, CrewAI, LlamaIndex, Agno, Vercel AI SDK, LiteLLM, n8n, Zapier and Make.
The project sits on an open-source base: the Scrapegraph-ai repository claims 27.3k GitHub stars, and the home page advertises 250M+ pages extracted and 1M+ users. A SOC 2 Type 2 audit was announced completed on 3 July 2026. Advertised use cases span price monitoring, lead generation, market research, real-estate tracking and AI agent tooling.
What it does
- Extract structured JSON from a URL, raw HTML or Markdown using a natural-language prompt, with an optional JSON Schema to lock the output shape
- Turn any page into clean Markdown, HTML, a summary, links, images, screenshots or a branding analysis, several formats in a single call
- Search the web and extract structured data from the results in one request, with geographic targeting
- Crawl an entire site or a section and return the requested formats page by page, as an asynchronous job with progress tracking
- Monitor a page on a cron schedule, record the diffs and fire a webhook when something changes
- Process PDFs page by page, with a configurable page cap
- Give an AI agent live web access through the official MCP server, a CLI or the Python and JavaScript SDKs
When to use ScrapeGraphAI / When not to
A quick filter to help you decide if ScrapeGraphAI is the right fit.
When to use ScrapeGraphAI
- Backend developers who need to feed an application or a production pipeline with web data through a single API call, without running headless browsers or proxy pools.
- Data and ETL engineers, who use crawl and extract to load warehouses, vector databases and RAG ingestion pipelines.
- AI agent builders, served by the official MCP server, the Python and JavaScript SDKs and the LangChain, LangGraph, CrewAI, LlamaIndex, Agno and Vercel AI SDK integrations.
- E-commerce and pricing teams running the advertised price-monitoring use case across Amazon, eBay and Shopify stores, with scheduled checks and webhook alerts on price or stock drops.
- Market and competitive intelligence analysts aggregating reviews, ratings and customer sentiment across many sites, alongside sales teams building lead lists and real-estate teams tracking Zillow, Redfin and local listing sites.
When not to use ScrapeGraphAI
- Anyone whose target is a private page they hold no valid access to: the vendor says so itself, and the terms forbid bypassing logins, paywalls, CAPTCHAs, rate limits, robots restrictions or anti-bot systems without authorisation.
- Teams that only need a one-off manual copy of a handful of pages, where an API key and a billing account cost more effort than the job is worth.
- Teams that already run a stable in-house scraper over a small fixed set of sources, for whom AI extraction adds spend without removing maintenance.
- Non-technical users expecting a point-and-click interface or a one-click spreadsheet export: there is no mobile app, and the output is an API response rather than a downloadable dashboard export.
- Organisations that would keep confidential or regulated personal data on the free plan, since free-plan data feeds model research and training and the only opt-out is upgrading to a paid tier. The services are also not intended for anyone under 18.
How to use ScrapeGraphAI
A typical end-to-end flow, from setup to results.
- Create an account through the sign-up page: 500 credits are granted, with no credit card required
- Copy your API key from the dashboard; it travels in the SGAI-APIKEY HTTP header
- Pick the right service: scrape when you already know the URL, extract when the output must keep a stable shape, search when the source URLs are unknown, crawl to walk a site, monitor to watch it over time
- Install an SDK (pip install scrapegraph-py or npm install scrapegraph-js), call the API with cURL, or wire up the just-scrape CLI
- For scrape, pass the URL and the list of formats you want (markdown, links, screenshot, json and so on) in a single call
- For extract, pass a natural-language prompt and, if the shape matters, a JSON Schema such as a Pydantic model
- Touch fetchConfig only when the target demands it (js mode, wait, scrolls, country, stealth): the vendor advises starting from the default fetch to keep latency and credit spend under control
- For monitor, supply the URL, a cron expression and a webhookUrl; the service records the ticks and the diffs
- Estimate spend with the credit calculator on the home page or the standalone price calculator before scaling up
- For an AI agent, install the official MCP server (Claude Desktop, Claude Code, Codex, Smithery) or go no-code through n8n, Zapier or Make
Pros & Cons
Pros
- No CSS or XPath selectors to write and maintain — the AI adapts when a layout changes — and no proxies, rate limiting or scraping infrastructure to host
- One call can return several formats at once, instead of scraping the same page repeatedly
- Credit pricing published in full, down to the per-endpoint formula, with two online calculators to model spend before committing
- Permanent free plan of 500 credits with no credit card, plus one-time credit packs that never expire and stack on top of a subscription
- Broad integration ecosystem: MCP, LangChain, LangGraph, CrewAI, LlamaIndex, Agno, Vercel AI SDK, LiteLLM, n8n, Zapier and Make
- SOC 2 Type 2 completed, with the report available on request for procurement reviews
- Visible open-source base on GitHub (27.3k stars), a public API status page and a dense changelog of 131 dated entries from January 2024 to July 2026
Cons
- Free-plan data becomes training data: prompts, target content, outputs, logs and feedback feed model research, evaluation, training and fine-tuning. Section 9 of the terms names the only opt-out as stopping use of the free plan and moving to a paid tier — it is not a product setting.
- No postal address is published anywhere: no legal notice, no address in the terms, a contact page reduced to an email and a Calendly link, and a footer limited to a copyright line.
- Section 18 never names the governing law or the forum, referring only to “the jurisdiction where ScrapeGraphAI, Inc. is established”. With no address published, you cannot know which court would hear a dispute — and liability is capped at the greater of three months of fees and USD 100.
- The privacy policy is dated 15 March 2024, fifteen months before the terms of 4 June 2026, and reads as a generic template: it announces the collection of passport details, marital status and social security numbers, along with advertising cookies and remarketing, none of which matches the API product the terms describe.
- No Article 27 EU representative, no Data Protection Officer, no published subprocessor list, no named transfer mechanism, and no retention period expressed in figures — the policy says data is kept “as long as necessary” and stops there.
- Legal responsibility for scraping is transferred wholesale to the customer: the vendor states it gives no advice on whether a given target may lawfully be scraped, and the terms restate the point in three separate places.
- Developer-only surface: no consumer interface, no mobile app, and the output is an API response. No interface or processing language is declared anywhere and the site is English only. Subscription credits expire under plan terms, do not roll over between cycles unless separately agreed in writing, and are non-refundable.
Pricing & Plans
ScrapeGraphAI operates a permanent free plan of 500 one-time API credits, with no credit card required. The lowest paid entry point is the Starter plan at USD 20 per month, or USD 204 per year (equivalent to USD 17 per month). One-time credit packs are sold separately from USD 5 for 1,000 credits, and Enterprise pricing is quoted on request.
- 500 one-time API credits
- 10 requests per minute
- 1 monitor
- 1 concurrent crawl. No credit card required.
- 10
- 000 credits per month
- 100 requests per minute
- 5 monitors
- 3 concurrent crawls. Published unit cost: USD 0.0020 per scrape and USD 0.0100 per AI extraction.
- 100
- 000 credits per month
- 500 requests per minute
- 25 monitors
- 15 concurrent crawls
- basic proxy rotation. Published unit cost: USD 0.0010 per scrape and USD 0.0050 per AI extraction.
- 750
- 000 credits per month
- 5
- 000 requests per minute
- 100 monitors
- 50 concurrent crawls
- advanced proxy rotation and priority support. Published unit cost: USD 0.0007 per scrape and USD 0.0033 per AI extraction.
- ad hoc credits
- custom limits
- dedicated support and an SLA guarantee.
- Small at USD 5 for 1
- 000 credits (USD 5.00 per 1
- 000)
- Medium at USD 40 for 10
- 000 credits (USD 4.00 per 1
- 000)
- Large at USD 150 for 50
- 000 credits (USD 3.00 per 1
- scrape costs 1 credit for markdown
- 2 for a screenshot and 25 for branding
- extract costs 5 plus the stealth modifier
- search costs 2 per result without a prompt and 5 with one
- crawl costs 2 to start plus the per-page scrape cost
- monitor costs the format price plus 5 credits when a change is detected
- PDFs cost 1 credit per processed page
- with a 25-page default cap.
- the rendering mode (auto
- fast or js) does not change the price
- while the stealth toggle adds 5 credits.
Data, GDPR & hosting
A consolidated view of how ScrapeGraphAI handles your data.
GDPR overview
The privacy policy carries a named GDPR section for EU and EEA residents, listing access, rectification, erasure, objection, restriction, portability and withdrawal of consent, exercised via support@scrapegraphai.com. ScrapeGraphAI positions itself as Data Controller and its providers as Data Processors, and states that personal data is transferred to and processed in the United States, acceptance of the policy standing as agreement to that transfer. The gaps are substantial: no Article 27 EU representative, no named Data Protection Officer, no transfer mechanism cited — neither Standard Contractual Clauses nor the Data Privacy Framework — and no published subprocessor list, the terms redirecting you to contact the vendor for the DPA, enterprise privacy terms and subprocessor details. The policy is dated 15 March 2024, fifteen months before the terms of 4 June 2026. CalOPPA and CCPA sections exist, and the vendor says it neither sells nor rents personal data.
Who owns the data?
The terms define three layers. Customer Data — your prompts, URLs, schemas, files, credentials and configuration — stays yours; Target Content is the third-party material fetched on your instruction; Outputs are the summaries, extracted fields, screenshots, markdown, JSON, links and images returned to you. ScrapeGraphAI claims ownership of neither your Customer Data nor your Outputs. You grant it the rights it needs to host, process, transmit, display, troubleshoot, secure and improve the service and to provide support. The service gives you no ownership over third-party websites, their data or their content. Conversely, ScrapeGraphAI, Inc. and its licensors keep the service itself: software, APIs, workflows, models, interfaces, documentation and trademarks. Feedback you send may be used freely, without compensation and without confidentiality.
Reuse rights
You may reuse Outputs without asking ScrapeGraphAI for permission: the contract reserves no ownership over them for the vendor. What the contract does not grant is any right in the underlying source. Section 4 places on you every right, permission, notice, consent and legal basis for the sites, URLs, data and instructions you submit, together with compliance with applicable law, third-party terms of service, robots.txt and other machine-readable restrictions, intellectual property, privacy rights, contractual restrictions and platform rules. ScrapeGraphAI states plainly that it provides no legal advice on whether a given scraping, crawling, monitoring or extraction activity is permitted. Outputs may also be incomplete, inaccurate, delayed or based on third-party content that changed without notice, so section 3 requires you to review them before relying on them, publishing them, or feeding them into commercial, legal, financial, medical, employment, housing, credit or insurance decisions. Two hard limits remain under section 5: you may not resell or sublicense the service in a way that obscures end-user responsibility, and you may not use it to build a competing service by copying its APIs, systems, documentation or user experience.
Data retention & training
Hosting summary
ScrapeGraphAI hosts and processes data in the United States. The privacy policy states that information, including personal data, may be transferred to and maintained on computers located outside your jurisdiction, and that for anyone located outside the United States the data is transferred to the United States and processed there; consenting to the policy and then submitting information counts as agreement to that transfer. Beyond the country, nothing is named: no region, no data centre, no cloud provider and no EU data-residency option. No transfer mechanism is cited, neither Standard Contractual Clauses nor the Data Privacy Framework, and no subprocessor list is published — the terms direct you to contact the vendor for those details. What is documented is the control layer: the SOC 2 Type 2 scope announced on 3 July 2026 covers encryption of data in transit and at rest, cloud configuration, network controls, logging and monitoring, under a shared-responsibility model. For reference only, the domain's IP address (216.24.57.1) geolocates to San Francisco, United States, on AS397273 Render; that is an infrastructure observation with anycast flagged on the record, not a declaration by the vendor.
Things to keep in mind
Risks and trade-offs to weigh before adopting ScrapeGraphAI.
- Free-plan data becomes training data. Prompts, target content, outputs, logs and feedback submitted on the free plan feed model research, evaluation, training and fine-tuning, and the terms themselves forbid submitting confidential information or regulated personal data there. It is easy to prototype on the free tier and forget what went through it.
- The tool lowers the cost of scraping so far that mass collection of public personal data becomes trivially cheap. Lead generation — LinkedIn profiles, Twitter accounts, contact details at scale — is a use case advertised on the home page. Legality and proportionality are yours to judge, every single time.
- Compliance with robots.txt and with target sites' terms is left entirely to you, while the service ships a stealth toggle capable of getting past anti-bot protections, something the terms contractually forbid without authorisation. The technical capability and the contractual permission do not overlap.
- The vendor gives no opinion on whether a given scrape is lawful, so the legal risk transfers wholesale to the customer — and in a dispute you would face no published postal address, no named jurisdiction and a liability cap set at the greater of three months of fees and USD 100.
- Outputs can be incomplete, inaccurate or stale. The terms explicitly warn against relying on them without review in credit, employment, housing, insurance or health decisions, which are exactly the settings where an unreviewed JSON field does the most damage.
- LLM-driven extraction can drift silently. When a page changes, a selector-based scraper breaks loudly; an AI extractor may keep returning plausible but wrong values with no parsing error to warn you, and automation bias makes that failure mode expensive.
- Personal data travels to and is processed in the United States with no transfer mechanism named and no EU representative appointed, and no retention period is published for any category of data.
Setup & Integrations
Technical difficulty
Moderate, and squarely technical. This is an API product: the shortest path is to create an account, copy an API key and run a three-line cURL, and the vendor advertises setup in five minutes with no proxies, no selectors and no infrastructure to configure. Official Python (scrapegraph-py) and JavaScript (scrapegraph-js) SDKs and a just-scrape CLI cover developers. Non-developers have a genuine no-code route through the official n8n, Zapier and Make connectors, and agent builders install the MCP server in one command. Advanced fetchConfig settings (js mode, wait, scrolls, cookies, country, stealth) are optional and only needed for stubborn targets.
Deployment
Integrations
Behind ScrapeGraphAI
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement ScrapeGraphAI.
Frequently asked questions
What exactly does ScrapeGraphAI do?
Is there a free plan, and what does the cheapest paid plan cost?
Is my data used to train models?
Where is my data hosted?
Does it handle JavaScript pages and single-page applications?
How fast is it?
Is there an MCP server for AI agents?
What security and contractual documents can I obtain?
Is there an offer for startups?
How do I contact the company, and is there a minimum age?
Should you pick ScrapeGraphAI?
ScrapeGraphAI is a technically accomplished, well-documented product. Five clearly separated services — scrape, extract, search, crawl and monitor — sit behind one API; pricing is published down to the per-endpoint credit formula; two calculators let you model spend before committing; the changelog runs to 131 dated entries between January 2024 and July 2026; there is a public API status page; and a SOC 2 Type 2 audit was announced completed on 3 July 2026. The positioning is owned, not apologised for: a developer tool, not a consumer interface. Value for money reads directly in the published unit costs, from USD 0.0020 per scrape on Starter down to USD 0.0007 on Pro, with AI extraction between USD 0.0100 and USD 0.0033, and a permanent free plan of 500 credits makes evaluation cheap. Two reservations deserve to travel with that verdict. First, product transparency does not extend to corporate transparency: no postal address anywhere on the site, no incorporation date, a governing law described only as the jurisdiction where the company is established without ever naming it, and a generic privacy policy from March 2024 that does not line up with terms of June 2026. Second, the free plan is paid for in data, since free-plan prompts, target content and outputs feed model research and training and the only way out is a paid tier. Legal responsibility for scraping also stays entirely with the customer, a point the terms restate three times while the vendor declines to advise on whether a given target may lawfully be scraped. For a technical team with its own view on source legality, ScrapeGraphAI is a strong and honestly priced option. For a procurement function that needs a named jurisdiction, an Article 27 representative and figures for data retention, the paperwork is not there yet.
- Choosing a selection results in a full page refresh.
- Opens in a new window.