FlowRunner
PricingContact
Theme
Start Free

ScrapeGraphAI

AI

Extract structured data from any web page with a plain-English prompt through ScrapeGraphAI's LLM-powered scraping API. Agents enforce JSON output schemas, run search-and-scrape queries across the web, convert pages to LLM-ready Markdown, and poll long-running jobs by request ID.

6 actions API key available
The weekly competitor pricing sweep starts on schedule
Get Credits confirms the account balance covers the full run at AI extraction rates
Smart Scraper extracts plan names, prices, and limits from each tracked pricing page, forced into a JSON Output Schema
Get Smart Scraper Result polls the pages that processed asynchronously until every result lands
The agent diffs the fresh extraction against last week's snapshot and flags every changed value
An analyst confirms the flagged changes against the live pages before anything downstream updates
Confirmed changes write to the tracking sheet and the sales battlecards
A what-moved summary posts to the competitive channel with source URLs attached

What This Integration Enables

ScrapeGraphAI moves web extraction from code to language. The traditional scraper is a maintenance liability: selectors coupled to someone else's markup, breaking without notice, owned by whoever wrote them. ScrapeGraphAI's LLM-driven approach takes a plain-English prompt and an optional JSON schema instead, which means the extraction survives redesigns and the person who maintains it does not need to read HTML. For FlowRunner agents, that turns the public web into one more queryable system. - Smart Scraper pulls structured fields from any page by prompt, with schema-enforced output shapes ready for downstream steps - Search Scraper answers a question across the web in one call, merging results and returning the reference URLs used - Markdownify converts articles and documentation into clean, LLM-ready Markdown for knowledge pipelines - Asynchronous jobs are polled by request ID inside the flow, so downstream steps receive finished data - Get Credits makes spend visible before a large run, and [human-in-the-loop](/concepts/human-in-the-loop) checkpoints sit between extracted data and the records it changes

Without FlowRunner

Scrapers are code that rots Every tracked site means CSS selectors that break silently the next time a div gets renamed
Web research is a browser marathon Competitive and market questions get answered by whoever has an afternoon to burn on tabs
Scraped data goes straight to decisions Extracted numbers flow into sheets unreviewed, and nobody remembers which ones were wrong

With FlowRunner

Extraction is a prompt, not a parser Describe what you want in plain English, constrain it with a schema, and the same request survives redesigns
The web answers on a schedule Search and scrape runs merge structured answers with the reference URLs that produced them
Machine collection, human confirmation Changed values pause for an analyst's check before they touch battlecards or prices

Use Case Scenarios

Competitor pricing that watches itself

The competitive team tracks a dozen pricing pages. Weekly, the agent runs Smart Scraper against each with a schema for plan name, monthly price, and usage limits, then diffs against the stored snapshot in [Google Sheets](/integrations/google-sheets). Unchanged weeks close silently. Changed values post to [Slack](/integrations/slack) with the old number, the new number, and the URL; an analyst confirms before the battlecard updates. Sales walks into calls with pricing that was verified this week, not last quarter.

Account research that arrives before the discovery call

When a new deal reaches qualification in [HubSpot](/integrations/hubspot), the agent runs Search Scraper with a prompt covering the account's funding, headcount signals, tech stack mentions, and recent news, capped at a sensible Number of Results. The merged structured answer, with its reference URLs, is written to the deal record as a research note. The rep spends preparation time reading, not searching, and every claim in the note carries the link it came from.

A knowledge base that ingests the web cleanly

The operations team maintains an internal knowledge base in [Notion](/integrations/notion) that references vendor documentation and industry guidance scattered across the web. The agent runs Markdownify on each source page, polls Get Markdownify Result for the longer ones, and files the clean Markdown, stripped of navigation and ads, into the right section with its source URL. When a source page changes, the refreshed conversion is diffed and queued for an editor's review rather than silently replacing the old version.

Human-in-Loop Highlight

The gate on this page sits between Smart Scraper's output and the systems that act on it, and it exists because LLM extraction fails differently than a broken selector. A dead CSS path returns nothing; a prompt against a redesigned page can return something plausible and wrong, a decoy price from a promo banner, an annual figure read as monthly. When the weekly sweep flags that a competitor's mid tier dropped by a third, the flow does not update the battlecards or trigger the repricing discussion. It posts the old value, the new value, and the source URL, and waits: "Three price changes detected across twelve pages. Confirm against the live pages before publishing?" The analyst clicks through, confirms two, rejects one misread bundle offer. Collection runs at machine scale; the moment extracted data becomes company truth runs through a person. Get Credits closes the loop on the other finite resource, keeping large runs from launching into an empty balance.

Agent processes routinely
Detects exception requiring judgment
Clear match Continues automatically
Ambiguous Routes to human via preferred channel
Human decides
Agent resumes with decision

Agent Capabilities

6 actions

Scraping

5
  • Smart Scraper Extracts structured data from a web page with an LLM-driven extraction prompt. Takes a public Website URL or raw Website HTML plus a natural-language User Prompt, an optional JSON Output Schema to force a fixed response shape, and a Number of Scrolls for infinite-scroll content. Returns the extracted data directly, or a request_id and status for asynchronous pages.
  • Get Smart Scraper Result Retrieves the result of a submitted Smart Scraper job by request ID. The polling step behind pages processed asynchronously.
  • Search Scraper Searches the web for a query and extracts structured data from the top results in a single call, with a Number of Results between 3 and 20. Returns the merged structured answer plus the reference_urls used, optionally schema-constrained.
  • Markdownify Converts a web page into clean, LLM-ready Markdown, stripping navigation, ads, and boilerplate. Returns Markdown directly or a request_id for asynchronous processing.
  • Get Markdownify Result Retrieves the result of a submitted Markdownify job by request ID once processing completes.

Account

1
  • Get Credits Returns the remaining API credits on the account. Credit cost varies by operation and options, with AI extraction and stealth mode consuming more, so flows check this before large runs and alert the team when the balance runs low.

Frequently Asked Questions

What can FlowRunner do with ScrapeGraphAI?

FlowRunner agents can run Smart Scraper, Get Smart Scraper Result, and Search Scraper in ScrapeGraphAI, plus 3 more actions.

Does connecting ScrapeGraphAI to FlowRunner require OAuth?

No. ScrapeGraphAI connects to FlowRunner with an API key, no OAuth flow required.

Can ScrapeGraphAI trigger a FlowRunner workflow automatically?

ScrapeGraphAI doesn't currently expose triggers in FlowRunner. It connects as an action step inside workflows started by another trigger.

Start building with ScrapeGraphAI

$100 in credits. No card required. Connect in minutes.