FlowRunner
PricingContact
Theme
Start Free

Bright Data

Developer Tools

Collect web data at scale through Bright Data. Agents trigger Web Scraper dataset collections and download result snapshots, fetch bot-protected pages through the Web Unlocker, and inspect the proxy and unlocker zones on the account.

6 actions API key available
The weekly competitor price check comes due on schedule
Trigger Dataset Collection starts the product dataset run with the tracked URL list as input records
Get Collection Progress polls the snapshot until the status reads ready, or surfaces a failure early
Download Snapshot retrieves the records and archives the raw data to [S3](/integrations/s3) before the retention window closes
The agent diffs prices and availability against last week's archive and flags moves past the threshold
The pricing owner reviews the flagged changes before any repricing decision propagates
The digest with the week's significant movers posts to the pricing channel in [Slack](/integrations/slack)

What This Integration Enables

Bright Data is web data collection at industrial scale, and the operating model matters more than the scale: datasets are managed scrapers, so the endless maintenance war against changing markup belongs to the platform, not to your team. FlowRunner agents drive the whole lifecycle: trigger collection runs with input records or discovery modes, poll progress, download completed snapshots, and route the records to storage and analysis. The Web Unlocker covers the point retrievals, single pages that block ordinary requests, and zone inspection keeps the account's routing auditable. Scale is also exactly why collection runs get scoped by a [human-in-the-loop](/concepts/human-in-the-loop) before they spend. - Trigger Web Scraper dataset runs from known URL lists or discovery modes - Poll collection progress and download completed snapshots as JSON records - Fetch bot-protected pages through Web Unlocker zones, raw or parsed - Audit the proxy and unlocker zones configured on the account - Keep a durable archive of every collection, because a snapshot that expired is a bill with no deliverable

Without FlowRunner

Scrapers rot In-house parsers break on every markup change, and fixing them is nobody's job
Blocked pages end the story A CAPTCHA or bot wall turns a data question into an engineering project
Data dies in snapshots Results sit undownloaded until the retention window quietly deletes them

With FlowRunner

Collection is a managed dataset You send input records; maintaining the extraction against site changes is Bright Data's problem
Protected pages are reachable The Web Unlocker handles proxies, CAPTCHAs, and bot detection behind one request
Snapshots archive themselves Flows download and store every completed run to durable storage on completion

Use Case Scenarios

Marketplace monitoring that survives site redesigns

The team tracks listings across a marketplace that redesigns quarterly and blocks casual scrapers year-round. The agent triggers the listings dataset weekly with the watch list as input records, polls to ready, downloads the snapshot, and loads the records into [Google Sheets](/integrations/google-sheets) for the category managers while the raw archive lands in [S3](/integrations/s3). When the site changes its markup, the dataset keeps returning the same structured fields, which is the entire point of paying for managed collection. The category managers see prices and stock states in columns they recognize, on the schedule they chose, without ever learning what a selector is.

Mapping a market with discovery mode, deliberately

Instead of collecting known URLs, a Discover By run finds records by category or keyword, which means the input does not bound the output. The agent stages the discovery parameters, and a person approves the scope before Trigger Dataset Collection fires, because a discovery run's record count, and its bill, is decided by what exists out there, not by what you listed. Results download on completion, dedupe against prior snapshots, and feed the market map the strategy team maintains. Each quarter's run compares against the last, so the map shows not just who exists but who appeared, who vanished, and who changed category.

The single page that will not load

A compliance check needs one competitor's terms page, and it sits behind bot detection. The agent resolves the right zone with List Active Zones, fetches the page through Send Unlocker Request, and posts the extracted content to the requesting channel in [Slack](/integrations/slack) with the retrieval timestamp. List Snapshots keeps the collection history auditable, so when someone asks where a number came from six weeks later, the answer is a snapshot id with a timestamp, not a shrug. Point retrievals stay rare and deliberate, which is how they should be.

Human-in-Loop Highlight

Trigger Dataset Collection in a Discover By mode is spend with an open upper bound: the run collects whatever matches, billed per record, and a broad category can return an order of magnitude more records than anyone budgeted. So discovery runs never start on an agent's judgment alone; the scope goes to a person with the parameters laid out and the prior run's record count beside them for comparison, and the run fires on approval. The other clock in this connector is quieter: snapshots are retained for a limited window, about 16 days, and a collection nobody downloads is money spent on data that deletes itself. FlowRunner flows close that gap structurally, archiving every ready snapshot to durable storage as a workflow step rather than a human memory, so the spend always produces an artifact someone can use later.

Agent processes routinely
Detects exception requiring judgment
Clear match Continues automatically
Ambiguous Routes to human via preferred channel
Human decides
Agent resumes with decision

Agent Capabilities

6 actions

Web Scraper

4
  • Trigger Dataset Collection Starts an asynchronous collection run for a dataset, taking an array of input records matching the dataset's expected shape, or a Discover By mode to find records instead of collecting known URLs. Returns the snapshot id that keys the rest of the lifecycle.
  • Get Collection Progress Returns a snapshot's current status by id, moving from running to ready or failed. The polling step between trigger and download.
  • Download Snapshot Downloads a completed snapshot's records as JSON. Retention is limited to roughly 16 days, so flows archive on ready rather than on request.
  • List Snapshots Lists a dataset's snapshots with ids, status, and creation times. The audit trail over what was collected and when, and the recovery path to a run's id after the fact.

Web Unlocker

1
  • Send Unlocker Request Fetches a target page through the Web Unlocker, which handles proxies, CAPTCHAs, and bot detection automatically. Returns raw HTML by default or parsed data when a structured format is selected. The tool for pages that refuse standard requests.

Zones

1
  • List Active Zones Lists the account's active zones with name and type, unblocker, datacenter, or residential. The lookup behind Send Unlocker Request and the audit view of how the account routes traffic.

Frequently Asked Questions

What can FlowRunner do with Bright Data?

FlowRunner agents can run Trigger Dataset Collection, Get Collection Progress, and Download Snapshot in Bright Data, plus 3 more actions.

Does connecting Bright Data to FlowRunner require OAuth?

No. Bright Data connects to FlowRunner with an API key, no OAuth flow required.

Can Bright Data trigger a FlowRunner workflow automatically?

Bright Data doesn't currently expose triggers in FlowRunner. It connects as an action step inside workflows started by another trigger.

Start building with Bright Data

$100 in credits. No card required. Connect in minutes.