Diffbot
AITurn any public web page into structured data with Diffbot. Agents extract clean articles, products, and discussions from URLs, auto-detect page types, query the Knowledge Graph with DQL, and enrich people and companies from a name, URL, or email.
What This Integration Enables
Diffbot's proposition is that the public web is already a database; it just needs a schema imposed on it. The extraction side does that per page: hand an agent a URL and it comes back with a typed object, an article with author and full text, a product with price and availability, a discussion with threaded posts and sentiment, or an auto-detected type when you do not know what you have. The Knowledge Graph side does it at world scale: organizations and people as queryable entities with relationships. FlowRunner agents run both, and the [human-in-the-loop](/concepts/human-in-the-loop) gate sits exactly where probabilistic matching meets your systems of record. - Extract clean articles, products, and discussions from any public URL - Auto-detect unknown page types and route to the right extractor in one call - Query organizations and people with DQL over the Knowledge Graph - Enrich partial leads into full records from a name, URL, email, or phone - Gate every probabilistic match before it becomes a fact in your CRM
Without FlowRunner
With FlowRunner
Use Case Scenarios
Competitor prices, read like data
On a schedule, the agent runs Extract Product against the tracked competitor product pages and logs title, price, currency, and availability to [Google Sheets](/integrations/google-sheets). Extract Discussion runs against the same products' review pages, capturing new reviews with sentiment. Price moves past a threshold and review themes that keep recurring both surface in [Slack](/integrations/slack), so merchandising reacts to the market instead of discovering it quarterly. When a competitor redesigns their product pages, the extractor keeps returning the same fields, which is the difference between typed extraction and a parser someone owns.
The reading pile that files itself
Team members drop article links into a channel all day. The agent picks up each URL, runs Extract Article, and files the clean title, author, date, and full text into [Notion](/integrations/notion), tagged and searchable, with a summary posted back to the thread. Unknown links route through Analyze URL first, so a product page or a forum thread lands in the right shape instead of failing as a non-article. The knowledge base builds itself while the team keeps talking, and the archive is full text, not a graveyard of links that will rot.
A target list defined by criteria, not by a CSV
The ICP is specific: organizations in a named industry above a revenue floor in particular geographies. The agent expresses exactly that as a DQL query through Query Knowledge Graph, pages through the matching entities, and enriches the shortlist with Enhance Entity before writing target accounts to the CRM. When the ICP changes, the list changes by editing a query string, and the delta posts to the sales channel with the criteria that produced it. Disagreements about targeting become disagreements about a visible query, which is a far better argument to have than one about a spreadsheet nobody can reproduce.
Human-in-Loop Highlight
Enhance Entity returns the best match, and best is a probability, not a promise. Feed it a common company name with thin identifying attributes and it can resolve to the wrong organization entirely, at which point the CRM write that follows stamps another company's revenue, size, and industry onto a real lead. Every downstream decision, routing, scoring, the first line of the outreach email, then inherits the error with total confidence and a professional tone. So enrichment flows treat match quality as a gate: attribute agreement gets scored, strong matches flow through, and weak ones stop for a person who can look at the resolved entity next to the original lead and say yes or no. Enrichment at machine speed is the point; being confidently wrong in the CRM at machine speed is the failure mode the gate exists to prevent, and the review takes seconds precisely because the agent has already laid out the evidence.
Agent Capabilities
6 actionsAutomatic Extraction
4- Analyze URL Fetches a public page, auto-detects its type, article, product, discussion, image, event, or list, and routes it to the best-fit extractor. The right first call when the content type is unknown; the detected type comes back with the extracted objects.
- Extract Article Extracts a clean, structured article from a news, blog, or long-form page: title, author, publication date, full text and normalized HTML, tags, top image, and discussion metadata.
- Extract Product Extracts structured product data from an e-commerce page: title, price, currency, availability, brand, SKU, specifications, images, and offer details.
- Extract Discussion Extracts threaded discussions, comment sections, forum threads, and review lists as individual posts with author, date, text, and nested replies, plus aggregate sentiment and participant counts.
Knowledge Graph
2- Query Knowledge Graph Queries the Diffbot Knowledge Graph with DQL to retrieve organizations, people, articles, products, and their relationships, with hit counts and pagination. The primitive that turns an ICP definition into a live account list.
- Enhance Entity Resolves a partially known Person or Organization against the Knowledge Graph from whatever attributes exist, name, URL, email, phone, employer, location, and returns a fully populated entity with firmographics, biographic details, social profiles, and related entities. More input attributes mean better matches, and the weak ones wait for a human.
Frequently Asked Questions
What can FlowRunner do with Diffbot?
FlowRunner agents can run Analyze URL, Extract Article, and Extract Product in Diffbot, plus 3 more actions.
Does connecting Diffbot to FlowRunner require OAuth?
No. Diffbot connects to FlowRunner with an API key, no OAuth flow required.
Can Diffbot trigger a FlowRunner workflow automatically?
Diffbot doesn't currently expose triggers in FlowRunner. It connects as an action step inside workflows started by another trigger.
Start building with Diffbot
$100 in credits. No card required. Connect in minutes.