Diffbot
Extracts structured data from web pages and crawls entire sites into datasets without hallucinations.
Diffbot is a developer of machine learning and computer vision algorithms and public APIs for extracting data from web pages / web scraping to create a knowledge base.First sentence of the English Wikipedia article
83/100
- Entity25/25
- Machine-readable21/35
- Retrieval access20/20
- Footprintnot read
- Diffbot83
- Data medianof 9 rated69
- Atlas medianof 537 rated61
- Retrieval botsopenverify
Whether the bots that decide citations are let in. Training bots are a separate list and are not counted here.
- llms.txtyesverify
A plain-text index at /llms.txt. Optional; read by some tools.
- Organization schemano
schema.org Organization markup on the homepage, which is what a resolver reads a company's name and links from.
- Wikidata itemQ17052069Wikipedia article
Matched on the official website, never on the name. The substrate under the knowledge graph an engine resolves names against.
- Server-rendered words409swept
What a reader without JavaScript receives from the homepage.
- Identity score60/100swept
The analyzer's crawl-only section, 0-100: identity files, schema, crawler policy, sitemap. Not the analyzer score.
- Founded2011P571
- HeadquartersMenlo ParkP159
- IndustryInternetP452
60/100
identity score
82
86
74
69
67
66
Not measured with the analyzer.
The full run asks four engines about this company by name. It runs on request, not on a schedule.
What is Diffbot?
Extracts structured data from web pages and crawls entire sites into datasets without hallucinations.
What is Diffbot's Atlas Readiness?
83 of 100 on 2026-09-15, computed from four parts — entity, machine-readable, retrieval access, footprint — out of 80 of 100 points — footprint not read. It places #2 of 9 rated Data tools.
Does Diffbot let AI retrieval bots in?
Yes. On 2026-09-15, robots.txt at diffbot.com disallowed none of the retrieval crawlers we check.
Does Diffbot publish an llms.txt?
Yes: diffbot.com/llms.txt was present on 2026-09-15.
Is Diffbot a Wikidata entity?
Yes: Q17052069, matched on its official website.
How to read a row
- swept
- Read by a crawler: HTTP and parsing, no model call. Every row, every week.
- measured
- The full analyzer run, asked of four engines. Shown only where the company agreed to show it.
Can a retrieval engine find, resolve and quote this site? Four parts, each read by the sweep, each shown beside the total.
Entity 25 · Machine-readable 35 · Retrieval access 20 · Footprint 20
A part that was never read is left out and the total is rescaled to what was. It is never counted as zero.