Skip to content

Data9 companies

9 companies, 9 swept. The median identity score is 90 of 100. Rows are ordered by readiness; rows not yet rated sit last.

9 companies · 9 swept · last sweep 2026-09-16

Identity scores in Dataswept9 swept

90/100

median

0identity score, 0-100100
9 swept

Ranked by readiness

  • #1

    Apify

    apify.com

    Marketplace of web scraping and data extraction tools for AI agents and applications.

    bots in llms.txt

    86/100

    #1 of 9 in Data
  • #2

    Diffbot

    diffbot.com

    Extracts structured data from web pages and crawls entire sites into datasets without hallucinations.

    bots in llms.txt

    83/100

    #2 of 9 in Data
  • #3

    Bright Data

    brightdata.com

    Provides web scraping APIs, proxy infrastructure, and datasets for extracting structured data from websites at scale.

    bots partly llms.txt

    82/100

    #3 of 9 in Data
  • #4

    Browserbase

    browserbase.com

    Provides browser automation and web interaction APIs for agents to navigate, extract data, and complete tasks on websites.

    bots in llms.txt

    74/100

    #4 of 9 in Data
  • #5

    Zyte

    zyte.com

    Provides web scraping APIs and managed data extraction services with proxy rotation and ban handling.

    bots partly llms.txt

    69/100

    #5 of 9 in Data
  • #6

    Tavily

    tavily.com

    Provides real-time web search and content extraction API for AI agents to access and reason over live web data.

    bots partly llms.txt

    67/100

    #6 of 9 in Data
  • #7

    SerpApi

    serpapi.com

    Provides structured data from search engines and shopping sites via API.

    bots partly llms.txt

    66/100

    #7 of 9 in Data
  • #8

    Exa

    exa.ai

    Provides web search API and data indexing for AI agents with real-time web crawling and structured data extraction.

    bots partly llms.txt

    61/100

    #8 of 9 in Data
  • #9

    Firecrawl

    firecrawl.dev

    Searches, scrapes, and interacts with websites to provide clean data for AI agents.

    bots partly llms.txt

    47/100

    #9 of 9 in Data

How to read a row

swept
Read by a crawler: HTTP and parsing, no model call. Every row, every week.
measured
The full analyzer run, asked of four engines. Shown only where the company agreed to show it.

Can a retrieval engine find, resolve and quote this site? Four parts, each read by the sweep, each shown beside the total.

Entity 25 · Machine-readable 35 · Retrieval access 20 · Footprint 20

A part that was never read is left out and the total is rescaled to what was. It is never counted as zero.

Measure your own site