Retrieval (RAG)16 companies
16 companies, 16 swept. The median identity score is 83 of 100. Rows are ordered by readiness; rows not yet rated sit last.
16 companies · 16 swept · last sweep 2026-09-16
83/100
median
Ranked by readiness
- #1
Cohere
cohere.comProvides generative language models and retrieval optimization for enterprise AI applications.
bots partly✓ llms.txt92/100
#1 of 16 in Retrieval (RAG) - #2
Unstructured
unstructured.ioConverts unstructured documents into clean structured data for AI applications.
bots in— llms.txt87/100
#2 of 16 in Retrieval (RAG) - #3
LangChain
langchain.comBuilds and deploys agents with observability, evaluation, and monitoring tools.
bots in— llms.txt81/100
#3 of 16 in Retrieval (RAG) - #4
Pinecone
pinecone.ioStores and retrieves vector embeddings at scale for retrieval-augmented generation and agent knowledge systems.
bots partly✓ llms.txt76/100
#4 of 16 in Retrieval (RAG) - #5
Qdrant
qdrant.techStores and searches vectors at scale with metadata filtering, hybrid search, and reranking.
bots in✓ llms.txt76/100
#5 of 16 in Retrieval (RAG) - #6
Jina AI
jina.aiConverts URLs to markdown and provides multimodal embeddings and reranking for search and retrieval systems.
bots in— llms.txt74/100
#6 of 16 in Retrieval (RAG) - #7
Reducto
reducto.aiParses documents into structured data with layout preservation, table extraction, and bounding box citations.
bots in✓ llms.txt74/100
#7 of 16 in Retrieval (RAG) - #8
LlamaIndex
llamaindex.aiParses and extracts structured data from complex documents using agentic OCR and layout-aware processing.
bots in— llms.txt70/100
#8 of 16 in Retrieval (RAG) - #9
Zilliz
zilliz.comManages vector data for AI applications with real-time search, discovery, and analytics on a single platform.
bots partly✓ llms.txt70/100
#9 of 16 in Retrieval (RAG) - #10
Chiri
chiri.aiEmbeds engineering teams to build production AI applications and agents integrated with your existing software and data.
bots in✓ llms.txt67/100
#10 of 16 in Retrieval (RAG) - #11
turbopuffer
turbopuffer.comVector and full-text search database built on object storage, scaling to 256TB per index with sub-10ms latency.
bots in✓ llms.txt63/100
#11 of 16 in Retrieval (RAG) - #12
Weaviate
weaviate.ioVector database that stores, indexes, and searches high-dimensional vectors for retrieval-augmented generation and semantic search.
bots partly✓ llms.txt63/100
#12 of 16 in Retrieval (RAG) - #13
Vectara
vectara.comBuilds retrieval-augmented generation agents with policy enforcement and hallucination detection across on-premises, VPC, and SaaS deployments.
bots partly— llms.txt59/100
#13 of 16 in Retrieval (RAG) - #14
Voyage AI
voyageai.comGenerates embeddings and reranks search results to improve retrieval quality for unstructured data.
bots in— llms.txt49/100
#14 of 16 in Retrieval (RAG) - #15
Chroma
chroma.comManufactures optical filters and beamsplitters for imaging, spectroscopy, and scientific applications.
bots partly— llms.txt40/100
#15 of 16 in Retrieval (RAG) - #16
LlamaParse
llamaparse.comNo line written yet.
bots in— llms.txt36/100
#16 of 16 in Retrieval (RAG)
How to read a row
- swept
- Read by a crawler: HTTP and parsing, no model call. Every row, every week.
- measured
- The full analyzer run, asked of four engines. Shown only where the company agreed to show it.
Can a retrieval engine find, resolve and quote this site? Four parts, each read by the sweep, each shown beside the total.
Entity 25 · Machine-readable 35 · Retrieval access 20 · Footprint 20
A part that was never read is left out and the total is rescaled to what was. It is never counted as zero.