Evaluation47 of 5720 tools
47 tools filed under Evaluation by what their own homepage shows, ranked by who uses them among those still maintained.
A tag is one of thirty-odd words from a closed list, chosen by the same reader from the same page as the line — never a keyword the vendor supplied.
Ranked by use
Mercor
mercor.comConnects enterprises with expert networks to evaluate AI models, generate training data, and deploy AI agents.
Arize
arize.comObserves agent behavior, evaluates performance with evals, and suggests improvements through automated debugging.
Braintrust
braintrust.devObservability platform that traces agent execution, evaluates outputs, and discovers production patterns to improve quality.
micro1
micro1.aiProvides expert human data, evaluation benchmarks, and reinforcement learning environments for training and testing frontier AI models and agents.
Galileo
galileo.aiCaptures groundtruth data, builds and auto-tunes evaluations, and distills them into production guardrails for AI agents and RAG systems.
Patronus
patronus.aiGenerates simulated digital environments and benchmarks for training and evaluating AI agents.
FetchSandbox MCP
fetchsandbox.comRuns deterministic tests for coding agents by verifying API integrations fail before fixes and pass after.
Coasty
coasty.aiEvaluates computer use agents on real operating systems with real software, graded on task outcomes.
Clinician AI Assist
thepromoter.comOrganizes patient findings and simulates clinical trajectories to support diagnostic reasoning for healthcare professionals.
SearchQ
searchq.comRoutes questions to optimal AI models, fact-checks answers across multiple models, and synthesizes verified responses with configurable privacy levels.
Hlido
hlido.euTests AI tools by hand against vendor claims and publishes scored evidence.
Hemelion
hemelion.comGuides decisions between options and tests whether repeated responses fit new situations.
Megaton
megaton.aiBenchmarks and ranks video generation models across quality dimensions like physics, animation, and prompt adherence.
oqoqo
oqoqo.aiBuilds and runs evaluations for AI agents on real-world tasks in sandboxed cloud environments.
Agent Arena
arena42.aiHosts competitive benchmarking tournaments where autonomous AI agents compete on real-world tasks and earn rewards.
crixpix
crixpix.comReviews and compares AI tools with hands-on testing, benchmarks, and pricing breakdowns.
ProveIt Hiring Innovations Inc.
proveit.meSends code review and system design assessments to engineering candidates, scores submissions against role-specific rubrics.
GlobalMatch
globalmatch.techConducts video interviews tailored to candidate resumes and provides explainable skill scores with reskilling recommendations.
ClientCoded
clientcoded.comTests conversational agents with adversarial scenarios and monitors production conversations in real time.
kodwai
kodwai.comScores developers on how well they direct AI coding agents through real challenges, not test passage.
Agent Checker
agentchecker.aiAudits websites to identify where AI agents fail to complete tasks like search, signup, and checkout.
Promptyx
promptyx.techOrganizes, versions, and evaluates prompts and workflows across multiple language models.
CriteriaBot
criteriabot.ioEvaluates content against custom criteria using a consensus panel of AI models, returning true/false verdicts.
Aequitas AI
aequitas.worldVerifies AI decisions in regulated industries by checking claims mathematically and certifying verdicts or declining when confidence is insufficient.
InterviewSkool
interviewskool.comNo description yet.
EasyEnv
easyenv.ioEvaluates engineering candidates through live and take-home coding interviews in real production environments.
AgentVet.ai
agentvet.aiDiscovers, benchmarks, and rates AI agents through independent testing and user reviews.
Atom Foundry
atomfoundry.devMeasures how AI shopping agents discover, evaluate, and recommend e-commerce stores across product categories.
SecNav
secnavpro.comVerifies security professionals through high-fidelity mission simulations and forensic telemetry tracking.
SmartAssess
smartassess.inConducts, evaluates, and ranks job candidates through automated AI interviews with real-time analytics and anti-cheating detection.
CAFE
cafe-ai.deRuns factorial experiments on RAG agents and LLM chains to identify which components drive quality.
Orinyx
orinyx.ioVerifies clinical AI medication recommendations against FDA labeling and clinical pharmacology before and after deployment.
Roleplay
roleplay.shTests AI agents for vulnerability to social engineering attacks through simulated pressure, authority, and policy bypass scenarios.
TrueCode
truecode.co.inEvaluates coding ability by monitoring real debugging work in an IDE, scoring judgment and verification rather than just test results.
Redline AI
tryredlineai.coRed teams AI agents before release and blocks prompt injection and tool abuse in production.
CodeVerdict
codeverdict.ioEvaluates take-home coding assessments by running submitted code in a sandbox and scoring against requirements.
Agentic Diaries
agenticdiaries.comMeasures gaps between what AI agents know internally and what they communicate to users.
Alva Labs
alvalabs.ioEvaluates job candidates using structured assessments, psychometric tests, and interview tools to standardize hiring decisions.
since 2017
Arena AI: The Official AI Ranking & LLM Leaderboard
arena.aiRanks and compares large language models through crowdsourced blind battles.
Botate
botate.botStructures debates between AI models to reach decisions through analyst, critic, and judge roles.
Empromptu AI
empromptu.aiBuilds production-ready AI features for regulated industries with accuracy monitoring and automatic model improvement.
Eval-X
eval-x.comEvaluates software engineers on real-world coding tasks in full IDEs while monitoring their AI interactions and decision-making patterns.
Hard Look
gethardlook.comSubmits project descriptions to domain experts who identify risks and flaws through structured interrogation.
Hume EVI 2
hume.aiEvaluates voice AI models by measuring emotional expression, naturalness, and human preference across 48 emotions and 50 languages.
AgentsFrom $3/moLangWatch
langwatch.aiTests and evaluates AI agents through simulation, tracing, and scoring before production deployment.
PromptQuorum
promptquorum.comSends one prompt to 25 AI models simultaneously and scores responses for consensus and hallucination risk.
Savyre
savyre.comEvaluates job candidates through real-world coding assessments with AI scoring and proctoring.
How we read a tool
No votes, no reviews, no vendor claims. Every week a crawler reads each tool's own site and a few public registries, and the words on the card are bands over what it read.
- Use: visits to the site (estimated), installs from npm and PyPI, presence in Chrome's usage report.
- Activity: the newest release on GitHub, npm or PyPI; the newest dated page on the site; open roles on a public jobs board.
- Price: the vendor's own pricing page, read with the date. A figure is printed only when it is on that page.
- The line and the use cases are written by us from the homepage, one row at a time, and refused when they repeat the vendor's marketing.
How AI assistants read each site — the reading vendors ask us about — is on each tool's own visibility page.