Skip to content

Evaluation47 of 5720 tools

47 tools filed under Evaluation by what their own homepage shows, ranked by who uses them among those still maintained.

A tag is one of thirty-odd words from a closed list, chosen by the same reader from the same page as the line — never a keyword the vendor supplied.

Ranked by use

By tag
  • Mercor

    mercor.com

    Connects enterprises with expert networks to evaluate AI models, generate training data, and deploy AI agents.

    Establishedsince 2023

    #149 of 3977
  • Arize

    arize.com

    Observes agent behavior, evaluates performance with evals, and suggests improvements through automated debugging.

    LLM operationsFrom $50/mo

    Established

    #174 of 3977
  • Braintrust

    braintrust.dev

    Observability platform that traces agent execution, evaluates outputs, and discovers production patterns to improve quality.

    LLM operationsFrom $249/mo

    Establishedsince 2022

    #262 of 3977
  • micro1

    micro1.ai

    Provides expert human data, evaluation benchmarks, and reinforcement learning environments for training and testing frontier AI models and agents.

    HRContact sales

    Established

    #273 of 3977
  • Galileo

    galileo.ai

    Captures groundtruth data, builds and auto-tunes evaluations, and distills them into production guardrails for AI agents and RAG systems.

    LLM operationsFrom $100/mo

    Established

    #285 of 3977
  • Patronus

    patronus.ai

    Generates simulated digital environments and benchmarks for training and evaluating AI agents.

    LLM operationsFrom $25/mo

    Established

    #399 of 3977
  • FetchSandbox MCP

    fetchsandbox.com

    Runs deterministic tests for coding agents by verifying API integrations fail before fixes and pass after.

    EstablishedActively maintained

    #453 of 3977
  • Coasty

    coasty.ai

    Evaluates computer use agents on real operating systems with real software, graded on task outcomes.

    EmergingActively maintained

    #648 of 3977
  • Clinician AI Assist

    thepromoter.com

    Organizes patient findings and simulates clinical trajectories to support diagnostic reasoning for healthcare professionals.

    Emerging

    #806 of 3977
  • SearchQ

    searchq.com

    Routes questions to optimal AI models, fact-checks answers across multiple models, and synthesizes verified responses with configurable privacy levels.

    SecurityFree plan

    EmergingActively maintained

    #1033 of 3977
  • Hlido

    hlido.eu

    Tests AI tools by hand against vendor claims and publishes scored evidence.

    EmergingActively maintained

    #1084 of 3977
  • Hemelion

    hemelion.com

    Guides decisions between options and tests whether repeated responses fit new situations.

    EmergingMaintained

    #1219 of 3977
  • Megaton

    megaton.ai

    Benchmarks and ranks video generation models across quality dimensions like physics, animation, and prompt adherence.

    EmergingActively maintained

    #1243 of 3977
  • oqoqo

    oqoqo.ai

    Builds and runs evaluations for AI agents on real-world tasks in sandboxed cloud environments.

    Emerging

    #1290 of 3977
  • Agent Arena

    arena42.ai

    Hosts competitive benchmarking tournaments where autonomous AI agents compete on real-world tasks and earn rewards.

    EmergingActively maintained

    #1342 of 3977
  • crixpix

    crixpix.com

    Reviews and compares AI tools with hands-on testing, benchmarks, and pricing breakdowns.

    EmergingActively maintained

    #1384 of 3977
  • ProveIt Hiring Innovations Inc.

    proveit.me

    Sends code review and system design assessments to engineering candidates, scores submissions against role-specific rubrics.

    HRFree

    Emerging

    #1486 of 3977
  • GlobalMatch

    globalmatch.tech

    Conducts video interviews tailored to candidate resumes and provides explainable skill scores with reskilling recommendations.

    EmergingActively maintained

    #1524 of 3977
  • ClientCoded

    clientcoded.com

    Tests conversational agents with adversarial scenarios and monitors production conversations in real time.

    Emerging

    #1655 of 3977
  • kodwai

    kodwai.com

    Scores developers on how well they direct AI coding agents through real challenges, not test passage.

    EmergingActively maintained

    #1758 of 3977
  • Agent Checker

    agentchecker.ai

    Audits websites to identify where AI agents fail to complete tasks like search, signup, and checkout.

    AutomationFrom £39/mo

    EmergingMaintained

    #2516 of 3977
  • Promptyx

    promptyx.tech

    Organizes, versions, and evaluates prompts and workflows across multiple language models.

    EmergingActively maintained

    #2596 of 3977
  • CriteriaBot

    criteriabot.io

    Evaluates content against custom criteria using a consensus panel of AI models, returning true/false verdicts.

    Emerging

    #2672 of 3977
  • Aequitas AI

    aequitas.world

    Verifies AI decisions in regulated industries by checking claims mathematically and certifying verdicts or declining when confidence is insufficient.

    EmergingActively maintained

    #2890 of 3977
  • InterviewSkool

    interviewskool.com

    No description yet.

    EmergingMaintained

    #2966 of 3977
  • EasyEnv

    easyenv.io

    Evaluates engineering candidates through live and take-home coding interviews in real production environments.

    EmergingActively maintained

    #3006 of 3977
  • AgentVet.ai

    agentvet.ai

    Discovers, benchmarks, and rates AI agents through independent testing and user reviews.

    EmergingActively maintained

    #3181 of 3977
  • Atom Foundry

    atomfoundry.dev

    Measures how AI shopping agents discover, evaluate, and recommend e-commerce stores across product categories.

    EmergingActively maintained

    #3284 of 3977
  • SecNav

    secnavpro.com

    Verifies security professionals through high-fidelity mission simulations and forensic telemetry tracking.

    EmergingActively maintained

    #3309 of 3977
  • SmartAssess

    smartassess.in

    Conducts, evaluates, and ranks job candidates through automated AI interviews with real-time analytics and anti-cheating detection.

    Emerging

    #3551 of 3977
  • CAFE

    cafe-ai.de

    Runs factorial experiments on RAG agents and LLM chains to identify which components drive quality.

    Emerging

    #3770 of 3977
  • Orinyx

    orinyx.io

    Verifies clinical AI medication recommendations against FDA labeling and clinical pharmacology before and after deployment.

    EmergingActively maintained

    #3780 of 3977
  • Roleplay

    roleplay.sh

    Tests AI agents for vulnerability to social engineering attacks through simulated pressure, authority, and policy bypass scenarios.

    EmergingActively maintained

    #3785 of 3977
  • TrueCode

    truecode.co.in

    Evaluates coding ability by monitoring real debugging work in an IDE, scoring judgment and verification rather than just test results.

    Emerging

    #3794 of 3977
  • Redline AI

    tryredlineai.co

    Red teams AI agents before release and blocks prompt injection and tool abuse in production.

    EmergingActively maintained

    #3802 of 3977
  • CodeVerdict

    codeverdict.io

    Evaluates take-home coding assessments by running submitted code in a sandbox and scoring against requirements.

    EmergingActively maintained

    #3909 of 3977
  • Agentic Diaries

    agenticdiaries.com

    Measures gaps between what AI agents know internally and what they communicate to users.

    not read yet

  • Alva Labs

    alvalabs.io

    Evaluates job candidates using structured assessments, psychometric tests, and interview tools to standardize hiring decisions.

    since 2017

  • Arena AI: The Official AI Ranking & LLM Leaderboard

    arena.ai

    Ranks and compares large language models through crowdsourced blind battles.

    not read yet

  • Botate

    botate.bot

    Structures debates between AI models to reach decisions through analyst, critic, and judge roles.

    not read yet

  • Empromptu AI

    empromptu.ai

    Builds production-ready AI features for regulated industries with accuracy monitoring and automatic model improvement.

    not read yet

  • Eval-X

    eval-x.com

    Evaluates software engineers on real-world coding tasks in full IDEs while monitoring their AI interactions and decision-making patterns.

    not read yet

  • Hard Look

    gethardlook.com

    Submits project descriptions to domain experts who identify risks and flaws through structured interrogation.

    not read yet

  • Hume EVI 2

    hume.ai

    Evaluates voice AI models by measuring emotional expression, naturalness, and human preference across 48 emotions and 50 languages.

    AgentsFrom $3/mo

    not read yet

  • LangWatch

    langwatch.ai

    Tests and evaluates AI agents through simulation, tracing, and scoring before production deployment.

    not read yet

  • PromptQuorum

    promptquorum.com

    Sends one prompt to 25 AI models simultaneously and scores responses for consensus and hallucination risk.

    not read yet

  • Savyre

    savyre.com

    Evaluates job candidates through real-world coding assessments with AI scoring and proctoring.

    not read yet

How we read a tool

No votes, no reviews, no vendor claims. Every week a crawler reads each tool's own site and a few public registries, and the words on the card are bands over what it read.

  • Use: visits to the site (estimated), installs from npm and PyPI, presence in Chrome's usage report.
  • Activity: the newest release on GitHub, npm or PyPI; the newest dated page on the site; open roles on a public jobs board.
  • Price: the vendor's own pricing page, read with the date. A figure is printed only when it is on that page.
  • The line and the use cases are written by us from the homepage, one row at a time, and refused when they repeat the vendor's marketing.

How AI assistants read each site — the reading vendors ask us about — is on each tool's own visibility page.