Skip to content

Оценка моделей45 из 5723 инструментов

45 инструментов с тегом «Оценка моделей» по тому, что показывает их главная страница, по тому, кто ими пользуется, среди тех, что продолжают развивать.

Тег - одно из трёх десятков слов закрытого списка; его выбирает тот же читатель по той же странице, что и строку описания. Никогда - ключевое слово от вендора.

По использованию

По тегу
  • Mercor

    mercor.com

    Connects enterprises with expert networks to evaluate AI models, generate training data, and deploy AI agents.

    Устоявшийсяс 2023

    #149 из 3952
  • Arize

    arize.com

    Observes agent behavior, evaluates performance with evals, and suggests improvements through automated debugging.

    Устоявшийся

    #174 из 3952
  • Braintrust

    braintrust.dev

    Observability platform that traces agent execution, evaluates outputs, and discovers production patterns to improve quality.

    Устоявшийсяс 2022

    #262 из 3952
  • micro1

    micro1.ai

    Provides expert human data, evaluation benchmarks, and reinforcement learning environments for training and testing frontier AI models and agents.

    HRПо запросу

    Устоявшийся

    #273 из 3952
  • Galileo

    galileo.ai

    Captures groundtruth data, builds and auto-tunes evaluations, and distills them into production guardrails for AI agents and RAG systems.

    Устоявшийся

    #285 из 3952
  • Patronus

    patronus.ai

    Generates simulated digital environments and benchmarks for training and evaluating AI agents.

    Устоявшийся

    #399 из 3952
  • FetchSandbox MCP

    fetchsandbox.com

    Runs deterministic tests for coding agents by verifying API integrations fail before fixes and pass after.

    УстоявшийсяАктивно развивается

    #453 из 3952
  • Coasty

    coasty.ai

    Evaluates computer use agents on real operating systems with real software, graded on task outcomes.

    Набирает аудиториюАктивно развивается

    #650 из 3952
  • Clinician AI Assist

    thepromoter.com

    Organizes patient findings and simulates clinical trajectories to support diagnostic reasoning for healthcare professionals.

    Набирает аудиторию

    #805 из 3952
  • SearchQ

    searchq.com

    Routes questions to optimal AI models, fact-checks answers across multiple models, and synthesizes verified responses with configurable privacy levels.

    БезопасностьЕсть бесплатный план

    Набирает аудиториюАктивно развивается

    #1029 из 3952
  • Hlido

    hlido.eu

    Tests AI tools by hand against vendor claims and publishes scored evidence.

    Набирает аудиториюАктивно развивается

    #1080 из 3952
  • Hemelion

    hemelion.com

    Guides decisions between options and tests whether repeated responses fit new situations.

    Набирает аудиториюПоддерживается

    #1215 из 3952
  • Megaton

    megaton.ai

    Benchmarks and ranks video generation models across quality dimensions like physics, animation, and prompt adherence.

    Набирает аудиториюАктивно развивается

    #1239 из 3952
  • oqoqo

    oqoqo.ai

    Builds and runs evaluations for AI agents on real-world tasks in sandboxed cloud environments.

    Набирает аудиторию

    #1287 из 3952
  • Agent Arena

    arena42.ai

    Hosts competitive benchmarking tournaments where autonomous AI agents compete on real-world tasks and earn rewards.

    Набирает аудиториюАктивно развивается

    #1339 из 3952
  • crixpix

    crixpix.com

    Reviews and compares AI tools with hands-on testing, benchmarks, and pricing breakdowns.

    Набирает аудиториюАктивно развивается

    #1380 из 3952
  • ProveIt Hiring Innovations Inc.

    proveit.me

    Sends code review and system design assessments to engineering candidates, scores submissions against role-specific rubrics.

    HRБесплатно

    Набирает аудиторию

    #1481 из 3952
  • GlobalMatch

    globalmatch.tech

    Conducts video interviews tailored to candidate resumes and provides explainable skill scores with reskilling recommendations.

    Набирает аудиториюАктивно развивается

    #1519 из 3952
  • ClientCoded

    clientcoded.com

    Tests conversational agents with adversarial scenarios and monitors production conversations in real time.

    Набирает аудиторию

    #1649 из 3952
  • kodwai

    kodwai.com

    Scores developers on how well they direct AI coding agents through real challenges, not test passage.

    Набирает аудиториюАктивно развивается

    #1750 из 3952
  • Agent Checker

    agentchecker.ai

    Audits websites to identify where AI agents fail to complete tasks like search, signup, and checkout.

    Набирает аудиториюПоддерживается

    #2497 из 3952
  • Promptyx

    promptyx.tech

    Organizes, versions, and evaluates prompts and workflows across multiple language models.

    Набирает аудиториюАктивно развивается

    #2578 из 3952
  • CriteriaBot

    criteriabot.io

    Evaluates content against custom criteria using a consensus panel of AI models, returning true/false verdicts.

    Набирает аудиторию

    #2653 из 3952
  • Aequitas AI

    aequitas.world

    Verifies AI decisions in regulated industries by checking claims mathematically and certifying verdicts or declining when confidence is insufficient.

    Набирает аудиториюАктивно развивается

    #2868 из 3952
  • InterviewSkool

    interviewskool.com

    Описания пока нет.

    Набирает аудиториюПоддерживается

    #2944 из 3952
  • EasyEnv

    easyenv.io

    Evaluates engineering candidates through live and take-home coding interviews in real production environments.

    Набирает аудиториюАктивно развивается

    #2984 из 3952
  • AgentVet.ai

    agentvet.ai

    Discovers, benchmarks, and rates AI agents through independent testing and user reviews.

    Набирает аудиториюАктивно развивается

    #3159 из 3952
  • Atom Foundry

    atomfoundry.dev

    Measures how AI shopping agents discover, evaluate, and recommend e-commerce stores across product categories.

    Набирает аудиториюАктивно развивается

    #3262 из 3952
  • SecNav

    secnavpro.com

    Verifies security professionals through high-fidelity mission simulations and forensic telemetry tracking.

    Набирает аудиториюАктивно развивается

    #3286 из 3952
  • SmartAssess

    smartassess.in

    Conducts, evaluates, and ranks job candidates through automated AI interviews with real-time analytics and anti-cheating detection.

    Набирает аудиторию

    #3525 из 3952
  • CAFE

    cafe-ai.de

    Runs factorial experiments on RAG agents and LLM chains to identify which components drive quality.

    Набирает аудиторию

    #3743 из 3952
  • Orinyx

    orinyx.io

    Verifies clinical AI medication recommendations against FDA labeling and clinical pharmacology before and after deployment.

    Набирает аудиториюАктивно развивается

    #3753 из 3952
  • Roleplay

    roleplay.sh

    Tests AI agents for vulnerability to social engineering attacks through simulated pressure, authority, and policy bypass scenarios.

    Набирает аудиториюАктивно развивается

    #3758 из 3952
  • TrueCode

    truecode.co.in

    Evaluates coding ability by monitoring real debugging work in an IDE, scoring judgment and verification rather than just test results.

    Набирает аудиторию

    #3767 из 3952
  • Redline AI

    tryredlineai.co

    Red teams AI agents before release and blocks prompt injection and tool abuse in production.

    Набирает аудиториюАктивно развивается

    #3775 из 3952
  • CodeVerdict

    codeverdict.io

    Evaluates take-home coding assessments by running submitted code in a sandbox and scoring against requirements.

    Набирает аудиториюАктивно развивается

    #3883 из 3952
  • Alva Labs

    alvalabs.io

    Evaluates job candidates using structured assessments, psychometric tests, and interview tools to standardize hiring decisions.

    с 2017

  • Arena AI: The Official AI Ranking & LLM Leaderboard

    arena.ai

    Ranks and compares large language models through crowdsourced blind battles.

  • Botate

    botate.bot

    Structures debates between AI models to reach decisions through analyst, critic, and judge roles.

    ещё не прочитан

  • Empromptu AI

    empromptu.ai

    Builds production-ready AI features for regulated industries with accuracy monitoring and automatic model improvement.

  • Eval-X

    eval-x.com

    Evaluates software engineers on real-world coding tasks in full IDEs while monitoring their AI interactions and decision-making patterns.

  • Hard Look

    gethardlook.com

    Submits project descriptions to domain experts who identify risks and flaws through structured interrogation.

    ещё не прочитан

  • Hume EVI 2

    hume.ai

    Evaluates voice AI models by measuring emotional expression, naturalness, and human preference across 48 emotions and 50 languages.

    АгентыОт $3/мес

    ещё не прочитан

  • LangWatch

    langwatch.ai

    Tests and evaluates AI agents through simulation, tracing, and scoring before production deployment.

  • Savyre

    savyre.com

    Evaluates job candidates through real-world coding assessments with AI scoring and proctoring.

Как мы читаем инструмент

Без голосов, без отзывов, без заявлений вендора. Каждую неделю краулер читает собственный сайт инструмента и несколько публичных реестров, а слова на карточке - полосы над тем, что он прочитал.

  • Использование: визиты на сайт (оценка), установки из npm и PyPI, присутствие в отчёте Chrome об использовании.
  • Активность: последний релиз на GitHub, npm или PyPI; самая свежая датированная страница сайта; открытые вакансии на публичной доске.
  • Цена: собственная страница цен вендора, прочитанная с датой. Число печатается, только если оно есть на той странице.
  • Строку и сценарии пишем мы сами по главной странице, по одной на строку, и отбраковываем, когда они повторяют маркетинг вендора.

Как AI-ассистенты читают каждый сайт - чтение, о котором нас спрашивают вендоры, - на отдельной странице видимости каждого инструмента.