Оценка моделей45 из 5723 инструментов
45 инструментов с тегом «Оценка моделей» по тому, что показывает их главная страница, по тому, кто ими пользуется, среди тех, что продолжают развивать.
Тег - одно из трёх десятков слов закрытого списка; его выбирает тот же читатель по той же странице, что и строку описания. Никогда - ключевое слово от вендора.
По использованию
Mercor
mercor.comConnects enterprises with expert networks to evaluate AI models, generate training data, and deploy AI agents.
Arize
arize.comObserves agent behavior, evaluates performance with evals, and suggests improvements through automated debugging.
Braintrust
braintrust.devObservability platform that traces agent execution, evaluates outputs, and discovers production patterns to improve quality.
micro1
micro1.aiProvides expert human data, evaluation benchmarks, and reinforcement learning environments for training and testing frontier AI models and agents.
Galileo
galileo.aiCaptures groundtruth data, builds and auto-tunes evaluations, and distills them into production guardrails for AI agents and RAG systems.
Patronus
patronus.aiGenerates simulated digital environments and benchmarks for training and evaluating AI agents.
FetchSandbox MCP
fetchsandbox.comRuns deterministic tests for coding agents by verifying API integrations fail before fixes and pass after.
Coasty
coasty.aiEvaluates computer use agents on real operating systems with real software, graded on task outcomes.
Clinician AI Assist
thepromoter.comOrganizes patient findings and simulates clinical trajectories to support diagnostic reasoning for healthcare professionals.
SearchQ
searchq.comRoutes questions to optimal AI models, fact-checks answers across multiple models, and synthesizes verified responses with configurable privacy levels.
Hlido
hlido.euTests AI tools by hand against vendor claims and publishes scored evidence.
Hemelion
hemelion.comGuides decisions between options and tests whether repeated responses fit new situations.
Megaton
megaton.aiBenchmarks and ranks video generation models across quality dimensions like physics, animation, and prompt adherence.
oqoqo
oqoqo.aiBuilds and runs evaluations for AI agents on real-world tasks in sandboxed cloud environments.
Agent Arena
arena42.aiHosts competitive benchmarking tournaments where autonomous AI agents compete on real-world tasks and earn rewards.
crixpix
crixpix.comReviews and compares AI tools with hands-on testing, benchmarks, and pricing breakdowns.
ProveIt Hiring Innovations Inc.
proveit.meSends code review and system design assessments to engineering candidates, scores submissions against role-specific rubrics.
GlobalMatch
globalmatch.techConducts video interviews tailored to candidate resumes and provides explainable skill scores with reskilling recommendations.
ClientCoded
clientcoded.comTests conversational agents with adversarial scenarios and monitors production conversations in real time.
kodwai
kodwai.comScores developers on how well they direct AI coding agents through real challenges, not test passage.
Agent Checker
agentchecker.aiAudits websites to identify where AI agents fail to complete tasks like search, signup, and checkout.
Promptyx
promptyx.techOrganizes, versions, and evaluates prompts and workflows across multiple language models.
CriteriaBot
criteriabot.ioEvaluates content against custom criteria using a consensus panel of AI models, returning true/false verdicts.
Aequitas AI
aequitas.worldVerifies AI decisions in regulated industries by checking claims mathematically and certifying verdicts or declining when confidence is insufficient.
InterviewSkool
interviewskool.comОписания пока нет.
EasyEnv
easyenv.ioEvaluates engineering candidates through live and take-home coding interviews in real production environments.
AgentVet.ai
agentvet.aiDiscovers, benchmarks, and rates AI agents through independent testing and user reviews.
Atom Foundry
atomfoundry.devMeasures how AI shopping agents discover, evaluate, and recommend e-commerce stores across product categories.
SecNav
secnavpro.comVerifies security professionals through high-fidelity mission simulations and forensic telemetry tracking.
SmartAssess
smartassess.inConducts, evaluates, and ranks job candidates through automated AI interviews with real-time analytics and anti-cheating detection.
CAFE
cafe-ai.deRuns factorial experiments on RAG agents and LLM chains to identify which components drive quality.
Orinyx
orinyx.ioVerifies clinical AI medication recommendations against FDA labeling and clinical pharmacology before and after deployment.
Roleplay
roleplay.shTests AI agents for vulnerability to social engineering attacks through simulated pressure, authority, and policy bypass scenarios.
TrueCode
truecode.co.inEvaluates coding ability by monitoring real debugging work in an IDE, scoring judgment and verification rather than just test results.
Redline AI
tryredlineai.coRed teams AI agents before release and blocks prompt injection and tool abuse in production.
CodeVerdict
codeverdict.ioEvaluates take-home coding assessments by running submitted code in a sandbox and scoring against requirements.
Alva Labs
alvalabs.ioEvaluates job candidates using structured assessments, psychometric tests, and interview tools to standardize hiring decisions.
с 2017
Arena AI: The Official AI Ranking & LLM Leaderboard
arena.aiRanks and compares large language models through crowdsourced blind battles.
Botate
botate.botStructures debates between AI models to reach decisions through analyst, critic, and judge roles.
Empromptu AI
empromptu.aiBuilds production-ready AI features for regulated industries with accuracy monitoring and automatic model improvement.
Eval-X
eval-x.comEvaluates software engineers on real-world coding tasks in full IDEs while monitoring their AI interactions and decision-making patterns.
Hard Look
gethardlook.comSubmits project descriptions to domain experts who identify risks and flaws through structured interrogation.
Hume EVI 2
hume.aiEvaluates voice AI models by measuring emotional expression, naturalness, and human preference across 48 emotions and 50 languages.
АгентыОт $3/месLangWatch
langwatch.aiTests and evaluates AI agents through simulation, tracing, and scoring before production deployment.
Savyre
savyre.comEvaluates job candidates through real-world coding assessments with AI scoring and proctoring.
Как мы читаем инструмент
Без голосов, без отзывов, без заявлений вендора. Каждую неделю краулер читает собственный сайт инструмента и несколько публичных реестров, а слова на карточке - полосы над тем, что он прочитал.
- Использование: визиты на сайт (оценка), установки из npm и PyPI, присутствие в отчёте Chrome об использовании.
- Активность: последний релиз на GitHub, npm или PyPI; самая свежая датированная страница сайта; открытые вакансии на публичной доске.
- Цена: собственная страница цен вендора, прочитанная с датой. Число печатается, только если оно есть на той странице.
- Строку и сценарии пишем мы сами по главной странице, по одной на строку, и отбраковываем, когда они повторяют маркетинг вендора.
Как AI-ассистенты читают каждый сайт - чтение, о котором нас спрашивают вендоры, - на отдельной странице видимости каждого инструмента.