✍️Writing & Content29🎨Image Generation35🎬Video & Animation70🎵Audio & Music53💬Chatbots & Assistants46💻Coding & Development222📈Marketing & SEO70Productivity173🎯Design & UI/UX62📊Data & Analytics54📚Education & Research28💼Business & Finance65🏥Healthcare & Wellness18🔍Search & Knowledge15🤖AI Agent Infrastructure85🛡️AI Security & Testing11🧊3D & Spatial19🔎SEO Tools18🏡Real Estate4🗃️Data Extraction17🧠ADHD & Focus Tools9
Listed in AI Agent Infrastructure with 88 other toolsPart of 1248+ curated AI tools on AISO
EvalsHub logo

EvalsHub

LLM-as-a-judge evaluation platform with custom rubrics, regression catching, red-teaming and CI/CD integration.

freemiumStarter free at $0/mo with 5,000 trace spans, 3 projects, 3 experiments, 1 GB storage and 7-day retention. Pro at $39/mo with 50,000 spans, unlimited experiments and projects, 10 GB storage, unlimited retention, CI/CD integration, red-team testing, A/B prompt tests, auto evals and custom judges. Enterprise is custom, quoted on the page as typically $500–$3,000+/mo.View full pricing →

Visit EvalsHub

https://evalshub.ai

About EvalsHub

EvalsHub is an AI quality-assurance platform built around LLM-as-a-judge scoring, aimed at teams still catching regressions through manual spot-checks. You define rubrics as natural-language criteria with weights and thresholds — accuracy matched against ground truth, hallucination held above a confidence bar — and evaluations run continuously against your data, comparing models and flagging regressions before a release rather than after a user finds them. Results are deterministic scores rather than impressions, which is the stated point: the site frames it as bringing traditional engineering rigour to generative output, so you can compare GPT-, Claude- and Llama-family responses on the same rubric and see which passed and which hallucinated. Alongside evaluation there is an adversarial testing surface that red-teams the model automatically: heuristic and LLM-based detection of prompt injection hidden in user input, stress testing against evolving persona-based jailbreaks and DAN-style bypasses, and verification of content filtering, PII leakage and internal policy compliance. Tracing, datasets and experiments are the underlying units — spans, AI-generated dataset rows, experiments and projects — and CI/CD integration puts the whole thing in the release path. Pricing is published in full: a genuinely usable free tier, a $39/mo Pro tier that unlocks red-teaming, A/B prompt tests, online auto-evals and custom LLM judges, and a scoped enterprise tier.

Key Features

Natural-language rubrics with weights and thresholds
LLM-as-a-judge scoring tailored to specific use cases
Automatic regression detection and cross-model comparison
Red-team suite for prompt injection, jailbreaks and PII leakage
CI/CD integration and online auto-evals
AI-generated dataset rows and trace-span based experiments

Tags

llm-evalsllm-as-judgered-teamingprompt-testingobservability
🏷️

Is this your tool?

Claim your listing to get a Featured badge, edit your description, and stand out from competitors. All plans include a permanent dofollow backlink to your site.

Claim Now →

ChatGPT already recommends EvalsHub. Does it recommend yours?

If you're building in AI Agent Infrastructure, run a free AI-visibility scan on your own product — we ask ChatGPT across 5 prompt angles and score how often you get named. ~30 seconds, no signup, no card.

Stay updated on AI Agent Infrastructure tools — join our weekly newsletter

One concise email with fresh launches, trending picks, and featured standouts.

Agent connectivity: not yet verified