ai-evaluation
5 servers · 183★ total
Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
MCP server for human-in-the-loop surveys, A/B preference tests, ratings, and rankings. Get real human feedback inside Claude Code, Claude Desktop, Cursor, Windsurf, and any MCP client — powered by Datapoint AI.
Diagnose your AI agents in production. Extract policies from prompts, evaluate traces, generate diagnostic reports.
Build self-evolving AI agent harnesses with portable harness units, artifact-aware testing, trace-backed diagnosis, and evidence-gated promotion.