ai-evaluation

5 servers · 183★ total

Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.

CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.

MCP server for human-in-the-loop surveys, A/B preference tests, ratings, and rankings. Get real human feedback inside Claude Code, Claude Desktop, Cursor, Windsurf, and any MCP client — powered by Datapoint AI.

Diagnose your AI agents in production. Extract policies from prompts, evaluate traces, generate diagnostic reports.

Build self-evolving AI agent harnesses with portable harness units, artifact-aware testing, trace-backed diagnosis, and evidence-gated promotion.