evaluation
10 servers · 2,184★ total
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.
Harness-oriented agent system framework for production-grade LLM agent applications
MCP server for comprehensive AI testing, evaluation, and quality assurance
An MCP server for dataset generation, fine-tuning, RL, and evaluation of LLMs — directly from your coding agent.
University for AI agents. 92 courses, 4400+ scenarios, any model via OpenRouter. Auto-training loops generate per-model SKILL.md documents. Works with Claude Code, OpenClaw, Cursor, Windsurf. No fine-tuning required.
The agent eval standard for MCP — score output quality, catch safety failures, enforce cost budgets
Evaluate how LLM agents actually use your MCP server — recording proxy, claude/codex, optional LLM judge, HTML report.
Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.
Python and MCP tools for testing falsifiable claims and recording MATCH, DRIFT, or UNVERIFIABLE outcomes against supplied evidence.