evaluation

10 servers · 2,184★ total

A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

Harness-oriented agent system framework for production-grade LLM agent applications

MCP server for comprehensive AI testing, evaluation, and quality assurance

An MCP server for dataset generation, fine-tuning, RL, and evaluation of LLMs — directly from your coding agent.

University for AI agents. 92 courses, 4400+ scenarios, any model via OpenRouter. Auto-training loops generate per-model SKILL.md documents. Works with Claude Code, OpenClaw, Cursor, Windsurf. No fine-tuning required.

The agent eval standard for MCP — score output quality, catch safety failures, enforce cost budgets

Evaluate how LLM agents actually use your MCP server — recording proxy, claude/codex, optional LLM judge, HTML report.

MCPLab - Test and evaluate MCP servers with LLMs

Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.

Python and MCP tools for testing falsifiable claims and recording MATCH, DRIFT, or UNVERIFIABLE outcomes against supplied evidence.