llm-evaluation
15 servers · 2,725★ total
Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.
Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.
Private, local-first AI assistant for Windows with permissioned tools, durable memory, and evidence-driven small-model improvements.
CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
Glass Box Framework — runtime constitutional verification for AI answers. Trust Cards with claim-level reasoning chains, formal ECS scoring, the 7-angle Glassbox Court red team, and deterministic audit logs. MCP-native.
Multi-model consensus debate via the filesystem. LLMs propose, peer-review, rebut, vote and synthesize a group-confirmed answer. CLI + MCP.
MCP server for human-in-the-loop surveys, A/B preference tests, ratings, and rankings. Get real human feedback inside Claude Code, Claude Desktop, Cursor, Windsurf, and any MCP client — powered by Datapoint AI.
Head-to-head benchmark comparing the official MCP to the MCP auto-created by Hintas.
Diagnose your AI agents in production. Extract policies from prompts, evaluate traces, generate diagnostic reports.
A standalone agent harness in Rust: a provider-agnostic loop, MCP tools, a path jail and prompt-injection interlock, sandboxed shell, scheduled triggers, and an eval rig that grades the trace.
A team-wide prompting coach. Scores every prompt sent through Claude Code, Claude.ai, VS Code or Copilot Chat on 5 dimensions in real time, coaches the weak ones, and feeds the strong ones into a team wiki that gets injected back into the next person's context.
Build full-stack applications with an autonomous IDE using an agent swarm architecture powered by the Kimi 2.6 model.
Python and MCP tools for testing falsifiable claims and recording MATCH, DRIFT, or UNVERIFIABLE outcomes against supplied evidence.
AI agent evaluation framework for multi-participant coordination tasks. Built with LangGraph, custom MCP tools, and LLM-as-a-Judge evaluation. MSc dissertation project (University of Edinburgh, 2025).