ai-safety

51 servers · 1,380★ total

The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable evidence.

Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.

Agent orchestration & security template featuring MCP tool building, agent2agent workflows, mechanistic interpretability on sleeper agents, and agent integration via CLI wrappers

Governance gateway for AI agents — bounded, auditable, session-aware control with MCP proxy, shell proxy & HTTP API. Works with Cursor, Claude Code, Codex, and any MCP-compatible agent.

[L0 CONSTITUTION] arifOS — constitutional MCP kernel. Law, identity, F1–F13, VAULT999. Judges but never executes. DITEMPA BUKAN DIBERI.

Make AI coding agents safe to scale autonomously: assign work, cap spend, enforce policy, verify output, roll back failures, learn from loops, and prove ROI across every repo.

Agent Identity Protocol - Zero-trust security layer for AI agents. Policy enforcement proxy for MCP with Human-in-the-Loop approval, DLP scanning, and audit logging.

Approval workflows for AI agents

MCTS (Model Context Threat Scanner) is a local-first security scanner for MCP servers -- static and live tool discovery, multiple analyzers, auditable risk scores, and JSON, SARIF, and HTML output. For authors and platform teams; CI-ready, no cloud API.

Noisegate: a differential privacy gateway that lets an untrusted LLM agent query sensitive data over MCP (Model Context Protocol), with a formal guarantee no individual's record can leak even if the agent is adversarial - enforcement lives in trusted code below the model, validated by a runnable attack gallery.

Audit any agent decision across its past, present, and future, on one typed graph.

The open-source safety layer for AI agents — block unsafe tool calls, require approval, enforce budgets, audit, replay.

Open-source observability and audit trail platform for AI agents. MCP-native, tamper-evident event logging, real-time dashboard.

LLM guardrails & prompt injection detection for Python. Auto-instruments LangChain, CrewAI, OpenAI, LiteLLM + 8 more frameworks. PII masking, toxicity detection, policy CI/CD. One line, zero code changes.

Open-source model inventory & governance. Discovers every model, rule, and pipeline across all your platforms as one immutable, agent-queryable graph — git for models.

🔪 Open-source safety firewall for AI agents. Intercepts tool calls before they execute, enforces YAML policies, and kills dangerous operations in real-time. Works with OpenAI, Anthropic, LangChain, and MCP. She doesn't guard. She kills.

Real VMs for your LLM. Provision, execute, validate. In seconds. LLM4Ops.

MCP EU AI Act Compliance Scanner - Open source tool to detect EU AI Act violations in codebases

Glass Box Framework — runtime constitutional verification for AI answers. Trust Cards with claim-level reasoning chains, formal ECS scoring, the 7-angle Glassbox Court red team, and deterministic audit logs. MCP-native.

Workflow guardrails for coding agents and LLM-assisted development: AGENTS.md, native hooks, MCP, SaneMaster, circuit breakers, and shared process checks.

MCP middleware that blocks dangerous AI agent actions using a simple YAML config

A safety layer for AI coding agents. CLAUDE.md/AGENTS.md generator, MCP runtime guardrail, pre-commit hook, GitHub Action.

Checks AI-generated code changes before merge: scope, validation, risk, review evidence, and optional Verity receipts.

Trust infrastructure for AI agents — constitution enforcement and cryptographic receipts. Python SDK.

MCP servers expose tools with no information about what they actually do at runtime. mcpsafetywarden sits between your agent and any MCP server, profiling tool behavior, blocking destructive calls, and running active security audits before you trust them in a workflow.

Local guardrails for AI coding agents. Wraps any MCP server and blocks destructive tool calls — DROP TABLE, rm -rf, force-push, unscoped UPDATE/DELETE — before they execute. Free, open-source, runs entirely on your machine.

Agentic security control plane for MCP and AI agent tool calls. MCP-native policy gateway with topology discovery and audit.

Secure email proxy for AI agents — content filtering, PII redaction, and prompt injection detection over MCP

An AI that administers Linux through typed, approval-gated, Ed25519-audited actions. The model selects the action; the daemon builds the command.

Authorization, delegation, provenance, and verifiable-audit engine for AI agents. MCP adapter published.

Tamper-evident audit log for Go: an append-only, hash-chained, optionally Ed25519-signed log/slog handler, with offline logx verify. Built on the standard library, zero third-party dependencies. Also a full slog toolkit (colored output, JSON, rotation, fan-out, CLI).

A local proxy that wraps your MCP servers and checks each tool call against policy and live state before it runs - allow, block, or request a refresh, with a reason the agent can act on.

The verification layer for autonomous agents — an independent, capital-aware verdict before an irreversible action (/review), a cryptographically signed proof after, and a public on-chain track record, losses included. Free EdgeProof backtest validation. Underneath: memory, sandboxed execution, marketplace. Pay in sats, USDC (x402), or card.

Belay is an open-source, local-first security layer for AI coding agents (Claude Code, Codex, Cursor, OpenClaw, Hermes Agent and MCP) that blocks dangerous commands, secret leaks, and prompt injection at the tool-call boundary in under 100ms — no LLM in the decision path by default, no cloud, no phone-home.

Runtime artifact existence & freshness verification for AI agent completion claims — a lightweight, zero-LLM MCP gate that source-binds 'done' claims to real files. MIT.

Audit all locally configured MCP servers for permission risks, prompt injection threats, and schema drift

The quality gate for the MCP ecosystem. Spec compliance, security scanning, and performance benchmarks for MCP servers.

Make every AI chart write a reviewable proposal. Draft protocol for agent-initiated FHIR writes (Proposal → Decision → Commit → Audit) with web + MCP implementations and a conformance suite. Runs over Medplum, HAPI, Firely, Aidbox, or Oystehr.

Agent Run Config — an open specification for declaring, packaging, securing, and sharing portable, governed AI agents. Like a Dockerfile for agents: one reviewable Agentfile for identity, tools, boundaries, policy, and OCI packaging.

Deterministic security proxy for MCP tool calls — iptables for MCP

Sunglasses for AI agents. Protection layer + neighborhood watch.

An AI organism - the human-AI interface layer. A safety reflex that refuses known lethal actions even when the model is fooled, and a memory that only keeps what actually worked. For code, research, writing, or a fleet of agents. 100% local.

Stop AI agents from doing dangerous things through MCP. Wraps any MCP server with allow/block/hold-for-approval guardrails.

Deterministic MCP Security Architecture. FrozenNamespace as Root of Trust for Model Context Protocol tool verification

Runtime governance for AI agents — 57 MCP tools, embedded engine, no key required. Decision gates, forensic audit, EU AI Act / NIST / SOC 2 mapping. Model-agnostic.

Prompt-injection defenses for Claude Code. A PreToolUse Bash hook blocks compositional credential-exfiltration shapes (secret read plus network, env dump to network, remote script to shell, reverse shells). A sanitizing MCP server wraps untrusted URLs and files in sentinels, strips invisible unicode, flags jailbreaks.

Trust infrastructure for AI agents — constitution enforcement and cryptographic receipts. TypeScript SDK.

Runtime guardrails for AI agent MCP tools. Blocks dangerous tool calls at runtime with a default-deny policy proxy.

A curated, daily-updated list of awesome resources, tools, SDKs, papers, and projects for Anthropic & Claude AI

Porcupine: A Safe Autonomous AI Agent

MCP server for AI security intelligence. Check any MCP server for supply-chain threats before installing -- from Claude, Cursor, or Windsurf.