token-optimization
49 servers · 3,982★ total
Cut AI token costs 95%+ on code exploration. The leading MCP server for precise, symbol-level GitHub code retrieval via tree-sitter AST. Works with Claude Code, Cursor & any MCP client. 313B+ tokens saved.
An MCP server that executes Python code in isolated rootless containers with optional MCP server proxying. Implementation of Anthropic's and Cloudflare's ideas for reducing MCP tool definitions context bloat.
The leading, most token-efficient MCP server for documentation exploration and retrieval via structured section indexing
Token-efficient MCP CLI client. 97% less tokens on tool discovery, 40-60% on results. Zero deps. Cross-platform. Works with every AI agent.
Context compression for MCP. Same upstream call, exact results recoverable.
Token-efficient MCP server for tabular data retrieval. Index CSV/Excel files, query rows, aggregate — 99%+ token savings vs raw file reads.
Context compression plugin for Claude Code. Automatically trims large tool output—JSON, YAML, stack traces, and logs—before it enters the context window.
OCTAVE protocol - structured AI communication with 3-20x token reduction. MCP server with lenient-to-canonical pipeline and schema validation.
Keep Claude Code context clean. Open-source toolkit: drift detection, re-read dedup, integrity scoring, AST-aware reads, 15 MCP tools. 62.6% measured savings, reproducible.
The AI-native wire format for structured data. 100% comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ lossless round-trips across 17 formats. Spec v3.4 Stable.
Smarter reads, safer edits. An MCP plugin that cuts token usage and catches editing mistakes before they hit disk. Supports Claude Code, Gemini CLI, GitHub Copilot, and Codex.
Run DeepSeek as a real sub-agent inside Claude Code / Codex CLI — DeepSeek gets its own 7-tool agent loop in a sandboxed workspace, not just a single LLM call.
MCP server for Unity Editor — 160 tools for scene, assets, animation, VFX, playtest & more
⚡ Cut LLM inference costs 80% with Programmatic Tool Calling. Instead of N tool call round-trips, generate JavaScript to orchestrate tools in Vercel Sandbox. Supports Anthropic, OpenAI, 100+ models via AI Gateway. Novel MCP Bridge for external service integration.
Token optimizer for Claude Code Gemini CLI & Qwen Code — compresses shell outputs via PreToolUse hook, tracks USD savings in TUI dashboard
Call the Google Antigravity CLI (agy) headlessly from any non-TTY context — Windows ConPTY + POSIX pty + MCP server. Fixes the empty-output bug (#76).
AI-powered MCP proxy for Playwright and Figma. Optimize Claude Code and AI agents with session recovery, context reduction and intelligent tool routing.
✈️ 24-tool MCP server for Claude Code: preflight checks for your prompts, cross-service context, session history search with LanceDB vectors, correction pattern learning, cost estimation
⚜️ An MCP server for context compaction and recycling in Claude Code
Token-saving code search for AI coding agents — TypeScript, JavaScript & Python (zero setup), C#, C++, Go & 30+ more, via a language-server or tree-sitter index instead of grep. Local-only, no IDE. Claude Code plugin + MCP + CLI.
MCP proxy: zero-code GCF adoption. Wraps any MCP server, converts JSON to GCF mid-flight. 53-71% fewer tokens. Works with any structured data.
Context-optimized MCP server for web scraping. Reduces LLM token usage by 70-90% through server-side CSS filtering and HTML-to-markdown conversion.
Persistent code-knowledge memory for AI chats. An MCP server that caches a codebase's structure and summaries in a local SQLite "brain" so sessions query it instead of re-reading files — token-cheap, 100% local.
MCP proxy server — 85-96% token savings via lazy tool loading & fuzzy search across 100+ MCP servers. npm i mcp-tool-search
CLI for Model Context Protocol (MCP) servers. Pipe, script and automate MCP tools from the Unix shell, and keep their tool schemas out of your LLM context.
Use less context in Claude, Codex, and Cursor. Local AI memory and verifiable work briefs for long-running tasks.
Git LFS for your LLM context — 20x Context Savings — an MCP server that keeps big documents out of the model and returns byte-exact quotes for a fraction of the tokens.
Tentra MCP — memory for AI coding agents. Persistent code graph + AI architecture diagrams. 32 MCP tools. Works in Cursor, Claude Code, Codex, Windsurf.
Save 60-80% tokens when AI reads code — MCP server for token-efficient code navigation with AST-aware structural reading
Reduce Claude AI token consumption. Zero install to start. Python 3.7+ for auto-manifest generation.
GCF Python implementation. 100% LLM comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ round-trips verified. Zero dependencies.
GCF Rust implementation. 100% LLM comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ round-trips verified. Zero dependencies.
Local-first AI context engine combining a structural Atlas, verified Archive, episodic Animus memory, CLI, MCP, and focused context routing.
MATLAB/Simulink/Simscape MCP compression proxy — 14 domain-specific rules, 48-85% token reduction on simulation output
Compression Runtime for Universal eXecution — a 60–95% token reducer for AI coding agents (Claude Code, Cursor, Cline, …). Single Rust binary, SQLite, 11 layers, local-first.
Cut the tokens your IDE-to-LLM coding session is billed for. Stops re-sending files the model already has, places provider cache breakpoints where they cover the request, and compresses what needs sending — measured against real tokenizers, not character counts. CLI, VS Code, MCP server, proxy.
GCF Go implementation. 100% LLM comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ round-trips verified. Zero dependencies.
GCF TypeScript implementation. 100% LLM comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ round-trips verified. Zero dependencies.
Context compression for AI agents — cut token usage costs with retrievable CCR compression. Rust core + Python API, Claude Code & Codex plugin, MCP server. PyPI: furl-ctx
Context-cheap SerpApi search for AI agents: a CLI + Claude skill that costs ~0 standing tokens vs an MCP server (measured).
Token-Optimized Unified MCP Server for Gmail & Microsoft 365. 60% fewer tokens, 100% more power.
Rust core + MCP server for running many LLM agents in parallel without blowing your token budget — hard budget caps, automatic work dedup, and deterministic result merging.
Find the MCP schema issues eating your context window. 156 checks. Grade A+ through F. ESLint for MCP.
MCP server giving LLM agents a seven-verb API over papers, documents, code, state, patents, and cached web/Wolfram/YouTube tool calls
One persistent MCP terminal your AI drives — and launches other coding agents (Codex/Grok/Composer) into. SSH, containers, and REPLs nest as text you send in. tmux-backed, token-reduced reads, headless over MCP.
The local-first brain for coding agents: project memory + workflow intelligence + lossless ~50% context savings. MCP server + LLM proxy, zero cloud.