llama-cpp

11 servers · 352★ total

On-device memory layer for AI agents. Claude Code, Hermes and OpenClaw. Hooks + MCP server + hybrid RAG search.

Android keyboard with local AI (Ollama, Whisper, MCP) or cloud (Gemini, Groq, OpenAI)

Private, on-device AI desktop app — GGUF (llama.cpp) & MLX models, a local coding agent, RAG knowledge base, Deep Research, vision and voice. 100% offline, no account, no telemetry. Windows & macOS.

Agentic Android Open Source Project (AAOSP) — Android fork with native LLM system service, MCP-aware apps, and an agent-driven launcher. On-device Qwen 2.5 via llama.cpp. Apps declare tools in their manifest. The OS runs the model.

AI-Inference-Managed-by-AI: Go binary for managing AI inference on edge devices

Local‑first MCP server for multi‑repository semantic code search with Qdrant and llama. Turns your entire workspace into private context for AI coding assistants like Claude Code, Codex, Cursor, Copilot, Antigravity and Windsurf.

A self-hosted, observable orchestrator for multi-agent workflows on your repos. Compose from skills and memory, run on events or cron, dispatch between agents.

MCP server that delegates Claude Code subagents to local models (LiteLLM/llama.cpp/Ollama), DeepSeek, AWS Bedrock, or any OpenAI/Anthropic-compatible backend — without losing the Claude Code orchestrator session.

Native C++17 AI agent runtime — local-first, embeddable, MCP-native. Tool-calling, memory, and an async agent loop that runs where Python can't: ROS2, game threads, drones, HFT.

MCP server for hot-swapping llama.cpp models in Claude Code - launchctl (macOS) + systemd (Linux)

🏭 Coding agent CLI engineered to make any LLM — local or cloud, frontier or 7B — actually useful. 16 providers, multi-tab sessions, automatic key+model failover, plan-mode review, MCP, and recovery loops for models that misbehave.