local-llm

50 servers · 50,147★ total

47,109★

Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps

A curated list of OpenClaw resources, tools, skills, tutorials & articles. OpenClaw (formerly Moltbot / Clawdbot) — open-source self-hosted AI agent for WhatsApp, Telegram, Discord & 50+ integrations.

Harness the power of local LLMs with this TUI MCP Client for Ollama. Featuring all core MCP primitives (tools, prompts, resources), agent mode, multi-server, model switching, streaming responses, human-in-the-loop, thinking mode, model params config, system prompts, and saved preferences.

Curated real-world use cases for Hermes Agent — the self-improving AI agent from Nous Research. Backed by primary sources.

Persistent session memory for AI coding agents — local-first, with on-device inference, associative recall, and drift detection. Works with Claude Code, Cursor, and Codex.

131★

Evidence-gated runner for Codex, Claude Code, OpenCode, and local coding agents. Routes tasks into scoped DAG lanes with replayable artifacts.

Memory that learns what works.

MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.

Extend the Ollama API with dynamic AI tool integration from multiple MCP (Model Context Protocol) servers. Fully compatible, transparent, and developer-friendly, ideal for building powerful local LLM applications, AI agents, and custom chatbots

Android keyboard with local AI (Ollama, Whisper, MCP) or cloud (Gemini, Groq, OpenAI)

Live Web Access for Your Local AI — Tunable Search & Clean Content Extraction

Swift 6 agent SDK: type-safe tools, streaming, cloud + on-device inference via MLX on Apple Silicon

Private, on-device AI desktop app — GGUF (llama.cpp) & MLX models, a local coding agent, RAG knowledge base, Deep Research, vision and voice. 100% offline, no account, no telemetry. Windows & macOS.

RamiBot v3.8.0 is a local-first AI security operations platform integrating multi-LLM support, a dynamic red/blue team skill pipeline, MCP tool orchestration, Docker terminal access, Tor proxy management, and an auto-integrated Kali-based tool server (rami-kali) for controlled, extensible offensive and defensive workflows

SDK for Android — gives local LLMs structured access to phone data (calendar, contacts, SMS, files). Everything stays on device.

Open-source, self-hosted AI workspace for local and open-weight LLMs with vLLM, LiteLLM, autonomous agents, MCP tools, deep research, artifacts, and BYOK.

Local agent infrastructure in one stdlib-only Go binary

Provider-agnostic AI dev pipeline: clarify → plan → build → review → PR across your repos, mixing LLM providers per stage with your own keys. No vendor lock-in, no markup.

AI agent engine with 22 chat channels, 65 tools, and 15 skills. Self-hosted, local-first, works with any LLM provider.

AI-Inference-Managed-by-AI: Go binary for managing AI inference on edge devices

Private, local-first AI assistant for Windows with permissioned tools, durable memory, and evidence-driven small-model improvements.

The all-in-one, airgapped Local AI. Ypipe bundles a high-performance inference engine, specialized models, and pre-configured MCP servers into a single, self-contained Java executable. Zero-dependency orchestration for autonomous agent workflows with 100% data sovereignty. No cloud, no Python hell - just industrial-grade local intelligence.

A standalone terminal agent in a single C file. Model agnostic (13 providers + local hako), skill-driven, tool permission gated. No node, no python.

ChatGPT-like AI that runs 100% locally on your hardware. No subscriptions, no cloud, complete privacy. Multi-agent swarm + 10 MCP tools + hybrid RAG vector DB + . Runs on one GPU (RTX 5090 recommended)

MCP server for delegating mechanical tasks to local LLMs via Ollama. Claude does the thinking, your local model does the grunt work.

Talk to your data locally 💬📊. A private AI Data Analyst built with the Model Context Protocol (MCP), Ollama, and SQLite. Turn natural language into SQL queries without data leaving your machine. Includes a Dockerized Streamlit UI

Self-hosted Compound AI System for sovereign environments. Features token-saving complexity routing, deterministic MCP tools for math/logic, configurable expert templates, causal/reinforcement learning, agentic loops, and a federated Neo4j GraphRAG.

MCP server that delegates Claude Code subagents to local models (LiteLLM/llama.cpp/Ollama), DeepSeek, AWS Bedrock, or any OpenAI/Anthropic-compatible backend — without losing the Claude Code orchestrator session.

Local LLM web search agent powered by Ollama — self-hosted, source-grounded research with MCP tools, citations, and verified PDF reports.

Practical AI & more workflows for coding, MCP, local AI, productivity, research, media, and ops.

Local-first AI coding and automation agent with automatic Ollama setup, hardware-aware models, RAG, MCP, and a friendly terminal UI.

🗓️ AI Calendar Assistant powered by local LLM (Ollama) + Google Calendar API + MCP

Pulse — Hermes-style self-improving AI agent. Reliability-first rebuild with evaluated skill self-evolution, multi-agent team orchestration, dialectic user modeling, and fully self-hosted default stack (Ollama + SQLite FTS5). Apache 2.0 license.

MCP server for hot-swapping llama.cpp models in Claude Code - launchctl (macOS) + systemd (Linux)

A self-hosted, containerized platform for AI agents, exposed as Capability Packs — schema-validated, one-shot JSON tools — and native MCP. The defining metric is ≥90% pack success on 7B–30B-class open-weight models, something no frontier-targeting competitor is optimizing for.

No-code, local-first studio for building and running production ready AI agents. Inference runs on your machine — llama.cpp, MLX, or vLLM — or any OpenAI-compatible API you register

Quiet MCP executor for Claude Code and Codex. Offload edits, tests, refactors, and commits to cheaper local or flat-rate models.

The AI memory layer with 17 intelligence tools — catches contradictions, predicts burnout, tracks relationships. Model-agnostic, self-hosted

Personal AI-agent skills, workflows, prompts, and setup guides for local LLMs and MCP-powered assistants.

Universal cyber assistant: ~80 audited Kali tools + an expert skills library, driven by any model over MCP or direct run

Self-hosted AI agent framework with permission modes, audit logging and MCP. Runs on your machine, in Docker, or on a Raspberry Pi.

Web frontend for LM Studio — browser access, adaptive memory, multi-user auth, MCP tools

🏭 Coding agent CLI engineered to make any LLM — local or cloud, frontier or 7B — actually useful. 16 providers, multi-tab sessions, automatic key+model failover, plan-mode review, MCP, and recovery loops for models that misbehave.

Daedalus is a standalone terminal-based AI coding assistant that runs on your machine. It connects to local LLM servers (LM Studio, Ollama, llama.cpp, vLLM) or remote providers (OpenAI, Groq, OpenRouter, Anthropic), routes requests intelligently, and gives your AI agent access to your file system, terminal, git, web search, and codebase indexing

MCP routing proxy for Claude Desktop — local Ollama model switching via LiteLLM

Local MCP server with safe tools for files, Git, browser checks, prompt improvement, skill routing, Notion, Obsidian, RAG, and personal AI workflows.

Model-agnostic multi-agent orchestration built on CrewAI. YAML-driven agent and MCP catalogs, dynamic LLM planning, sessions, and pluggable execution — mix Ollama, OpenAI, Anthropic, HuggingFace, and your own tools without rewriting orchestration logic.

Self-hosted, extensible AI agent platform with governed MCP tools, persistent project memory, durable workflows, capability packs, autonomous execution, and distributed local-model compute.

MCP server that offloads cheap work from your cloud LLM agent to a local Ollama model — summaries, drafts, extractions, first-pass reviews — at zero cloud cost.