retrieval-augmented-generation

42 servers · 5,702★ total

Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.

NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

Pharos — local-first agentic RAG for your team's document library: multi-format ingest, hybrid retrieval, enterprise ACL, dual HTTP + MCP exits.

Open-source CLI for semantic Taiwan legal judgment retrieval. Search judgments, package them for your own AI (Claude/ChatGPT), and run a bundle-level citation check. Bring your own LLM; retrieval-only.

On-device memory layer for AI agents. Claude Code, Hermes and OpenClaw. Hooks + MCP server + hybrid RAG search.

🔥🔥🔥 Enterprise AI middleware, alternative to unifyapps, n8n, lyzr

Cross-platform persistent memory MCP for Codex, Gemini CLI, Claude Code, and other local MCP hosts. 36 cited neuroscience mechanisms, local-first SQLite/PostgreSQL, hybrid retrieval, decay-based consolidation, and reproducible benchmarks. Claude adds optional automatic lifecycle hooks.

Biologically-inspired persistent memory engine for Claude Code. 26 cognitive subsystems, Hopfield networks, predictive coding, causal discovery, successor representations, all running locally over SQLite.

Local-first, auto-injecting context + memory layer for Claude Code, Codex and MCP. Per-prompt hook injection, explicit freshness/Trust Contract, AST indexing, call graphs, hybrid search, wiki + architecture dashboard. Everything stays on your machine.

Self-hosted AI knowledge base with hybrid semantic search (pgvector + FTS + RRF), MCP server, multi-provider LLM inference (Ollama, OpenAI, OpenRouter, llama.cpp), multimodal ingestion (vision, audio transcription, speaker diarization), and knowledge graph. Rust + PostgreSQL.

Self-hosted RAG engine for AI coding assistants. Ingests technical docs & code repositories locally with structure-aware chunking. Serves grounded context via MCP to prevent hallucinations in software development workflows.

Simple RAG implementation from scratch using MCP, focusing on Perception, Memory, Decision and Action

Open-source, self-hosted knowledge backend for AI agents — hybrid search (vector + keyword), MCP server, 5 connectors, Docker-ready

Curated catalog and dataset of 50 cloneable GenAI apps and Hugging Face Spaces across RAG, AI agents, LangGraph, MCP, and more

A curated list of open-source alternatives to proprietary AI tools — LLMs, agents, RAG, coding assistants, image/video/voice generation, and local deployment. Updated July 2026.

Citable research retrieval for AI agents across papers, books, patents, Wikipedia, and public discussions.

Lightning-fast RAG for AI agents. ONNX-powered, 4-layer fusion, MCP server. No PyTorch.

Zero-infrastructure, local-first memory for AI agents. One SQLite file — no Postgres, Neo4j, or cloud. Bi-temporal supersession (updates, never deletes), four-channel hybrid retrieval, and a drop-in MCP server. Runs offline on cheap models. Apache-2.0. Bring your own benchmark.

Lightweight RAG server for the Model Context Protocol: ingest source code, docs, build a vector index, and expose search/citations to LLMs via MCP tools.

Developer acceleration layer for enterprise RAG + agent pipelines. 32 production-style flagships on LangChain, LangGraph, CrewAI, Temporal, OpenAI Agents SDK, pydantic-ai. Permission-aware retrieval, chunk-level citations, schema-enforced output. Built on Model Context Protocol (MCP).

RAG for researchers: page-level citations from your personal library, LLM access via MCP. Ask your entire archive, get answers with the intelligence of leading AI.

AletheiaDB: The cognitive memory engine for AI agents with fact supersession.

the memory layer of your codebase. Knows the WHY, the WHAT, the WHERE-IT-BREAKS.

The 21 agentic design patterns, distilled to concepts (no code) — a design-phase reference for building AI agents. Works as a Claude Code skill.

Musubi (結び) — Ai Agent shared memory and thought layer. The braiding of threads between presences.

RAG MCP Registry Finder (RAGMap): subregistry API + MCP server + Firestore ingestion

Create unlimited AI chatbot agents for your website — powered by OpenAI-compatible LLMs, RAG, and MCP.

Local-first semantic search for your ChatGPT, Claude, and AI conversation history. Python CLI + MCP server for Claude Desktop. Uses sentence-transformers + ChromaDB. No API keys, nothing leaves your machine. Part of the Resonant ecosystem.

MCP server for agent memory over HexxlaDB—ring retrieval, embeddings + lexical search, seams, facets, YAML persistence policy, localhost HTTP transport.

Retrieval-grounded prompt refinement as a Python library, CLI, and MCP server

Agentic RAG Chatbot using Model Context Protocol (MCP) to answer queries from PDFs, Word, and PPT files.

Rust-native universal retrieval layer for humans and AI agents — search, extract, rank, and synthesize the web. CLI + REST API + MCP server, with adaptive extraction (CEP), token-budgeted packing, multi-signal ranking, and cited research.

Turn any repo into a Code Knowledge Graph your coding agent can reason over — symbols, API routes, ORM models, architecture decisions & git history as a typed, provenance-tracked graph, served over MCP. Built on AgentForge.

Local-first RAG engine with cited answers from your own docs. GraphRAG, sealed sharing, MCP-native, VS Code Copilot integration. Nothing leaves your machine.

Shared context substrate for AI agents. Retrieval that learns what's useful. Runs local or cloud.

Self-hosted documentation RAG pipeline that serves semantic search to LLM agents over the Model Context Protocol.

Fast, local semantic search over web content for AI agents. Hybrid BM25 + potion-retrieval-32M embeddings, cross-page dedup, token-budget mode, MCP server, SearXNG bridge. ~90% fewer tokens than raw web_fetch.

Repository containing practical exercises and notebooks focused on AI application development and experimentation.

RE-call — Retrieval-Augmented Self-Recall: RAG over an AI agent's own memory that knows when it doesn't know (gap detection, freshness, anti-re-litigation). PostgreSQL + pgvector, hybrid retrieval + RRF.

Portable hybrid (BM25 + dense + RRF) retrieval engine and a label-free evaluation harness — extracted from a personal AI-assistant memory index and decoupled to run on any source tree.

A verifiable engineering knowledge engine that turns technical documents, CAD models, and equipment data into structured knowledge for humans and AI agents.

Cloudflare AI Search toolkit: MCP server, streaming /ask Worker, docs widget, and git-to-R2 corpus sync. AGPL.