multimodal

17 servers · 3,376★ total

Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.

TeleMem is a high-performance drop-in replacement for Mem0, featuring semantic deduplication, long-term dialogue memory, and multimodal video reasoning.

MCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

Gemini Vision & Image Generation MCP for Claude Desktop and Claude Code

A unified Model Context Protocol server for MiniMax CLI (mmx)

Swift 6 agent SDK: type-safe tools, streaming, cloud + on-device inference via MLX on Apple Silicon

Private, on-device AI desktop app — GGUF (llama.cpp) & MLX models, a local coding agent, RAG knowledge base, Deep Research, vision and voice. 100% offline, no account, no telemetry. Windows & macOS.

Self-hosted AI knowledge base with hybrid semantic search (pgvector + FTS + RRF), MCP server, multi-provider LLM inference (Ollama, OpenAI, OpenRouter, llama.cpp), multimodal ingestion (vision, audio transcription, speaker diarization), and knowledge graph. Rust + PostgreSQL.

Zero-dependency Python SDK for WeChat Bot (iLink protocol) — text, image, video messaging & MCP server | 零依赖微信机器人 Python SDK,支持多媒体消息和 MCP 服务器

MCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude Code, Codex, Kimi, opencode, PI.

Token-lean web microfetch for LLM agents: any URL → clean markdown via CLI, MCP server, and Claude Code plugin. Real browser-cookie auth, passkeys, anti-bot reach, on-by-default prompt-injection defense, plus on-device multimodal ASR/OCR. A single Rust binary — not a browser.

54-tool MCP server for multimodal automation the only open-source MCP with full PPTX/DOCX/PDF lifecycle, plus FFmpeg, GIMP, Inkscape, Blender, FreeCAD, MATLAB, Godot integration. 7 MCP services orchestrated in one workspace.

MCP server for reading WeChat (微信) Official Account articles with native multimodal output - images and video keyframes as content blocks, not URLs.

本地视觉理解 MCP 服务,为 AI agent 提供识图能力:通过 vision / ocr / list_models / providers 四个 MCP 工具描述截图、界面、图表与照片。

Specialized AI agents for DeFi, crypto, blockchain, Web3, DeFAI. Deep expertise in yield optimization, smart contracts, portfolio management. Multi-agent collaboration (Agent Teams), MCP (Model Context Protocol) integrations, universal JSON API. Open-source, no vendor lock-in, 18 languages.

A token-efficient, provider-agnostic vision layer for reasoning models via MCP.

Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.