vision

10 servers · 251★ total

多模型视觉理解 MCP 服务器,为不支持图片理解的 AI 编码模型提供视觉能力:分析截图、报错、UI 与文档,可接入多家主流视觉大模型。Multi-model vision MCP server that adds image understanding to AI coding models without native vision — analyze screenshots, errors, UI and documents via major vision LLM providers.

MCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字

MCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude Code, Codex, Kimi, opencode, PI.

Persistent visual memory for AI agents — capture screenshots, embed with CLIP ViT-B/32, compare, recall. MCP server + Rust core library.

An MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.

DS-VISION V3: auditable visual evidence for coding agents via MCP (visual evidence for agents without native vision).

本地视觉理解 MCP 服务,为 AI agent 提供识图能力:通过 vision / ocr / list_models / providers 四个 MCP 工具描述截图、界面、图表与照片。

AI image generation, vision analysis, reference editing and sprite processing MCP for Cocos Creator game development, powered by OXOXOS.

Codex Vision is a macOS-only Codex plugin that lets a local Codex session request live camera frames through an explicit MCP tool. It packages a native AVFoundation capture app, compliant plugin metadata, install scripts, and docs so Codex can clone, build, install, and use laptop-camera vision without cloud camera services or browser hacks. Local.