kitlau86/agent-vision-mcp
View on GitHub ↗An MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.
10 ★3 forksTypeScriptUpdated 3mo ago
Topics
- Stars
- 10★
- Forks
- 3
- Language
- TypeScript
- License
- MIT
- Created
- 2026-06-04
- Last push
- 2026-06-05