ocr
22 servers · 4,546★ total
A Model Context Protocol server for converting almost anything to Markdown
Give your AI agent eyes for PDFs — structured text, tables, OCR, visual evidence, and page-level citations via MCP. Native Rust, local-first.
Desktop automation MCP server — computer use for any AI agent: control screen, windows, mouse/keyboard, and Chrome via Model Context Protocol (stdio)
MCP server for computer use & browser automation - screenshot, OCR, click, type, find_text, Chrome/Electron CDP, template matching. macOS, Windows & Android. Works with Claude, Cursor, and any MCP client.
MCP server that lets Claude Code and other AI agents work through large PDFs, and whole folders of them, without overflowing context: hybrid semantic + keyword search, selective page reading, tables, images, OCR, chart data, and multi-column/CJK layouts.
Zotero AI plugin Research assistant for Zotero 9. Chat with your library, run federated scholarly search, RAG, OCR, systematic reviews, and manage cloud storage. Includes standalone MCP, Agentic capabilities, and skills library.
MCP server that turns any video — YouTube, Instagram, TikTok, Loom, X, Vimeo, direct URLs, local files — into transcripts, key frames, OCR text, and metadata for AI agents.
Convert PDF, Word, PowerPoint, HTML, email & 40+ formats to clean Markdown — and back. Built for LLMs, RAG & Python pipelines, with a built-in MCP server.
Don't write a bug report — record it. Local MCP: screen recording → transcript, frames, OCR, wall-clock evidence for coding agents.
Give an LLM agent control of the macOS desktop — a single self-contained Rust binary (no Python/JS stack) that speaks MCP: screenshots, Set-of-Mark targeting, Apple Vision OCR, mouse & keyboard.
Full Mistral AI MCP server — OCR, Voxtral audio, Codestral FIM, durable workflows, document extraction — for Claude Code, Cursor, Windsurf, Zed, Mistral Connectors
MCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude Code, Codex, Kimi, opencode, PI.
Give AI eyes and hands on your desktop. Open-source MCP server for desktop automation — screenshots, UI control, browser automation, OCR. Works with Claude, Cursor, and any MCP client. macOS + Windows.
MCP server for OCR using native Tesseract (C++), built with Node.js, delivering up to 10x faster performance than tesseract.js and integrable with ChatGPT Desktop.
Token-lean web microfetch for LLM agents: any URL → clean markdown via CLI, MCP server, and Claude Code plugin. Real browser-cookie auth, passkeys, anti-bot reach, on-by-default prompt-injection defense, plus on-device multimodal ASR/OCR. A single Rust binary — not a browser.
An MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.
Desktop AI agent for Windows — an architecture that expands what any model can do. Image/video OCR, offline speech transcription, an agent-driven browser, real .xlsx/.pptx/.docx output, MCP, and a delegated team of agents. Runs on your machine, with any provider.
本地视觉理解 MCP 服务,为 AI agent 提供识图能力:通过 vision / ocr / list_models / providers 四个 MCP 工具描述截图、界面、图表与照片。
Turn social-media creators' videos into a searchable, curated knowledge base — Next.js download/manage app + GPU transcription (faster-whisper) + RAG
One MCP server, 270 tools, replaces 75+ — files, shell, GitHub, Android, desktop vision, security & forensics, persistent memory. Smart context-saving tool loading, safe-delete guardrails, and an orphan-proof self-healing process supervisor. Works with Claude Desktop, Cursor, Windsurf, and any MCP client.
Self-hosted browser automation you own. Record a web task once, replay it forever — as a workflow, a REST API, or an MCP tool. Runs on your box: one SQLite file, no cloud, no telemetry.
Research intake that reaches the hard places: web, video, papers, scanned PDFs, browser, OCR, and audio into structured research packets. DOM extraction and change tracking built in; provenance rides along on every item.