multimodal-ai
6 servers · 1,911★ total
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
Wan 3.0 API Python SDK and MCP server for AI video generation: text-to-video, image-to-video, multimodal references, uploads, and async job polling.
面向智能座舱的云边协同 AI Agent:端侧毫秒级车控,云端声明式 Multi-Agent 与 Skill/DAG 编排,支持 S2S 实时语音、声纹多用户、视觉与 HMI;LLM 只负责理解与规划,VAL 负责确定性安全执行。
Agent-first CLI and MCP bridge for 2,000+ AI models, multimodal generation, sandboxes, and agent workflows—one command, one account.
Community MCP vision bridge for Xiaomi MiMo Vision, enabling image understanding for text-only LLM agents.
Companion source code for AI APIs in Practice, a visual engineering book on multimodal AI, agents, RAG, evaluation, governance, and production-ready products. Book DOI: 10.5281/zenodo.21431944 · Software DOI: 10.5281/zenodo.21419466