AI Video Editor Pipeline with Vision LLM Models
-
Updated
Apr 11, 2026 - Python
AI Video Editor Pipeline with Vision LLM Models
Claude Code skill — 给无视觉能力的 LLM(DeepSeek/o1/o3)外挂看图能力。Vision proxy for text-only LLMs, supports OpenAI/Anthropic/DashScope/Qwen3-VL.
A lightweight Model Context Protocol (MCP) server for retrieving local and remote images for LLM vision models, featuring metadata extraction and configurable availability retries.
DSH plugin: vision_read — route image reading to a dedicated vision model (e.g. Kimi K3) so text-only agents can see images
Alt'Ollama: The app for generating alt-text for image/s using LLMs with image processing support.
Native macOS video frame extractor built on AVFoundation — no ffmpeg, precise timestamps, for feeding frames to LLM vision models.
AI-first smart home on Home Assistant: Gemini vision at the front door, Home Front Command alert choreography, mmWave presence, Telegram command center. Real, redacted screenshots + interactive demo.
LLM Vision integration for Home Assistant using Google Gemini
Home Assistant Card to display the LLM Vision Timeline
To associate your repository with the llm-vision topic, visit your repo's landing page and select "manage topics."