Skip to content
#

vision-model

Here are 36 public repositories matching this topic...

The self-hosted AI workstation. Autonomous screen agents, 3-tier neural routing, parallel agent swarms, video generation, 4K/8K upscaling, RAG, voice interface, 70+ tool execution engine — all running locally on your hardware.

  • Updated Jul 29, 2026
  • Python

AI-powered browser automation agent using a dual-LLM architecture. The orchestrator (qwen3-vl-32k) creates execution plans from screenshots, while the executor (llama3.1:8b) translates steps into browser actions using an accessibility tree for reliable element selection. Local, private, powered by Ollama.

  • Updated Dec 13, 2025
  • JavaScript

Локальная мульти-агентная система на self-hosted LLM (Ollama). 3 канала (Telegram/консоль/MAX), 4 типа моделей (LLM/embed/vision/voice), мультимодальность (текст/PDF/фото/голос), роли Planner/Executor/Critic с рефлексией. Инструменты: почта, Яндекс.Диск, web, OCR, cron-планировщик. Семантическая память (sqlite-vec). Coverage 88%. Без облачных API.

  • Updated Jul 28, 2026
  • Python

AI video editing skill that watches your footage before cutting. Vision model analyzes every frame, scores usability, drops bad takes, reorders by narrative, then renders with ffmpeg. For Claude Code, Cola, OpenClaw. 会看画面的 AI 自动剪辑 Skill——先用视觉模型看懂素材再决定剪哪几秒。

  • Updated Jul 26, 2026
  • Python

🔍 A CLIP-powered image similarity finder built with Streamlit — upload a query image and find the most visually similar matches from a gallery using deep visual embeddings.

  • Updated Jul 27, 2025
  • Python

Improve this page

Add a description, image, and links to the vision-model topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the vision-model topic, visit your repo's landing page and select "manage topics."

Learn more