agent-vision-toolkit
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
- Type: Skill
- Imported from GitHub
- Popularity: 1.2K GitHub stars
- License: MIT
- Source: https://github.com/Anionex/agent-vision-toolkit
- Repository: https://github.com/anionex/agent-vision-toolkit
- Tags: agent, agent-skills, claude-code, codex, computer-use, deepseek, dsh-plugin, glm, harness-engineering, multimodal, opencode, text-only-llm
- Updated: 2026-10-01
More skills
- skills — Public repository for Agent Skills
- ponytail — Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- agent-skills — Production-grade engineering skills for AI coding agents.
- open-design — 🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becom…
- claude-mem — Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and inject…
- Understand-Anything — Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about…
- archify — Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion a…
- awesome-claude-skills — A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows