The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网第一个 DeepSeek Harness 视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
为 DeepSeek Harness 增加外挂识图模型:圆形鲸鱼按钮、发送图片识图自动回传、模型自主截图+识图工具、多协议自动适配、小白一键安装(未装 Node.js 自动下载)
DeepSeek Harness 插件:DeepSeek 大脑 + 自动识图。GUI 附加图片自动经 OpenAI 兼容 VLM 转译成文字后交给 DeepSeek 作答;支持百炼/智谱/OpenRouter 等任意 OpenAI 兼容端点(默认 qwen3.7-flash),无 key 自动探测本地 Ollama(图片不出本机);安装时有一问式确认
DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.
Image understanding, OCR, and persistent visual evidence for text-only DeepSeek Harness models
本地 OCR 插件:让纯文本生成 LLM 也能读懂图片 | Local OCR plugin: give text-only generative LLMs the ability to read images
Fully local document intelligence for DeepSeek Harness. Parse PDF, Office files, images, and scanned documents with offline OCR.
Native interactive visual-reasoning plugin for DeepSeek Harness: precise pixel grounding (SOM grid / zoom / annotate / measure / diff / color / OCR) + MiMo V2.5 multimodal backend, zero external MCP servers.
PaddleOCR skills for DeepSeek Harness with native tools and GUI configuration
GLM-4.6V 图像理解 MCP:识图/OCR/图表解析,原生接入 DeepSeek Harness(dsh-mcp-client),也兼容 Codex/Cline 等
DSH 插件:让纯文本模型也能看图。Web 端直接粘贴图片即可发送,无需指定图片路径;模型自主调用视觉技能查看,多模态模型原生直通,零skill绑定。
Codex-style attachment formats for the DeepSeek Harness Web GUI: PDF text-layer extraction, Office text extraction, scanned-PDF OCR, long-document spill + index cards, image-to-PNG.
Local OCR for DeepSeek Harness — read text from screenshots on your Mac with Apple's Vision framework. No API key, no upload.
Give a text-only dsh model eyes: pasted images recognized into text via an OpenAI-compatible vision endpoint.
Unlimited-OCR for DeepSeek Harness with a native tool and GUI configuration
Matter-aware legal workspace dashboard and document agent tools for DeepSeek Harness
让纯文本模型通过桌面豆包看见聊天图片的 DeepSeek Harness 宿主插件(CDP 桥接,全预设生效,识别可取消)
DSH 自建插件集合:微信桥接器 + GUI 微信入口补丁,一键安装
On-device macOS OCR and Apple Vision for DeepSeek Harness — one native plugin with a bundled Skill.
Paste images into DeepSeek Harness with a four-model vision race, OCR, and an automatic text bridge.
EagleEye MCP — pixel-accurate visual toolbox for Agents (screenshot, measure, OCR, regression)
零修改、零切换的 DeepSeek Harness 视觉能力插件:纯文本模型粘贴即读图片,云端 + 本地 Ollama 双后端自动切换,ModLens v2 风格结构化证据输出。
让 DeepSeek Harness 获得"看图"能力,自动识别图片真实格式,经任意 OpenAI 兼容视觉模型返回详细文字描述。
该仓库暂未提供项目说明。
Vision for text-only LLMs in DeepSeek Harness (DSH): describe images / OCR / VQA via free Gemini & GLM vision APIs
该仓库暂未提供项目说明。
DSH bundle: Qwen multimodal bridge — vision (qwen3-vl), speech-to-text (qwen3-asr), text-to-image (qwen-image), for DeepSeek Harness
Give text-only DeepSeek models eyes — a DeepSeek Harness plugin that transparently converts chat images into OCR text + vision-model descriptions before they reach the LLM. Configure vision backends (GLM-4V, Qwen-VL, Gemini, Ollama…) right in the Models settings page; multi-backend fallback chain, double-layer caching, no config files.