返回目录

GITHUB TOPIC

multimodal

27 个项目包含此标签

27 个项目

GitHub Topic 精确匹配

文件与数据技能精选

modlens

liustack

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网第一个 DeepSeek Harness 视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

1.4kTypeScript2026-02-22
文件与数据插件

Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

71JavaScript2026-08-14
文件与数据插件

DeepSeek Harness 插件:DeepSeek 大脑 + 自动识图。GUI 附加图片自动经 OpenAI 兼容 VLM 转译成文字后交给 DeepSeek 作答;支持百炼/智谱/OpenRouter 等任意 OpenAI 兼容端点(默认 qwen3.7-flash),无 key 自动探测本地 Ollama(图片不出本机);安装时有一问式确认

7JavaScript2026-08-13
文件与数据插件

Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel

5TypeScript2026-08-14
文件与数据插件

DeepEye vision plugin for DeepSeek Harness (DSH): image description, OCR, VQA, UI layout, and clipboard analysis.

4TypeScript2026-08-14
文件与数据插件

给 DeepSeek 安装一双眼睛和一支画笔:会话里直接贴截图/图片,GLM 视觉模型先精确转写图片内容(报错信息、代码、界面逐字保留),然后 DeepSeek 继续处理你的问题——同一轮完成,全程无感;需要配图时,DeepSeek 自动调用文生图后端出图并显示在会话中。

3TypeScript2026-08-14
文件与数据插件

dsh-vision

237229953-create

DSH plugin: text-only models (e.g. DeepSeek-V4) automatically see images via a vision model. Official surface-replace, cache-friendly, human transcript untouched. 纯文本模型自动识图桥

2JavaScript2026-08-14
模型与 MCP插件

dsh-qwen-mm

RRRosmontis

Qwen-MM-Plugins integration bundle for DeepSeek Harness (dsh) — multimodal MCP tools (vision, OCR, ASR, search, video, Blender, FreeCAD) + image attachment bridge. 让 DeepSeek Harness 原生支持多模态。

2TypeScript2026-08-14
模型与 MCP插件

DeepSeek Harness 的 Qwen-MM-Plugins 集成插件:12 个多模态 MCP 工具(视觉/OCR/定位/ASR/音视频)、Web 设置页(粘贴 Qwen API Key 即用)、内置技能与一键安装器

2JavaScript2026-08-14
文件与数据插件

A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).

2JavaScript2026-08-14
生活娱乐插件

dsh-guide-dog

AtropinolTT

Guide Dog for DSH — MiniMax multimodal plugin: image/video/music/speech generation & vision tools, hard-metric voice mode (host event-driven auto TTS, session-safe playback), voice input (mic → faster-whisper → composer), recorder page, two-layer settings.

1JavaScript2026-08-14
开发工具插件

deepsee

chang416

Vision + smart model routing for DeepSeek Harness. Gemini sees. DeepSeek codes.

1TypeScript2026-08-15
文件与数据插件

Paste images into DeepSeek Harness with a four-model vision race, OCR, and an automatic text bridge.

1JavaScript2026-08-14
文件与数据技能

给纯文本 LLM 一双慧眼。 一个 DeepSeek Harness(DSH)原生 skill + 零依赖 Python CLI, 为 DeepSeek 等纯文本模型补上图像理解与文档解析(OCR、表格、公式、PDF → Markdown), 使用免费额度优先的三方多模态 API,国内网络直连、无需代理。

1Python2026-08-14
文件与数据插件

MiniMax multimodal capability hub for DeepSeek Harness (DSH): image understanding (VLM), text/image-to-video, speech, music, audio cover, web search, quota — one mmx_multimodal model tool wrapping the mmx-cli.

1JavaScript2026-08-14
生活娱乐插件

dsh-voice

zhuiyueya

Voice for DeepSeek Harness(dsh) — speech-to-text input + read-aloud TTS for text-only DeepSeek, zero API key.

1JavaScript2026-08-14
文件与数据插件

dsh-vision-relay

junhongchashui

零修改、零切换的 DeepSeek Harness 视觉能力插件:纯文本模型粘贴即读图片,云端 + 本地 Ollama 双后端自动切换,ModLens v2 风格结构化证据输出。

0JavaScript2026-08-14
模型与 MCP插件

vision_kit

Seom-ingit

Make your AI agent a math tutor. Structured extraction of vectors, matrices & geometry from math figures, with dimension-consistency + geometric self-check. Vision plugins for DeepSeek Harness, opencode (MCP) & CLI. Verify, don't believe.

0Python2026-08-14
文件与数据插件

DSH bundle: Qwen multimodal bridge — vision (qwen3-vl), speech-to-text (qwen3-asr), text-to-image (qwen-image), for DeepSeek Harness

0JavaScript2026-08-14
文件与数据插件

dsh-visionary

zhuiyueya

Give text-only DeepSeek models eyes — a DeepSeek Harness plugin that transparently converts chat images into OCR text + vision-model descriptions before they reach the LLM. Configure vision backends (GLM-4V, Qwen-VL, Gemini, Ollama…) right in the Models settings page; multi-backend fallback chain, double-layer caching, no config files.

0JavaScript2026-08-14