dsh-gemini-pool
qikairo7
多账号 Google Gemini 提供商(DSH 插件):订阅额度即用即扣,账号池自动调度;主模型没有视觉也能贴图看图(池内 Gemini 自动转述),另有生图工具与中英双语设置页
DSH-PLUGIN STORE / LIVE CATALOG
聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。
70 个项目,匹配「multimodal」
qikairo7
多账号 Google Gemini 提供商(DSH 插件):订阅额度即用即扣,账号池自动调度;主模型没有视觉也能贴图看图(池内 Gemini 自动转述),另有生图工具与中英双语设置页
reimu-create
DSH plugin: text-only models (e.g. DeepSeek-V4) automatically see images via a vision model. Official surface-replace, cache-friendly, human transcript untouched. 纯文本模型自动识图桥
shaoqiuyuavailable
Scene-aware vision routing layer for DeepSeek Harness (dsh): decides which engine/backend an image should go to (chat/UI/table/code) before other vision plugins route it. Switch-gated, front-loaded, never touches other plugins' tools. 场景级识图路由层。
zouyuanqing
Native interactive visual-reasoning plugin for DeepSeek Harness: precise pixel grounding (SOM grid / zoom / annotate / measure / diff / color / OCR) + MiMo V2.5 multimodal backend, zero external MCP servers.
173787247
Save a Windows clipboard image into a WSL file for multimodal chat.
2472786266-spec
DSH DevKit: multimodal gallery + multi-agent supervision console (DeepSeek Harness dynamic Cordis plugin)
AlloyPlane
该仓库暂未提供项目说明。
Harvey-Will
DeepSeek Harness 图像理解插件 · 8 种分析模式 · 支持任意API接口 · 内置免费视觉模型 | DeepSeek Harness vision plugin · 8 analysis modes · works with any OpenAI- or Anthropic-compatible API · built-in free vision model
Leeminjing
Give text-only DeepSeek models on-demand vision: upload images, DeepSeek answers by calling a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).
TwistedRiCen
DSH-native Vision Evidence bridge for text-only reasoning models with native image attachments and strict multi-image validation.
WardLu
Open-source MCP vision server that gives text-only LLMs and AI agents image understanding, OCR, visual analysis, UI inspection, and multimodal capabilities.
aijunjiang
Give your DSH agent eyes via any OpenAI-compatible vision model - 11 provider presets (Doubao/Qwen-VL/GLM-V/OpenAI/Gemini/Ollama...), capability checkboxes that inject live prompt guidance, and analysis of images the user drops into the chat; the agent writes its own observation prompt, and base64 never enters its context.
ankye
Give your DeepSeek Harness agent eyes. dsh-client-vision is a screen-capture + external image-recognition plugin for DeepSeek Harness: the agent takes a screenshot (or points at any image), hands it to a vision-capable model through a pluggable channel, and gets back plain text it can actually act on — no multimodal model required.
kanchengw
Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, and GUI automation, through layered evidence memory and cache.
wulusai2333
DeepSeek Harness (DSH) native plugin — describe_image tool: a vision bridge (image → mimo-v2.5 → text description) over the ctx.fs / ctx.credentials seams
xisheng687
Bring your Grok subscription into DSH as an ACP subagent, extending native images with audio and video tools.
xsoc1
Eyes for text-only DeepSeek: view_image tool (local Ollama or any OpenAI-compatible VLM) + chat image-attachment bridge — paste/drop images in the chat and the model can see them.
yauntyour
DSH 多模态输入插件:为不同类型的文件(图片 / 视频 / 音频 / 文本)配置独立的处理模型链,在文件进入会话模型之前,先用预设模型把它处理成 Prompt Tokens(文本),再交给会话模型。插件在 DSH 设置中新增独立的 Multimodal 页面。
EmmanuelMartinez
Attach an image to the DeepSeek Harness composer and the session model switches itself to a vision model — then restores your previous model when the image is removed. Material Design 3 button, native file picker, thumbnails. MIT.
Lab-sku
明眸 VisionBridge - 自研视觉桥:瞎子模型收图时自动调用视觉模型识别
baldovinmarques391-design
GLM Vision plugin for DSH: image translation for non-multimodal models via GLM-4V-Flash
ch1bug
Xiaomi MiMo-powered voice for DeepSeek Harness: browser 🎤/🧠/🔊 UI, voice_transcribe/voice_understand/voice_speak tools, configurable voice map (preset/voicedesign/voiceclone). Fork of zhuiyueya/dsh-voice (MIT), Settings pattern from Anionex/dsh-vision-toolkit (MIT).