WeKnora
Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:gugu123a/dsh-tool-see-image
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
Give your text-only model eyes.
see_image routes an image to a configurable vision model (default: Zhipu GLM-4V-Flash, free) and relays its description back to your DeepSeek Harness session.
🎖️ Featured in the community awesome-deepseek-harness-plugins list.
DSH's default text-only model (e.g. deepseek-v4-flash) can't see images. This plugin gives it eyes via a small, free vision model — no local GPU, no image-editing, no changes to your model.
see_image tool — pass any image path + a question; get a text description back./chat/completions endpoint (Zhipu GLM-4V-Flash by default, or SiliconFlow / Qwen2.5-VL, …).ctx.fs, respecting DSH's sandbox / observation policy.You: "Look at this image" ──► Text-only model (no vision)
│ calls see_image(path, question)
▼
This plugin (Host plane)
│ 1. ctx.fs resolves & reads the image (sandbox/observation policy aware)
│ 2. encodes it as a base64 data URL
│ 3. POST {baseURL}/chat/completions (OpenAI-compatible)
▼
Vision model (GLM-4V-Flash)
│ text description
▼
Text-only model ──► reports to you
Copy the plugin into your profile directory, e.g.
$DSH_HOME/profiles/web/plugins/dsh-tool-see-image/
($DSH_HOME is usually ~/.dsh).
Declare the dependency in $DSH_HOME/profiles/web/package.json:
"dsh-tool-see-image": "file:plugins/dsh-tool-see-image"
Then run pnpm install (creates a junction to the source under
profiles/node_modules).
Compose it into the profile in $DSH_HOME/profiles/web/cordis.patch.yml:
- insert:
- id: tool-see-image
name: 'dsh-tool-see-image'
config:
baseURL: https://open.bigmodel.cn/api/paas/v4
apiKeyEnv: ZHIPU_API_KEY
model: glm-4v-flash
Set your API key: create one at the Zhipu (bigmodel) console
(format id.secret), then set the environment variable (Windows example):
setx ZHIPU_API_KEY "your-key"
Restart your terminal, then restart dsh web (the web profile does not
hot-reload patch layers yet — tested).
Verify: in a new session the see_image tool should appear. Try it:
Use see_image to look at path/to/your/image.png
| Key | Default | Description |
|---|---|---|
baseURL |
https://open.bigmodel.cn/api/paas/v4 |
OpenAI-compatible endpoint; the plugin appends /chat/completions |
apiKeyEnv |
ZHIPU_API_KEY |
Env var name that holds the API key |
model |
glm-4v-flash |
Vision model id (free on Zhipu) |
maxTokens |
1024 |
Max output tokens. Note: glm-4v-flash caps at 1024 (higher returns 400 max_tokens参数非法; raise it if you switch to a bigger model) |
timeoutMs |
60000 |
Request timeout |
maxBytes |
15728640 (15MB) |
Per-image size limit |
prompt |
(Chinese detailed-description instruction) | Default question; the question argument takes precedence |
To use a different vision API, change these three keys, e.g. SiliconFlow:
config:
baseURL: https://api.siliconflow.cn/v1
apiKeyEnv: SILICONFLOW_API_KEY
model: Qwen/Qwen2.5-VL-32B-Instruct
- insert: ... tool-see-image ... block from cordis.patch.yml;Remove-Item profiles\node_modules\dsh-tool-see-image;profiles/web/package.json dependencies;dsh web.{ name, inject, Config, apply }, same shape as every DSH tool plugin;inject: ["tools", "fs"] — the tool registry and the sandboxed file service
are both Host-global services;ctx.fs (sandbox/observation policy applied), never raw
node:fs;required: true for required,
omit the required key for optional (required: false is rejected by
defineTool);exec.signal cancellation; errors are
model-readable.test/mount-test.mjs mounts this plugin line under a real Cordis Loader
(timer + system-prompt + tools + this plugin) and asserts that see_image
lands in the tool registry. Run:
$env:DSH_CHECKOUT="<your dsh install root, containing node_modules/@deepseek-ai>"
node test/mount-test.mjs
Expected output: tools.schemas() 含 see_image: true and === MOUNT TEST PASS ===.
The script temporarily links the plugin into the checkout's node_modules
(Windows junction / other-platform symlink) and cleans up afterwards — no
hardcoded local paths.
triz-workflow.png (a DSH Web GUI screenshot) →
HTTP 200 in ~6.8s, correctly read the UI text (search box / MCP settings /
Fetch / Filesystem / Sequential-Thinking).max_tokens cap is 1024 (the default of 2048 caused
a 400; the default has been fixed).Beyond the see_image tool (which reads an image by path), this repo ships a
patch script that lets you paste images straight into the chat box and
have them auto-converted to text:
$env:DSH_CHECKOUT="<your dsh install root, containing node_modules/@deepseek-ai>"
node scripts/patch-dsh-image-relay.mjs # apply (idempotent, auto-backup)
node scripts/patch-dsh-image-relay.mjs --check # status
node scripts/patch-dsh-image-relay.mjs --revert # rollback
pm2 restart dsh-web # then hard-refresh the browser
It patches three DSH packages (host-apiproxy, llm-deepseek, client-ui) so that:
【图片:...】 description from GLM-4V-Flash instead of raw pixels;~/.dsh/cache/image-relay/), and failures
degrade gracefully within 8s.Requires ZHIPU_API_KEY. The image cache itself is portable and stays under
~/.dsh/cache/image-relay/; only the DSH checkout must be supplied explicitly.
Re-run the script after any npx dsh upgrade —
the patch is lost when the npm cache is refreshed.
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: image-recognition、vision。