deepseek-harness
deepseek-ai
DeepSeek Harness: Everything is a Plugin.
PROJECT TOPICS
PROJECT README
The plugin follows the DSH language setting (Chinese and English in DSH 0.2.0-rc.2), including configuration, status messages and plugin-list metadata. Language-pack locales use the host fallback chain. Switching languages preserves unsaved settings; there is no separate plugin language selector.
Open Plugins → Installed → dsh-tool-vision from the homepage sidebar to configure and save this plugin. The page uses the official plugins.bundle.config interface, without a duplicate entry in global Settings. Web and Desktop share the page. This version requires DSH 0.2.0-rc.2 or a later 0.2.x host; existing configuration is retained.
Choose Automatic images and previews to enable image bridging, automatic admission and thumbnails together, or Vision tools only for explicit tool calls. Enter the endpoint, key and model on the main page; model overrides and request tuning live under Advanced. Existing mixed switch values appear as Custom combination and are retained until you change them. Selecting a mode edits a draft; click Save to apply it. Only edited fields are written, blank keys are preserved, and rejected writes keep the draft.
GitHub: Scorp1o117/dsh-tool-vision · npm: dsh-tool-vision
Part of the DeepSeek Harness Enhancement Suite — Vision · Soul/Persona · Long-term Memory · Plugin Marketplace.
External vision model for DeepSeek Harness.
DSH 0.1.1 adds native image input for DeepSeek's vision catalog. This plugin
remains useful when you want a separate OpenAI-compatible vision endpoint,
pixel-level image tools, screenshots, or a text-model bridge. The harness
derives every model request strictly from the session log (llm/stream
requests must equal the durable derivation — the agent-loop invariant), so the
bridge keeps its conversion inside that durable path:
inspect_image tool — sends an image (local file, or http(s) URL) to
any OpenAI-compatible /chat/completions endpoint that supports
image_url content parts, and returns the vision model's textual answer
into the agent loop.inspect_image
hints before they enter the durable log, on the agent/pre-step
waterfall (the one seam where the harness lets a plugin replace the
messages of a proposed step). Images already logged by an older version
are repaired lazily with a surface replace on the session's first
pre-step. Only models listed in multimodalModels receive image blocks
directly; a model's declared inputModalities are never consulted,
because profiles routinely declare input: [text, image] on text-only
models just to pass the harness's prompt-admission check.inspect_image.tool-vision namespace (API endpoint, write-only key, model, bridge
options) in the active Profile patch; changes hot-apply without a restart. The API
key lives in that patch. Mount by package name
(name: 'dsh-tool-vision') so the web client bundle is discovered.Verified with DSH 0.1.7-rc.2 (Web) and 0.2.0-rc.2 (Desktop runtime) in isolated profiles. The Desktop app uses its own desktop profile. Other DSH prereleases remain unverified.
Use the Desktop-installed dsh command (Application → Manage dsh Command), or the app’s Plugins page. Then install into the Desktop profile:
dsh plugin --profile desktop add dsh-tool-vision@0.9.7
Restart the Desktop app to load the client bundle. Desktop keeps its profile under $DSH_HOME/profiles/desktop.
sha256:... no longer create NTFS alternate data streams and zero-byte visible files.bridgeExportDir.0.1.5-rc.3 Profile.Mount in a profile patch ($DSH_HOME/profiles/<name>/cordis.patch.yml):
- insert:
- id: tool-vision
name: 'dsh-tool-vision' # after: pnpm add dsh-tool-vision in the profile
config:
baseURL: 'https://api.openai.com/v1'
apiKeyEnv: 'VISION_API_KEY'
model: 'gpt-4o-mini'
Or load it from a local path without npm:
- id: tool-vision
name: './plugins/dsh-tool-vision/index.js'
| Field | Default | Meaning |
|---|---|---|
enabled |
true |
Master switch (v0.8.0). Off unregisters everything this plugin contributes — inspect_image, the 14 vision_* tools, the image bridge, the preview route and the image-capability declaration. The settings section stays mounted so the switch can turn it back on. Hot-applies; no dsh restart. |
baseURL |
https://api.openai.com/v1 |
OpenAI-compatible API base URL. |
apiKey |
'' |
API key (takes precedence over env). |
apiKeyEnv |
VISION_API_KEY |
Env var holding the key. |
model |
gpt-4o-mini |
Vision model id. |
maxTokens |
1024 |
Max output tokens. |
timeoutMs |
60000 |
Per-request timeout. |
maxImageBytes |
10MB |
Largest accepted local image. |
description |
default | Tool description shown to the model. |
bridgeTextOnly |
true |
Bridge pasted images to text hints on models that cannot see images. |
bridgeExportDir |
temp | Export dir for bridged images (os.tmpdir()/dsh-vision-bridge). Filenames use a portable hash of the attachment ID. |
multimodalModels |
[] |
Model list (comma-separated). Each entry is matched case-insensitively against the full id, its bare id after the last /, and provider/id, with * / ? globs (*vl*, deepseek/*). What the list means is set by the mode below. |
multimodalListMode |
whitelist |
List mode (v0.9.0). whitelist: listed models receive image blocks directly (the historical behaviour). blacklist: listed models are forced through the bridge — the correction layer for a model that claims image support it does not have. off: the list is ignored. An unknown value falls back to whitelist. |
autoDetectMultimodal |
true |
Auto-detect (v0.9.0). Decide from the current route's own declared inputModalities, then combine with the list (whitelist unions, blacklist subtracts). On by default: a text-only route is bridged, a multimodal one is treated like a whitelist member and gets images directly. The declaration is always read before this plugin's admission wrap, so bridgeAutoImage can never feed its own claim back in as evidence. Set false for the hand-maintained "list only" behaviour. |
probeResults |
{} |
Measured verdicts (v0.9.0): "provider/model" → "yes"/"no", written by the vision_probe_model tool — do not edit by hand. A measurement outranks a declaration (a real request beats a claim) but not multimodalModels (explicit human intent has the last word). |
bridgePreview |
true |
Inline preview for bridged images: thumbnail above the hint text in the user bubble (click to zoom). |
bridgePreviewScanIntervalMs |
2000 |
Fallback scan interval for the preview scanner (ms); 0 disables the fallback. |
bridgePreviewHideHint |
true |
Hide the bridged hint text once the preview image has loaded (kept on failure — safe degradation). |
bridgeAutoImage |
true |
While the bridge is on, report image input capability for every model to the host admission gate, so pasted images are accepted on text-only models without hand-editing provider configs. |
sendSessionHeader |
true |
Send a stable session-id header on vision requests. OpenCode Go and similar gateways require x-opencode-session (one stable id per conversation); requests without it may error from 2026-09-06. |
sessionHeaderName |
x-opencode-session |
Header name carrying the session id. |
sessionId |
'' |
Fixed session id for calls without a dsh session context; empty = auto (current dsh session id, else a stable per-process random id). |
bridgeAutoImage is disabled, declare
image input on the models you paste images onto, so the harness admits
image messages (pi-ai style):llm-pi-ai:
providers:
your-provider:
models:
- id: deepseek-v4-flash
input: [text, image]
- id: tool-vision
name: 'dsh-tool-vision'
config:
multimodalListMode: whitelist # default: listed models get images directly
multimodalModels: ['mimo-v2.5', 'grok-4.5']
Then pasting an image while on a text-only model stores a hint like
[User sent an image, exported to: <path>. Inspect it with the inspect_image tool...]
in the transcript (the pasted image no longer renders as pixels in that
message), and the agent inspects it through the configured vision endpoint.
Why not
llm/stream? The harness freezes every request and the agent-loop invariant fails any request whose messages diverge from the session-log derivation (log-reconstruction desync), and this cordis waterfall'snext()cannot replace request arguments. Theagent/pre-stepwaterfall is the supported seam: its decision messages become the durable log, so the invariant stays satisfied.
Key resolution order: config.apiKey → process.env[apiKeyEnv] →
process.env.OPENAI_API_KEY.
On text-only models, pasted images become [User sent an image...] hint
text in the transcript. With bridgePreview enabled (default), the browser
half renders those hints as inline thumbnails in the display layer only:
Esc to close;bridgePreviewScanIntervalMs);bridgePreviewHideHint on, the hint text is
hidden once the image has loaded, leaving just the image; on load failure
the text stays (safe degradation — never "no image AND no text");\u200b[bridge]), so ordinary user text that happens to contain
"exported to:" is never misidentified;inspect_image chain are untouched.Preview images are served by the same-origin loopback route
/plugins/dsh-tool-vision/image: read-only access to the bridge export
directory, localhost-only Host, image extensions only, ≤ 20MB per file,
path-traversal protected.
inspect_image| Arg | Required | Meaning |
|---|---|---|
path |
✅ | Image path (absolute, or relative to the current workspace) or http(s) URL. |
question |
– | Optional specific question about the image. |
detail |
– | auto / low / high resolution hint. |
Example endpoints (baseURL):
https://api.openai.com/v1 — gpt-4o, gpt-4o-minihttps://dashscope.aliyuncs.com/compatible-mode/v1 — qwen-vl-plus, qwen-vl-maxhttps://open.bigmodel.cn/api/paas/v4 — glm-4v-flash (free tier), glm-4v-plushttps://api.moonshot.cn/v1 — moonshot-v1-8k-vision-previewhttp://localhost:11434/v1 — llama3.2-vision (no key)Note for users
- This plugin is a standard profile bundle (
dsh.bundle.patch):dsh plugin --profile web add dsh-tool-visioninstalls and mounts it in one step — no manualcordis.patch.ymledits needed.- Settings changes hot-apply (no restart needed).
- Version 0.6.3 and newer require DSH
0.1.0-rc.7or newer and are tested against0.1.0-rc.7,0.1.0-rc.8, and0.1.1-rc.1.- DSH
0.1.0-rc.6users must pindsh-tool-vision@0.6.1, the last release carrying the legacy settings-allowlist compatibility patch.
14 vision_* tools driven by the same configured endpoint as
inspect_image (baseURL/apiKey/model) — no provider chain, no local models,
no extra settings:
| Tool | Purpose |
|---|---|
vision_describe |
Image Q&A / multi-image comparison (optional structured JSON) |
vision_ground |
Locate a target and return its ORIGINAL-pixel bounding box |
vision_detect |
Enumerate elements (buttons, inputs, icons…) with numbered boxes |
vision_crop |
Crop a pixel region to a PNG artifact |
vision_pixel_diff |
Per-pixel comparison: ratio, worst regions, heatmap, report |
vision_colors |
Dominant-color quantization for palette matching |
vision_ocr |
Verbatim text transcription (letters only — not scene analysis) |
vision_long_screenshot_ocr |
Chunked long-screenshot transcription into Markdown |
vision_trace |
Potrace vectorization into colored SVG (worker-thread, safe) |
vision_extract_foreground |
Solid-background removal → transparent PNG |
vision_html_screenshot |
Headless render of a local .html (network blocked) |
vision_screenshot |
Desktop capture (privacy-gated: enable desktopScreenshot in settings; Win: PowerShell / macOS: screencapture / Linux: import/scrot) |
vision_present |
Publish a generated image to the user via the host attachment store |
vision_materialize |
Copy an attachment/local image into the workspace as a real path |
Quality & safety details:
VISION_CONTENT_FILTERED
instead of a generic backend error.<workspace>/.dsh-tool-vision/.Requires sharp / potrace / puppeteer-core (declared as optional
dependencies: a failed platform install never blocks the plugin; missing ones
degrade lazily with an install hint and never break other tools).
vision_screenshot is privacy-sensitive and therefore not registered by
default — set desktopScreenshot: true in the tool-vision settings to
enable desktop capture.
The bridge answers one question: can the current model see images directly? v0.9.0 splits it into two independent inputs.
base = autoDetectMultimodal ? (route declares image) : {}
off → direct = base the list takes no part
whitelist → direct = base ∪ list the list only adds
blacklist → direct = base \ list the list only subtracts
A list hit always wins: the list is explicit user intent, so it outranks the model's own declaration — which is what makes it a usable correction layer.
A blacklist never degrades into "everything unlisted is direct". Its base set is the auto-detected one; with auto-detection off that base set is empty, so an unlisted model is still bridged. That is deliberate: the alternative lets one typo push images at a text-only endpoint.
Matching: mimo-v2.5, xiaomi/mimo-v2.5 and commandcode/xiaomi/mimo-v2.5
all address the same route; * / ? are globs; matching is case-insensitive.
v0.8.1 compared ids literally, so this README's own mimo-v2.5 example silently
did nothing on a route spelled xiaomi/mimo-v2.5 — fixed here, and the fix only
ever adds models to the direct set (no entry that used to force a model direct
stops doing so).
In the panel (Plugins → dsh-tool-vision Model):
llm.listProviders() + llm.listModels()), grouped by provider
and labelled with whether the route declares image input. Tick to add,
untick to remove; the text field above still takes globs by hand. Both edit the
same draft, persisted by Save — the picker head flags it as unsaved until then.provider / model, whether images go direct
or through the bridge, and why (list hit / auto-detect / default).Why not just a
<datalist>: a native datalist stays invisible until the user focuses the field and types, which made v0.9.0's first cut look like a dead panel. The list is now always visible, with the datalist kept as a typing aid.Unticking removes the entry that actually matched — if
mimo-v2.5in the list is what coversxiaomi/mimo-v2.5, unticking dropsmimo-v2.5rather than inventing a full id. Which entry hit is computed server-side with the same matcher the bridge uses (matchedEntriesin the payload), so the panel can never display a state that disagrees with the decision.
Read path (the easiest thing to get wrong here): autoDetectMultimodal MUST
read the value from before resolveModelInfo was wrapped, or the "image
support" that bridgeAutoImage stamps onto every model becomes evidence for
itself. unwrappedResolveModelInfo() enforces that, with a dedicated regression
test.
Candidates come from the plugin's own loopback route:
GET /plugins/dsh-tool-vision/models (loopback Host only, read-only, no-store).
It returns provider/model ids and one declared-capability boolean — no keys and
no endpoint addresses. It is registered on the plugin fiber rather than the
master switch's child fiber, so the panel keeps working while the plugin is off.
⚠️
inputModalitiesis a declaration, not a guarantee — profiles commonly setinput: [text, image]on text-only models just to pass the admission gate. Upstreamdsh-llm-pi-aimakes the same call for undeclared models, and its source says why: the two wrong answers do not cost the same. Under-claiming refuses the image before it is attached and names the model; over-claiming admits one the provider rejects mid-turn, after the message is already durable.That is why detection is on by default with three backstops: (1) the first time a route is promoted purely by its own declaration, the log says so and names the fix; (2) the panel always shows the current route, the decision and the reason; (3) listing that model under
blacklistmode forces the bridge back on.
npm test # server-side unit tests (no extra dependencies)
npm run test:render # panel render test (needs devDependencies)
npm run test:render loads the real client bundle in jsdom, drives the
real registration path (apply → slots.register → the component), feeds it
from the real server route handler, and asserts on the real DOM and the real
settings writes.
It is a separate command and deliberately not part of npm test: it needs
react / react-dom / jsdom, and a DOM test that silently skips when a
dependency is missing is a false comfort. Install with
npm i -D react@18 react-dom@18 jsdom.
Its reason to exist is specific: v0.9.0's first cut rendered the model list only
into a native <datalist> — every server-side unit test passed while the panel
looked completely dead. Nothing below the DOM can catch that class of bug.
vision_probe_model (v0.9.0)Each request defaults to a 60-second deadline and 2048 output tokens. Response
bodies are limited to 1 MiB while streaming; an oversized body stops the read
and yields unknown. Authentication, rate-limit and server errors remain
unknown even if their messages mention image input. Only HTTP 400, 415 or
422 image-rejection errors yield a negative verdict from an error response.
Every other signal here rests on what a model says about itself. This tool sends a real image to the route and reports what it does — the only ground truth in the plugin.
base = autoDetect ? route declares image : {}
a measured verdict (if any) overrides base measurement > claim
a list hit (if any) overrides everything human intent > measurement
Why one request is not enough (each of these was learned the hard way):
content when the thinking budget eats
max_tokens; the default is 2048 and the reader falls back to
reasoning_content.meituan/LongCat-2.0:free returned HTTP 200 for the image
request and answered "I can't see any image." Only an endpoint that
actively rejects the image part is a conclusive negative.So a probe ends in one of three verdicts: yes (control passed, both colors
correct), no (the endpoint rejected the image part, or answered without
reading it), or unknown (network, auth or protocol trouble — never turned
into a capability claim).
How it gets a route's endpoint and credential (no new configuration):
llm.listConfigurableProviders() names the provider's settings namespace and
path → settings.get(ns) resolves baseURL/apiKeyEnv/api →
credentials.resolve(apiKeyEnv) yields the secret (.value) — the same path
dsh-llm-pi-ai uses for a real call. Read-only, and neither the endpoint nor
the key ever appears in a probe result.
Linking: a verdict is written to probeResults and takes effect in the
bridge decision immediately — yes sends images directly, no forces the
bridge back on — regardless of the list mode. (That is also why it does not
silently write into multimodalModels: a blacklist list means the opposite,
so an automatic entry there would produce exactly the wrong result.) Each row in
the panel's picker carries a badge:
measured: reads images / measured: no image reading (blue / red), shown in
preference to the declaration;declares image badge;measured (a real image was read / NOT read).A route on an unknown protocol (e.g.
openai-responses) returnsunknownwith a reason, rather than guessing confidently withchat/completions.
The bug. v0.9.2's installDispatchImageAdmission satisfied DSH core's
LlmService.generate by declaring image on adapterCall.model.inputModalities.
However, dsh-llm-pi-ai's adapter executes a second internal guard inside its
streaming path (streamWithSnapshot):
const model = this.modelOf(snapshot, options.provider, options.model);
if (containsImage && !model.input.includes("image"))
throw new LlmError(`pi-ai model "${model.id}" does not support image input`, "UNSUPPORTED_CONTENT");
this.modelOf resolves from the adapter's own snapshot.models catalog, which
defaults to input: ["text"] unless the user explicitly declared
input: [text, image] in the Profile patch. When admitted by v0.9.2, the raw image
reached streamWithSnapshot and triggered UNSUPPORTED_CONTENT. Since the image
was already in the session's durable transcript, every subsequent turn failed.
The fix. On direct routes, installDispatchImageAdmission now additionally:
adapter.modelOf (when present) to include "image" in resolved.input;snapshot.models.getModel(route, model).input)
with "image";adapter.modelOf and reverts modified input arrays upon dispose.The bug. Every capability signal this plugin collects — probeResults,
multimodalModels, autoDetectMultimodal — decided whether the bridge
intercepts an image. None of them decided whether the model receives it.
LlmService.generate resolves modalities from the adapter, not from the plugin:
const adapterCall = await adapter.prepareCall(provider, model, signal);
modelInfo = this.normalizeModelInfo(registration, model, adapterCall.model);
if (modelInfo.inputModalities !== undefined
&& !modelInfo.inputModalities.includes("image")
&& projectedMessages.some((message) => contentHasImage(message.content)))
projectedMessages = projectImagesForTextModel(projectedMessages);
projectImagesForTextModel rewrites every image block into
[image omitted because this model accepts text only; attachment sha256:…]
before the adapter is called. So on a route the plugin had measured as
image-capable — where the bridge therefore stepped aside and let the image
through — the image was still destroyed, by a check that reads the adapter's
declaration and nothing else. The resolveModelInfo wrap (bridgeAutoImage) is
admission: it decides who may offer an image, and it never changes what the
adapter streams. Nothing in the plugin reached the layer that does.
The symptom is a route that probes yes, is listed in multimodalModels, and
still answers [image omitted …].
The fix. installDispatchImageAdmission wraps llm.registration — the
accessor the core itself calls — so every adapter, including one registered
later, answers prepareCall through the same routeDirectDecision the bridge
uses. One precedence rule, two seams, so the two can never disagree about a
route. When the decision is "bridge", nothing changes: the bridge already turned
the image into an inspect_image hint, and the core's projection stays as the
correct fallback for any image that still reaches dispatch.
Three properties worth stating, because each one is a test:
routeDirectDecision is now the single implementation of the precedence rule,
shared by the bridge and the dispatch path — two copies of it would drift, and
this one has three inputs.
Master switch. enabled, plus a one-click button at the top of the section
(Disable all / Re-enable). Registrations are effects on the cordis fiber
that makes them, so the plugin now puts every tool, the image bridge, the
preview route and the image-capability declaration in a child fiber: turning
the switch off disposes it, and all 15 tools leave the model's tool list
together. The settings section stays on the parent fiber, so the switch can turn
the plugin back on. No dsh restart.
Save-path fix. The form used to submit its 18 fields as parallel
scope.set()/unset() calls. Each write carries its own revision fence, a fence
behind the Host document is refused with settings/conflict, and a refused
write still resolves — the scope's contract is "settle after the write and any
recovery read", not "throw on refusal". The section therefore reported "Saved"
while the edits silently reverted, which reads as "settings cannot be saved at
all".
Writes are now one atomic mutate(), so the whole batch shares one fence and
one persistence decision, and the section is inspected after the write settles:
"Saved" only when the change is really there, otherwise "Write did not take
effect" plus a reload of the form. Hosts without mutate() fall back to
sequential writes (each waits for its predecessor, keeping the revision chain
intact) — never parallel.
Also removes three if (typeof scope.load === "function") scope.load() guards.
The SettingsScope seam has never had load() — it is getSnapshot /
subscribe / mutate / set / unset, and reads ride the shared describe
mirror driven by the Host's settings/document-updated. Those guards were dead
code that read like a refresh which never happened, and they made the missing
write verification look intentional.
inspect_image.agent/pre-step, so switching to a
multimodal model later does not turn it back into an image block. (The other
direction — multimodal to text-only — is repaired automatically by
repairLoggedImages.)MIT — bridge preview & integration: xing666173. Pixel vision tools ported from dsh-vision-router (© ysr666, MIT) with gratitude.
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: 无有效分类标签。