deepseek-harness
deepseek-ai
DeepSeek Harness: Everything is a Plugin.
ankye/dsh-client-vision
Give your DeepSeek Harness agent eyes. dsh-client-vision is a screen-capture + external image-recognition plugin for DeepSeek Harness: the agent takes a screenshot (or points at any image), hands it to a vision-capable model through a pluggable channel, and gets back plain text it can actually act on — no multimodal model required.
PROJECT TOPICS
PROJECT README
English | 中文
Give your DeepSeek Harness agent eyes. dsh-client-vision is a screen-capture + external image-recognition plugin for DeepSeek Harness: the agent takes a screenshot (or points at any image), hands it to a vision-capable model through a pluggable channel, and gets back plain text it can actually act on — no multimodal model required.
This revision requires DeepSeek Harness core 0.1.2-alpha.5 or later within the 0.1.x line. It uses the Settings service API introduced in that core release.
fullscreen / window (with live window enumeration) / region / interactive — grab the browser, a game window, or one corner of the screen.gpt channel ships ready to use; adding Claude, Gemini, or a local model is one analyze() implementation + one registry line — the three tools never change.credentials store (VISION_GPT_API_KEY) — never in settings files, logs, or the conversation transcript.code, standard, cordis, minimal — every agent sees the tools. No preset switching.pnpm publish / tarball).| Tool | What it does |
|---|---|
take_screenshot |
Capture the screen: fullscreen (primary display), window (by id from list_windows), region (x, y, width, height), interactive (user selection), android (adb device/emulator), or ios (booted simulator). Returns the PNG path + dimensions. |
list_windows |
Enumerate on-screen windows (id, app, title) — macOS CGWindowList, Windows Get-Process main handles, Linux X11 (wmctrl/xprop) — pick the browser or game window to capture. |
analyze_image |
Submit an image (a path, or the most recent screenshot) to the configured vision channel and return a plain-text description. |
view_image |
One-shot "look at this": capture the screen (or use image_path) and recognize it through the active channel. The screenshot is rendered as an image card in the Web conversation, while the model context receives only the plain-text description — the image bytes never enter the model context. |
| Platform | Capture backend | Window enumeration | Extra requirements |
|---|---|---|---|
| macOS | screencapture (system) |
Swift CGWindowList |
Screen Recording permission on first use |
| Windows | PowerShell System.Drawing (system) |
Get-Process main window handles |
PowerShell System.Drawing |
| Linux | ImageMagick import |
wmctrl + xprop |
ImageMagick (convert/identify), wmctrl, x11-utils |
mode=interactive (system selection UI) is macOS-only; on Windows and Linux
use mode=region with explicit coordinates.
| Mode | What it captures | Requirements |
|---|---|---|
android |
A connected Android device or emulator screen | adb on PATH with a device online (adb devices); works from any host. With several devices online, pass device=<serial>. |
ios |
The booted iOS simulator | macOS host with Xcode (xcrun simctl) |
vision namespace)Configured in Settings → Plugins → Plugin configuration → Vision:
| Field | Meaning |
|---|---|
Endpoint (baseUrl) |
Domain + optional path prefix; /chat/completions is appended. e.g. https://api.example.com/v1 |
| Channel | The active recognition backend (currently gpt). |
| Model | gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra |
| API key | Stored through the harness credentials service as VISION_GPT_API_KEY; the literal never leaves your machine. |
| Channel | Backend | Model | API key |
|---|---|---|---|
gpt |
OpenAI-compatible /chat/completions |
gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra |
required (e.g. VISION_GPT_API_KEY) |
zhipu |
Zhipu GLM-4V, OpenAI-compatible /chat/completions |
glm-4v-plus / glm-4v-flash |
required (e.g. VISION_ZHIPU_API_KEY) |
ollama |
local Ollama /api/chat (default http://localhost:11434) |
llava / llava-llama3 / bakllava / moondream / qwen2-vl / minicpm-v (or any installed vision model) |
none |
Pick the channel in Settings → Plugins → Vision; the model dropdown follows
the channel and the API-key control is hidden for ollama. For ollama the
base URL defaults to http://localhost:11434 and the model to llava when
left blank.
model → analyze_image(image, prompt)
│ reads vision.channel
▼
channels/<id>/analyze() ← one implementation per backend
│
gpt: POST {baseUrl}/chat/completions (image_url data URL)
claude / gemini / local: … ← add yours here
Adding a channel is deliberately small:
// src/channels/<id>/index.ts
export async function myAnalyze(ctx, call): Promise<string> {
// call.imageB64, call.mime, call.prompt, call.config, call.signal
return await fetchYourVisionApi(...)
}
// src/channels/index.ts — one registry line
export const channels = {
gpt: { label: 'GPT', analyze: gptAnalyze },
myChannel: { label: 'My Channel', analyze: myAnalyze },
}
The tools (take_screenshot / list_windows / analyze_image) and their schemas never change.
dsh plugin add installs the packages into your profile; each package declares dsh.bundle, so the rows mount automatically — no patch rows, no repo edits.
0.1.2-alpha.5 or later within the 0.1.x line, plus dsh and pnpm on PATH.a. From this repository (recommended until published to npm):
dsh plugin --profile web add \
file:/path/to/dsh-client-vision/packages/tool-vision \
file:/path/to/dsh-client-vision/packages/ui-vision
b. Tarball:
cd packages/tool-vision && npm pack
cd packages/ui-vision && npm pack
dsh plugin --profile web add file:/path/to/deepseek-ai-dsh-tool-vision-0.1.0-rc.7.tgz \
file:/path/to/deepseek-ai-dsh-client-ui-vision-0.1.0-rc.7.tgz
c. npm registry (after publishing):
dsh plugin --profile web add @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
A
[WARN] Issues with peer dependenciesmessage is expected and safe to ignore — the peers come from your deployment's own bundles at runtime.
node -e "console.log(JSON.stringify(require(process.env.HOME + '/.dsh/profiles/web/package.json').dsh.profile.bundles))"
# should list dsh-tool-vision and dsh-client-ui-vision
Restart the harness, then Settings → Plugins → Plugin configuration → Vision: set the endpoint, model, and your own API key (VISION_GPT_API_KEY), save.
Ask the agent to "look at the screen" — it should call take_screenshot → analyze_image and describe what it sees.
dsh plugin --profile web remove @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
If you run a fork of deepseek-harness (not the official deployment), you can drop the packages into the monorepo instead:
cp -R packages/tool-vision <harness>/packages/vision/tool-vision
cp -R packages/ui-vision <harness>/packages/client/ui-vision
Then add both to apps/cli/package.json (workspace:^), add ./packages/vision/tool-vision to tsconfig.host.json and ./packages/client/ui-vision to tsconfig.client.json, pnpm install, build (tsdown host + client passes), and restart.
take_screenshot / list_windows / analyze_image.@deepseek-ai/dsh-tools, …) resolve from your deployment. lib/ ships prebuilt, so npm pack works immediately.tsconfig.json files are standalone; the harness monorepo's build pipeline (including the client-bundle tsdown.config.ts) applies in Option A..credentials.yaml.MIT
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: 无有效分类标签。