WeKnora
Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:GXX182/dsh-vision-bridge
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
A vision layer for text-first models in DeepSeek Harness
简体中文 · Install · Configure · Security
dsh-vision-bridge is an installable DeepSeek Harness bundle that adds image understanding to text-first model routes. It preserves the Harness model list, adds a small glasses control to eligible models, and delegates image analysis to a separately configured vision provider.
The plugin supports Gemini-native, OpenAI-compatible Chat Completions/Responses, and Anthropic-compatible Messages APIs. The selected vision provider returns bounded text analysis; image blocks are never forwarded directly to the active upstream model.
| Control | Result |
|---|---|
| Gray glasses | Vision Bridge is disabled for that model. |
| Blue glasses | Vision Bridge is enabled for that model. |
| Click glasses | Toggle and remember the preference only; the selected model does not change. |
| Click model name/row | Select the model. Blue routes through Vision Bridge; gray uses the normal upstream route. |
| Hover glasses | Show the vision provider and model that will perform image understanding. |
The glasses control appears only when a matching bridge route exists and the upstream model is text-only or has unknown image capability. The preference is stored per upstream model in the local Harness client.
vision_bridge; the tool reads the latest session image through Harness attachment services.Explicit workspace paths remain supported through image_paths; they are resolved with Harness filesystem policy.
0.1.0-rc.5 or a compatible 0.1.x release^22.19 or >=24The backward-compatible default profile uses GOOGLE_API_KEY, the Gemini native endpoint, and gemini-3.6-flash.
Install the highest semantic-version release tag:
dsh plugin --profile web add "github:GXX182/dsh-vision-bridge#semver:*"
#semver:* selects the newest matching GitHub version tag. Pin an exact tag such as #v0.2.0 when reproducible installs are required.
If Harness is started with npx:
npx @deepseek-ai/dsh plugin --profile web add "github:GXX182/dsh-vision-bridge#semver:*"
npx @deepseek-ai/dsh plugin --profile web list
npx @deepseek-ai/dsh web
The persistent web profile is stored under ~/.dsh/profiles/web unless DSH_HOME is changed.
npm install
npm run build
dsh plugin --profile web add .
dsh --profile web --dump-config
dsh --profile web
The config dump should contain a dsh-vision-bridge layer and a vision-bridge row.
Open Settings → Plugins → Plugin configuration → Image understanding.
Each provider profile contains:
auto, Gemini, OpenAI compatible, or Anthropic compatible);Adding a provider first verifies its model-list endpoint. The credential is stored through Harness credential services; the complete key is never returned to the browser. Switching providers immediately updates the glasses tooltip, including the selected provider name and model.
With apiFormat: auto, complete endpoint paths take priority, followed by official hosts and version paths:
:generateContent, /v1beta, or generativelanguage.googleapis.com → Gemini native/v1/messages or api.anthropic.com → Anthropic compatible/chat/completions or /responses → OpenAI compatibleSet the format explicitly when an ambiguous relay uses Gemini or Anthropic semantics. The plugin never probes several protocols by resending the same image.
The schema defaults work without editing the patch. To override them, replace the inserted row's complete config in the profile cordis.patch.yml:
- id: vision-bridge
config:
bridgeProvider: deepseek-vision-bridge
upstreamProvider: deepseek-official
apiKeyEnv: GOOGLE_API_KEY
apiFormat: auto
baseURL: https://generativelanguage.googleapis.com/v1beta
model: gemini-3.6-flash
maxImages: 8
maxImageBytes: 8388608
maxTotalImageBytes: 12582912
maxQuestionChars: 8000
maxOutputTokens: 4096
maxResponseBytes: 524288
maxAnswerBytes: 131072
timeoutMs: 90000
Conversation attachments normally require no explicit tool instruction. For a workspace file, ask the agent:
Use
vision_bridgeto inspectscreens/settings.png. List the visible controls and validation errors.
Code Mode can call await tools.vision_bridge(...). Omit image arguments for the latest conversation attachment, use attachment_ids for specific session images, or use image_paths for workspace files.
apiFormat is explicit.llm/stream middleware observes both the bridge request and its delegated upstream request.npm install
npm run verify
npm pack --dry-run
Built lib/ artifacts are intentionally committed for direct GitHub installation.
MIT
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: vision、vision-ai、vision-api、vision-language、vision-language-model、vision-language-models、visionos。