WeKnora
Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:AlloyPlane/dsh-eye-vision
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
Give text-only DeepSeek Harness models eyes — image understanding, OCR, and UI analysis through any OpenAI-compatible multimodal API.
Fork of dsh-free-vision (MIT) v1.0.1, with:
CUSTOM_MODEL_NAME is now forwarded, so arbitrary OpenAI-compatible endpoints (GPT-4o, Qwen-VL, GLM-4V, youtu-vita, vLLM, Ollama…) actually startYou send an image path
→ agent calls the image_understand tool
→ plugin spawns the luma-mcp vision engine (child process)
→ engine calls your multimodal API
→ text description comes back to the text-only model
The main model never needs image input support. Paste the image path, get answers.
image_understand tool registered on ctx.tools, visible to every session in the profilecustom provider — bring your own multimodal APILUMA_ALLOWED_DIRS) — read images from your workspace# from a local checkout
dsh plugin --profile web add D:/xd/dsh-eye-vision
# once published
dsh plugin --profile web add dsh-eye-vision
Restart dsh web. The tool appears as image_understand.
Settings file: ~/.dsh/free-vision.json (same path as upstream for drop-in compatibility):
{
"modelProvider": "custom",
"baseURLs": { "custom": "https://your-api.example.com/v1" },
"modelName": "your-vision-model",
"apiKey": "sk-...",
"allowedDirs": "D:/workspace",
"toolName": "image_understand"
}
Or use environment variables (fallback chain: settings file > env):
| Provider | Key env | Base URL env |
|---|---|---|
| custom | CUSTOM_API_KEY |
CUSTOM_BASE_URL + CUSTOM_MODEL_NAME |
| qwen | DASHSCOPE_API_KEY |
QWEN_BASE_URL |
| volcengine | VOLCENGINE_API_KEY |
VOLCENGINE_BASE_URL |
| siliconflow | SILICONFLOW_API_KEY |
SILICONFLOW_BASE_URL |
| zhipu | ZHIPU_API_KEY |
ZHIPU_BASE_URL |
| hunyuan | HUNYUAN_API_KEY |
HUNYUAN_BASE_URL |
allowedDirs: semicolon/comma-separated extra roots the engine may read images from (default: engine CWD + home directory).
看图:D:/path/to/screenshot.png
OCR:D:/path/to/document.png
UI 分析:D:/path/to/design.png (task_type: ui)
Tool arguments: image_source (local path / http(s) URL / data URI), prompt, task_type (auto|general|ocr|ui|debug|describe). PNG/JPG/WebP/GIF up to ~10 MB.
cd dsh-eye-vision
pnpm install # installs luma-mcp engine + MCP SDK
pnpm test
The engine patches in scripts/patch-luma.mjs re-apply automatically on install (idempotent, pinned to luma-mcp 1.7.1).
MIT — see LICENSE. Upstream: dsh-free-vision (MIT) by FuzzySoul.
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: image-understanding、multimodal、ocr、vision。