deepseek-harness
deepseek-ai
DeepSeek Harness: Everything is a Plugin.
PROJECT TOPICS
PROJECT README
Skill-trigger evaluation plugin for DeepSeek Harness (DSH).
An LLM judge recreates the exact DSH skill catalog prompt and decides, for each test query, whether the target skill should be triggered. The plugin reports accuracy, precision, recall, false-positive/negative rates, and a confusion matrix — a reproducible measure of how reliably a skill description routes matching queries (and how often it over- or under-triggers).
From the repo root:
dsh plugin --profile web add ./dsh-skill-eval
Then configure the judge model route in your profile/overlay cordis.patch.yml:
- id: skill-eval
config:
provider: <provider-id>
model: <model-name>
The provider must be registered in the DSH LLM runtime (the same one your profile uses for chat). The plugin validates the route at startup and warns if the provider is not yet registered.
Slash command:
/skill-eval <skill-name> [test-file]
Model-callable tool:
run_skill_eval(skill_name="<skill-name>", test_file="examples/dsh-plugin-eval.json")
test-file is optional; it defaults to examples/dsh-plugin-eval.json inside
the plugin package. Relative paths resolve against the plugin package directory.
A JSON array of { query, should_trigger } objects:
[
{ "query": "add a tool to the harness that persists across restarts", "should_trigger": true },
{ "query": "help me write a Python script for this CSV", "should_trigger": false }
]
category is optional and reserved for future use.
ctx.skills.snapshot).<system-reminder> +
<available_skills> + normalized/truncated/escaped descriptions).YES/NO answer.The evaluation measures the judge model's routing accuracy for the given
skill description. Swap provider/model in the config to test other judges.
npm run check # syntax check for every JS file
npm test # node:test, including official catalog fidelity and mock ctx tests
npm run smoke # 51 pure-function smoke assertions
npm pack --dry-run # published file list check
bash scripts/mount-smoke.sh # real DSH mount smoke in a scratch home
The catalog fidelity fixture pins the official dsh-tool-skill@0.1.0-rc.6
template. After a DSH upgrade, refresh the fixture from a local official
install and review the diff:
node scripts/refresh-catalog-fixture.mjs <path-to-dsh-tool-skill/lib/index.js>
index.js — plugin entry: registers the run_skill_eval tool and the
/skill-eval command.runner.js — catalog reproduction, judge LLM call, verdict parsing.catalog.js — pure functions: catalog message rendering and verdict parsing.llm-helpers.js — dependency-free BlockAssembler, createUserMessage,
deepFreeze (mirrors the official dsh-llm pattern).parser.js — test-case JSON loading and validation.metrics.js — confusion matrix, metrics, and markdown report formatting.examples/ — default test cases.CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: 无有效分类标签。