sandbase-harness
sandbaseai
Local-first, self-hosted AI agent runtime and MCP bridge with sandboxed sessions, memory, credentials, audit/replay, and a local Console.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:cat552/dsh-agent-quality-diagnosis
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
English | 中文
dsh-agent-quality-diagnosis is a DSH Web plugin that turns the current session's event log into an execution-quality report for agent work. It complements DSH Trajectory: Trajectory shows the event ledger, while this plugin answers whether anything still blocks delivery and what concrete action closes it.
The plugin is an MVP and currently provides:
质量诊断 tab registered through dsh-better-sidebar;ctx.sessions.get(sessionId).events and builds a tool-call-level report;ready, needs_action, or blocked status;defaultVerificationCommand Host config value for project-specific mutation verification;agent-loop, subagent scheduling, tool permissions, or model calls.Install dsh-better-sidebar in the web profile first:
pnpm dsh plugin --profile web add dsh-better-sidebar@latest
Install this plugin from the repository path:
pnpm dsh plugin --profile web add link:/path/to/your/dsh-agent-quality-diagnosis
Restart or refresh DSH Web, then open 质量诊断 in the sidebar.
The gallery shows the main report, summary metrics, finding details, and copyable actions/evidence.
| Overview | Report summary |
|---|---|
![]() |
![]() |
| Finding detail | Actions and evidence |
|---|---|
![]() |
![]() |
Optional Host config:
- id: agent-quality-diagnosis
name: 'dsh-agent-quality-diagnosis'
config:
defaultVerificationCommand: 'pnpm run check'
Host session-event reports detect:
The Web runtime fallback detects:
The report is a heuristic diagnostic signal, not an absolute quality score. Each finding carries user-readable impact, an action, confidence, and expandable evidence so users can judge false positives.
Overall status only considers open findings. A failed tool call with a later successful retry, a failed verification with a later passing verification, or a risky operation with an allowed-once approval record moves into the resolved section. Duplicate attempts and non-sensitive risky operations are observations unless they remain failed or lack the authorization evidence required for delivery.
Risky operations keep their original command as a related command, not as a command to rerun. The agent prompt asks for authorization, impact, and rollback explanation before any further high-risk operation.
All analysis runs locally inside the plugin. The plugin does not upload the session log and does not change execution behavior. If DSH session-event fields change, update src/analyzer/trace-builder.ts; rule logic stays under src/rules/.
The plugin carries its own lightweight checks because it lives outside the main workspace package list:
pnpm run check
pnpm exec tsc -p tsconfig.json
pnpm exec vitest run --config vitest.config.ts
node --check lib\index.js
node --check lib\client.js
Generate a Markdown report from the included sample sessions:
pnpm run demo:report
pnpm run demo:report -- fixtures/missing-verification-session.json
The fixtures cover a normal session, a failed tool call, and a mutation without later verification. They also include resolved examples for a failed tool retry and a failed verification followed by a passing verification.
Use the linked web profile, start DSH Web, then open 质量诊断 on a session with tool calls:
$env:DSH_HOME='<path-to-your-dsh-home>'
pnpm dsh --profile web
The tab should show Host session-event metrics when the Host route is available. If the route fails, it should show the Web runtime fallback and the Host error note.
After fixing an open issue, run refresh again. The tab should show the closed finding under 本次已关闭 when the previous cached report had that finding open.
Expanded evidence rows should provide a copyable locator that can be pasted into Trajectory search or a bug note.
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: agent-skills、agent-monitoring、agent-observability、developer-tools、diagnostics、multi-agent。