返回目录
其他 插件

dsh-reset-handoff

nicecx/dsh-reset-handoff

DSH never restarts itself: host plugin that hands reset requests to an external ops agent (e.g. Hermes) via a versioned JSON protocol — preflight → gate → restart → health-check → recover → deliver back

Stars
0
Forks
0
Issues
0
更新
22 天前

PROJECT TOPICS

项目标签

INSTALL REFERENCE

安装参考

未验证
dsh plugin --profile web add github:nicecx/dsh-reset-handoff

该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。

PROJECT README

README

dsh-reset-handoff

DSH never restarts itself. A host plugin that hands reset requests to an external ops agent over a versioned JSON protocol — preflight-snapshot → restart → health-check → recover — then delivers the result back to the requesting session after reboot.

Why

A long-lived DeepSeek Harness (DSH) instance needs to restart for many reasons: reload plugins/config, apply settings, recover from a wedged state. But the agent inside DSH should not restart DSH itself:

  • restarting kills the very process that issued it, so the agent has no chance to see the outcome;
  • the agent cannot see what business is running (other live sessions, the relay channels, pending jobs);
  • if the reboot fails, nobody is left to diagnose and recover.

The safe pattern is a handoff: DSH writes a request, a separate, independent ops agent (here: Hermes Agent) reads it, runs the restart with preflight/health/recovery, and writes a result that DSH reads back after it comes up.

How it works

[DSH]  agent calls reset_handoff(reason)
         │  writes request.json (JSON protocol)
         │  (optional) triggers the external executor
         ▼
[ext]  ops agent reads request.json
         │  1. preflight — snapshot live sessions, relay state, pending jobs
         │  2. GATE     — pre-restart maturity gate (see below)
         │  3. restart  — restart the dsh web service (macOS launchd)
         │  4. health   — poll http://127.0.0.1:3080 until 200 (with timeout)
         │  5. recover  — verify relay/auth-proxy self-heal, list interrupted sessions
         │  6. result   — write result.json (status done/failed + per-stage detail)
         ▼
[DSH]  after reboot, the plugin reads result.json and delivers a readable
         summary back into the requesting session (followup), so the agent
         that asked can resume its interrupted work.

Pre-restart maturity gate

The reference executor refuses to restart unless it is safe to do so. It checks:

  1. No pending approvals/questions — relay pending.json has an empty pending list (a restart would otherwise drop the approval stack).
  2. Enough free disk — at least MIN_FREE_DISK_MB (default 500 MB).
  3. Cooldown — at least RESTART_COOLDOWN_SEC (default 60 s) since the previous restart, to break crash loops.

If any condition fails, the executor writes result.json with status: "failed", restart: { ok: false, gated: true }, and a gate array listing each failed condition with its detail — and does not restart. The requesting agent (or user) sees the exact reason and decides when it is safe to retry.

Tools

Tool Purpose
reset_handoff(reason, scope?) Submit a reset request to the external ops agent. Never restarts DSH in-process.
reset_status() Query the latest request and its result (read-only).

Both tools are registered host-wide, so every session's agent can call them when a reset is needed.

Protocol (v1)

The plugin and the executor are decoupled — they only share two JSON files under ~/.dsh/reset-handoff/ (override with DSH_RESET_HANDOFF_DIR):

request.json (written by DSH):

{
  "schema": "dsh-reset-handoff/request",
  "version": 1,
  "id": "<uuid>",
  "requestedAt": "2026-08-30T12:00:00+08:00",
  "reason": "重新加载插件配置",
  "sessionId": "<requesting session id>",
  "requester": "dsh-reset-handoff",
  "scope": { "restartDshWeb": true, "healthCheck": true, "recoverInterrupted": true }
}

result.json (written by the executor):

{
  "schema": "dsh-reset-handoff/result",
  "version": 1,
  "requestId": "<uuid>",
  "status": "done",
  "startedAt": "...",
  "finishedAt": "...",
  "preflight": { "liveSessions": ["..."], "relay": { }, "hermesJobs": ["..."] },
  "restart": { "ok": true },
  "health": { "ok": true, "checks": [ { "name": "dsh-web http :3080", "ok": true, "detail": "200" } ] },
  "recovery": { "resumed": ["..."], "report": "..." },
  "recoveryAction": {
    "ok": true,
    "attempts": [ { "attempt": 1, "restart": { "ok": true }, "time": "..." } ],
    "diag": { "keyErrors": [], "logTail": "..." }
  },
  "gate": [ { "name": "relay 无待审批/待回答诉求", "ok": true } ]
}

Executor recovery contract (the part that makes "recover DSH itself" real): if the health check fails after restart, the executor must attempt recovery, not just report failure:

  1. Diagnose — read the dsh web error log tail and extract key errors (loader failures, missing deps like undici, EADDRINUSE multi-instance, OOM).
  2. Retry — restart up to MAX_RESTART_ATTEMPTS (default 3) times with a cooldown between attempts.
  3. Observe — after each restart, wait an initialization window (default 120 s) before judging success.
  4. Report — write the outcome in recoveryAction (attempts + diag), so the requesting agent and the human see why it failed and how many tries were made.

Any executor that reads/writes these two files can drive the reset — Hermes, a custom script, a cloud function. The protocol is the contract.

Install

dsh plugin --profile <profile> add github:<owner>/dsh-reset-handoff

Optional executor trigger: configure triggerCommand so reset_handoff also wakes the external agent (default: none — the executor may poll request.json instead). See cordis.patch.yml for the config shape.

Executor (Hermes example)

A reference executor is included under hermes/reset_agent.py (pure Python, no deps). It is meant to live inside a Hermes profile (reset-agent) and be triggered by hermes cron run <job>:

python3 reset_agent.py            # run the five-step flow
python3 reset_agent.py --dry-run  # print the flow, don't restart

Requirements

  • DeepSeek Harness with the web profile (host plugin).
  • The external executor must be able to restart the dsh web service (macOS launchctl kickstart -k com.dsh.web, or equivalent for your OS/init).

Ops guardrails (read this before restarting anything)

Learned the hard way from a real 7-hour restart loop (2026-08-30). These rules are mandatory for any agent that manages a DSH host:

  1. Never create suicide/unconditional restart jobs. No launchctl submit jobs containing kickstart -k, no kill -9 on the DSH port, no unconditional restart logic. Restart only via the reset_handoff tool (which goes through the executor's gate) or DSH's own mechanism.
  2. Verify plugin dependencies before restart. The DSH loader resolves from ~/.dsh/profiles/web/node_modules — a missing transitive dep (e.g. undici) makes the whole plugin tree fail to load. Confirm deps exist and the tree loads cleanly before restarting.
  3. Health-check first, observe after. Before any restart: curl the port, check for single instance (lsof -i :3080). After restart: wait a 2-minute observation window and confirm the PID is stable before proceeding.
  4. Make plugins degrade gracefully. Missing config / bad fields should fall back to defaults with a friendly error, so calling agents never feel the need to edit plugin source or kill services to work around bugs.

License

MIT

CLASSIFICATION EVIDENCE

分类依据

项目类型插件
功能分类其他
规则置信度

系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: 无有效分类标签。