sandbase-harness
sandbaseai
Local-first, self-hosted AI agent runtime and MCP bridge with sandboxed sessions, memory, credentials, audit/replay, and a local Console.
Aik358/dsh-anchored-monitor
Real-time chain-of-thought anchoring monitor & intervention plugin for DeepSeek Harness: three-band (spec/mixed/react) fingerprint detection, L1 hint / L2 reset / L3 restart interventions, liquid-glass web overlay + rheostat bar, standalone monitor process with JSONL experiment logs.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:Aik358/dsh-anchored-monitor
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
English | 简体中文
In one sentence (一句话说清): it's a whip for DeepSeek V4 Pro. When the model falls from the focused, high-capability mode ("We need…" / "I will…") into the scattered, low-focus mode ("let me…"), the whip cracks — and pulls it back. 这是给 DeepSeek V4 Pro 加的一根鞭子——当它从「We need / I will」的高专注、高能力模式, 跌落到「let me」的发散、低专注、低效率模式时,就抽它一鞭,让它改回去。
Real-time chain-of-thought anchoring monitor & intervention for DeepSeek Harness. Watch the we / let's / let me fingerprint of every reasoning block, stay in the spec band, and pull the model back automatically when the trajectory drifts.

Don't touch a thing and it's simply a live gauge of how hard your model is thinking right now. Thinking intensity, three-band state and the ECG-style curve, tucked into the corner of your screen — a real-time visualization of your model's thinking efficiency / capability intensity.
Open the settings page and you can tune every parameter — and one-click Reset to defaults if you ever change too much.
You don't need to force-install dsh-anchored-standard.
The monitor never checks a preset's name. It reads every session's reasoning
blocks (~/.dsh/sessions/*/events.jsonl) and scores them purely by the
we / let's / let me fingerprint (spec band < 0.2, mixed 0.2–0.5, react ≥ 0.5).
So any preset that implements the anchored-standard discipline — anchor the
first turn with the Minimal persona (46-character persona + bash/str_replace_editor),
then plan in a collective "we" style without "let me" — pairs with this plugin.
Renamed / derivative presets work too: 梁神模式 (Liang-god Mode), liangshen,
and any other preset built on the same anchoring idea. The L2 reset payload is
fully self-contained (the Minimal 46-character persona + the dual tools ship in
this plugin's own config), so nothing is imported from another repo at runtime.
⚠️ One caveat: if the session was never anchored in the first place, it stays in the react band and L2 keeps firing — that's a fight loop, not monitoring. Keep interventions OFF (monitor-only) unless you're on the model this whip was tuned for.
When to turn interventions on / off — the whip is tuned for DeepSeek V4 Pro 0813:
This exact advice is shown right inside the plugin's panels:
ℹ Tip: keep interventions on only for DeepSeek V4 Pro 0813; turn them off (monitor-only) for other models.
DeepSeek V4 Pro conditions heavily on what the first request shows it. The
community measured the consequence: the official Minimal preset (46-character
persona + bash/str_replace_editor) anchors a collective "we" trajectory
and scores 99/96 on Project2, while the full Standard preset anchors an
actor-style "let me" trajectory and scores 91. Behavior is path-committed:
once anchored, expanding the tool catalog perturbs at most one reasoning block —
the mode never flips back on its own.
Anchored presets (dsh-anchored-standard) solve the bootstrap. This plugin solves what happens after anchoring: external factors (imperative hints, oversized injected context) can drag the trajectory out of the spec band, and nothing detects or repairs that drift — until now.
we, let's, we'll, we need, …) versus the
react marker (let me), over a sliding window.persona_ratio < 0.2 → spec, 0.2–0.5 → mixed (the unstable transition
band), ≥ 0.5 → react.| Level | Trigger | Action |
|---|---|---|
| L1 hint | entering the mixed band | injects a suggestive hint (never imperative — commands flip we→let me) |
| L2 reset | entering the react band | next request gets the 46-char Minimal persona + bash/str_replace_editor only; monitor window/baseline reset |
| L3 restart | L2 retries exhausted | recommends restarting the session |
llm/stream and pushes
reasoning-delta chunks (1s throttle), so the charts and the bar move
while the model thinks — not just after each turn. Counts are additive, so
window aggregates stay exact./api/push-text), so it can flag the two
degradations the reasoning fingerprint alone can't see: text_leak (thinking-style
prose leaking into the visible body — high text volume AND a high
let me/(we+let me) ratio) and streaming_stall (reasoning goes quiet while
text keeps flowing). Alert-only, never auto-intervenes; per-session cot
counters + guard_triggered events show up in the panel / dashboard.bash/str_replace_editor); L3
applies the same soft restart plus restart advice.config/*.yaml, validated against
config/schema.json). JSONL experiment logs, offline replay and grid-search
calibration scripts included.# 1) install the web plugin into your web profile (dsh CLI = pnpm forwarder)
dsh plugin --profile web add @a9i5k4/dsh-anchored-monitor
# 2) start the monitor process (default profile = production-safe `default`;
# `demo` is only for the accelerated L1→L2→L3 demo)
npx anchored-monitor
# 3) restart DeepSeek Harness (host bundle) and refresh the web GUI
You should now see 锚定监控 / Anchored Monitor in the left sidebar footer. Click it to open the glass panel; click the floating bar to expand/collapse.
To change the monitor address: edit ~/.dsh/anchored-monitor.json or
POST /api/anchored-monitor/config with { "monitorUrl": "http://127.0.0.1:9301" }.
git clone https://github.com/Aik358/dsh-anchored-monitor.git
cd dsh-anchored-monitor
npm install && npm run build
npm run demo:generate # 300 synthetic reasoning blocks
npm run dev -- --profile demo # monitor + dashboard on :9301
npm run demo:feed # live-feed the blocks (watch L1→L2 cascade)
Copy preset/ into ~/.dsh/.agent-presets/anchored-monitor to let the
harness agent push its own reasoning blocks to the monitor and execute the
L1/L2/L3 interventions inside the agent loop (pair it with an anchoring preset
such as dsh-anchored-standard for the first-round anchor).
The monitor process exposes (default http://127.0.0.1:9301):
| Method | Path | Description |
|---|---|---|
| GET | /api/overview |
sessions + selected snapshot + tail events in one call |
| GET | /api/sessions |
session summaries |
| GET | /api/sessions/:id |
full snapshot (history / interventions / baseline) |
| GET | /api/events?sessionId=&limit= |
tail of the experiment JSONL |
| POST | /api/push |
push a reasoning block {sessionId, text, sequence?, timestamp?} |
| POST | /api/push-text |
push a visible text chunk {sessionId, text, sequence?, timestamp?} (CoT guards) |
| POST | /api/sessions/:id/ack |
acknowledge an intervention |
| POST | /api/sessions/:id/reset |
trigger a manual L2 reset |
| GET | /api/stream |
SSE event stream |
The web plugin proxies these through /api/anchored-monitor/* (loopback-only).
| Command | Description |
|---|---|
npm run demo:generate |
generate a synthetic session JSONL (labelled, for calibration) |
npm run demo:feed |
live-feed blocks into a running monitor |
npm run replay -- --file x.jsonl |
offline replay with a summary report |
npm run calibrate -- --file x.jsonl |
grid-search window/weights/thresholds |
npm run preview:build |
snapshot the dashboard into a standalone HTML |
See docs/experiment-params.md for the full parameter reference. Everything is YAML — no hardcoded tuning.
What happens after an intervention — does the conversation stop? No. Every intervention auto-continues the task: L1 injects a hint and the next turn keeps going; L2 stops the running turn (soft restart) and immediately re-enters context — the next request continues the task under the Minimal persona + bootstrap pair; L3 does the same and adds restart advice.
Why did the chart feel slow / frozen?
Before 0.2.0, data only arrived once per finished turn. The host now streams
reasoning-delta chunks to the monitor every ~1s, so the curve moves while
the model thinks.
How do I disable interventions?
Use the panel-header switch (persisted), or set intervention.enabled: false
in the settings — monitoring continues, interventions stop.
Does L2 truncate the model's context? No. L2 replaces only the next request's persona section with the 46-char Minimal sentence and narrows the visible tool catalog to the bootstrap pair. All conversation history is preserved; the monitor only resets its own fingerprint statistics (invisible to the model). The plan-mode and all other sections are kept, because dropping them causes re-exploration amnesia (measured by dsh-router-standard).
Why is the transition band treated as a warning? The mixed band (0.2–0.5) is the training-distribution gap: measured scores are lower than either stable band. Entering it triggers the L1 hint; entering the react band escalates.
Is it safe to run?
The monitor is read-only with respect to the model: it consumes reasoning text
and sends intervention signals. All HTTP routes are loopback-only. Reasoning
text may be sensitive — the experiment log is local by default; rotate/disable
it in experiment_log.
Streaming waterfall rule — llm/stream is a stream-passthrough waterfall:
listeners earlier in the chain iterate the return value of the listeners after
them. Therefore:
llm/stream listener MUST be a plain function that
returns an async generator. Never declare it async — the generator gets
wrapped in a Promise and upstream for await consumers crash with
next(...) is not a function or its return value is not async iterable.for await (const chunk of await next()) —
await first, then iterate; safe regardless of what downstream returns.agent/pre-step / system-prompt/assemble are value-passing events;
async + await next() is correct there.A violation took down every model request with zero logged events — if all
sessions suddenly fail after a bundle reorder or a new plugin, audit the
llm/stream chain first.
Single-intervention-executor rule — L1/L2/L3 must have exactly one
executor. The Web plugin (host half) owns interventions and is the default;
the agent preset (preset/) ships with handleInterventions: false and only
pushes reasoning. Enabling both would double-register
agent/pre-step / system-prompt/assemble and fire L2 resets twice.
Built on the measured results of these community projects (shallow-cloned in
../references for audit):
we/let's/let me) and the E1/E1.5 hint-wording experimentsbandOf)MIT © 2026 Aik358
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: monitoring。