返回目录
Agent 与会话 插件

dsh-agent-memory

Culeot/dsh-agent-memory

Cross-session long-term memory plugin for DeepSeek Harness (DSH)

Stars
8
Forks
0
Issues
0
更新
12 天前

PROJECT TOPICS

项目标签

INSTALL REFERENCE

安装参考

未验证
dsh plugin --profile web add github:Culeot/dsh-agent-memory

该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。

PROJECT README

README

dsh-agent-memory

Cross-session long-term memory plugin for DeepSeek Harness (DSH).

Memory panel

中文文档

What it does

Remembers facts, preferences, decisions, and lessons across sessions, so a new session starts with what earlier sessions already learned. Storage lives in $DSH_HOME/storages/memory.json (plain JSON — inspectable, git-friendly). No MCP servers, no vector database, no embedding API, no extra runtime dependencies.

Features

  • Cross-session persistence — write in session A, recall in session B; new sessions pick up where old ones left off.
  • Relevance-threshold injection — before each model step, the plugin recalls memories matching the user's current message, filters by minimum relevance score (injectMinScore), and injects up to injectCount that pass the threshold. Low-relevance memories are discarded — you only pay for what actually matters.
  • Multi-dimensional relevance — scoring combines semantic similarity, task relevance, pattern matching, causal linkage, and recency (not just keyword overlap). A memory is used when ≥2 dimensions match, reducing noise from coincidental keyword hits.
  • Anti-decay rules — the protocol section includes self-reinforcing rules: task-start recall re-reads core rules, recurring mistakes auto-solidify as lessons, and language drift triggers self-correction.
  • Self-correction loop — when the same error (same code/message) fires repeatedly (lessonizeAfter, default 2), the plugin nudges the agent to write it as an importance-3 lesson; the lesson then auto-injects on related topics, preventing recurrence. User corrections are covered by the memory protocol (write the lesson right away).
  • Built-in hygiene — capacity cap with lowest-value eviction first; near-duplicate entries merge instead of piling up (Jaccard ≥ 0.7); optional TTL expiry; deleting importance-3 records requires an explicit confirm.
  • Chinese-friendly search — Chinese text is indexed by bigrams plus single-char fallback plus a BM25 term-frequency signal, English by words, plus exact substring matching. Works without a tokenizer or any ML dependency.
  • No unrelated association — weak matches (pure single-char coincidence, filler-only queries like "ok了吗") score zero and are never injected; only substring/tag/bigram-strength matches surface. Saves tokens on short chatter.
  • Optional semantic channel — a second, independent scorer: a locally-trained sentence encoder running through ONNX. Dormant until you point semanticModelDir at a model you trained yourself (no model ships with the plugin). Once present it fuses with lexical scoring. See Semantic retrieval.
  • Explainable recall — every recall hit carries reasons (substring/tag/bigram/BM25/importance/recency/access signal breakdown), so both you and the agent can audit why a memory surfaced.
  • Memory panel (Web UI) — a top-level "记忆" entry in Settings with stats, search, kind filters, and direct create / edit / delete of memories; changes apply immediately, dark mode included.
  • memory_sediment — batch-persist facts/decisions/lessons at session wind-down (≤3 entries per call, cooldown-guarded); the agent summarizes what it already has in context, so there is zero extra model cost.
  • Native integration — uses DSH's own storage domain (ctx.storageDomain), tool registry, and agent lifecycle hooks, so it stays compatible with official releases.

Tools

Tool Purpose
memory_remember Store a durable memory (content, kind, tags, scope, importance, optional TTL).
memory_recall Search memory by keywords; ranks by relevance, importance, recency, past usage.
memory_index Browse the inventory with kind/tag/scope filters, paginated, title-level.
memory_forget Delete by id or tags; importance-3 records need confirm: true.
memory_import Import memories from a JSONL/JSON file through the write chain (not by editing the store file) — safe bulk loading.
memory_reload Reopen the store from disk after an external edit; merges external changes without a restart.
memory_sediment Batch-persist several memories at once (session wind-down), with an entry cap and cooldown guard.

Kinds: fact | preference | decision | lesson | todo | note. Scopes: user (applies everywhere) or project (this project only).

Semantic retrieval (optional) — you train the model yourself

No model file ships with this plugin. Lexical scoring on its own is complete and needs nothing else; the semantic channel is a second opinion you can add later.

It is worth adding, but not for the reason you'd expect. On a 382-query held-out set built from a real memory bank:

Channel MRR Recall@1
lexical only 0.7248 63.4%
semantic only 0.7121 61.3%
both, fused 0.8432 76.4%

Read the middle row again: semantic alone loses to lexical. No single model you could download makes retrieval better by itself. The gain comes from fusing two channels that fail in different places — lexical owns exact terms, paths and command names; semantic owns paraphrase.

Why no model is bundled

  • Size — 393 MB base model + 391 MB trained + 486 MB quantized ONNX. That is not something to put in an npm package.
  • It would be the wrong model anyway — a retriever learns what your queries and your memories look like. Ours are Chinese technical notes queried as keyword stacks (语音插件 语音播报 speak tts). Yours almost certainly differ, and a mismatched retriever is a silent downgrade.
  • Optional means optional — a missing @huggingface/transformers, a missing model directory, or a failed vector computation all degrade quietly to lexical scoring and record why. A memory plugin that refuses to start because an optional accelerator is misconfigured would be worse than having no accelerator at all.

Full walkthrough

See training/README.md — environment setup, downloading the base model, held-out splitting, two-track training-pair generation, contrastive training, ONNX export with a correctness check, and wiring the result into the plugin.

The short version:

pip install torch transformers onnxruntime numpy huggingface_hub

python training/prep.py                              # split the memory bank, hold out 20%
python training/gen_data.py --kw-per 4 --llm-per 6   # build query/memory pairs
python training/train.py --epochs 30 --batch 32      # contrastive training
python training/export_onnx.py                       # → work/model-onnx/

Then, in your profile: pnpm add @huggingface/transformers, point semanticModelDir at work/model-onnx, restart DSH.

All six scripts in training/ take their paths from environment variables (MEM_TRAIN_WORK, MEM_TRAIN_MEMORY, MEM_TRAIN_BASE, MEM_TRAIN_LLM), so they run from any directory. An NVIDIA GPU helps (RTX 4080 Laptop 12 GB: ~2.4 min for 306 memories × 30 epochs) but CPU works — just slower. Budget ~2 GB of disk.

External-modification protection

The store file is loaded once at startup and written as a whole by the running process (single-writer model). To prevent an in-process write from silently wiping external edits:

  • every write checks the file fingerprint first — a mismatch (another process or a script edited it) refuses the write with a clear error instead of overwriting;
  • memory_import is the supported way to bulk-load data (goes through the write chain, file and memory stay in sync);
  • if you still edit memory.json by hand or copy it in, call memory_reload to merge it back (or restart).

Troubleshooting: panel reports HTTP 405

A transport failure for /dsh-memory-read/…: HTTP 405 in the memory panel means the RPC channel was never mounted — the request fell through to the SPA static fallback, which only serves GET/HEAD. Root cause is in DSH 0.1.5-rc.2 itself: connection.rpc.handle() resolves webServer from the calling context to register its route, and throws cannot get property "webServer" without inject when that lookup fails, silently dropping the channel (still unfixed in 0.1.6-alpha.2). This plugin therefore tries the official handle() first and, on failure, registers its own prefix routes through ctx.inject(['connection', 'webServer']), reusing connection.requestRejection() for the Host/Origin fence and browser authentication.

A DSH restart is required — routes are mounted at startup. Verify with:

curl -i -X POST http://127.0.0.1:3080/dsh-memory-read/stats   # 401 = route mounted (no cookie), 405 = stale process

Repeat suppression

Per-step injection skips when the recalled set is unchanged from the previous step, so consecutive messages about the same topic don't re-inject the same block — the「相关记忆」notice appears on topic change, not on every message.

Install & enable

# 1. Add the dependency to your profile
cd ~/.dsh/profiles/<name>
npm install dsh-agent-memory
# 1. add the dependency to your profile
cd ~/.dsh/profiles/<name>
pnpm add dsh-agent-memory@file:/path/to/dsh-agent-memory
# 2. add one row to your agent preset (~/.dsh/.agent-presets/<preset>/agent.cordis.yml)
- id: memory
  name: 'dsh-agent-memory'
# 3. restart DSH — the four tools appear in new sessions

No preset? Mount it on the host plane instead, in ~/.dsh/profiles/<name>/cordis.patch.yml:

- insert:
    - id: memory
      name: 'dsh-agent-memory'

Requires the storage trio already present in the profile (dsh-storage, dsh-storage-json, dsh-storage-domain — the web profile ships with it).

Configuration

All options are optional.

Store

Option Default Meaning
maxRecords 400 Capacity cap; lowest-value records evicted first.
maxContentChars 2000 Max content length per record.
mergeSimilarity 0.7 Near-duplicate merge threshold.
recencyHalfLifeDays 90 Freshness half-life, in days.
protocolSection true Inject the memory protocol prompt section.
recallContentMax 400 Max content chars returned per recall hit.

Injection

Option Default Meaning
injectEnabled true Per-step injection of relevant memories (via agent/pre-step).
injectCount 3 Max memories injected per step (0 disables).
injectMinScore 1.0 Minimum relevance score to inject (0 = no threshold, just rank).
injectMaxChars 120 Max chars per injected memory summary.
injectBudgetRatio 0.5 Inject only candidates scoring within this fraction of the top hit — a query with one obvious answer injects one memory, not three (0 = off, historical floor-and-cap behaviour).
noMatchThreshold 0.18 Below this top score, recall returns noMatch and injects nothing. Offering a weak memory is not neutral — it spends attention and misleads.

Ranking

Option Default Meaning
boostSlope 0.1 Importance boost slope: 1 + (importance − 1) × slope.

Semantic channel (needs a model you trained — see above)

Option Default Meaning
semanticEnabled true Master switch; a no-op while semanticModelDir is empty.
semanticModelDir '' Directory holding config.json, tokenizer.json, onnx/*.onnx. Empty = channel off.
semanticCacheDir '' Vector-cache directory. Empty = $DSH_HOME/storages/memory-semantic.
semanticWeight 0.3 Semantic weight in the fusion: relevance = w × semantic + (1 − w) × lexical.
semanticMin 0.5 Cosine floor below which a semantic score doesn't count as a match.

Lessons & sediment

Option Default Meaning
lessonizeEnabled true Auto-nudge to solidify repeated errors as lessons.
lessonizeAfter 2 Same error fingerprint occurrences before nudging.
sedimentMaxEntries 3 Max memories per memory_sediment call.
sedimentCooldownMs 300000 Cooldown between sediment calls (5 min).

Uninstall & troubleshooting

  • Uninstall: pnpm remove dsh-agent-memory in the profile, delete the preset/patch row. Data stays in memory.json and is restored on reinstall.
  • Tools missing: check the row exists, the dependency is installed, and DSH was restarted.
  • Storage errors: the memory plugin needs the storage trio; add it to the patch if your profile lacks it.
  • Corrupted memory.json: it is plain JSON — fix it by hand or delete it (deleting resets memory).

Development

npm install && npm run build && npm test   # build + 31 unit tests
npm run smoke                              # real-machine headless round-trip check

License

MIT

CLASSIFICATION EVIDENCE

分类依据

项目类型插件
功能分类Agent 与会话
规则置信度高

系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: memory。