OpenViking
volcengine
Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:stuarthu/dsh-crew
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
Run work in DeepSeek Harness (dsh) as a small crew of role agents.
Your own dsh session becomes the product manager (PM). The PM is the only one who talks to you. It writes down what "done" means, asks you to confirm it, then starts an architect to design the work, engineers to write the code, and reviewers to judge both. The roles never talk to each other — they share work through files on disk, and the PM passes messages.
Version 0.7.0. PM, researcher, architect, engineer, QA, code reviewer, security reviewer, doc reviewer — plus a language and stack you approve before any work starts, a QA suite that stays on disk and runs again, a written change request for every scope or contract change, pushing with your permission, CI watching, and picking a job up after a crash.
dsh keeps model-facing tools in an agent preset, not in your profile. dsh-crew follows that, and splits itself in two:
| Piece | Where | Why there |
|---|---|---|
| PM rules | host plane (your profile) | They need no tools, so they work in every session, on any preset |
| Role tools | the crew agent preset |
A role's allow/deny list is checked against the preset when a child starts, so the names must be defined in the same place |
The first dsh start after installing writes the preset into
$DSH_HOME/.agent-presets/crew. Start a session on the crew preset to get
the roles. In a session on another preset the PM still behaves like a PM,
notices it has no role tools, and offers to either move to the crew preset or do
the job itself.
The crew preset is dsh's own standard preset with one change: subagent,
subagent_fork, workflow, ralph and the product subagents are gone, and the
crew roles are there instead. So inside it, a crew role is the only way to
start an agent.
dsh has three hard rules about agents, and the design follows them:
| dsh rule | What it means here |
|---|---|
| A message can only go to a direct child | Every role is a direct child of the PM, so the PM can reach all of them |
A child answers only its direct parent (report) |
Every answer comes back to the PM |
| Two children cannot talk to each other at all | Roles share work through files, not chat |
If the architect started the engineers, the PM could not reach the engineers at
all. So only the PM starts agents. Two independent guards enforce that: every
role is denied the delegation tools, and each role tool has maxDepth: 1, so a
crew child cannot start another crew child.
A role is not a prompt the PM pastes in. It is a real delegation tool built from
@deepseek-ai/dsh-tool-subagent:
| Role | Tool | Persona | Tools |
|---|---|---|---|
| Researcher | crew_researcher |
roles/researcher.md |
only read, glob, grep, write, web_search — no shell |
| Architect | crew_architect |
roles/architect.md |
everything except the crew tools |
| Engineer | crew_engineer |
roles/engineer.md |
everything except the crew tools |
| QA | crew_qa |
roles/qa.md |
everything except the crew tools — it must run the software |
| Code reviewer | crew_code_reviewer |
roles/code-reviewer.md |
only read, glob, grep |
| Security reviewer | crew_security_reviewer |
roles/security-reviewer.md |
only read, glob, grep |
| Doc reviewer | crew_doc_reviewer |
roles/doc-reviewer.md |
only read, glob, grep |
So a code reviewer cannot change a file, even if it decides it wants to. The persona is locked in as that child's own system prompt.
The reviewer uses an allow list, and two live tests are the reason:
write and edit denied, it created a file anyway with
echo hello > file. A shell is a file-writing tool.workflow,
ralph and a set of desktop-control MCP tools — every one of them a way out.A deny list cannot name what a deployment has not installed yet. An allow list does not have to. The PM pastes the diff into the review task and runs any command the reviewer asks for.
Two more guards sit under the filters:
maxDepth: 1 on every role tool — only the root PM can start a role, and
it names no tool at all, so no preset change can weaken it.workflow, ralph or a bare
subagent.Role personas are plain markdown in roles/. Copy one into
~/.dsh/crew/roles/ under the same name and it replaces the shipped version.
One limit: prompt text may not contain {{ — dsh would read it as a variable,
and the plugin fails at startup naming the file.
Role tool filters and per-role models are configured where the roles live: the
dsh-crew-roles row in ~/.dsh/.agent-presets/crew/agent.cordis.yml.
That file sits inside the installed preset, and a new version of dsh-crew
replaces the whole folder. Your edited files are kept beside the new ones as
agent.cordis.yml.bak, and the boot log names them — but the settings do not
come back by themselves. After an upgrade, copy your changes into the new file.
The PM sorts your ask into a lane: ask (answer only), quick (do it), or
team (the full flow). If the size is unclear, it asks you.
It asks which language to use. It never guesses.
It grills you — one question per turn, each with a recommended answer. It
waits for your answer before it asks the next one, and never sends you a list
of questions. It looks up every fact it can in the repository first. For anything bigger than a
quick look it starts a crew_researcher, which writes findings with a source
for every answer, so you are only asked what the files cannot answer.
It settles the language and stack, and you approve it. If the repository already has one, that is the stack. The PM reads the manifest, the lock file, the test folder and the CI workflow, says what it found, and you confirm it in one line.
When the choice is real — an empty repository, a new service — the PM starts a
crew_researcher first. The researcher reports what this kind of project is
normally built with today, with a source for every claim, what the machine
already has, and what each option costs. It lists the options and is not allowed
to recommend one.
Then the PM recommends one, names the runner-up and why it lost, and writes a Language and stack section into the document: language and version, package manager, framework, database, and the test framework with the exact test command. That last one matters most. Engineers write their tests with it and QA writes its cases with it, so it has to be one choice, not five.
You confirm the stack together with the document. After that it moves only through a CRD — a change request document, explained further down. An engineer may pick freely among the libraries the project already has; adding a brand-new dependency goes back to the PM.
It picks the document and writes it: a DoD (definition of done,
~/.dsh/crew/jobs/<job-slug>/dod.md) for small work, a PRD (product
requirements document, docs/design/prd.md) for a real product. It says
which it picked, and one word switches it. You confirm before any work
starts.
A PRD is cut into milestones — three to six stops, each one something you
can look at and judge, written in your words rather than in code words. M1
is the proof of concept: the thinnest real path through the riskiest part,
running for real. You confirm the milestone list on its own, because it decides
when you get a say.
It turns your job name into a short slug — lowercase letters, digits and -,
nothing else — tells you the slug it picked, and creates a crew/<job-slug>
branch with it. Then, for PRD work, it starts crew_architect, which writes
the high level design, the decision records and the task breakdown. It also
splits the system into
modules — reusing what the repository already has before it invents anything
new — and when two or more modules talk to each other it writes one contract
file per boundary in docs/design/api/: how the two sides talk (in-process
call, HTTP, gRPC, events, and so on), the data format, every call with its
inputs, output and errors, and how the shape stops a caller getting it wrong.
It picks the style, not the library. Those contracts are what let two
engineers build the two sides at the same time, because crew roles cannot talk
to each other, so each contract also names one test per side — the callee
proves it answers what the file says, the caller tests against a stub built
from the file. And when there is a boundary, the first task is a walking
skeleton: one engineer builds the thinnest real path across the riskiest
boundary, alone, before anything runs in parallel. A contract that does not fit
is cheapest to fix there. It also puts every task under one of your
milestones — it cannot add, rename or reorder them. Then crew_doc_reviewer
must pass all of it before a single line of code is written.
It runs one crew_engineer per task, one milestone at a time. Two engineers
run together only when their file lists do not overlap, and never across a
milestone line. Every
engineer works test first: it writes one unit test, runs it, checks that it
fails for the right reason, then writes the smallest code that makes it pass.
Its report has to show you the failing run and then the passing run. Every one
of those tests is a real file in your project's test suite, named in the task
row and committed with the code — never a command someone ran once in a shell.
An engineer that believes a test cannot come first must ask the PM before it
writes any code.
Each finished task is checked in order: code review (correctness, then the
tests that drove the change, then reuse, simpler code, readability and this
repository's own style — the reviewer may hold up a task on those, but only if
it shows the exact replacement it wants; otherwise the finding is optional) → security review, only when the change touches
the network, login, secrets, files outside the project, the shell, user input,
customer data or a new dependency → QA, which writes its test plan from
the document before reading the code, then turns every case into a real test
file under docs/qa/<task-id>/, in your project's own test framework,
with a run.sh beside it. The plan itself is single-use and stays out of your
repository — once the cases exist they say the same thing in a form that runs,
so the plan is dropped with the job. The one part of it that must not be lost
is "what I could not test here, and why": QA writes that into
docs/qa/gaps.md, a standing list about your product's testability that later
jobs shorten. bash docs/qa/run-all.sh runs every task's cases,
and QA runs it on every task it checks — so a case written in an earlier task
guards the new one. An old case that starts failing is a blocking regression,
and nobody is allowed to edit it green. Your QA suite grows with the project
and outlives the job — and the PM wires that command into your project's own
default test command, so it runs without anyone remembering it. If your test
runner cannot see the folder, the PM adds the one line that makes it visible;
"the cases are not runnable" is a problem the PM brings to you, not a place to
stop. Round two of any review
only re-checks the blocking findings; after the round limit the PM brings the
disagreement to you.
The PM commits — engineers never touch git. It stages only the files that task
owns, never git add -A.
Milestone review — the PM stops and asks you. When every task in the milestone has passed those checks and is committed, the PM reports what works now, the exact commands to try it yourself, what is deliberately not there yet, the test results, and where shipping stands. Then you say: ship this milestone, go on without shipping, change something, or stop — one question, four answers. A change that touches the PRD sends the plan back through the architect and the doc reviewer before code starts again. No milestone begins until you have answered the one before it. Small DoD work has no milestones — it is one piece of work with one report at the end.
A milestone you ship gets two plans, and their shape is looked up, not guessed. These plans are not alike. An npm package cannot un-publish a version. A mobile app waits for a store review. A web service rolls back by redeploying. A database schema needs a migration that is safe to run twice.
So the PM starts a crew_researcher and asks what those two plans hold for
your project type, with a source and a date for every claim. It reads what
your repository already does first: the workflows, the changelog, the tags, any
release script. Then it writes two files:
docs/release/<milestone>-release.md — the version and the rule behind
it, the release notes, the exact steps and who approves each one, what must be
true before you start, how you check afterwards that it worked, and how to
undo it. If it cannot be undone, the plan says so in those words.docs/release/<milestone>-upgrade.md — who is upgrading and from which
versions, every breaking change and what the user must do about it, the
migration steps and whether they are safe to run twice, what happens to
someone who skips a version, how to go back and what data that loses, how long
it takes, and what goes offline while it runs.A milestone you are not shipping gets no plan. It gets a gap list instead: one honest paragraph naming what is still missing before it could ship. That list gets shorter as milestones pass. And approving a plan is not approving a push — every push and publish still needs its own yes, every time.
The PM updates the repository README to match what was built. README.md is
always English. If you chose another language for the job, it keeps a second
file beside it — README-zh.md, README-ja.md — saying the same thing. If
nothing a reader would notice changed, it leaves the README alone and tells
you so.
A last crew_doc_reviewer pass over every document the job produced, the
README included. It checks that the documents can be worked from, that they
stay consistent (one name per idea, one shape, and the language files saying
the same thing), and that they read easily for someone about 14 years old
whose first language is not English — by counting things like sentence
length, idioms and unexplained terms, not by taste. It may hold up the job on
wording, but only if it writes the replacement sentence itself.
Push and CI, if you allow it. The PM checks there is a remote, a
workflow and a working gh first. Then it asks you — before every push,
including a re-push after a fix. It pushes only what you said yes to — a
crew/* branch, main, or a release tag — watches the run, and sends a red
CI's real error text back to the engineer that owns those files.
Merge and clean up, only when you ask for it. The PM merges the
crew/<job-slug> branch into main itself. It asks you three separate
times — once for the merge, once for pushing main, once for deleting the
branch — and one yes never covers the next thing. The merge is never
squashed, so your one commit per task and its test-first proof stay readable
in the history. Before it pushes main it tells you whether that push would
start a workflow that publishes, naming the file it read — and if you still
say yes, it pushes. It offers to delete the branch only after it has proved
the work is merged and really on the remote, including that the remote
branch holds nothing main does not: git push origin --delete has no
protection of its own. With trustRootAgent: false that remote delete is
refused on purpose; the PM then hands you the command to run yourself
instead of retrying. A work branch that simply stays is a normal ending too.
Fixing a bug can come back to you as a question, and every choice gets written down. This can happen at any time inside step 8. An engineer fixing a bug — a defect QA found, a blocking review finding, or one it hit itself — first finds at least two ways that would really work. If the ways only differ in wording, it picks one and says in its report which ones it compared. If the difference would stay in the code, it stops. The difference stays in the code when any one of these six is different between the ways:
docs/design/api/;When it stops, it hands the PM the cause of the bug and every way it found, each with the files it would change, what it costs and where it would hurt later, plus the one it recommends. Then the PM decides by the same line a CRD uses: a difference you can see, it asks you about right away; a difference that stays inside the code, it decides itself and names at the next milestone review. New features and refactors do not go this way.
The decision is written down before any code is built, and it holds
every option. It goes in one place, whatever the size of the job: an
ADR — a decision record in docs/decisions/adr/. On PRD work an
architect may write it; small work has no architect, so the PM writes it
itself. The ADR holds the cause of the bug, every option with what it costs,
where it would hurt later and why it lost, which one was taken, who
decided, and the reason. The options section quotes the engineer's own
question file word for word — the PM adds only the decision and the
reason, so it cannot quietly reshape the options into a case for the choice
it already made, and a pointer like "options: see Q-03" is not allowed
because that file is dropped with the job. Every
ADR is written for you: a reader who has never seen the code must be able
to tell the options apart, and the recommended one is marked. The design
does not stop and wait for you to pick — the architect keeps going on its
own recommendation, and at the milestone review the PM puts the options
from every ADR of that milestone in front of you. You may overturn any of
them; that is a CRD, and the tasks already built the old way are done
again.
The crew is flat: the PM talks to each role, and two roles can never talk to each other. So a message reaches one role and dies there. That is why the crew talks through documents — a role's report points at the file it wrote, and the PM's answer points at the document it changed and that document's new version. Two engineers building two sides of the same boundary read the same file, and a role started tomorrow reads what a role started an hour ago read.
On top of that, every change request gets its own file. If anyone — you, a
role, or the PM itself — asks for something that changes what you get (the scope,
an acceptance check, the milestone list) or how two modules talk (a boundary
contract), the PM writes docs/decisions/crd/NNNN-<short-name>.md first: who asked,
what they want, why, which documents and tasks it touches, what it costs, and the
decision with its reason. Nothing is built from an undecided one, and a rejected
one is kept, so you can see later what was asked for and refused.
Who decides which:
Small questions do not become CRDs — a role's question that the files can answer is just a note in the job folder, and a review finding about code is a review finding. Only scope and contracts, the two things that cost real work to redo.
A decision about how gets an ADR instead, whatever the size of the job. One
question tells the two apart: did someone ask for this? If someone did — you,
QA, a review — it is a change request and gets a CRD. If nobody did, and the crew
ran into a choice while doing the work, it is an ADR in docs/decisions/adr/.
Nothing else decides where it lands: not how big the job was, and not whether it
had an architect. Small work has no architect, so the PM writes the ADR itself.
Where a document lives depends on how long it lives. What outlives the job is
in your repository, under docs/, and each folder name says what it holds: the
PRD and the design in docs/design/ (with one contract file per module boundary
in docs/design/api/), the decision records and the change requests in
docs/decisions/ (adr/ and crd/), QA's runnable cases and its standing list
of what no test can check in docs/qa/, the release and upgrade plans for each milestone you ship in
docs/release/, and the researcher's answers in docs/research/.
What belongs to this one job is outside the repository, in
~/.dsh/crew/jobs/<job-slug>/, so your git status stays clean: the job state
(state.json), the DoD (dod.md), QA's test plans, and the Q- question files a
role leaves for the PM. That whole folder is dropped when the job ends, and a test
run's output was never a file at all.
Before anything is dropped, the durable half moves out. This is a real step at
the end of a job, and it happens after the PM's closing summary to you — not the
moment the acceptance checks turn green, because the thinking usually carries on
past that point. A rule the crew must keep next time goes into principles.md, a
decision about how into an ADR, a decision about what or a contract into a CRD,
this change's reasons and its real test numbers into the commit message, and QA's
"what I could not test here, and why" into docs/qa/gaps.md. "Not needed any
more" has to be earned. It is the same reason an ADR copies the engineer's options
in instead of pointing at the file they came from.
The state file alone is not enough — the next session has to know about it. So dsh-crew reads the job folder on every turn and, when something is unfinished, puts a short note in front of the PM:
Unfinished crew work: 1 job left in /home/you/.dsh/crew/jobs.
- "add-sso-login" in /home/you/project (branch crew/add-sso-login):
5 of 9 tasks done, 2 blocked. Last change 2026-08-18 09:12.
The PM must tell you before it does anything else, and ask one question: carry on
or start clean. It never does either without your answer. A job belonging to a
different folder is ignored, and a state file it cannot read is reported rather
than counted as finished. Turn the whole thing off with resumeNotice: false.
host/git-guard.js inspects every shell command. Your own session is the root
agent, and it is trusted: its git and publishing commands pass straight through.
Every crew role is a child agent, and the guard refuses, from a child:
git push of main, master, trunk, develop, HEAD, or with no branch
named;--mirror, --all, or force push;npm/pnpm/yarn/bun publish, npm dist-tag, gh release create;crew/push-ok-flow branch, a push-okay.md file and a push-ok.bak backup
are not mistaken for the approval file.A child's other branch push needs a one-shot approval that you create:
mkdir -p ~/.dsh/crew && touch ~/.dsh/crew/push-ok
The guard deletes that file as soon as one push uses it. One approval, one push.
A refused child is told to ask you, and is not shown those two commands. Only your own session sees them, so the agent that was just refused does not also get the recipe.
Set trustRootAgent: false to guard your own session exactly like every child.
approvalFile must name a file, never a folder. A value like ~/.dsh/crew/
would leave crew as the protected name, so the guard refuses to load and the
message tells you to name the file itself.
Three honest limits, and each one is a real hole, not a caveat. It reads
command text, so it is a strong seat belt, not a locked door: a command
hidden inside a script file could still slip past, and so does a file name the
shell builds from pieces. It cuts the other way too — a command that only
mentions the approval file by name is refused as well, your own session
included, so a commit message with push-ok in it will not run. And it reads
bash and pwsh only. A role that can write files — the engineer can —
could create the approval file as a plain file write, and the guard never sees
that call. Nothing here stops that; dsh's own approval prompt for writing a
file is what stops it. Your dsh approval prompts stay the real gate.
And the publishing scan is GitHub-only. The guard reads .github/workflows
and understands GitHub's on: push: shape. GitLab, CircleCI, Jenkins and Azure
Pipelines are outside it, on purpose: stretching GitHub's trigger rules
half-way onto another CI system would produce false alarms, and a false alarm
is worse than no alarm, because it teaches you to say yes without reading. So
the guard is a GitHub-only backstop for child agents. The wider check is the
PM's own judgement: in step 17 it also reads .gitlab-ci.yml,
.circleci/config.yml, Jenkinsfile and azure-pipelines.yml when they
exist, and tells you what it found before it pushes main.
dsh plugin --profile tui add dsh-crew # or --profile web
Then restart dsh. Starting it writes the crew preset into
$DSH_HOME/.agent-presets/crew (a crew folder somebody else wrote is left
alone). Pick the Crew preset for a session to get the roles.
To check the plugin without dsh:
npm test # replays the guard rules and the mount, then runs every QA case; no dsh needed
This repository's own CI runs npm test on every push, and publishes only when a
v* tag is pushed. One gap is worth knowing about: tools/verify-mount.mjs
skips its role-tool half on any machine without @deepseek-ai/dsh-tool-subagent
installed, and CI is such a machine. It says out loud which half it skipped, so a
green run means "everything a public runner can check", not "everything".
Everything is optional. Settings live in the plane they belong to.
PM and guard — the dsh-crew-core and dsh-crew-git-guard rows in your
profile's cordis.patch.yml:
| Setting | Default | What it does |
|---|---|---|
rolesDir |
~/.dsh/crew/roles |
Your own role markdown files replace the shipped ones, by file name |
limits.liveAgents |
20 |
Crew agents awake at the same time |
limits.reviewRounds |
3 |
Review rounds before the PM asks you to decide |
installPreset |
true |
Write the crew preset into $DSH_HOME/.agent-presets |
jobsDir |
~/.dsh/crew/jobs |
Where job state lives, and what the crash notice reads |
resumeNotice |
true |
Put unfinished jobs in front of the PM at session start |
enabled (guard) |
true |
Turn the git guard off — not recommended |
trustRootAgent (guard) |
true |
Trust your own session (the PM) with any git or publish command |
approvalFile |
~/.dsh/crew/push-ok |
One-shot push approval file. A file path, never a folder — a trailing slash fails at startup |
Roles — the dsh-crew-roles row in
~/.dsh/.agent-presets/crew/agent.cordis.yml:
| Setting | Default | What it does |
|---|---|---|
rolesDir |
~/.dsh/crew/roles |
Same override folder, for the role personas |
roleAllow |
reviewers: read, glob, grep; researcher: read, glob, grep, write, web_search |
Only these tools for that role; everything else is closed |
roleDeny |
architect, engineer, QA: the crew tools | Everything except these for that role |
roleModels |
session model | Per-role provider and model |
MIT
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: multi-agent。