ouroboros
Q00
Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
DSH-PLUGIN STORE / LIVE CATALOG
聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。
9 个项目,匹配「evaluation」
Q00
Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
edonadei
Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
BiBoyang
DSH 插件评测工具:YAML 用例驱动真实 agent 回归评测 + baseline 对比 PASS/WARN/FAIL 门禁|Regression eval harness for DeepSeek Harness plugins
timwhitez
Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.
KirschBluteX
Evidence-driven engineering workflows for Codex and DeepSeek Harness, backed by deterministic routing and behavior evaluations.
Apageoflove
DeepSeek Harness 插件:同一个任务下对比多个模型配置,跑完给出 Pareto 排名和实验报告
Leeaoyin
Structured, reusable skill modules for AI coding agents — covering engineering workflows, reliability evaluation, and production readiness.
Web0926
DeepSeek Harness plugin that validates and ranks 3/5 independent coding-agent patches before approval-gated apply.