WeKnora
Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
DSH-PLUGIN STORE / LIVE CATALOG
聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。
23 个项目,匹配「evaluation」
Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Q00
Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
EverMind-AI
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
edonadei
Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.
cofy-x
Secure, reproducible sandboxes for AI agent evaluation, training, and data synthesis.
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
BiBoyang
DSH 插件评测工具:YAML 用例驱动真实 agent 回归评测 + baseline 对比 PASS/WARN/FAIL 门禁|Regression eval harness for DeepSeek Harness plugins
dttxorg
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness — strict evidence, health-checked failover, original pixels, and Token accounting.
timwhitez
Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.
KirschBluteX
Evidence-driven engineering workflows for Codex and DeepSeek Harness, backed by deterministic routing and behavior evaluations.
Apageoflove
DeepSeek Harness 插件:同一个任务下对比多个模型配置,跑完给出 Pareto 排名和实验报告
Leeaoyin
Structured, reusable skill modules for AI coding agents — covering engineering workflows, reliability evaluation, and production readiness.
Web0926
DeepSeek Harness plugin that validates and ranks 3/5 independent coding-agent patches before approval-gated apply.
dsh-plugin-evaluation
Open evaluation datasets, test cases, and metrics for DSH plugins.
hccccc01333
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
istarwyh
DeepSeek Harness plugin and Harbor template for reproducible Agent evaluation, self-evolution, and controlled promotion.
young-tim
Reproducible DSH profile and patch experiment matrices with reports and policy gates
hqa-shu
Work in progress: an independent conversation-review plugin for DeepSeek Harness and local Codex sessions. Evidence-aware feedback, goal-drift analysis, and a dedicated review panel. 正在开发中。
tbxy09
Controlled request-surface replay and regression workbench for DeepSeek Harness
CatheadOwl
Agent eval framework over dsh headless runs: case runner, session-trace assertions, and a scripted mock-LLM layer for plugin intent tests.
Harzva
DSH-native multi-runtime baseline, ablation, and reproducible evaluation control plane
acosmi
Durable, bounded lifecycle supervisor with scheduled evaluation for live DeepSeek Harness sessions (community plugin)
alison-xx
Visual workflows and multi-model evaluation for DeepSeek Harness