oh-my-knowledge
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
DSH-PLUGIN STORE / LIVE CATALOG
聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。
7 个项目,匹配「agent-evaluation」
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
timwhitez
Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.
Web0926
DeepSeek Harness plugin that validates and ranks 3/5 independent coding-agent patches before approval-gated apply.
hccccc01333
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
hqa-shu
Work in progress: an independent conversation-review plugin for DeepSeek Harness and local Codex sessions. Evidence-aware feedback, goal-drift analysis, and a dedicated review panel. 正在开发中。
CatheadOwl
Agent eval framework over dsh headless runs: case runner, session-trace assertions, and a scripted mock-LLM layer for plugin intent tests.
Harzva
DSH-native multi-runtime baseline, ablation, and reproducible evaluation control plane