DSH-PLUGIN STORE / LIVE CATALOG

DSH插件商店

聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。

已收录
4638
功能分类
11
更新时间
08/19 09:47

5 个项目,匹配「llm-eval」

开发工具 完整应用

ouroboros

Q00

Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

agent-os agentic-ai ai-agent ai-coding-agent
Agent 与会话 插件

Visual workflows and multi-model evaluation for DeepSeek Harness

agent-workflow deepseek deepseek-harness dsh-plugin
学习研究 插件

dsh-plugin-evaluation-standards

dsh-plugin-evaluation

Open evaluation datasets, test cases, and metrics for DSH plugins.

benchmarks deepseek-harness deepseek-harness-plugin dsh
文件与数据 插件

dsh-design-qa

sunxin-ai

Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.

benchmark deepseek-harness design-qa design-review
学习研究 插件

Sovereign, agent-driven LLM evaluation harness forked from deepseek-harness (dsh) — Cordis plugin architecture, EntheAI backends, and multi-model benchmarking

agent-harness ai-agents cordis deepseek