DSH-PLUGIN STORE / LIVE CATALOG

DSH插件商店

聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。

已收录
6553
功能分类
11
更新时间
10/07 02:01

23 个项目,匹配「evaluation」

文件与数据 渠道适配

WeKnora

Tencent

Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.

agent agentic ai chatbot
开发工具 技能

ouroboros

Q00

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

agent-os agentic-ai ai-agent ai-coding-agent
Agent 与会话 技能

SkillCorpus

EverMind-AI

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

agent-memory agent-skills ai-agents benchmark
开发工具 技能

caliper

edonadei

Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.

ai-agents claude-code cli codex
安全与治理 基础设施

axern

cofy-x

Secure, reproducible sandboxes for AI agent evaluation, training, and data synthesis.

agent-sandbox agentic-infrastructure ai-agents cloud-native
开发工具 技能

oh-my-knowledge

lizhiyao

OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.

agent-evaluation ai benchmark bootstrap-ci
开发工具 插件

dsh-eval-harness

BiBoyang

DSH 插件评测工具:YAML 用例驱动真实 agent 回归评测 + baseline 对比 PASS/WARN/FAIL 门禁|Regression eval harness for DeepSeek Harness plugins

deepseek-harness dsh dsh-plugin evaluation
文件与数据 插件

deepseekeyes

dttxorg

Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness — strict evidence, health-checked failover, original pixels, and Token accounting.

ai-agent ai-evaluation auditable-ai browser-automation
开发工具 插件

dsh-self-evolving

timwhitez

Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.

agent-evaluation ai-agents cordis deepseek
开发工具 技能

engineer-software

KirschBluteX

Evidence-driven engineering workflows for Codex and DeepSeek Harness, backed by deterministic routing and behavior evaluations.

agent-skills agent-workflow ai-coding-agent codex-plugin
开发工具 插件

DSH-arena

Apageoflove

DeepSeek Harness 插件:同一个任务下对比多个模型配置,跑完给出 Pareto 排名和实验报告

ai-agent arena benchmarking cordis
开发工具 技能

dr-agent-skills

Leeaoyin

Structured, reusable skill modules for AI coding agents — covering engineering workflows, reliability evaluation, and production readiness.

coding dsh dsh-plugin harness-engineering
开发工具 完整应用

DeepSeek Harness plugin that validates and ranks 3/5 independent coding-agent patches before approval-gated apply.

agent-evaluation ai-agents coding-agent deepseek-harness
学习研究 插件

dsh-plugin-evaluation-standards

dsh-plugin-evaluation

Open evaluation datasets, test cases, and metrics for DSH plugins.

benchmarks deepseek-harness deepseek-harness-plugin dsh
学习研究 插件

dsh-eval

hccccc01333

Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.

agent-evaluation benchmark deepseek-harness dsh
其他 待识别

DeepSeek Harness plugin and Harbor template for reproducible Agent evaluation, self-evolution, and controlled promotion.

dsh-plugin
其他 插件

dsh-profile-lab

young-tim

Reproducible DSH profile and patch experiment matrices with reports and policy gates

deepseek deepseek-harness dsh-plugin evaluation
其他 插件

dsh-review-mode

hqa-shu

Work in progress: an independent conversation-review plugin for DeepSeek Harness and local Codex sessions. Evidence-aware feedback, goal-drift analysis, and a dedicated review panel. 正在开发中。

agent-evaluation ai-agents codex deepseek-harness
部署运维 插件

Controlled request-surface replay and regression workbench for DeepSeek Harness

agent-observability deepseek deepseek-harness dsh
其他 插件

dsh-eval

CatheadOwl

Agent eval framework over dsh headless runs: case runner, session-trace assertions, and a scripted mock-LLM layer for plugin intent tests.

agent-evaluation deepseek-harness dsh-plugin
文件与数据 插件

DSH-native multi-runtime baseline, ablation, and reproducible evaluation control plane

ablation-study agent-evaluation deepseek-harness dsh
其他 待识别

Durable, bounded lifecycle supervisor with scheduled evaluation for live DeepSeek Harness sessions (community plugin)

deepseek-harness dsh-plugin
Agent 与会话 插件

Visual workflows and multi-model evaluation for DeepSeek Harness

agent-workflow deepseek deepseek-harness dsh-plugin