DSH-PLUGIN STORE / LIVE CATALOG

DSH插件商店

聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。

已收录
6553
功能分类
11
更新时间
10/07 02:01

25 个项目,匹配「benchmark」

学习研究 技能

GLM-5.3-Flash × J-Space capability realization — benchmark presentation of the J-Space Cognition Suite

agent-skills ai-agent benchmark deepseek
Agent 与会话 技能

SkillCorpus

EverMind-AI

Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.

agent-memory agent-skills ai-agents benchmark
模型与 MCP 插件

flameox

morluto

Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.

benchmarking coding-agents cordis debugging
安全与治理 插件

openguardrails

openguardrails

The vendor-neutral protocol for AI agent safety & security — and the neutral benchmark that ranks the vendors.

agents ai-safety ai-security dsh-plugin
文件与数据 插件

dsh-design-qa

sunxin-ai

Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.

benchmark deepseek-harness design-qa design-review
开发工具 插件

SmartMoney-Cub

myc0576

Read-only trading journal and review harness: Jev typed judgments, agent integration, and a reproducible finance benchmark. No orders, no advice.

ai-agent ai-agents backtesting decision-logging
安全与治理 索引目录

AI agents for pentesting, code audit, fuzzing, vulnerability discovery, and reverse engineering — harnesses, sandboxes, security MCP servers, benchmarks, and evals.

agentic-ai ai-agents ai-hacking ai-pentesting
开发工具 技能

oh-my-knowledge

lizhiyao

OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.

agent-evaluation ai benchmark bootstrap-ci
其他 插件

Continual self-evolution plugin for DeepSeek Harness: versioned, auditable, rollback-safe harness state refined from session trajectories, with a benchmark-driven validation loop.

ai-agent deepseek-harness deepseek-harness-plugin dsh
部署运维 基础设施

dsh-budget

PerryLink

Cost governance for DeepSeek Harness: aggregated token/cost metering per model, session and day, budget caps with threshold alerts and over-limit policies, carbon footprint estimation, per-model latency benchmarks, a Settings budget tab, and the /budget command

ai-agent ai-agents budget carbon-footprint
学习研究 插件

dsh-excel-chat

hccccc01333

dsh-excel-chat — talk to Excel in DeepSeek Harness: create, edit, repair, and verify spreadsheets by conversation (cells, formulas, styles, filters, tables, charts); every edit is auto-validated.

agent benchmark deepseek-harness dsh-plugin
开发工具 完整应用

dsh-codex-mode

xiaosu19

Codex, adaptive Codex PTC, and optional Codex Harness agent presets for DeepSeek Harness, with published benchmarks

agent-preset codex codex-harness coding-agent
开发工具 插件

smokinggun

morluto

Help your agents find the smoking gun they're looking for. Optimization evidence for agents: find complexity hotspots.

ai-agents benchmarks code-optimization code-quality
学习研究 插件

DeepSeek Harness tools for reproducing the ml-quant-trading protocol v1 benchmark.

benchmark deepseek-harness dsh-plugin mlquant
Agent 与会话 插件

dsh-plan-lattice

1052326311

Execution-time drift firewall for long-running DeepSeek Harness agents. Real-Harness tests: unsafe stale mutations 12/12 native -> 0/12; valid controls 7/7 both; post-SIGKILL unsafe continuation 2/2 -> 0/2.

agent-harness agent-orchestration agent-planning agent-safety
开发工具 插件

DSH-arena

Apageoflove

DeepSeek Harness 插件:同一个任务下对比多个模型配置,跑完给出 Pareto 排名和实验报告

ai-agent arena benchmarking cordis
其他 待识别

Benchmark-driven self-evolution for DeepSeek Harness · 冻结基准上的 Agent Profile 自我进化:评测 → 候选 → 严格接受/回滚

dsh dsh-plugin
开发工具 插件

dsh-benchmark

dongsheng123132

Deterministic revision-pinned benchmarks and regression evidence for DeepSeek Harness

ai-agent benchmark deepseek-harness dsh
Agent 与会话 插件

dsh-ops-kit

LeslieWylie

A reusable DeepSeek Harness bundle for evidence-driven memory, orchestration, benchmark operations, and plugin release workflows.

agent-tools deepseek-harness dsh dsh-plugin
学习研究 插件

该仓库暂未提供项目说明。

agent benchmark deepseek-harness dsh
学习研究 插件

dsh-plugin-evaluation-standards

dsh-plugin-evaluation

Open evaluation datasets, test cases, and metrics for DSH plugins.

benchmarks deepseek-harness deepseek-harness-plugin dsh
学习研究 插件

dsh-eval

hccccc01333

Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.

agent-evaluation benchmark deepseek-harness dsh
学习研究 插件

dsh-models-radar

hi-fangj

Model capability radar plugin for the DeepSeek Harness Web GUI

cordis deepseek-harness dsh dsh-plugin
学习研究 插件

Evidence-backed, type-aware quality scorecards for DeepSeek Harness plugins.

benchmark deepseek-harness dsh-plugin plugin-quality