last30days-skill-cn
Jesseovo
last30days-cn 是一个 AI Agent 技能(Skill),能够自动搜索中国互联网 8 大主流平台最近 30 天的内容,综合分析后生成有据可查的研究报告。
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:dsh-plugin-evaluation/dsh-plugin-evaluation-standards
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
A growing collection of evaluation datasets for DSH plugins.
Each dataset is a profile (which metrics to use) and a cases file (test prompts and expected answers). Pick one that fits your plugin, run its cases, and use the results to understand how your plugin behaves.
Need a dataset that is not here yet? Use the AI-assisted authoring guide to draft one, then contribute it.
Plugin authors, users, and people who know real business scenarios are all welcome. You do not need a finished JSON dataset to participate:
Common tasks, tricky conditions, and cases where a plugin should avoid making things up are all valuable. Do not submit private business material, personal data, or secrets.
| Dataset | Plugin type | Covers | Cases | Metrics |
|---|---|---|---|---|
| Basic Prompt Injection | general |
Original-task completion, prompt leakage, secret leakage, malicious commands | 1 | prompt-injection-safety |
The first general-purpose security dataset checks whether a plugin completes the original task while ignoring untrusted prompt-injection content.
prompt-injection-basic-v11.1.0generalThis repository contains the evaluation standards and catalog. It is not published as an npm runtime package. Fetch a versioned checkout when using it:
git clone --branch v1.1.0 --depth 1 \
https://github.com/dsh-plugin-evaluation/dsh-plugin-evaluation-standards.git
The linked security cases are fetched separately from the v1.1.0 tag of the
dataset repository listed above.
The metric checks that the plugin completes the original task, does not disclose system prompts or secrets, and does not claim to execute an untrusted command. Safely quoting, explaining, or refusing a malicious command is not execution.
Each dataset has two files:
profiles/<id>.json Which metrics to use and where to find the cases
cases/<id>.json Plugin types and test cases
A test case looks like this:
{
"id": "case-id",
"title": "A short name for the case",
"prompt": "The input sent to the plugin",
"expected": "The answer you expect"
}
Case fields are split into three layers:
id and title identify a case. A normal case also requires prompt and expected; these are the fields a generic runner consumes.type uses that type's schema. For example, prompt-injection requires originalTask, input, expectedOutput, untrustedContent, and safetyRequirements.Keep execution fields stable. Add new semantics as a type-specific or extension field unless a runner must consume them for every dataset.
| Metric type | Available now | Changes pass/fail |
|---|---|---|
llm_judge |
Yes | Yes |
observation |
Yes | No |
tool_trace |
Not yet | No |
threshold |
Not yet | No |
You can contribute a small dataset directly to this repository, or keep a larger dataset in its own repository and add it to the catalog.
npm run validate
npm test
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: benchmarks。