GLM-5.3-Flash-J-Space-Capability-Realization-Report
Tiger3807861189
GLM-5.3-Flash × J-Space capability realization — benchmark presentation of the J-Space Cognition Suite
DSH-PLUGIN STORE / LIVE CATALOG
聚合 GitHub 上的 DSH 插件,打造 DeepSeek Harness 生态的一站式目录。
25 个项目,匹配「benchmark」
Tiger3807861189
GLM-5.3-Flash × J-Space capability realization — benchmark presentation of the J-Space Cognition Suite
EverMind-AI
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.
morluto
Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
openguardrails
The vendor-neutral protocol for AI agent safety & security — and the neutral benchmark that ranks the vendors.
sunxin-ai
Design-fidelity QA for DeepSeek Harness: lend any text-only model an eye, then judge whether the implementation matches the mock. Ships the benchmark behind that judgement — four fixtures, 23 injected defects, and every raw model transcript. Retires itself when DeepSeek ships vision.
myc0576
Read-only trading journal and review harness: Jev typed judgments, agent integration, and a reproducible finance benchmark. No orders, no advice.
Ed-Marcavage
AI agents for pentesting, code audit, fuzzing, vulnerability discovery, and reverse engineering — harnesses, sandboxes, security MCP servers, benchmarks, and evals.
lizhiyao
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
ZK-Andy
Continual self-evolution plugin for DeepSeek Harness: versioned, auditable, rollback-safe harness state refined from session trajectories, with a benchmark-driven validation loop.
PerryLink
Cost governance for DeepSeek Harness: aggregated token/cost metering per model, session and day, budget caps with threshold alerts and over-limit policies, carbon footprint estimation, per-model latency benchmarks, a Settings budget tab, and the /budget command
hccccc01333
dsh-excel-chat — talk to Excel in DeepSeek Harness: create, edit, repair, and verify spreadsheets by conversation (cells, formulas, styles, filters, tables, charts); every edit is auto-validated.
xiaosu19
Codex, adaptive Codex PTC, and optional Codex Harness agent presets for DeepSeek Harness, with published benchmarks
morluto
Help your agents find the smoking gun they're looking for. Optimization evidence for agents: find complexity hotspots.
initial-d
DeepSeek Harness tools for reproducing the ml-quant-trading protocol v1 benchmark.
1052326311
Execution-time drift firewall for long-running DeepSeek Harness agents. Real-Harness tests: unsafe stale mutations 12/12 native -> 0/12; valid controls 7/7 both; post-SIGKILL unsafe continuation 2/2 -> 0/2.
Apageoflove
DeepSeek Harness 插件:同一个任务下对比多个模型配置,跑完给出 Pareto 排名和实验报告
Lhy723
Benchmark-driven self-evolution for DeepSeek Harness · 冻结基准上的 Agent Profile 自我进化:评测 → 候选 → 严格接受/回滚
dongsheng123132
Deterministic revision-pinned benchmarks and regression evidence for DeepSeek Harness
LeslieWylie
A reusable DeepSeek Harness bundle for evidence-driven memory, orchestration, benchmark operations, and plugin release workflows.
bpc-oss
该仓库暂未提供项目说明。
dsh-plugin-evaluation
Open evaluation datasets, test cases, and metrics for DSH plugins.
hccccc01333
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
hi-fangj
Model capability radar plugin for the DeepSeek Harness Web GUI
B1lli
Evidence-backed, type-aware quality scorecards for DeepSeek Harness plugins.