开发工具
完整应用
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.
agent-os
agentic-ai
ai-agent
ai-coding-agent
开发工具
技能
Know if your agent skill actually works. A lightweight evaluation harness that tracks a success rate across Claude Code, Codex, Pi, and Hermes.
ai-agents
claude-code
cli
codex
开发工具
技能
Evidence-driven engineering workflows for Codex and DeepSeek Harness, backed by deterministic routing and behavior evaluations.
agent-skills
agent-workflow
ai-coding-agent
codex-plugin
开发工具
插件
DSH 插件评测工具:YAML 用例驱动真实 agent 回归评测 + baseline 对比 PASS/WARN/FAIL 门禁|Regression eval harness for DeepSeek Harness plugins
deepseek-harness
dsh
dsh-plugin
evaluation
开发工具
插件
Local-first experiment and evaluation workbench plugin for DeepSeek Harness (DSH).
ai-agent
arena
benchmarking
cordis
开发工具
技能
Structured, reusable skill modules for AI coding agents — covering engineering workflows, reliability evaluation, and production readiness.
coding
dsh
dsh-plugin
harness-engineering
开发工具
插件
该仓库暂未提供项目说明。
deepseek-harness
dsh-plugin
evaluation
regression-testing
开发工具
插件
Evidence-first, crash-resumable self-evolution engine for DeepSeek Harness and Harbor.
agent-evaluation
ai-agents
cordis
deepseek