last30days-skill-cn
Jesseovo
last30days-cn 是一个 AI Agent 技能(Skill),能够自动搜索中国互联网 8 大主流平台最近 30 天的内容,综合分析后生成有据可查的研究报告。
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:8b-is/deepsiper-enthea
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
English | 中文
Deepsiper Enthea (deepsiper-enthea) is a sovereign, agent-driven LLM evaluation harness forked from deepseek-ai/deepseek-harness (dsh 0.1.0-rc.7). It provides end-to-end multi-model orchestration, self-hosted and sovereign backend integration, Cordis-powered extensible plugin pipelines, and JSON-RPC automation for modern LLM evaluation workflows.
tool-eval, eval-entheai, and custom benchmarking metrics.Everything in Deepsiper Enthea is a composable Cordis plugin.
┌─────────────────────────────────────┐
│ OpenCode / JSON-RPC / CLI / Web │
└──────────────────┬──────────────────┘
│
┌──────────────────▼──────────────────┐
│ Cordis Kernel (Context & DI) │
└────┬──────────────┬───────────────┬─┘
│ │ │
┌──────────────▼─────┐ ┌──────▼──────┐ ┌──────▼──────────────┐
│ Sovereign Backends│ │ Eval Plugins│ │ Sandboxed Tool Seams│
│ (EntheAI / Local) │ │ (tool-eval) │ │ (Landlock / Bash) │
└────────────────────┘ └─────────────┘ └─────────────────────┘
>=22.19.0 or >=24.0.0, TypeScript 6 (Strict ESM)tsdown / rolldown + tsc project references^22.19.0 || >=24.0.0pnpm >= 11.0.0# Clone repository
git clone https://github.com/8b-is/deepsiper-enthea.git
cd deepsiper-enthea
# Install dependencies and build harness
pnpm install
pnpm build
---
## Empirical AST & Logic Benchmark Suite
Deepsiper Enthea includes an automated multi-case AST evaluation harness (`examples/eval-entheai/driver/benchmark_sweep.py`) verifying mathematical precision, algorithm syntax, and zero reward-hacking across isolated execution environments:
| Benchmark Task | Category | Trials | Pass Rate | Mean Latency |
|---|---|---|---|---|
| **FizzBuzz Logic & 18 Edge Cases** | Logic Verification | 2 / 2 | **100.0%** | ~9.4 s |
| **$O(\log N)$ Matrix Power Fibonacci** | Mathematical Exponentiation | 2 / 2 | **100.0%** | ~14.8 s |
| **Kademlia 256-bit XOR Metric Distance** | Distributed DHT Routing | 2 / 2 | **100.0%** | ~10.5 s |
| **BitLinear $\{-1, 0, +1\}$ Quantizer** | Ternary Weight Mapping | 2 / 2 | **100.0%** | ~11.0 s |
| **AST Invariant & Pure Syntax Validator** | Syntax Tree Verification | 2 / 2 | **100.0%** | ~9.5 s |
### Run the Benchmark Sweep
```sh
# Execute the full multi-case benchmark suite
python3 examples/eval-entheai/driver/benchmark_sweep.py
https://axiomquant.orghttps://github.com/8b-is/classroom-sota-traininghttps://github.com/peterlodri-sec/etherhive · https://etherhive.vaked.devhttps://vaked.devhttps://peterl.dev@0xp3t3rl.bsky.socialMIT © 8b-is & DeepSeek AI contributors. Third-party dependency notices are listed in THIRD_PARTY_NOTICES.md.
Genesis Seal: 7c242080f5f821e5eaf563fe2208d60632c451687baf65f4fe8e4a0d226e3ecf · WE. {-1, 0, +1}. <3
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: evaluation-benchmark。