返回目录
学习研究 插件

cancer-meta-pipeline

ReGMeIoN/cancer-meta-pipeline

AI-assisted, reproducible pipeline that turns a paper list into publication-ready extraction tables for cancer systematic reviews

Stars
1
Forks
0
Issues
0
更新
5 天前

PROJECT TOPICS

项目标签

INSTALL REFERENCE

安装参考

未验证
dsh plugin --profile web add github:ReGMeIoN/cancer-meta-pipeline

该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。

PROJECT README

README

cancer-meta-pipeline

A reproducible, AI-assisted pipeline that turns a list of papers into publication-ready extraction tables for cancer systematic reviews / meta-analyses.

Give it three things — a paper list, a screening protocol written in plain language, and an abstrackr project id — and it runs the whole literature workflow: normalise → deduplicate → compile a machine-readable rubric → title/abstract screening in batches → write decisions back to the platform → full-text chase → data extraction → Table 1 / Table 2.

It ships as a DeepSeek Harness Skill: a SKILL.md plus plain Python scripts. No framework, no database, no service to run.


What it does (8 stages, each with a human checkpoint)

# Stage Output
1 Ingest — CSV / Excel / RIS / BibTeX / EndNote / bare PMID list work/records_local.csv + import report
2 Pull platform — full export with citation_id, status, tags work/records_all.csv, status snapshot
3 Compile rubric — plain-language protocol → E-code table + inclusion clauses work/_compiled/rubric.json, PROMPT_batch.md, PROMPT_fulltext.md
4 Batch screening — LLM subagents judge title/abstract in 250-record batches work/batches/ → work/decisions/
5 Validate & merge — one central validator, never self-written checks work/validation_report.md, screening_decisions.csv
6 Submit — dry-run first, audit tags, rollback path platform labels + logs/submit_audit.log
7 Full text — three-channel open-access chase (Europe PMC / Unpaywall / PMC) article/oa/*.txt, needs-a-human list
8 Extract — Table 1 (characteristics + reported effects), Table 2 (effect sizes) out/table1_*, out/table2_*, .docx, coverage_audit.md

Every stage ends with a checkpoint: the skill stops and asks the reviewer before moving on.

Design rules baked in

  • Batches of 200–300 records (small batches multiply subagent startup cost).
  • Subagents may not write their own scripts or validators — one central validator.
  • Prompts live in files, so a hundred batches share one identical wording.
  • Everything is traceable: every decision carries a reason + a verbatim quote; every platform write leaves an audit tag (ft:corrected-from-*) and a documented rollback.
  • E-code ⇒ excluded is enforced in code: a study that reports only a composite outcome (sensitivity-only) or is a Mendelian randomisation study (triangulation stream) can never enter the primary synthesis, even if a subagent labels it "Include".

Requirements

  • Python 3.10+ (standard library only; pypdf for PDF→text, python-docx for the Word table)
  • Optional: R + metafor for the pooling step (see reference/HANDOFF_to_analysis.md)
  • Optional: abstrackr-api-toolkit for the platform stage, and an ABSTRACKR_HOME directory holding your credentials

Quick start

projects/<name>/
├── config.json      # platform project id + batch sizes
├── protocol.md      # your PECO + inclusion clauses + E-code table (see templates/)
└── input/           # the paper list, in any supported format
$skill = '<this repo>'
$py    = 'python'
$env:PYTHONIOENCODING = 'utf-8'
$proj  = '<absolute path to projects/<name>>'

& $py "$skill\scripts\ingest.py"             --project $proj
& $py "$skill\scripts\pull_platform.py"      --project $proj
& $py "$skill\scripts\compile_rubric.py"     --project $proj      # ⛳ confirm the E-code table
& $py "$skill\scripts\prep_batches.py"       --project $proj --size 250
#   ... one subagent per batch, using work/_compiled/PROMPT_batch.md ...
& $py "$skill\scripts\validate_decisions.py" --project $proj
& $py "$skill\scripts\merge_decisions.py"    --project $proj
& $py "$skill\scripts\submit_decisions.py"   --project $proj      # dry-run, then --execute
& $py "$skill\scripts\verify_write.py"       --project $proj
& $py "$skill\scripts\build_tables.py"       --project $proj --audit

Offline smoke test (fictional data, no network, no platform):

& $py "$skill\scripts\ingest.py"             --project "$skill\examples\smoke"
& $py "$skill\scripts\compile_rubric.py"     --project "$skill\examples\smoke"
& $py "$skill\scripts\prep_batches.py"       --project "$skill\examples\smoke" --size 2
& $py "$skill\scripts\build_tables.py"       --project "$skill\examples\smoke" --audit

Layout

SKILL.md                 entry point: stages, hard rules, checkpoints
stages/DETAILS.md        step-by-step operating notes
scripts/                 13 scripts, one per stage (+ common.py)
templates/               protocol / rubric / prompt / config / handover templates
reference/PITFALLS.md    26 real pitfalls (read this before running anything)
reference/checklist.md   per-stage deliverables and the questions to ask
examples/                fictional smoke projects (CSV / RIS / BibTeX)

Read this first

reference/PITFALLS.md is the most valuable file here — every entry is a real failure from a real review (platform throttling, silently truncated snapshots, BOM-broken JSON, a case-sensitive regex that quietly disabled a central rule, metafor's I² already being a percentage, and more).

License

Not yet chosen — add one before relying on this in production.

CLASSIFICATION EVIDENCE

分类依据

项目类型插件
功能分类学习研究
规则置信度高

系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: academic、data-extraction、literature-screening、research-automation、research-software、research-tools、scientific-workflow、systematic-review。