WeKnora
Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:ZJU-REAL/Polaris
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
Autonomous, end-to-end AI research: from literature to a reviewed paper.
Powered by a long-running agent core that plans, executes, and self-verifies its own work, turning every task into a resumable, auditable, human-gated run.
English · 简体中文
Polaris runs the entire research lifecycle as a single web application: literature survey, idea generation, idea review, experiment building on real GPU servers, LaTeX paper writing, and paper review. It is built for a research lab, with multi-user access, RBAC, and invite-code registration, and it treats every long task as a Voyage: a persisted, resumable, human-gated agent run that can span hours or days without losing state.
[!NOTE] Polaris is not a chatbot wrapper. The heavy lifting (crawling, parsing, deduplication, metric parsing, citation matching) is deterministic code. LLMs are reserved for the judgement calls: scoring, synthesis, drafting, and review. This split keeps runs cheap, reproducible, and auditable.
A 2-minute tour of the platform: the six-stage pipeline, the Voyage agent core, a real experiment run, and PolarisBuddy.
https://github.com/user-attachments/assets/388972c1-7ffa-45f2-94c4-07f388379ba2
A guest account on a running instance, for looking around: sign in at
http://101.37.174.109:8080 with the username guest and the password zjuguest123.
The account is for demonstration only: it is read-only and cannot call any model. It reaches every screen, the admin views included, but nothing it does changes state — creating, editing, deleting and uploading are all refused, and no LLM call will run, whether from chat, compilation or the assistant. Lab members' details and the registration codes are hidden from it as well. It is there to show what the platform looks like, not to do work on it.
Polaris models research as six stages. Each stage produces durable artifacts that the next stage consumes, and every hand-off can pause at a human approval gate.
flowchart LR
L["Literature<br/>Research Wiki"]
I["Idea<br/>Idea Forge"]
R["Idea Review<br/>Elo debate"]
X["Experiment<br/>GPU / SSH"]
W["Paper Writing<br/>LaTeX"]
V["Paper Review<br/>Citation check"]
S(["Submission"])
L --> I --> R
R -->|promotion gate| X
X --> W --> V
V -->|submission gate| S
classDef stage fill:#eaf1ff,stroke:#2f6bff,stroke-width:1px,color:#10233f;
classDef gate fill:#fff3e0,stroke:#f59e0b,stroke-width:1px,color:#5b3b00;
class L,I,R,X,W,V stage;
class S gate;
| Stage | What Polaris actually does |
|---|---|
| Literature | The Research Wiki ingests papers from OpenAlex, Semantic Scholar, and arXiv. Cold start snowballs citations from anchor papers and scores relevance against the direction library's inclusion config — statement, goals, scope and exclusions, written through a structured AI interview — then extracts full text (PyMuPDF) and compiles a cross-linked wiki page (TL;DR, method, reusable ideas, concept backlinks). There is one wiki per paper, shared platform-wide: the compile prompt carries no library statement or rubric, so the same paper never reads differently depending on where you opened it, and a concept is promoted only once two papers cite it. New arXiv work arrives through the daily feed, the single entry point libraries sync from; incremental sync with watermark resume, pgvector semantic search, research digests, and Obsidian vault sync. |
| Idea | Idea Forge runs multi-signal gap analysis over the knowledge base (concept co-occurrence holes, extracted paper limitations, trend velocity, survey gaps) to drive retrieval-planned idea generation. Ideas are scored on four axes (novelty, feasibility, operability, impact), deduplicated semantically, and funneled to a candidate pool. A deep Research Proposal builder then hardens the winner with a plan-execute-verify loop. |
| Idea Review | Configurable-persona reviewer agents debate pairwise; a judge produces an Elo tournament ranking. Lab members join the discussion live over WebSocket, and their comments enter the agent context as first-class input. |
| Experiment | The Experiment Lab uses per-user, Fernet-encrypted SSH credentials to reach the lab's GPU servers. An experiment Voyage asks intake questions first, plans the study, passes a compute-budget check, writes code, runs a smoke test, launches runs with streamed logs and live metric curves, then auto-iterates: parse metrics, reflect, then improve, debug, or stop — repairing failures under a time budget rather than a fixed retry count. It keeps a file-based memory it reads and writes across steps, and when it is genuinely stuck it asks the user instead of failing. A console gives each run a task map and a terminal you can talk to mid-stream. Figures are generated and VLM-checked. |
| Paper Writing | The Paper Writer opens a multi-file LaTeX project (NeurIPS, ICLR, ACL templates) with a CodeMirror 6 editor, real-time collaborative editing (CRDT), and server-side tectonic compilation to a live PDF preview. An agent drafts section by section, but experiment numbers may only come from real ExperimentRun metrics and citations must map to real knowledge-base entries. One click refreshes the references and wires the bibliography into the main TeX file. |
| Paper Review | Line-by-line citation verification (existence: exact, minor, or fabricated; support: supported, partial, or unsupported) plus deterministic fact-checking of every number against the experiment record, then multi-perspective top-venue reviewer agents and a meta-review. A fabricated citation forces a non-pass. |
Research tasks are long-running by nature: a cold-start literature backfill takes hours, an experiment runs for days. Polaris's central abstraction is that every complex task is a Voyage: a resumable, auditable run driven by a persisted three-part loop.
| Component | Role |
|---|---|
| Navigator | Planning. Decomposes a goal into a step plan with sub-goals, dependencies, and budget. In loop mode it edits the plan incrementally as evidence arrives, rather than replanning from scratch. |
| Helm | Execution. Runs a single step (LLM calls, tool calls, SSH remote ops, literature-API queries) and returns an observation. |
| Sextant | Self-verification. Checks each step against structured acceptance criteria (exit code, artifact exists, schema valid, metric threshold, count, LLM rubric). Deterministic checks run first; failures feed diagnostics back to Navigator, and repeated failure escalates to a human gate. |
[!IMPORTANT] A Voyage is backed by a persistent state machine (
planning -> executing -> verifying -> ...). If a worker crashes mid-run, the Voyage resumes from its last checkpoint after a health check. Budgets are attached to the run and auto-pause it when exceeded; every plan, action, and verdict is retained and replayable in the UI.
Not every task needs the full cognitive loop. A shared Runtime shell (state machine, checkpointing, gates, budget, cancellation, event streaming) serves all task kinds, while the Brain (the full plan-execute-verify loop) activates only for open-ended kinds such as experiments. Predictable pipelines (wiki compile, idea review, paper drafting) run on fixed templates instead of being over-orchestrated.
[[wikilinks]] and frontmatter.chat, plan (research-only, propose before acting), and goal (loop toward
an objective) modes. It runs the same Navigator / Helm / Sextant split as a Voyage, so its steps are
verified rather than merely produced, and it can hand work to subagents. Its greeting is stitched from
real SQL counts rather than the model, it searches every library you can see, and it carries page
context and persistent per-user memory. It is off unless the account may call a model.guidance, rubric, persona, and workflow packs injected at named points
into agent prompts, enabled globally, with a publish-approve-install-rate marketplace; each Voyage
snapshots the skills it used for reproducibility. Agent skills take the SKILL.md shape with
three-level progressive disclosure — a one-line description in the catalog, the body fetched by
skill_load as a tool result, attachments read on demand — so the model decides what to load and the
prompt prefix stays cacheable.| Layer | Technology |
|---|---|
| Frontend | React 18 + TypeScript 5 + Vite 5, TanStack Query for all server state, CodeMirror 6, Yjs (CRDT), react-pdf, KaTeX |
| Desktop | Electron shell (macOS / Windows / Linux) that reuses the web bundle over an app:// protocol; all heavy state stays on the remote server |
| Backend | FastAPI (fully async) + SQLAlchemy 2 + Alembic + fastapi-users (JWT) |
| Task queue | ARQ (Redis broker); every long task runs off the request thread |
| Data | PostgreSQL 16 with pgvector (embedding spaces isolated per model, so vectors never mix) + Redis 7 |
| Remote execution | asyncssh to GPU servers; SSH keys encrypted at rest with Fernet |
| LaTeX | tectonic, server-side, with a cached macro volume |
| LLM | Multi-provider abstraction (OpenAI-compatible and Anthropic) with a DB model-routing table |
| Deployment | Docker Compose (postgres, redis, api, worker, frontend) |
Polaris ships as a desktop app for macOS, Windows, and Linux. Download the installer from
Releases — .dmg / .zip (macOS, universal),
.exe / portable .zip (Windows), .AppImage / .deb (Linux), built by CI on every v* tag. The
app checks for updates and applies them without a restart where it can.
The builds are neither signed nor notarized, so each platform needs to be told once that the app is
safe to run: on macOS xattr -dr com.apple.quarantine /Applications/Polaris.app (or right-click →
Open); on Windows choose More info → Run anyway past SmartScreen; on Linux the AppImage needs
libnss3 libgtk-3-0 libasound2, and --no-sandbox under Ubuntu 24.04+ AppArmor.
Since v0.4.0 the app is an offline, single-machine build: it ships its own Python backend and bootstraps it on first launch (SQLite, no Docker, no login, no server address). The first launch downloads a Python toolchain and dependencies, which can take a few minutes; later launches start immediately. If the local engine cannot start, the app falls back to asking for a Polaris server address, which is also how you connect to a shared multi-user server.
[!NOTE] v0.3.x and earlier are remote-only shells: they always ask for a server address on first run. They cannot update themselves into the local-engine build — download a v0.4.0+ installer from Releases and install it over the old one.
To build it yourself:
make desktop-deps # install the shell's dependencies (once)
make desktop-dev # build the frontend and run the shell (app:// protocol)
make desktop-dist # package an unsigned installer for the current platform
See docs/desktop.md for the process model, the IPC contract, and packaging notes.
[!TIP] Docker Compose is the recommended way to run Polaris, in development and in production. It needs only Docker and Docker Compose installed, with no local Python, Node, or database. See docs/deployment.md for production deployment.
cp .env.example .env # set provider keys and secrets
make dev # full stack via docker compose, hot reload
Local development without Docker (falls back to SQLite):
make backend-dev # venv + uvicorn on :8000
make frontend-dev # npm install + vite dev on :5173
Common tasks:
make migrate # alembic upgrade head
make test # backend pytest + frontend build
make lint # ruff check + tsc --noEmit
Deploy from pre-built images on Docker Hub (tricktreat/polaris-{api,worker,frontend}, published by
CI on every v* tag) — no local build needed:
cp .env.example .env # set POLARIS_ENV=prod, POLARIS_IMAGE_TAG, secrets, and an LLM key
docker compose --env-file .env -f docker/docker-compose.yml pull
docker compose --env-file .env -f docker/docker-compose.yml up -d
docker compose -f docker/docker-compose.yml exec api alembic upgrade head # required on first run
The frontend is served at http://<host>:8080. The worker container is required (it runs all long
tasks), and the first-run migration is mandatory (Postgres tables are not auto-created). Pass
--env-file .env so Compose reads POLARIS_IMAGE_TAG (default latest) / POLARIS_IMAGE_PREFIX
(default tricktreat) from the repo-root .env.
For building locally instead, bind mounts, backups, and restricted networks, see docs/deployment.md.
Full documentation lives in docs/:
src/
backend/ FastAPI app (package: app) and ARQ worker (package: worker)
app/
api/ thin routers
services/ business logic (ingest, wiki, ideas, review, experiments, manuscripts, ...)
models/ SQLAlchemy models
agents/voyage/ the Voyage engine (navigator, helm, sextant, tool loop, per-domain actions)
core/ config, db, queue (ARQ), events (SSE), llm/ abstraction
tools/, mcp/ read-only tool registry and the external MCP server
frontend/ React + Vite (src/features/ has one folder per product area)
desktop/ Electron shell that wraps the web bundle (macOS / Windows / Linux)
docker/ Dockerfiles and compose (base, dev override, prod overlay)
docs/ English project documentation
See docs/architecture.md for the full design.
See CONTRIBUTING.md. In short: one feature is one branch is one pull request, branched
from the latest origin/main, with English conventional-commit messages, and main stays a read-only
fast-forward mirror of origin/main.
Licensed under the Apache License, Version 2.0. See LICENSE for the full text.
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: auto-research。