WeKnora
Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
PROJECT TOPICS
INSTALL REFERENCE
dsh plugin --profile web add github:A3Boy/dsh-web-tools
该命令指向仓库当前默认分支;尚无绑定当前 commit 的完整验证结果。
PROJECT README
Empower DeepSeek Harness with unified search and deep content extraction across the open web and social platforms.
Native-Capability Adaptation Across 8 Web Providers · SearchHints Semantic Compilation · Multi-Source Resilience · Xiaohongshu & Twitter / X Retrieval
English | 简体中文
When web access depends on a single provider, exhausted quota, rate limits, or timeouts can interrupt retrieval. Wrapping multiple APIs with a naive proxy often flattens them to a lowest common denominator, failing to leverage each provider's specialized search modes, categories, freshness, domain rules, and extraction capabilities.
dsh-web-tools connects Exa, Tavily, Firecrawl, Parallel, Brave, You.com, Jina, SearXNG, Xiaohongshu, and Twitter / X to DSH’s standard web_search / web_fetch interface.
While keeping the tool interface unified, dsh-web-tools normalizes search intents through SearchHints and compiles queries into native provider-specific parameters. This leverages each engine's native categories, freshness filters, domain policies, regional targeting, and content extraction, while maximizing uptime through multi-key allocation, automated failover, and dedicated browser profiles.
Key Highlights:
web_search / web_fetch contracts while compiling unified search intent into provider-specific native parameters instead of reducing every backend to the same lowest-common-denominator feature set.web_search / web_fetch without requiring new tool declarations, and includes a session-level "Search Mode" toggle. DSH Agent
│
web_search / web_fetch
│
▼
SearchHints
topic / freshness / domains
locale / cleanQuery ...
│
Provider-specific
compilation
│
┌────────┬─────────┬───────────┬──────────┐
│ Exa │ Tavily │ Firecrawl │ Parallel │ ...
│ │ │ │ │
│category│ topic │ github │objective │
│ date │ chunks │ research │ policy │
│domains │ time │ tbs │ queries │
└────────┴─────────┴───────────┴──────────┘
The tool interface and search semantics are unified for the DSH agent, but underlying provider capabilities are not.
dsh-web-tools expresses query intent through SearchHints and crafts dedicated requests for each provider, rather than compressing all backends into a single set of lowest-common-denominator parameters.
For example, when searching for "AI coding references from the past week":
github category, research queries to research, and applies tbs freshness filters;objective, clean search_queries, and source_policy;boost_domains for soft domain preference.All adaptations run through deterministic code without invoking an extra LLM call.
publication / news / financial report), ISO-8601 date ranges, and domain constraints.github category, academic queries to research, and supports tbs, domain constraints, and clean Markdown extraction.objective soft-steering + clean search_queries) and source_policy domain/freshness filters.basic / advanced / fast / ultra-fast search depths; basic, advanced, and fast support chunks_per_source, with native news / finance topics, date ranges, country, and domain constraints.pd/pw/pm/py freshness filters, country, and search language.boost_domains soft-weighting, freshness presets, and geo/language targeting.categories (it/science/news) and time_range.web_fetch): Automatically routes to provider-native scraping backends (Exa /contents, Tavily /extract, Firecrawl /scrape, Parallel /v1/extract, You.com /v1/contents, Jina Reader) when available; seamlessly falls back to the built-in generic HTTP fetcher (powered by local Defuddle Markdown parsing with SSRF/DNS protection) for SearXNG-only / Brave-only setups or when native extractors fail.Unlike general search engines, the plugin connects directly to native social platform sessions:
Native Isolated Browser Architecture:
Xiaohongshu:
web_fetch): Uses the dedicated browser profile to extract structured __INITIAL_STATE__ data with DOM fallback while preserving signed xsec_token URLs. When the page loads comment data, top-level comments and returned nested replies are extracted individually.web_search): Uses the signed-in browser and the real search controls on /explore, entering only the cleaned topic query rather than platform names or site: operators. It distinguishes login walls, security verification, and a genuinely signed-out session. Operators can temporarily disable this path with XHS_NATIVE_SEARCH=0 while diagnosing browser issues.Twitter / X:
from:, since:, and until: operators.Each detail fetch includes at most 30 comments or replies, with up to 800 characters per entry. Additional pages or nested replies are explicitly marked as truncated to keep model context bounded.
Capability Boundary: For Xiaohongshu notes where the primary content resides in images, the fetcher returns the title, text description, engagement metrics, image count, and loaded comments; it does not perform optical character recognition (OCR) on image text or automatically paginate through all comments.
Agents use 小红书: or X: as a platform-routing prefix. The prefix selects the platform and is removed before the native search runs; for example, 小红书: DeepSeek Harness enters only DeepSeek Harness into Xiaohongshu's search box.
a1 and web_session, then performs a stabilized live /explore check in the interactive browser. A visible login wall invalidates the old session and restores the sign-in action; a wall appearing only after search submission is reported directly as search-restricted rather than being hidden behind indexed web results.
Retry-After headers, skipping rate-limited providers immediately.web_search supports Ordered, Round-Robin, and Random starting-provider selection. web_fetch always follows the deterministic fetch-capable chain.web_search or web_fetch call before an answer. A failed call still counts as an attempt, and the agent is instructed to disclose what could not be verified.HTTP_PROXY, HTTPS_PROXY, and NO_PROXY, with automatic loopback bypass.# Install
dsh plugin --profile web add github:A3Boy/dsh-web-tools
# Update
dsh plugin --profile web update dsh-web-tools
# Remove
dsh plugin --profile web remove dsh-web-tools
Restart dsh web and navigate to Settings → Web Search.
| Provider | Search | Fetch / Extract | Key Integrations & Adaptations | Quota Inspection |
|---|---|---|---|---|
| Exa | Yes | Yes, /contents |
Semantic retrieval (auto / fast / deep), category mappings, query-aware highlights, ISO-8601 date ranges |
Dashboard only |
| Tavily | Yes | Yes, /extract |
Search depths (basic / advanced / fast / ultra-fast); the first three support chunking, plus news / finance, date, country, and domain parameters |
Official API |
| Firecrawl | Yes | Yes, /scrape |
Technical queries map to github, academic queries to research; supports tbs, domain constraints, and clean Markdown extraction |
Official API |
| Parallel | Yes | Yes, /v1/extract |
Agent dual-layer semantic search (advanced / basic / turbo), objective soft-steering, source_policy |
Dashboard only |
| Brave Search | Yes | — | Preferred LLM Context pre-extraction with pd/pw/pm/py freshness, country/lang filters, auto fallback to Classic Search |
Response headers |
| You.com | Yes | Yes, /v1/contents |
Snippet highlights, native boost_domains soft-weighting, freshness and country/lang filters, markdown endpoint |
Official API |
| Jina | Yes | Yes, Reader | Search query noise filtering, ReaderLM-v2 high-precision markdown, token budget and truncation control | Best-effort |
| SearXNG | Yes | — | Open-source self-hosted metasearch with categories and supported time_range values; the adapter requires no API key |
Instance-defined |
New installations default to Exa; existing installations keep their saved provider configuration.
| Scenario | Recommended Provider | Notes |
|---|---|---|
| Social Discovery & Content Detail | Xiaohongshu / Twitter / X | Both platforms provide signed-in native search, detail retrieval, and returned comments or replies |
| Semantic Search / Technical Documentation | Exa | Semantic modes plus category, date, domain, and highlight parameters |
| Pre-Extracted Search Context | Brave Search | Prefers LLM Context with Classic Search fallback |
| Configurable Search Depth and Extraction | Tavily / Parallel | Provider-native depth modes and content extraction endpoints |
| Content-to-Markdown Extraction | Firecrawl / Jina | Main-content filtering, scraping, and Reader conversion |
| Freshness, Region, and Domain Preference | You.com | Freshness, locale, and boost_domains parameters |
| Self-Hosted Metasearch | SearXNG | Uses your SearXNG endpoint; the adapter requires no API key |
pnpm install # Install dependencies
pnpm test # Run test suite
pnpm run typecheck # Type checking
pnpm run build # Build bundle into lib/
When using Docker or self-hosted SearXNG, if you encounter no usable provider or HTTP 403 Forbidden errors, check the following:
settings.yml and add - json under search.formats:search:
formats:
- html
- json # Required for API queries
Restart your SearXNG container (docker restart searxng) to apply changes.
Base URL field (e.g. http://127.0.0.1:8080) and click "Test Search" to verify connectivity.If cache issues occur after upgrading via local path or symlinks, reinstall from the profile directory:
cd ~/.dsh/profiles/web && pnpm install
Exa and Parallel balances are checked in their provider dashboards; Brave quota information comes from actual search response headers. Quota display is informational and does not affect routing or fallback.
Xiaohongshu and Twitter / X each use a dedicated local browser profile. The plugin does not export raw cookies to configuration, logs, or third-party relays; the browser sends them through normal authenticated requests to the corresponding platform domains.
MIT © A3Boy
CLASSIFICATION EVIDENCE
系统优先读取 GitHub Topics,再与站内分类词典和词根规则比对。当前命中: web-search。