Documentation compiler for coding agents

Turn any docs site into context your agent can use.

DocHarvest captures public documentation, removes page noise, and writes a local corpus with source URLs, search, snapshots, and exports.

docharvest capture https://docs.openalgo.in/v/v2.0 --rag --pdf

One command. Local files. No account required.

Read the source

Live capture specimen

The output starts as a source-backed index.

251 rows
FILTER:

14 of 14 rows · 4 captures · 23 pages

  • 10book.md:10 3. [Providers](pages/docs/latest/providers.md) pi.dev
  • 20book.md:20 Guides and references for configuring and extending Pi. pi.dev
  • 28book.md:28 Search documentation pi.dev
  • 72book.md:72 Authenticate with `/login` for subscription providers, or set an API key such as `ANTHROPIC_API_KEY` before starting pi. pi.dev
  • 82book.md:82 * [Providers](latest/providers) - subscription and API-key setup for built-in providers. pi.dev
  • 100book.md:100 * [Custom models](latest/models) - add model entries for supported provider APIs. pi.dev
  • 101book.md:101 * [Custom providers](latest/custom-provider) - implement custom APIs and OAuth flows. pi.dev
  • 116book.md:116 * [Environment variables](latest/environment-variables) - Pi process configuration and session metadata available to bash tools. pi.dev
  • 141book.md:141 Guides and references for configuring and extending Pi. pi.dev
  • 149book.md:149 Search documentation pi.dev
  • 184book.md:184 Configure via `package.json`: pi.dev
  • 195book.md:195 Change `name`, `configDir`, and `bin` field for your fork. Affects CLI banner, config paths, and environment variable names. pi.dev
  • 203book.md:203 **Always use `src/config.ts`** for package assets: pi.dev
  • 248book.md:248 Guides and references for configuring and extending Pi. pi.dev

Rows are read from committed capture data, not a mock terminal.

How it works

One capture, several useful destinations.

The product has one job: make documentation easier for people and agents to trust.

  1. 01

    Capture

    Give DocHarvest a documentation URL. It detects the platform, stays inside the documentation path, and handles client-rendered sites when you opt into rendering.

    docharvest capture <url> --scope /docs
  2. 02

    Compile

    The engine writes page files with source URLs and content hashes, then builds a book, an llms.txt manifest, search records, and optional exports.

    pages/ book.md llms.txt search.db
  3. 03

    Use anywhere

    Read the files locally, search the corpus, open a PDF, load RAG JSONL, or connect an MCP client to the same library.

    docharvest mcp

Product walkthrough

See the capture before you install it.

These are real desktop surfaces from the shipped GUI. The browser showcase explains the same flow without pretending to be the application.

DocHarvest Capture Studio showing provider detection, download progress, and live capture logs
Capture Studio: detection, progress, and emitted files in one view.

Capture

The tool tells you what it found.

Provider detection, page counts, skipped URLs, and warnings remain visible while the capture runs.

  • 01

    Paste a docs URL

  • 02

    Confirm scope and provider

  • 03

    Open the local corpus

Inspect and export

The library keeps the work reusable.

Search captured domains, open their folders, compare snapshots, and send the same files to a RAG pipeline or a PDF reader.

~/.docharvest/docs/<domain>/
DocHarvest Document Library showing captured domains, providers, pages, and export actions
Document Library: one place to search, open, diff, and export captures.

Outputs

A capture is a folder, not a black box.

Every artifact has a filename, a consumer, and enough provenance to inspect what happened.

DocHarvest output artifacts
ArtifactUseEvidence inside
pages/One Markdown file per captured pagesource_url + content_hash frontmatter
book.md / docs.mdOne combined handbook with a table of contentslocal output / global library
llms.txtA compact index for AI discoverycaptured titles, paths, and source URLs
search.dbSQLite FTS5 BM25 search indexpage and heading records
exports/*_rag.jsonlChunked records for vector storesmetadata, headings, and hashes
exports/*_handbook.pdfPrintable offline handbookpure-Python fpdf2 export
versions/v*.mdSemver snapshots for change reviewunified diff and changelog inputs

Connectivity

One local source for people and tools.

The capture does not disappear into a hosted dashboard. It stays in the library you can inspect, copy, index, or expose through MCP.

12 tools · 2 resources · 2 prompts

MCP clientConnect a client over stdio

.cursor/mcp.json or claude_desktop_config.json

{
  "mcpServers": {
    "docharvest": {
      "command": "uvx",
      "args": ["docharvest", "mcp"]
    }
  }
}
Agent SkillInstall the retrieval instructions

.agents/skills, .cursor/skills, or .claude/skills

docharvest skill install docharvest -o .agents/skills
Local consumerUse the files directly

~/.docharvest/docs/<domain>/

pages/
book.md
llms.txt
search.db
versions/

Supported platforms

Detect the platform. Keep the content.

Dedicated detectors handle the documentation systems most teams publish with. A generic provider remains available for custom portals and SPAs.

DocHarvest dedicated documentation platform detectors
DetectorPriorityReference site
GitBook100https://docs.openalgo.in/v/v2.0/api-reference
Mintlify90https://docs.anthropic.com/en/docs
Docusaurus80https://reactnative.dev/docs/getting-started
Nextra75https://swr.vercel.app/docs/getting-started
VitePress72https://vitepress.dev/guide/what-is-vitepress
MkDocs70https://squidfunk.github.io/mkdocs-material/
ReadMe.io65https://docs.readme.com/reference
ReadTheDocs60https://docs.readthedocs.io/en/stable/
Generic HTML and SPA0Fallback for portals without a dedicated detector

8 dedicated detectors plus a generic fallback. Detection is automatic.

Proof and benchmarks

Numbers with a stated context.

These values describe a reference capture and the current repository. They are not a promise about every source site.

Reference capture
673 pages

in 18.2 seconds, about 37 pages per second

Measured reduction
~83%

fewer tokens than the raw HTML comparison

Agent surface
12 tools

plus resources and prompts over stdio

Repository verification
920 of 929

8 currently failing in the root suite, 2026-10-06

Source: `uv run pytest` (2026-10-06) and the reference capture in docs/SEO_GUIDE.md §3. Test counts are regenerated by docs/scripts/sync-stats.mjs; benchmark values are reference measurements, not guarantees.

FAQ

Before you capture a site.

Short answers about access, privacy, output scope, and agent connections.

Isn't this just another web scraper? How is it different from curl, Scrapy, or BeautifulSoup?

Basic scrapers dump messy HTML soup loaded with 40KB+ of cookie banners, navigation menus, and fragmented code blocks with broken indentation. DocHarvest is an engineered documentation compiler: it automatically detects frameworks (GitBook, Mintlify, Docusaurus, Nextra), probes native .md raw endpoints directly, locks crawls strictly to doc subpaths, injects cryptographic SHA-256 YAML frontmatter, compiles unified book.md handbooks, exports pure-Python PDFs, and indexes everything into an embedded SQLite FTS5 BM25 search database.

Why choose DocHarvest over Firecrawl, Jina Reader, or cloud scraper APIs?

Cloud scraping APIs charge per-page fees ($0.01 - $0.05/page) that quickly escalate on 1,000+ page libraries, require active internet connections, send your proprietary internal docs to third-party servers, and do not provide local search libraries, PDF generation, or semver diff engines. DocHarvest is 100% free, open-source (MIT), runs locally on your machine, and has zero network telemetry.

Which AI coding agents and IDEs support DocHarvest FastMCP?

DocHarvest's FastMCP v2 server is fully standard-compliant over stdio and ships ready-made configs for 14 documented clients: Claude Code, Claude Desktop, Cursor, Windsurf, VS Code, JetBrains, Zed, Cline, Continue.dev, Kiro, OpenCode, Oh My Pi (omp.sh), Antigravity / Gemini CLI, and OpenAI Codex CLI.

Does it work with client-rendered JavaScript Single-Page Applications (SPAs)?

Yes. Modern documentation SPAs (GitBook, Mintlify, Docusaurus, Nextra, VitePress) publish underlying raw .md endpoints and sitemaps that DocHarvest probes first. For purely client-rendered SPAs (like omp.sh), DocHarvest includes an opt-in Playwright headless rendering engine (--render) to execute client-side JavaScript before compilation.

What dependencies are needed for PDF export? Do I need WeasyPrint or wkhtmltopdf?

Zero external C-dependencies! DocHarvest uses a custom layout engine built on pure-Python fpdf2. It generates styled, syntax-highlighted printable PDF handbooks with page numbers and table of contents out of the box on Windows, macOS, and Linux.