Getting started¶
This tutorial installs the VERA CLI, converts a PDF or Markdown file into a
portable .vera archive, and searches it.
VERA is currently pre-1.0 and experimental. Preserve source documents and expect
API changes before a stable release. Release 0.3.x is the software, CLI, and
Python API version; the .vera archive format remains 0.2, so existing
archives stay compatible.
Requirements¶
- Python 3.10 or newer
- A local PDF or Markdown file
- Windows, macOS, or Linux
Install from PyPI¶
Install the CLI and its dependencies:
That installs vera-doc, vera-ingest, vera-ingest-pymupdf, and
vera-embed-openai as well.
Add MCP support with:
Verify that the console script is available:
If vera is not on PATH, invoke the same CLI as a Python module:
Library-only installs¶
# Storage and search only
python -m pip install "vera-doc>=0.3.0"
# PDF conversion and viewer helpers
python -m pip install "vera-ingest>=0.3.0" "vera-ingest-pymupdf>=0.3.0"
# MCP server package
python -m pip install "vera-mcp>=0.3.0"
Install from source¶
Contributors can clone the repository and install workspace packages:
git clone https://github.com/dkylewillis/vera.git
cd vera
python -m pip install ./packages/vera-doc ./packages/vera-ingest ./packages/vera-ingest-pymupdf ./packages/vera-embed-openai ./packages/vera-cli
Or with uv:
Convert a PDF or Markdown file¶
The default hashing embedding model is local and has no machine-learning
dependency. The resulting archive contains parsed pages or Markdown structure,
chunks, embeddings, the keyword index, citation metadata, figures (PDF), and
the original source file.
If the output path is omitted, VERA uses the input filename with a .vera
suffix:
VERA validates the completed archive before publishing it. Image-based pages
with little or no native text are OCR-processed locally by default via
vera-ingest-pymupdf. English language data is bundled, so default OCR works
offline without another installation. Selecting another language with
--ocr-language requires its Tesseract data — list or fetch curated packs
with vera ocr-languages, or pass --ocr-allow-download during convert. Use
--ocr off to disable recognition or --ocr force to process every page. A
failed conversion does not replace an existing destination.
Inspect and validate¶
Inspect the archive:
Validate integrity:
Both commands accept --json for machine-readable output. Text-mode
inspect omits the pipeline ocr diagnostics bag (OCR pages, Docling
recovery); use --json when you need those fields.
Search with citations¶
Each result includes page range and heading path for grounded citations. Add
--figures, --regions, or --context-chunks 1 when you need visual or
neighboring context.