Skip to content

Getting started

This tutorial installs the VERA CLI, converts a PDF or Markdown file into a portable .vera archive, and searches it.

VERA is currently pre-1.0 and experimental. Preserve source documents and expect API changes before a stable release. Release 0.3.x is the software, CLI, and Python API version; the .vera archive format remains 0.2, so existing archives stay compatible.

Requirements

  • Python 3.10 or newer
  • A local PDF or Markdown file
  • Windows, macOS, or Linux

Install from PyPI

Install the CLI and its dependencies:

python -m pip install "vera>=0.3.2"

That installs vera-doc, vera-ingest, vera-ingest-pymupdf, and vera-embed-openai as well. Add MCP support with:

python -m pip install "vera[mcp]>=0.3.2"

Verify that the console script is available:

vera --help

If vera is not on PATH, invoke the same CLI as a Python module:

python -m vera_cli --help

Library-only installs

# Storage and search only
python -m pip install "vera-doc>=0.3.0"

# PDF conversion and viewer helpers
python -m pip install "vera-ingest>=0.3.0" "vera-ingest-pymupdf>=0.3.0"

# MCP server package
python -m pip install "vera-mcp>=0.3.0"

Install from source

Contributors can clone the repository and install workspace packages:

git clone https://github.com/dkylewillis/vera.git
cd vera
python -m pip install ./packages/vera-doc ./packages/vera-ingest ./packages/vera-ingest-pymupdf ./packages/vera-embed-openai ./packages/vera-cli

Or with uv:

uv sync --extra dev
uv run python -m vera_cli --help

Convert a PDF or Markdown file

vera convert "manual.pdf" "manual.vera"
vera convert "notes.md" "notes.vera"

The default hashing embedding model is local and has no machine-learning dependency. The resulting archive contains parsed pages or Markdown structure, chunks, embeddings, the keyword index, citation metadata, figures (PDF), and the original source file.

If the output path is omitted, VERA uses the input filename with a .vera suffix:

vera convert "manual.pdf"

VERA validates the completed archive before publishing it. Image-based pages with little or no native text are OCR-processed locally by default via vera-ingest-pymupdf. English language data is bundled, so default OCR works offline without another installation. Selecting another language with --ocr-language requires its Tesseract data — list or fetch curated packs with vera ocr-languages, or pass --ocr-allow-download during convert. Use --ocr off to disable recognition or --ocr force to process every page. A failed conversion does not replace an existing destination.

Inspect and validate

Inspect the archive:

vera inspect "manual.vera"

Validate integrity:

vera validate "manual.vera"

Both commands accept --json for machine-readable output. Text-mode inspect omits the pipeline ocr diagnostics bag (OCR pages, Docling recovery); use --json when you need those fields.

Search with citations

vera search "manual.vera" "stormwater detention requirements" --top-k 5 --json

Each result includes page range and heading path for grounded citations. Add --figures, --regions, or --context-chunks 1 when you need visual or neighboring context.

Next steps