Convert documents¶
vera convert turns PDFs into portable .vera archives. Conversion parses
page layout, detects headings and figures, creates citation-ready chunks,
computes embeddings, builds the FTS5 keyword index, and normally stores the
original PDF.
Convert one PDF¶
Omit the output to create a same-named archive:
Single-file conversion builds and validates a temporary sibling archive before atomically replacing the output path. If parsing, writing, or validation fails, VERA removes the temporary file and preserves any existing output.
Conversion uses selective OCR by default. Native-text pages keep the fast PyMuPDF extraction path; image-dominant pages with little or no text are recognized locally with Tesseract. A PDF that still produces no chunks fails with a message that it may be scanned and requires OCR.
OCR¶
Automatic OCR is the default:
OCR runs only on pages that are mostly a scanned image and have too little native text to search reliably. Blank pages are skipped. Mixed PDFs can therefore use native extraction on ordinary pages and OCR on scanned pages in one conversion.
Controls:
--ocr autoselects scanned pages (default);--ocr offnever invokes OCR;--ocr forceOCRs every page, replacing native text extraction;--ocr-language engselects Tesseract language data;--ocr-dpi 300controls recognition resolution.
VERA bundles the official tessdata_fast English model and passes it directly
to PyMuPDF's Tesseract integration. Default English OCR therefore works
offline without installing Tesseract or configuring the system. Other
languages are not bundled; install their .traineddata files and set
TESSDATA_PREFIX when selecting them with --ocr-language. OCR failures name
the page and language and preserve any existing destination.
OCR text is stored as ordinary paragraph blocks with page bounding boxes, so search results and highlight regions work normally. This first OCR path targets scanned prose. It does not reconstruct scanned tables, forms, or complex multi-column reading order. Use an external layout-aware OCR tool for those documents; Docling is a candidate for a future optional parser if representative corpus testing shows that need.
Convert a directory¶
Directory conversion writes each archive beside its PDF:
Discover nested PDFs:
Existing archives are validated before they are skipped. A malformed existing archive is reported separately and is not silently preserved as a successful skip. Replace existing outputs explicitly:
Do not provide a single output path for directory conversion.
For a machine-readable batch report:
The report distinguishes discovered, converted, valid existing skips,
malformed existing outputs, and conversion failures. malformed_existing
entries include input, output, and validation issues. Batch conversion
continues after an individual PDF fails and exits nonzero if any conversion
failed or malformed existing output was found.
Embedding models¶
The default model is hashing:
It is deterministic, local, and requires no machine-learning package.
For neural embeddings, install the optional dependency and name a Sentence Transformers model:
python -m pip install "vera-cli>=0.2.2" "vera-doc[ml]>=0.2.2"
vera convert "input.pdf" --model sentence-transformers/all-MiniLM-L6-v2
The model name and vector dimension are recorded in the archive. Search uses
the recorded model, so the ml extra must also be installed on machines that
search an archive created with a Sentence Transformers model.
Use only hashing, vera-hashing-384, all-MiniLM-L6-v2, or a
sentence-transformers/... name. An unrecognized model name falls
back to hashing while retaining the requested name in metadata; that can make
the archive difficult to query consistently on another machine.
Chunking options¶
Defaults:
--chunk-size 500--overlap 75
Example:
Chunks never span pages, preserving page-precise citations. Larger chunks carry more context but may reduce retrieval precision; smaller chunks are more specific but may separate related clauses. Evaluate changes against a representative query set before adopting non-default values.
Parser¶
vera-ingest currently supports the pymupdf parser:
Other parser names currently fail.
Storing the source PDF¶
The original PDF is stored by default, enabling later export and document viewing. To omit it:
An archive created this way remains searchable, but:
vera exportcannot restore the source;- validation reports the missing original document as a warning;
- viewers cannot obtain the original PDF from the archive.
Verify conversion¶
After conversion:
Inspect confirms the source, page and chunk counts, parser, and embedding model. Validate checks SQLite integrity, required tables and metadata, embedding counts, FTS consistency, and the stored source document.
Python equivalent¶
from vera_ingest import convert
path = convert(
"input.pdf",
"output.vera",
model="hashing",
chunk_size=500,
overlap=75,
store_original=True,
ocr_mode="auto",
ocr_language="eng",
ocr_dpi=300,
)
print(path)
See Python API for more.