Examples and recipes¶
These recipes use the vera console script. Substitute
python -m vera_cli if the console script is not on PATH.
Convert, index, and search a library¶
Start with a folder of PDFs or Markdown files. Conversion writes .vera
files beside their sources:
vera convert ./library --recursive
vera index build ./library --recursive
vera search ./library "stormwater detention" --mode keyword --pretty
This needs no model download or API key. For meaning-based search, install the neural runtime and replace the hashing embeddings by reconverting:
python -m pip install "vera-doc[ml]"
vera convert ./library --recursive --overwrite --model sentence-transformers:all-MiniLM-L6-v2
vera index update ./library
vera search ./library "how can we prevent downstream flooding?" --mode semantic --pretty
The first use downloads the model; it can run offline once cached.
--overwrite is needed here because batch conversion otherwise skips
unchanged sources even when the requested embedding model differs.
Tag a document and filter before the result limit:
vera convert manual.pdf ./library/manual.vera --metadata project=riverpark --metadata type=manual
vera index update ./library
vera search ./library "stormwater detention" --where project=riverpark --where type=manual --top-k 5 --pretty
The last conversion uses the default hashing model. Add --model if you want
that archive to use neural embeddings too. Libraries can contain both.
To try a custom provider, install the two-file example from a repository clone:
python -m pip install ./examples/embedding-plugin
vera convert manual.pdf manual-custom.vera --model custom:all-MiniLM-L6-v2
vera search manual-custom.vera "how can we prevent downstream flooding?" --mode semantic --pretty
See creating an embedding provider for the complete source and packaging walkthrough.
Convert and search one document¶
vera convert "ordinance.pdf" "ordinance.vera" --model hashing
vera convert "notes.md" "notes.vera" --model hashing
vera convert "memo.docx" "memo.vera" --parser docling --model hashing
vera convert "notes.html" "notes.vera" --model hashing
vera inspect "ordinance.vera"
vera validate "ordinance.vera"
vera search "ordinance.vera" "minimum parking required for restaurant" --mode hybrid --top-k 5 --json
vera get "ordinance.vera" "chunk_0001" --json
Convert a scanned PDF¶
Automatic mode uses native text where available and OCRs image-based low-text pages:
vera convert "scanned-manual.pdf" "scanned-manual.vera" --ocr auto --ocr-language eng
vera search "scanned-manual.vera" "emergency shutdown procedure" --json --regions
Use --ocr force only when automatic detection misses a scanned page.
English language data is bundled. For another Tesseract language:
vera ocr-languages list fra --json
vera ocr-languages download fra --json
vera convert "scanned-manual.pdf" "scanned-manual.vera" --ocr-language fra
--ocr-allow-download fetches the same curated pack during convert. Codes
outside the registry still need a manual TESSDATA_PREFIX install.
Get JSON with surrounding context¶
vera search "ordinance.vera" \
"restaurant parking requirements" \
--mode hybrid \
--top-k 5 \
--context-chunks 1 \
--json
For a readable version of the same context, replace --json with --pretty:
vera search "ordinance.vera" \
"restaurant parking requirements" \
--mode hybrid \
--top-k 5 \
--context-chunks 1 \
--pretty
PowerShell equivalent:
vera search "ordinance.vera" `
"restaurant parking requirements" `
--mode hybrid `
--top-k 5 `
--context-chunks 1 `
--json
Find an exact identifier¶
Confirm that EL-A appears literally in the result text. Keyword fallback can
remove punctuation and broaden short identifiers.
Find a figure or chart¶
The CLI returns metadata and captions, not image bytes.
Write stored rasters when you need to look at them:
Each object then includes path. Attach that file. MCP clients should call
vera_get_figure with the asset_id instead of writing files.
Tables extracted as selectable text are markdown in the hit, not figure attachments.
Get source highlight regions¶
The returned top-left-origin bounding boxes can be scaled onto a page viewer.
Convert and index a nested library¶
vera convert "./proposals" --recursive --json
vera index build "./proposals" --recursive --exclude "archive/**" --include "active/**" --json
vera search "./proposals" "termination clause" --where company=GRID --top-k 10 --json
Check malformed_existing after conversion, then skipped_files and
skipped_semantic_model_groups after search. These diagnostics identify broken
archives and unavailable or incompatible semantic model groups without
preventing healthy documents or keyword matches in the same library from being
converted, inspected, or searched.
After adding or replacing documents:
index status exits 1 when the index is stale or missing while still printing
a JSON report.
Compare several documents¶
Each corpus result includes file. Group findings and citations by source
archive rather than merging conflicting provisions.
Export the embedded source¶
Evaluate retrieval changes¶
Create queries.json:
[
{
"query": "restaurant parking",
"expected_pages": [42, 43],
"expected_terms": ["parking"],
"note": "Parking schedule"
}
]
Run all search modes:
The command exits 1 if any expected answer is missed. See Evaluate retrieval quality for hit rules and metric guidance.
Search from Python¶
from vera_doc import VeraDocument
doc = VeraDocument.open("ordinance.vera")
try:
for result in doc.search("restaurant parking", mode="hybrid", top_k=5):
print(result.score, result.page_start, result.heading_path)
print(result.text)
finally:
doc.close()
Search a library from Python¶
from vera_doc import VeraCorpus
with VeraCorpus.open("./proposals", recursive=True) as corpus:
for result in corpus.search("termination clause", top_k=10):
print(result.file, result.page_start, result.text[:100])
See the getting-started tutorial, search guide, and Python API for details.