Search documents¶
VERA supports semantic, keyword, and hybrid search over the content already stored in an archive. Search is local and does not require a retrieval server.
Basic search¶
The default mode is hybrid and the default result limit is 10:
Use --json for scripts and agents:
Use --pretty for clean Markdown-like context that can be read directly or
pasted into a prompt:
vera search "manual.vera" "stormwater detention requirements" \
--top-k 5 --context-chunks 1 --pretty
Pretty output includes each result's archive (for corpus searches), heading,
source filename, page or page range, and complete text. Requested neighboring
chunks are labeled as previous, matching, and following context. --figures
also adds figure captions and pages. Scores, chunk ids, and region coordinates
remain available through --json. --pretty and --json are mutually
exclusive.
Choose a search mode¶
Hybrid¶
Hybrid combines min-max-normalized semantic and keyword rankings with equal weights by default. The Python API can change the blend:
results = document.search(
"when must runoff be detained?",
mode="hybrid",
semantic_weight=0.7,
keyword_weight=0.3,
)
print(results[0].citation.page_start, results[0].citation.heading_path)
Use hybrid for most questions, especially when both concepts and document terminology matter.
Keyword¶
Keyword search uses SQLite FTS5. Use it for exact phrases, section numbers, identifiers, table labels, and known terminology.
Keyword search first runs the raw query as an FTS5 MATCH. If that returns
no rows — including when SQLite reports an FTS syntax error such as an
unbalanced quote — VERA retries with safe_fts_query: whitespace-split
tokens, keep letters/digits/_, drop other punctuation, append *, join
with OR. storm-water detention! becomes stormwater* OR detention*.
Queries that contain no alphanumeric tokens (!!!, "") produce an empty
fallback and return no hits. Runtime FTS failures (locked database,
malformed file, missing chunks_fts) are not treated as syntax errors
and still raise. The collection index uses the same helper.
Punctuation-stripping can merge a hyphenated code such as EL-A into
ELA*. Confirm that the literal identifier appears in the returned text
before treating the hit as an exact match.
Semantic¶
Semantic search compares the query embedding with stored chunk embeddings. Use it when the wording is likely to differ from the document.
Set up meaning-based search¶
The default hashing embedder is based on shared words; it does not learn
meaning or synonyms. Selecting --mode semantic alone does not upgrade those
vectors. Convert the source with a neural model first:
python -m pip install "vera-doc[ml]"
vera convert manual.pdf manual-semantic.vera --model sentence-transformers:all-MiniLM-L6-v2
vera search manual-semantic.vera "how can we prevent downstream flooding?" --mode semantic --pretty
The first conversion downloads the model. Once cached, the model can run offline. Semantic and hybrid search resolve the model recorded in the archive to embed the query; the archive stores document vectors, not model weights. Keyword search needs no embedding runtime or API key.
To change the model for an existing library, reconvert with
vera convert ./library --recursive --overwrite --model sentence-transformers:all-MiniLM-L6-v2,
then run vera index update ./library. The --overwrite flag replaces batch
outputs: unchanged source files would otherwise be skipped, even if --model
changed. Building an index alone does not replace the stored embeddings.
Other models use the same --model provider:model-id syntax. See
embedding setup or the
minimal embedding plugin.
Interpret results¶
Every result contains:
chunk_idscoretextpage_startandpage_endheading_pathsource_filenamedocument_id
In Python these citation fields also appear on result.citation. CLI, MCP, and
desktop JSON flatten the same keys onto the result object — there is no nested
citation field in those payloads.
Scores rank results within a search. They are not probabilities or confidence
values, and scores from different queries or modes should not be compared as
though they share one scale. Desktop Ask additionally drops weak hits with a
relative quality cutoff against the top score; CLI, MCP, and the Search view
return the unfiltered ranked list.
Treat the text and its location as evidence. A citation should include the source filename, page or page range, and heading when available:
Filter before top_k¶
Scope a search with stored metadata rather than by dropping hits from JSON.
--where is applied before top_k, so rank and result counts stay honest.
vera search "./library" "adding capacity" --where company=GRID --json
vera search "./library" "adding capacity" --where company=GRID,PWRX --json
vera search "./library" "adding capacity" \
--where company=GRID --where document_type=filings --json
Distinct --where keys are AND. Comma-separated values for one key are IN.
Repeated flags for the same key union the IN set. Values coerce like
--pipeline-option (digit-only integers, boolean words, otherwise strings;
dotted tokens such as 3.10 stay strings). A missing key fails the predicate.
List-valued stored metadata is not an IN clause. Stamp those
keys at convert time:
--include and --exclude choose files by relative path (discovery), not
metadata. --include on a single-file search exits 2. Convert --metadata
tags match everywhere. Convert-owned archive headers such as
source_file_name match indexed directory search; single-file and fallback
search evaluate chunk metadata only. Desktop Search and Ask do not expose
--where; use the CLI or MCP.
To reload one stored chunk by id — for example to verify that a quoted span is
still in the chunk body — use vera get FILE CHUNK_ID --json (MCP:
vera_get_chunk). Searching again is not a substitute: rank can change, and
keyword search can miss a short quote. get returns the same citation fields
as a search hit, without score.
Include neighboring chunks¶
A result may begin after a definition or end before an exception. Include neighboring chunks:
Each result gains before_chunks and after_chunks. These chunks are ordered
in document sequence and carry their own citation fields.
Find figures and page regions¶
Add figure metadata:
Add source block bounding boxes:
See Figures and highlight regions for the coordinate contract and limitations.
Improve a weak search¶
If results are too broad:
- add the governing action:
requirements,definition,exceptions; - add a section, district, facility type, or threshold;
- switch to keyword mode for exact language.
If results are sparse:
- remove one constraint;
- use a likely synonym;
- search the parent concept;
- switch to semantic mode;
- increase
--top-k.
Split compound questions into separate searches. For comprehensive research, use several targeted queries and synthesize only claims supported by the retrieved text.
Empty and failed searches¶
An empty successful search returns results: [] and exits 0. It means no
candidate was returned for that query and mode; it does not prove the concept
is absent from the document.
Missing paths, unreadable archives, unavailable embedding dependencies, and directories with no archives generally exit nonzero and write an error to stderr. Check the process exit code before parsing JSON.
For a directory containing both healthy and malformed archives, search
continues across the healthy subset and returns exit 0. Inspect the top-level
skipped_files array for excluded paths and validation reasons. A directory
with no discoverable archives is still an error.
Search a library¶
Pass a directory instead of a file to search multiple archives as one corpus:
See Document libraries for recursive discovery, exclusions, collection indexes, stale-index fallback, and mixed embedding models.