Skip to content

Search documents

VERA supports semantic, keyword, and hybrid search over the content already stored in an archive. Search is local and does not require a retrieval server.

vera search "manual.vera" "stormwater detention requirements"

The default mode is hybrid and the default result limit is 10:

vera search "manual.vera" "stormwater detention requirements" --mode hybrid --top-k 10

Use --json for scripts and agents:

vera search "manual.vera" "stormwater detention requirements" --top-k 5 --json

Use --pretty for clean Markdown-like context that can be read directly or pasted into a prompt:

vera search "manual.vera" "stormwater detention requirements" \
  --top-k 5 --context-chunks 1 --pretty

Pretty output includes each result's archive (for corpus searches), heading, source filename, page or page range, and complete text. Requested neighboring chunks are labeled as previous, matching, and following context. --figures also adds figure captions and pages. Scores, chunk ids, and region coordinates remain available through --json. --pretty and --json are mutually exclusive.

Choose a search mode

Hybrid

vera search "manual.vera" "when must runoff be detained?" --mode hybrid

Hybrid combines min-max-normalized semantic and keyword rankings with equal weights by default. The Python API can change the blend:

results = document.search(
    "when must runoff be detained?",
    mode="hybrid",
    semantic_weight=0.7,
    keyword_weight=0.3,
)
print(results[0].citation.page_start, results[0].citation.heading_path)

Use hybrid for most questions, especially when both concepts and document terminology matter.

Keyword

vera search "manual.vera" "\"Section 4.2\"" --mode keyword

Keyword search uses SQLite FTS5. Use it for exact phrases, section numbers, identifiers, table labels, and known terminology.

Keyword search first runs the raw query as an FTS5 MATCH. If that returns no rows — including when SQLite reports an FTS syntax error such as an unbalanced quote — VERA retries with safe_fts_query: whitespace-split tokens, keep letters/digits/_, drop other punctuation, append *, join with OR. storm-water detention! becomes stormwater* OR detention*. Queries that contain no alphanumeric tokens (!!!, "") produce an empty fallback and return no hits. Runtime FTS failures (locked database, malformed file, missing chunks_fts) are not treated as syntax errors and still raise. The collection index uses the same helper.

Punctuation-stripping can merge a hyphenated code such as EL-A into ELA*. Confirm that the literal identifier appears in the returned text before treating the hit as an exact match.

Semantic

vera search "manual.vera" "how does the site reduce peak flow?" --mode semantic

Semantic search compares the query embedding with stored chunk embeddings. Use it when the wording is likely to differ from the document.

The default hashing embedder is based on shared words; it does not learn meaning or synonyms. Selecting --mode semantic alone does not upgrade those vectors. Convert the source with a neural model first:

python -m pip install "vera-doc[ml]"
vera convert manual.pdf manual-semantic.vera --model sentence-transformers:all-MiniLM-L6-v2
vera search manual-semantic.vera "how can we prevent downstream flooding?" --mode semantic --pretty

The first conversion downloads the model. Once cached, the model can run offline. Semantic and hybrid search resolve the model recorded in the archive to embed the query; the archive stores document vectors, not model weights. Keyword search needs no embedding runtime or API key.

To change the model for an existing library, reconvert with vera convert ./library --recursive --overwrite --model sentence-transformers:all-MiniLM-L6-v2, then run vera index update ./library. The --overwrite flag replaces batch outputs: unchanged source files would otherwise be skipped, even if --model changed. Building an index alone does not replace the stored embeddings.

Other models use the same --model provider:model-id syntax. See embedding setup or the minimal embedding plugin.

Interpret results

Every result contains:

  • chunk_id
  • score
  • text
  • page_start and page_end
  • heading_path
  • source_filename
  • document_id

In Python these citation fields also appear on result.citation. CLI, MCP, and desktop JSON flatten the same keys onto the result object — there is no nested citation field in those payloads.

Scores rank results within a search. They are not probabilities or confidence values, and scores from different queries or modes should not be compared as though they share one scale. Desktop Ask additionally drops weak hits with a relative quality cutoff against the top score; CLI, MCP, and the Search view return the unfiltered ranked list.

Treat the text and its location as evidence. A citation should include the source filename, page or page range, and heading when available:

(manual.pdf, p. 117, Chapter 4 > Detention Design)

Filter before top_k

Scope a search with stored metadata rather than by dropping hits from JSON. --where is applied before top_k, so rank and result counts stay honest.

vera search "./library" "adding capacity" --where company=GRID --json
vera search "./library" "adding capacity" --where company=GRID,PWRX --json
vera search "./library" "adding capacity" \
  --where company=GRID --where document_type=filings --json

Distinct --where keys are AND. Comma-separated values for one key are IN. Repeated flags for the same key union the IN set. Values coerce like --pipeline-option (digit-only integers, boolean words, otherwise strings; dotted tokens such as 3.10 stay strings). A missing key fails the predicate. List-valued stored metadata is not an IN clause. Stamp those keys at convert time:

vera convert filing.md archives/src_aaa.vera --metadata company=GRID --json

--include and --exclude choose files by relative path (discovery), not metadata. --include on a single-file search exits 2. Convert --metadata tags match everywhere. Convert-owned archive headers such as source_file_name match indexed directory search; single-file and fallback search evaluate chunk metadata only. Desktop Search and Ask do not expose --where; use the CLI or MCP.

To reload one stored chunk by id — for example to verify that a quoted span is still in the chunk body — use vera get FILE CHUNK_ID --json (MCP: vera_get_chunk). Searching again is not a substitute: rank can change, and keyword search can miss a short quote. get returns the same citation fields as a search hit, without score.

Include neighboring chunks

A result may begin after a definition or end before an exception. Include neighboring chunks:

vera search "manual.vera" "detention requirements" --context-chunks 1 --json

Each result gains before_chunks and after_chunks. These chunks are ordered in document sequence and carry their own citation fields.

Find figures and page regions

Add figure metadata:

vera search "manual.vera" "pipe sizing chart" --figures --json

Add source block bounding boxes:

vera search "manual.vera" "detention requirements" --regions --json

See Figures and highlight regions for the coordinate contract and limitations.

If results are too broad:

  • add the governing action: requirements, definition, exceptions;
  • add a section, district, facility type, or threshold;
  • switch to keyword mode for exact language.

If results are sparse:

  • remove one constraint;
  • use a likely synonym;
  • search the parent concept;
  • switch to semantic mode;
  • increase --top-k.

Split compound questions into separate searches. For comprehensive research, use several targeted queries and synthesize only claims supported by the retrieved text.

Empty and failed searches

An empty successful search returns results: [] and exits 0. It means no candidate was returned for that query and mode; it does not prove the concept is absent from the document.

Missing paths, unreadable archives, unavailable embedding dependencies, and directories with no archives generally exit nonzero and write an error to stderr. Check the process exit code before parsing JSON.

For a directory containing both healthy and malformed archives, search continues across the healthy subset and returns exit 0. Inspect the top-level skipped_files array for excluded paths and validation reasons. A directory with no discoverable archives is still an error.

Search a library

Pass a directory instead of a file to search multiple archives as one corpus:

vera search "./library" "termination clause" --json

See Document libraries for recursive discovery, exclusions, collection indexes, stale-index fallback, and mixed embedding models.