Skip to content

VERA Collection Indexes

VERA keeps one portable .vera archive per source document. A collection index is a disposable acceleration artifact for searching many archives as one library; it is never the source of truth for text, citations, figures, or the original document.

Local index layout

vera index build <root> --recursive creates:

<root>/
  .vera-index/
    current.json
    generations/
      generation-<id>/
        index.sqlite3
        vectors-<model-hash>-<dimension>.npy
  ... nested .vera files ...

SQLite stores:

  • persisted discovery settings and index version
  • root-relative file paths, fingerprints, and source metadata
  • chunk-to-file references
  • a unified FTS5 keyword index
  • vector-group manifests

The NumPy files store normalized, contiguous float32 matrices grouped by embedding model and dimension. Semantic queries run as one batched matrix operation per model group. Results from mixed models are rank-fused, then the winning chunk IDs are resolved against their source .vera files. Index builds L2-normalize each non-zero vector regardless of the source archive's declared normalization policy, preserving cosine-similarity ranking. If a query embedder cannot be loaded or has a different dimension, skipped_semantic_model_groups reports the omitted group and error; keyword search remains available.

Index status and the desktop Library Info view report the active generation, build and verification times, database/vector/total storage, indexed and discovered document counts, exclusions, skipped archives, and embedding coverage. Model groups are listed separately with model name, vector dimension, document count, chunk count, and vector-file size. Status JSON indexed_chunks is the number of chunk rows written into FTS and vector matrices; source_chunks is the chunk-row count from archives that were successfully indexed. Current builds keep these equal. Indexes created before source-chunk coverage was recorded omit source_chunks and treat indexed chunks as the source total until they are rebuilt.

See Library index structure for diagrams of the generation layout, SQLite relationships, vector mapping, and indexed search path.

Lifecycle and fallback

  • vera index build indexes into a unique temporary sibling (.vera-index.build-<uuid>/), validates it, then takes an exclusive .vera-index/build.lock to move the tree under generations/ and replace current.json. A successful publish deletes every other generation directory. There is no separate garbage-collection command. Concurrent readers can keep the previous generation open during the pointer swap; cleanup uses ignore_errors=True, so a Windows process that still has files open may leave remnants rather than fail the build. Two builds may index in parallel; only publication is serialized, and the last successful publish wins.
  • A directory with no discovered .vera files raises unstructured No .vera files found in ... (exit 1, no JSON). Recursive discovery is off by default.
  • vera index update rebuilds with the saved recursive and exclusion settings and reports added, changed, moved, and removed archives.
  • The index persists across app and CLI restarts. Opening a library checks its freshness but does not rebuild it; rebuilds occur only through an explicit build or update action.
  • vera index status compares the manifest with the current library and checks file content hashes (verify_hashes defaults to true), the SQLite database, and vector matrix shapes. JSON sets verified_at when hashes ran.
  • vera search <root> uses a fresh index automatically. A missing or stale index falls back to direct corpus search using the saved discovery settings. Automatic searches and VeraCorpus.open use a fast size/mtime freshness check (verify_hashes=false; verified_at is null). The desktop index badge uses the same fast check; Inspect on a library folder refreshes with full hashes. A same-size, same-mtime byte change is not stale until that full verification.

Skipped archives are recorded so they stay visible in build reports without making an otherwise valid index permanently stale. Build/update JSON lists them as invalid (validation or open failure) or incompatible (a chunk vector length that does not match the declared embedding dimension). vera index status --json repeats those rows in skipped_files with relative paths, category, and reason. If every discovered archive is skipped, build raises No valid .vera files could be indexed with no JSON report. Folder inspection uses the skip manifest when the index is fresh and does not reopen known-invalid archives. The desktop app also reads library counts and source metadata directly from a fresh index when opening a folder, avoiding a full validation scan of every archive. Its explicit Inspect action still reopens and validates all archives when a current health check is needed.

Performance baseline

The deterministic benchmark in dev/benchmarks/benchmark_corpus.py generated 100 archives with 100 chunks each (10,000 chunks total). On the development Windows machine, three hybrid searches produced:

  • direct fan-out median: 0.195 seconds
  • local index median: 0.009 seconds
  • index build: 1.071 seconds
  • expected-file hit rate: 100% for both paths
  • traced Python peak: 2.0 MB fan-out and 0.3 MB indexed
  • index size: 21.2 MB for a 36.5 MB synthetic library

This synthetic result is hardware-specific, but it demonstrates that contiguous local search removes most per-file SQLite and vector-deserialization overhead. Run the benchmark with proposal-like document and chunk counts before choosing a remote backend.

Optional external backends

External adapters remain outside the first implementation. Introduce one only when measurement demonstrates at least one of these requirements:

  • the exact local matrix search misses the application's latency target
  • vector matrices no longer fit comfortably on the serving machine
  • multiple processes or hosts must update and query one shared index
  • replication, tenant isolation, remote filtering, or service-level backups are required

A future backend contract should be additive and small:

class CollectionSearchBackend(Protocol):
    def build(self, root: str, files: list[str], config: dict) -> dict: ...
    def update(self, root: str) -> dict: ...
    def status(self, root: str) -> dict: ...
    def search(self, query: str, mode: str, top_k: int) -> list[IndexHit]: ...
    def close(self) -> None: ...

Candidate adapters belong in an optional vera-index package:

  • sqlite-vec for local ANN search when its packaging and Windows extension behavior are acceptable
  • FAISS or HNSW for a larger local in-process index
  • Qdrant for shared, concurrent, remotely operated collections

Every adapter must return root-relative .vera paths and chunk IDs. For Qdrant, those values and the embedding model/dimension belong in each point's payload. The adapter may duplicate vectors and filter metadata, but complete source text and citation geometry continue to come from .vera archives.