VERA Collection Indexes¶
VERA keeps one portable .vera archive per source document. A collection index
is a disposable acceleration artifact for searching many archives as one
library; it is never the source of truth for text, citations, figures, or the
original document.
Local index layout¶
vera index build <root> --recursive creates:
<root>/
.vera-index/
current.json
generations/
generation-<id>/
index.sqlite3
vectors-<model-hash>-<dimension>.npy
... nested .vera files ...
SQLite stores:
- persisted discovery settings and index version
- root-relative file paths, fingerprints, and source metadata
- chunk-to-file references
- a unified FTS5 keyword index
- vector-group manifests
The NumPy files store normalized, contiguous float32 matrices grouped by
embedding model and dimension. Semantic queries run as one batched matrix
operation per model group. Results from mixed models are rank-fused, then the
winning chunk IDs are resolved against their source .vera files.
Index builds L2-normalize each non-zero vector regardless of the source
archive's declared normalization policy, preserving cosine-similarity ranking.
If a query embedder cannot be loaded or has a different dimension,
skipped_semantic_model_groups reports the omitted group and error; keyword
search remains available.
Index status and the desktop Library Info view report the active generation,
build and verification times, database/vector/total storage, indexed and
discovered document counts, exclusions, skipped archives, and embedding
coverage. Model groups are listed separately with model name, vector dimension,
document count, chunk count, and vector-file size. Status JSON
indexed_chunks is the number of chunk rows written into FTS and vector
matrices; source_chunks is the chunk-row count from archives that were
successfully indexed. Current builds keep these equal. Indexes created before
source-chunk coverage was recorded omit source_chunks and treat indexed
chunks as the source total until they are rebuilt.
See Library index structure for diagrams of the generation layout, SQLite relationships, vector mapping, and indexed search path.
Lifecycle and fallback¶
vera index buildindexes into a unique temporary sibling (.vera-index.build-<uuid>/), validates it, then takes an exclusive.vera-index/build.lockto move the tree undergenerations/and replacecurrent.json. A successful publish deletes every other generation directory. There is no separate garbage-collection command. Concurrent readers can keep the previous generation open during the pointer swap; cleanup usesignore_errors=True, so a Windows process that still has files open may leave remnants rather than fail the build. Two builds may index in parallel; only publication is serialized, and the last successful publish wins.- A directory with no discovered
.verafiles raises unstructuredNo .vera files found in ...(exit 1, no JSON). Recursive discovery is off by default. vera index updaterebuilds with the saved recursive and exclusion settings and reports added, changed, moved, and removed archives.- The index persists across app and CLI restarts. Opening a library checks its freshness but does not rebuild it; rebuilds occur only through an explicit build or update action.
vera index statuscompares the manifest with the current library and checks file content hashes (verify_hashesdefaults to true), the SQLite database, and vector matrix shapes. JSON setsverified_atwhen hashes ran.vera search <root>uses a fresh index automatically. A missing or stale index falls back to direct corpus search using the saved discovery settings. Automatic searches andVeraCorpus.openuse a fast size/mtime freshness check (verify_hashes=false;verified_atis null). The desktop index badge uses the same fast check; Inspect on a library folder refreshes with full hashes. A same-size, same-mtime byte change is not stale until that full verification.
Skipped archives are recorded so they stay visible in build reports without
making an otherwise valid index permanently stale. Build/update JSON lists
them as invalid (validation or open failure) or incompatible (a chunk
vector length that does not match the declared embedding dimension).
vera index status --json repeats those rows in skipped_files with
relative paths, category, and reason. If every discovered archive is
skipped, build raises No valid .vera files could be indexed with no JSON
report. Folder inspection uses the skip manifest when the index is fresh and
does not reopen known-invalid archives.
The desktop app also reads library counts and source metadata directly from a
fresh index when opening a folder, avoiding a full validation scan of every
archive. Its explicit Inspect action still reopens and validates all
archives when a current health check is needed.
Performance baseline¶
The deterministic benchmark in dev/benchmarks/benchmark_corpus.py generated 100
archives with 100 chunks each (10,000 chunks total). On the development Windows
machine, three hybrid searches produced:
- direct fan-out median: 0.195 seconds
- local index median: 0.009 seconds
- index build: 1.071 seconds
- expected-file hit rate: 100% for both paths
- traced Python peak: 2.0 MB fan-out and 0.3 MB indexed
- index size: 21.2 MB for a 36.5 MB synthetic library
This synthetic result is hardware-specific, but it demonstrates that contiguous local search removes most per-file SQLite and vector-deserialization overhead. Run the benchmark with proposal-like document and chunk counts before choosing a remote backend.
Optional external backends¶
External adapters remain outside the first implementation. Introduce one only when measurement demonstrates at least one of these requirements:
- the exact local matrix search misses the application's latency target
- vector matrices no longer fit comfortably on the serving machine
- multiple processes or hosts must update and query one shared index
- replication, tenant isolation, remote filtering, or service-level backups are required
A future backend contract should be additive and small:
class CollectionSearchBackend(Protocol):
def build(self, root: str, files: list[str], config: dict) -> dict: ...
def update(self, root: str) -> dict: ...
def status(self, root: str) -> dict: ...
def search(self, query: str, mode: str, top_k: int) -> list[IndexHit]: ...
def close(self) -> None: ...
Candidate adapters belong in an optional vera-index package:
sqlite-vecfor local ANN search when its packaging and Windows extension behavior are acceptable- FAISS or HNSW for a larger local in-process index
- Qdrant for shared, concurrent, remotely operated collections
Every adapter must return root-relative .vera paths and chunk IDs. For
Qdrant, those values and the embedding model/dimension belong in each point's
payload. The adapter may duplicate vectors and filter metadata, but complete
source text and citation geometry continue to come from .vera archives.