Skip to content

vera-doc

vera-doc publishes the vera_doc Python package. It owns the portable SQLite format implementation, typed records, transactional CRUD, embeddings, search, corpus queries, and rebuildable library indexes.

It intentionally does not parse PDFs, perform OCR, provide the CLI, expose MCP tools, or implement the desktop application.

Install

From PyPI:

python -m pip install "vera-doc>=0.3.0"

From a repository checkout:

python -m pip install ./packages/vera-doc

Python 3.10 or newer is required. The default hashing embedder works locally without a model download or API key. MiniLM neural embeddings use the ml extra (Sentence Transformers), or the optional onnx extra (ONNX Runtime) plus a MiniLM ONNX snapshot. MiniLM prefers ONNX Runtime when a graph is present and falls back to Sentence Transformers otherwise. The Windows installer freezes ONNX Runtime and vendors a VERA-exported all-MiniLM-L6-v2 graph. Other Hub Sentence Transformers models always use the ml extra. Additional providers can be registered with register_embedder or discovered through the vera.embedders entry-point group (optional vera.embedder_descriptors for schema-driven options); unknown model specs raise UnknownEmbeddingModelError. See Creating an embedding provider plugin.

Start here

from vera_doc import ChunkRecord, VeraDocument

with VeraDocument.create("knowledge.vera") as document:
    document.add([
        ChunkRecord(
            id="chunk-1",
            text="The minimum pipe diameter is 12 inches.",
            metadata={"source_filename": "manual.pdf", "page_start": 42},
        )
    ])

with VeraDocument.open("knowledge.vera") as document:
    results = document.search(text="minimum pipe size", top_k=5)

Documentation

API reference