VERA Architecture¶
Package boundaries¶
VERA is a mono-repo of independently installable packages. Dependencies move inward toward the storage and search engine:
vera-ingest-pymupdf ─┐
vera-ingest-docling ─┤
vera-embed-openai ───┤
vera-ingest ─────────┼──> vera-doc
vera ────────────────┤
vera-app ────────────┤
vera-mcp ────────────┘
vera-lab (dev only) ─┘
No package may make vera-doc depend on extraction, a user interface, MCP,
evaluation tooling, or a source-file format. vera-lab is a contributor-only
leaf (workspace dev extra); it is not shipped in releases or the packaged
desktop sidecar.
vera-doc¶
vera-doc publishes the vera_doc Python package. It is an embedded vector
database backed by one portable SQLite .vera file.
It owns:
- the current
.veraschema and format validation; - immutable
ChunkRecord,AttachmentRecord, and query-result objects; - transactional chunk and attachment CRUD;
- embedding generation and storage;
- keyword, semantic, and hybrid retrieval;
- corpus search and rebuildable
.vera-index/library indexes.
It does not parse, clean, OCR, or chunk source content. It does not know what a
PDF is. Attachments are opaque bytes, and chunk/archive metadata is
JSON-compatible caller data. Format 0.1 is historical; vera-doc reads 0.2 archives only.
vera-ingest¶
vera-ingest publishes vera_ingest. It owns the provider-neutral ingest
core:
- ingest-pipeline registry and descriptors;
- shared types (
IngestRequest/IngestResult, pages/blocks); - reusable chunking helpers (
build_chunks_from_blocksfor structured layout;chunk_pages/detect_headingfor custom page-text pipelines); - atomic single-file and batch conversion workflows;
- reader helpers for pages, figures, regions, and source export.
PDF parsing and OCR live in plugin packages that register through
vera.ingest_pipelines. It emits final ChunkRecord objects and writes them
through VeraDocument. vera-doc never imports vera_ingest. Plugins are
named after the parsing engine (pymupdf, docling), not the file type;
planned extra formats and Markdown grounding stay in ingest conventions
(see Additional source formats and visual grounding)
and do not change the 0.2 archive schema.
vera-ingest-pymupdf¶
vera-ingest-pymupdf registers the default pymupdf pipeline:
- PDF parsing and table extraction (PyMuPDF + pdfplumber);
- selective OCR and bundled Tesseract English data;
- heading detection and sliding-window chunk construction (whitespace-split words);
- mapping pages, regions, figures, and provenance to chunk metadata.
vera and vera-app depend on it so conversion works out of the box.
vera-ingest-docling¶
vera-ingest-docling is the official Docling pipeline that registers
docling / docling:hybrid. CLI users install it with
pip install "vera[docling]>=0.3.0" or uv sync --extra docling. The
0.3.x desktop app does not freeze or list it; Convert lists PyMuPDF and the
bundled Markdown pipeline. Extra
ingest plugins are ordinary pip packages in the same environment. See
vera-ingest-docling.
vera-embed-openai¶
vera-embed-openai is the official OpenAI embeddings plugin. It registers
openai under vera.embedders with stdlib urllib (no OpenAI SDK).
vera and vera-app depend on it, and the packaged sidecar freezes it
beside PyMuPDF. Secrets stay in OPENAI_API_KEY. See
vera-embed-openai.
vera¶
vera publishes the vera console script and vera_cli module (source in
packages/vera-cli). It owns argument parsing, text/JSON formatting, exit
codes, and retrieval evaluation. Conversion commands compose vera-ingest
(+ pipeline plugins) with vera-doc. vera-cli remains a PyPI compatibility
alias.
vera-mcp¶
vera-mcp publishes vera_mcp and the vera-mcp console script. It is a
thin MCP adapter over vera-doc search/storage APIs and vera-ingest.viewer
helpers. The optional vera[mcp] extra installs it for vera mcp.
vera-app¶
vera-app owns the Electron/React desktop application, Python sidecar, LLM
providers, sessions, and application state. It depends on vera-doc,
vera-ingest, vera-ingest-pymupdf, and vera-embed-openai
(including
viewer helpers), not on the CLI package. Packaged conversions keep one frozen
sidecar for search, Ask, PyMuPDF, Markdown ingest, and OpenAI embeddings. MiniLM (all-MiniLM-L6-v2) runs on
ONNX Runtime; npm run app:dev vendors that graph before launch. When no
MiniLM ONNX graph is present, MiniLM falls back to Sentence Transformers
(ml extra), which is how CLI installs run it.
Extra ingest and embedding plugins are pip packages in that same environment.
vera-lab¶
vera-lab is a contributor layout lab. It depends on vera-ingest and
PyMuPDF, runs a pipeline (or reads an archive) into a normalized view model,
and writes a self-contained HTML report with page overlays, lint, and stats.
Nothing in the release path depends on it. See
vera-lab.
Core Python API¶
Applications that already have chunks need only vera-doc:
from vera_doc import ChunkRecord, VeraDocument
with VeraDocument.create("knowledge.vera") as document:
document.add(
[
ChunkRecord(
id="requirements-1",
text="The minimum pipe diameter is 12 inches.",
metadata={"source": "manual.pdf", "page": 42},
)
]
)
with VeraDocument.open("knowledge.vera") as document:
results = document.search(text="minimum pipe size", top_k=5)
The write API accepts only final chunks and optional opaque attachments:
from vera_doc import AttachmentRecord, AttachmentRef, ChunkRecord
source = AttachmentRecord(
id="source",
media_type="application/pdf",
filename="manual.pdf",
data=pdf_bytes,
)
record = ChunkRecord(
id="chunk-1",
text="Ready-made chunk text.",
attachments=(AttachmentRef("source", role="source"),),
)
Extraction is explicitly composed:
Repository layout¶
packages/
vera-doc/
src/vera_doc/
document.py
models.py
embeddings.py
collection.py
corpus.py
vera-ingest/
src/vera_ingest/
convert.py
chunking.py
pipeline.py
descriptors.py
vera-ingest-pymupdf/
src/vera_ingest_pymupdf/
parser.py
pipeline.py
tessdata/
vera-ingest-docling/
src/vera_ingest_docling/
mapping.py
converter.py
recovery.py
pipeline.py
options.py
vera-embed-openai/
src/vera_embed_openai/
options.py
provider.py
vera-cli/
src/vera_cli/
commands.py
evaluate.py
vera-mcp/
src/vera_mcp/
server.py
vera-app/
The root uv workspace links packages as editable dependencies. Published packages use normal version constraints; source copying and Git submodules are not dependency mechanisms.
Format compatibility¶
New archives and conversions write VERA 0.2 only. The chunk-oriented current
format is documented in vera-spec-v0.2.md. The older
0.1 document/page/block schema remains in
vera-spec-v0.1.md for historical reference. That schema is
deprecated and is no longer read by vera-doc.
Repository strategy¶
Keep the packages in one GitHub repository while schema, extraction, CLI, and desktop changes commonly need atomic integration work. Separate repositories only when ownership, release cadence, governance, or deployment constraints actually diverge. Before a split, publish versioned wheels and test dependents against both minimum and latest supported package versions.