vera-ingest¶
vera-ingest publishes the vera_ingest Python package. It depends on
vera-doc and owns the ingest-pipeline registry, shared descriptors and
types, conversion orchestration, reusable chunking helpers, and
ingest-produced viewer helpers (result_payload, format_search_context,
figure and region readers).
PDF parsing and OCR live in provider plugins that register through the
vera.ingest_pipelines entry-point group. The default
vera-ingest-pymupdf package provides the pymupdf
pipeline; Markdown ingest ships inside vera-ingest as provider markdown;
vera-ingest-docling provides Docling's
hybrid chunker (PDF plus search-only DOCX/PPTX/XLSX/HTML). Extra plugins are pip packages in the same environment as
the CLI or desktop sidecar.
Pipelines return a normalized IngestResult. Shared convert() writes
validated archives through one atomic path and emits ready-made ChunkRecord
values plus optional attachments via VeraDocument. Viewer helpers under
vera_ingest.viewer read those ingest conventions back out for CLI, MCP, and
app consumers. format_search_context is the shared renderer behind CLI
--pretty and MCP pretty: true; import it from vera_ingest.viewer (it is
not in vera_ingest.__all__).
Install¶
From PyPI:
For PDF conversion, also install a pipeline plugin (vera and vera-app
depend on vera-ingest-pymupdf by default):
From a repository checkout:
Python 3.10 or newer is required. Core vera-ingest does not pull in PyMuPDF
or pdfplumber; those arrive with the pipeline plugin.
Start here¶
from vera_ingest import convert
convert(
"manual.pdf",
"manual.vera",
parser="pymupdf",
pipeline_options={"ocr_mode": "auto"},
model="hashing",
store_original=True,
)
convert("notes.md", "notes.vera", model="hashing")
Pass embedding_function= for a custom embedder, or use a
provider:model-id model spec resolved by vera_doc.get_embedder. New callers
should pass parser, pipeline_options, and embedder settings
(model / embedding_function / embedder_options); legacy
kwargs such as chunk_size and ocr_mode remain compatibility aliases
forwarded only when explicitly provided. Omitted aliases mean the pipeline's
own default.
Concepts¶
- Pipeline registry discovers installed providers via entry points
(
vera.ingest_pipelines) or in-processregister_ingest_pipeline(). Specs resolve asprovider[:variant]with no silent fallback. Broken entry points are logged and collected bylist_ingest_pipeline_load_errors()instead of hiding other providers. Registry and descriptor APIs are experimental and may change before 1.0. - Pipeline-owned config keeps typed defaults, validation, and field
descriptors inside each ingest plugin; shared convert passes a thin
IngestRequestwith opaquepipeline_options. - Chunking helpers remain available for providers that want sliding-window
behavior (whitespace-split words, not characters) without owning the writer.
First-party pipelines use
build_chunks_from_blocks(or their own chunker).chunk_pagesanddetect_headingstay public for custom pipelines that only have page text. Parsers emitParsedBlock; convert withIngestBlock.from_parsedbefore returningIngestResult.blocks. - Atomic conversion validates a temporary archive before publishing it.
Documentation¶
- Convert documents — OCR, chunking, embedding, and batch conversion.
- Creating an ingest pipeline plugin — write and register a new pipeline provider.
- Additional source formats and visual grounding — Markdown ingest, plugin naming, and Markdown/PDF viewer surfaces.
- Figures and regions — extracted visual metadata and schema storage map.
- Conversion recipes — single files, scans, and nested libraries.
- Python conversion example.
API reference¶
vera_ingest— curated public conversion, ingest-pipeline registry, page/block types, and chunking interfaces.