Skip to content

vera_ingest package

Provider-neutral ingest registry, shared types, chunking helpers, and conversion to .vera archives. PDF parsing/OCR live in plugins such as vera-ingest-pymupdf.

vera_ingest

Source ingestion, chunking, and conversion adapters for VERA.

Modules:

Classes:

  • Chunk –

    Text segment produced by the extraction chunker.

  • IngestBlock –

    Normalized layout block produced by an ingest pipeline.

  • IngestChunk –

    Readable chunk text and its normalized provenance.

  • IngestRequest –

    Thin shared request passed to ingest pipelines.

  • IngestOptions –

    Deprecated Tesseract-shaped options retained for compatibility.

  • IngestResult –

    Normalized bundle consumed by VERA's shared archive writer.

  • PipelineDescriptor –

    Metadata describing an installed ingest pipeline provider/variant.

  • PipelineOptions –

    Base for an ingest pipeline's typed, validated settings.

  • UnknownIngestPipelineError –

    Raised when an ingest pipeline spec cannot be resolved.

  • ParsedBlock –

    Layout block from a parser before a stable block_id is assigned.

  • ParsedPage –

    Single page extracted from a source document.

Functions:

Attributes:

  • IngestPipeline –

    A pipeline that normalizes a source document into an ingest bundle.

IngestPipeline module-attribute

IngestPipeline = Callable[
    [str, IngestRequest], IngestResult
]

A pipeline that normalizes a source document into an ingest bundle.

A pipeline is any callable matching this signature — a plain function, or an object implementing __call__ if it needs to hold state. There is no base class to inherit from::

def create_pipeline(variant: str = "") -> IngestPipeline:
    def ingest(source_path: str, options: IngestRequest) -> IngestResult:
        ...
    return ingest

For compatibility with pre-0.3.x plugins, an object exposing a callable ingest(self, source_path, options) method is also accepted; see :func:invoke_ingest_pipeline.

Chunk dataclass

Chunk(
    text: str,
    page_start: int,
    page_end: int,
    heading_path: str,
    token_count: int,
    block_ids: list[str] = list(),
)

Text segment produced by the extraction chunker.

Attributes:

  • text (str) –

    Chunk content.

  • page_start (int) –

    First page number spanned by the chunk.

  • page_end (int) –

    Last page number spanned by the chunk.

  • heading_path (str) –

    Heading breadcrumb at chunk start.

  • token_count (int) –

    Whitespace-split word count.

  • block_ids (list[str]) –

    Source layout block identifiers.

text instance-attribute

text: str

page_start instance-attribute

page_start: int

page_end instance-attribute

page_end: int

heading_path instance-attribute

heading_path: str

token_count instance-attribute

token_count: int

block_ids class-attribute instance-attribute

block_ids: list[str] = field(default_factory=list)

IngestBlock dataclass

IngestBlock(
    block_id: str,
    page_number: int,
    block_type: str,
    text: str,
    bbox: tuple[float, float, float, float] | None = None,
    heading_level: int | None = None,
    image_bytes: bytes | None = None,
    image_ext: str = "",
    regions: list[dict[str, Any]] = list(),
)

Normalized layout block produced by an ingest pipeline.

block_id must be stable for the same source and pipeline. Bounding boxes use page points with a top-left origin.

Methods:

  • from_parsed –

    Copy parser output into an archive-ready block with a stable id.

Attributes:

block_id instance-attribute

block_id: str

page_number instance-attribute

page_number: int

block_type instance-attribute

block_type: str

text instance-attribute

text: str

bbox class-attribute instance-attribute

bbox: tuple[float, float, float, float] | None = None

heading_level class-attribute instance-attribute

heading_level: int | None = None

image_bytes class-attribute instance-attribute

image_bytes: bytes | None = None

image_ext class-attribute instance-attribute

image_ext: str = ''

regions class-attribute instance-attribute

regions: list[dict[str, Any]] = field(default_factory=list)

from_parsed classmethod

from_parsed(
    block_id: str,
    block: ParsedBlock,
    *,
    regions: list[dict[str, Any]] | None = None,
) -> IngestBlock

Copy parser output into an archive-ready block with a stable id.

Use this when a parser (including parse_pdf_structured) returns :class:ParsedBlock values. Pipelines that mint IDs while parsing can construct :class:IngestBlock directly.

IngestChunk dataclass

IngestChunk(
    chunk_id: str,
    text: str,
    page_start: int,
    page_end: int,
    heading_path: str,
    token_count: int,
    block_ids: list[str] = list(),
    embedding_text: str | None = None,
    metadata: dict[str, Any] = dict(),
)

Readable chunk text and its normalized provenance.

embedding_text may provide contextualized text for embedding while leaving text readable and keyword-searchable in the archive.

Attributes:

chunk_id instance-attribute

chunk_id: str

text instance-attribute

text: str

page_start instance-attribute

page_start: int

page_end instance-attribute

page_end: int

heading_path instance-attribute

heading_path: str

token_count instance-attribute

token_count: int

block_ids class-attribute instance-attribute

block_ids: list[str] = field(default_factory=list)

embedding_text class-attribute instance-attribute

embedding_text: str | None = None

metadata class-attribute instance-attribute

metadata: dict[str, Any] = field(default_factory=dict)

IngestRequest dataclass

IngestRequest(
    variant: str = "",
    cancel: Any | None = None,
    pipeline_options: dict[str, Any] = dict(),
)

Thin shared request passed to ingest pipelines.

Provider-specific settings live in pipeline_options. Chunking and OCR defaults are owned by each pipeline, not by this shared request.

Attributes:

variant class-attribute instance-attribute

variant: str = ''

cancel class-attribute instance-attribute

cancel: Any | None = None

pipeline_options class-attribute instance-attribute

pipeline_options: dict[str, Any] = field(
    default_factory=dict
)

IngestOptions dataclass

IngestOptions(
    chunk_size: int = 500,
    overlap: int = 75,
    ocr_mode: str = "auto",
    ocr_language: str = "eng",
    ocr_dpi: int = 300,
    variant: str = "",
    cancel: Any | None = None,
    pipeline_options: dict[str, Any] = dict(),
)

Deprecated Tesseract-shaped options retained for compatibility.

Prefer :class:IngestRequest with pipeline_options. Constructing this type still works for older plugins and tests; call :meth:to_request before invoking modern pipelines.

Methods:

  • to_request –

    Convert compatibility fields into a thin :class:IngestRequest.

Attributes:

chunk_size class-attribute instance-attribute

chunk_size: int = 500

overlap class-attribute instance-attribute

overlap: int = 75

ocr_mode class-attribute instance-attribute

ocr_mode: str = 'auto'

ocr_language class-attribute instance-attribute

ocr_language: str = 'eng'

ocr_dpi class-attribute instance-attribute

ocr_dpi: int = 300

variant class-attribute instance-attribute

variant: str = ''

cancel class-attribute instance-attribute

cancel: Any | None = None

pipeline_options class-attribute instance-attribute

pipeline_options: dict[str, Any] = field(
    default_factory=dict
)

to_request

to_request() -> IngestRequest

Convert compatibility fields into a thin :class:IngestRequest.

Tesseract-shaped aliases (ocr_language, ocr_dpi) are not stuffed into pipeline_options; put them on pipeline_options (or use :func:~vera_ingest.pipeline.prepare_pipeline_options) so non-Tesseract pipelines keep their own language/DPI defaults.

IngestResult dataclass

IngestResult(
    pages: list[ParsedPage],
    blocks: list[IngestBlock],
    chunks: list[IngestChunk],
    parser_name: str,
    parser_version: str,
    chunking_strategy: str,
    diagnostics: dict[str, Any] = dict(),
)

Normalized bundle consumed by VERA's shared archive writer.

Attributes:

pages instance-attribute

pages: list[ParsedPage]

blocks instance-attribute

blocks: list[IngestBlock]

chunks instance-attribute

chunks: list[IngestChunk]

parser_name instance-attribute

parser_name: str

parser_version instance-attribute

parser_version: str

chunking_strategy instance-attribute

chunking_strategy: str

diagnostics class-attribute instance-attribute

diagnostics: dict[str, Any] = field(default_factory=dict)

PipelineDescriptor dataclass

PipelineDescriptor(
    provider: str,
    variant: str,
    spec: str,
    label: str,
    description: str = "",
    installed: bool = True,
    capabilities: PipelineCapabilities = PipelineCapabilities(),
    fields: tuple[PipelineField, ...] = (),
    notes: tuple[str, ...] = (),
)

Metadata describing an installed ingest pipeline provider/variant.

Methods:

Attributes:

provider instance-attribute

provider: str

variant instance-attribute

variant: str

spec instance-attribute

spec: str

label instance-attribute

label: str

description class-attribute instance-attribute

description: str = ''

installed class-attribute instance-attribute

installed: bool = True

capabilities class-attribute instance-attribute

capabilities: PipelineCapabilities = field(
    default_factory=PipelineCapabilities
)

fields class-attribute instance-attribute

fields: tuple[PipelineField, ...] = ()

notes class-attribute instance-attribute

notes: tuple[str, ...] = ()

field_keys

field_keys() -> set[str]

defaults

defaults() -> dict[str, Any]

as_dict

as_dict() -> dict[str, Any]

Serialize for sidecar/CLI JSON clients.

PipelineOptions

Bases: OptionsBase

Base for an ingest pipeline's typed, validated settings.

Subclass alongside @dataclass(frozen=True)::

@dataclass(frozen=True)
class MyOptions(PipelineOptions):
    chunk_size: int = field(default=2000, metadata={"label": "Chunk size"})

MyOptions.from_mapping(raw) validates a raw pipeline_options dict field by field. For each field, its own default value's type (not its static annotation, which may be a string under from __future__ import annotations) picks the validator:

  • a bool default uses :func:~vera_doc.option_parsing.require_bool;
  • an int default uses :func:~vera_doc.option_parsing.require_bounded_int with metadata["minimum"] / metadata["maximum"] when those are numbers (otherwise the value must be non-negative);
  • a str default with metadata["choices"] and no metadata["allow_custom"] uses :func:~vera_doc.option_parsing.require_choice restricted to those choices' values;
  • any other str default uses :func:~vera_doc.option_parsing.require_string (free text).

A field of any other type (for example float) is not supported; override from_mapping for a class with such a field instead.

Two class attributes customize behavior without any of that:

  • options_label sets the name used in error messages (default: the class name with a trailing Options dropped, so PyMuPDFOptions reads as "PyMuPDF").
  • ignored_keys names legacy pipeline_options keys to silently accept and drop instead of rejecting as unknown — for compatibility aliases shared with another pipeline that don't apply to this one.

Methods:

Attributes:

options_label class-attribute

options_label: str = ''

ignored_keys class-attribute

ignored_keys: frozenset[str] = frozenset()

enforce_step class-attribute

enforce_step: bool = False

options_mapping_label class-attribute

options_mapping_label: str = 'pipeline_options'

from_mapping classmethod

from_mapping(raw: Mapping[str, Any] | None = None) -> Any

UnknownIngestPipelineError

Bases: ValueError

Raised when an ingest pipeline spec cannot be resolved.

ParsedBlock dataclass

ParsedBlock(
    page_number: int,
    block_type: str,
    text: str,
    bbox: tuple[float, float, float, float] | None = None,
    heading_level: int | None = None,
    image_bytes: bytes | None = None,
    image_ext: str = "",
)

Layout block from a parser before a stable block_id is assigned.

First-party pipelines (and most custom parsers) emit these, then convert with :meth:IngestBlock.from_parsed before returning :class:IngestResult. Pipelines that mint IDs while parsing can construct :class:IngestBlock directly.

Attributes:

page_number instance-attribute

page_number: int

block_type instance-attribute

block_type: str

text instance-attribute

text: str

bbox class-attribute instance-attribute

bbox: tuple[float, float, float, float] | None = None

heading_level class-attribute instance-attribute

heading_level: int | None = None

image_bytes class-attribute instance-attribute

image_bytes: bytes | None = None

image_ext class-attribute instance-attribute

image_ext: str = ''

ParsedPage dataclass

ParsedPage(
    page_number: int,
    width: float | None,
    height: float | None,
    text: str,
)

Single page extracted from a source document.

Attributes:

page_number instance-attribute

page_number: int

width instance-attribute

width: float | None

height instance-attribute

height: float | None

text instance-attribute

text: str

batch_convert

batch_convert(
    directory: str | None = None,
    *,
    paths: Sequence[str] | None = None,
    recursive: bool = False,
    overwrite: bool = False,
    model: str = "hashing",
    embedding_function: EmbeddingFunction | None = None,
    parser: str | None = None,
    chunk_size: int | None = None,
    overlap: int | None = None,
    store_original: bool = True,
    ocr_mode: str | None = None,
    ocr_language: str | None = None,
    ocr_dpi: int | None = None,
    ocr_download: bool | None = None,
    pipeline_options: dict[str, Any] | None = None,
    embedder_options: dict[str, Any] | None = None,
    progress: Callable[[int, int, str], None] | None = None,
    cancel: Any | None = None,
    metadata: Mapping[str, Any] | None = None,
) -> dict[str, Any]

Convert source files from a directory scan or an explicit path list.

Parameters:

  • directory (str | None, default: None ) –

    Root directory to scan when paths is omitted. Discovery uses the selected pipeline's source_formats, or every installed pipeline when parser is omitted.

  • paths (Sequence[str] | None, default: None ) –

    Explicit source file paths to convert. When set, directory discovery is skipped and recursive is ignored.

  • recursive (bool, default: False ) –

    When True, scan subdirectories (directory mode only).

  • overwrite (bool, default: False ) –

    When True, replace existing .vera outputs. When False, skip a sibling archive only when it validates and its stored source_file_hash matches the current source file. Stale or hash-less archives are reconverted. Same-stem sources that would write the same .vera path fail instead of overwriting each other, including when overwrite is true.

  • model (str, default: 'hashing' ) –

    Embedding model spec passed to :func:convert.

  • embedding_function (EmbeddingFunction | None, default: None ) –

    Optional custom embedder passed to :func:convert.

  • parser (str | None, default: None ) –

    Ingest pipeline spec passed to :func:convert. None (the default) selects a pipeline per file from its extension.

  • chunk_size (int | None, default: None ) –

    Compatibility alias passed to :func:convert. None means the pipeline default.

  • overlap (int | None, default: None ) –

    Compatibility alias passed to :func:convert. None means the pipeline default.

  • store_original (bool, default: True ) –

    Whether to embed originals passed to :func:convert.

  • ocr_mode (str | None, default: None ) –

    Compatibility OCR mode alias passed to :func:convert. None means the pipeline default.

  • ocr_language (str | None, default: None ) –

    Compatibility OCR language alias passed to :func:convert. None means the pipeline default.

  • ocr_dpi (int | None, default: None ) –

    Compatibility OCR DPI alias passed to :func:convert. None means the pipeline default.

  • ocr_download (bool | None, default: None ) –

    Compatibility OCR download alias passed to :func:convert. None means the pipeline default.

  • pipeline_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned options passed to :func:convert.

  • embedder_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned embedding options passed to :func:convert.

  • progress (Callable[[int, int, str], None] | None, default: None ) –

    Optional (current, total, filename) callback.

  • cancel (Any | None, default: None ) –

    Optional cancellation token.

  • metadata (Mapping[str, Any] | None, default: None ) –

    Extra keys stamped onto every archive and chunk in this run.

Returns:

  • dict[str, Any] –

    A report dict with converted, skipped, failed, and related

  • dict[str, Any] –

    fields. directory is the scan root, or the common parent of

  • dict[str, Any] –

    paths.

Raises:

  • NotADirectoryError –

    When directory is not a directory.

  • FileNotFoundError –

    When a path in paths is missing.

  • ValueError –

    When neither directory nor paths is usable.

  • ReservedMetadataKeyError –

    When metadata uses a reserved key.

  • UnknownEmbeddingModelError –

    When model cannot be resolved.

describe_ingest_pipeline

describe_ingest_pipeline(
    spec: str = "pymupdf",
) -> PipelineDescriptor

Return metadata for an installed pipeline without instantiating it.

get_ingest_pipeline

get_ingest_pipeline(
    spec: str = "pymupdf",
) -> IngestPipeline

Resolve and cache an installed pipeline.

Resolution is strict: an unknown provider or variant raises instead of falling back to another installed pipeline.

invoke_ingest_pipeline

invoke_ingest_pipeline(
    pipeline: IngestPipeline,
    source_path: str,
    request: IngestRequest,
) -> IngestResult

Call pipeline, accepting both a bare callable and a legacy .ingest() object.

list_ingest_pipelines

list_ingest_pipelines() -> list[str]

Return sorted installed provider names.

list_ingest_pipeline_descriptors

list_ingest_pipeline_descriptors() -> list[
    PipelineDescriptor
]

Return descriptors for each installed provider using its default variant.

prepare_pipeline_options

prepare_pipeline_options(
    *,
    spec: str,
    pipeline_options: dict[str, Any] | None = None,
    legacy_options: dict[str, Any] | None = None,
) -> dict[str, Any]

Merge legacy convert kwargs with explicit pipeline options.

When a pipeline publishes descriptor fields, only those legacy keys are forwarded so PyMuPDF defaults such as overlap and ocr_dpi do not leak into plugins that omit them. Tesseract-shaped aliases (ocr_language, ocr_dpi, ocr_download) are forwarded only when capabilities.ocr_engine is "tesseract", so Docling/RapidOCR keeps its own ocr_language default instead of inheriting eng. Undescribed plugins receive the remaining compatibility bag. Explicit pipeline_options always win.

register_ingest_pipeline

register_ingest_pipeline(
    provider: str,
    factory: Callable[[str], IngestPipeline] | None = None,
    *,
    replace: bool = False,
) -> (
    Callable[
        [Callable[[str], IngestPipeline]],
        Callable[[str], IngestPipeline],
    ]
    | None
)

Register a provider factory called with the requested variant.

Called with both arguments, this registers factory immediately and returns None, as before. Omit factory to use it as a decorator instead — handy for local experiments, notebooks, and tests that would otherwise need a separate factory function and a separate call::

@register_ingest_pipeline("myexperiment")
def create_pipeline(variant: str = "") -> IngestPipeline:
    return MyPipeline()

register_ingest_pipeline_descriptor

register_ingest_pipeline_descriptor(
    provider: str,
    factory: Callable[[str], PipelineDescriptor]
    | None = None,
    *,
    replace: bool = False,
) -> (
    Callable[
        [Callable[[str], PipelineDescriptor]],
        Callable[[str], PipelineDescriptor],
    ]
    | None
)

Register a descriptor factory called with the requested variant.

Also usable as a decorator when factory is omitted — see :func:register_ingest_pipeline.

chunk_pages

chunk_pages(
    pages: list[ParsedPage],
    chunk_size: int = 500,
    overlap: int = 75,
) -> list[Chunk]

Sliding-window chunker over ParsedPage.text (whitespace-split words).

Public helper for custom pipelines that only have page text. First-party pipelines (PyMuPDF, Docling) do not call this; they chunk structured layout with :func:build_chunks_from_blocks or their own chunker.

build_chunks_from_blocks

build_chunks_from_blocks(
    blocks: list[tuple[str, ParsedBlock | IngestBlock]],
    chunk_size: int = 500,
    overlap: int = 75,
) -> list[Chunk]

Heading-aware chunking over structured layout blocks.

Accepts (block_id, block) pairs of either :class:ParsedBlock or :class:IngestBlock. First-party PyMuPDF uses this path; custom pipelines that only have page text should use :func:chunk_pages instead.

detect_heading

detect_heading(text: str, current: str) -> str

Return the first short heading-like line in text, else current.

Public helper for custom pipelines that chunk page text with :func:chunk_pages. First-party pipelines use structured heading blocks via :func:build_chunks_from_blocks instead.

figures

figures(
    document: VeraDocument,
    page_start: int | None = None,
    page_end: int | None = None,
    include_data: bool = False,
    attachment_ids: Iterable[str] | None = None,
) -> list[dict[str, Any]]

Return figure attachments produced during ingest.

figures_for

figures_for(
    document: VeraDocument,
    result: QueryResult,
    include_data: bool = False,
) -> list[dict[str, Any]]

Return figure attachments linked to a query result.

get_page

get_page(
    document: VeraDocument, page_number: int
) -> dict[str, Any] | None

Return ingest-provided viewer data for one page.

get_blocks

get_blocks(
    document: VeraDocument, page_number: int | None = None
) -> list[dict[str, Any]]

Return ingest-provided layout blocks.

get_chunk_regions

get_chunk_regions(
    document: VeraDocument, chunk_id: str
) -> list[dict[str, Any]]

Return ingest-provided highlight regions for a chunk.

regions_for

regions_for(
    document: VeraDocument, result: QueryResult
) -> list[dict[str, Any]]

Return ingest-provided highlight regions for a query result.

chunk_payload

chunk_payload(
    record: ChunkRecord,
    *,
    document: VeraDocument | None = None,
    include_figures: bool = False,
    include_regions: bool = False,
    include_figure_data: bool = False,
    figure_data_urls: bool = False,
) -> dict[str, Any]

Flatten a stored chunk for CLI/MCP JSON (metadata keys at the top level).

Citation fields such as page_start and heading_path sit beside chunk_id and text. The embedding vector and retrieval scores are omitted. Optional figure and region enrichment matches :func:result_payload.

result_payload

result_payload(
    result: QueryResult,
    *,
    document: VeraDocument | None = None,
    include_figures: bool = False,
    include_regions: bool = False,
    include_figure_data: bool = False,
    figure_data_urls: bool = False,
) -> dict[str, Any]

Flatten a search hit for CLI/MCP/app JSON (metadata keys at the top level).

Optional figure and region enrichment uses ingest viewer helpers. Sidecar callers can set figure_data_urls to replace raw figure bytes with a data_url instead of forking the serializer. Extra as_dict() keys such as corpus file are preserved.

get_source_document

get_source_document(
    document: VeraDocument,
) -> AttachmentRecord

Return the attachment identified as the archive's source document.

export_figures

export_figures(
    document: VeraDocument,
    directory: str | PathLike[str],
    *,
    asset_ids: Iterable[str] | None = None,
    page_start: int | None = None,
    page_end: int | None = None,
) -> list[dict[str, Any]]

Write figure attachments under directory and return metadata plus paths.

Output names are {asset_id}.{ext}. ext comes from the stored mime type or filename. Requested ids that are missing or not figure attachments raise ValueError so a source PDF id cannot leak. Raw data is never included in the returned dicts.

export_source_document

export_source_document(
    document: VeraDocument,
    path: str | PathLike[str] | None = None,
) -> str

Write the source attachment to disk and return its path.

The stored filename is used as Path(...).name only. Absolute names and .. segments are rejected. When path is omitted the file is written under the current working directory; when path is a directory the file stays under that directory. An explicit file path is the caller's chosen output location.

Conversion writes through VeraDocument. Viewer helpers interpret ingest-produced attachments and metadata. Shared convert accepts opaque pipeline_options on a thin IngestRequest; pipelines own typed defaults and descriptors. Prefer IngestRequest / pipeline_options over the deprecated IngestOptions compatibility bag. See the conversion guide and figures and regions.

convert

Classes:

Functions:

  • convert –

    Convert a source document into a validated .vera archive.

  • batch_convert –

    Convert source files from a directory scan or an explicit path list.

ReservedMetadataKeyError

Bases: ValueError

Caller metadata collided with a reserved convert or format key.

convert

convert(
    input_path: str,
    output_path: str,
    *,
    model: str = "hashing",
    embedding_function: EmbeddingFunction | None = None,
    parser: str | None = None,
    chunk_size: int | None = None,
    overlap: int | None = None,
    store_original: bool = True,
    ocr_mode: str | None = None,
    ocr_language: str | None = None,
    ocr_dpi: int | None = None,
    ocr_download: bool | None = None,
    pipeline_options: dict[str, Any] | None = None,
    embedder_options: dict[str, Any] | None = None,
    cancel: Any | None = None,
    metadata: Mapping[str, Any] | None = None,
) -> str

Convert a source document into a validated .vera archive.

Parses the file, chunks extracted text, embeds chunks, and writes the result through :class:~vera_doc.document.VeraDocument. The archive is validated before the temporary file is published atomically.

New callers should pass parser, pipeline_options, and embedder settings (model / embedding_function / embedder_options). chunk_size, overlap, ocr_mode, ocr_language, ocr_dpi, and ocr_download remain compatibility aliases for CLI and sidecar callers; they are forwarded only when explicitly provided and the selected pipeline advertises them (Tesseract OCR aliases do not leak to Docling). Omitted aliases mean "use the pipeline's own default" (for example a plugin chunk_size of 2000 is not overwritten by 500). The CLI still passes its argparse defaults when invoked from the command line.

Parameters:

  • input_path (str) –

    Source document path (PDF, Markdown, or another format advertised by an installed ingest pipeline).

  • output_path (str) –

    Destination .vera path.

  • model (str, default: 'hashing' ) –

    Embedding model spec (default "hashing"). Ignored when embedding_function is provided. Accepts provider:model-id or built-in legacy aliases.

  • embedding_function (EmbeddingFunction | None, default: None ) –

    Optional custom embedder satisfying :class:~vera_doc.EmbeddingFunction. When omitted, model is resolved via :func:~vera_doc.get_embedder before parsing begins.

  • parser (str | None, default: None ) –

    Ingest pipeline spec in provider[:variant] form. None (the default) selects an installed pipeline from the file extension. An explicit spec must advertise that extension.

  • chunk_size (int | None, default: None ) –

    Compatibility alias forwarded only when explicitly provided and the selected pipeline advertises a chunk_size field. None (the default) means the pipeline default.

  • overlap (int | None, default: None ) –

    Compatibility alias forwarded only when explicitly provided and advertised by the selected pipeline (PyMuPDF, Markdown). Ignored by Docling. None means the pipeline default.

  • store_original (bool, default: True ) –

    When True, embed the original file as an attachment.

  • ocr_mode (str | None, default: None ) –

    Compatibility OCR mode alias when explicitly provided and advertised by the pipeline. None means the pipeline default.

  • ocr_language (str | None, default: None ) –

    Tesseract OCR language alias (PyMuPDF). Forwarded only when explicitly provided and the selected pipeline's ocr_engine is "tesseract". None means the pipeline default.

  • ocr_dpi (int | None, default: None ) –

    Compatibility OCR DPI alias when explicitly provided and advertised (PyMuPDF). None means the pipeline default.

  • ocr_download (bool | None, default: None ) –

    Compatibility alias (PyMuPDF only) allowing on-demand, checksum-verified download of missing Tesseract language data. None means the pipeline default.

  • pipeline_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned options. These override compatibility aliases for the same keys.

  • embedder_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned embedding options forwarded to :func:~vera_doc.get_embedder when embedding_function is omitted.

  • cancel (Any | None, default: None ) –

    Optional cancellation token with raise_if_cancelled().

  • metadata (Mapping[str, Any] | None, default: None ) –

    Extra keys stamped onto archive metadata and every chunk. Reserved ingest, citation, and format keys are rejected.

Returns:

  • str –

    The output_path string.

Raises:

  • FileNotFoundError –

    When input_path does not exist.

  • ValueError –

    When no searchable text is extracted, or parser does not support the source file type.

  • ReservedMetadataKeyError –

    When metadata uses a reserved key.

  • UnknownIngestPipelineError –

    When parser cannot be resolved.

  • UnknownEmbeddingModelError –

    When model cannot be resolved.

batch_convert

batch_convert(
    directory: str | None = None,
    *,
    paths: Sequence[str] | None = None,
    recursive: bool = False,
    overwrite: bool = False,
    model: str = "hashing",
    embedding_function: EmbeddingFunction | None = None,
    parser: str | None = None,
    chunk_size: int | None = None,
    overlap: int | None = None,
    store_original: bool = True,
    ocr_mode: str | None = None,
    ocr_language: str | None = None,
    ocr_dpi: int | None = None,
    ocr_download: bool | None = None,
    pipeline_options: dict[str, Any] | None = None,
    embedder_options: dict[str, Any] | None = None,
    progress: Callable[[int, int, str], None] | None = None,
    cancel: Any | None = None,
    metadata: Mapping[str, Any] | None = None,
) -> dict[str, Any]

Convert source files from a directory scan or an explicit path list.

Parameters:

  • directory (str | None, default: None ) –

    Root directory to scan when paths is omitted. Discovery uses the selected pipeline's source_formats, or every installed pipeline when parser is omitted.

  • paths (Sequence[str] | None, default: None ) –

    Explicit source file paths to convert. When set, directory discovery is skipped and recursive is ignored.

  • recursive (bool, default: False ) –

    When True, scan subdirectories (directory mode only).

  • overwrite (bool, default: False ) –

    When True, replace existing .vera outputs. When False, skip a sibling archive only when it validates and its stored source_file_hash matches the current source file. Stale or hash-less archives are reconverted. Same-stem sources that would write the same .vera path fail instead of overwriting each other, including when overwrite is true.

  • model (str, default: 'hashing' ) –

    Embedding model spec passed to :func:convert.

  • embedding_function (EmbeddingFunction | None, default: None ) –

    Optional custom embedder passed to :func:convert.

  • parser (str | None, default: None ) –

    Ingest pipeline spec passed to :func:convert. None (the default) selects a pipeline per file from its extension.

  • chunk_size (int | None, default: None ) –

    Compatibility alias passed to :func:convert. None means the pipeline default.

  • overlap (int | None, default: None ) –

    Compatibility alias passed to :func:convert. None means the pipeline default.

  • store_original (bool, default: True ) –

    Whether to embed originals passed to :func:convert.

  • ocr_mode (str | None, default: None ) –

    Compatibility OCR mode alias passed to :func:convert. None means the pipeline default.

  • ocr_language (str | None, default: None ) –

    Compatibility OCR language alias passed to :func:convert. None means the pipeline default.

  • ocr_dpi (int | None, default: None ) –

    Compatibility OCR DPI alias passed to :func:convert. None means the pipeline default.

  • ocr_download (bool | None, default: None ) –

    Compatibility OCR download alias passed to :func:convert. None means the pipeline default.

  • pipeline_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned options passed to :func:convert.

  • embedder_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned embedding options passed to :func:convert.

  • progress (Callable[[int, int, str], None] | None, default: None ) –

    Optional (current, total, filename) callback.

  • cancel (Any | None, default: None ) –

    Optional cancellation token.

  • metadata (Mapping[str, Any] | None, default: None ) –

    Extra keys stamped onto every archive and chunk in this run.

Returns:

  • dict[str, Any] –

    A report dict with converted, skipped, failed, and related

  • dict[str, Any] –

    fields. directory is the scan root, or the common parent of

  • dict[str, Any] –

    paths.

Raises:

  • NotADirectoryError –

    When directory is not a directory.

  • FileNotFoundError –

    When a path in paths is missing.

  • ValueError –

    When neither directory nor paths is usable.

  • ReservedMetadataKeyError –

    When metadata uses a reserved key.

  • UnknownEmbeddingModelError –

    When model cannot be resolved.

batch_convert

batch_convert(
    directory: str | None = None,
    *,
    paths: Sequence[str] | None = None,
    recursive: bool = False,
    overwrite: bool = False,
    model: str = "hashing",
    embedding_function: EmbeddingFunction | None = None,
    parser: str | None = None,
    chunk_size: int | None = None,
    overlap: int | None = None,
    store_original: bool = True,
    ocr_mode: str | None = None,
    ocr_language: str | None = None,
    ocr_dpi: int | None = None,
    ocr_download: bool | None = None,
    pipeline_options: dict[str, Any] | None = None,
    embedder_options: dict[str, Any] | None = None,
    progress: Callable[[int, int, str], None] | None = None,
    cancel: Any | None = None,
    metadata: Mapping[str, Any] | None = None,
) -> dict[str, Any]

Convert source files from a directory scan or an explicit path list.

Parameters:

  • directory (str | None, default: None ) –

    Root directory to scan when paths is omitted. Discovery uses the selected pipeline's source_formats, or every installed pipeline when parser is omitted.

  • paths (Sequence[str] | None, default: None ) –

    Explicit source file paths to convert. When set, directory discovery is skipped and recursive is ignored.

  • recursive (bool, default: False ) –

    When True, scan subdirectories (directory mode only).

  • overwrite (bool, default: False ) –

    When True, replace existing .vera outputs. When False, skip a sibling archive only when it validates and its stored source_file_hash matches the current source file. Stale or hash-less archives are reconverted. Same-stem sources that would write the same .vera path fail instead of overwriting each other, including when overwrite is true.

  • model (str, default: 'hashing' ) –

    Embedding model spec passed to :func:convert.

  • embedding_function (EmbeddingFunction | None, default: None ) –

    Optional custom embedder passed to :func:convert.

  • parser (str | None, default: None ) –

    Ingest pipeline spec passed to :func:convert. None (the default) selects a pipeline per file from its extension.

  • chunk_size (int | None, default: None ) –

    Compatibility alias passed to :func:convert. None means the pipeline default.

  • overlap (int | None, default: None ) –

    Compatibility alias passed to :func:convert. None means the pipeline default.

  • store_original (bool, default: True ) –

    Whether to embed originals passed to :func:convert.

  • ocr_mode (str | None, default: None ) –

    Compatibility OCR mode alias passed to :func:convert. None means the pipeline default.

  • ocr_language (str | None, default: None ) –

    Compatibility OCR language alias passed to :func:convert. None means the pipeline default.

  • ocr_dpi (int | None, default: None ) –

    Compatibility OCR DPI alias passed to :func:convert. None means the pipeline default.

  • ocr_download (bool | None, default: None ) –

    Compatibility OCR download alias passed to :func:convert. None means the pipeline default.

  • pipeline_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned options passed to :func:convert.

  • embedder_options (dict[str, Any] | None, default: None ) –

    Explicit provider-owned embedding options passed to :func:convert.

  • progress (Callable[[int, int, str], None] | None, default: None ) –

    Optional (current, total, filename) callback.

  • cancel (Any | None, default: None ) –

    Optional cancellation token.

  • metadata (Mapping[str, Any] | None, default: None ) –

    Extra keys stamped onto every archive and chunk in this run.

Returns:

  • dict[str, Any] –

    A report dict with converted, skipped, failed, and related

  • dict[str, Any] –

    fields. directory is the scan root, or the common parent of

  • dict[str, Any] –

    paths.

Raises:

  • NotADirectoryError –

    When directory is not a directory.

  • FileNotFoundError –

    When a path in paths is missing.

  • ValueError –

    When neither directory nor paths is usable.

  • ReservedMetadataKeyError –

    When metadata uses a reserved key.

  • UnknownEmbeddingModelError –

    When model cannot be resolved.