vera_ingest package¶
Provider-neutral ingest registry, shared types, chunking helpers, and
conversion to .vera archives. PDF parsing/OCR live in plugins such as
vera-ingest-pymupdf.
vera_ingest
¶
Source ingestion, chunking, and conversion adapters for VERA.
Modules:
-
convert–
Classes:
-
Chunk–Text segment produced by the extraction chunker.
-
IngestBlock–Normalized layout block produced by an ingest pipeline.
-
IngestChunk–Readable chunk text and its normalized provenance.
-
IngestRequest–Thin shared request passed to ingest pipelines.
-
IngestOptions–Deprecated Tesseract-shaped options retained for compatibility.
-
IngestResult–Normalized bundle consumed by VERA's shared archive writer.
-
PipelineDescriptor–Metadata describing an installed ingest pipeline provider/variant.
-
PipelineOptions–Base for an ingest pipeline's typed, validated settings.
-
UnknownIngestPipelineError–Raised when an ingest pipeline spec cannot be resolved.
-
ParsedBlock–Layout block from a parser before a stable
block_idis assigned. -
ParsedPage–Single page extracted from a source document.
Functions:
-
batch_convert–Convert source files from a directory scan or an explicit path list.
-
describe_ingest_pipeline–Return metadata for an installed pipeline without instantiating it.
-
get_ingest_pipeline–Resolve and cache an installed pipeline.
-
invoke_ingest_pipeline–Call
pipeline, accepting both a bare callable and a legacy.ingest()object. -
list_ingest_pipelines–Return sorted installed provider names.
-
list_ingest_pipeline_descriptors–Return descriptors for each installed provider using its default variant.
-
prepare_pipeline_options–Merge legacy convert kwargs with explicit pipeline options.
-
register_ingest_pipeline–Register a provider factory called with the requested variant.
-
register_ingest_pipeline_descriptor–Register a descriptor factory called with the requested variant.
-
chunk_pages–Sliding-window chunker over
ParsedPage.text(whitespace-split words). -
build_chunks_from_blocks–Heading-aware chunking over structured layout blocks.
-
detect_heading–Return the first short heading-like line in
text, elsecurrent. -
figures–Return figure attachments produced during ingest.
-
figures_for–Return figure attachments linked to a query result.
-
get_page–Return ingest-provided viewer data for one page.
-
get_blocks–Return ingest-provided layout blocks.
-
get_chunk_regions–Return ingest-provided highlight regions for a chunk.
-
regions_for–Return ingest-provided highlight regions for a query result.
-
chunk_payload–Flatten a stored chunk for CLI/MCP JSON (metadata keys at the top level).
-
result_payload–Flatten a search hit for CLI/MCP/app JSON (metadata keys at the top level).
-
get_source_document–Return the attachment identified as the archive's source document.
-
export_figures–Write figure attachments under
directoryand return metadata plus paths. -
export_source_document–Write the source attachment to disk and return its path.
Attributes:
-
IngestPipeline–A pipeline that normalizes a source document into an ingest bundle.
IngestPipeline
module-attribute
¶
IngestPipeline = Callable[
[str, IngestRequest], IngestResult
]
A pipeline that normalizes a source document into an ingest bundle.
A pipeline is any callable matching this signature — a plain function, or an
object implementing __call__ if it needs to hold state. There is no base
class to inherit from::
def create_pipeline(variant: str = "") -> IngestPipeline:
def ingest(source_path: str, options: IngestRequest) -> IngestResult:
...
return ingest
For compatibility with pre-0.3.x plugins, an object exposing a callable
ingest(self, source_path, options) method is also accepted; see
:func:invoke_ingest_pipeline.
Chunk
dataclass
¶
Chunk(
text: str,
page_start: int,
page_end: int,
heading_path: str,
token_count: int,
block_ids: list[str] = list(),
)
Text segment produced by the extraction chunker.
Attributes:
-
text(str) –Chunk content.
-
page_start(int) –First page number spanned by the chunk.
-
page_end(int) –Last page number spanned by the chunk.
-
heading_path(str) –Heading breadcrumb at chunk start.
-
token_count(int) –Whitespace-split word count.
-
block_ids(list[str]) –Source layout block identifiers.
IngestBlock
dataclass
¶
IngestBlock(
block_id: str,
page_number: int,
block_type: str,
text: str,
bbox: tuple[float, float, float, float] | None = None,
heading_level: int | None = None,
image_bytes: bytes | None = None,
image_ext: str = "",
regions: list[dict[str, Any]] = list(),
)
Normalized layout block produced by an ingest pipeline.
block_id must be stable for the same source and pipeline. Bounding boxes
use page points with a top-left origin.
Methods:
-
from_parsed–Copy parser output into an archive-ready block with a stable id.
Attributes:
-
block_id(str) – -
page_number(int) – -
block_type(str) – -
text(str) – -
bbox(tuple[float, float, float, float] | None) – -
heading_level(int | None) – -
image_bytes(bytes | None) – -
image_ext(str) – -
regions(list[dict[str, Any]]) –
regions
class-attribute
instance-attribute
¶
from_parsed
classmethod
¶
from_parsed(
block_id: str,
block: ParsedBlock,
*,
regions: list[dict[str, Any]] | None = None,
) -> IngestBlock
Copy parser output into an archive-ready block with a stable id.
Use this when a parser (including parse_pdf_structured) returns
:class:ParsedBlock values. Pipelines that mint IDs while parsing can
construct :class:IngestBlock directly.
IngestChunk
dataclass
¶
IngestChunk(
chunk_id: str,
text: str,
page_start: int,
page_end: int,
heading_path: str,
token_count: int,
block_ids: list[str] = list(),
embedding_text: str | None = None,
metadata: dict[str, Any] = dict(),
)
Readable chunk text and its normalized provenance.
embedding_text may provide contextualized text for embedding while
leaving text readable and keyword-searchable in the archive.
Attributes:
-
chunk_id(str) – -
text(str) – -
page_start(int) – -
page_end(int) – -
heading_path(str) – -
token_count(int) – -
block_ids(list[str]) – -
embedding_text(str | None) – -
metadata(dict[str, Any]) –
metadata
class-attribute
instance-attribute
¶
IngestRequest
dataclass
¶
IngestRequest(
variant: str = "",
cancel: Any | None = None,
pipeline_options: dict[str, Any] = dict(),
)
Thin shared request passed to ingest pipelines.
Provider-specific settings live in pipeline_options. Chunking and OCR
defaults are owned by each pipeline, not by this shared request.
Attributes:
-
variant(str) – -
cancel(Any | None) – -
pipeline_options(dict[str, Any]) –
IngestOptions
dataclass
¶
IngestOptions(
chunk_size: int = 500,
overlap: int = 75,
ocr_mode: str = "auto",
ocr_language: str = "eng",
ocr_dpi: int = 300,
variant: str = "",
cancel: Any | None = None,
pipeline_options: dict[str, Any] = dict(),
)
Deprecated Tesseract-shaped options retained for compatibility.
Prefer :class:IngestRequest with pipeline_options. Constructing this
type still works for older plugins and tests; call :meth:to_request before
invoking modern pipelines.
Methods:
-
to_request–Convert compatibility fields into a thin :class:
IngestRequest.
Attributes:
-
chunk_size(int) – -
overlap(int) – -
ocr_mode(str) – -
ocr_language(str) – -
ocr_dpi(int) – -
variant(str) – -
cancel(Any | None) – -
pipeline_options(dict[str, Any]) –
pipeline_options
class-attribute
instance-attribute
¶
to_request
¶
to_request() -> IngestRequest
Convert compatibility fields into a thin :class:IngestRequest.
Tesseract-shaped aliases (ocr_language, ocr_dpi) are not
stuffed into pipeline_options; put them on pipeline_options
(or use :func:~vera_ingest.pipeline.prepare_pipeline_options) so
non-Tesseract pipelines keep their own language/DPI defaults.
IngestResult
dataclass
¶
IngestResult(
pages: list[ParsedPage],
blocks: list[IngestBlock],
chunks: list[IngestChunk],
parser_name: str,
parser_version: str,
chunking_strategy: str,
diagnostics: dict[str, Any] = dict(),
)
Normalized bundle consumed by VERA's shared archive writer.
Attributes:
-
pages(list[ParsedPage]) – -
blocks(list[IngestBlock]) – -
chunks(list[IngestChunk]) – -
parser_name(str) – -
parser_version(str) – -
chunking_strategy(str) – -
diagnostics(dict[str, Any]) –
diagnostics
class-attribute
instance-attribute
¶
PipelineDescriptor
dataclass
¶
PipelineDescriptor(
provider: str,
variant: str,
spec: str,
label: str,
description: str = "",
installed: bool = True,
capabilities: PipelineCapabilities = PipelineCapabilities(),
fields: tuple[PipelineField, ...] = (),
notes: tuple[str, ...] = (),
)
Metadata describing an installed ingest pipeline provider/variant.
Methods:
-
field_keys– -
defaults– -
as_dict–Serialize for sidecar/CLI JSON clients.
Attributes:
-
provider(str) – -
variant(str) – -
spec(str) – -
label(str) – -
description(str) – -
installed(bool) – -
capabilities(PipelineCapabilities) – -
fields(tuple[PipelineField, ...]) – -
notes(tuple[str, ...]) –
capabilities
class-attribute
instance-attribute
¶
PipelineOptions
¶
Bases: OptionsBase
Base for an ingest pipeline's typed, validated settings.
Subclass alongside @dataclass(frozen=True)::
@dataclass(frozen=True)
class MyOptions(PipelineOptions):
chunk_size: int = field(default=2000, metadata={"label": "Chunk size"})
MyOptions.from_mapping(raw) validates a raw pipeline_options dict
field by field. For each field, its own default value's type (not its
static annotation, which may be a string under from __future__ import
annotations) picks the validator:
- a
booldefault uses :func:~vera_doc.option_parsing.require_bool; - an
intdefault uses :func:~vera_doc.option_parsing.require_bounded_intwithmetadata["minimum"]/metadata["maximum"]when those are numbers (otherwise the value must be non-negative); - a
strdefault withmetadata["choices"]and nometadata["allow_custom"]uses :func:~vera_doc.option_parsing.require_choicerestricted to those choices' values; - any other
strdefault uses :func:~vera_doc.option_parsing.require_string(free text).
A field of any other type (for example float) is not supported;
override from_mapping for a class with such a field instead.
Two class attributes customize behavior without any of that:
options_labelsets the name used in error messages (default: the class name with a trailingOptionsdropped, soPyMuPDFOptionsreads as"PyMuPDF").ignored_keysnames legacypipeline_optionskeys to silently accept and drop instead of rejecting as unknown — for compatibility aliases shared with another pipeline that don't apply to this one.
Methods:
Attributes:
-
options_label(str) – -
ignored_keys(frozenset[str]) – -
enforce_step(bool) – -
options_mapping_label(str) –
UnknownIngestPipelineError
¶
Bases: ValueError
Raised when an ingest pipeline spec cannot be resolved.
ParsedBlock
dataclass
¶
ParsedBlock(
page_number: int,
block_type: str,
text: str,
bbox: tuple[float, float, float, float] | None = None,
heading_level: int | None = None,
image_bytes: bytes | None = None,
image_ext: str = "",
)
Layout block from a parser before a stable block_id is assigned.
First-party pipelines (and most custom parsers) emit these, then convert
with :meth:IngestBlock.from_parsed before returning
:class:IngestResult. Pipelines that mint IDs while parsing can
construct :class:IngestBlock directly.
Attributes:
-
page_number(int) – -
block_type(str) – -
text(str) – -
bbox(tuple[float, float, float, float] | None) – -
heading_level(int | None) – -
image_bytes(bytes | None) – -
image_ext(str) –
ParsedPage
dataclass
¶
Single page extracted from a source document.
Attributes:
-
page_number(int) – -
width(float | None) – -
height(float | None) – -
text(str) –
batch_convert
¶
batch_convert(
directory: str | None = None,
*,
paths: Sequence[str] | None = None,
recursive: bool = False,
overwrite: bool = False,
model: str = "hashing",
embedding_function: EmbeddingFunction | None = None,
parser: str | None = None,
chunk_size: int | None = None,
overlap: int | None = None,
store_original: bool = True,
ocr_mode: str | None = None,
ocr_language: str | None = None,
ocr_dpi: int | None = None,
ocr_download: bool | None = None,
pipeline_options: dict[str, Any] | None = None,
embedder_options: dict[str, Any] | None = None,
progress: Callable[[int, int, str], None] | None = None,
cancel: Any | None = None,
metadata: Mapping[str, Any] | None = None,
) -> dict[str, Any]
Convert source files from a directory scan or an explicit path list.
Parameters:
-
directory(str | None, default:None) –Root directory to scan when
pathsis omitted. Discovery uses the selected pipeline'ssource_formats, or every installed pipeline whenparseris omitted. -
paths(Sequence[str] | None, default:None) –Explicit source file paths to convert. When set, directory discovery is skipped and
recursiveis ignored. -
recursive(bool, default:False) –When
True, scan subdirectories (directory mode only). -
overwrite(bool, default:False) –When
True, replace existing.veraoutputs. WhenFalse, skip a sibling archive only when it validates and its storedsource_file_hashmatches the current source file. Stale or hash-less archives are reconverted. Same-stem sources that would write the same.verapath fail instead of overwriting each other, including whenoverwriteis true. -
model(str, default:'hashing') –Embedding model spec passed to :func:
convert. -
embedding_function(EmbeddingFunction | None, default:None) –Optional custom embedder passed to :func:
convert. -
parser(str | None, default:None) –Ingest pipeline spec passed to :func:
convert.None(the default) selects a pipeline per file from its extension. -
chunk_size(int | None, default:None) –Compatibility alias passed to :func:
convert.Nonemeans the pipeline default. -
overlap(int | None, default:None) –Compatibility alias passed to :func:
convert.Nonemeans the pipeline default. -
store_original(bool, default:True) –Whether to embed originals passed to :func:
convert. -
ocr_mode(str | None, default:None) –Compatibility OCR mode alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_language(str | None, default:None) –Compatibility OCR language alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_dpi(int | None, default:None) –Compatibility OCR DPI alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_download(bool | None, default:None) –Compatibility OCR download alias passed to :func:
convert.Nonemeans the pipeline default. -
pipeline_options(dict[str, Any] | None, default:None) –Explicit provider-owned options passed to :func:
convert. -
embedder_options(dict[str, Any] | None, default:None) –Explicit provider-owned embedding options passed to :func:
convert. -
progress(Callable[[int, int, str], None] | None, default:None) –Optional
(current, total, filename)callback. -
cancel(Any | None, default:None) –Optional cancellation token.
-
metadata(Mapping[str, Any] | None, default:None) –Extra keys stamped onto every archive and chunk in this run.
Returns:
-
dict[str, Any]–A report dict with
converted,skipped,failed, and related -
dict[str, Any]–fields.
directoryis the scan root, or the common parent of -
dict[str, Any]–paths.
Raises:
-
NotADirectoryError–When
directoryis not a directory. -
FileNotFoundError–When a path in
pathsis missing. -
ValueError–When neither
directorynorpathsis usable. -
ReservedMetadataKeyError–When
metadatauses a reserved key. -
UnknownEmbeddingModelError–When
modelcannot be resolved.
describe_ingest_pipeline
¶
describe_ingest_pipeline(
spec: str = "pymupdf",
) -> PipelineDescriptor
Return metadata for an installed pipeline without instantiating it.
get_ingest_pipeline
¶
get_ingest_pipeline(
spec: str = "pymupdf",
) -> IngestPipeline
Resolve and cache an installed pipeline.
Resolution is strict: an unknown provider or variant raises instead of falling back to another installed pipeline.
invoke_ingest_pipeline
¶
invoke_ingest_pipeline(
pipeline: IngestPipeline,
source_path: str,
request: IngestRequest,
) -> IngestResult
Call pipeline, accepting both a bare callable and a legacy .ingest() object.
list_ingest_pipelines
¶
Return sorted installed provider names.
list_ingest_pipeline_descriptors
¶
list_ingest_pipeline_descriptors() -> list[
PipelineDescriptor
]
Return descriptors for each installed provider using its default variant.
prepare_pipeline_options
¶
prepare_pipeline_options(
*,
spec: str,
pipeline_options: dict[str, Any] | None = None,
legacy_options: dict[str, Any] | None = None,
) -> dict[str, Any]
Merge legacy convert kwargs with explicit pipeline options.
When a pipeline publishes descriptor fields, only those legacy keys are
forwarded so PyMuPDF defaults such as overlap and ocr_dpi do not
leak into plugins that omit them. Tesseract-shaped aliases
(ocr_language, ocr_dpi, ocr_download) are forwarded only when
capabilities.ocr_engine is "tesseract", so Docling/RapidOCR keeps
its own ocr_language default instead of inheriting eng.
Undescribed plugins receive the remaining compatibility bag. Explicit
pipeline_options always win.
register_ingest_pipeline
¶
register_ingest_pipeline(
provider: str,
factory: Callable[[str], IngestPipeline] | None = None,
*,
replace: bool = False,
) -> (
Callable[
[Callable[[str], IngestPipeline]],
Callable[[str], IngestPipeline],
]
| None
)
Register a provider factory called with the requested variant.
Called with both arguments, this registers factory immediately and
returns None, as before. Omit factory to use it as a decorator
instead — handy for local experiments, notebooks, and tests that would
otherwise need a separate factory function and a separate call::
@register_ingest_pipeline("myexperiment")
def create_pipeline(variant: str = "") -> IngestPipeline:
return MyPipeline()
register_ingest_pipeline_descriptor
¶
register_ingest_pipeline_descriptor(
provider: str,
factory: Callable[[str], PipelineDescriptor]
| None = None,
*,
replace: bool = False,
) -> (
Callable[
[Callable[[str], PipelineDescriptor]],
Callable[[str], PipelineDescriptor],
]
| None
)
Register a descriptor factory called with the requested variant.
Also usable as a decorator when factory is omitted — see
:func:register_ingest_pipeline.
chunk_pages
¶
chunk_pages(
pages: list[ParsedPage],
chunk_size: int = 500,
overlap: int = 75,
) -> list[Chunk]
Sliding-window chunker over ParsedPage.text (whitespace-split words).
Public helper for custom pipelines that only have page text. First-party
pipelines (PyMuPDF, Docling) do not call this; they chunk structured
layout with :func:build_chunks_from_blocks or their own chunker.
build_chunks_from_blocks
¶
build_chunks_from_blocks(
blocks: list[tuple[str, ParsedBlock | IngestBlock]],
chunk_size: int = 500,
overlap: int = 75,
) -> list[Chunk]
Heading-aware chunking over structured layout blocks.
Accepts (block_id, block) pairs of either :class:ParsedBlock or
:class:IngestBlock. First-party PyMuPDF uses this path; custom pipelines
that only have page text should use :func:chunk_pages instead.
detect_heading
¶
Return the first short heading-like line in text, else current.
Public helper for custom pipelines that chunk page text with
:func:chunk_pages. First-party pipelines use structured heading blocks
via :func:build_chunks_from_blocks instead.
figures
¶
figures(
document: VeraDocument,
page_start: int | None = None,
page_end: int | None = None,
include_data: bool = False,
attachment_ids: Iterable[str] | None = None,
) -> list[dict[str, Any]]
Return figure attachments produced during ingest.
figures_for
¶
figures_for(
document: VeraDocument,
result: QueryResult,
include_data: bool = False,
) -> list[dict[str, Any]]
Return figure attachments linked to a query result.
get_page
¶
get_page(
document: VeraDocument, page_number: int
) -> dict[str, Any] | None
Return ingest-provided viewer data for one page.
get_blocks
¶
get_blocks(
document: VeraDocument, page_number: int | None = None
) -> list[dict[str, Any]]
Return ingest-provided layout blocks.
get_chunk_regions
¶
get_chunk_regions(
document: VeraDocument, chunk_id: str
) -> list[dict[str, Any]]
Return ingest-provided highlight regions for a chunk.
regions_for
¶
regions_for(
document: VeraDocument, result: QueryResult
) -> list[dict[str, Any]]
Return ingest-provided highlight regions for a query result.
chunk_payload
¶
chunk_payload(
record: ChunkRecord,
*,
document: VeraDocument | None = None,
include_figures: bool = False,
include_regions: bool = False,
include_figure_data: bool = False,
figure_data_urls: bool = False,
) -> dict[str, Any]
Flatten a stored chunk for CLI/MCP JSON (metadata keys at the top level).
Citation fields such as page_start and heading_path sit beside
chunk_id and text. The embedding vector and retrieval scores are
omitted. Optional figure and region enrichment matches
:func:result_payload.
result_payload
¶
result_payload(
result: QueryResult,
*,
document: VeraDocument | None = None,
include_figures: bool = False,
include_regions: bool = False,
include_figure_data: bool = False,
figure_data_urls: bool = False,
) -> dict[str, Any]
Flatten a search hit for CLI/MCP/app JSON (metadata keys at the top level).
Optional figure and region enrichment uses ingest viewer helpers. Sidecar
callers can set figure_data_urls to replace raw figure bytes with a
data_url instead of forking the serializer. Extra as_dict() keys such
as corpus file are preserved.
get_source_document
¶
get_source_document(
document: VeraDocument,
) -> AttachmentRecord
Return the attachment identified as the archive's source document.
export_figures
¶
export_figures(
document: VeraDocument,
directory: str | PathLike[str],
*,
asset_ids: Iterable[str] | None = None,
page_start: int | None = None,
page_end: int | None = None,
) -> list[dict[str, Any]]
Write figure attachments under directory and return metadata plus paths.
Output names are {asset_id}.{ext}. ext comes from the stored mime
type or filename. Requested ids that are missing or not figure attachments
raise ValueError so a source PDF id cannot leak. Raw data is never
included in the returned dicts.
export_source_document
¶
export_source_document(
document: VeraDocument,
path: str | PathLike[str] | None = None,
) -> str
Write the source attachment to disk and return its path.
The stored filename is used as Path(...).name only. Absolute names and
.. segments are rejected. When path is omitted the file is written
under the current working directory; when path is a directory the file
stays under that directory. An explicit file path is the caller's chosen
output location.
Conversion writes through VeraDocument. Viewer helpers
interpret ingest-produced attachments and metadata. Shared convert accepts
opaque pipeline_options on a thin IngestRequest; pipelines own typed
defaults and descriptors. Prefer IngestRequest / pipeline_options over the
deprecated IngestOptions compatibility bag. See the
conversion guide and
figures and regions.
convert
¶
Classes:
-
ReservedMetadataKeyError–Caller
metadatacollided with a reserved convert or format key.
Functions:
-
convert–Convert a source document into a validated
.veraarchive. -
batch_convert–Convert source files from a directory scan or an explicit path list.
ReservedMetadataKeyError
¶
Bases: ValueError
Caller metadata collided with a reserved convert or format key.
convert
¶
convert(
input_path: str,
output_path: str,
*,
model: str = "hashing",
embedding_function: EmbeddingFunction | None = None,
parser: str | None = None,
chunk_size: int | None = None,
overlap: int | None = None,
store_original: bool = True,
ocr_mode: str | None = None,
ocr_language: str | None = None,
ocr_dpi: int | None = None,
ocr_download: bool | None = None,
pipeline_options: dict[str, Any] | None = None,
embedder_options: dict[str, Any] | None = None,
cancel: Any | None = None,
metadata: Mapping[str, Any] | None = None,
) -> str
Convert a source document into a validated .vera archive.
Parses the file, chunks extracted text, embeds chunks, and writes the
result through :class:~vera_doc.document.VeraDocument. The archive is
validated before the temporary file is published atomically.
New callers should pass parser, pipeline_options, and embedder
settings (model / embedding_function / embedder_options).
chunk_size, overlap, ocr_mode, ocr_language, ocr_dpi,
and ocr_download remain compatibility aliases for CLI and sidecar
callers; they are forwarded only when explicitly provided and the
selected pipeline advertises them (Tesseract OCR aliases do not leak
to Docling). Omitted aliases mean "use the pipeline's own default"
(for example a plugin chunk_size of 2000 is not overwritten by 500).
The CLI still passes its argparse defaults when invoked from the command
line.
Parameters:
-
input_path(str) –Source document path (PDF, Markdown, or another format advertised by an installed ingest pipeline).
-
output_path(str) –Destination
.verapath. -
model(str, default:'hashing') –Embedding model spec (default
"hashing"). Ignored whenembedding_functionis provided. Acceptsprovider:model-idor built-in legacy aliases. -
embedding_function(EmbeddingFunction | None, default:None) –Optional custom embedder satisfying :class:
~vera_doc.EmbeddingFunction. When omitted,modelis resolved via :func:~vera_doc.get_embedderbefore parsing begins. -
parser(str | None, default:None) –Ingest pipeline spec in
provider[:variant]form.None(the default) selects an installed pipeline from the file extension. An explicit spec must advertise that extension. -
chunk_size(int | None, default:None) –Compatibility alias forwarded only when explicitly provided and the selected pipeline advertises a
chunk_sizefield.None(the default) means the pipeline default. -
overlap(int | None, default:None) –Compatibility alias forwarded only when explicitly provided and advertised by the selected pipeline (PyMuPDF, Markdown). Ignored by Docling.
Nonemeans the pipeline default. -
store_original(bool, default:True) –When
True, embed the original file as an attachment. -
ocr_mode(str | None, default:None) –Compatibility OCR mode alias when explicitly provided and advertised by the pipeline.
Nonemeans the pipeline default. -
ocr_language(str | None, default:None) –Tesseract OCR language alias (PyMuPDF). Forwarded only when explicitly provided and the selected pipeline's
ocr_engineis"tesseract".Nonemeans the pipeline default. -
ocr_dpi(int | None, default:None) –Compatibility OCR DPI alias when explicitly provided and advertised (PyMuPDF).
Nonemeans the pipeline default. -
ocr_download(bool | None, default:None) –Compatibility alias (PyMuPDF only) allowing on-demand, checksum-verified download of missing Tesseract language data.
Nonemeans the pipeline default. -
pipeline_options(dict[str, Any] | None, default:None) –Explicit provider-owned options. These override compatibility aliases for the same keys.
-
embedder_options(dict[str, Any] | None, default:None) –Explicit provider-owned embedding options forwarded to :func:
~vera_doc.get_embedderwhenembedding_functionis omitted. -
cancel(Any | None, default:None) –Optional cancellation token with
raise_if_cancelled(). -
metadata(Mapping[str, Any] | None, default:None) –Extra keys stamped onto archive metadata and every chunk. Reserved ingest, citation, and format keys are rejected.
Returns:
-
str–The
output_pathstring.
Raises:
-
FileNotFoundError–When
input_pathdoes not exist. -
ValueError–When no searchable text is extracted, or
parserdoes not support the source file type. -
ReservedMetadataKeyError–When
metadatauses a reserved key. -
UnknownIngestPipelineError–When
parsercannot be resolved. -
UnknownEmbeddingModelError–When
modelcannot be resolved.
batch_convert
¶
batch_convert(
directory: str | None = None,
*,
paths: Sequence[str] | None = None,
recursive: bool = False,
overwrite: bool = False,
model: str = "hashing",
embedding_function: EmbeddingFunction | None = None,
parser: str | None = None,
chunk_size: int | None = None,
overlap: int | None = None,
store_original: bool = True,
ocr_mode: str | None = None,
ocr_language: str | None = None,
ocr_dpi: int | None = None,
ocr_download: bool | None = None,
pipeline_options: dict[str, Any] | None = None,
embedder_options: dict[str, Any] | None = None,
progress: Callable[[int, int, str], None] | None = None,
cancel: Any | None = None,
metadata: Mapping[str, Any] | None = None,
) -> dict[str, Any]
Convert source files from a directory scan or an explicit path list.
Parameters:
-
directory(str | None, default:None) –Root directory to scan when
pathsis omitted. Discovery uses the selected pipeline'ssource_formats, or every installed pipeline whenparseris omitted. -
paths(Sequence[str] | None, default:None) –Explicit source file paths to convert. When set, directory discovery is skipped and
recursiveis ignored. -
recursive(bool, default:False) –When
True, scan subdirectories (directory mode only). -
overwrite(bool, default:False) –When
True, replace existing.veraoutputs. WhenFalse, skip a sibling archive only when it validates and its storedsource_file_hashmatches the current source file. Stale or hash-less archives are reconverted. Same-stem sources that would write the same.verapath fail instead of overwriting each other, including whenoverwriteis true. -
model(str, default:'hashing') –Embedding model spec passed to :func:
convert. -
embedding_function(EmbeddingFunction | None, default:None) –Optional custom embedder passed to :func:
convert. -
parser(str | None, default:None) –Ingest pipeline spec passed to :func:
convert.None(the default) selects a pipeline per file from its extension. -
chunk_size(int | None, default:None) –Compatibility alias passed to :func:
convert.Nonemeans the pipeline default. -
overlap(int | None, default:None) –Compatibility alias passed to :func:
convert.Nonemeans the pipeline default. -
store_original(bool, default:True) –Whether to embed originals passed to :func:
convert. -
ocr_mode(str | None, default:None) –Compatibility OCR mode alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_language(str | None, default:None) –Compatibility OCR language alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_dpi(int | None, default:None) –Compatibility OCR DPI alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_download(bool | None, default:None) –Compatibility OCR download alias passed to :func:
convert.Nonemeans the pipeline default. -
pipeline_options(dict[str, Any] | None, default:None) –Explicit provider-owned options passed to :func:
convert. -
embedder_options(dict[str, Any] | None, default:None) –Explicit provider-owned embedding options passed to :func:
convert. -
progress(Callable[[int, int, str], None] | None, default:None) –Optional
(current, total, filename)callback. -
cancel(Any | None, default:None) –Optional cancellation token.
-
metadata(Mapping[str, Any] | None, default:None) –Extra keys stamped onto every archive and chunk in this run.
Returns:
-
dict[str, Any]–A report dict with
converted,skipped,failed, and related -
dict[str, Any]–fields.
directoryis the scan root, or the common parent of -
dict[str, Any]–paths.
Raises:
-
NotADirectoryError–When
directoryis not a directory. -
FileNotFoundError–When a path in
pathsis missing. -
ValueError–When neither
directorynorpathsis usable. -
ReservedMetadataKeyError–When
metadatauses a reserved key. -
UnknownEmbeddingModelError–When
modelcannot be resolved.
batch_convert
¶
batch_convert(
directory: str | None = None,
*,
paths: Sequence[str] | None = None,
recursive: bool = False,
overwrite: bool = False,
model: str = "hashing",
embedding_function: EmbeddingFunction | None = None,
parser: str | None = None,
chunk_size: int | None = None,
overlap: int | None = None,
store_original: bool = True,
ocr_mode: str | None = None,
ocr_language: str | None = None,
ocr_dpi: int | None = None,
ocr_download: bool | None = None,
pipeline_options: dict[str, Any] | None = None,
embedder_options: dict[str, Any] | None = None,
progress: Callable[[int, int, str], None] | None = None,
cancel: Any | None = None,
metadata: Mapping[str, Any] | None = None,
) -> dict[str, Any]
Convert source files from a directory scan or an explicit path list.
Parameters:
-
directory(str | None, default:None) –Root directory to scan when
pathsis omitted. Discovery uses the selected pipeline'ssource_formats, or every installed pipeline whenparseris omitted. -
paths(Sequence[str] | None, default:None) –Explicit source file paths to convert. When set, directory discovery is skipped and
recursiveis ignored. -
recursive(bool, default:False) –When
True, scan subdirectories (directory mode only). -
overwrite(bool, default:False) –When
True, replace existing.veraoutputs. WhenFalse, skip a sibling archive only when it validates and its storedsource_file_hashmatches the current source file. Stale or hash-less archives are reconverted. Same-stem sources that would write the same.verapath fail instead of overwriting each other, including whenoverwriteis true. -
model(str, default:'hashing') –Embedding model spec passed to :func:
convert. -
embedding_function(EmbeddingFunction | None, default:None) –Optional custom embedder passed to :func:
convert. -
parser(str | None, default:None) –Ingest pipeline spec passed to :func:
convert.None(the default) selects a pipeline per file from its extension. -
chunk_size(int | None, default:None) –Compatibility alias passed to :func:
convert.Nonemeans the pipeline default. -
overlap(int | None, default:None) –Compatibility alias passed to :func:
convert.Nonemeans the pipeline default. -
store_original(bool, default:True) –Whether to embed originals passed to :func:
convert. -
ocr_mode(str | None, default:None) –Compatibility OCR mode alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_language(str | None, default:None) –Compatibility OCR language alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_dpi(int | None, default:None) –Compatibility OCR DPI alias passed to :func:
convert.Nonemeans the pipeline default. -
ocr_download(bool | None, default:None) –Compatibility OCR download alias passed to :func:
convert.Nonemeans the pipeline default. -
pipeline_options(dict[str, Any] | None, default:None) –Explicit provider-owned options passed to :func:
convert. -
embedder_options(dict[str, Any] | None, default:None) –Explicit provider-owned embedding options passed to :func:
convert. -
progress(Callable[[int, int, str], None] | None, default:None) –Optional
(current, total, filename)callback. -
cancel(Any | None, default:None) –Optional cancellation token.
-
metadata(Mapping[str, Any] | None, default:None) –Extra keys stamped onto every archive and chunk in this run.
Returns:
-
dict[str, Any]–A report dict with
converted,skipped,failed, and related -
dict[str, Any]–fields.
directoryis the scan root, or the common parent of -
dict[str, Any]–paths.
Raises:
-
NotADirectoryError–When
directoryis not a directory. -
FileNotFoundError–When a path in
pathsis missing. -
ValueError–When neither
directorynorpathsis usable. -
ReservedMetadataKeyError–When
metadatauses a reserved key. -
UnknownEmbeddingModelError–When
modelcannot be resolved.