vera-ingest examples¶
These examples use the public vera_ingest API directly. For equivalent
shell workflows, use vera convert.
Convert one PDF¶
from vera_ingest import convert
output = convert(
"manual.pdf",
"manual.vera",
model="hashing",
store_original=True,
pipeline_options={
"chunk_size": 500,
"overlap": 75,
"ocr_mode": "auto",
},
)
print(output)
The function returns the output path after validating and atomically publishing
the archive. Legacy kwargs such as chunk_size= and ocr_mode= remain
compatibility aliases; explicit pipeline_options win for matching keys.
Force OCR¶
from vera_ingest import convert
convert(
"scanned-manual.pdf",
"scanned-manual.vera",
pipeline_options={
"ocr_mode": "force",
"ocr_language": "eng",
"ocr_dpi": 300,
},
)
Use forced OCR only when automatic detection misses scanned content.
Convert a directory¶
from vera_ingest import batch_convert
report = batch_convert(
"./proposals",
recursive=True,
model="hashing",
ocr_mode="auto",
)
print("converted:", report["converted"])
print("skipped existing:", report["skipped_existing"])
print("failed:", report["failed"])
print("malformed existing:", report["malformed_existing"])
Batch conversion continues after per-file failures. skipped_existing lists
valid archives whose stored source_file_hash still matches the current source file.
Same-stem sources that would write the same .vera path appear in errors
instead of overwriting each other. Check
both failed and malformed_existing before treating the batch as successful.
See Convert documents for every supported option and its filesystem behavior.