Skip to content

Document

This is an automatically generated API reference of the VERA storage and search engine (vera_doc.document).

document

Classes:

  • VeraDocument –

    An embedded storage and search engine backed by one portable .vera file.

  • EmbeddingFunction –

    Structural protocol for embedders used when writing and querying archives.

  • DuplicateRecordError –

    Raised when add() receives an ID that already exists.

  • RecordNotFoundError –

    Raised when a requested chunk or attachment does not exist.

  • ReadOnlyError –

    Raised when a write is attempted on a read-only database.

VeraDocument

VeraDocument(
    path: Path,
    conn: Connection,
    *,
    mode: OpenMode,
    embedding_function: EmbeddingFunction | None,
)

Bases: _DocumentRecordsMixin, _DocumentSearchMixin, _DocumentAttachmentsMixin

An embedded storage and search engine backed by one portable .vera file.

Use :meth:create to initialize a new archive and :meth:open to access an existing one. Archives support CRUD on :class:~vera_doc.models.ChunkRecord objects, optional binary attachments, and semantic, keyword, or hybrid search.

Example
from vera_doc import ChunkRecord, VeraDocument

with VeraDocument.create("example.vera") as document:
    document.add([ChunkRecord(id="1", text="Hello world.")])

with VeraDocument.open("example.vera") as document:
    results = document.search(text="hello", top_k=5)

Methods:

  • create –

    Create a new empty .vera archive at path.

  • open –

    Open an existing .vera archive.

  • set_metadata –

    Replace caller-controlled archive metadata.

  • inspect –

    Return archive metadata and record counts.

  • validate –

    Check archive integrity.

  • transaction –

    Run a batch of writes in a single SQLite transaction.

  • format_metadata –

    Return the archive format header as a key/value mapping.

  • iter_raw_chunks –

    Yield chunk text, metadata, model name, dimension, and raw vector bytes.

  • close –

    Close the database connection.

  • __enter__ –
  • __exit__ –
  • put_attachments –

    Store opaque binary attachments.

  • get_attachment –

    Return a stored attachment by ID.

  • write_attachment –

    Copy attachment bytes to dest without building an in-memory record.

  • attachment_metadata –

    Return attachment descriptors without reading binary payloads.

  • attachments –

    Return stored attachments, optionally filtered by metadata equality.

  • delete_attachment –

    Delete a stored attachment.

  • get –

    Fetch chunk records by ID and/or metadata filter.

  • search –

    Search chunk records.

  • add –

    Insert new chunk records.

  • upsert –

    Insert or replace chunk records atomically.

  • delete –

    Delete chunk records by ID and/or metadata filter.

Attributes:

  • path –
  • mode (OpenMode) –
  • metadata (dict[str, Any]) –

    Caller-controlled archive metadata as a mutable dict.

path instance-attribute

path = path

mode property

mode: OpenMode

metadata property

metadata: dict[str, Any]

Caller-controlled archive metadata as a mutable dict.

create classmethod

create(
    path: str | PathLike[str],
    *,
    embedding_function: EmbeddingFunction | None = None,
    model: str = "hashing",
    embedding_normalization: EmbeddingNormalization
    | None = None,
    metadata: JsonObject | None = None,
    overwrite: bool = False,
) -> VeraDocument

Create a new empty .vera archive at path.

Parameters:

  • path (str | PathLike[str]) –

    Destination file path.

  • embedding_function (EmbeddingFunction | None, default: None ) –

    Custom embedder. When omitted, model selects the default embedder.

  • model (str, default: 'hashing' ) –

    Default embedding model name (for example "hashing").

  • embedding_normalization (EmbeddingNormalization | None, default: None ) –

    Stored-vector normalization policy. When omitted, use the embedder's declared policy or "unknown".

  • metadata (JsonObject | None, default: None ) –

    Caller-controlled JSON metadata stored in the archive.

  • overwrite (bool, default: False ) –

    When False (default), raise :class:FileExistsError if path already exists.

Returns:

Raises:

  • FileExistsError –

    When the target exists and overwrite is false.

open classmethod

open(
    path: str | PathLike[str],
    *,
    mode: OpenMode = "read",
    embedding_function: EmbeddingFunction | None = None,
) -> VeraDocument

Open an existing .vera archive.

Parameters:

  • path (str | PathLike[str]) –

    Path to an existing archive.

  • mode (OpenMode, default: 'read' ) –

    "read" (default) opens SQLite read-only; "write" allows mutations.

  • embedding_function (EmbeddingFunction | None, default: None ) –

    Embedder used for write-mode searches and record writes. When omitted in write mode, the model recorded in the archive is used.

Returns:

Raises:

  • FileNotFoundError –

    When path does not exist.

  • ValueError –

    When the archive format version is unsupported or the embedder dimension does not match the archive.

set_metadata

set_metadata(metadata: JsonObject) -> None

Replace caller-controlled archive metadata.

Parameters:

  • metadata (JsonObject) –

    JSON-compatible mapping stored in the archive header.

Raises:

inspect

inspect() -> dict[str, Any]

Return archive metadata and record counts.

Returns:

  • dict[str, Any] –

    A dict with path, format_version, embedding_model,

  • dict[str, Any] –

    chunks, attachments, and related fields.

validate

validate() -> dict[str, Any]

Check archive integrity.

Returns:

  • dict[str, Any] –

    A report dict with an ok boolean and an issues list.

transaction

transaction() -> Iterator[VeraDocument]

Run a batch of writes in a single SQLite transaction.

Yields:

  • VeraDocument –

    This database handle for use inside the with block.

Raises:

  • RuntimeError –

    When nested transactions are attempted.

  • ReadOnlyError –

    When the database is opened read-only.

format_metadata

format_metadata() -> dict[str, str]

Return the archive format header as a key/value mapping.

iter_raw_chunks

iter_raw_chunks() -> Iterator[dict[str, Any]]

Yield chunk text, metadata, model name, dimension, and raw vector bytes.

This is the bulk-read path used by the library index builder. It avoids constructing :class:ChunkRecord objects and does not load attachments.

close

close() -> None

Close the database connection.

__enter__

__enter__() -> VeraDocument

__exit__

__exit__(*exc: object) -> None

put_attachments

put_attachments(
    attachments: Iterable[AttachmentRecord],
    *,
    upsert: bool = False,
) -> None

Store opaque binary attachments.

Parameters:

  • attachments (Iterable[AttachmentRecord]) –

    Attachment payloads to insert.

  • upsert (bool, default: False ) –

    When True, replace existing attachments with the same ID.

Raises:

get_attachment

get_attachment(attachment_id: str) -> AttachmentRecord

Return a stored attachment by ID.

Parameters:

  • attachment_id (str) –

    Attachment identifier.

Returns:

Raises:

write_attachment

write_attachment(
    attachment_id: str,
    dest: str | PathLike[str],
    *,
    chunk_size: int = _ATTACHMENT_WRITE_CHUNK,
    on_chunk: Callable[[int], None] | None = None,
) -> int

Copy attachment bytes to dest without building an in-memory record.

Prefers SQLite incremental blob I/O when the interpreter provides Connection.blobopen so a large PDF is not materialized as a Python bytes object. Older Pythons fall back to a single SELECT.

Parameters:

  • attachment_id (str) –

    Attachment identifier.

  • dest (str | PathLike[str]) –

    Destination file path. Parent directories are created.

  • chunk_size (int, default: _ATTACHMENT_WRITE_CHUNK ) –

    Write size in bytes.

  • on_chunk (Callable[[int], None] | None, default: None ) –

    Optional callback invoked with bytes written so far before each chunk, including a final call after the last write. Useful for cooperative cancellation.

Returns:

  • int –

    Number of bytes written.

Raises:

  • RecordNotFoundError –

    When no attachment exists with that ID.

  • ValueError –

    When chunk_size is less than 1.

attachment_metadata

attachment_metadata(
    ids: Iterable[str] | None = None,
    *,
    where: Mapping[str, Any] | None = None,
) -> list[dict[str, Any]]

Return attachment descriptors without reading binary payloads.

Parameters:

  • ids (Iterable[str] | None, default: None ) –

    Specific attachment IDs. When omitted, inspect all attachments.

  • where (Mapping[str, Any] | None, default: None ) –

    Exact equality filter on top-level attachment metadata keys.

Returns:

  • list[dict[str, Any]] –

    Matching descriptors in storage order, or requested ID order when

  • list[dict[str, Any]] –

    ids is provided. Each descriptor includes size (payload

  • list[dict[str, Any]] –

    length in bytes) and omits data.

attachments

attachments(
    *, where: Mapping[str, Any] | None = None
) -> list[AttachmentRecord]

Return stored attachments, optionally filtered by metadata equality.

Parameters:

  • where (Mapping[str, Any] | None, default: None ) –

    Exact equality filter on top-level attachment metadata keys.

Returns:

delete_attachment

delete_attachment(attachment_id: str) -> None

Delete a stored attachment.

Parameters:

  • attachment_id (str) –

    Attachment identifier.

Raises:

  • RecordNotFoundError –

    When no attachment exists with that ID.

  • ValueError –

    When the attachment is still referenced by a chunk.

  • ReadOnlyError –

    When the database is opened read-only.

get

get(
    ids: Iterable[str] | None = None,
    *,
    where: Mapping[str, Any] | None = None,
    limit: int | None = None,
) -> list[ChunkRecord]

Fetch chunk records by ID and/or metadata filter.

Parameters:

  • ids (Iterable[str] | None, default: None ) –

    Specific chunk IDs to retrieve. When omitted, all matching records are returned subject to where and limit.

  • where (Mapping[str, Any] | None, default: None ) –

    Filter on top-level metadata keys. Scalars are exact equality; a list, tuple, or set value is IN. Distinct keys are AND.

  • limit (int | None, default: None ) –

    Maximum number of records to return.

Returns:

  • Matching ( list[ChunkRecord] ) –

    class:~vera_doc.models.ChunkRecord objects in storage order.

search

search(
    text: str | None = None,
    *,
    vector: Sequence[float] | None = None,
    mode: SearchMode = "hybrid",
    where: Mapping[str, Any] | None = None,
    top_k: int = 10,
    context_chunks: int = 0,
    semantic_weight: float = DEFAULT_HYBRID_SEMANTIC_WEIGHT,
    keyword_weight: float = DEFAULT_HYBRID_KEYWORD_WEIGHT,
) -> list[QueryResult]

Search chunk records.

Parameters:

  • text (str | None, default: None ) –

    Query string for semantic or keyword search. Required unless vector is supplied for semantic mode.

  • vector (Sequence[float] | None, default: None ) –

    Pre-computed query vector for semantic search.

  • mode (SearchMode, default: 'hybrid' ) –

    "hybrid" (default), "semantic", or "keyword".

  • where (Mapping[str, Any] | None, default: None ) –

    Filter on top-level metadata keys, applied before top_k. Scalars are exact equality; a list, tuple, or set value is IN. Distinct keys are AND.

  • top_k (int, default: 10 ) –

    Maximum number of results to return.

  • context_chunks (int, default: 0 ) –

    Number of adjacent stored chunks to include.

  • semantic_weight (float, default: DEFAULT_HYBRID_SEMANTIC_WEIGHT ) –

    Hybrid blend weight for semantic scores.

  • keyword_weight (float, default: DEFAULT_HYBRID_KEYWORD_WEIGHT ) –

    Hybrid blend weight for keyword scores.

Returns:

  • Ranked ( list[QueryResult] ) –

    class:~vera_doc.models.QueryResult objects.

add

add(records: Iterable[ChunkRecord]) -> None

Insert new chunk records.

Parameters:

  • records (Iterable[ChunkRecord]) –

    Records to insert. IDs must not already exist.

Raises:

upsert

upsert(records: Iterable[ChunkRecord]) -> None

Insert or replace chunk records atomically.

Parameters:

  • records (Iterable[ChunkRecord]) –

    Records to insert or update by ID.

Raises:

delete

delete(
    ids: Iterable[str] | None = None,
    *,
    where: Mapping[str, Any] | None = None,
    delete_all: bool = False,
) -> int

Delete chunk records by ID and/or metadata filter.

Parameters:

  • ids (Iterable[str] | None, default: None ) –

    Specific chunk IDs to delete.

  • where (Mapping[str, Any] | None, default: None ) –

    Filter on top-level metadata keys. Scalars are exact equality; a list, tuple, or set value is IN. Distinct keys are AND.

  • delete_all (bool, default: False ) –

    Required to delete every chunk when ids and where are omitted.

Returns:

  • int –

    The number of records deleted.

Raises:

  • ValueError –

    When neither ids nor where is given and delete_all is false.

EmbeddingFunction

Bases: Protocol

Structural protocol for embedders used when writing and querying archives.

Implementations may also expose a normalization attribute ("l2", "none", or "unknown"). Embedders without it are recorded as "unknown".

Attributes:

  • model_name (str) –

    Identifier stored in archive metadata.

  • dimension (int) –

    Vector length expected by the database.

Methods:

  • embed –

    Embed a batch of texts.

model_name instance-attribute

model_name: str

dimension instance-attribute

dimension: int

embed

embed(texts: list[str]) -> ndarray | list[ndarray]

Embed a batch of texts.

DuplicateRecordError

Bases: ValueError

Raised when add() receives an ID that already exists.

RecordNotFoundError

Bases: KeyError

Raised when a requested chunk or attachment does not exist.

ReadOnlyError

Bases: PermissionError

Raised when a write is attempted on a read-only database.

See QueryResult for the shape of search hits.