Skip to content

API Reference

LineageStore is the public entry point, available from the top-level ancestree package. Everything below is a method on it, apart from the two node types at the end.

The store is grouped here by what you are trying to do rather than alphabetically.

Section What it covers
Creating a store Opening a store and setting its rules and policy
Recording work Writing nodes, artifacts and metadata
Searching and querying Finding nodes and walking lineage
Visualisation The static snapshot and the live explorer
Maintenance Pruning, compacting, backing up and exporting
Introspection Direct SQL and store-level statistics
Working with nodes The two node types you receive

Creating a store

A store is a directory holding one SQLite database. Rules, generation triggers and the reuse_identical/delta policy are written in at creation and read back on every later open, so reopening by path is enough. See Caveats for what cannot be changed afterwards.

ancestree.LineageStore

LineageStore(root, rules=None, gen_triggers=None, reuse_identical=None, delta=None)

Orchestrates the lineage and interactions of a data pipeline.

The entire durable store is one SQLite database (<root>/ancestree.db) holding metadata and deduplicated artifact chunks. A store can be recreated from its root at any time; rules and policy persist in the database itself.

Opens (creating if needed) the store at root.

Parameters:

Name Type Description Default
root Path | str

The store's directory; the database lives inside it.

required
rules dict[str, list[str]] | None

Allowed transitions, e.g. {"clean": ["ingest"]}. Persisted at creation; immutable afterwards.

None
gen_triggers list[str] | None

Step types that start a new generation. Persisted at creation; immutable afterwards.

None
reuse_identical bool | None

A node with the same parents and content as an existing one is bound to that node instead of being written again (persisted at creation; defaults to True).

None
delta bool | None

Layer-2 storage policy — a new chunk similar to one already stored is kept as a delta against it. Layer-1 chunking and exact chunk deduplication always run, whatever this is set to (persisted at creation; defaults to True).

None

Raises:

Type Description
SchemaError

If root holds a 0.1.x store, or a database this ancestree did not write (see db.schema.ensure_schema).

Source code in src/ancestree/store.py
def __init__(
    self,
    root: Path | str,
    rules: dict[str, list[str]] | None = None,
    gen_triggers: list[str] | None = None,
    reuse_identical: bool | None = None,
    delta: bool | None = None,
) -> None:
    """Opens (creating if needed) the store at `root`.

    Args:
        root: The store's directory; the database lives inside it.
        rules: Allowed transitions, e.g. ``{"clean": ["ingest"]}``.
            Persisted at creation; immutable afterwards.
        gen_triggers: Step types that start a new generation.
            Persisted at creation; immutable afterwards.
        reuse_identical: A node with the same parents and content as an
            existing one is bound to that node instead of being written
            again (persisted at creation; defaults to True).
        delta: Layer-2 storage policy — a new chunk similar to one
            already stored is kept as a delta against it. Layer-1
            chunking and exact chunk deduplication always run,
            whatever this is set to (persisted at creation; defaults
            to True).

    Raises:
        SchemaError: If `root` holds a 0.1.x store, or a database this
            ancestree did not write (see `db.schema.ensure_schema`).
    """
    self.root = Path(root)
    _refuse_legacy_root(self.root)
    self.root.mkdir(parents=True, exist_ok=True)
    self._manager = ConnectionManager(self.root / DB_FILENAME)
    self._metadata = MetadataStore(self._manager)
    config = self._load_or_create_config(
        rules, gen_triggers, reuse_identical, delta
    )
    self.rules: dict[str, list[str]] = config["rules"]
    self.gen_triggers: list[str] = config["gen_triggers"]
    self.reuse_identical: bool = config["reuse_identical"]
    self.delta: bool = config["delta"]
    self._chunks = ChunkStore(self._manager, delta=self.delta)
    self._engine = RuleEngine(self.rules, self.gen_triggers)
    self._pruner = Pruner(self._manager, self._metadata)
    # Resources are released when close() is called, when the store is
    # garbage collected, or at interpreter exit — whichever comes first.
    # The finalizer callback must not reference `self` (that would keep
    # the store alive forever), so the resources travel in this dict.
    self._resources: dict[str, Any] = {
        "manager": self._manager,
        "chunks": self._chunks,
        "live_server": None,
    }
    self._finalizer = weakref.finalize(self, LineageStore._release, self._resources)
    adopted = sweep_orphan_scratch(
        self.root, self._manager, self._metadata, self._chunks
    )
    if adopted:
        warnings.warn(
            f"Adopted {len(adopted)} orphaned node(s) from a crashed "
            f"session as unhealthy: {adopted}. Their partial artifacts "
            "are preserved as evidence.",
            UserWarning,
            stacklevel=2,
        )

Recording work

The one context manager everything else depends on. On clean exit the node commits atomically; if the block raises, partial work is kept and flagged healthy=False; an untouched node is discarded with a warning.

create_node

create_node(step_type, parent=None)

Creates a new node while enforcing lineage rules.

Yields a recording handle: write artifacts with node / "file" and attach metadata with node.add_meta. On clean exit the node is committed atomically; if the block raises, whatever was written is persisted with healthy=False (partial work is evidence). An untouched node is discarded with a warning.

Parameters:

Name Type Description Default
step_type str

The type of pipeline step being performed.

required
parent ParentArg

A Node/handle/node_id, or a list of them for a join. Every parent must exist in this store. None creates a root.

None

Raises:

Type Description
InvalidTransition

If the transition is not permitted.

ValueError

For an unusable step_type or unknown parent.

Source code in src/ancestree/store.py
@contextmanager
def create_node(
    self, step_type: str, parent: ParentArg = None
) -> Iterator[RecordingNode]:
    """Creates a new node while enforcing lineage rules.

    Yields a recording handle: write artifacts with ``node / "file"``
    and attach metadata with ``node.add_meta``. On clean exit the node
    is committed atomically; if the block raises, whatever was written
    is persisted with ``healthy=False`` (partial work is evidence). An
    untouched node is discarded with a warning.

    Args:
        step_type: The type of pipeline step being performed.
        parent: A Node/handle/node_id, or a list of them for a join.
            Every parent must exist in this store. None creates a root.

    Raises:
        InvalidTransition: If the transition is not permitted.
        ValueError: For an unusable step_type or unknown parent.
    """
    validate_step_type(step_type)
    parents = self._resolve_parents(parent)
    self._engine.validate(step_type, [p.step_type for p in parents])
    generation = self._engine.generation_for(
        step_type, [p.generation for p in parents]
    )

    node_id = uuid.uuid4().hex[:8]
    while (
        self._metadata.exists(node_id)
        or (self.root / ".scratch" / node_id).exists()
    ):
        node_id = uuid.uuid4().hex[:8]

    parent_id = [p.node_id for p in parents]
    workspace = NodeWorkspace(
        self.root, node_id, step_type, parent_id, generation=generation
    )
    handle = RecordingNode(node_id, step_type, generation, parent_id, workspace)

    start = time.monotonic()
    try:
        yield handle
    except BaseException:
        # Keep partial work: anything recorded before the failure
        # persists, flagged unhealthy. An untouched node leaves no trace.
        self._finalize(
            handle, workspace, healthy=False, duration=time.monotonic() - start
        )
        raise

    if not self._finalize(
        handle, workspace, healthy=True, duration=time.monotonic() - start
    ):
        warnings.warn(
            f"Node '{handle.node_id}' (step_type='{step_type}') was "
            "discarded: no artifacts were written and no metadata was "
            "added. Write at least one file or call node.add_meta() to "
            "persist the node.",
            UserWarning,
            # 1=here, 2=contextlib.__exit__, 3=user `with` statement.
            stacklevel=3,
        )

Searching and querying

find and latest search the whole store; lineage, ancestors, children and from_parent move around the graph. Filters match structural attributes and searchable metadata in one namespace, and a callable is treated as a predicate.

find

find(**filters)

Nodes matching every filter, oldest first. Filters match structural attributes (step_type, generation, healthy, ...) and searchable metadata by equality; pass a callable for a predicate (it receives the stored value, or None when the key is absent).

Examples:

>>> store.find(step_type="model")
>>> store.find(accuracy=lambda a: a is not None and a > 0.9)
Source code in src/ancestree/store.py
def find(self, **filters: Any) -> list[Node]:
    """Nodes matching every filter, oldest first. Filters match
    structural attributes (step_type, generation, healthy, ...) and
    searchable metadata by equality; pass a callable for a predicate
    (it receives the stored value, or None when the key is absent).

    Examples:
        >>> store.find(step_type="model")
        >>> store.find(accuracy=lambda a: a is not None and a > 0.9)
    """
    return self._nodes_for_ids(self._metadata.find(**filters))

latest

latest(**filters)

The most recently created node matching the filters, or None.

Source code in src/ancestree/store.py
def latest(self, **filters: Any) -> Node | None:
    """The most recently created node matching the filters, or None."""
    node_id = self._metadata.most_recent(self._metadata.find(**filters))
    return None if node_id is None else self.get(node_id)

get

get(node)

Resolves a node_id (or an existing Node/handle) into a Node record. Returns None for None, "none", or an unknown id.

Source code in src/ancestree/store.py
def get(self, node: NodeLike) -> Node | None:
    """Resolves a node_id (or an existing Node/handle) into a Node
    record. Returns None for None, "none", or an unknown id."""
    if isinstance(node, Node):
        return node
    node_id = self._node_id_of(node)
    if node_id is None:
        return None
    record = self._metadata.get(node_id)
    return None if record is None else self._to_node(record)

lineage

lineage(node)

The node's full ancestry plus itself, oldest first; every node appears once, after all of its parents.

Raises:

Type Description
NodeNotFound

If the node is not in this store.

Source code in src/ancestree/store.py
def lineage(self, node: NodeLike) -> list[Node]:
    """The node's full ancestry plus itself, oldest first; every node
    appears once, after all of its parents.

    Raises:
        NodeNotFound: If the node is not in this store.
    """
    node_id = self._node_id_of(node)
    if node_id is None:
        return []
    return self._nodes_for_ids(self._metadata.lineage(node_id))

ancestors

ancestors(node, **filters)

find restricted to the node's lineage: ancestors (plus the node itself) matching every filter, in lineage order.

Source code in src/ancestree/store.py
def ancestors(self, node: NodeLike, **filters: Any) -> list[Node]:
    """``find`` restricted to the node's lineage: ancestors (plus the
    node itself) matching every filter, in lineage order."""
    node_id = self._node_id_of(node)
    if node_id is None:
        return []
    matches = set(self._metadata.find(**filters))
    return self._nodes_for_ids(
        [nid for nid in self._metadata.lineage(node_id) if nid in matches]
    )

children

children(node)

The direct children of the node (empty for unknown ids).

Source code in src/ancestree/store.py
def children(self, node: NodeLike) -> list[Node]:
    """The direct children of the node (empty for unknown ids)."""
    node_id = self._node_id_of(node)
    if node_id is None or not self._metadata.exists(node_id):
        return []
    return self._nodes_for_ids(self._metadata.children(node_id))

from_parent

from_parent(node, filename)

Shortcut for the previous step's outputs: matching artifact paths from the node's parent(s), in parent order. Empty when the node or its parents cannot be resolved.

Source code in src/ancestree/store.py
def from_parent(self, node: NodeLike, filename: str) -> list[Path]:
    """Shortcut for the previous step's outputs: matching artifact
    paths from the node's parent(s), in parent order. Empty when the
    node or its parents cannot be resolved."""
    if isinstance(node, (Node, RecordingNode)):
        parent_id: Sequence[str] = list(node.parent_id)
    else:
        record = (
            self._metadata.get(self._node_id_of(node) or "")
            if node is not None
            else None
        )
        if record is None:
            return []
        parent_id = record.parent_id
    paths: list[Path] = []
    for parent in parent_id:
        if self._metadata.exists(parent):
            paths.extend(self._artifact_paths(parent, filename))
    return paths

Visualisation

Two ways to look at a store. export_graph writes a self-contained, view-only file you can share; serve_graph serves the searchable explorer, with diffs and the runs table, on localhost.

export_graph

export_graph(dest=None, include_artifacts=True)

Renders the store into one shareable HTML file: a view-only snapshot: the lineage graph plus click-to-view metadata (search lives in the live server, so the query grammar exists once).

Small images inline as data URIs; other artifacts are copied beside the file (<name>_files/) so links work offline. Pass include_artifacts=False for a metadata-only snapshot of a store with huge artifacts. Defaults to <root>/interactive_pipeline.html; returns the written path.

Source code in src/ancestree/store.py
def export_graph(
    self,
    dest: Path | str | None = None,
    include_artifacts: bool = True,
) -> Path:
    """Renders the store into one shareable HTML file: a **view-only
    snapshot**: the lineage graph plus click-to-view metadata (search
    lives in the live server, so the query grammar exists once).

    Small images inline as data URIs; other artifacts are copied
    beside the file (``<name>_files/``) so links work offline. Pass
    ``include_artifacts=False`` for a metadata-only snapshot of a
    store with huge artifacts. Defaults to
    ``<root>/interactive_pipeline.html``; returns the written path.
    """
    return export_static(self, dest=dest, include_artifacts=include_artifacts)

serve_graph

serve_graph(port=0, block=False, open_browser=True)

Serves the searchable explorer on 127.0.0.1 and returns its URL.

By default, this method runs non-blocking (block=False) and automatically opens the live application graph in your default web browser (open_browser=True). Calling it again replaces the running background server, so re-running a notebook cell restarts the explorer rather than leaking a second one.

Parameters:

Name Type Description Default
port int

Port to bind to. Defaults to 0 (assigns an ephemeral, free port).

0
block bool

If True, blocks execution using a loop until a KeyboardInterrupt is received.

False
open_browser bool

If True, automatically launches the application in the system browser.

True

Returns:

Name Type Description
str str

The local URL running the application server.

Source code in src/ancestree/store.py
def serve_graph(
    self, port: int = 0, block: bool = False, open_browser: bool = True
) -> str:
    """Serves the searchable explorer on ``127.0.0.1`` and returns its URL.

    By default, this method runs non-blocking (``block=False``) and automatically
    opens the live application graph in your default web browser (``open_browser=True``).
    Calling it again replaces the running background server, so re-running a
    notebook cell restarts the explorer rather than leaking a second one.

    Args:
        port: Port to bind to. Defaults to 0 (assigns an ephemeral, free port).
        block: If True, blocks execution using a loop until a KeyboardInterrupt is received.
        open_browser: If True, automatically launches the application in the system browser.

    Returns:
        str: The local URL running the application server.
    """
    import webbrowser

    from .web.server import start_server

    # Re-running the call (a notebook cell) replaces the previous
    # background server instead of leaking it until close().
    previous = self._resources.get("live_server")
    if previous is not None:
        self._resources["live_server"] = None
        previous.close()

    handle = start_server(self, port=port)
    url = handle.url
    print(
        f"Ancestree explorer running at: {url}  (Store close will terminate background thread)"
    )

    # Automatically launch browser window/tab if requested
    if open_browser:
        # A short delay can sometimes be useful if your local backend
        # requires split-second initialization, though start_server handles it.
        webbrowser.open(url)

    if not block:
        self._resources["live_server"] = (
            handle  # Automatically closed when store.close() runs
        )
        return url

    try:
        while True:
            handle._thread.join(1)
    except KeyboardInterrupt:
        print("\nStopping Ancestree explorer server...")
    finally:
        handle.close()

    return url

Maintenance

Deleting, reclaiming space, and getting data out. prune defaults to a dry run. backup is the safe way to copy a store that is open; see Caveats for why copying ancestree.db by hand is not.

prune

prune(node, dry_run=True, compact=True)

Deletes a node and the descendants it solely supports, and reclaims the space they occupied.

A descendant is removed only when EVERY one of its parents is also being removed. A child still reachable from an unpruned branch survives, and its edge to the pruned parent disappears via the foreign-key cascade. Preview with the default dry_run=True, which never deletes and never compacts.

Parameters:

Name Type Description Default
node NodeLike

The node to prune.

required
dry_run bool

Preview only (the default). Nothing is deleted.

True
compact bool

Reclaim the freed chunk space afterwards (the default). Pass False when pruning many nodes in a loop and call compact() once at the end. Compaction scans the whole chunk pool, so doing it per node is wasted work.

True

Returns:

Type Description
list[Node]

The nodes that were (or would be) deleted, deepest first.

Source code in src/ancestree/store.py
def prune(
    self, node: NodeLike, dry_run: bool = True, compact: bool = True
) -> list[Node]:
    """Deletes a node and the descendants it solely supports, and
    reclaims the space they occupied.

    A descendant is removed only when EVERY one of its parents is also
    being removed. A child still reachable from an unpruned branch
    survives, and its edge to the pruned parent disappears via the
    foreign-key cascade. Preview with the default ``dry_run=True``,
    which never deletes and never compacts.

    Args:
        node: The node to prune.
        dry_run: Preview only (the default). Nothing is deleted.
        compact: Reclaim the freed chunk space afterwards (the
            default). Pass False when pruning many nodes in a loop and
            call ``compact()`` once at the end. Compaction scans the
            whole chunk pool, so doing it per node is wasted work.

    Returns:
        The nodes that were (or would be) deleted, deepest first.
    """
    node_id = self._node_id_of(node)
    if node_id is None:
        return []
    doomed = self._pruner.plan(node_id)
    nodes = self._nodes_for_ids(doomed)  # fetched before deletion
    if not dry_run and doomed:
        self._pruner.delete(doomed)
        if compact:
            self.compact()
    return nodes

compact

compact()

Reclaims space: deletes chunks no artifact references (a delta base still in use survives, the one-hop closure of AD5), then returns freed pages to the OS via incremental_vacuum and truncates the WAL.

prune calls this for you, so it is only needed directly after a batch of prune(..., compact=False) calls, or to tidy a store pruned by an older version. It does not touch the session read cache (<root>/.cache/), which is transient and cleaned up automatically when the store closes.

Returns:

Type Description
int

The number of chunks removed.

Source code in src/ancestree/store.py
def compact(self) -> int:
    """Reclaims space: deletes chunks no artifact references (a delta
    base still in use survives, the one-hop closure of AD5), then
    returns freed pages to the OS via ``incremental_vacuum`` and
    truncates the WAL.

    ``prune`` calls this for you, so it is only needed directly after a
    batch of ``prune(..., compact=False)`` calls, or to tidy a store
    pruned by an older version. It does not touch the session read
    cache (``<root>/.cache/``), which is transient and cleaned up
    automatically when the store closes.

    Returns:
        The number of chunks removed.
    """
    return compact_chunks(self._manager)

backup

backup(dest)

Writes a consistent, self-contained copy of the whole store, safe to take while the store is open and being written to.

A live store is not one file: WAL journalling keeps recent commits in ancestree.db-wal until a checkpoint, so copying ancestree.db out from under an open store silently loses everything since the last one. This goes through SQLite's online backup API, which reads through the WAL and produces a fully checkpointed single file, which is the correct way to back a store up without closing it.

Parameters:

Name Type Description Default
dest Path | str

A path ending in .db, written as that file; any other path is treated as a store root and given an ancestree.db inside it, so LineageStore(dest) opens the copy directly. An existing destination is overwritten.

required

Returns:

Type Description
Path

The path of the database file written.

Raises:

Type Description
ValueError

If the destination is this store's own database.

Source code in src/ancestree/store.py
def backup(self, dest: Path | str) -> Path:
    """Writes a consistent, self-contained copy of the whole store,
    safe to take while the store is open and being written to.

    A *live* store is not one file: WAL journalling keeps recent commits
    in ``ancestree.db-wal`` until a checkpoint, so copying
    ``ancestree.db`` out from under an open store silently loses
    everything since the last one. This goes through SQLite's online
    backup API, which reads through the WAL and produces a fully
    checkpointed single file, which is the correct way to back a store up
    without closing it.

    Args:
        dest: A path ending in ``.db``, written as that file; any other
            path is treated as a store root and given an
            ``ancestree.db`` inside it, so ``LineageStore(dest)`` opens
            the copy directly. An existing destination is overwritten.

    Returns:
        The path of the database file written.

    Raises:
        ValueError: If the destination is this store's own database.
    """
    dest_path = Path(dest)
    target = dest_path if dest_path.suffix == ".db" else dest_path / DB_FILENAME
    if target.resolve() == self._manager.db_path.resolve():
        raise ValueError(
            "Refusing to back a store up onto itself; choose a different "
            "destination."
        )
    target.parent.mkdir(parents=True, exist_ok=True)
    destination = sqlite3.connect(target)
    try:
        self._manager.read().backup(destination)
    finally:
        destination.close()
    return target

export_metadata

export_metadata(dest=None)

Writes grep-able JSON sidecars: one meta.json per node under <dest>/<node_id>/ (default <root>/export), holding the node's structural facts, provenance, metadata envelopes and artifact digests. The database remains the source of truth; this is the file-portability escape hatch (AD9), so lineage stays legible to grep even if ancestree is uninstalled tomorrow.

Returns:

Type Description
Path

The export directory.

Source code in src/ancestree/store.py
def export_metadata(self, dest: Path | str | None = None) -> Path:
    """Writes grep-able JSON sidecars: one ``meta.json`` per node under
    ``<dest>/<node_id>/`` (default ``<root>/export``), holding the
    node's structural facts, provenance, metadata envelopes and
    artifact digests. The database remains the source of truth; this
    is the file-portability escape hatch (AD9), so lineage stays
    legible to grep even if ancestree is uninstalled tomorrow.

    Returns:
        The export directory.
    """
    dest_dir = Path(dest) if dest is not None else self.root / "export"
    for node_id in self._metadata.all_node_ids():
        record = self._metadata.get(node_id)
        if record is None:
            continue
        document = {
            "node_id": record.node_id,
            "step_type": record.step_type,
            "generation": record.generation,
            "parent_id": list(record.parent_id),
            "created_utc": record.created_utc,
            "healthy": record.healthy,
            "duration_seconds": record.duration_seconds,
            "size_bytes": record.size_bytes,
            "content_hash": record.content_hash,
            "provenance": self._to_node(record).provenance,
            "metadata": self._metadata_envelopes(node_id),
            "artifacts": {
                relpath: {"size": artifact.size, "sha256": artifact.sha256}
                for relpath, artifact in sorted(
                    self._chunks.artifact_manifest(node_id).items()
                )
            },
        }
        out = dest_dir / node_id / "meta.json"
        out.parent.mkdir(parents=True, exist_ok=True)
        out.write_text(json.dumps(document, indent=2))
    return dest_dir

close

close()

Releases the store's connections and wipes the session read cache. Idempotent; reopen by constructing a new LineageStore. Runs automatically when the store is garbage collected or the interpreter exits, so calling it explicitly is optional.

Source code in src/ancestree/store.py
def close(self) -> None:
    """Releases the store's connections and wipes the session read
    cache. Idempotent; reopen by constructing a new LineageStore.
    Runs automatically when the store is garbage collected or the
    interpreter exits, so calling it explicitly is optional."""
    self._finalizer()

Introspection

The escape hatches, for questions the API above does not cover.

sql

sql(query, params=())

Runs a read-only SELECT over the documented schema (blueprint section 6) and returns the rows. The connection is opened read-only with PRAGMA query_only, so writes are impossible by construction, so the escape hatch can never corrupt invariants.

Examples:

>>> store.sql("SELECT step_type, count(*) FROM node GROUP BY 1")
Source code in src/ancestree/store.py
def sql(self, query: str, params: Sequence[Any] = ()) -> list[sqlite3.Row]:
    """Runs a read-only SELECT over the documented schema (blueprint
    section 6) and returns the rows. The connection is opened read-only
    with ``PRAGMA query_only``, so writes are impossible by
    construction, so the escape hatch can never corrupt invariants.

    Examples:
        >>> store.sql("SELECT step_type, count(*) FROM node GROUP BY 1")
    """
    conn = sqlite3.connect(f"file:{self._manager.db_path}?mode=ro", uri=True)
    try:
        conn.row_factory = sqlite3.Row
        conn.execute("PRAGMA query_only = ON")
        return conn.execute(query, tuple(params)).fetchall()
    finally:
        conn.close()

stats

stats()

Store-level numbers that make deduplication visible: node/ artifact/chunk counts, the bytes the artifacts add up to versus the bytes the chunk pool actually holds, the dedup ratio (artifact_bytes ÷ chunk_stored_bytes; higher is better) and the store's size on disk.

database_bytes counts the database file plus its write-ahead log: mid-session most recently written bytes live in the WAL, so counting only ancestree.db understates real disk usage by orders of magnitude until the next checkpoint.

Source code in src/ancestree/store.py
def stats(self) -> dict[str, Any]:
    """Store-level numbers that make deduplication visible: node/
    artifact/chunk counts, the bytes the artifacts add up to versus the
    bytes the chunk pool actually holds, the dedup ratio
    (``artifact_bytes`` ÷ ``chunk_stored_bytes``; higher is better) and
    the store's size on disk.

    ``database_bytes`` counts the database file **plus its write-ahead
    log**: mid-session most recently written bytes live in the WAL, so
    counting only ``ancestree.db`` understates real disk usage by
    orders of magnitude until the next checkpoint.
    """
    conn = self._manager.read()
    nodes = conn.execute("SELECT count(*) AS n FROM node").fetchone()["n"]
    art = conn.execute(
        "SELECT count(*) AS n, COALESCE(SUM(size), 0) AS bytes FROM artifact"
    ).fetchone()
    chunks = conn.execute(
        "SELECT count(*) AS n, COALESCE(SUM(length), 0) AS plain, "
        "COALESCE(SUM(LENGTH(data)), 0) AS stored FROM chunk"
    ).fetchone()
    stored = int(chunks["stored"])
    artifact_bytes = int(art["bytes"])
    return {
        "nodes": int(nodes),
        "artifacts": int(art["n"]),
        "chunks": int(chunks["n"]),
        "artifact_bytes": artifact_bytes,
        "chunk_plain_bytes": int(chunks["plain"]),
        "chunk_stored_bytes": stored,
        "database_bytes": self._database_bytes(),
        "dedup_ratio": round(artifact_bytes / stored, 3) if stored else None,
    }

Working with nodes

You never construct a node yourself. LineageStore.create_node yields a recording handle, which is the object you write artifacts and metadata through inside the with block. The store's search and lineage methods return immutable Node records for everything already persisted.

ancestree.Node dataclass

One persisted pipeline step: an immutable, hashable record.

Structural facts are attributes; metadata holds the user's entries; artifacts() and / return readable paths (reassembled on demand from the store; a node is a database row, not a directory).

Source code in src/ancestree/domain/node.py
@dataclass(frozen=True)
class Node:
    """One persisted pipeline step: an immutable, hashable record.

    Structural facts are attributes; ``metadata`` holds the user's entries;
    ``artifacts()`` and ``/`` return readable paths (reassembled on demand
    from the store; a node is a database row, not a directory).
    """

    node_id: str
    step_type: str = field(compare=False)
    generation: int = field(compare=False)
    #: Ordered parent ids — a tuple, since a node may be a join.
    parent_id: tuple[str, ...] = field(compare=False)
    created_utc: str = field(compare=False)
    healthy: bool = field(compare=False)
    duration_seconds: float | None = field(compare=False)
    size_bytes: int = field(compare=False)
    content_hash: str | None = field(compare=False, repr=False)
    _provenance: dict[str, Any] = field(compare=False, repr=False)
    _store: LineageStore = field(compare=False, repr=False)

    @property
    def metadata(self) -> dict[str, dict[str, Any]]:
        """The node's user metadata: each key maps to its envelope
        ``{'value', 'data_type', 'group', 'searchable'}``. Structural facts
        (step_type, generation, ...) are attributes, not metadata."""
        return self._store._metadata_envelopes(self.node_id)

    @property
    def provenance(self) -> dict[str, Any]:
        """Who/what/how produced this node: user, python_version, platform,
        git_commit, git_dirty, git_branch."""
        return dict(self._provenance)

    def artifacts(self, contains: str = "*") -> list[Path]:
        """The node's artifact files as readable paths, reassembled from
        the store on demand (served from the session read cache).

        Args:
            contains: A glob pattern, or a plain substring matched
                case-insensitively anywhere in the filename. Defaults to
                all files.
        """
        return self._store._artifact_paths(self.node_id, contains)

    def __truediv__(self, relative: str | Path) -> Path:
        """Read-side ``/``: a readable path for one artifact.

        Raises:
            ArtifactNotFound: If the node has no artifact at that path.
        """
        return self._store._artifact_path(self.node_id, Path(relative).as_posix())

    def __repr__(self) -> str:
        return (
            f"Node(node_id={self.node_id!r}, step_type={self.step_type!r}, "
            f"generation={self.generation})"
        )

metadata property

The node's user metadata: each key maps to its envelope {'value', 'data_type', 'group', 'searchable'}. Structural facts (step_type, generation, ...) are attributes, not metadata.

provenance property

Who/what/how produced this node: user, python_version, platform, git_commit, git_dirty, git_branch.

artifacts(contains='*')

The node's artifact files as readable paths, reassembled from the store on demand (served from the session read cache).

Parameters:

Name Type Description Default
contains str

A glob pattern, or a plain substring matched case-insensitively anywhere in the filename. Defaults to all files.

'*'
Source code in src/ancestree/domain/node.py
def artifacts(self, contains: str = "*") -> list[Path]:
    """The node's artifact files as readable paths, reassembled from
    the store on demand (served from the session read cache).

    Args:
        contains: A glob pattern, or a plain substring matched
            case-insensitively anywhere in the filename. Defaults to
            all files.
    """
    return self._store._artifact_paths(self.node_id, contains)

__truediv__(relative)

Read-side /: a readable path for one artifact.

Raises:

Type Description
ArtifactNotFound

If the node has no artifact at that path.

Source code in src/ancestree/domain/node.py
def __truediv__(self, relative: str | Path) -> Path:
    """Read-side ``/``: a readable path for one artifact.

    Raises:
        ArtifactNotFound: If the node has no artifact at that path.
    """
    return self._store._artifact_path(self.node_id, Path(relative).as_posix())

ancestree.domain.node.RecordingNode

The mutable handle a create_node block writes through.

node / "file.csv" returns a real, ready-to-write path inside the node's transient scratch directory; add_meta records metadata to be persisted at block exit. After the block closes the handle stays valid as a parent reference (it carries the node_id).

Source code in src/ancestree/domain/node.py
class RecordingNode:
    """The mutable handle a ``create_node`` block writes through.

    ``node / "file.csv"`` returns a real, ready-to-write path inside the
    node's transient scratch directory; ``add_meta`` records metadata to be
    persisted at block exit. After the block closes the handle stays valid
    as a parent reference (it carries the node_id).
    """

    def __init__(
        self,
        node_id: str,
        step_type: str,
        generation: int,
        parent_id: list[str],
        workspace: NodeWorkspace,
    ) -> None:
        self.node_id = node_id
        self.step_type = step_type
        self.generation = generation
        self.parent_id = list(parent_id)
        self._workspace = workspace
        self._entries: dict[str, PreparedEntry] = {}

    def __truediv__(self, relative: str | Path) -> Path:
        """Write-side ``/``: a ready-to-write path inside the node's
        scratch directory (intermediate directories are created).

        Raises:
            ValueError: If the path escapes the node's directory.
        """
        return self._workspace.resolve(relative)

    def add_meta(
        self,
        key: str,
        value: Any,
        group: str | None = "General",
        data_type: str = "auto",
        searchable: bool = True,
    ) -> None:
        """Attaches a piece of metadata to the node.

        Searchable entries can be matched by the store's query methods;
        everything shows in the web graph. Adding an existing key
        overwrites the previous entry.

        Args:
            key: The name of the metadata entry.
            value: The value to store. Must be JSON-serialisable (numpy/
                pandas values are coerced with a warning).
            group: A heading to group related entries under in the web
                graph. Defaults to "General".
            data_type: How the value renders: 'auto' (infer), 'image',
                'link', 'table' (DataFrame), 'json', 'code', or 'text'.
            searchable: False for display-only metadata.
        """
        entry = prepare_entry(
            key, value, group=group, data_type=data_type, searchable=searchable
        )
        self._entries[key] = self._normalise_artifact_reference(entry)

    def _normalise_artifact_reference(self, entry: PreparedEntry) -> PreparedEntry:
        """Rewrites an image/link value that points into this node's
        scratch directory to the node-relative artifact path, so the
        reference stays valid once the scratch is gone and the bytes live
        in the store. URLs pass through untouched."""
        if entry.data_type not in ("image", "link"):
            return entry
        text = str(entry.value)
        if text.startswith(("http://", "https://")):
            return entry
        scratch = self._workspace.path.resolve()
        candidate = Path(text)
        if candidate.is_absolute():
            try:
                return replace(
                    entry, value=candidate.resolve().relative_to(scratch).as_posix()
                )
            except ValueError:
                return replace(entry, value=candidate.as_posix())
        return replace(entry, value=candidate.as_posix())

    def artifacts(self, contains: str = "*") -> list[Path]:
        """The files written so far, as native scratch paths, matching the
        record's ``artifacts()`` semantics so code inside the block reads
        what it just wrote at native speed."""
        files = dict(self._workspace.files())
        return [files[rel] for rel in filter_relpaths(sorted(files), contains)]

    def __repr__(self) -> str:
        return (
            f"RecordingNode(node_id={self.node_id!r}, "
            f"step_type={self.step_type!r}, generation={self.generation})"
        )

add_meta(key, value, group='General', data_type='auto', searchable=True)

Attaches a piece of metadata to the node.

Searchable entries can be matched by the store's query methods; everything shows in the web graph. Adding an existing key overwrites the previous entry.

Parameters:

Name Type Description Default
key str

The name of the metadata entry.

required
value Any

The value to store. Must be JSON-serialisable (numpy/ pandas values are coerced with a warning).

required
group str | None

A heading to group related entries under in the web graph. Defaults to "General".

'General'
data_type str

How the value renders: 'auto' (infer), 'image', 'link', 'table' (DataFrame), 'json', 'code', or 'text'.

'auto'
searchable bool

False for display-only metadata.

True
Source code in src/ancestree/domain/node.py
def add_meta(
    self,
    key: str,
    value: Any,
    group: str | None = "General",
    data_type: str = "auto",
    searchable: bool = True,
) -> None:
    """Attaches a piece of metadata to the node.

    Searchable entries can be matched by the store's query methods;
    everything shows in the web graph. Adding an existing key
    overwrites the previous entry.

    Args:
        key: The name of the metadata entry.
        value: The value to store. Must be JSON-serialisable (numpy/
            pandas values are coerced with a warning).
        group: A heading to group related entries under in the web
            graph. Defaults to "General".
        data_type: How the value renders: 'auto' (infer), 'image',
            'link', 'table' (DataFrame), 'json', 'code', or 'text'.
        searchable: False for display-only metadata.
    """
    entry = prepare_entry(
        key, value, group=group, data_type=data_type, searchable=searchable
    )
    self._entries[key] = self._normalise_artifact_reference(entry)

artifacts(contains='*')

The files written so far, as native scratch paths, matching the record's artifacts() semantics so code inside the block reads what it just wrote at native speed.

Source code in src/ancestree/domain/node.py
def artifacts(self, contains: str = "*") -> list[Path]:
    """The files written so far, as native scratch paths, matching the
    record's ``artifacts()`` semantics so code inside the block reads
    what it just wrote at native speed."""
    files = dict(self._workspace.files())
    return [files[rel] for rel in filter_relpaths(sorted(files), contains)]

__truediv__(relative)

Write-side /: a ready-to-write path inside the node's scratch directory (intermediate directories are created).

Raises:

Type Description
ValueError

If the path escapes the node's directory.

Source code in src/ancestree/domain/node.py
def __truediv__(self, relative: str | Path) -> Path:
    """Write-side ``/``: a ready-to-write path inside the node's
    scratch directory (intermediate directories are created).

    Raises:
        ValueError: If the path escapes the node's directory.
    """
    return self._workspace.resolve(relative)