Ancestree
Data lineage tracking for Python, built on the standard library. Track each step of a pipeline, enforce valid transitions, and explore the result as an interactive graph.
Features
-
Interactive graphs
One call renders the pipeline as a self-contained HTML file. Open it in any browser, share it as-is, click a node for its metadata and artifacts.
-
Rule enforcement
Rules are optional, but once declared, invalid transitions raise immediately rather than being recorded after the fact.
-
Metadata does double duty
Metadata is searchable by value or predicate, and also decides how each entry renders in the explorer.
-
No dependencies
Pure standard library. Nothing to pin, nothing to conflict with, runs wherever Python 3.9+ runs.
-
Crash-safe by design
Nodes are created in a context manager. If the code fails, partial work is kept and flagged unhealthy; untouched nodes are discarded.
-
One SQLite file
Metadata, lineage and deduplicated artifact bytes all sit in a single
ancestree.db. No server, nothing to configure. Back it up by copying one file, or query it withstore.sql(...).
Quick Start
How it works
A LineageStore is a directory holding one SQLite database. Every node is a row, with its artifacts stored as deduplicated, content-addressed chunks alongside its metadata and lineage.
my_store/
├── ancestree.db # the entire store: nodes, metadata, artifacts
├── interactive_pipeline.html # generated web graph (optional snapshot)
├── .scratch/ # a node's files, only while its block runs
└── .cache/ # artifacts reassembled for reading, per session
Only ancestree.db holds anything you cannot regenerate. The dotted directories are working space: .scratch/ holds a node's files while its with block runs and empties when the node commits, .cache/ holds artifacts reassembled for reading and clears when the session ends. Deleting either at rest costs nothing. While a store is open SQLite also keeps ancestree.db-wal and -shm beside the database, so read Caveats before backing one up.
Every write is a transaction, so a node is committed whole or not at all. prune() deletes a branch and reclaims the space, export() writes meta.json sidecars, and store.sql(...) queries a documented schema directly.
Track, search and visualise
Metadata does double duty
Metadata is both a search index and the instruction set for how a node displays in the web graph. Each entry appears under its group heading, and data_type controls how the value renders. It defaults to auto, which infers the type from the value, and can be overridden.
with store.create_node(step_type="model", parent=parent) as node:
fig.savefig(node / "confusion.png")
node.add_meta("accuracy", 0.94, group="Metrics") # searchable, shown as text
node.add_meta(
"confusion_matrix",
node / "confusion.png", # rendered inline as a figure
data_type="auto",
group="Figures",
)
node.add_meta(
"notes",
"rerun after fix", # display-only, excluded from search
searchable=False,
)
Metadata is not needed to expose files: every artifact appears as a clickable link under the node's Artifacts heading. Use data_type="image" to display a figure inline.
Next steps
- Work through the Examples, including a machine learning workflow.
- See the API Reference for
LineageStoreandNode. - Read the Benchmarks for what deduplication saves and what each operation costs.