Getting started with ancestree¶
ancestree records the lineage of the artifacts a pipeline produces: every step is
saved as a node that knows which node it came from, what files it wrote and what
metadata describes it.
This notebook covers the essentials — defining a ruleset, creating a store, writing a node, looking it back up and rendering the result — using nothing more than a single synthetic array.
1. Imports¶
import shutil
import tempfile
from pathlib import Path
import numpy as np
import pandas as pd
import ancestree
# The store lives in a temp folder here so the executed notebook leaves no
# litter behind; point it at any real directory in your own work.
WORKDIR = Path(tempfile.mkdtemp(prefix="ancestree-basics-"))
2. Define a ruleset¶
The ruleset describes how a pipeline is allowed to progress. Each key is a step type and
its value lists the step types it may follow, where [None] marks a valid starting point.
Separately, gen_triggers names the step(s) that begin a new generation — a fresh
run of the pipeline.
You only need to declare this once: it is persisted inside the store's database and read back automatically whenever the store is reopened by path.
RULES = {
"import data": [None], # a pipeline starts by importing data
}
3. Create the store¶
The store lives in a folder of its own. Its name can be a str or a Path. Passing the
rules and gen_triggers here persists the configuration; from then on
ancestree.LineageStore(WORKDIR / "basic_usage") is enough to reopen it. The entire
store — metadata, lineage and artifact bytes — is the single ancestree.db
file inside that folder.
store = ancestree.LineageStore(WORKDIR / "basic_usage", rules=RULES)
4. Write a node¶
store.create_node(...) opens a node as a context manager. Inside the block, writing
files uses familiar syntax — the store handles pathing via node / "filename" —
and node.add_meta(...) attaches searchable, displayable metadata. Here a pandas DataFrame
is stored as metadata (it renders as a table in the web graph) alongside a saved array.
with store.create_node(step_type="import data") as node:
# Save a file
X = np.random.randint(0, 100, (5, 5))
np.save(node / "samples.npy", X)
# Add some metadata
node.add_meta("something", pd.DataFrame({"key": [1, 2, 3], "other": [4, 5, 6]}))
5. Look the node back up¶
latest() returns the last node written to the store (it was
get_most_recent_node() in 0.1.x).
n = store.latest()
n
Node(node_id='8f4dc169', step_type='import data', generation=0)
6. Visualise the store¶
export_graph() writes a self-contained, interactive HTML page showing every node,
its lineage and its metadata.
store.export_graph()
PosixPath('/var/folders/xf/n_m7ztrx4x577r3935n_1m1w0000gn/T/ancestree-basics-ow0hvdfl/basic_usage/interactive_pipeline.html')
# No store.close() needed — the store releases its connections and read cache
# automatically when the kernel exits (or when it is garbage-collected). We only
# remove the temp folder here so the executed notebook leaves nothing behind.
shutil.rmtree(WORKDIR, ignore_errors=True)