Tracking an ML experiment with ancestree¶
This notebook walks through a small but complete machine-learning workflow and uses
ancestree to record the lineage of every artifact it produces.
We take the classic Iris dataset and fan it out across a grid of modelling choices:
- two scalers —
RobustScalerandStandardScaler - three embeddings — UMAP, t-SNE and PCA
- two clusterers — K-Means and HDBSCAN
That is 2 × 3 × 2 = 12 clustering results, each with its own provenance. Instead of
juggling file names by hand, we let ancestree capture how each node was derived from
its parent, so the whole experiment stays queryable once it has finished running.
1. Imports¶
The usual scientific-Python stack, plus ancestree itself.
import tempfile
from pathlib import Path
import matplotlib
import numpy as np
import pandas as pd
matplotlib.use("Agg") # headless: the figures go into the store, not a window
import matplotlib.pyplot as plt
import umap
from sklearn.cluster import HDBSCAN, KMeans
from sklearn.datasets import load_iris
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.metrics import calinski_harabasz_score
from sklearn.preprocessing import RobustScaler, StandardScaler
import ancestree
WORKDIR = Path(tempfile.mkdtemp(prefix="ancestree-ml-"))
2. Configure the lineage store¶
A LineageStore is the on-disk record of the experiment. Two settings shape it:
rulesdeclare which step types may follow which. Each key is a step type and its value lists the permitted parents, where[None]marks a valid pipeline root. This keeps provenance honest — aclusteringnode can only ever hang off anembedding.gen_triggersname the step types that begin a new generation, which is howancestreegroups successive runs of the pipeline.
The configuration is persisted inside the store's database, so it is picked up
automatically next time — reopening with ancestree.LineageStore(...) on the same
path is enough. dedup and chunk are persisted the same way now: set them once at
creation and every later open uses the stored policy.
rules = {
"raw data": [None], # a pipeline starts from raw data
"scaling": ["raw data"], # scaling follows the raw data
"embedding": ["scaling"], # embeddings are computed on scaled data
"clustering": ["embedding"], # clusters are found in the embedding space
}
triggers = ["scaling", "embedding", "clustering"]
store = ancestree.LineageStore(
WORKDIR / "demo", rules=rules, gen_triggers=triggers, reuse_identical=True, delta=True
)
3. Run the pipeline¶
Each step is wrapped in a store.create_node(...) context manager. Inside the block we
write artifacts straight to the node — the store resolves paths for us via
node / "file" — and attach searchable metadata with node.add_meta(...). Passing
parent= links a node to the one it was derived from, so the nested loops below build the
full provenance tree as a side effect of simply running the experiment.
scalers = {"robust": RobustScaler(), "standard": StandardScaler()}
embedders = ["umap", "tsne", "pca"]
clusterers = ["kmeans", "hdbscan"]
scale_nodes = {}
embed_nodes = {}
# Raw data ingestion
with store.create_node("raw data") as root_node:
iris = load_iris()
df_raw = pd.DataFrame(iris.data, columns=iris.feature_names)
df_raw.to_csv(root_node / "data.csv", index=False)
# Two different scaling steps
for s_name, s_obj in scalers.items():
with store.create_node("scaling", parent=root_node) as s_node:
scaled_data = s_obj.fit_transform(df_raw)
pd.DataFrame(scaled_data).to_csv(s_node / "scaled.csv", index=False)
s_node.add_meta("scaler", s_name, group="Config")
scale_nodes[s_name] = (s_node, s_obj.fit_transform(df_raw))
# Three different dimensionality reductions steps
for s_name, (s_node, scaled_data) in scale_nodes.items():
for e_type in embedders:
with store.create_node("embedding", parent=s_node) as e_node:
if e_type == "pca":
embedder = PCA(n_components=2)
elif e_type == "tsne":
embedder = TSNE(n_components=2, perplexity=30)
elif e_type == "umap":
embedder = umap.UMAP(n_components=2)
embedding = embedder.fit_transform(scaled_data)
np.save(e_node / "map.npy", embedding)
e_node.add_meta("method", e_type, group="Config")
e_node.add_meta("scaler", s_name, group="Config")
embed_nodes[(s_name, e_type)] = (e_node, embedding)
# Two different clustering methods
for (s_name, e_type), (e_node, embedding) in embed_nodes.items():
for c_type in clusterers:
with store.create_node("clustering", parent=e_node) as c_node:
if c_type == "kmeans":
model = KMeans(n_clusters=3, n_init="auto")
elif c_type == "hdbscan":
model = HDBSCAN(min_cluster_size=5)
labels = model.fit_predict(embedding)
score = calinski_harabasz_score(embedding, labels)
c_node.add_meta("clusterer", c_type, group="Config")
c_node.add_meta("CH_score", round(score, 3), group="Metrics")
fig, ax = plt.subplots(figsize=(6, 4))
ax.scatter(embedding[:, 0], embedding[:, 1], c=labels, cmap="turbo")
ax.set_title(f"{s_name} → {e_type} → {c_type}")
fig.savefig(c_node / "clustering.png", bbox_inches="tight")
plt.close(fig)
c_node.add_meta("cluster_plot", c_node / "clustering.png", group="Results")
/Users/js/coding-projects/ancestree/.venv/lib/python3.12/site-packages/sklearn/cluster/_hdbscan/hdbscan.py:722: FutureWarning: The default value of `copy` will change from False to True in 1.10. Explicitly set a value for `copy` to silence this warning. warn( /Users/js/coding-projects/ancestree/.venv/lib/python3.12/site-packages/sklearn/cluster/_hdbscan/hdbscan.py:722: FutureWarning: The default value of `copy` will change from False to True in 1.10. Explicitly set a value for `copy` to silence this warning. warn( /Users/js/coding-projects/ancestree/.venv/lib/python3.12/site-packages/sklearn/cluster/_hdbscan/hdbscan.py:722: FutureWarning: The default value of `copy` will change from False to True in 1.10. Explicitly set a value for `copy` to silence this warning. warn( /Users/js/coding-projects/ancestree/.venv/lib/python3.12/site-packages/sklearn/cluster/_hdbscan/hdbscan.py:722: FutureWarning: The default value of `copy` will change from False to True in 1.10. Explicitly set a value for `copy` to silence this warning. warn( /Users/js/coding-projects/ancestree/.venv/lib/python3.12/site-packages/sklearn/cluster/_hdbscan/hdbscan.py:722: FutureWarning: The default value of `copy` will change from False to True in 1.10. Explicitly set a value for `copy` to silence this warning. warn( /Users/js/coding-projects/ancestree/.venv/lib/python3.12/site-packages/sklearn/cluster/_hdbscan/hdbscan.py:722: FutureWarning: The default value of `copy` will change from False to True in 1.10. Explicitly set a value for `copy` to silence this warning. warn(
4. Visualise the lineage¶
export_graph() renders the whole tree — nodes, parent/child edges, metadata
and embedded figures — as a self-contained, interactive HTML page. For searching and
diffing there is also the live explorer: store.serve_graph().
store.export_graph()
PosixPath('/var/folders/xf/n_m7ztrx4x577r3935n_1m1w0000gn/T/ancestree-ml-fk9vo3t2/demo/interactive_pipeline.html')
5. Querying the experiment¶
Because every node carries searchable metadata, we can interrogate the experiment after
it has run. find (find_node in 0.1.x) accepts a callable for any field to express a
predicate — here we pull back every clustering result whose Calinski–Harabasz score
came out below 1000.
# Pass a lambda for richer queries than exact matching
bad_nodes = store.find(CH_score=lambda x: x is not None and x < 1000)
bad_nodes
[Node(node_id='949e411b', step_type='clustering', generation=3), Node(node_id='a74b7b61', step_type='clustering', generation=3), Node(node_id='2cb055cb', step_type='clustering', generation=3), Node(node_id='f03cccbc', step_type='clustering', generation=3), Node(node_id='7dd3a3c0', step_type='clustering', generation=3), Node(node_id='01d117fe', step_type='clustering', generation=3), Node(node_id='d13f1484', step_type='clustering', generation=3)]
Plain keyword arguments match exactly, and several can be combined to narrow the search. Unpacking a dict keeps things tidy when the query is assembled programmatically.
# Chain search terms; supply them with dictionary unpacking
search_vars = {"method": "pca", "step_type": "embedding"}
nodes_of_interest = store.find(**search_vars)
nodes_of_interest
[Node(node_id='3aa55ac3', step_type='embedding', generation=2), Node(node_id='9a1c5309', step_type='embedding', generation=2)]
6. Walking a single lineage¶
latest() returns the last node written, and lineage() walks back up its ancestry to
the root (get_most_recent_node() and get_lineage() in 0.1.x). Below we print the
artifacts recorded at each step, from the raw data all the way down to the final
clustering image.
last_node = store.latest()
nodes_in_lineage = store.lineage(last_node)
for n in nodes_in_lineage:
print(n.artifacts())
[PosixPath('/var/folders/xf/n_m7ztrx4x577r3935n_1m1w0000gn/T/ancestree-ml-fk9vo3t2/demo/.cache/36853-8ee6aed5/17f7db84/data.csv')]
[PosixPath('/var/folders/xf/n_m7ztrx4x577r3935n_1m1w0000gn/T/ancestree-ml-fk9vo3t2/demo/.cache/36853-8ee6aed5/b65a8bd3/scaled.csv')]
[PosixPath('/var/folders/xf/n_m7ztrx4x577r3935n_1m1w0000gn/T/ancestree-ml-fk9vo3t2/demo/.cache/36853-8ee6aed5/9a1c5309/map.npy')]
[PosixPath('/var/folders/xf/n_m7ztrx4x577r3935n_1m1w0000gn/T/ancestree-ml-fk9vo3t2/demo/.cache/36853-8ee6aed5/d13f1484/clustering.png')]
Reference: the rest of the store API¶
A quick tour of the other methods for when you come back to an existing store.
# Reopen an existing store (config is read back from the database automatically)
# store = ancestree.LineageStore(WORKDIR / "demo")
# Search
# store.find(step_type="embedding")
# store.ancestors(node="0d33f882", step_type="embedding") # was find_in_lineage
# store.latest() # was get_most_recent_node
# Fetch by id and navigate the tree
# store.get(node="0d33f882") # was get_node
# store.lineage(node="0d33f882") # was get_lineage
# store.children(node="0d33f882") # was get_child_nodes
# Pull a parent's artifact into the current step
# store.from_parent(node="0d33f882", filename="*.csv")
# Ask the database anything (read-only SQL over a documented schema)
# store.sql("SELECT step_type, count(*) FROM node GROUP BY 1")
# Maintenance: preview a deletion, then reclaim the space
# store.prune("0d33f882") # dry run by default