Skip to content

CU21 — The diagnosis, which says it is an opinion

from somatize.health import diagnose, overlaid, alerts
from somatize.torch import Trainer, Audit
t = Trainer(g, objective=..., optimizer=..., auditing=Audit(every=10, inside=True),
watching=[Recorder(store, summarising=["loss"])])
t.fit(data, epochs=5)
found = diagnose(store, run=t.run) # read back, nothing runs again
alerts(found) # the loud one, cards in a cell
overlaid(g, store, run=t.run, inside=True) # and where, on the graph

The third of the three things CU19 split observability into. The first was the declaration drawn, the second the record of what happened, and this one is neither: it is an opinion about the record.

A crate with no dependencies at all, not even the core’s

Section titled “A crate with no dependencies at all, not even the core’s”

health/ takes numbers and gives back flags. It does not measure, it has no clock, it never touches a store. That is what turns CU19’s invariant from an aspiration into a test:

a diagnosis has to be reproducible from the stored record, without training again.

Change a bound and ask again. The record has not moved, so an argument about a threshold costs a scan instead of an afternoon of GPU. It is the shape study/ already had: pure, deterministic, hashable, and the loop that owns a tensor stays in Python.

The taxonomy is inherited, and how it reads is the knowledge

Section titled “The taxonomy is inherited, and how it reads is the knowledge”

DEAD and SATURATED read the maximum over a window and never the mean. A layer that dies one step in four is dead, and an average hides it. Dormant is not dead — two findings, not one bound with two names.

Three come from the literature. STALLED and OVERSTEPPING from the update-to-weight ratio, which the original measured and never said anything about, and which lands a healthy layer at about 1e-3. LOSING_PLASTICITY is a conjunction on purpose: weights growing, or units going quiet, one at a time, is a network that is training.

A threshold that was measured and did not survive it

Section titled “A threshold that was measured and did not survive it”

NARROWING is in the vocabulary and off by default. The published monitor’s certificate is the deviation from a healthy baseline, and one run has none. Measured: healthy runs sit at 0.69–0.71 and a destabilised one at 0.43–0.86, which overlap. The measurement is in soma-health/tests/narrowing.py, the metric is recorded and drawn, and the alarm was not invented. Thresholds is data, so whoever does have a baseline sets the bound and gets the finding.

Measuring is level 2’s, and thresholds never go near it

Section titled “Measuring is level 2’s, and thresholds never go near it”

Trainer(..., auditing=True) hooks the nodes and emits health facts through the same watching= CU20 built. A threshold baked into the measurement would make disagreeing with it cost another training run.

Audit(inside=True) looks inside a node, because a node is often a whole architecture and this node is unhealthy is not an answer when it is twenty layers. Findings are keyed node.path.to.submodule, and the audit’s scope is the same scope the drawing uses — which is what makes what is measured has a box true rather than hopeful.

A node is opened up, and an architecture is a graph

Section titled “A node is opened up, and an architecture is a graph”

architecture(g, x) traces what a node is made of: fx where it can, because it sees the operations that are not modules and a residual connection is exactly one; a real forward where it cannot, saying so, because a residual that is missing looks like a residual that is not there.

g.figure(inside=...) draws the node’s box as a frame — the shape a Wave and a Remote already are — and lays the inside out by what feeds what, so a skip runs down a gutter and enters from the side. The rules that make it readable: a kind decides the silhouette, not a class name; a composite everybody recognises is one box and depth= opens it; blocks that are the same block collapse to ×N; the shape is written on the layer, because that is the only thing that makes a bottleneck a picture; and every number says what it is4 batch · 16 steps · 24 dim, never 4×16×24.

Two of those grew after CU22 — see After CU22 — a block is a box.

Findings are coloured by family — numeric, signal, activation, step, capacity, data — with a legend of the ones on the figure. Six alarms that all look the same are one alarm.

Health gets a channel of its own: the fill goes on saying where a node runs and the outline turns red. On a graph spread over three machines, where does this run is the answer somebody came for, and taking the fill for a second fact would have cost it.

gantt is the timeline. Every fact carries how far into the forward it began, so a Wave draws as overlapping bars and a remote slice sits inside the round trip it arrived under. An offset into a slice is a fact about the slice; two wall clocks would not have composed.

And a third question, which is not about the network at all

Section titled “And a third question, which is not about the network at all”

somatize.data.contribution shuffles one input and scores again; the drop is what that input was worth. health asks whether a network is learning; this asks whether it is learning what you meant, which no amount of looking at a gradient will ever say.

It exists because of a real project: symptom channels for a mental-health condition, months spent on the architecture, and the predictive signal was in the self-disclosure and not in the presence of symptoms. IGNORED_INPUT is the finding that would have said so in an afternoon; SOLE_RELIANCE is the other end of the same worry.

Shuffled and not zeroed: a zero is a value, and what is being asked about is the correspondence with the answer.

The static half, before a GPU is spent: signal propagation and dynamical isometry at initialisation, where a normalisation layer is missing, and the zero-cost proxies. Deferred rather than refused, and with a caveat already written down — synflow correlates 0.76 with parameter count, which is close to saying it measures size. If it does not separate when measured, it ships off with its measurement beside it, the way NARROWING did. (CU22, and that is exactly what happened to it.)

The verdict (soma-health/tests/unit/verdict.rs)

  • a gradient too small to train on says so, and one too big to step on, and it cannot be both
  • a layer that dies one step in four is dead, and one that is merely sparse is not
  • a layer pinned where the derivative is nothing says so
  • a node moving too little next to its own weights says so, and one moving so much it forgets where it was
  • a healthy ratio is near a thousandth and says nothing
  • dead channels are counted and are not the same as a dead layer
  • a channel alive and never asked for is its own finding
  • two groups carrying the same information leak
  • a collapsing update says nothing at the default bound, because it is off — and says so for whoever has a baseline to set it against
  • an update that is merely low rank all along is not narrowing
  • losing plasticity needs all three signs at once, and any one alone is ordinary
  • what stops a run is read first
  • the same numbers answer differently under other thresholds
  • a node nobody measured is not called healthy and is not flagged

The vocabulary (soma-health/tests/unit/flag.rs)

  • a flag that counts something says how many, and its name is stable whatever it counts
  • every flag says what to do about it

The data (soma-health/tests/unit/leaning.rs)

  • shares add up to one, so they read as how much of what matters
  • an input the model is not using says so
  • two inputs that share the work say nothing
  • one input carrying everything is worth knowing before it goes missing
  • a model that loses nothing whatever you take away is using none of it
  • an input the model does better without keeps its negative
  • one input alone says nothing, because there is nothing to compare
  • the bounds are data here too

Measured and read back (soma-python/tests/test_health.py)

  • a diagnosis is taken from the record and not from the run
  • the same record answers differently under other thresholds
  • a threshold nobody has is refused by name
  • a deep sigmoid stack starves its early layers
  • the update ratio lands a healthy layer near a thousandth
  • a block whose relu cuts everything off is dead, and one pinned at the far end of its range is saturated
  • a gradient too big to step on explodes
  • a healthy shallow stack raises nothing
  • a node with no weights is not diagnosed at all
  • a run that is not audited says nothing about health
  • a cadence measures fewer steps and says the same kind of thing
  • auditing does not change what the network computes

What a node is made of (soma-python/tests/test_health.py)

  • a node says what it is made of
  • a skip connection is an edge and not an order
  • a bottleneck is visible in the shapes
  • a module fx cannot trace is still drawn, and says how
  • what is drawn is a superset of what is measured
  • a composite everybody recognises is one box, and depth counts composites opened and not names
  • blocks that are the same block collapse to one and a count
  • what comes after a stack is not adopted by its last block
  • a tensor nobody holds cannot invent an edge
  • a shape says what each of its numbers is, and something that did not change the shape keeps the names
  • a recurrent cell says its output and not its hidden state

What was measured survives the steps that did not measure it (soma-python/tests/test_health.py)

  • a number taken on the snapshot cadence is still in seen after steps that did not take it

Half of what an audit measures is an SVD and runs every snapshot steps — eff_rank, group_cka, update_rank. seen promises the latest of each, and the first draft read that as the latest fact, replacing the whole dict on every step: so a run that did not happen to end on a snapshot step dropped all of them, and LEAKAGE and NARROWING could fire only by arithmetic luck. The fix is one word — merge rather than replace — and what found it was writing an example whose entire point was a leaking pair of branches, which is the third real bug the examples have paid for.

Drawn (soma-python/tests/test_figure.py)

  • a node with an inside becomes a frame around it, and one without is drawn exactly as before
  • a layer is drawn by what it is and not by its name
  • a layer that narrows is drawn wide at the top
  • what feeds what decides the rows, and a skip jumps one
  • every layer sits inside the box it belongs to, named after the node it is in
  • the shape is written on the layer, because a bottleneck is shapes
  • a flag on a layer marks that layer and not the others
  • the overlay marks the node without taking the fill
  • a branch of a wave can be opened too

The data layer (soma-python/tests/test_data.py)

  • an input the model is not using is found in one afternoon
  • and the shares say how lopsided it is
  • two channels that both carry it say nothing
  • nothing is trained and nothing is changed
  • shuffling keeps the channel and breaks only what it lines up with
  • an opaque is unwrapped and wrapped again
  • something that is not a batch is left alone
  • only the inputs asked for are tried
  • no data is nothing and not a failure