The Learning Algorithm of a Holon is Natural Gradient
Vanchurin's programme derives physical and biological law from learning dynamics. It leaves one central quantity undetermined: the functional relation between the metric tensor on trainable space and the covariance of gradient noise, parametrised as a power law with three named regimes — (stochastic gradient), (efficient learning, Adam-like, conjectured to underlie biological complexity) and (natural gradient, the "quantum regime"). He states that direct empirical estimates of remain unavailable and calls obtaining them a formidable challenge.
In UHM this quantity is not free. Both factors are forced by the axioms — by Axiom 2 (Bures is the unique monotone Riemannian metric), by the canonical dissipator whose Lindblad operators are the atoms of the subobject classifier plus the seven Fano projectors. Their relation is therefore a theorem, not a parameter:
on the population sector of a holon of dimension , where is the projector onto the tangent space. This is verbatim Vanchurin's own criterion for : substituting into his Eq. (7.5) gives . Hence
Equation numbers throughout refer to:
- [V1] V. Vanchurin, Geometric framework for biological evolution, arXiv:2603.15198 — Eqs. (4.7), (6.3), (6.7)–(6.9), (7.1)–(7.6). This is the source of the programme in the form used here.
- [V2] V. Vanchurin, Geometric Learning Dynamics, arXiv:2504.14728, Biological Cybernetics (2026) — Eqs. (2.2), (2.8), (2.9), (5.5).
Two notational hazards, stated once. [V1] uses for the learning exponent and as a coordinate index; [V2] uses for the exponent. We follow [V1] and write . And his is a learning rate; our is the dissipator rate — they are unrelated, and every on this page is ours.
Four further exact results follow. The canonical Fano layer multiplies the metric-weighted noise trace by the universal rational factor — independent of the state — so that the block layer carries exactly of the noise, and it introduces a computable anisotropy. Noise magnitude is tied to purity, . The dissipator annihilates every diagonal state, so on populations it is pure noise without drift — which is what makes a genuinely centred covariance in the sense of his (7.4). And there is a strict sector split: on the decohered manifold, where coherences instead contract deterministically at .
The central consequence for the learning-theoretic programme: UHM predicts that a viable learning system implements natural gradient descent, not the regime conjectured for biological complexity. Its efficiency comes not from a different exponent but from the Fano-induced anisotropy of the noise and from confinement of purity to the window . This is a sharp, falsifiable disagreement, and §9 gives a metric-free test statistic for it, together with the sample size it needs.
1. The two frameworks and their dictionary
Vanchurin's covariant learning framework (henceforth VL) is built on: trainable variables , a loss/fitness , a metric on trainable space, a noise covariance (his 6.7), and the covariant update (his 6.3). The learning algorithm is defined by the functional relation (his 6.8–6.9).
UHM is built on five axioms: reality as an -topos over density matrices, the Bures Grothendieck topology, dimension , a scale , and a Page–Wootters decomposition (the last derivable, T-87). Its dynamics is the triad forced by LGKS-completeness (T-57):
The dictionary is exact on every line that matters:
| VL object | UHM object | Status of the link |
|---|---|---|
| trainable state | coherence matrix | structural |
| loss | free-energy functional | T-39e (variational ) |
| metric | Bures/SLD metric — unique monotone metric (Petz; T-187) | forced, not chosen |
| noise covariance | covariance of one-step Kraus increments of | forced by T-41/T-59, relative to the canonical Kraus resolution (§7.2) |
| covariant descent | regeneration toward | T-39f–h |
| emergent time = block index | (Page–Wootters) | T-38b, T-87 |
| maximum-entropy identity (his 4.7) | derived here as a dynamical theorem (§4) | this page |
| multi-level structure | fractal holon, contraction | T-72, CC-5 |
| cellular network | Fano plane , unique optimal BIBD | T-41i |
2. Setup: the noise covariance of a holon
Let be the state of a holon and let the canonical dissipator have Lindblad set
Every result on this page depends on the line set only through the design parameters — which axis-triples realise the lines never enters. Axis-labelled corollaries (such as the third-point screen quoted in §5) do require the canonical translate set of the corpus, with — see selection rules.
In the quantum-jump (unravelling) picture, during a jump occurs with probability , taking to . Define the noise covariance as the covariance of the one-step increment :
This is Vanchurin's with raised indices — the covariance of temporal changes, his (7.4). The index placement is load-bearing, and it is the one thing a reader should check first: is a tangent vector (a change in ), so its covariance is contravariant, and his (7.4) states exactly . Get this backwards and the derived exponent flips sign.
His (7.4) is a centred covariance, while the definition above is a second moment. They agree only if the mean increment vanishes — which it does, exactly, and that is §7.1. Without it the bridge would not close.
Restricted to a state diagonal in the canonical basis, , all objects live on the population sector, the tangent space , with projector where .
2.1 Normalisation of the metric, stated explicitly
Monotone metrics are defined only up to a normalisation convention, and since the headline identity carries a numerical constant, the convention must be pinned down rather than assumed. On commuting (diagonal) perturbations every monotone metric is proportional to the classical Fisher–Rao metric, so write
The Bures value is , not : from , the diagonal terms have , giving . The SLD quantum Fisher information is Bures and reduces to the classical Fisher metric exactly (Braunstein–Caves; T-187 Char-III); it is the normalisation used in the natural-gradient literature.
What depends on and what does not. The exponent , the Fano factor , the noise–purity law and the effective rank are all independent of . Only the absolute constants move:
| quantity | Bures () | SLD-QFI () |
|---|---|---|
| per mode |
Bures is quoted as primary, because that is what Axiom 2 forces. All tests in §9 are stated so as not to depend on the choice.
2.2 A jump process, not a diffusion
One more thing bounds what the bridge claims. VL's (6.6) posits white noise, , which reads as a diffusion; the unravelling above is a compound Poisson process whose jumps are not small — each takes to a corner of the simplex. What the definition computes is therefore the quadratic-variation rate, i.e. the noise intensity, and for a zero-mean compound Poisson process the increment over a window has covariance exactly . So the identification with his (7.4) is exact at the level of the second cumulant and only there: the higher cumulants of a holon's jumps are not Gaussian. Every prediction on this page is a second-moment statement, which is why they survive; the practical consequence — increments must be aggregated over windows long enough for the central limit theorem to apply — is spelled out in §9.1.
3. The Kraus covariance is exactly multinomial
For the atomic part of the canonical dissipator and any diagonal state ,
i.e. exactly times the covariance matrix of a single-trial multinomial distribution with parameters .
Proof. For we have , hence , and , so as a vector in the population sector. Substituting,
using and . The bracket is .
Machine check: over 2000 random states (Dirichlet concentration ), the residual is .
The matrix is simultaneously the single-trial multinomial covariance and the inverse of the Fisher metric on the simplex. That coincidence is the whole engine of §4, and saying so up front is more honest than deriving it as though it were a surprise.
4. The exact natural-gradient identity
On commuting perturbations the Bures/SLD metric reduces to times the classical Fisher–Rao metric (§2.1).
On the tangent space ,
exactly, in Bures normalisation. Substituting into Vanchurin's (7.5) gives precisely . The holon therefore realises : natural gradient descent.
Proof. By §3 and ,
Since , the rank-one term is annihilated: , leaving , which is at .
Why this settles the exponent. VL's diagnostic (7.5) reads , and he spells out the two readable ends: and . The theorem is the second of these, with a constant. The constant carries units of rate — has dimensions of (state)/time while does not — so it fixes the time unit exactly as his (6.8) implicitly does; the functional relation is linear, which is what asserts.
Machine check: for 200 random states the spectrum of is constant ( at , ) with maximal spread ; and .
4.1 What this predicts about data
Combine the theorem with VL's own maximum-entropy identity (his 4.7, where is the static covariance of the population):
The covariance of temporal changes is proportional to the static covariance of the population, with the same shape and a single scalar of proportionality that is a pure rate. In [V1]'s own terms both sides are the single-trial multinomial covariance: the abundant, already-measured object (, his 7.1) and the never-measured object (, his 7.4) are predicted to be the same matrix up to one number. For that is a 20-parameter shape prediction with one free scalar; for a real genome it is enormously over-determined. This — not the constant — is the practically useful content of the identity.
As mathematics the two theorems above are elementary: for the multinomial family the Fisher metric is the inverse of the single-trial covariance, a standard fact of information geometry. No depth is claimed for the computation. The content is the forcing: UHM leaves no freedom in either factor — by Axiom 2, and the noise by the Lindblad set being the classifier atoms (T-41g–i) rather than a modelling choice. VL's §4 imposes by a maximum-entropy argument on the population covariance (his 4.3–4.7: isotropic constraint spherical distribution covariance identity in local coordinates); here the same identity emerges for the dynamical noise without being postulated. The results in §5–§7 are not elementary and carry the new content.
UHM proves, as a dynamical theorem, the identity that VL adopts as a maximum-entropy postulate — and simultaneously fixes the learning exponent that VL leaves open.
5. The Fano correction is the universal factor 11/9
With the full canonical dissipator (atomic + Fano) at matched per-channel rate , for every state,
independently of and of . Equivalently: the block layer carries exactly of the total metric-weighted noise, at every state. In Bures normalisation and .
Proof. For a Fano line with : , , , and the post-jump vector is with . Then
The metric-weighted quadratic form evaluates in closed form:
Hence
where is the number of lines and because every point lies on exactly lines (BIBD). Therefore , while . The ratio is and the total factor , with no dependence on and no dependence on .
Machine check: 300 random states, ratio , deviation from at most ; block fraction .
General BIBD form. For a design on points with blocks and replication , the same computation gives
Verified on three designs — (dev. ), (dev. ), (dev. ). One caveat the general form hides: the block layer is separately trace-preserving only when , since . This holds for and but fails for , where and the channel must be renormalised. The Fano plane is thus doubly distinguished: optimal and self-normalising.
Anisotropy. While the trace ratio is universal, the spectrum of is isotropic only at the maximally mixed state, where it equals in all six directions. Away from it the spectrum spreads, in units of :
| state | spectrum | anisotropy |
|---|---|---|
| , | ||
| (window centre), |
The first row is checked to against the exact fractions; each spectrum sums to , as the universal trace requires. The Fano layer is therefore what turns exact natural gradient into an anisotropic, preconditioned natural gradient — the structural analogue of the adaptive preconditioners (Adam/AdaBelief, his 6.9) that VL associates with , but obtained without leaving .
Per-axis refinement: the third-point screen (2026-08-07). The block
layer does one more exact thing that the trace cannot see: it screens
the third point of each line. The sensitivity of the population
covariance to is suppressed, relative to every
axis outside the pair, precisely when is the third point of the line
through — at by exactly , analytic and
machine-confirmed to six digits in both and
; for an abstract line set the screen moves to its third
points, as labelling invariance demands. On random states the third
point is the least-sensitive outside axis in of cases.
Corpus statement and proof sketch:
the suppression corollary;
instrument shadow_marks.py alongside the reproduction script.
6. Noise magnitude is fixed by purity
and , where is the purity and . Neither depends on the metric normalisation.
Proof. Immediate from §3: ; and by expanding the square and using and .
Consequences inside the UHM viability structure (all with , ):
| State | ||
|---|---|---|
| maximally mixed | ||
| viability threshold | ||
| centre of the conscious window | ||
| upper edge of the window |
So the conscious window — derived in UHM from independent considerations (T-124) — corresponds to a learning-noise band . In VL language: a viable learner is one whose gradient noise sits in a narrow band; too much noise (high entropy, ) and structure dissolves; too little () and the learner freezes at a fixed point. This is the same "Goldilocks" statement VL conjectures for the origin of life, but here it is a derived interval with explicit endpoints — and, being metric-free, it is the most directly testable of the predictions.
7. Zero drift, and the sector split
7.1 The dissipator is pure noise on populations
The canonical dissipator annihilates every diagonal state: , equivalently . Hence the second moment is the centred covariance of his (7.4), and all population drift is carried by .
Proof. and , since each point lies on lines; and . So .
Machine check: over 200 states, and exactly.
This is the clean structural match to VL's Langevin split: his (6.3) carries the deterministic ascent and (6.5)–(6.6) the zero-mean noise. In UHM the split is not modelled but forced by the triad — is the drift, is the noise, and they do not overlap on the population sector.
7.2 What the Fano layer is, and is not
The full canonical dissipator equals total dephasing at rate : its generator has spectrum exactly , and it coincides identically with the atomic-only dissipator at rescaled rate .
Machine check: atomic spectrum , full spectrum , and .
Two different Kraus resolutions of the same generator give different . So is well defined only relative to the canonical Kraus resolution, and that the classifier atoms and the seven Fano projectors are the physical channels is an ontological commitment of UHM (L-unification, T-41g–i), not a consequence of the dynamics. A reader who rejects that commitment keeps §3, §4, §6 and §7.1, and loses §5.
The same fact fixes the correct operational statement of the factor, which any experiment must respect:
- at matched per-channel rate : ratio — this is the comparison of a seven-channel and a fourteen-channel holon;
- at matched generator (i.e. matched observed decoherence rate, ): ratio — this is the comparison of two resolutions of one physical process.
Both are exact. Quoting for the second comparison would be wrong.
7.3 The sector split
All fourteen canonical Lindblad operators are diagonal in the classifier basis. Hence on the decohered manifold (diagonal ) every increment is diagonal and exactly, while coherences contract deterministically at (T-59, confirmed by the superoperator above).
Proof. and are diagonal; conjugation of a diagonal by a diagonal operator is diagonal; hence and are diagonal, so all contributions lie in the population sector.
Off the decohered manifold this is only approximate, and the honest statement is quadratic suppression rather than identical vanishing: with coherences of size inserted into an otherwise diagonal state, at and at — a ratio of for a factor in , i.e. exactly . Since the coherences themselves decay at , decays at . The decohered manifold is the attractor, so always; it is not identically zero everywhere.
Interpretation, and a genuine tension with VL. UHM predicts a two-sector structure: a population sector with stochastic, natural-gradient learning (, anisotropic through Fano), and a coherence sector that is noiseless and purely deterministic at .
VL calls the "quantum regime" because it yields Schrödinger-like dynamics on the trainables from a discrete shift symmetry [V2]. In UHM the sector is the population (diagonal) sector — the classical degrees of freedom — while literal quantum coherence lives in the sector where vanishes. These are two different senses of "quantum", and conflating them would be the easiest way to overclaim agreement. What UHM actually says is that both are present and complementary: emergent unitary -dynamics on populations (Page–Wootters, T-38b/T-87) and a strictly separate coherence sector. That is a refinement of his dichotomy rather than a restatement of it, and the sector split is the part that is independently falsifiable.
8. Effective rank for a holon
VL notes that empirical genotype covariance spectra decay as a power law with effective rank (his 7.3) of order against . For a holon the same quantity has an exact closed form. With over the population sector, §6 gives immediately
verified to against the numerical spectrum on 300 random states.
This is a two-moment identity, and it makes explicit that depends on as well as . At equal purity:
| one dominant six equal | two dominant five equal | |
|---|---|---|
| (maximum, ) | — | |
A search over all with gives , the upper end rising to as along the one-dominant family. So the viable band compresses to – of maximum only along the one-dominant family; over the whole window it can fall to . The falsifiable statement is the closed form, which is stronger than any interval: it ties two measured moments of to two measured moments of the state, with no free parameter beyond .
9. Predictions and how to falsify them
All statements below are consequences of the results above; none is fitted. The column "needs ?" matters in practice: predictions that do not require estimating a metric are immune to the normalisation question of §2.1 and are much easier to measure.
| # | Prediction | Value | Needs ? | Test |
|---|---|---|---|---|
| P1 | Learning exponent of a viable holon | exactly (natural gradient), not | no | Sphericity of — §9.1 |
| P2 | Shape identity with the static covariance | (with his 4.7) | no | Compare the covariance of temporal changes with the static population covariance; predicted equal up to one scalar |
| P3 | Universal Fano factor | at matched per-channel rate; at matched generator; block fraction | yes | Resolve the same process at site and block level, holding the declared normalisation fixed |
| P4 | Noise–purity law | ; viable band | no | Zero-intercept regression of on ; slope |
| P5 | Second moment / effective rank | exactly | no | Two moments of against two moments of the state |
| P6 | Sector split | on the decohered manifold, off it; coherences decay at | no | Interferometry / process tomography |
| P7 | Absolute normalisation | (Bures) or (SLD-QFI) | yes | Fixes from one measurement, then over-determines the rest |
| P8 | Network topology of the "cellular" layer | For with complete pairwise coverage the coordinating design is forced to be ; it is also the design for which the block layer is self-normalising () | no | Any architecture claiming optimal pairwise coordination must reproduce Fano incidence |
9.1 The metric-free discriminator, and the sample size it needs
VL's (7.5) degenerates at : the exponent diverges, and indeed carries no information about at all. That degeneracy is exactly what makes a clean test possible, because the two hypotheses make opposite shape claims about a single measured object:
- is spherical on (all six eigenvalues equal);
- , fully computable from the measured .
Neither requires estimating . The statistic is the classical sphericity (Mauchly) likelihood ratio on the whitened estimate,
with to test and to test . The theory supplies the separating power exactly: of the eigenvalues of , giving at and , , at , , .
The asymptotic sample size is a mirage. The naive gives –, but at those sizes the calibration fails badly: the simulated size of a nominal test is at , at , at , at . Power at is everywhere in the window (, , at for , , ). Use independent windows; below the test is anti-conservative and should not be reported.
Increments must be aggregated. A single idealised jump has only seven outcomes, so is parametrised by six numbers rather than twenty-one and is not Wishart; the calibration does not apply to it. Aggregate over windows with , where is the total jump rate, so the central limit theorem applies; the window covariance is then exactly (Campbell). Because is state-dependent, the window must sit between two timescales, ; whether such a window exists in a given system is an empirical question about the ratio of regeneration to dissipation rates.
Note finally that the discriminating power vanishes at the maximally mixed state (): there, and are indistinguishable in principle. The test needs a state away from uniformity — and the viability window is precisely where UHM says a conscious holon sits. The theory therefore predicts that the systems it claims as holons are exactly the systems in which its own central claim is measurable.
The sharpest disagreement. VL conjectures that biological complexity — and possibly the origin of life — corresponds to the intermediate regime , the "efficient learning" regime of Adam/AdaBelief (his 6.9). UHM says: a viable learner sits at with an anisotropic preconditioner supplied by the Fano layer, and its sophistication is carried by the anisotropy and the purity window, not by the exponent. Both claims live in the same measurable object; §9.1 decides between them with of order a hundred windows of data.
The disagreement extends to [V2]'s tri-regime classification, which
places classical biological evolution at . A toy model with no
UHM in it says otherwise, on both sides of the Langevin split. The
noise side: neutral Wright–Fisher resampling of individuals over
types has increment covariance exactly
— the tangent-space inverse
Fisher metric (a fact with a classical address: Antonelli–Strobeck
1977 already described genetic drift as diffusion in exactly this
geometry, the simplex becoming a Fisher sphere of radius two), so his
own criterion (7.5) reads identically,
with as the sole constant (simulated at , : pooled
shape to , quasi-static window to = sampling error,
Mauchly against the spherical prediction's critical ). The
drift side is classical: replicator dynamics under selection is
natural-gradient ascent of mean fitness in the Shahshahani–Fisher
metric (Shahshahani 1979; Hofbauer–Sigmund 1998). Whatever realises
or in biology must therefore break the
multinomial character of reproduction noise — a positive, checkable
question about mechanism rather than a preference among exponents.
The algebra on both sides is half a century old; the new content is
only the collision it creates inside [V2]'s own classification.
The thirty-line simulation is
wf_toy_a1.py.
The derivation is at the level of a holon's internal state dynamics. Applying it to population genetics assumes that an evolving population is itself a holon — which is exactly what UHM's scale-invariance theorem (T-72, with contraction ) asserts, but which remains a bridge assumption [И] rather than a theorem about biology. If the identification is rejected, P1 still stands for physical holons and for engineered learning systems built on the UHM core.
10. What each framework gives the other
UHM → VL.
-
Determination of the free function. is no longer a modelling choice: with the exact constants above.
-
Derivation of the maxent postulate's dynamical counterpart. The relation — his condition (7.5), for him a modelling choice among three — is here a theorem. It is the dynamical analogue of the static identity that his §4 obtains by a maximum-entropy argument; granting that static identity as well, the two together give the sharp form , which relates a routinely measured object to an unmeasured one.
-
A preferred basis. VL inherits einselection's open question ("why this interaction?"); UHM fixes the basis as the atoms of the subobject classifier, rigid up to (T-42a), with the seven directions functionally unique (7/7 minimality).
-
A forced network. The "cellular network" that coordinates agents is, at with complete pairwise coverage and optimal block size, uniquely the Fano plane (T-41c, T-41i) — which is moreover the case in which the block layer needs no renormalisation.
-
A ceiling on self-reference, and a second one on breadth. The ladder across levels gives : a self-learning system cannot nest reflection indefinitely — a structural limit absent from VL. A second bound runs the other way. A holon types exactly non-overlapping channels, and a node that coordinates others spends one on each, so no node addresses more than subordinates. The two bounds multiply: one holarchy reaches at most addressed contexts (T-304). This is not the dimension count of §9 — that is how large a level's state space is, this is how many situations it can tell apart — and the distinction matters, because VL's networks scale without either.
-
A variable that must not be trainable. VL's organising distinction is trainable against non-trainable, and it leaves open which is which for any given quantity. For one quantity the answer is now measured: addressing. A policy that decides which subsystem handles what, optimised against the same loss as the task it routes, never holds still long enough for a subsystem to specialise; splitting under such a policy was measured to be worth almost nothing, while the same split under an address declared once and frozen recovered the gain, and least-loaded declaration recovered the rest. Stability is the precondition and balance the multiplier (HOLARCH §9). The engineering reading is sharp and falsifiable in VL's own terms: the routing variable belongs on the non-trainable side, and a framework that trains it is paying for a freedom that costs more than it buys.
-
A reason why learning never completes. Lawvere incompleteness (T-55: ) implies a permanent gap between the state and its self-model; in UHM this gap is the source of a positive vacuum energy (T-71). A self-learning universe is, provably, a universe that cannot finish learning itself.
VL → UHM.
- An operational reading of : regeneration as covariant gradient ascent on fitness, with the loss identified as negative Malthusian fitness — a bridge to measurable biology. UHM returns the favour by proving the drift/noise split (§7.1) that his Langevin form assumes.
- An experimental target: his proposal to measure the covariance of temporal changes is precisely the measurement that tests P1–P5, and his (7.5) supplies the criterion in a form UHM can meet exactly.
- A macroscopic derivation route: the cell/agent split of the Self-Learning Universe (memory-cost vs processing-cost) is a physically motivated variational principle whose UHM counterpart — Gap curvature (T-73) and from the spectral action (T-74) — can be compared term by term.
Three ways to live in one geometry. A note that sharpens where the two frameworks stand relative to machine learning. Learning, in every case on this page, is motion along an information geometry; the cases differ only in how the step is obtained. In machine learning the step is approximated: the posterior over a continuum of weights is intractable, so it is chased by iterated local moves (Adam's second-moment preconditioner is a diagonal — VL's regime). In evolution the step is generated by physics: multinomial reproduction noise at finite population size has covariance equal to the inverse Fisher metric exactly (the Wright–Fisher toy above), so is not chosen by any optimiser — it falls out of discrete copying. And in UHM's cognitive layer the step is computed in closed form: counter-based conjugate Bayes (KT estimators, Willems mixtures) is mirror descent whose step is exact, i.e. the gradient is integrated rather than iterated — which is why the silicon contains no descent procedure at all while living in the same geometry, and why the physical layer's dissipative flow (a BKM/Bures step) still sits squarely in the family. Approximated, generated, computed: one geometry, three ways of taking the step.
The exponent as a phase of knowledge. A further sharpening, born from the goal-currency experiments in the silicon. Inside one living system the exponent is not a fixed property but a phase of its knowledge. The unseen is computed: every tact of conjugate counting is an exact natural-gradient step — . The proven is executed: once a road has been established by the repetition law, the agent replays it without recomputing any geometry — the step is blind to informational distinguishability, which is precisely the regime. In the silicon this is the golden-path executor, and it is not a defect but a load-bearing economy: in the transfer bench a third session takes three levels in 79 ticks instead of 1107 by executing proven traces. The boundary between the phases is settledness itself. Read this way, VL's classification acquires an architectural translation: "classical biology at " says that observed large-scale biology is dominated by the execution of fossilised decisions — DNA is evolution's golden path — while the per-generation sampling tact stays at by the multinomial identity above; the crossover scale between the two phases is where the algorithmic layer of heredity lives, and it is measurable.
The autodidaxy curve: knowing more never kills learning more. The phase picture makes one falsifiable prediction about any single learner: if the proven only narrows where computation happens (rather than extinguishing the computing organ), then no amount of accumulated knowledge should destroy the ability to keep learning. The text side of the silicon put this to an instrument. One bitwise-CTW base is pretrained on a prefix of length , then lives on a fixed common window of text; the level falls monotonically with (2.416 down to 2.000 bits per byte) — the base helps the life it precedes — and the content-controlled learning progress , priced against a frozen twin of the same base on the same two half-windows, stays positive at every (+0.278 down to +0.109). There is no ossification from erudition: the organ survives arbitrary amounts of inheritance. The frozen twin itself pays 2.220 against the living 2.000 — freezing costs real bits — and a small live first-order ledger mixed over the frozen base claws back only five percent of that gap: a light residual organ softens the freeze but does not replace continued learning of the deep base. In the phase language: execution is an economy, freezing is a tax, and the computing phase is indestructible by knowledge alone.
11. Reproducibility
Every number on this page is printed by
full_predictions.py (deterministic, seeds
20260806 / 20260807 / 20260808; NumPy + SciPy); the metric
normalisation runs as its own §0 block inside the same script, and the
per-axis third-point screen of §5 is printed by its companion
shadow_marks.py (seeds 20260807 / 20260808).
| § | Check | Result |
|---|---|---|
| 0 | Bures Fisher on the commuting sector | at |
| 1 | Multinomial identity, 2000 random states | |
| 2 | Natural-gradient identity, 200 random states; both normalisations; his (7.5) | spread ; |
| 3 | Fano factor, 300 random states; anisotropy; three BIBDs | ; ; |
| 4 | Both noise moments at four reference states | exact match |
| 5 | generator; ; on and off the diagonal | ; ; |
| 6 | closed form; window range | ; |
| 7 | Zero drift | ; |
| 8 | Sphericity test: separating power, simulated size and power; Campbell window identity | |
| 9 | Third-point screen: gap ; outside-axis argmin share | six digits; – |
A condensed statement of this bridge, set against Deutsch's many-worlds programme, is §2.9 of Many-Worlds (Everett–Deutsch) and UHM.
Corpus cross-references: T-41 (Fano channel family), T-42a (-rigidity), T-55 (Lawvere incompleteness), T-57 (triadic completeness), T-59 (), T-71 (vacuum energy), T-72 (scale invariance), T-73/T-74 (Gap curvature, spectral action), T-87 (A5 derivable), T-124 (conscious window), T-187 (why Bures). Registry rows: T-293, T-294, T-295; the per-axis screen lives under T-298; the composition ceiling and the addressing regime under T-304.
12. Summary
Vanchurin's framework asks: which learning algorithm does nature implement? UHM answers, without adjustable parameters:
Nature — at the level of a viable holon — performs natural gradient descent on the Bures geometry, preconditioned by the Fano design, in a purity band that keeps its effective rank well below maximum, on a sector strictly complementary to the one where quantum coherence lives.
The first of these is measurable without ever estimating a metric, and needs about a hundred windows of data. If the covariance of temporal changes is ever measured the way Vanchurin proposes, these are the numbers to compare against.