Skip to main content

The Learning Algorithm of a Holon is Natural Gradient

Vanchurin's programme derives physical and biological law from learning dynamics. It leaves one central quantity undetermined: the functional relation g(κ)g(\kappa) between the metric tensor on trainable space and the covariance of gradient noise, parametrised as a power law g=κag=\kappa^{a} with three named regimes — a=0a=0 (stochastic gradient), a=12a=\tfrac12 (efficient learning, Adam-like, conjectured to underlie biological complexity) and a=1a=1 (natural gradient, the "quantum regime"). He states that direct empirical estimates of κ\kappa remain unavailable and calls obtaining them a formidable challenge.

In UHM this quantity is not free. Both factors are forced by the axioms — gg by Axiom 2 (Bures is the unique monotone Riemannian metric), κ\kappa by the canonical dissipator whose Lindblad operators are the atoms of the subobject classifier plus the seven Fano projectors. Their relation is therefore a theorem, not a parameter:

κat  =  γ4Ng1ΠgκatΠ  =  γ4NΠ\kappa^{\uparrow\uparrow}_{\text{at}} \;=\; \frac{\gamma}{4N}\,g^{-1} \qquad\Longleftrightarrow\qquad \Pi\,g\,\kappa_{\text{at}}\,\Pi \;=\; \frac{\gamma}{4N}\,\Pi

on the population sector of a holon of dimension N=7N=7, where Π\Pi is the projector onto the tangent space. This is verbatim Vanchurin's own criterion for a=1a=1: substituting a=1a=1 into his Eq. (7.5) gives g1=g1κg1=κg^{-1}=g^{-1}\kappa g^{-1}=\kappa^{\uparrow\uparrow}. Hence

a=1(natural gradient), exactly.a = 1 \quad\text{(natural gradient), exactly.}
Primary sources

Equation numbers throughout refer to:

  • [V1] V. Vanchurin, Geometric framework for biological evolution, arXiv:2603.15198 — Eqs. (4.7), (6.3), (6.7)–(6.9), (7.1)–(7.6). This is the source of the g(κ)g(\kappa) programme in the form used here.
  • [V2] V. Vanchurin, Geometric Learning Dynamics, arXiv:2504.14728, Biological Cybernetics (2026) — Eqs. (2.2), (2.8), (2.9), (5.5).

Two notational hazards, stated once. [V1] uses aa for the learning exponent and α\alpha as a coordinate index; [V2] uses α\alpha for the exponent. We follow [V1] and write aa. And his γ\gamma is a learning rate; our γ\gamma is the dissipator rate — they are unrelated, and every γ\gamma on this page is ours.

Four further exact results follow. The canonical Fano layer multiplies the metric-weighted noise trace by the universal rational factor 11/911/9 — independent of the state — so that the block layer carries exactly 2/112/11 of the noise, and it introduces a computable anisotropy. Noise magnitude is tied to purity, Trκ=γN(1P)\operatorname{Tr}\kappa= \frac{\gamma}{N}(1-P). The dissipator annihilates every diagonal state, so on populations it is pure noise without drift — which is what makes κ\kappa a genuinely centred covariance in the sense of his (7.4). And there is a strict sector split: κcoh=0\kappa^{\text{coh}}=0 on the decohered manifold, where coherences instead contract deterministically at λdeco=5γ/21\lambda_{\text{deco}}=5\gamma/21.

The central consequence for the learning-theoretic programme: UHM predicts that a viable learning system implements natural gradient descent, not the a=12a=\tfrac12 regime conjectured for biological complexity. Its efficiency comes not from a different exponent but from the Fano-induced anisotropy of the noise and from confinement of purity to the window P(2/7,3/7]P\in(2/7,3/7]. This is a sharp, falsifiable disagreement, and §9 gives a metric-free test statistic for it, together with the sample size it needs.


1. The two frameworks and their dictionary

Vanchurin's covariant learning framework (henceforth VL) is built on: trainable variables qq, a loss/fitness F\mathcal F, a metric gαr,βsg_{\alpha r,\beta s} on trainable space, a noise covariance καr,βs=αϕβϕ\kappa_{\alpha r,\beta s}=\langle\partial_\alpha\phi\, \partial_\beta\phi\rangle (his 6.7), and the covariant update q˙=g1F\dot q = g^{-1}\nabla\mathcal F (his 6.3). The learning algorithm is defined by the functional relation g(κ)g(\kappa) (his 6.8–6.9).

UHM is built on five axioms: reality as an \infty-topos over density matrices, the Bures Grothendieck topology, dimension N=7N=7, a scale ω0\omega_0, and a Page–Wootters decomposition (the last derivable, T-87). Its dynamics is the triad forced by LGKS-completeness (T-57):

LΩ[Γ]=i[H,Γ]Aut+DΩ[Γ]Fano dissipator+R[Γ]regeneration.\mathcal L_\Omega[\Gamma] = \underbrace{-i[H,\Gamma]}_{\text{Aut}} + \underbrace{\mathcal D_\Omega[\Gamma]}_{\text{Fano dissipator}} + \underbrace{\mathcal R[\Gamma]}_{\text{regeneration}} .

The dictionary is exact on every line that matters:

VL objectUHM objectStatus of the link
trainable state qqcoherence matrix ΓD(C7)\Gamma\in\mathcal D(\mathbb C^7)structural
loss F-\mathcal Ffree-energy functional F[φ;Γ]\mathcal F[\varphi;\Gamma]T-39e (variational φ\varphi)
metric ggBures/SLD metricunique monotone metric (Petz; T-187)forced, not chosen
noise covariance κ\kappacovariance of one-step Kraus increments of DΩ\mathcal D_\Omegaforced by T-41/T-59, relative to the canonical Kraus resolution (§7.2)
covariant descent q˙=g1F\dot q=g^{-1}\nabla\mathcal Fregeneration R\mathcal R toward ρ\rho_*T-39f–h
emergent time = block indexτZ7\tau\in\mathbb Z_7 (Page–Wootters)T-38b, T-87
maximum-entropy identity g1=cg^{-1}=c (his 4.7)derived here as a dynamical theorem (§4)this page
multi-level structurefractal holon, contraction cF=1/3=1/QR(7)c_F=1/3=1/\lvert\mathrm{QR}(7)\rvertT-72, CC-5
cellular networkFano plane PG(2,2)\mathrm{PG}(2,2), unique optimal BIBD(7,3,1)(7,3,1)T-41i

2. Setup: the noise covariance of a holon

Let ρ\rho be the state of a holon and let the canonical dissipator have Lindblad set

K  =  {Lk=kk}k=17atomic    {Lp=13Πp}pPG(2,2)Fano,Πp=ipii.\mathcal K \;=\; \underbrace{\{L_k=|k\rangle\langle k|\}_{k=1}^{7}}_{\text{atomic}} \;\cup\; \underbrace{\Bigl\{L_p=\tfrac{1}{\sqrt3}\Pi_p\Bigr\}_{p\in\mathrm{PG}(2,2)}}_{\text{Fano}}, \qquad \Pi_p=\sum_{i\in p}|i\rangle\langle i| .
Labelling invariance

Every result on this page depends on the line set only through the design parameters (v,k,r,b)=(7,3,3,7)(v,k,r,b)=(7,3,3,7) — which axis-triples realise the lines never enters. Axis-labelled corollaries (such as the third-point screen quoted in §5) do require the canonical translate set of the corpus, {1,2,4}+k(mod7)\{1,2,4\}+k \pmod 7 with U=6,O=7U=6,\,O=7 — see selection rules.

In the quantum-jump (unravelling) picture, during dtdt a jump ee occurs with probability pe=γdtNTr(LeρLe)p_e=\frac{\gamma\,dt}{N}\operatorname{Tr} (L_e\rho L_e^\dagger), taking ρ\rho to ρe=LeρLe/Tr(LeρLe)\rho_e=L_e\rho L_e^\dagger/ \operatorname{Tr}(L_e\rho L_e^\dagger). Define the noise covariance as the covariance of the one-step increment Δe=ρeρ\Delta_e=\rho_e-\rho:

κ  :=  1dtepe  ΔeΔe.\kappa \;:=\; \frac{1}{dt}\sum_{e} p_e\;\Delta_e\otimes\Delta_e .

This is Vanchurin's καr,βs\kappa^{\alpha r,\beta s} with raised indices — the covariance of temporal changes, his (7.4). The index placement is load-bearing, and it is the one thing a reader should check first: Δe\Delta_e is a tangent vector (a change in λ\lambda), so its covariance is contravariant, and his (7.4) states exactly κ=g1κg1=q˙q˙q˙q˙\kappa^{\uparrow\uparrow}=g^{-1}\kappa_{\downarrow\downarrow}g^{-1} =\langle\dot q\dot q\rangle-\langle\dot q\rangle\langle\dot q\rangle. Get this backwards and the derived exponent flips sign.

His (7.4) is a centred covariance, while the definition above is a second moment. They agree only if the mean increment vanishes — which it does, exactly, and that is §7.1. Without it the bridge would not close.

Restricted to a state diagonal in the canonical basis, ρ=diag(λ)\rho=\operatorname{diag}(\lambda), all objects live on the population sector, the tangent space T={xRN:1 ⁣x=0}T=\{x\in\mathbb R^N:\mathbf 1^{\!\top}x=0\}, with projector Π=I1NJ\Pi=I-\tfrac1N J where J=11 ⁣J=\mathbf 1\mathbf 1^{\!\top}.

2.1 Normalisation of the metric, stated explicitly

Monotone metrics are defined only up to a normalisation convention, and since the headline identity carries a numerical constant, the convention must be pinned down rather than assumed. On commuting (diagonal) perturbations every monotone metric is proportional to the classical Fisher–Rao metric, so write

gT  =  cgdiag(1/λ),cg=14 (Bures),cg=1 (SLD-QFI = Fisher–Rao).g\big|_T \;=\; c_g\,\operatorname{diag}(1/\lambda), \qquad c_g=\tfrac14\ \text{(Bures)},\qquad c_g=1\ \text{(SLD-QFI = Fisher–Rao)} .

The Bures value is 14\tfrac14, not 12\tfrac12: from DB2=2(1F)=12jkjdρk2/(λj+λk)D_B^2=2(1-F)=\tfrac12\sum_{jk}\lvert\langle j|d\rho|k\rangle\rvert^2/ (\lambda_j+\lambda_k), the diagonal terms have λj+λk=2λj\lambda_j+\lambda_k=2\lambda_j, giving DB2=14j(dλj)2/λjD_B^2=\tfrac14\sum_j (d\lambda_j)^2/\lambda_j. The SLD quantum Fisher information is 4×4\times Bures and reduces to the classical Fisher metric exactly (Braunstein–Caves; T-187 Char-III); it is the normalisation used in the natural-gradient literature.

What depends on cgc_g and what does not. The exponent a=1a=1, the Fano factor 11/911/9, the noise–purity law and the effective rank are all independent of cgc_g. Only the absolute constants move:

quantityBures (cg=14c_g=\tfrac14)SLD-QFI (cg=1c_g=1)
ΠgκatΠ\Pi g\kappa_{\text{at}}\Piγ4NΠ=γ28Π\frac{\gamma}{4N}\Pi=\frac{\gamma}{28}\PiγNΠ=γ7Π\frac{\gamma}{N}\Pi=\frac{\gamma}{7}\Pi
Tr(gκat)\operatorname{Tr}(g\kappa_{\text{at}})3γ14\frac{3\gamma}{14}6γ7\frac{6\gamma}{7}
Tr(gκfull)\operatorname{Tr}(g\kappa_{\text{full}})11γ42\frac{11\gamma}{42}22γ21\frac{22\gamma}{21}
per mode11γ252\frac{11\gamma}{252}11γ63\frac{11\gamma}{63}

Bures is quoted as primary, because that is what Axiom 2 forces. All tests in §9 are stated so as not to depend on the choice.

2.2 A jump process, not a diffusion

One more thing bounds what the bridge claims. VL's (6.6) posits white noise, ϕϕ=Cδ(tt)\langle\phi\phi'\rangle=C\,\delta(t-t'), which reads as a diffusion; the unravelling above is a compound Poisson process whose jumps are not small — each takes λ\lambda to a corner eke_k of the simplex. What the definition computes is therefore the quadratic-variation rate, i.e. the noise intensity, and for a zero-mean compound Poisson process the increment over a window Δt\Delta t has covariance exactly κΔt\kappa\,\Delta t. So the identification with his (7.4) is exact at the level of the second cumulant and only there: the higher cumulants of a holon's jumps are not Gaussian. Every prediction on this page is a second-moment statement, which is why they survive; the practical consequence — increments must be aggregated over windows long enough for the central limit theorem to apply — is spelled out in §9.1.


3. The Kraus covariance is exactly multinomial

Theorem [Т]

For the atomic part of the canonical dissipator and any diagonal state ρ=diag(λ)\rho=\operatorname{diag}(\lambda),

κat  =  γN(diag(λ)λλ ⁣),\kappa_{\text{at}} \;=\; \frac{\gamma}{N}\Bigl( \operatorname{diag}(\lambda)-\lambda\lambda^{\!\top}\Bigr),

i.e. exactly γ/N\gamma/N times the covariance matrix of a single-trial multinomial distribution with parameters λ\lambda.

Proof. For Lk=kkL_k=|k\rangle\langle k| we have LkρLk=λkkkL_k\rho L_k^\dagger= \lambda_k|k\rangle\langle k|, hence Tr(LkρLk)=λk\operatorname{Tr}(L_k\rho L_k^\dagger)=\lambda_k, pk=γdtλk/Np_k=\gamma\,dt\,\lambda_k/N and ρk=kk\rho_k=|k\rangle\langle k|, so Δk=ekλ\Delta_k=e_k-\lambda as a vector in the population sector. Substituting,

κat=γNkλk(ekλ)(ekλ) ⁣=γN[kλkekek ⁣λλ ⁣λλ ⁣+λλ ⁣],\kappa_{\text{at}}=\frac{\gamma}{N}\sum_k \lambda_k (e_k-\lambda)(e_k-\lambda)^{\!\top} =\frac{\gamma}{N}\Bigl[\sum_k\lambda_k e_ke_k^{\!\top} -\lambda\lambda^{\!\top}-\lambda\lambda^{\!\top} +\lambda\lambda^{\!\top}\Bigr],

using kλk=1\sum_k\lambda_k=1 and kλkek=λ\sum_k\lambda_k e_k=\lambda. The bracket is diag(λ)λλ ⁣\operatorname{diag}(\lambda)-\lambda\lambda^{\!\top}. \blacksquare

Machine check: over 2000 random states (Dirichlet concentration [0.4,4]\in[0.4,4]), the residual is 2.78×10172.78\times10^{-17}.

The matrix diagλλλ ⁣\operatorname{diag}\lambda-\lambda\lambda^{\!\top} is simultaneously the single-trial multinomial covariance and the inverse of the Fisher metric on the simplex. That coincidence is the whole engine of §4, and saying so up front is more honest than deriving it as though it were a surprise.


4. The exact natural-gradient identity

On commuting perturbations the Bures/SLD metric reduces to cgc_g times the classical Fisher–Rao metric (§2.1).

Theorem (natural gradient) [Т]

On the tangent space TT,

ΠgκatΠ  =  γ4NΠ,i.e.κat=γ4Ng1\Pi\,g\,\kappa_{\text{at}}\,\Pi \;=\; \frac{\gamma}{4N}\,\Pi , \qquad\text{i.e.}\qquad \kappa^{\uparrow\uparrow}_{\text{at}}=\frac{\gamma}{4N}\,g^{-1}

exactly, in Bures normalisation. Substituting a=1a=1 into Vanchurin's (7.5) gives precisely g1=g1κg1=κg^{-1}=g^{-1}\kappa g^{-1}= \kappa^{\uparrow\uparrow}. The holon therefore realises a=1a=1: natural gradient descent.

Proof. By §3 and g=cgdiag(1/λ)g=c_g\operatorname{diag}(1/\lambda),

gκat=cgγNdiag(1/λ)(diagλλλ ⁣)=cgγN(I1λ ⁣).g\,\kappa_{\text{at}} =\frac{c_g\gamma}{N}\operatorname{diag}(1/\lambda) \bigl(\operatorname{diag}\lambda-\lambda\lambda^{\!\top}\bigr) =\frac{c_g\gamma}{N}\bigl(I-\mathbf 1\lambda^{\!\top}\bigr).

Since Π1=0\Pi\mathbf 1=0, the rank-one term is annihilated: Π(1λ ⁣)Π=(Π1)(λ ⁣Π)=0\Pi(\mathbf 1\lambda^{\!\top})\Pi=(\Pi\mathbf 1)(\lambda^{\!\top}\Pi)=0, leaving ΠgκatΠ=cgγNΠ\Pi g\kappa_{\text{at}}\Pi=\frac{c_g\gamma}{N}\Pi, which is γ4NΠ\frac{\gamma}{4N}\Pi at cg=14c_g=\tfrac14. \blacksquare

Why this settles the exponent. VL's diagnostic (7.5) reads g1=(g1κg1)a/(2a1)g^{-1}=\bigl(g^{-1}\kappa g^{-1}\bigr)^{a/(2a-1)}, and he spells out the two readable ends: a=0g1=Ia=0\Rightarrow g^{-1}=I and a=1g1=g1κg1a=1\Rightarrow g^{-1}=g^{-1}\kappa g^{-1}. The theorem is the second of these, with a constant. The constant carries units of rate — κ\kappa has dimensions of (state)2^2/time while gg does not — so it fixes the time unit exactly as his (6.8) implicitly does; the functional relation is linear, which is what a=1a=1 asserts.

Machine check: for 200 random states the spectrum of ΠgκatΠ\Pi g\kappa_{\text{at}}\Pi is constant =0.035714285714=0.035714285714 (=1/28=γ/4N=1/28=\gamma/4N at γ=1\gamma=1, N=7N=7) with maximal spread 8.33×10178.33\times10^{-17}; and maxκatγ4Ng1=1.73×1017\max\lvert\kappa^{\uparrow\uparrow}_{\text{at}}-\frac{\gamma}{4N} g^{-1}\rvert=1.73\times10^{-17}.

4.1 What this predicts about data

Combine the theorem with VL's own maximum-entropy identity (his 4.7, g1=cg^{-1}=c where cc is the static covariance of the population):

κ  =  γ4N  c.\kappa^{\uparrow\uparrow}\;=\;\frac{\gamma}{4N}\;c .

The covariance of temporal changes is proportional to the static covariance of the population, with the same shape and a single scalar of proportionality that is a pure rate. In [V1]'s own terms both sides are the single-trial multinomial covariance: the abundant, already-measured object (cc, his 7.1) and the never-measured object (κ\kappa^{\uparrow\uparrow}, his 7.4) are predicted to be the same matrix up to one number. For N=7N=7 that is a 20-parameter shape prediction with one free scalar; for a real genome it is enormously over-determined. This — not the constant — is the practically useful content of the identity.

How deep is this, honestly

As mathematics the two theorems above are elementary: for the multinomial family the Fisher metric is the inverse of the single-trial covariance, a standard fact of information geometry. No depth is claimed for the computation. The content is the forcing: UHM leaves no freedom in either factor — gg by Axiom 2, and the noise by the Lindblad set being the classifier atoms (T-41g–i) rather than a modelling choice. VL's §4 imposes g1=cg^{-1}=c by a maximum-entropy argument on the population covariance (his 4.3–4.7: isotropic constraint \Rightarrow spherical distribution \Rightarrow covariance \propto identity in local coordinates); here the same identity emerges for the dynamical noise without being postulated. The results in §5–§7 are not elementary and carry the new content.

UHM proves, as a dynamical theorem, the identity that VL adopts as a maximum-entropy postulate — and simultaneously fixes the learning exponent that VL leaves open.


5. The Fano correction is the universal factor 11/9

Theorem [Т]

With the full canonical dissipator (atomic + Fano) at matched per-channel rate γ/N\gamma/N, for every state,

Tr ⁣(gκfull)Tr ⁣(gκat)  =  1+29  =  119,\frac{\operatorname{Tr}\!\left(g\,\kappa_{\text{full}}\right)} {\operatorname{Tr}\!\left(g\,\kappa_{\text{at}}\right)} \;=\; 1+\frac29 \;=\; \frac{11}{9},

independently of λ\lambda and of cgc_g. Equivalently: the block layer carries exactly 2/112/11 of the total metric-weighted noise, at every state. In Bures normalisation Tr(gκat)=3γ14\operatorname{Tr}(g\kappa_{\text{at}})= \frac{3\gamma}{14} and Tr(gκfull)=11γ42\operatorname{Tr}(g\kappa_{\text{full}})= \frac{11\gamma}{42}.

Proof. For a Fano line pp with sp=ipλis_p=\sum_{i\in p}\lambda_i: LpρLp=13diag(λi[ip])L_p\rho L_p^\dagger=\tfrac13\operatorname{diag}(\lambda_i[i\in p]), Tr=sp/3\operatorname{Tr}=s_p/3, pe=γdtNsp3p_e=\frac{\gamma dt}{N}\frac{s_p}{3}, and the post-jump vector is upu_p with (up)i=λi[ip]/sp(u_p)_i=\lambda_i[i\in p]/s_p. Then

κF=γNpsp3(upλ)(upλ) ⁣.\kappa_{\text{F}}=\frac{\gamma}{N}\sum_{p}\frac{s_p}{3} \,(u_p-\lambda)(u_p-\lambda)^{\!\top}.

The metric-weighted quadratic form evaluates in closed form:

(upλ) ⁣g(upλ)=cg[ipλi(1sp1)2+ipλi]=cg[(1sp)2sp+(1sp)]=cg(1sp)sp.(u_p-\lambda)^{\!\top}g\,(u_p-\lambda) =c_g\Bigl[\sum_{i\in p}\lambda_i\Bigl(\tfrac1{s_p}-1\Bigr)^{2} +\sum_{i\notin p}\lambda_i\Bigr] =c_g\Bigl[\frac{(1-s_p)^2}{s_p}+(1-s_p)\Bigr] =\frac{c_g(1-s_p)}{s_p}.

Hence

Tr(gκF)=γNpsp3cg(1sp)sp=cgγ3Np(1sp)=cgγ3N(br),\operatorname{Tr}(g\kappa_{\text F}) =\frac{\gamma}{N}\sum_p\frac{s_p}{3}\cdot\frac{c_g(1-s_p)}{s_p} =\frac{c_g\gamma}{3N}\sum_p(1-s_p) =\frac{c_g\gamma}{3N}\bigl(b-r\bigr),

where b=7b=7 is the number of lines and psp=riλi=r=3\sum_p s_p=r\sum_i\lambda_i=r=3 because every point lies on exactly r=3r=3 lines (BIBD(7,3,1)(7,3,1)). Therefore Tr(gκF)=4cgγ21\operatorname{Tr}(g\kappa_{\text F})=\frac{4c_g\gamma}{21}, while Tr(gκat)=cgγN(N1)=18cgγ21\operatorname{Tr}(g\kappa_{\text{at}})=\frac{c_g\gamma}{N}(N-1)= \frac{18c_g\gamma}{21}. The ratio is 4/18=2/94/18=2/9 and the total factor 11/911/9, with no dependence on λ\lambda and no dependence on cgc_g. \blacksquare

Machine check: 300 random states, ratio [1.222222222222,1.222222222222]\in[1.222222222222,\,1.222222222222], deviation from 11/911/9 at most 4.4×10164.4\times10^{-16}; block fraction 2/11=0.1818181818182/11=0.181818181818.

General BIBD form. For a design (v,k,1)(v,k,1) on v=Nv=N points with bb blocks and replication rr, the same computation gives

Tr(gκblocks)Tr(gκat)=brk(v1).\frac{\operatorname{Tr}(g\kappa_{\text{blocks}})} {\operatorname{Tr}(g\kappa_{\text{at}})}=\frac{b-r}{k\,(v-1)} .

Verified on three designs — PG(2,2)(7,3,1)7336=29\mathrm{PG}(2,2)\,(7,3,1)\to\frac{7-3} {3\cdot6}=\frac29 (dev. 1.4×10161.4\times10^{-16}), AG(2,3)(9,3,1)13\mathrm{AG}(2,3)\,(9,3,1)\to\frac13 (dev. 1.7×10161.7\times10^{-16}), PG(2,3)(13,4,1)316\mathrm{PG}(2,3)\,(13,4,1)\to\frac3{16} (dev. 5.6×10175.6\times10^{-17}). One caveat the general form hides: the block layer is separately trace-preserving only when r=kr=k, since pLpLp=rkI\sum_p L_p^\dagger L_p= \frac rk I. This holds for (7,3,1)(7,3,1) and (13,4,1)(13,4,1) but fails for (9,3,1)(9,3,1), where r/k=4/3r/k=4/3 and the channel must be renormalised. The Fano plane is thus doubly distinguished: optimal and self-normalising.

Anisotropy. While the trace ratio is universal, the spectrum of ΠgκfullΠ\Pi g\kappa_{\text{full}}\Pi is isotropic only at the maximally mixed state, where it equals 119γ4N\frac{11}{9}\cdot\frac{\gamma}{4N} in all six directions. Away from it the spectrum spreads, in units of γ/4N\gamma/4N:

statespectrumanisotropy
I/7I/7{119}×6\{\tfrac{11}9\}^{\times6}1.00001.0000
λ=(12,112×6)\lambda=(\tfrac12,\tfrac1{12}^{\times6}), P=724P=\tfrac7{24}{1312,1312,119,119,119,32}\{\tfrac{13}{12},\tfrac{13}{12},\tfrac{11}9,\tfrac{11}9,\tfrac{11}9,\tfrac32\}1.38461.3846
P=514P=\tfrac5{14} (window centre), λ=(47,114×6)\lambda=(\tfrac47,\tfrac1{14}^{\times6}){1615,1615,119,119,119,2315}\{\tfrac{16}{15},\tfrac{16}{15},\tfrac{11}9,\tfrac{11}9,\tfrac{11}9,\tfrac{23}{15}\}1.43751.4375

The first row is checked to 1.78×10151.78\times10^{-15} against the exact fractions; each spectrum sums to 6119=2236\cdot\frac{11}9=\frac{22}3, as the universal trace requires. The Fano layer is therefore what turns exact natural gradient into an anisotropic, preconditioned natural gradient — the structural analogue of the adaptive preconditioners (Adam/AdaBelief, his 6.9) that VL associates with a=12a=\tfrac12, but obtained without leaving a=1a=1.

Per-axis refinement: the third-point screen (2026-08-07). The block layer does one more exact thing that the trace cannot see: it screens the third point of each line. The sensitivity of the population covariance κij\kappa_{ij} to λx\lambda_x is suppressed, relative to every axis outside the pair, precisely when xx is the third point of the line through (i,j)(i,j) — at I/7I/7 by exactly γ/189\gamma/189, analytic and machine-confirmed to six digits in both κOE\kappa_{OE} and κOU\kappa_{OU}; for an abstract line set the screen moves to its third points, as labelling invariance demands. On random states the third point is the least-sensitive outside axis in 92%\approx92\% of cases. Corpus statement and proof sketch: the κ0\kappa_0 suppression corollary; instrument shadow_marks.py alongside the reproduction script.


6. Noise magnitude is fixed by purity

Theorem [Т]

Trκat=γN(1P)\operatorname{Tr}\kappa_{\text{at}}=\frac{\gamma}{N}(1-P) and Trκat2=γ2N2(P+P22S3)\operatorname{Tr}\kappa_{\text{at}}^2=\frac{\gamma^2}{N^2} (P+P^2-2S_3), where P=Trρ2P=\operatorname{Tr}\rho^2 is the purity and S3=iλi3S_3=\sum_i\lambda_i^3. Neither depends on the metric normalisation.

Proof. Immediate from §3: Tr[diagλλλ ⁣]=iλiiλi2=1P\operatorname{Tr}[\operatorname{diag}\lambda-\lambda\lambda^{\!\top}]= \sum_i\lambda_i-\sum_i\lambda_i^2=1-P; and Tr[(diagλλλ ⁣)2]=P2S3+P2\operatorname{Tr}[(\operatorname{diag}\lambda-\lambda\lambda^{\!\top})^2] =P-2S_3+P^2 by expanding the square and using Tr(diagλλλ ⁣)=S3\operatorname{Tr}(\operatorname{diag}\lambda\cdot\lambda\lambda^{\!\top}) =S_3 and Tr(λλ ⁣λλ ⁣)=P2\operatorname{Tr}(\lambda\lambda^{\!\top}\lambda \lambda^{\!\top})=P^2. \blacksquare

Consequences inside the UHM viability structure (all with γ=1\gamma=1, N=7N=7):

StatePPTrκat\operatorname{Tr}\kappa_{\text{at}}
maximally mixed I/7I/71/71/76/49=0.1224496/49=0.122449
viability threshold PcritP_{\text{crit}}2/72/75/49=0.1020415/49=0.102041
centre of the conscious window5/145/149/98=0.0918379/98=0.091837
upper edge of the window3/73/74/49=0.0816334/49=0.081633

So the conscious window P(2/7,3/7]P\in(2/7,3/7] — derived in UHM from independent considerations (T-124) — corresponds to a learning-noise band Trκ[4γ/49,5γ/49)\operatorname{Tr}\kappa\in[4\gamma/49,\,5\gamma/49). In VL language: a viable learner is one whose gradient noise sits in a narrow band; too much noise (high entropy, P1/7P\to1/7) and structure dissolves; too little (P1P\to1) and the learner freezes at a fixed point. This is the same "Goldilocks" statement VL conjectures for the origin of life, but here it is a derived interval with explicit endpoints — and, being metric-free, it is the most directly testable of the predictions.


7. Zero drift, and the sector split

7.1 The dissipator is pure noise on populations

Theorem [Т]

The canonical dissipator annihilates every diagonal state: DΩ[diagλ]=0\mathcal D_\Omega[\operatorname{diag}\lambda]=0, equivalently epeΔe=0\sum_e p_e\Delta_e=0. Hence the second moment is the centred covariance of his (7.4), and all population drift is carried by R\mathcal R.

Proof. kLkρLk=diagλ=ρ\sum_k L_k\rho L_k^\dagger=\operatorname{diag}\lambda=\rho and pLpρLp=13pdiag(λi[ip])=133λ=ρ\sum_p L_p\rho L_p^\dagger=\tfrac13\sum_p \operatorname{diag}(\lambda_i[i\in p])=\tfrac13\cdot3\lambda=\rho, since each point lies on r=3r=3 lines; and kTr(LkρLk)=pTr(LpρLp)=1\sum_k\operatorname{Tr} (L_k\rho L_k^\dagger)=\sum_p\operatorname{Tr}(L_p\rho L_p^\dagger)=1. So epeΔe=γN(2ρ2ρ)=0\sum_e p_e\Delta_e=\frac{\gamma}{N}(2\rho-2\rho)=0. \blacksquare

Machine check: maxepeΔe=3.64×1017\max\lvert\sum_e p_e\Delta_e\rvert=3.64\times10^{-17} over 200 states, and maxLfull[diagλ]=0.00\max\lvert\mathcal L_{\text{full}} [\operatorname{diag}\lambda]\rvert=0.00 exactly.

This is the clean structural match to VL's Langevin split: his (6.3) carries the deterministic ascent and (6.5)–(6.6) the zero-mean noise. In UHM the split is not modelled but forced by the triad — R\mathcal R is the drift, DΩ\mathcal D_\Omega is the noise, and they do not overlap on the population sector.

7.2 What the Fano layer is, and is not

Theorem [Т]

The full canonical dissipator equals total dephasing at rate 5γ/215\gamma/21: its 49×4949\times49 generator has spectrum exactly {0×7,(5γ/21)×42}\{0^{\times7},(-5\gamma/21)^{\times42}\}, and it coincides identically with the atomic-only dissipator at rescaled rate 5γ/35\gamma/3.

Machine check: atomic spectrum {0×7,(γ/7)×42}\{0^{\times7},(-\gamma/7)^{\times42}\}, full spectrum {0×7,(5γ/21)×42}\{0^{\times7},(-5\gamma/21)^{\times42}\}, and maxLfull(γ)Lat(5γ/3)=2.78×1017\max\lvert\mathcal L_{\text{full}}(\gamma)-\mathcal L_{\text{at}} (5\gamma/3)\rvert=2.78\times10^{-17}.

warning
κ\kappa is not unravelling-invariant, and the Fano layer is invisible in the master equation

Two different Kraus resolutions of the same generator give different κ\kappa. So κ\kappa is well defined only relative to the canonical Kraus resolution, and that the classifier atoms and the seven Fano projectors are the physical channels is an ontological commitment of UHM (L-unification, T-41g–i), not a consequence of the dynamics. A reader who rejects that commitment keeps §3, §4, §6 and §7.1, and loses §5.

The same fact fixes the correct operational statement of the 11/911/9 factor, which any experiment must respect:

  • at matched per-channel rate γ/N\gamma/N: ratio =11/9=11/9 — this is the comparison of a seven-channel and a fourteen-channel holon;
  • at matched generator (i.e. matched observed decoherence rate, γat=5γ/3\gamma_{\text{at}}=5\gamma/3): ratio =11/15=11/15 — this is the comparison of two resolutions of one physical process.

Both are exact. Quoting 11/911/9 for the second comparison would be wrong.

7.3 The sector split

Theorem [Т]

All fourteen canonical Lindblad operators are diagonal in the classifier basis. Hence on the decohered manifold (diagonal ρ\rho) every increment Δe\Delta_e is diagonal and κcoh=0\kappa^{\text{coh}}=0 exactly, while coherences contract deterministically at λdeco=5γ3N=5γ21\lambda_{\text{deco}}=\frac{5\gamma}{3N}=\frac{5\gamma}{21} (T-59, confirmed by the 49×4949\times49 superoperator above).

Proof. Lk=kkL_k=|k\rangle\langle k| and Lp=Πp/3L_p=\Pi_p/\sqrt3 are diagonal; conjugation of a diagonal ρ\rho by a diagonal operator is diagonal; hence ρe\rho_e and Δe\Delta_e are diagonal, so all contributions lie in the population sector. \blacksquare

Off the decohered manifold this is only approximate, and the honest statement is quadratic suppression rather than identical vanishing: with coherences of size cc inserted into an otherwise diagonal state, maxκcoh=1.079×104\max\lvert\kappa^{\text{coh}}\rvert=1.079\times10^{-4} at c=0.02c=0.02 and 2.698×1032.698\times10^{-3} at c=0.10c=0.10 — a ratio of 25.025.0 for a factor 55 in cc, i.e. exactly O(c2)O(c^2). Since the coherences themselves decay at 5γ/215\gamma/21, κcoh\kappa^{\text{coh}} decays at 10γ/2110\gamma/21. The decohered manifold is the attractor, so κcoh0\kappa^{\text{coh}}\to0 always; it is not identically zero everywhere.

Interpretation, and a genuine tension with VL. UHM predicts a two-sector structure: a population sector with stochastic, natural-gradient learning (a=1a=1, anisotropic through Fano), and a coherence sector that is noiseless and purely deterministic at 5γ/215\gamma/21.

VL calls a=1a=1 the "quantum regime" because it yields Schrödinger-like dynamics on the trainables from a discrete shift symmetry [V2]. In UHM the a=1a=1 sector is the population (diagonal) sector — the classical degrees of freedom — while literal quantum coherence lives in the sector where κ\kappa vanishes. These are two different senses of "quantum", and conflating them would be the easiest way to overclaim agreement. What UHM actually says is that both are present and complementary: emergent unitary τ\tau-dynamics on populations (Page–Wootters, T-38b/T-87) and a strictly separate coherence sector. That is a refinement of his dichotomy rather than a restatement of it, and the sector split is the part that is independently falsifiable.


8. Effective rank for a holon

VL notes that empirical genotype covariance spectra decay as a power law with effective rank reff=ζ(α)2/ζ(2α)r_{\text{eff}}=\zeta(\alpha)^2/\zeta(2\alpha) (his 7.3) of order 10210^2 against dim(g)109\dim(g)\sim10^9. For a holon the same quantity has an exact closed form. With reff=(Trκ)2/Trκ2r_{\text{eff}}= (\operatorname{Tr}\kappa)^2/\operatorname{Tr}\kappa^2 over the population sector, §6 gives immediately

reff=(1P)2P+P22S3,S3=iλi3,r_{\text{eff}}=\frac{(1-P)^2}{P+P^2-2S_3},\qquad S_3=\sum_i\lambda_i^3 ,

verified to 6.66×10156.66\times10^{-15} against the numerical spectrum on 300 random states.

The effective rank is not a function of purity alone

This is a two-moment identity, and it makes explicit that reffr_{\text{eff}} depends on S3S_3 as well as PP. At equal purity:

PPone dominant ++ six equaltwo dominant ++ five equal
1/71/76.00006.0000 (maximum, =N1=N-1)
2/72/74.22474.22473.08543.0854
5/145/1427/7=3.857127/7=3.85712.18582.1858
3/73/73.59313.59311.50471.5047

A search over all λ\lambda with P(2/7,3/7]P\in(2/7,3/7] gives reff[1.4719,4.1262]r_{\text{eff}}\in[1.4719,\,4.1262], the upper end rising to 4.22474.2247 as P2/7+P\to2/7^+ along the one-dominant family. So the viable band compresses reffr_{\text{eff}} to 60\approx6070%70\% of maximum only along the one-dominant family; over the whole window it can fall to 25%\approx25\%. The falsifiable statement is the closed form, which is stronger than any interval: it ties two measured moments of κ\kappa to two measured moments of the state, with no free parameter beyond γ\gamma.


9. Predictions and how to falsify them

All statements below are consequences of the results above; none is fitted. The column "needs gg?" matters in practice: predictions that do not require estimating a metric are immune to the normalisation question of §2.1 and are much easier to measure.

#PredictionValueNeeds gg?Test
P1Learning exponent of a viable holona=1a=1 exactly (natural gradient), not a=12a=\tfrac12noSphericity of κ\kappa^{\uparrow\uparrow} — §9.1
P2Shape identity with the static covarianceκ=γ4Nc\kappa^{\uparrow\uparrow}=\frac{\gamma}{4N}c (with his 4.7)noCompare the covariance of temporal changes with the static population covariance; predicted equal up to one scalar
P3Universal Fano factor11/911/9 at matched per-channel rate; 11/1511/15 at matched generator; block fraction 2/112/11yesResolve the same process at site and block level, holding the declared normalisation fixed
P4Noise–purity lawTrκ=γN(1P)\operatorname{Tr}\kappa=\frac{\gamma}{N}(1-P); viable band [4γ/49,5γ/49)[4\gamma/49,5\gamma/49)noZero-intercept regression of Trκ^\operatorname{Tr}\hat\kappa on 1P^1-\hat P; slope =γ/N=\gamma/N
P5Second moment / effective rankreff=(1P)2P+P22S3r_{\text{eff}}=\frac{(1-P)^2}{P+P^2-2S_3} exactlynoTwo moments of κ^\hat\kappa against two moments of the state
P6Sector splitκcoh=0\kappa^{\text{coh}}=0 on the decohered manifold, O(ρcoh2)O(\lVert\rho_{\text{coh}}\rVert^2) off it; coherences decay at 5γ/215\gamma/21noInterferometry / process tomography
P7Absolute normalisationTr(gκfull)=11γ42\operatorname{Tr}(g\kappa_{\text{full}})=\frac{11\gamma}{42} (Bures) or 22γ21\frac{22\gamma}{21} (SLD-QFI)yesFixes γ\gamma from one measurement, then over-determines the rest
P8Network topology of the "cellular" layerFor N=7N=7 with complete pairwise coverage the coordinating design is forced to be PG(2,2)\mathrm{PG}(2,2); it is also the design for which the block layer is self-normalising (r=kr=k)noAny N=7N=7 architecture claiming optimal pairwise coordination must reproduce Fano incidence

9.1 The metric-free discriminator, and the sample size it needs

VL's (7.5) degenerates at a=12a=\tfrac12: the exponent a/(2a1)a/(2a-1) diverges, and indeed κ=g1/a2=I\kappa^{\uparrow\uparrow}=g^{1/a-2}=I carries no information about gg at all. That degeneracy is exactly what makes a clean test possible, because the two hypotheses make opposite shape claims about a single measured object:

  • a=12a=\tfrac12 \Longrightarrow κ\kappa^{\uparrow\uparrow} is spherical on TT (all six eigenvalues equal);
  • a=1a=1 \Longrightarrow κdiagλλλ ⁣\kappa^{\uparrow\uparrow}\propto \operatorname{diag}\lambda-\lambda\lambda^{\!\top}, fully computable from the measured λ\lambda.

Neither requires estimating gg. The statistic is the classical sphericity (Mauchly) likelihood ratio on the whitened estimate,

T=MlndetTA(TrA/6)6,A=Σ01/2κ^Σ01/2,df=d(d+1)21=20,T=-M\ln\frac{\det{}_T A}{\bigl(\operatorname{Tr}A/6\bigr)^{6}}, \qquad A=\Sigma_0^{-1/2}\,\hat\kappa\,\Sigma_0^{-1/2}, \qquad \mathrm{df}=\tfrac{d(d+1)}2-1=20 ,

with Σ0=Π\Sigma_0=\Pi to test a=12a=\tfrac12 and Σ0=diagλλλ ⁣\Sigma_0= \operatorname{diag}\lambda-\lambda\lambda^{\!\top} to test a=1a=1. The theory supplies the separating power exactly: Δ(λ)=6ln(AM/GM)\Delta(\lambda)=6\ln(\text{AM}/\text{GM}) of the eigenvalues of κ\kappa^{\uparrow\uparrow}, giving Δ=0\Delta=0 at I/7I/7 and 0.81590.8159, 1.04651.0465, 1.23841.2384 at P=2/7P=2/7, 5/145/14, 3/73/7.

Two practical warnings, both measured rather than asserted

The asymptotic sample size is a mirage. The naive Mncp/ΔM\gtrsim\text{ncp}/\Delta gives M17M\gtrsim172626, but at those sizes the χ202\chi^2_{20} calibration fails badly: the simulated size of a nominal 5%5\% test is 10.0%10.0\% at M=25M=25, 8.0%8.0\% at M=50M=50, 6.1%6.1\% at M=100M=100, 5.5%5.5\% at M=200M=200. Power at M=100M=100 is 100%100\% everywhere in the window (77.3%77.3\%, 88.2%88.2\%, 93.1%93.1\% at M=25M=25 for P=2/7P=2/7, 5/145/14, 3/73/7). Use M100M\approx100 independent windows; below M50M\approx50 the test is anti-conservative and should not be reported.

Increments must be aggregated. A single idealised jump has only seven outcomes, so κ^=diagp^p^λ ⁣λp^ ⁣+λλ ⁣\hat\kappa=\operatorname{diag}\hat p- \hat p\lambda^{\!\top}-\lambda\hat p^{\!\top}+\lambda\lambda^{\!\top} is parametrised by six numbers rather than twenty-one and is not Wishart; the df=20\mathrm{df}=20 calibration does not apply to it. Aggregate over windows with ΛΔt1\Lambda\Delta t\gg1, where Λ=2γ/N\Lambda=2\gamma/N is the total jump rate, so the central limit theorem applies; the window covariance is then exactly κΔt\kappa\,\Delta t (Campbell). Because κ\kappa is state-dependent, the window must sit between two timescales, 1/ΛΔtτdrift1/\Lambda\ll\Delta t\ll\tau_{\text{drift}}; whether such a window exists in a given system is an empirical question about the ratio of regeneration to dissipation rates.

Note finally that the discriminating power vanishes at the maximally mixed state (Δ=0\Delta=0): there, a=1a=1 and a=12a=\tfrac12 are indistinguishable in principle. The test needs a state away from uniformity — and the viability window P(2/7,3/7]P\in(2/7,3/7] is precisely where UHM says a conscious holon sits. The theory therefore predicts that the systems it claims as holons are exactly the systems in which its own central claim is measurable.

The sharpest disagreement. VL conjectures that biological complexity — and possibly the origin of life — corresponds to the intermediate regime a=12a=\tfrac12, the "efficient learning" regime of Adam/AdaBelief (his 6.9). UHM says: a viable learner sits at a=1a=1 with an anisotropic preconditioner supplied by the Fano layer, and its sophistication is carried by the anisotropy and the purity window, not by the exponent. Both claims live in the same measurable object; §9.1 decides between them with of order a hundred windows of data.

The disagreement extends to [V2]'s tri-regime classification, which places classical biological evolution at a=0a=0. A toy model with no UHM in it says otherwise, on both sides of the Langevin split. The noise side: neutral Wright–Fisher resampling of MM individuals over dd types has increment covariance exactly (diagppp ⁣)/M(\operatorname{diag}p-pp^{\!\top})/M — the tangent-space inverse Fisher metric (a fact with a classical address: Antonelli–Strobeck 1977 already described genetic drift as diffusion in exactly this geometry, the simplex becoming a Fisher sphere of radius two), so his own criterion (7.5) reads a=1a=1 identically, with 1/M1/M as the sole constant (simulated at d=3d=3, M=105M=10^5: pooled shape to 1.9%1.9\%, quasi-static window to 6.4%6.4\% = sampling error, Mauchly 506506 against the spherical prediction's critical 5.995.99). The drift side is classical: replicator dynamics under selection is natural-gradient ascent of mean fitness in the Shahshahani–Fisher metric (Shahshahani 1979; Hofbauer–Sigmund 1998). Whatever realises a=0a=0 or a=12a=\tfrac12 in biology must therefore break the multinomial character of reproduction noise — a positive, checkable question about mechanism rather than a preference among exponents. The algebra on both sides is half a century old; the new content is only the collision it creates inside [V2]'s own classification. The thirty-line simulation is wf_toy_a1.py.

Level-matching caveat

The derivation is at the level of a holon's internal state dynamics. Applying it to population genetics assumes that an evolving population is itself a holon — which is exactly what UHM's scale-invariance theorem (T-72, with contraction cF=1/3c_F=1/3) asserts, but which remains a bridge assumption [И] rather than a theorem about biology. If the identification is rejected, P1 still stands for physical holons and for engineered learning systems built on the UHM core.


10. What each framework gives the other

UHM → VL.

  1. Determination of the free function. g(κ)g(\kappa) is no longer a modelling choice: a=1a=1 with the exact constants above.

  2. Derivation of the maxent postulate's dynamical counterpart. The relation g1κg^{-1}\propto\kappa^{\uparrow\uparrow} — his a=1a=1 condition (7.5), for him a modelling choice among three — is here a theorem. It is the dynamical analogue of the static identity g1=cg^{-1}=c that his §4 obtains by a maximum-entropy argument; granting that static identity as well, the two together give the sharp form κ=γ4Nc\kappa^{\uparrow\uparrow}=\frac{\gamma}{4N}c, which relates a routinely measured object to an unmeasured one.

  3. A preferred basis. VL inherits einselection's open question ("why this interaction?"); UHM fixes the basis as the atoms of the subobject classifier, rigid up to G2=Aut(O)G_2=\operatorname{Aut}(\mathbb O) (T-42a), with the seven directions functionally unique (7/7 minimality).

  4. A forced network. The "cellular network" that coordinates agents is, at N=7N=7 with complete pairwise coverage and optimal block size, uniquely the Fano plane (T-41c, T-41i) — which is moreover the case r=kr=k in which the block layer needs no renormalisation.

  5. A ceiling on self-reference, and a second one on breadth. The PcritP_{\text{crit}} ladder across levels gives SADmax=3\mathrm{SAD}_{\max}=3: a self-learning system cannot nest reflection indefinitely — a structural limit absent from VL. A second bound runs the other way. A holon types exactly (72)=21\binom{7}{2}=21 non-overlapping channels, and a node that coordinates others spends one on each, so no node addresses more than 2121 subordinates. The two bounds multiply: one holarchy reaches at most 213=926121^{3}=9261 addressed contexts (T-304). This is not the dimension count dimCogn=7n+1\dim\mathrm{Cog}_n=7^{\,n+1} of §9 — that is how large a level's state space is, this is how many situations it can tell apart — and the distinction matters, because VL's networks scale without either.

  6. A variable that must not be trainable. VL's organising distinction is trainable against non-trainable, and it leaves open which is which for any given quantity. For one quantity the answer is now measured: addressing. A policy that decides which subsystem handles what, optimised against the same loss as the task it routes, never holds still long enough for a subsystem to specialise; splitting under such a policy was measured to be worth almost nothing, while the same split under an address declared once and frozen recovered the gain, and least-loaded declaration recovered the rest. Stability is the precondition and balance the multiplier (HOLARCH §9). The engineering reading is sharp and falsifiable in VL's own terms: the routing variable belongs on the non-trainable side, and a framework that trains it is paying for a freedom that costs more than it buys.

  7. A reason why learning never completes. Lawvere incompleteness (T-55: ThUHMΩ\mathrm{Th}_{\text{UHM}}\subsetneq\Omega) implies a permanent gap between the state and its self-model; in UHM this gap is the source of a positive vacuum energy (T-71). A self-learning universe is, provably, a universe that cannot finish learning itself.

VL → UHM.

  1. An operational reading of R\mathcal R: regeneration as covariant gradient ascent on fitness, with the loss identified as negative Malthusian fitness — a bridge to measurable biology. UHM returns the favour by proving the drift/noise split (§7.1) that his Langevin form assumes.
  2. An experimental target: his proposal to measure the covariance of temporal changes is precisely the measurement that tests P1–P5, and his (7.5) supplies the criterion in a form UHM can meet exactly.
  3. A macroscopic derivation route: the cell/agent split of the Self-Learning Universe (memory-cost vs processing-cost) is a physically motivated variational principle whose UHM counterpart — Gap curvature Curv2=ω02γij2Gap2\lVert\mathrm{Curv}\rVert^2=\omega_0^2 \lVert\gamma_{ij}\rVert^2\mathrm{Gap}^2 (T-73) and VGapV_{\text{Gap}} from the spectral action (T-74) — can be compared term by term.

Three ways to live in one geometry. A note that sharpens where the two frameworks stand relative to machine learning. Learning, in every case on this page, is motion along an information geometry; the cases differ only in how the step is obtained. In machine learning the step is approximated: the posterior over a continuum of weights is intractable, so it is chased by iterated local moves (Adam's second-moment preconditioner is a diagonal g1/2g^{-1/2} — VL's a=1/2a=1/2 regime). In evolution the step is generated by physics: multinomial reproduction noise at finite population size has covariance equal to the inverse Fisher metric exactly (the Wright–Fisher toy above), so a=1a=1 is not chosen by any optimiser — it falls out of discrete copying. And in UHM's cognitive layer the step is computed in closed form: counter-based conjugate Bayes (KT estimators, Willems mixtures) is mirror descent whose step is exact, i.e. the gradient is integrated rather than iterated — which is why the silicon contains no descent procedure at all while living in the same geometry, and why the physical layer's dissipative flow (a BKM/Bures step) still sits squarely in the a=1a=1 family. Approximated, generated, computed: one geometry, three ways of taking the step.

The exponent as a phase of knowledge. A further sharpening, born from the goal-currency experiments in the silicon. Inside one living system the exponent aa is not a fixed property but a phase of its knowledge. The unseen is computed: every tact of conjugate counting is an exact natural-gradient step — a=1a=1. The proven is executed: once a road has been established by the repetition law, the agent replays it without recomputing any geometry — the step is blind to informational distinguishability, which is precisely the a=0a=0 regime. In the silicon this is the golden-path executor, and it is not a defect but a load-bearing economy: in the transfer bench a third session takes three levels in 79 ticks instead of 1107 by executing proven traces. The boundary between the phases is settledness itself. Read this way, VL's classification acquires an architectural translation: "classical biology at a=0a=0" says that observed large-scale biology is dominated by the execution of fossilised decisions — DNA is evolution's golden path — while the per-generation sampling tact stays at a=1a=1 by the multinomial identity above; the crossover scale between the two phases is where the algorithmic layer of heredity lives, and it is measurable.

The autodidaxy curve: knowing more never kills learning more. The phase picture makes one falsifiable prediction about any single learner: if the proven only narrows where computation happens (rather than extinguishing the computing organ), then no amount of accumulated knowledge should destroy the ability to keep learning. The text side of the silicon put this to an instrument. One bitwise-CTW base is pretrained on a prefix of length T0T_0, then lives on a fixed common window of text; the level falls monotonically with T0T_0 (2.416 down to 2.000 bits per byte) — the base helps the life it precedes — and the content-controlled learning progress LPLP^{*}, priced against a frozen twin of the same base on the same two half-windows, stays positive at every T0T_0 (+0.278 down to +0.109). There is no ossification from erudition: the a=1a=1 organ survives arbitrary amounts of a=0a=0 inheritance. The frozen twin itself pays 2.220 against the living 2.000 — freezing costs real bits — and a small live first-order ledger mixed over the frozen base claws back only five percent of that gap: a light residual organ softens the freeze but does not replace continued learning of the deep base. In the phase language: execution is an economy, freezing is a tax, and the computing phase is indestructible by knowledge alone.


11. Reproducibility

Every number on this page is printed by full_predictions.py (deterministic, seeds 20260806 / 20260807 / 20260808; NumPy + SciPy); the metric normalisation runs as its own §0 block inside the same script, and the per-axis third-point screen of §5 is printed by its companion shadow_marks.py (seeds 20260807 / 20260808).

§CheckResult
0Bures =14=\tfrac14 Fisher on the commuting sector0.24997277560.2499727756 at ε=105\varepsilon=10^{-5}
1Multinomial identity, 2000 random states2.78×10172.78\times10^{-17}
2Natural-gradient identity, 200 random states; both normalisations; his (7.5)spread 8.33×10178.33\times10^{-17}; 1.73×10171.73\times10^{-17}
3Fano factor, 300 random states; anisotropy; three BIBDs4.4×10164.4\times10^{-16}; 1.78×10151.78\times10^{-15}; 1.7×1016\le1.7\times10^{-16}
4Both noise moments at four reference statesexact match
549×4949\times49 generator; Lfull(γ)=Lat(5γ/3)\mathcal L_{\text{full}}(\gamma)=\mathcal L_{\text{at}}(5\gamma/3); κcoh\kappa^{\text{coh}} on and off the diagonal{0×7,(5γ/21)×42}\{0^{\times7},(-5\gamma/21)^{\times42}\}; 2.78×10172.78\times10^{-17}; O(c2)O(c^2)
6reffr_{\text{eff}} closed form; window range6.66×10156.66\times10^{-15}; [1.4719,4.1262][1.4719,4.1262]
7Zero drift3.64×10173.64\times10^{-17}; 0.000.00
8Sphericity test: separating power, simulated size and power; Campbell window identityM100M\approx100
9Third-point screen: I/7I/7 gap γ/189\equiv\gamma/189; outside-axis argmin sharesix digits; 919196%96\%

A condensed statement of this bridge, set against Deutsch's many-worlds programme, is §2.9 of Many-Worlds (Everett–Deutsch) and UHM.

Corpus cross-references: T-41 (Fano channel family), T-42a (G2G_2-rigidity), T-55 (Lawvere incompleteness), T-57 (triadic completeness), T-59 (λdeco=5γ/21\lambda_{\text{deco}}=5\gamma/21), T-71 (vacuum energy), T-72 (scale invariance), T-73/T-74 (Gap curvature, spectral action), T-87 (A5 derivable), T-124 (conscious window), T-187 (why Bures). Registry rows: T-293, T-294, T-295; the per-axis screen lives under T-298; the composition ceiling and the addressing regime under T-304.


12. Summary

Vanchurin's framework asks: which learning algorithm does nature implement? UHM answers, without adjustable parameters:

κat=γ4Ng1    a=1,Tr(gκfull)Tr(gκat)=119,Trκ=γN(1P),\kappa^{\uparrow\uparrow}_{\text{at}}=\frac{\gamma}{4N}\,g^{-1} \;\Longrightarrow\; a=1, \qquad \frac{\operatorname{Tr}(g\kappa_{\text{full}})} {\operatorname{Tr}(g\kappa_{\text{at}})}=\frac{11}{9}, \qquad \operatorname{Tr}\kappa=\frac{\gamma}{N}(1-P), reff=(1P)2P+P22S3,κcohdiag=0.r_{\text{eff}}=\frac{(1-P)^2}{P+P^2-2S_3}, \qquad \kappa^{\text{coh}}\big|_{\text{diag}}=0 .

Nature — at the level of a viable holon — performs natural gradient descent on the Bures geometry, preconditioned by the Fano design, in a purity band that keeps its effective rank well below maximum, on a sector strictly complementary to the one where quantum coherence lives.

The first of these is measurable without ever estimating a metric, and needs about a hundred windows of data. If the covariance of temporal changes is ever measured the way Vanchurin proposes, these are the numbers to compare against.