Research portal

Mathematical note

History-conditioned modular succession: mathematical contract

math/history-conditioned-modular-succession.md

Edition
Site v0.3.0 · continuous main snapshot
Source revision
ec2865b0eac15148675c629981a545632b3571c5
Extent
3,875 words
Public route
https://www.cordana.dev/math/history-conditioned-modular-succession/
Mapped records1 mapped record

Direct repository links only; no document-level evidence status is implied.

  • Status: frontier notation and experiment design; no result
  • Audit: history-conditioned modular succession and priority effects
  • Promotion boundary: this note creates no claim, principle, candidate, protocol, or fixture identifier
  • Purpose: define fixed-task-and-eligibility order estimands, resource identities, causal-cut contrasts, endpoint units, multiplicity control, and kill boundaries before any large implementation is proposed

Scope

The mathematical question is whether a randomized sequence changes a final learned system after the presented task multiset, eligible module identity set, and all declared resources are held fixed. Realized active, consolidated, merged, or retired module state may differ and is an endpoint. The notation does not assume that an order effect exists or that any effect is ecological in mechanism.

Microbial abundance, model accuracy, allocated parameters, examples, seconds, bytes, and joules remain different quantities. No conversion among them is introduced here.

Indices, objects, and units

SymbolDefinitionUnit
KKnumber of distinct task blocks and eligible module identities in the bounded design; initially K=4K=4count
kkpaired admission-unit index, k{1,,K}k\in\{1,\ldots,K\}count
\ellexecuting-module index when tasks and modules are evaluated separatelycount
j,jj,j'sequence-position indices, each in {1,,K}\{1,\ldots,K\}count
rrpaired random-seed indexcount
aalearning-method indexcategory
qqendpoint indexcategory
h,hh,h'distinct mechanism-factor indices when used togethercategory
T\mathcal Tfrozen multiset of task blocksset-valued
TkT_kchecksummed task block with identity kkdata object
MkM_keligible module identity paired with TkT_k before randomizationtyped identity
Ak=(Tk,Mk)A_k=(T_k,M_k)frozen task/module admission unittyped pair
ΠK\Pi_Kset of all permutations of KK identitiesset-valued
π\pione randomized permutation in ΠK\Pi_Kdimensionless mapping
πj\pi_jpaired admission-unit identity shown at sequence position jjcount
z\mathbf zvector of binary mechanism interventionsdimensionless vector
sa,js_{a,j}complete learned and structural state for method aa after position jjtyped state
θa,j\theta_{a,j}trainable parameter state within sa,js_{a,j}parameter vector
ρa,j\rho_{a,j}router/admission state within sa,js_{a,j}typed state
ωa,j\omega_{a,j}optimizer state within sa,js_{a,j}typed state
ma,jm_{a,j}replay or external learning-memory state within sa,js_{a,j}typed state
ca,jc_{a,j}capacity-allocation state within sa,js_{a,j}parameter/byte vector
Ya,π,z,r,qY_{a,\pi,\mathbf z,r,q}observed endpoint qq for one method, order, intervention cell, and seedendpoint-specific
μa,π,z,q\mu_{a,\pi,\mathbf z,q}seed expectation of Ya,π,z,r,qY_{a,\pi,\mathbf z,r,q} under the frozen generatorendpoint-specific
NkdataN^{\mathrm{data}}_kexogenous examples presented from task kkcount
Na,π,kaccN^{\mathrm{acc}}_{a,\pi,k\ell}examples from task TkT_k accepted by executing module MM_\ellcount
Ua,π,kU_{a,\pi,k\ell}optimizer-update applications to MM_\ell attributed to examples from TkT_k under the frozen attribution rulecount
CatotC^{\mathrm{tot}}_atotal installed parameter-capacity ceilingparameters
CaactC^{\mathrm{act}}_amaximum simultaneously active parameter ceilingparameters
BapeakB^{\mathrm{peak}}_apeak live attributed storagebytes
ttelapsed wall time from the registered run startseconds
EaE_acalibrated externally measured run energy at the declared boundaryjoules
PaP_amean measured power over a declared intervalwatts

Every typed state is serialized or hashed at each task boundary. A variable that affects later updates but is absent from sa,js_{a,j} is an unmeasured shared state and invalidates the corresponding causal cut.

Fixed-task-multiset and module-eligibility identity

For the initial design, freeze a one-to-one admission map g(Tk)=Mkg(T_k)=M_k before order assignment and define Ak=(Tk,Mk)A_k=(T_k,M_k). Each pair appears exactly once. A valid permutation satisfies

π:{1,,K}{1,,K}\pi:\{1,\ldots,K\}\rightarrow\{1,\ldots,K\}

as a bijection, and therefore

j=1K{Tπj}=T,\biguplus_{j=1}^{K}\{T_{\pi_j}\}=\mathcal T,

where \biguplus denotes multiset union. Equivalently, for every identity kk,

j=1K1{πj=k}=1.\sum_{j=1}^{K}\mathbb 1\{\pi_j=k\}=1.

Here 1{}\mathbb 1\{\cdot\} is a dimensionless indicator. If a later design repeats blocks, the right-hand side becomes a frozen multiplicity nkn_k that is identical for every order. Adding or deleting a presented task, eligible module identity, initialization image, example, label, augmentation, or response opportunity violates the order estimand. Realized activation, consolidation, merge, or retirement does not violate it; those states remain measured outcomes.

Let Ik,rI_{k,r} be the stored initialization image for paired module MkM_k and seed rr. It is drawn before π\pi is assigned, so

Ik,r(π)=Ik,r(π)I_{k,r}(\pi)=I_{k,r}(\pi')

for all compared orders π\pi and π\pi'. Every other random stream is keyed by its typed identity and paired seed rather than by global call order. Otherwise a permutation could change initialization, dropout, augmentation, router, or fault draws and the contrast would not isolate order.

Let ξk,r\boldsymbol\xi_{k,r} denote the frozen stream bundle keyed by admission unit AkA_k, paired seed rr, stream type, and within-identity draw index. A sequence position selects that identity's bundle; it does not create a new position-keyed bundle.

The method-specific state transition is written abstractly as

sa,j=Fa ⁣(sa,j1,Aπj,z,ξπj,r),s_{a,j} = F_a\!\left( s_{a,j-1},A_{\pi_j},\mathbf z,\boldsymbol\xi_{\pi_j,r} \right),

where FaF_a is the frozen transition implementation and ξπj,r\boldsymbol\xi_{\pi_j,r} is the paired random-input bundle for the identity occupying position jj. Position-indexed exogenous disturbances are prohibited unless they are a separately registered, randomized factor. This notation permits parameters, routers, optimizer moments, replay, and structure to carry history; it does not require all methods to contain every component.

Exposure, capacity, optimizer, evaluator, and budget identities

Exogenous exposure

For each task kk, let Dk\mathcal D_k be the frozen checksummed sequence of examples, labels, augmentations, and interaction outcomes. Exogenous parity requires

Na,π,kdata=NkdataN^{\mathrm{data}}_{a,\pi,k} = N^{\mathrm{data}}_k

for every method aa and order π\pi in a matched comparison. The equality is in counts and content: equal counts with different examples are not equal exposure.

Routed acceptance may be an endogenous mechanism. The accepted-exposure ledger therefore retains both task identity kk and executing-module identity \ell. Define the task-total accepted count and its dimensionless exposure ratio as

Na,π,kacc==1KNa,π,kacc,xa,π,k=Na,π,kaccNkdata,xa,π,k0.N^{\mathrm{acc}}_{a,\pi,k\cdot} = \sum_{\ell=1}^{K}N^{\mathrm{acc}}_{a,\pi,k\ell}, \qquad x_{a,\pi,k\cdot} = \frac{N^{\mathrm{acc}}_{a,\pi,k\cdot}} {N^{\mathrm{data}}_k}, \qquad x_{a,\pi,k\cdot}\ge0.

The ratio can exceed one only when the registered method duplicates or fans out one presented example to more than one module; every duplicate remains charged. Module-total accepted exposure is

Na,π,acc=kNa,π,kacc.N^{\mathrm{acc}}_{a,\pi,\cdot\ell} = \sum_k N^{\mathrm{acc}}_{a,\pi,k\ell}.

The exposure-equalizer intervention freezes the complete task-by-executing-module targets NkN^{\star}_{k\ell} and UkU^{\star}_{k\ell} such that

Na,π,kacc=Nk,Ua,π,k=UkN^{\mathrm{acc}}_{a,\pi,k\ell}=N^{\star}_{k\ell}, \qquad U_{a,\pi,k\ell}=U^{\star}_{k\ell}

for every task kk, executing module \ell, and evaluated order in that intervention level. Thus equal task totals cannot hide a different routing allocation. Mixed-task update attribution and any fan-out rule are frozen before order assignment. Rejected, downweighted, or duplicate work remains charged even when it is not applied as an update.

Let ba,π,kb_{a,\pi,k} be module kk's birth time in seconds from run start and let Ta,πwallT^{\mathrm{wall}}_{a,\pi} be the registered run duration in seconds. Its final wall-clock age is

Aa,π,kfinal=Ta,πwallba,π,k[s].A^{\mathrm{final}}_{a,\pi,k} = T^{\mathrm{wall}}_{a,\pi}-b_{a,\pi,k} \quad [\mathrm{s}].

The pre-instantiation cut sets all ba,π,k=0b_{a,\pi,k}=0 while freezing inactive state. The exposure equalizer changes the NkaccN^{\mathrm{acc}}_{k\ell} and UkU_{k\ell} matrices, not bb. This is why age and accepted exposure can be crossed independently.

Capacity

Let Ca,π,k(t)C_{a,\pi,k}(t) be parameters allocated to module kk at time tt, and let Aa,π,k(t){0,1}A_{a,\pi,k}(t)\in\{0,1\} indicate that the module is active. Valid runs satisfy

k=1KCa,π,k(t)Catot\sum_{k=1}^{K}C_{a,\pi,k}(t)\le C^{\mathrm{tot}}_a

and

k=1KAa,π,k(t)Ca,π,k(t)Caact\sum_{k=1}^{K}A_{a,\pi,k}(t)C_{a,\pi,k}(t) \le C^{\mathrm{act}}_a

for all measured tt. Every term is a parameter count. Storage for optimizer state, router state, replay, checkpoints, and metadata is counted separately in bytes; parameter count is not treated as bytes without the registered representation width.

For the position-blind reservation cut, frozen quotas CkC_k^{\star} obey

k=1KCk=Catot,Ca,π,k(t)Ck.\sum_{k=1}^{K}C_k^{\star}=C^{\mathrm{tot}}_a, \qquad C_{a,\pi,k}(t)\le C_k^{\star}.

Unused reserved capacity cannot be borrowed. Otherwise the later order could change effective total opportunity while nominal capacity remained fixed.

Optimizer and evaluator

Let Ua,πtotU^{\mathrm{tot}}_{a,\pi} be all optimizer updates, including replay, router, consolidation, recovery, and failed-attempt updates. Let Va,πevalV^{\mathrm{eval}}_{a,\pi} be evaluator calls. Within a method-specific order contrast,

Ua,πtot=Ua,πtot,Va,πeval=Va,πevalU^{\mathrm{tot}}_{a,\pi}=U^{\mathrm{tot}}_{a,\pi'}, \qquad V^{\mathrm{eval}}_{a,\pi}=V^{\mathrm{eval}}_{a,\pi'}

for every compared π\pi and π\pi', unless a count is itself a declared endpoint under a common ceiling. If early stopping is allowed, unused budget is reported; it is not silently transferred to tuning.

Across different methods, the optimizer mechanism may differ. Equality then means equal allowed update, search, evaluator, precision, and stopping budgets, not pretending that EWC, OGD, replay, PBT, and ordinary SGD execute identical operations. Method-specific work remains visible in the complete ledger.

Complete budget vector

Let Ndata=kNkdataN^{\mathrm{data}}=\sum_kN^{\mathrm{data}}_k be total presented exposure. Total accepted task-by-module exposure is

Nacc=kNa,π,kacc.N^{\mathrm{acc}} = \sum_k\sum_\ell N^{\mathrm{acc}}_{a,\pi,k\ell}.

Let NfwdN^{\mathrm{fwd}} and NbwdN^{\mathrm{bwd}} be forward- and backward-evaluation counts, and let BreadB^{\mathrm{read}}, BwriteB^{\mathrm{write}}, and BpeakB^{\mathrm{peak}} be attributed read, written, and peak-live byte counts. Let TworkerT^{\mathrm{worker}} be summed provisioned worker time in seconds, TwallT^{\mathrm{wall}} be elapsed run time in seconds, and EE be calibrated external energy in joules. UtotU^{\mathrm{tot}} and VevalV^{\mathrm{eval}} retain the definitions above.

Define

Ba,π=[Ndata,Nacc,Utot,Nfwd,Nbwd,Veval,Bread,Bwrite,Bpeak,Tworker,Twall,E]a,π.\mathbf B_{a,\pi} = \left[ N^{\mathrm{data}}, N^{\mathrm{acc}}, U^{\mathrm{tot}}, N^{\mathrm{fwd}}, N^{\mathrm{bwd}}, V^{\mathrm{eval}}, B^{\mathrm{read}}, B^{\mathrm{write}}, B^{\mathrm{peak}}, T^{\mathrm{worker}}, T^{\mathrm{wall}}, E \right]_{a,\pi}.

The first six entries are counts; the next three are bytes; TworkerT^{\mathrm{worker}} and TwallT^{\mathrm{wall}} are seconds; EE is joules. This vector is never summed directly. Paired order arms require componentwise equality within frozen tolerances or are compared on a preregistered Pareto frontier.

Mean measured power is derived only when a calibrated energy interval of duration Δt>0\Delta t>0 seconds exists:

P=EΔt.P=\frac{E}{\Delta t}.

Because J/s=W\mathrm{J}/\mathrm{s}=\mathrm{W}, PP is in watts. Operations, parameters, or bytes cannot replace EE in this equation.

Potential outcomes and order estimands

For endpoint qq, method aa, intervention vector z\mathbf z, and order π\pi, define the seed expectation

μa,π,z,q=Er[Ya,π,z,r,q],\mu_{a,\pi,\mathbf z,q} = \mathbb E_r \left[ Y_{a,\pi,\mathbf z,r,q} \right],

where the expectation is over the frozen seed generator, not over all possible tasks or machines.

Pairwise order effect

For two preregistered orders π\pi and π\pi', the controlled order effect is

τa,q(π,π;z)=μa,π,z,qμa,π,z,q.\tau_{a,q}(\pi,\pi';\mathbf z) = \mu_{a,\pi,\mathbf z,q} - \mu_{a,\pi',\mathbf z,q}.

Its unit is the endpoint unit. A positive value is beneficial only when higher values of endpoint qq are defined as better. Costs and errors retain their natural lower-is-better direction rather than having signs silently reversed.

The paired estimator over RR seeds is

τ^a,q(π,π;z)=1Rr=1R(Ya,π,z,r,qYa,π,z,r,q),\widehat\tau_{a,q}(\pi,\pi';\mathbf z) = \frac{1}{R} \sum_{r=1}^{R} \left( Y_{a,\pi,\mathbf z,r,q} -Y_{a,\pi',\mathbf z,r,q} \right),

where RR is a dimensionless seed count.

Order-distribution sensitivity

Let ΠK=K!|\Pi_K|=K!. The permutation-average endpoint is

μˉa,z,q=1K!πΠKμa,π,z,q.\bar\mu_{a,\mathbf z,q} = \frac{1}{K!} \sum_{\pi\in\Pi_K} \mu_{a,\pi,\mathbf z,q}.

The between-order variance is

Va,z,qorder=1K!πΠK(μa,π,z,qμˉa,z,q)2.V^{\mathrm{order}}_{a,\mathbf z,q} = \frac{1}{K!} \sum_{\pi\in\Pi_K} \left( \mu_{a,\pi,\mathbf z,q} -\bar\mu_{a,\mathbf z,q} \right)^2.

This is the finite-population variance for an order drawn uniformly from all K!K! permutations. If only m<K!m<K! orders are sampled uniformly without replacement, the preregistered sample estimator uses denominator m1m-1 and reports its sampling uncertainty; it is not silently substituted for the complete enumeration. If qq is a dimensionless score, the variance is squared score units. The order range is

Ra,z,qorder=maxπΠKμa,π,z,qminπΠKμa,π,z,q,R^{\mathrm{order}}_{a,\mathbf z,q} = \max_{\pi\in\Pi_K}\mu_{a,\pi,\mathbf z,q} - \min_{\pi\in\Pi_K}\mu_{a,\pi,\mathbf z,q},

in the endpoint unit. The range is descriptive and selection-biased as an estimate of a future best order; it cannot replace simultaneous pairwise intervals.

Position effect

For task identity kk and position jj, define

ηa,k,j,q(z)=1(K1)!πΠK:πj=kμa,π,z,q.\eta_{a,k,j,q}(\mathbf z) = \frac{1}{(K-1)!} \sum_{\pi\in\Pi_K:\,\pi_j=k} \mu_{a,\pi,\mathbf z,q}.

The position contrast

ηa,k,j,q(z)ηa,k,j,q(z)\eta_{a,k,j,q}(\mathbf z)-\eta_{a,k,j',q}(\mathbf z)

compares the same task at two positions averaged over all orders of the other tasks. It is not an “early-arrival law” unless it replicates across protected task families and survives the mechanism cuts.

For a plot or table that compares every position with the task's mean over positions, define

ηˉa,k,,q(z)=1Kj=1Kηa,k,j,q(z),ca,k,j,q(z)=ηa,k,j,q(z)ηˉa,k,,q(z).\bar\eta_{a,k,\cdot,q}(\mathbf z) = \frac{1}{K} \sum_{j'=1}^{K} \eta_{a,k,j',q}(\mathbf z), \qquad c_{a,k,j,q}(\mathbf z) = \eta_{a,k,j,q}(\mathbf z) - \bar\eta_{a,k,\cdot,q}(\mathbf z).

Thus jca,k,j,q(z)=0\sum_j c_{a,k,j,q}(\mathbf z)=0. If the only changed factor is the capacity-reservation cut, its position-specific interaction is

Γa,k,j,q(cap,pos)(zcap)=ca,k,j,q(zcap=1,zcap)ca,k,j,q(zcap=0,zcap).\Gamma^{(\mathrm{cap,pos})}_{a,k,j,q}(\mathbf z_{-\mathrm{cap}}) = c_{a,k,j,q}(z_{\mathrm{cap}}=1,\mathbf z_{-\mathrm{cap}}) - c_{a,k,j,q}(z_{\mathrm{cap}}=0,\mathbf z_{-\mathrm{cap}}).

Hypothetical centered position contrasts for shared and reserved capacity, with the resulting capacity-cut interaction

The figure is an algebraic reading aid. It substitutes constructed percentage- point contrasts for ca,k,j,qc_{a,k,j,q} and plots their constructed difference as Γa,k,j,q(cap,pos)\Gamma^{(\mathrm{cap,pos})}_{a,k,j,q}. The values are not measurements, estimates, predictions, recommended effect sizes, or evidence that a capacity mechanism exists. Its editable specification is the history-conditioned-position-contrast entry.

Optimized-order advantage

Let gg be a frozen order-selection algorithm, let π^g\widehat\pi_g be its selected order using discovery information only, and let Πrand\Pi^{\mathrm{rand}} be the uniform distribution over admissible orders. The held-out optimized advantage is

Δa,g,qopt=μa,π^g,z,qEπΠrand[μa,π,z,q].\Delta^{\mathrm{opt}}_{a,g,q} = \mu_{a,\widehat\pi_g,\mathbf z,q} - \mathbb E_{\pi\sim\Pi^{\mathrm{rand}}} \left[ \mu_{a,\pi,\mathbf z,q} \right].

The quality contrast is incomplete until the order optimizer's pilot examples, similarity or curvature measurements, candidate sequences, evaluator calls, failed runs, bytes, seconds, and joules are appended to Ba,π^g\mathbf B_{a,\widehat\pi_g}. Selecting the observed maximum from confirmation orders is not gg and is not a valid optimized-order estimate.

Factorial mechanism estimands

Define the binary intervention vector

z=[zage,zexp,zcap,zstate,zfac,zlock]{0,1}6,\mathbf z = \left[ z_{\mathrm{age}}, z_{\mathrm{exp}}, z_{\mathrm{cap}}, z_{\mathrm{state}}, z_{\mathrm{fac}}, z_{\mathrm{lock}} \right] \in\{0,1\}^{6},

where 00 is the natural-history level and 11 is the causal-cut level defined in the audit. Let zh\mathbf z_{-h} denote all factors except mechanism hh.

For one order π\pi, the controlled main contrast of mechanism hh at fixed zh\mathbf z_{-h} is

Δa,π,q(h)(zh)=μa,π,(zh=1,zh),qμa,π,(zh=0,zh),q.\Delta^{(h)}_{a,\pi,q}(\mathbf z_{-h}) = \mu_{a,\pi,(z_h=1,\mathbf z_{-h}),q} - \mu_{a,\pi,(z_h=0,\mathbf z_{-h}),q}.

The order-by-mechanism interaction for two orders is

Γa,q(h)(π,π;zh)=τa,q ⁣(π,π;zh=1,zh)τa,q ⁣(π,π;zh=0,zh).\Gamma^{(h)}_{a,q}(\pi,\pi';\mathbf z_{-h}) = \tau_{a,q}\!\left( \pi,\pi';z_h=1,\mathbf z_{-h} \right) - \tau_{a,q}\!\left( \pi,\pi';z_h=0,\mathbf z_{-h} \right).

Γ(h)\Gamma^{(h)} has the endpoint unit. If cutting a path reduces an order contrast toward zero, that is evidence that the measured effect depends on the cut under the frozen design. It is not proof that mechanism hh is the only mediator.

For distinct mechanisms hh and hh', the two-factor interaction at one order is the inclusion--exclusion contrast

Δa,π,q(h,h)=μ11μ10μ01+μ00,\Delta^{(h,h')}_{a,\pi,q} = \mu_{11}-\mu_{10}-\mu_{01}+\mu_{00},

where the two subscripts give (zh,zh)(z_h,z_{h'}) and all remaining factors are fixed. Higher-order interactions use the same inclusion--exclusion rule.

Decomposition limits

The factorial is randomized, but a unique additive causal decomposition is not generally available because:

  1. capacity changes which exposures are accepted;
  2. exposure changes optimizer and shared-state trajectories;
  3. shared state changes routing and therefore capacity demand;
  4. facilitation changes both information and subsequent work;
  5. lock-in changes the set of later admissible states; and
  6. nonlinearity allows higher-order interactions.

Consequently,

τa,q(π,π;0)hΓa,q(h)(π,π;zh)\tau_{a,q}(\pi,\pi';\mathbf 0) \ne \sum_h \Gamma^{(h)}_{a,q}(\pi,\pi';\mathbf z_{-h})

in general. The right side also depends on the levels at which the other factors are fixed. “Percent mediated” is not reported unless a separate causal model supplies and defends the required cross-world assumptions.

Ordinary-scheduling negative control

Let JkJ_k denote a task job that consumes the same declared compute and I/O but does not update parameters, routers, optimizer state, replay, normalization, or structure. Let σπ\sigma_\pi be its completion order under scheduler input π\pi. Scheduling may change a service endpoint L(σπ)L(\sigma_\pi) such as latency in seconds.

Separately, let uku_k be the complete checksummed learned update record produced for task kk. Apply all records after job completion in one canonical order κ\kappa to obtain

sa,Kcanon=uκKuκ1(sa,0).s^{\mathrm{canon}}_{a,K} = u_{\kappa_K}\circ\cdots\circ u_{\kappa_1}(s_{a,0}).

If sa,Kcanons^{\mathrm{canon}}_{a,K} and all final capability endpoints are identical across input permutations while only L(σπ)L(\sigma_\pi) differs, the detected effect is ordinary scheduling under this control. If update records themselves depend on the live history, canonical replay is diagnostic rather than an oracle; that dependence must be attributed to one or more serialized state paths.

Learning endpoints

Assume higher task score is better. Let Qa,π,kpostQ_{a,\pi,k}^{\mathrm{post}} be task kk's held-out score immediately after its acquisition block, and let Qa,π,kfinalQ_{a,\pi,k}^{\mathrm{final}} be its score after all KK blocks at the frozen retention horizon. Both retain the declared task score unit.

Backward transfer is

BWTa,π,k=Qa,π,kfinalQa,π,kpost.\operatorname{BWT}_{a,\pi,k} = Q_{a,\pi,k}^{\mathrm{final}} -Q_{a,\pi,k}^{\mathrm{post}}.

Positive values indicate improvement and negative values indicate forgetting. The nonnegative forgetting magnitude is

Fa,π,k=max(0,Qa,π,kpostQa,π,kfinal).F_{a,\pi,k} = \max\left( 0, Q_{a,\pi,k}^{\mathrm{post}} -Q_{a,\pi,k}^{\mathrm{final}} \right).

Let QkfloorQ_k^{\mathrm{floor}} be the preregistered protected score floor for task kk. The protected shortfall is

Ha,πprot=max1kKmax(0,QkfloorQa,π,kfinal).H_{a,\pi}^{\mathrm{prot}} = \max_{1\le k\le K} \max\left( 0, Q_k^{\mathrm{floor}} -Q_{a,\pi,k}^{\mathrm{final}} \right).

This has the score unit and cannot be cancelled by high performance on another task. The worst-task score is

Qa,πmin=min1kKQa,π,kfinal.Q_{a,\pi}^{\min} = \min_{1\le k\le K} Q_{a,\pi,k}^{\mathrm{final}}.

For newcomer task kk, let QkQ_k^{\star} be a frozen competence threshold and let na,π,kn_{a,\pi,k}^{\star} be the first accepted-example count at which the threshold is met and remains met for the frozen confirmation window. If the threshold is never met, na,π,kn_{a,\pi,k}^{\star} is right-censored at the task budget rather than deleted. Forward facilitation is evaluated through this sample-complexity endpoint and a matched no-history baseline, not through final score alone.

Structural and routing endpoints

Let pa,π,(t)p_{a,\pi,\ell}(t) be the fraction of routed load assigned to executing module \ell at time tt, with

pa,π,(t)0,=1Kpa,π,(t)=1p_{a,\pi,\ell}(t)\ge0, \qquad \sum_{\ell=1}^{K}p_{a,\pi,\ell}(t)=1

when at least one module is eligible. Router entropy is

Ha,πroute(t)==1Kpa,π,(t)logpa,π,(t),H^{\mathrm{route}}_{a,\pi}(t) = -\sum_{\ell=1}^{K} p_{a,\pi,\ell}(t) \log p_{a,\pi,\ell}(t),

in nats when the natural logarithm is used. Terms with p=0p=0 contribute zero. Low entropy is not by itself specialization or quality.

For a frozen task-by-module evaluation matrix Ga,π,kG_{a,\pi,k\ell}, where row kk is task identity and column \ell is module identity, define a dimensionless specialization contrast for module \ell as

Sa,π,=maxkGa,π,k1K1kkGa,π,k,S_{a,\pi,\ell} = \max_k G_{a,\pi,k\ell} - \frac{1}{K-1} \sum_{k'\ne k^{\star}_{\ell}} G_{a,\pi,k'\ell},

where kk^{\star}_{\ell} is the task attaining the maximum under a frozen tie rule. This metric is meaningful only if GG uses a common dimensionless score. If tasks have different native units, their raw matrix is reported and no such subtraction is allowed.

Admission time, commitment time, unlock count, retirement count, allocated capacity, accepted load, dropped load, and lineage remain separate endpoints. An expert label or high SS does not prove independent function, causal necessity, or efficient routing.

Service and lifecycle endpoints

If NdueN^{\mathrm{due}} predictions are due and NacceptedN^{\mathrm{accepted}} are correct, current, integrity-valid, and within the frozen deadline, accepted service is

Sacc=NacceptedNdue,S^{\mathrm{acc}} = \frac{N^{\mathrm{accepted}}}{N^{\mathrm{due}}},

a dimensionless fraction. Missing, late, stale, duplicated, and inaccurate predictions remain separate counts before aggregation.

Latency LiL_i for due item ii is completion time minus original due time, in seconds. Every due item remains in the denominator; missing completions receive the preregistered right-censor value. p50, p95, and p99 are reported with the exact quantile convention.

The mandatory result is a vector,

Va,π=[{Qa,π,kfinal}k=1K,Qa,πmin,Ha,πprot,{BWTa,π,k}k=1K,{na,π,k}k=1K,Sacc,L0.99,Ba,π].\mathbf V_{a,\pi} = \left[ \{Q_{a,\pi,k}^{\mathrm{final}}\}_{k=1}^{K}, Q_{a,\pi}^{\min}, H_{a,\pi}^{\mathrm{prot}}, \{\operatorname{BWT}_{a,\pi,k}\}_{k=1}^{K}, \{n_{a,\pi,k}^{\star}\}_{k=1}^{K}, S^{\mathrm{acc}}, L_{0.99}, \mathbf B_{a,\pi} \right].

The entries have different units and are not averaged. Method aa dominates method aa' only under preregistered direction and relevance margins, with no protected endpoint worse and at least one endpoint materially better.

Multiplicity and uncertainty

Let Hprimary\mathcal H_{\mathrm{primary}} be the frozen family of primary order, method, and order-by-mechanism contrasts. Its membership is committed before confirmation outcomes are opened.

For randomization inference, compute one test statistic for each hHprimaryh\in\mathcal H_{\mathrm{primary}} and, under the registered treatment re-randomizations, use the maximum absolute standardized statistic

Mperm=maxhHprimaryThperm.M^{\mathrm{perm}} = \max_{h\in\mathcal H_{\mathrm{primary}}} \left|T_h^{\mathrm{perm}}\right|.

The empirical distribution of MpermM^{\mathrm{perm}} supplies family-wise adjusted decisions and simultaneous intervals. If that procedure is computationally unavailable, Holm's step-down correction is the fallback. Unadjusted effect estimates and intervals remain visible, but they cannot carry the confirmatory decision.

Seeds are paired experimental units only for the generator they instantiate. Tasks, orders, endpoints, or repeated checkpoints from one run are not treated as independent sample-size multipliers. Generalization to a task population or machine population requires corresponding sampled levels and a second-family or second-machine replication.

The minimum relevant effect δqmin\delta_q^{\min} is declared in endpoint qq's native unit. “Statistically nonzero” without crossing δqmin\delta_q^{\min} does not keep the frontier alive.

Dimensional analysis checklist

  1. Task scores may be dimensionless proportions or native score units; the unit is declared before subtraction.
  2. Counts of examples, updates, parameters, modules, calls, and operations are dimensionless counts with different meanings and are not interchangeable.
  3. Storage and traffic are bytes; parameter counts become bytes only after representation width and metadata are included.
  4. Latency, worker time, and wall time are seconds but have different boundaries and remain separately named.
  5. Energy is joules and power is joules per second, or watts.
  6. Throughput is accepted items per second and is not an energy efficiency.
  7. Quality per joule is reported only beside its raw quality and energy components at matched task, quality floor, horizon, and boundary.
  8. Ecological abundance and artificial router load are not assigned a common unit merely to make an analogy.

Testable predictions

These are preregistrable hypotheses, not findings:

  1. In high-overlap task strata, the capacity reservation cut yields Γ(cap)\Gamma^{(\mathrm{cap})} opposite in sign to the natural order contrast if incumbency is carried by capacity pre-emption.
  2. If shared-state modification is a carrier, the magnitude of τ(π,π;z)\tau(\pi,\pi';\mathbf z) decreases when zstate=1z_{\mathrm{state}}=1 while private post-acquisition scores remain within their non-inferiority margins.
  3. If typed facilitation improves newcomer acquisition, setting zfac=1z_{\mathrm{fac}}=1 increases the newcomer threshold count nn^{\star} or right-censoring rate without a corresponding change under the sham-only scheduler control.
  4. If lock-in is useful rather than merely persistent, the irreversible arm improves protected delayed outcomes after the inducing module is removed and remains non-dominated after checkpoint, unlock, migration, and recovery costs enter B\mathbf B.
  5. If the effect is ordinary continual-learning interference, replay, EWC, or OGD matches the complete frontier without module-birth factors.
  6. If the effect is curriculum selection, a held-out optimized order yields the same frontier after its search budget is charged.
  7. VorderV^{\mathrm{order}} and the sign of position contrasts vary with feature overlap, readout conflict, horizon, and method; no universal early/late rule is predicted.

Numerical and contract checks

  1. Every stored permutation is a bijection and has the same task-multiset and eligible-module-set digests.
  2. Every task identity has identical exogenous example and update opportunity digests across orders.
  3. Capacity inequalities hold at every event, not only at final state.
  4. Every shared-state reset restores the same checksummed boundary image while retaining only the explicitly exempt private state.
  5. Sham facilitation payloads match byte count, availability time, parser work, and transport work, and are provably answer-free.
  6. Canonical replay applies exactly the same update-record multiset in the same canonical order for every scheduling control.
  7. Invalid, missing, censored, timed-out, and resource-exceeded runs remain in result tables and multiplicity families.
  8. The full budget vector is reconstructed from raw events before any method comparison.
  9. Half-precision, quantized, sparse, or parallel execution receives its own parameter, byte, work, and energy accounting; nominal operation equality is insufficient.
  10. No equation in this note authorizes a scientific, performance, or energy result before a registered executable package and sealed execution exist.

Kill boundary

This mathematical contract supplies no residual architecture contribution if:

  1. for every protected primary contrast on both task families, the registered simultaneous interval lies wholly inside (δqmin,+δqmin)(-\delta_q^{\min},+\delta_q^{\min}); the unobserved population condition τ<δqmin|\tau|<\delta_q^{\min} is the corresponding no-effect region, not an executable decision rule;
  2. canonical replay or the non-learning scheduler reproduces the effect;
  3. equal age plus equal accepted exposure removes it;
  4. no randomized causal cut changes it on fresh seeds;
  5. a complete conventional null is non-inferior at lower or equal lifecycle cost;
  6. results depend on selecting the best confirmation permutation;
  7. protected-task harm or resource ceilings are violated; or
  8. the second task family or machine replication fails.

Passing these equations and checks would justify evaluating evidence. It would not by itself establish a biological equivalence, universal mechanism, or novel architecture.