Research portal

Mathematical note

Opportunity- and history-qualified adaptive action

math/opportunity-history-qualified-action.md

Edition
Site v0.3.0 · continuous main snapshot
Source revision
ec2865b0eac15148675c629981a545632b3571c5
Extent
1,075 words
Public route
https://www.cordana.dev/math/opportunity-history-qualified-action/
Mapped records1 mapped record

Direct repository links only; no document-level evidence status is implied.

This note defines the comparison boundary for Fixture F-003. The fixture uses comparative-cognition tasks as diagnostic interventions, not as species rankings. Its central requirement is simple: performance can be compared only after the information, action, learning, and reward opportunities that produced it are explicit.

Episode contract

For episode ee, record

Ωe=(Xe,Se,Ae,He,Re,Ce,τe,Ue),\Omega_e=(X_e,S_e,A_e,H_e,R_e,C_e,\tau_e,U_e),

where XeX_e is task and apparatus state, SeS_e is the physically available sensory channel, AeA_e is the feasible realized-action set, HeH_e is training, rearing, prior-task, and social-exposure history, ReR_e is the reward, punishment, deprivation, and stopping rule, CeC_e is the intervention and control set, τe\tau_e is the relevant time in seconds, and UeU_e is the sampling unit such as episode, agent, dyad, or population. All fields except τe\tau_e are typed records or sets rather than unit-bearing scalars.

The learner receives the causally available prefix

Ie(t)={sSe:tsrecvt}HetRet,\mathcal I_e(t)= \{s\in S_e:t_s^{\mathrm{recv}}\le t\} \cup H_e^{\le t} \cup R_e^{\le t},

where tt and receipt time tsrecvt_s^{\mathrm{recv}} are seconds. Handler cues, apparatus sounds, demonstrator traces, residual odor, simulator metadata, and training-phase identifiers belong in SeS_e when they are actually available. Protocol intention is not a measurement of the input channel.

For requested action aa and plant version vv, feasibility is

Fv(a,Ωe){0,1},F_v(a,\Omega_e)\in\{0,1\},

and realized action follows

a~tpv(a~tat,Ae,Xe,He).\widetilde a_t\sim p_v(\widetilde a_t\mid a_t,A_e,X_e,H_e).

FvF_v is dimensionless. Requested and realized actions retain their native units, such as metres, radians, newtons, newton-metres, or a discrete tool identifier. A nominally identical task is not matched when one plant cannot sense while holding the required tool or cannot execute the required motion.

Opportunity-qualified performance

For method mm, success event Ye{0,1}Y_e\in\{0,1\}, and declared opportunity stratum ω\omega, define

Qm(ω)=Pr(Ye=1do(m),Ωeω).Q_m(\omega)= \Pr(Y_e=1\mid do(m),\Omega_e\in\omega).

QmQ_m is dimensionless. The paired contrast against baseline m0m_0 is

Δm,m0(ω)=Qm(ω)Qm0(ω).\Delta_{m,m_0}(\omega)=Q_m(\omega)-Q_{m_0}(\omega).

Report Δ\Delta by morphology, actual channel, history, reward schedule, demonstrator exposure, site, and intervention family. A pooled contrast is admissible only after heterogeneity is shown rather than assumed away.

Use a transport contrast to test whether the learned relation survives a controlled change gg:

Tg(m)=Qm(g(ω))Qm(ω).T_g(m)=Q_m(g(\omega))-Q_m(\omega).

TgT_g is dimensionless. Useful changes include material substitution, geometry change, causal inversion, perceptual-cue reversal, morphology or tool swap, history swap, demonstrator removal, partner turnover, and delayed future use. The sign is interpreted relative to a preregistered invariance: some changes should preserve performance; inversions should reverse the selected action.

Transfer lattice

Do not compress transfer into one difficulty axis. For task family kk, retain

Tk=(Tsurface,Tmaterial,Trelation,Thistory,Tplant,Tsocial,Tdelay).\mathbf T_k= (T_{\mathrm{surface}},T_{\mathrm{material}},T_{\mathrm{relation}}, T_{\mathrm{history}},T_{\mathrm{plant}},T_{\mathrm{social}}, T_{\mathrm{delay}}).

Every component is a dimensionless paired effect. Surface and material changes test perceptual or affordance generalization; relation changes test functional sensitivity; history and plant changes test dependence on acquisition and embodiment; social changes identify copied content; and delay changes test the lifetime of retained state. The vector prevents terminal success on one familiar apparatus from standing in for causal, prospective, or cross-plant transfer.

First-trial transfer after the held change is reported separately:

Tg(1)(m)=Ym,g,1Ym,g0,1.T_g^{(1)}(m)=Y_{m,g,1}-Y_{m,g_0,1}.

Later trial curves estimate adaptation, not prior transfer. Both are retained.

Social acquisition decomposition

A demonstration dd is represented as

d=(f,y,γ,ι,ρ,p),d=(f,y,\gamma,\iota,\rho,p),

where ff is action form, yy is end state, γ\gamma is trajectory, ι\iota is demonstrator identity and reliability, ρ\rho is the demonstrated functional relation, and pp is provenance. These are typed variables.

For component j{f,y,γ,ι,ρ}j\in\{f,y,\gamma,\iota,\rho\}, its causal uptake effect is

Ij=E[Ldo(dj=dj),dj]E[Ldo(dj=dj0),dj],I_j= \mathbb E[L\mid do(d_j=d_j'),d_{-j}] -\mathbb E[L\mid do(d_j=d_j^0),d_{-j}],

where LL is a declared dimensionless learner outcome and djd_{-j} holds the other components fixed. Ghost, result-only, novel-action, inefficient-action, and causal-relevance controls approximate these interventions. Matching the end state does not establish copying of action form.

Prospective and event memory

For stored event or resource ii,

ei=(wi,i,ti,qi,ci,pi),e_i=(w_i,\ell_i,t_i,q_i,c_i,p_i),

where wiw_i is content, i\ell_i location, tit_i time in seconds from a declared origin, qiq_i quality or perishability state, cic_i social or task context, and pip_i provenance. At future time tt, a conventional reservation null uses

Vi(t)=piavail(tei)ri(t)ki(t),V_i(t)=p_i^{\mathrm{avail}}(t\mid e_i)r_i(t)-k_i(t),

where availability probability is dimensionless and reward rir_i and retrieval/carrying cost kik_i use the same declared utility unit. A proposed prospective trace earns credit only beyond value tables, successor representations, POMDP planning, and retrieval with equal state and rollout budgets.

Bound-event memory is tested against factorized semantic and spatial state. For query qq over unique event eie_i, report exact-answer rate, calibration, update propagation, retrieval latency in seconds, and memory bytes. Subjective recollection is not an observable in this contract.

Costed uncertainty control

For optional observation or query zz with cost czc_z, the ordinary value-of-information null is

VOI(zb)=Ez ⁣[maxaE[U(a,θ)b,z]]maxaE[U(a,θ)b]cz,\operatorname{VOI}(z\mid b)= \mathbb E_z\!\left[\max_a\mathbb E[U(a,\theta)\mid b,z]\right] -\max_a\mathbb E[U(a,\theta)\mid b]-c_z,

where belief bb and latent state θ\theta are dimensionless, while utility UU and czc_z use the same declared unit. The policy acquires zz only when VOI is positive. Risk–coverage–cost surfaces must also include calibrated confidence, ensembles, conformal/selective prediction, learned difficulty cues, and response-strength policies.

Local, central, and mechanical contribution

For a compliant plant with nn local segments, let the central policy send message gtg_t and segment ii receive local observation oi,to_{i,t}. The realized control is

ui,t=πi(oi,t,gt,hi,t;v),u_{i,t}=\pi_i(o_{i,t},g_t,h_{i,t};v),

where local state hi,th_{i,t} is dimensionless, plant version is vv, and control ui,tu_{i,t} retains its native actuator unit. Compare this with a centralized policy receiving the same causally available observations and with a passive- mechanics arm receiving no learned local state.

Communication load over horizon [0,T][0,T] is

Bcomm(T)=m:tmTbytes(m)[bytes],B_{\mathrm{comm}}(T)=\sum_{m:t_m\le T}\operatorname{bytes}(m) \quad[\mathrm{bytes}],

and total control energy is

Econtrol=Esense+Ecentral+Elocal+Ecomm+Eact+Eadapt+Erecover[J].E_{\mathrm{control}}= E_{\mathrm{sense}}+E_{\mathrm{central}}+E_{\mathrm{local}} +E_{\mathrm{comm}}+E_{\mathrm{act}}+E_{\mathrm{adapt}} +E_{\mathrm{recover}} \quad[\mathrm{J}].

Peripheral credit requires a quality, recovery, latency, or energy gain after passive compliance, actuator work, sensor power, communication, and local hardware are charged. Anatomical distribution alone supplies no credit.

Outcome and lifecycle boundary

Keep the confirmatory result as a vector:

Y=(Q,T(1),T,Nunsafe,Nattempt,L50,L95,Bmemory,Bcomm,Hhuman,Elife).\mathbf Y= (Q,T^{(1)},\mathbf T,N_{\mathrm{unsafe}},N_{\mathrm{attempt}}, L_{50},L_{95},B_{\mathrm{memory}},B_{\mathrm{comm}},H_{\mathrm{human}}, E_{\mathrm{life}}).

QQ, T(1)T^{(1)}, and T\mathbf T are dimensionless; unsafe events and attempts are counts; latencies L50L_{50} and L95L_{95} are seconds; memory and communication are bytes; human effort HhumanH_{\mathrm{human}} is person-hours; and lifecycle energy ElifeE_{\mathrm{life}} is joules.

For method mm over all development and confirmatory work,

Elife(m)=Etrain+Edemonstrate+Esearch+Einfer+Einteract+Ecommunicate+Estore+Eadapt+Erecover[J].E_{\mathrm{life}}^{(m)}= E_{\mathrm{train}}+E_{\mathrm{demonstrate}}+E_{\mathrm{search}} +E_{\mathrm{infer}}+E_{\mathrm{interact}}+E_{\mathrm{communicate}} +E_{\mathrm{store}}+E_{\mathrm{adapt}}+E_{\mathrm{recover}} \quad[\mathrm{J}].

The fixture accepts the composed residual only if it improves a preregistered subset of literal outcomes without violating non-inferiority margins on the others, survives at least two task families, plant geometries, histories, model families, and hardware classes, and beats the complete conventional null stack at the same episodes, interventions, search, storage, communication, human effort, and energy. Otherwise retain the observation contract and retire the architectural story.

Editable system diagram: opportunity-history-qualified-action.mmd.