This note formalizes Fixture F-006 from the sports expertise, adaptive performance, and team coordination audit. It supplies a comparison contract for Candidate 002, Candidate 004, Candidate 006, Candidate 007, Candidate 009, Candidate 012, Candidate 014, and Candidate 019. It creates no new candidate.
Versioned episode state
For agent in episode , preserve
where:
- is the physical state, rules, task geometry, deadline, consequence, and hidden regime of episode ;
- is the information actually received by agent , including source, support, latency in seconds, occlusion, noise, loss, and calibration;
- is the feasible action set under the current body, actuator, equipment, authority, rate, range, and safety constraints;
- is the timestamped acquisition and selection history: practice, feedback, opponents, teammates, injury or faults, prior exclusions, and opportunities that were offered or withheld;
- is the teammate and opponent roster, role assignment, policy history, turnover event, and communication topology;
- is the feedback channel, its delay in seconds, its information in bits, and the party choosing when it is supplied;
- is the resource and fatigue state, with each component stored in its native unit rather than collapsed into a readiness score;
- is the damage or fault state, diagnostic uncertainty, protected capability envelope, and current return stage;
- is the selection and opportunity policy that determines access to training, roles, observation, intervention, and later outcome measurement;
- is the randomized intervention, control, counterfactual pair, and stopping rule;
- is the sampling unit, such as action, possession, episode, agent, dyad, team, site, season, or cohort; and
- is the complete resource ceiling: events, bytes, seconds, person-hours, joules, damage, unsafe events, replacements, and opportunity.
The comparison estimand for method and literal outcome is
where is measured in the registered unit for outcome . A contrast against baseline is uninterpretable when any element of differs without a registered intervention or adjustment.
Outcome firewall
No scalar “performance” score may replace the following vector:
The components and their units are:
| Symbol | Outcome | Required literal measurement and unit |
|---|---|---|
| anticipation | proper predictive score in bits per event, calibration error dimensionless, commitment latency in milliseconds | |
| physical interception | success probability dimensionless, endpoint error in metres, movement onset in milliseconds, unsafe-event probability dimensionless | |
| cue use | causal score change under a declared cue intervention, in bits per event or the registered task unit | |
| practice performance | literal task quality by attempt number and exposure time in seconds | |
| delayed retention | task quality after a delay in hours or days without the training scaffold | |
| transfer | source-to-target task quality and gap in the literal task unit | |
| exploration | action and outcome entropy in bits, action--outcome information in bits, coverage dimensionless, and later utility | |
| adaptability and recovery | perturbation loss, time to regain the envelope in seconds, overshoot, recurrence probability, and residual damage | |
| pacing | power in watts or action intensity in its declared unit as a time series, plus terminal task quality | |
| fatigue and readiness | task-specific capacity change, state-estimation error, calibrated admissibility, abstention, and recovery time | |
| staged return | false promotion, false withholding, stage dwell time in hours, recurrence, rollback, availability, and collateral loss | |
| coordination and shared information | team task quality, task-variable variance, compensation lag in seconds, belief log loss in bits per event, messages, bytes, cross-play, and repair latency | |
| deception and opponent adaptation | opponent log loss in bits per action, calibration, exploitability, regret, abstention utility, and adaptation time | |
| talent prediction | prospective calibration, false-negative recovery, later capability, attrition, opportunity received, and subgroup error | |
| complete efficiency | events, bytes, wall-seconds, person-hours, joules, equipment, damage, unsafe events, replacements, and opportunity cost as separate axes |
Representative distance is a vector
Let and be the training and target distributions. Register
where compares received information, feasible actions, deadline and consequence, teammate/opponent composition and policy, feedback, resource state, damage/return state, and selection/opportunity policy. Each is a declared divergence between the corresponding marginals or conditionals under and . It is dimensionless for a statistical divergence and has the registered ground-cost unit for optimal transport. No unreported weighted sum is a valid “representativeness” score.
For source stratum and target stratum , preserve the transfer matrix
where both and the transfer gap use the registered task unit. Report separately by cue, feasible action, opponent, feedback, resource, damage, and selection-policy changes.
Anticipation, interception, and cue use
For independent events, outcome , actual observation history available by occlusion time in seconds, and predictive distribution , define
where is log loss in bits per event. For information channel , the registered causal cue value is
also in bits per event. Here is measured under removal or neutralization of channel , not inferred from gaze or saliency. Predictive regulation must additionally retain false-alarm action, reserve debit, recovery, and cumulative exposure instead of treating cue value as the whole outcome (C-1494).
Physical coupling is reported separately as
where is dimensionless interception success, is endpoint error in metres, is movement onset in milliseconds, and is dimensionless unsafe-event probability. Label or joystick accuracy cannot substitute for .
Practice, retention, transfer, and exploration
Let be literal task quality after attempt and exposure time in seconds. Keep three estimands:
where is the scaffold-free retention delay in hours or days. The first target trial is frozen before any target update; later adaptation is a separate curve.
For action variable and reached-outcome variable , exploration is
where both entropies and mutual information are in bits, is coverage of the registered feasible region as a dimensionless fraction, is later transfer in its task unit, and is the separate cost vector. Higher action entropy without outcome information or later utility is not useful exploration.
After a perturbation at time , define recovery time
where and the stability horizon are in seconds, is task quality in its registered unit, and is the preregistered admissible quality envelope. Overshoot, recurrence, and damage are additional axes rather than hidden inside .
Resource state, pacing, and readiness
Let external power be in watts over event duration in seconds. The external work is
where is in joules. Metabolic, device, facility, embodied, and lifecycle energy use different boundaries and remain separate ledger rows.
A resource-qualified controller has the form
where is the commanded action or power target in its native unit, is causally received observation history, is the estimated resource/fatigue vector, is the estimated damage state, is remaining work in metres, seconds, events, or joules, is the opponent-policy estimate, is the teammate-state estimate, and is available feedback. Each estimate and channel receives its own ablation.
Readiness is a calibrated action envelope rather than a score:
where is the feasible action set, is the multidomain outcome vector over horizon in hours, is the registered safe envelope, is information available at decision time, and is the dimensionless tolerated risk. Empty envelopes require abstention or escalation.
Staged and reversible return
Let denote protected, modified, controlled, full-load, and adversarial operation. Promotion is admissible only if
where is the multidomain admissible envelope for stage , is the follow-up horizon in hours or days, and is its dimensionless risk tolerance. If the current envelope is violated, the gate must allow
Report false promotion, false withholding, dwell time in hours, recurrence, rollback count, availability, damage, and human adjudication hours separately.
Team coordination and shared information
For task variable and agent contribution , perturb agent by and estimate
where is lag in seconds and is task state. Compensation requires both a registered response in and reduced task-variable error. Here and are recorded in their native contribution units, is a registered perturbation in the unit of agent 's action or state, and has the product unit of the two contributions; correlation or synchrony without intervention is insufficient. Common-drive and edge-intervention controls are therefore mandatory before mapped synchrony can be credited with useful coordination (C-1495).
For teammate 's future action or intent and agent 's predictive belief , shared-information quality is
where is in bits per event, is information actually available to , and is the number of independent team events. Report it under message ablation, teammate turnover, role reassignment, and never-co-trained cross-play alongside messages, bytes, latency, repair, and task quality.
Deception and opponent adaptation
For opponent action , available history , and estimate , define
in bits per opponent action. For matched genuine and deceptive interventions,
where names a literal outcome and the difference retains its unit. Report calibration, confidence, exploitability, regret, abstention utility, and adaptation time separately under known, held-out, changing, and colluding opponents.
Selection, opportunity, and prospective prediction
Let denote selection, preselection evidence, postdecision opportunity in hours or task exposures, and later capability in its task unit. The prospective selection-policy estimand is
It cannot be estimated by comparing selected survivors with excluded agents when selection changes , coaching, opponents, follow-up, attrition, or injury exposure. Report prospective calibration in new cohorts, selection and opportunity rates, false-negative recovery, attrition, censoring, subgroup error, later capability, and complete development cost.
Complete efficiency and equal budgets
For method , retain lifecycle energy
where every term is in joules under one declared service interval. The terms denote training, inference, sensing, actuation, communication, facility, recovery, maintenance, and amortized embodied energy, respectively.
Human effort is
where each term is in person-hours and roles are reported separately. The complete cost vector is
where the four terms count events, environment or optimization steps, queries, and bytes; is wall time in seconds; is person-hours; is joules; counts unsafe events; is damage in a registered physical or severity unit; and is withheld opportunity in task exposures or person-hours.
Method is feasible only if
where is the preregistered componentwise ceiling with the same units. A complete efficiency claim requires non-inferiority on every protected outcome and a Pareto improvement on at least one preregistered resource axis. An over-budget run is infeasible, not a score to normalize afterward.
Confirmatory contrast and retirement
Let be the strongest mature baseline for track , selected on development data before confirmatory outcomes open. For protected outcome set , retain a residual only when
where is the preregistered improvement or non-inferiority margin in the unit of outcome , and is the dimensionless error budget. The contrast must survive actual-channel, feasible-action, history, opponent/team, feedback, resource, damage, selection, and complete-cost ablations on held-out task, model, site, and hardware strata. Otherwise retire the mechanism claim while preserving the measurement contract.