This note formalizes Fixture F-009 from the acoustics, hearing, and auditory-scene analysis audit. It supplies a hostile comparison boundary for Candidate 002, Candidate 006, Candidate 007, Candidate 009, Candidate 012, and Candidate 014.
Episode, operator, and action identity
For receiver , source , microphone or ear channel , and episode , preserve
where:
- is geometry, medium, boundaries, temperature in kelvins, relative humidity as a dimensionless fraction, flow in metres per second, occupancy, and time-varying room/material state;
- is every target and interfering source identity, waveform, position in metres, velocity in metres per second, directivity, emission time in seconds, and source-level record;
- is commanded and realized active emission: waveform, spectrum, duration in seconds, acoustic energy in joules, aim in degrees, repetition, and dose;
- is receiver/body/array position and orientation, aperture in metres, morphology, feasible motion, transducer response, gain, health, and saturation;
- is prior acoustic exposure, room/source history, adaptation, training, previous emissions and motions, and feedback with timestamps;
- is the observation operator: impulse responses, transfer functions, spatial/time support, sample rate in samples per second, quantization in bits, clock, latency in seconds, preprocessing, missingness, and selection;
- is calibration identity, uncertainty, reference pressure and distance, instrument class, traceability, and data vintage;
- is the target construct, literal outcome, deadline in seconds, and loss or utility in declared units;
- is the independent unit: waveform, frame, event, source, room, receiver, body, array, device, site, or population; and
- is the complete ceiling in samples, events, bytes, seconds, person-hours, joules, emissions, exposure, unsafe events, sensors, actuators, replacements, and opportunity.
For method and literal outcome , the estimand is
where uses the registered unit for outcome . A contrast against baseline does not isolate if aperture, calibration, room response, source distribution, observation receipt time, emission or motion authority, training scenes, or any binding resource differs without intervention.
Pressure, level, exposure, and clock contract
For acoustic pressure in pascals over interval in seconds,
where is RMS pressure in pascals, is the reference pressure in air, and is in decibels relative to . Frequency weighting, time weighting, band, location, orientation, and calibration are fields, not implied defaults.
Sound exposure and its level are
where is in pascal-squared seconds, is in decibels relative to , and reference duration . Exposure is not acoustic emission energy, electrical energy, or a universal risk model.
For channel clock , correct a device timestamp by
where device time , corrected time , and version- offset are in seconds; is the channel-clock index, is the dimensionless sample index, and is the dimensionless calibration version. Retain residual clock uncertainty in seconds, drift in seconds per second, capture time, receipt time, synchronization method, and calibration interval. At decision time , only observations with receipt time no later than are available.
Propagation, rooms, and time-varying mixing
At channel , use the time-varying mixture
where received waveform , source pressure , and noise are in pascals; is dimensionless source count; and delay are in seconds; and transfer kernel is in reciprocal seconds. A time-invariant room convolution is a special case. Source/receiver motion, changing boundaries, temperature, flow, occupancy, and adaptive emitters make the operator time-dependent.
For an ideal free-field point source with unchanged directivity and negligible absorption,
where distances are in metres and is in decibels. The comparison must expose directionality, near-field terms, boundaries, atmospheric absorption, scattering, and receiver orientation when present.
For a diffuse-field Sabine approximation,
where is decay time in seconds, room volume is in cubic metres, equivalent absorption area is in square metres, and has units seconds per metre. Report measured impulse responses and uncertainty by band, source, receiver, and support; this approximation is not an operator identity.
Masking, filterbank, compression, and events
A real gammatone-like channel is
where centre frequency and bandwidth are in hertz, filter order is dimensionless, phase is in radians, time is in seconds, and has units so is in reciprocal seconds. Fixed FFT, wavelet, mel, ERB, gammatone/gammachirp, modulation, and learned filterbanks remain competing representations.
A local normalized compression fit is
where magnitudes share one input unit, share one output unit, and exponent is dimensionless. Register frequency, level, state, attack/release time in seconds, saturation, and distortion; one exponent does not describe an adaptive cochlea or compressor globally.
For yes/no target detection,
where both terms are dimensionless probabilities, is the inverse standard-normal cumulative distribution, and sensitivity is dimensionless. Keep decision criterion, energetic overlap, informational uncertainty, spatial release, source grouping, and task history separate.
For event encoder threshold in the unit of channel response ,
where is channel identity, event time is in seconds, is the dimensionless event index, is the retained event record, is the previous retained-event time in seconds, and are dimensionless threshold and clock versions. Event rate is in events per second. Report missed sustained signals, false events, timestamp error, decoder cost, bytes, task information, and measured joules; event sparsity is not an energy unit.
ITD, ILD, aperture, correlation, and beamforming
For sound speed in metres per second and path-length difference in metres,
where interaural or interchannel time difference is in seconds. Clock offset, phase ambiguity, multipath, source extent, and head/array transfer functions are part of its uncertainty.
For left/right RMS pressures in pascals,
where interaural level difference is in decibels and is band-, time-, source-, and orientation-qualified. ITD, ILD, spectral cues, and head motion remain separate intervention axes.
For far-field array aperture in metres, frequency in hertz, and wavelength in metres,
where angular resolution is in radians. Geometry, beampattern, SNR, estimator, bandwidth, and resolution definition determine the coefficient; no learned estimator restores spatial frequencies excluded by the physical support without additional prior information.
For signals in pascals, define generalized cross-correlation
where are Fourier transforms in pascal-seconds, frequency is in hertz, delay and estimate are in seconds, is the feasible delay set in seconds, superscript is complex conjugation, and weighting has reciprocal pascal-squared-second-squared units so is in hertz after integration. The peak must carry a calibrated multimodal delay distribution under reverberation, not only an argmax.
For array snapshot in pascals, dimensionless steering vector , and noise covariance in pascal-squared units,
where is microphone count, superscript is conjugate transpose, is dimensionless, and output is in pascals. Steering mismatch, covariance error, short sample support, correlated sources, motion, and room change receive sealed interventions.
Reverberation, separation, and source identity
For direct component , early-reflection component , late component , and noise , all in pascals,
The early/late boundary is a registered time in seconds relative to the direct arrival. Waveform dereverberation, speech compensation, perceived distance, localization, and downstream intelligibility are distinct outcomes.
For estimated source and reference source with the same amplitude unit and sample count ,
where is the waveform inner product and SI-SDR is in decibels. Report unknown source count, permutation, causal identity, localization, intelligibility, calibration, perceptual/task quality, and residual mixture separately.
For predicted scene outcome and causally available acoustic observation , proper log loss is
where is independent event count and is in bits per event. Use it for detection, source count, assignment, localization bins, range bins, and abstention under held-out operators; do not pool incompatible outcomes.
Active emission, dose, and closed-loop action
For monostatic emission at and echo receipt at , both in seconds,
where estimated range is in metres and sound speed is in metres per second. Clock uncertainty, target motion, refraction, multipath, transducer ringing, waveform ambiguity, association, and detector threshold must propagate to range uncertainty.
For acoustic output power in watts over emission duration in seconds,
where is acoustic joules. Electrical input, transduction loss, body/sensor motion, sensing, compute, cooling, detectability, interference, and exposure remain separate axes.
At decision time , a complete acoustic policy is
where contains waveform, level, spectrum, aim, and timing; contains pose or head/body motion in metres, radians, and seconds; contains aperture, sample rate, and channel selection; is dimensionless or in declared decibels; is the downstream command in its native unit; is causally received history; is scene state; is operator/calibration state; is uncertainty; is the feasible action envelope; and is remaining budget in matched units. Attention- or efferent-like gain receives causal credit only when intervening on it changes its registered endpoint.
Hardware, lifecycle energy, and equal budgets
Lifecycle energy is
where every term is in joules over one service interval and denotes data acquisition, training, acoustic emission, sensing, inference, physical motion, communication, storage, calibration, maintenance, and amortized embodied energy. DSP, FPGA, ASIC, neuromorphic, CPU, GPU, and analog front ends are measured at identical physical boundaries and workload deadlines.
Human effort is
where every term is in person-hours and roles are reported separately.
The complete cost vector is
where the five terms count samples, emitted/received events, optimization or environment steps, queries, and bytes; is seconds; is person-hours; is joules; is pascal-squared-second exposure; counts unsafe events; uses a registered physical or severity unit; and is opportunity cost in task exposures or person-hours.
Method is feasible only when
where is the preregistered componentwise ceiling with identical units. Over-budget runs are infeasible; removed ablation resources stay unused.
Confirmatory contrast and retirement
Let be the strongest mature baseline for track , frozen on development data. Orient every protected endpoint as a preregistered benefit, so larger is better and a non-inferiority margin may be negative. For protected outcomes , retain a residual only when
where is an improvement or non-inferiority margin in the unit of outcome , and is the dimensionless error budget. The result must survive held-out rooms, sources, arrays, bodies, operators, scenes, source counts, interference policies, model families, sites, and hardware classes at equal complete cost. Otherwise retire the architectural residual while keeping the operator/action/measurement contract.