Research portal

Mathematical note

Operator-qualified active acoustic inference

math/operator-qualified-acoustic-inference.md

Edition
Site v0.3.0 · continuous main snapshot
Source revision
ec2865b0eac15148675c629981a545632b3571c5
Extent
1,919 words
Public route
https://www.cordana.dev/math/operator-qualified-acoustic-inference/
Mapped records8 mapped records

Direct repository links only; no document-level evidence status is implied.

This note formalizes Fixture F-009 from the acoustics, hearing, and auditory-scene analysis audit. It supplies a hostile comparison boundary for Candidate 002, Candidate 006, Candidate 007, Candidate 009, Candidate 012, and Candidate 014.

Episode, operator, and action identity

For receiver rr, source ss, microphone or ear channel mm, and episode ee, preserve

Ae=(Xe,Se,Ee,Re,He,Oe,Ce,Te,Ue,Be),\mathcal A_e=(X_e,S_e,E_e,R_e,H_e,O_e,C_e,T_e,U_e,B_e),

where:

  • XeX_e is geometry, medium, boundaries, temperature in kelvins, relative humidity as a dimensionless fraction, flow in metres per second, occupancy, and time-varying room/material state;
  • SeS_e is every target and interfering source identity, waveform, position in metres, velocity in metres per second, directivity, emission time in seconds, and source-level record;
  • EeE_e is commanded and realized active emission: waveform, spectrum, duration in seconds, acoustic energy in joules, aim in degrees, repetition, and dose;
  • ReR_e is receiver/body/array position and orientation, aperture in metres, morphology, feasible motion, transducer response, gain, health, and saturation;
  • HeH_e is prior acoustic exposure, room/source history, adaptation, training, previous emissions and motions, and feedback with timestamps;
  • OeO_e is the observation operator: impulse responses, transfer functions, spatial/time support, sample rate in samples per second, quantization in bits, clock, latency in seconds, preprocessing, missingness, and selection;
  • CeC_e is calibration identity, uncertainty, reference pressure and distance, instrument class, traceability, and data vintage;
  • TeT_e is the target construct, literal outcome, deadline in seconds, and loss or utility in declared units;
  • UeU_e is the independent unit: waveform, frame, event, source, room, receiver, body, array, device, site, or population; and
  • BeB_e is the complete ceiling in samples, events, bytes, seconds, person-hours, joules, emissions, exposure, unsafe events, sensors, actuators, replacements, and opportunity.

For method qq and literal outcome kk, the estimand is

Qq,k(A)=E ⁣[Ykdo(q),A],Q_{q,k}(\mathcal A)= \mathbb E\!\left[Y_k\mid do(q),\mathcal A\right],

where YkY_k uses the registered unit for outcome kk. A contrast against baseline bb does not isolate qq if aperture, calibration, room response, source distribution, observation receipt time, emission or motion authority, training scenes, or any binding resource differs without intervention.

Pressure, level, exposure, and clock contract

For acoustic pressure p(t)p(t) in pascals over interval TT in seconds,

prms=1T0Tp2(t)dt,Lp=20log10 ⁣(prmsp0),p_{\mathrm{rms}}= \sqrt{\frac{1}{T}\int_0^T p^2(t)\,dt}, \qquad L_p=20\log_{10}\!\left(\frac{p_{\mathrm{rms}}}{p_0}\right),

where prmsp_{\mathrm{rms}} is RMS pressure in pascals, p0=20μPap_0=20\,\mu\mathrm{Pa} is the reference pressure in air, and LpL_p is in decibels relative to 20μPa20\,\mu\mathrm{Pa}. Frequency weighting, time weighting, band, location, orientation, and calibration are fields, not implied defaults.

Sound exposure and its level are

Ep=0Tp2(t)dt,LE=10log10 ⁣(EpE0),E0=p02t0,E_p=\int_0^T p^2(t)\,dt, \qquad L_E=10\log_{10}\!\left(\frac{E_p}{E_0}\right), \qquad E_0=p_0^2t_0,

where EpE_p is in pascal-squared seconds, LEL_E is in decibels relative to E0E_0, and reference duration t0=1st_0=1\,\mathrm{s}. Exposure is not acoustic emission energy, electrical energy, or a universal risk model.

For channel clock cc, correct a device timestamp by

tˉc,n=tc,ndevδc,v,\bar t_{c,n}=t^{\mathrm{dev}}_{c,n}-\delta_{c,v},

where device time tc,ndevt^{\mathrm{dev}}_{c,n}, corrected time tˉc,n\bar t_{c,n}, and version-vv offset δc,v\delta_{c,v} are in seconds; cc is the channel-clock index, nn is the dimensionless sample index, and vv is the dimensionless calibration version. Retain residual clock uncertainty σc,v\sigma_{c,v} in seconds, drift in seconds per second, capture time, receipt time, synchronization method, and calibration interval. At decision time tt, only observations with receipt time no later than tt are available.

Propagation, rooms, and time-varying mixing

At channel mm, use the time-varying mixture

ym(t)=s=1Shm,s(t,τ)xs(tτ)dτ+nm(t),y_m(t)=\sum_{s=1}^{S} \int h_{m,s}(t,\tau)x_s(t-\tau)\,d\tau+n_m(t),

where received waveform ym(t)y_m(t), source pressure xs(t)x_s(t), and noise nm(t)n_m(t) are in pascals; SS is dimensionless source count; tt and delay τ\tau are in seconds; and transfer kernel hm,s(t,τ)h_{m,s}(t,\tau) is in reciprocal seconds. A time-invariant room convolution is a special case. Source/receiver motion, changing boundaries, temperature, flow, occupancy, and adaptive emitters make the operator time-dependent.

For an ideal free-field point source with unchanged directivity and negligible absorption,

ΔLp=20log10 ⁣(r1r2),\Delta L_p=20\log_{10}\!\left(\frac{r_1}{r_2}\right),

where distances r1,r2r_1,r_2 are in metres and ΔLp\Delta L_p is in decibels. The comparison must expose directionality, near-field terms, boundaries, atmospheric absorption, scattering, and receiver orientation when present.

For a diffuse-field Sabine approximation,

T600.161VA,T_{60}\approx0.161\frac{V}{A},

where T60T_{60} is decay time in seconds, room volume VV is in cubic metres, equivalent absorption area AA is in square metres, and 0.1610.161 has units seconds per metre. Report measured impulse responses and uncertainty by band, source, receiver, and support; this approximation is not an operator identity.

Masking, filterbank, compression, and events

A real gammatone-like channel is

gj(t)=ajtnj1e2πbjtcos(2πfjt+ϕj),t0,g_j(t)=a_jt^{n_j-1}e^{-2\pi b_jt} \cos(2\pi f_jt+\phi_j),\qquad t\ge0,

where centre frequency fjf_j and bandwidth bjb_j are in hertz, filter order njn_j is dimensionless, phase ϕj\phi_j is in radians, time tt is in seconds, and aja_j has units snj\mathrm{s}^{-n_j} so gjg_j is in reciprocal seconds. Fixed FFT, wavelet, mel, ERB, gammatone/gammachirp, modulation, and learned filterbanks remain competing representations.

A local normalized compression fit is

zz0=(xx0)γ,0<γ1,\frac{z}{z_0}=\left(\frac{x}{x_0}\right)^\gamma, \qquad 0<\gamma\le1,

where magnitudes x,x0x,x_0 share one input unit, z,z0z,z_0 share one output unit, and exponent γ\gamma is dimensionless. Register frequency, level, state, attack/release time in seconds, saturation, and distortion; one exponent does not describe an adaptive cochlea or compressor globally.

For yes/no target detection,

d=Φ1(Phit)Φ1(Pfalse alarm),d'=\Phi^{-1}(P_{\mathrm{hit}})- \Phi^{-1}(P_{\mathrm{false\ alarm}}),

where both PP terms are dimensionless probabilities, Φ1\Phi^{-1} is the inverse standard-normal cumulative distribution, and sensitivity dd' is dimensionless. Keep decision criterion, energetic overlap, informational uncertainty, spatial release, source grouping, and task history separate.

For event encoder threshold θj\theta_j in the unit of channel response uj(t)u_j(t),

ej,n=(j,tj,n,uj(tj,n),vθ,vc)whenuj(tj,n)uj(tj,n1)θj,e_{j,n}=\left(j,t_{j,n},u_j(t_{j,n}),v_{\theta},v_c\right) \quad\text{when}\quad |u_j(t_{j,n})-u_j(t_{j,n-1})|\ge\theta_j,

where jj is channel identity, event time tj,nt_{j,n} is in seconds, nn is the dimensionless event index, ej,ne_{j,n} is the retained event record, tj,n1t_{j,n-1} is the previous retained-event time in seconds, and vθ,vcv_{\theta},v_c are dimensionless threshold and clock versions. Event rate is in events per second. Report missed sustained signals, false events, timestamp error, decoder cost, bytes, task information, and measured joules; event sparsity is not an energy unit.

ITD, ILD, aperture, correlation, and beamforming

For sound speed csndc_{\mathrm{snd}} in metres per second and path-length difference Δr\Delta r in metres,

Δt=Δrcsnd,\Delta t=\frac{\Delta r}{c_{\mathrm{snd}}},

where interaural or interchannel time difference Δt\Delta t is in seconds. Clock offset, phase ambiguity, multipath, source extent, and head/array transfer functions are part of its uncertainty.

For left/right RMS pressures pL,pRp_L,p_R in pascals,

ILD=20log10 ⁣(pLpR),\mathrm{ILD}=20\log_{10}\!\left(\frac{p_L}{p_R}\right),

where interaural level difference is in decibels and is band-, time-, source-, and orientation-qualified. ITD, ILD, spectral cues, and head motion remain separate intervention axes.

For far-field array aperture DD in metres, frequency ff in hertz, and wavelength λ=csnd/f\lambda=c_{\mathrm{snd}}/f in metres,

ΔθλD,\Delta\theta\sim\frac{\lambda}{D},

where angular resolution Δθ\Delta\theta is in radians. Geometry, beampattern, SNR, estimator, bandwidth, and resolution definition determine the coefficient; no learned estimator restores spatial frequencies excluded by the physical support without additional prior information.

For signals y1(t),y2(t)y_1(t),y_2(t) in pascals, define generalized cross-correlation

R12(Ψ)(τ)=Ψ(f)Y1(f)Y2(f)ei2πfτdf,τ^=argmaxτTR12(Ψ)(τ),R_{12}^{(\Psi)}(\tau)= \int_{-\infty}^{\infty} \Psi(f)Y_1(f)Y_2^*(f)e^{i2\pi f\tau}\,df, \qquad \widehat{\tau}=\arg\max_{\tau\in\mathcal T}R_{12}^{(\Psi)}(\tau),

where Y1,Y2Y_1,Y_2 are Fourier transforms in pascal-seconds, frequency ff is in hertz, delay τ\tau and estimate τ^\widehat\tau are in seconds, T\mathcal T is the feasible delay set in seconds, superscript * is complex conjugation, and weighting Ψ(f)\Psi(f) has reciprocal pascal-squared-second-squared units so R12(Ψ)R_{12}^{(\Psi)} is in hertz after integration. The peak must carry a calibrated multimodal delay distribution under reverberation, not only an argmax.

For array snapshot yCM\mathbf y\in\mathbb C^M in pascals, dimensionless steering vector a\mathbf a, and noise covariance Rn=E[nnH]\mathbf R_n=\mathbb E[\mathbf n\mathbf n^H] in pascal-squared units,

wMVDR=Rn1aaHRn1a,z=wMVDRHy,\mathbf w_{\mathrm{MVDR}}= \frac{\mathbf R_n^{-1}\mathbf a} {\mathbf a^H\mathbf R_n^{-1}\mathbf a}, \qquad z=\mathbf w_{\mathrm{MVDR}}^H\mathbf y,

where MM is microphone count, superscript HH is conjugate transpose, wMVDR\mathbf w_{\mathrm{MVDR}} is dimensionless, and output zz is in pascals. Steering mismatch, covariance error, short sample support, correlated sources, motion, and room change receive sealed interventions.

Reverberation, separation, and source identity

For direct component dm(t)d_m(t), early-reflection component rmearly(t)r_m^{\mathrm{early}}(t), late component rmlate(t)r_m^{\mathrm{late}}(t), and noise nm(t)n_m(t), all in pascals,

ym(t)=dm(t)+rmearly(t)+rmlate(t)+nm(t).y_m(t)=d_m(t)+r_m^{\mathrm{early}}(t)+ r_m^{\mathrm{late}}(t)+n_m(t).

The early/late boundary is a registered time in seconds relative to the direct arrival. Waveform dereverberation, speech compensation, perceived distance, localization, and downstream intelligibility are distinct outcomes.

For estimated source s^\widehat{\mathbf s} and reference source sRN\mathbf s\in\mathbb R^N with the same amplitude unit and sample count NN,

starget=s^,ss22s,SI ⁣ ⁣ ⁣SDR=10log10starget22s^starget22,\mathbf s_{\mathrm{target}}= \frac{\langle\widehat{\mathbf s},\mathbf s\rangle} {\lVert\mathbf s\rVert_2^2}\mathbf s, \qquad \mathrm{SI\!\!-\!SDR}=10\log_{10} \frac{\lVert\mathbf s_{\mathrm{target}}\rVert_2^2} {\lVert\widehat{\mathbf s}-\mathbf s_{\mathrm{target}}\rVert_2^2},

where ,\langle\cdot,\cdot\rangle is the waveform inner product and SI-SDR is in decibels. Report unknown source count, permutation, causal identity, localization, intelligibility, calibration, perceptual/task quality, and residual mixture separately.

For predicted scene outcome znz_n and causally available acoustic observation ono_n, proper log loss is

Lq=1Nen=1Nelog2pq(znon),L_q=-\frac{1}{N_e}\sum_{n=1}^{N_e} \log_2p_q(z_n\mid o_n),

where NeN_e is independent event count and LqL_q is in bits per event. Use it for detection, source count, assignment, localization bins, range bins, and abstention under held-out operators; do not pool incompatible outcomes.

Active emission, dose, and closed-loop action

For monostatic emission at temitt_{\mathrm{emit}} and echo receipt at trecvt_{\mathrm{recv}}, both in seconds,

r^=csnd2(trecvtemit),\widehat r=\frac{c_{\mathrm{snd}}}{2} \left(t_{\mathrm{recv}}-t_{\mathrm{emit}}\right),

where estimated range r^\widehat r is in metres and sound speed csndc_{\mathrm{snd}} is in metres per second. Clock uncertainty, target motion, refraction, multipath, transducer ringing, waveform ambiguity, association, and detector threshold must propagate to range uncertainty.

For acoustic output power Pac(t)P_{\mathrm{ac}}(t) in watts over emission duration TeT_e in seconds,

Eemit=0TePac(t)dt,E_{\mathrm{emit}}=\int_0^{T_e}P_{\mathrm{ac}}(t)\,dt,

where EemitE_{\mathrm{emit}} is acoustic joules. Electrical input, transduction loss, body/sensor motion, sensing, compute, cooling, detectability, interference, and exposure remain separate axes.

At decision time tt, a complete acoustic policy is

(atemit,atbody,atsensor,atgain,attask)=πq ⁣(Ht,X^t,O^t,U^t,Atsafe,Bt),(a_t^{\mathrm{emit}},a_t^{\mathrm{body}},a_t^{\mathrm{sensor}}, a_t^{\mathrm{gain}},a_t^{\mathrm{task}})= \pi_q\!\left(\mathcal H_t,\widehat X_t,\widehat O_t, \widehat U_t,\mathcal A_t^{\mathrm{safe}},B_t\right),

where atemita_t^{\mathrm{emit}} contains waveform, level, spectrum, aim, and timing; atbodya_t^{\mathrm{body}} contains pose or head/body motion in metres, radians, and seconds; atsensora_t^{\mathrm{sensor}} contains aperture, sample rate, and channel selection; atgaina_t^{\mathrm{gain}} is dimensionless or in declared decibels; attaska_t^{\mathrm{task}} is the downstream command in its native unit; Ht\mathcal H_t is causally received history; X^t\widehat X_t is scene state; O^t\widehat O_t is operator/calibration state; U^t\widehat U_t is uncertainty; Atsafe\mathcal A_t^{\mathrm{safe}} is the feasible action envelope; and BtB_t is remaining budget in matched units. Attention- or efferent-like gain receives causal credit only when intervening on it changes its registered endpoint.

Hardware, lifecycle energy, and equal budgets

Lifecycle energy is

Eqlife=Eqdata+Eqtrain+Eqemit+Eqsense+Eqinfer+Eqmove+Eqcomm+Eqstore+Eqcal+Eqmaint+Eqemb,E_q^{\mathrm{life}}= E_q^{\mathrm{data}}+E_q^{\mathrm{train}}+E_q^{\mathrm{emit}}+ E_q^{\mathrm{sense}}+E_q^{\mathrm{infer}}+E_q^{\mathrm{move}}+ E_q^{\mathrm{comm}}+E_q^{\mathrm{store}}+E_q^{\mathrm{cal}}+ E_q^{\mathrm{maint}}+E_q^{\mathrm{emb}},

where every term is in joules over one service interval and denotes data acquisition, training, acoustic emission, sensing, inference, physical motion, communication, storage, calibration, maintenance, and amortized embodied energy. DSP, FPGA, ASIC, neuromorphic, CPU, GPU, and analog front ends are measured at identical physical boundaries and workload deadlines.

Human effort is

Hqhuman=Hqdesign+Hqrecord+Hqlabel+Hqlisten+Hqcal+Hqtune+Hqmonitor+Hqrepair,H_q^{\mathrm{human}}= H_q^{\mathrm{design}}+H_q^{\mathrm{record}}+H_q^{\mathrm{label}}+ H_q^{\mathrm{listen}}+H_q^{\mathrm{cal}}+H_q^{\mathrm{tune}}+ H_q^{\mathrm{monitor}}+H_q^{\mathrm{repair}},

where every HH term is in person-hours and roles are reported separately.

The complete cost vector is

Cq=(Nsample,Nevent,Nstep,Nquery,Nbyte,Twall,Hqhuman,Eqlife,Ep,Nunsafe,Dharm,Copp),\mathbf C_q=(N_{\mathrm{sample}},N_{\mathrm{event}},N_{\mathrm{step}}, N_{\mathrm{query}},N_{\mathrm{byte}},T_{\mathrm{wall}}, H_q^{\mathrm{human}},E_q^{\mathrm{life}},E_p, N_{\mathrm{unsafe}},D_{\mathrm{harm}},C_{\mathrm{opp}}),

where the five NN terms count samples, emitted/received events, optimization or environment steps, queries, and bytes; TwallT_{\mathrm{wall}} is seconds; HqhumanH_q^{\mathrm{human}} is person-hours; EqlifeE_q^{\mathrm{life}} is joules; EpE_p is pascal-squared-second exposure; NunsafeN_{\mathrm{unsafe}} counts unsafe events; DharmD_{\mathrm{harm}} uses a registered physical or severity unit; and CoppC_{\mathrm{opp}} is opportunity cost in task exposures or person-hours.

Method qq is feasible only when

CqB,\mathbf C_q\preceq\mathbf B,

where B\mathbf B is the preregistered componentwise ceiling with identical units. Over-budget runs are infeasible; removed ablation resources stay unused.

Confirmatory contrast and retirement

Let b(j)b^*(j) be the strongest mature baseline for track jj, frozen on development data. Orient every protected endpoint as a preregistered benefit, so larger QQ is better and a non-inferiority margin may be negative. For protected outcomes kPjk\in\mathcal P_j, retain a residual only when

Pr ⁣(Qq,kQb(j),k>δj,k for every kPj)1αj,\Pr\!\left( Q_{q,k}-Q_{b^*(j),k}>\delta_{j,k} \text{ for every }k\in\mathcal P_j \right)\ge1-\alpha_j,

where δj,k\delta_{j,k} is an improvement or non-inferiority margin in the unit of outcome kk, and αj\alpha_j is the dimensionless error budget. The result must survive held-out rooms, sources, arrays, bodies, operators, scenes, source counts, interference policies, model families, sites, and hardware classes at equal complete cost. Otherwise retire the architectural residual while keeping the operator/action/measurement contract.