The wideners are needed and too large, and they are multiplied when they should not be

Rendered from gate/RESULT-WIDENERS.md

Run: gate/widener_coverage.pygate/widener_coverage.json Date: 2026-09-22.

What I did

gate/RESULT-RECOVERY.md issued NS-0027 with an interval of [0.00, 199.42] — wider than its own value — and recorded a concern instead of acting on it: §5's base is the observed between-unit spread, which already contains whatever the units differ by, and INSTRUCTION.md Domain 5 then multiplies it by species and instrumentation wideners, which are allowances for carrying a value outward. The concern was that this double-counts. That is an empirical question, and an interval is judged empirically by one thing — coverage. A ± 1 SD interval should contain about 68.3% of the values it is meant to cover.

I measured coverage three ways on the register's own values: within a cohort, out of cohort, and for NS-0027 specifically. Every figure is leave-one-out — the interval is built from the other units and then asked whether it contains the held-out one. An in-sample interval is judged by a width the held-out value helped set, and control 2 measures how much that matters rather than assuming it: at n = 6, naive coverage runs 0.681 against a leave-one-out 0.583, so the correction is worth ten points.

The estimator is right, and two mis-specifications were caught on the way

control result
Leave-one-out coverage on Gaussian samples, where 0.683 is the answer 0.670
Naive exceeds leave-one-out at n = 6, so the correction does something 0.681 > 0.583

Two errors in this run's own design were caught and are recorded rather than quietly repaired.

The first version tested the wrong proposition. It built the interval from the pooled across-cohort spread and asked whether it covered its own members — but that set already contains the between-cohort variation a widener exists to supply. It would have reported the wideners unnecessary for a reason that had nothing to do with wideners. The pooled figure is retained in the output, labeled as what it is.

The second was a category error this standard exists to prevent. The target filter selected on state alone and pulled NS-0015 and NS-0016 — separation-scale ratios of 1.21 and 1.89 — into a comparison against a PCIst center of 63.97, reporting that they needed a ×3.6 widener. METHOD §3 forbids mixing the three scales in exactly these terms, and a run auditing the standard committed the error the standard was written to stop. Filtered on scale now.

Out of cohort, the wideners are needed — the base covers a third of targets

The widener's actual job: an interval built from one cohort's internal spread must cover a value from another. Source is the mouse awake cohort, six subjects, center 63.97, SD 17.48.

target value covered at base multiplier needed
NS-0021 human TMS-EEG 59.00 yes 0.28
NS-0023 human TMS benchmark 47.89 yes 0.92
NS-0019 rat intracranial 42.35 no 1.24
NS-0017 human intracranial 38.00 no 1.49
NS-0026 human grid, second cohort 35.83 no 1.61
NS-0012 human grid 32.26 no 1.81

Base alone covers 2 of 6, 33% against a nominal 68%. The wideners exist for a real gap and the direction is right. That half of the concern is answered against me.

But the prescribed multiplier over-covers, because two wideners are multiplied that are not independent

×1.61 reaches nominal coverage. INSTRUCTION 5 prescribes ×2.25 — species 1.50 times instrumentation 1.50 — and that covers 100%.

An interval that covers everything is not conservative; it is uninformative, and it is uninformative in the specific way that makes a scale impossible, because every reading's interval swallows every other reading's value. Note what ×1.61 sits next to: a single widener of 1.50 nearly achieves nominal on its own, while multiplying two gives total coverage.

The reason is that the two wideners are not independent events. Crossing species necessarily crosses instrumentation — there is no mouse recording taken on a human clinical grid — so the product counts one gap twice. §5 says "where more than one applies, the factors multiply," which is correct for wideners describing genuinely separate sources of error and wrong for wideners describing the same one from two angles.

Within cohort the same over-shoot appears. NS-0027's prescribed ×2.8125 gives 100% coverage where the recovery data needs ×1.24; the awake cohort's base already over-covers at 0.83 and needs ×0.98, meaning slightly less than one standard deviation.

set n CV coverage at base multiplier for nominal
mouse awake 6 0.27 0.83 0.98
mouse isoflurane 6 0.54 0.50 1.17
mouse recovery 4 0.80 0.25 1.24

What this does not establish

Six targets from one source cohort, and most of them human, so this measures a mouse-against-other gap and not a general one. At n = 6 the 68% quantile is the fifth of six sorted values, which is a coarse instrument for choosing a factor. The measured figures are therefore adequate to show that ×2.25 over-covers and inadequate to fix a replacement factor to two decimal places — so the amendment below changes the rule for combining wideners, which the data does support, and leaves the individual factors where they are.

What I recommend next

Adopt the combination rule the measurement supports: non-independent wideners combine as the maximum, not the product. Species extrapolation and instrumentation mismatch are the clear case, and the amendment names them as a dependent pair rather than inventing a general independence test the data cannot support. Under that rule NS-0027's multiplier falls from 2.8125 to 1.875 and the out-of-cohort multiplier from 2.25 to 1.50, which sits just under the 1.61 measured as nominal — conservative by a little rather than by 40%.

NS-0027 is not edited or recomputed. It applied the rule correctly as the rule stood, §8 keeps readings at the version they were issued under, and a superseding reading is the mechanism if one is wanted later. The gap between its interval and the amended rule stays visible in the register, which is the same treatment every other superseded practice has had.

Then run this again with a human source cohort once per-subject PRIOS values are computed, because one source cohort measuring outward is half a calibration.