Instruction
INSTRUCTION.mdContents
Version 1.8.0 · in force from 2026-09-21 · normative
companion to METHOD.md v0.4.6
Question or issue resolved
METHOD.md states what the Nooscope measures and what may
be said about the result. It does not state how a person executing it
resolves the judgments it leaves open, and there are more of those than
the document's tone suggests. Which channels are good. Which trials are
rejected. Whether a subject's implant is one geometry or two. Whether a
parameter differs enough from a documented set to count as a departure.
How wide the interval is, given that §5 says it must widen without
saying by how much. Which caveats a reading carries.
Every one of those is currently settled by whoever is at the bench. That is the failure mode this document exists to close, and it is a different failure mode from the one the register already guards against.
The reproduction gate (METHOD.md §9.1) asks whether this
pipeline can reproduce somebody else's published measurement. It is a
good gate and it is cheap, and it will find a broken pipeline in month
four rather than year three. But it is a test of the pipeline against
the world. It cannot detect the other failure: two competent
people running this method on the same deposit getting different
readings. Nothing in the register currently detects that, and
it is the failure that kills a measurement standard, because a standard
that does not produce the same answer twice is not a standard. It is a
house style.
The evidence that specification is the remedy is direct. In the published test of this approach, a written, context-specific implementation instruction moved interrater agreement from none to moderate and cut assessment time by roughly three quarters. Published agreement for unaided expert judgment on comparable rubrics runs from κ = 0.06 to κ = 0.44 — that is, from nothing to modest. This document is the Nooscope's instruction.
The rule that governs every threshold below. Where a
threshold is stated, the answer is determined by the threshold and the
rater has no discretion to override it. Disagreement with a threshold is
raised as a proposed amendment to this document under
GOVERNANCE.md, not resolved at the bench. A rater who
believes a threshold is wrong is very possibly right, and the way to be
right is to change the number for everybody and re-rate what it affects,
not to depart from it once.
Answer options throughout: Yes / Probably yes / Probably no / No / No information.
Domain 1 — Source admissibility
Whether this deposit can carry this reading at all. Assessed before anything is computed.
| # | Question | Threshold |
|---|---|---|
| 1.1 | Does the task named in source carry stimulation
events? |
≥ 1 stimulation event in the events file for that task. Zero = No, and the reading is not issuable as Type A or B. A deposit that advertises stimulation elsewhere does not satisfy this. |
| 1.2 | Is per-trial stimulation intensity recorded with units? | Present with units on ≥ 95% of retained trials = Yes. 80–94% = Probably yes, and a caveat is required. < 80% = No. |
| 1.3 | Are stimulation waveform and repetition rate recorded? | Both stated = Yes. Either absent = No, and METHOD.md
§2.1 requires the reading to carry the omission as a caveat. This is not
a formality: protocol moves the measure by a median factor of 2.07
(gate/RESULT-CHOCS.md), which is larger than wakefulness
against propofol. |
| 1.4 | Is the license an open license permitting redistribution of derived values? | A named license permitting it = Yes. "Available on request", "contact the authors", or silence = No. |
| 1.5 | Is the deposit retrievable at a stable identifier? | DOI or repository accession = Yes. A bare URL = Probably no. |
Domain 1 is not admissible if 1.1 or 1.4 is No. Both are hard gates; nothing later in this document can rescue a reading that fails either.
Domain 2 — Channel and trial admission
The largest source of silent divergence between two raters, because both the channel set and the rejection rule can be varied without anything visibly breaking.
| # | Question | Threshold |
|---|---|---|
| 2.1 | Is the good-channel set constant across every run combined into this reading? | The intersection of good channels across runs must be ≥
20 (MIN_CHANNELS),
provisional. Below 20 the reading is not issuable and
the pipeline raises rather than returning a value. The factor is marked
provisional in 1.8.0 because it carried no stated basis — the only
unmarked threshold in this instruction — while excluding every substrate
whose arrays are built below clinical scale. On the register's own data
sixteen spatially distributed channels reproduce the full-array value at
a median 1.014, interquartile 0.97–1.08, against 0.993 and 0.97–1.01 at
twenty; what degrades at low counts is distribution rather than count,
randomly chosen channels falling to 0.708 at eight
(gate/RESULT-ORGANOID.md). The threshold is not
moved here. That measurement is on mammalian cortex, §9.4.1
forbids carrying it to another substrate, and it was run by a party that
wanted the threshold lowered. A replacement pairing a count with a
spatial-distribution condition is proposed for the next cycle and is to
be decided against a non-cortical substrate. |
| 2.2 | Is the rejection rule the one recorded in the rejection
field, applied identically to every run? |
Any per-run variation = No, and the reading is not issuable. |
| 2.3 | What fraction of trials was rejected? | ≤ 25% = Yes. 26–50% = Probably yes, caveat required. > 50% = No, and the reading is not issuable: at that point the rejection rule is selecting the result. |
| 2.4 | Were any channels or trials excluded by a decision not expressible as a rule? | Any such exclusion = No. Hand-picking is the one rejection procedure this standard does not permit, because it cannot be audited or reproduced. |
| 2.5 | Does the rejection record state counts as well as the rule? | Both present = Yes. A rule without counts = No. A record stating "25 cells sampled" while the channel filter matched nothing is the defect this question exists to catch. |
Domain 3 — Recording geometry
METHOD.md §3.3. Assigned from the channel-type census,
never by the rater's impression of what the implant was.
| # | Question | Threshold |
|---|---|---|
| 3.1 | What geometry does the census return? | The channel type holding a plurality of the constant good-channel
set: ECOG → ecog_grid, SEEG →
seeg_depth, EEG → scalp_eeg,
planar microelectrode array → mea, penetrating silicon
probe with contacts along a shank → probe_linear. An
EEG census maps instead to epidural_array
where the deposit's device record names a surface array implanted under
the scalp; where that record is silent on whether the array was
implanted, the reading is not issuable rather than
defaulted to scalp_eeg, because the default would place an
implant and a non-invasive montage in one geometry. Assigned
mechanically from the census, not from the deposit's prose
description. |
| 3.5 | Does the census return a channel type outside that set? | Any type with no mapping = Yes, and the reading is not
issuable. The geometry set is closed (METHOD.md
§3.3) and is extended by amendment, never by naming an unfamiliar array
at the bench. |
| 3.2 | Does the subject carry a second geometry? | ≥ 8 channels of a second geometry type
(MIN_OTHER_GEOMETRY) = Yes, the subject is mixed. |
| 3.3 | Where 3.2 is Yes, may the reading proceed on the majority geometry alone? | No. ALLOW_GEOMETRY_SUBSET is false and
changing it is a MINOR version change to METHOD.md. A
mixed-implant subject is excluded from a single-geometry reading; it is
not silently reduced to its larger array. |
| 3.4 | Do all subjects combined into one reading share a geometry? | Any mismatch = No, and they are separate readings. Eight of the nine
readable deposits in the register are geometrically heterogeneous
(gate/RESULT-GEOMETRY.md), so this question will usually
bite. |
| 3.6 | Is the nominal contact pitch established from the source? | The deposit's electrode description or the manufacturer's
specification for the named array. Either = Yes. Neither = No, and the
reading is not issuable (METHOD.md §3.5)
unless the deposit states coordinates, in which case the measured median
nearest-neighbour distance over adjacent admitted contacts of the
majority group stands in and grain.basis records that it
did. A pitch is never estimated from the array's physical dimensions or
from a comparable implant. |
| 3.7 | Where the deposit states coordinates, does the measured pitch agree with the nominal? | Measured within 15% of nominal = Yes. Outside = No,
and the reading is not issuable until the discrepancy
is named: the coordinates are not in the space the deposit implies. The
factor is measured, not provisional — across the four
deposits in the register's human supply that declare coordinate units,
clinical grid and strip groups measure 9.00 to 10.05 mm against a
nominal 10 mm, inside 10% over 14 groups
(gate/RESULT-PITCH.md). |
| 3.8 | Do all subjects combined into one reading fall within one grain band? | Ratio of largest to smallest grain.pitch_mm ≤
1.3 = Yes. Above 1.3 = No, and they are separate readings.
Compared on the nominal figure, which is exact, and
never on the measured one, which carries estimation error — question 3.7
guards the measured value. The factor is measured: at a
1.26× pitch ratio the median reading moves 8%, below the register's
smallest published state contrast of 1.21×
(gate/RESULT-DECIMATION.md). |
| 3.9 | Is this value being compared as an absolute against a reading whose nominal pitch ratio exceeds 1.3? | Ratio above 1.3 in either direction = Yes, and the §7 cross-grain
caveat is required in the same visual field as the value, stating the
measured sensitivity rather than a general worry. A reading issued
before 0.3.9, which carries no grain, counts as outside the
band unless its source deposit states a nominal figure within it. |
Domain 4 — Parameter set and departure
| # | Question | Threshold |
|---|---|---|
| 4.1 | Does a documented parameter set exist for this geometry and species? | A named set in the reference implementation = Yes. None = No, and the reading is a departure in full and says so. |
| 4.2 | Does any parameter differ from that set? | Any difference, of any size, is a departure and
must be declared with a reason in
computation.departure_reason. There is no threshold of
triviality. |
| 4.3 | Is the response window free of stimulation artifact? | The window must exclude the amplifier saturation. For direct cortical stimulation the saturation runs to roughly 15 ms, and the declared departure window is [15, 300] ms. A reading computed on [0, 300] ms for direct cortical stimulation carries the saturation, runs about 16 points high, and is not issuable as a value. |
| 4.4 | Is the memory-bounding resample parameter set? | PCIst is O(T²) in the response window. resample set =
Yes. Unset on a long window = No; the run will exhaust memory rather
than return a wrong number, but the reading is not issuable until it is
set deliberately. |
| 4.5 | Is the scale declared, and is it the only scale in the comparison? | PCI-LZ (0–1), PCIst (unbounded, roughly 5–60) and sPCI (0–1) are three incompatible scales. Declared and unmixed = Yes. Any comparison spanning two = No, and it is a category error, not a caveat. |
Domain 5 — Interval width
METHOD.md §5 requires the interval to widen and names
six conditions. It does not say by how much, which leaves the single
most consequential number in the record to the rater's judgment. This
domain fixes it.
The base interval is ± 1 standard deviation of the between-unit spread actually observed — between cells, sessions or subjects as the reading's pairing unit dictates. It is never the standard error of the mean. The standard error describes how well the mean is determined; the wideners of §5 are systematic and are not reduced by measuring more cells, so an interval built on the standard error narrows as the evidence accumulates in exactly the case where it should not.
Each condition below scales the base. Where more than one
applies, independent wideners multiply and a dependent pair combines as
the larger of the two — added in 1.7.0 and measured
(gate/RESULT-WIDENERS.md).
Species extrapolation and instrumentation mismatch are declared a dependent pair: crossing species necessarily crosses instrumentation, because no mouse recording is taken on a human clinical grid, so multiplying them counts one gap twice. Measured on the register's own values, an interval built from one cohort and tested against six out-of-cohort values covers 33% at base — the wideners are needed — reaches the nominal 68% at ×1.61, and covers 100% at the ×2.25 the product prescribed. An interval that covers every other reading's value cannot support a scale.
No general independence test is offered, because six targets from one source cohort cannot support one. Any further pair is declared dependent by amendment, with its measurement.
| Condition | Factor | Basis |
|---|---|---|
| Small n — fewer than 5 subjects | × 1.25 | provisional |
| Instrumentation or cohort mismatch | × 1.50 | measured. Between-cohort variation with all known variables held constant is about 1.5× (NS-0012 against NS-0026) |
| Species extrapolation | × 1.50 | provisional |
| Preparation extrapolation — in vitro | × 2.00 | provisional |
| Proxy substitution — Type C, no perturbation delivered | × 2.00 | provisional |
| Parameter departure | × 1.25 | provisional |
A factor marked provisional is a placeholder with no measurement behind it. It is recorded as provisional on the face of the reading, and every provisional factor is listed for review each cycle. This is the honest state of the interval: the wideners are real, their sizes are mostly not yet measured, and saying so is better than either omitting them or implying a precision that does not exist. Replacing a provisional factor with a measured one is a MINOR version change and triggers re-derivation of the affected intervals.
| # | Question | Threshold |
|---|---|---|
| 5.1 | Is the base the between-unit standard deviation rather than the standard error? | SD = Yes. SEM = No, and the interval is wrong. |
| 5.2 | Is every applied widener named in interval.basis? |
All named = Yes. A widener applied but unnamed = No. |
| 5.3 | Does the interval bracket the value? | Yes required. A value outside its own interval is a defect, not a finding. |
| 5.4 | Does the interval span the empirical cutoff? | Where it does, the reading's own summary line must state that it does not resolve the question it was taken to answer. Being unable to tell is a result and is published like any other. |
Domain 6 — Type and modality tier
Both are assigned by decision tree. Neither admits discretion.
| # | Question | Threshold |
|---|---|---|
| 6.1 | Was a perturbation of stated modality and intensity delivered? | A stimulation record naming modality
and intensity with units = Yes. Either absent = No →
the reading is Type C or E and can never share the Type A scale. |
| 6.2 | Did the issuer compute the value from the underlying data? | A computation record with library, version and commit =
Yes. Absent = No → Type D, sharing the Type A scale only where the third
party's method and parameters match. |
| 6.3 | Is the value a ratio between two issued readings of the same subjects? | A derived_from naming two issued Type A or B readings
from one deposit = Yes → Type F, dimensionless, never on the Type A
scale. |
| 6.4 | Is the data the issuer's own? | source.access is the issuer's own deposit = Yes → Type
A. A third-party deposit under a named open license = No → Type B. |
| 6.5 | What is the measurement modality tier? | Assigned per METHOD.md §3.4 and displayed on the face
of the reading. The tier ceiling is applied without exception. |
The tier is not a restatement of the type. The type records where the data came from; the tier records what kind of evidentiary object the reading is. A Type D published value can be M1 or M4 depending on what the third party actually measured, and the difference matters more than the provenance does.
Domain 7 — Caveats
Free text is unbounded discretion, and the caveat field is where an inconvenient fact goes to be worded into harmlessness.
| # | Question | Threshold |
|---|---|---|
| 7.1 | Does the reading carry a caveat for every triggered condition? | Each of the following triggers a required caveat: any Probably yes or Probably no answer above; every applied interval widener; every provisional widener factor; any declared parameter departure; small n; species or preparation extrapolation; missing waveform or repetition rate. |
| 7.2 | Is each required caveat written from the stem for its condition? | The stems are fixed in schema/caveat-stems.json. Free
text may be added; it may never
replace a required caveat. |
| 7.3 | Does the caveat field name what a critic should attack first? | Required, and it must name one of: a specific widener, a specific parameter departure, a specific rejected channel or trial class, or a specific source limitation. A caveat list naming none of those = No. |
Fixing the stems makes caveat presence fully reproducible between two raters even though caveat wording cannot be. Presence is the part that matters.
Domain 8 — Rating, disagreement and adjudication
| # | Question | Threshold |
|---|---|---|
| 8.1 | Was the reading produced independently by two raters who did not confer? | Both ratings recorded in full, at domain level, before adjudication. |
| 8.2 | Are the pre-adjudication ratings retained? | Permanently. Most bodies adjudicate and discard the
disagreement. Retaining it is the only thing that makes the reliability
statistics in METHOD.md §9.3 verifiable rather than
asserted. |
| 8.3 | Is the adjudicator named, with reasoning recorded? | Named third party, decision and reasoning in the record. |
| 8.4 | Is assessment time recorded? | assessment_minutes present for both raters = Yes;
absent for either = No. Reported at cycle level. |
Conclusion
This instruction is the Nooscope's answer to a failure the register could not previously see. The reproduction gate tests the pipeline against the published world and will catch a pipeline that is broken. It is silent on whether the method is specified tightly enough that two people executing it arrive at the same place, and on the evidence from comparable rubrics the default answer to that is no — unaided expert agreement runs from κ = 0.06 to κ = 0.44, and the one published cross-organization comparison in the adjacent field found near-zero agreement.
Three things follow from adopting it.
It comes before the kill test, not after. The gate is scheduled for month four. This document should be in force first, because the gate will be executed by a person making every judgment above, and a gate whose result depends on who ran it does not discharge its purpose. The cost of writing it first is days; the cost of discovering afterwards that the gate was not reproducible is the gate.
The thresholds here are claims, and most of them are provisional. Six of the seven interval widening factors have no measurement behind them and are marked as such. Several domain thresholds are inherited from an adjacent field rather than derived here. That is the correct state for version 1.0 and the wrong state for version 2.0, and the mechanism for moving between them is the amendment process, not accumulated bench practice.
Low agreement is a defect in this document, not in the raters. Where a domain's agreement falls below κ = 0.40, the response is an amendment to the signalling questions and thresholds above. An instruction to raters to be more careful is not a response, because low agreement is evidence that a judgment is under-specified, and the remedy for under-specification is specification.