Instruction

Rendered from INSTRUCTION.md
Contents
  1. Question or issue resolved
  2. Domain 1 — Source admissibility
  3. Domain 2 — Channel and trial admission
  4. Domain 3 — Recording geometry
  5. Domain 4 — Parameter set and departure
  6. Domain 5 — Interval width
  7. Domain 6 — Type and modality tier
  8. Domain 7 — Caveats
  9. Domain 8 — Rating, disagreement and adjudication
  10. Conclusion

Version 1.8.0 · in force from 2026-09-21 · normative companion to METHOD.md v0.4.6


Question or issue resolved

METHOD.md states what the Nooscope measures and what may be said about the result. It does not state how a person executing it resolves the judgments it leaves open, and there are more of those than the document's tone suggests. Which channels are good. Which trials are rejected. Whether a subject's implant is one geometry or two. Whether a parameter differs enough from a documented set to count as a departure. How wide the interval is, given that §5 says it must widen without saying by how much. Which caveats a reading carries.

Every one of those is currently settled by whoever is at the bench. That is the failure mode this document exists to close, and it is a different failure mode from the one the register already guards against.

The reproduction gate (METHOD.md §9.1) asks whether this pipeline can reproduce somebody else's published measurement. It is a good gate and it is cheap, and it will find a broken pipeline in month four rather than year three. But it is a test of the pipeline against the world. It cannot detect the other failure: two competent people running this method on the same deposit getting different readings. Nothing in the register currently detects that, and it is the failure that kills a measurement standard, because a standard that does not produce the same answer twice is not a standard. It is a house style.

The evidence that specification is the remedy is direct. In the published test of this approach, a written, context-specific implementation instruction moved interrater agreement from none to moderate and cut assessment time by roughly three quarters. Published agreement for unaided expert judgment on comparable rubrics runs from κ = 0.06 to κ = 0.44 — that is, from nothing to modest. This document is the Nooscope's instruction.

The rule that governs every threshold below. Where a threshold is stated, the answer is determined by the threshold and the rater has no discretion to override it. Disagreement with a threshold is raised as a proposed amendment to this document under GOVERNANCE.md, not resolved at the bench. A rater who believes a threshold is wrong is very possibly right, and the way to be right is to change the number for everybody and re-rate what it affects, not to depart from it once.

Answer options throughout: Yes / Probably yes / Probably no / No / No information.


Domain 1 — Source admissibility

Whether this deposit can carry this reading at all. Assessed before anything is computed.

# Question Threshold
1.1 Does the task named in source carry stimulation events? ≥ 1 stimulation event in the events file for that task. Zero = No, and the reading is not issuable as Type A or B. A deposit that advertises stimulation elsewhere does not satisfy this.
1.2 Is per-trial stimulation intensity recorded with units? Present with units on ≥ 95% of retained trials = Yes. 80–94% = Probably yes, and a caveat is required. < 80% = No.
1.3 Are stimulation waveform and repetition rate recorded? Both stated = Yes. Either absent = No, and METHOD.md §2.1 requires the reading to carry the omission as a caveat. This is not a formality: protocol moves the measure by a median factor of 2.07 (gate/RESULT-CHOCS.md), which is larger than wakefulness against propofol.
1.4 Is the license an open license permitting redistribution of derived values? A named license permitting it = Yes. "Available on request", "contact the authors", or silence = No.
1.5 Is the deposit retrievable at a stable identifier? DOI or repository accession = Yes. A bare URL = Probably no.

Domain 1 is not admissible if 1.1 or 1.4 is No. Both are hard gates; nothing later in this document can rescue a reading that fails either.


Domain 2 — Channel and trial admission

The largest source of silent divergence between two raters, because both the channel set and the rejection rule can be varied without anything visibly breaking.

# Question Threshold
2.1 Is the good-channel set constant across every run combined into this reading? The intersection of good channels across runs must be ≥ 20 (MIN_CHANNELS), provisional. Below 20 the reading is not issuable and the pipeline raises rather than returning a value. The factor is marked provisional in 1.8.0 because it carried no stated basis — the only unmarked threshold in this instruction — while excluding every substrate whose arrays are built below clinical scale. On the register's own data sixteen spatially distributed channels reproduce the full-array value at a median 1.014, interquartile 0.97–1.08, against 0.993 and 0.97–1.01 at twenty; what degrades at low counts is distribution rather than count, randomly chosen channels falling to 0.708 at eight (gate/RESULT-ORGANOID.md). The threshold is not moved here. That measurement is on mammalian cortex, §9.4.1 forbids carrying it to another substrate, and it was run by a party that wanted the threshold lowered. A replacement pairing a count with a spatial-distribution condition is proposed for the next cycle and is to be decided against a non-cortical substrate.
2.2 Is the rejection rule the one recorded in the rejection field, applied identically to every run? Any per-run variation = No, and the reading is not issuable.
2.3 What fraction of trials was rejected? ≤ 25% = Yes. 26–50% = Probably yes, caveat required. > 50% = No, and the reading is not issuable: at that point the rejection rule is selecting the result.
2.4 Were any channels or trials excluded by a decision not expressible as a rule? Any such exclusion = No. Hand-picking is the one rejection procedure this standard does not permit, because it cannot be audited or reproduced.
2.5 Does the rejection record state counts as well as the rule? Both present = Yes. A rule without counts = No. A record stating "25 cells sampled" while the channel filter matched nothing is the defect this question exists to catch.

Domain 3 — Recording geometry

METHOD.md §3.3. Assigned from the channel-type census, never by the rater's impression of what the implant was.

# Question Threshold
3.1 What geometry does the census return? The channel type holding a plurality of the constant good-channel set: ECOGecog_grid, SEEGseeg_depth, EEGscalp_eeg, planar microelectrode array → mea, penetrating silicon probe with contacts along a shank → probe_linear. An EEG census maps instead to epidural_array where the deposit's device record names a surface array implanted under the scalp; where that record is silent on whether the array was implanted, the reading is not issuable rather than defaulted to scalp_eeg, because the default would place an implant and a non-invasive montage in one geometry. Assigned mechanically from the census, not from the deposit's prose description.
3.5 Does the census return a channel type outside that set? Any type with no mapping = Yes, and the reading is not issuable. The geometry set is closed (METHOD.md §3.3) and is extended by amendment, never by naming an unfamiliar array at the bench.
3.2 Does the subject carry a second geometry? 8 channels of a second geometry type (MIN_OTHER_GEOMETRY) = Yes, the subject is mixed.
3.3 Where 3.2 is Yes, may the reading proceed on the majority geometry alone? No. ALLOW_GEOMETRY_SUBSET is false and changing it is a MINOR version change to METHOD.md. A mixed-implant subject is excluded from a single-geometry reading; it is not silently reduced to its larger array.
3.4 Do all subjects combined into one reading share a geometry? Any mismatch = No, and they are separate readings. Eight of the nine readable deposits in the register are geometrically heterogeneous (gate/RESULT-GEOMETRY.md), so this question will usually bite.
3.6 Is the nominal contact pitch established from the source? The deposit's electrode description or the manufacturer's specification for the named array. Either = Yes. Neither = No, and the reading is not issuable (METHOD.md §3.5) unless the deposit states coordinates, in which case the measured median nearest-neighbour distance over adjacent admitted contacts of the majority group stands in and grain.basis records that it did. A pitch is never estimated from the array's physical dimensions or from a comparable implant.
3.7 Where the deposit states coordinates, does the measured pitch agree with the nominal? Measured within 15% of nominal = Yes. Outside = No, and the reading is not issuable until the discrepancy is named: the coordinates are not in the space the deposit implies. The factor is measured, not provisional — across the four deposits in the register's human supply that declare coordinate units, clinical grid and strip groups measure 9.00 to 10.05 mm against a nominal 10 mm, inside 10% over 14 groups (gate/RESULT-PITCH.md).
3.8 Do all subjects combined into one reading fall within one grain band? Ratio of largest to smallest grain.pitch_mm ≤ 1.3 = Yes. Above 1.3 = No, and they are separate readings. Compared on the nominal figure, which is exact, and never on the measured one, which carries estimation error — question 3.7 guards the measured value. The factor is measured: at a 1.26× pitch ratio the median reading moves 8%, below the register's smallest published state contrast of 1.21× (gate/RESULT-DECIMATION.md).
3.9 Is this value being compared as an absolute against a reading whose nominal pitch ratio exceeds 1.3? Ratio above 1.3 in either direction = Yes, and the §7 cross-grain caveat is required in the same visual field as the value, stating the measured sensitivity rather than a general worry. A reading issued before 0.3.9, which carries no grain, counts as outside the band unless its source deposit states a nominal figure within it.

Domain 4 — Parameter set and departure

# Question Threshold
4.1 Does a documented parameter set exist for this geometry and species? A named set in the reference implementation = Yes. None = No, and the reading is a departure in full and says so.
4.2 Does any parameter differ from that set? Any difference, of any size, is a departure and must be declared with a reason in computation.departure_reason. There is no threshold of triviality.
4.3 Is the response window free of stimulation artifact? The window must exclude the amplifier saturation. For direct cortical stimulation the saturation runs to roughly 15 ms, and the declared departure window is [15, 300] ms. A reading computed on [0, 300] ms for direct cortical stimulation carries the saturation, runs about 16 points high, and is not issuable as a value.
4.4 Is the memory-bounding resample parameter set? PCIst is O(T²) in the response window. resample set = Yes. Unset on a long window = No; the run will exhaust memory rather than return a wrong number, but the reading is not issuable until it is set deliberately.
4.5 Is the scale declared, and is it the only scale in the comparison? PCI-LZ (0–1), PCIst (unbounded, roughly 5–60) and sPCI (0–1) are three incompatible scales. Declared and unmixed = Yes. Any comparison spanning two = No, and it is a category error, not a caveat.

Domain 5 — Interval width

METHOD.md §5 requires the interval to widen and names six conditions. It does not say by how much, which leaves the single most consequential number in the record to the rater's judgment. This domain fixes it.

The base interval is ± 1 standard deviation of the between-unit spread actually observed — between cells, sessions or subjects as the reading's pairing unit dictates. It is never the standard error of the mean. The standard error describes how well the mean is determined; the wideners of §5 are systematic and are not reduced by measuring more cells, so an interval built on the standard error narrows as the evidence accumulates in exactly the case where it should not.

Each condition below scales the base. Where more than one applies, independent wideners multiply and a dependent pair combines as the larger of the two — added in 1.7.0 and measured (gate/RESULT-WIDENERS.md).

Species extrapolation and instrumentation mismatch are declared a dependent pair: crossing species necessarily crosses instrumentation, because no mouse recording is taken on a human clinical grid, so multiplying them counts one gap twice. Measured on the register's own values, an interval built from one cohort and tested against six out-of-cohort values covers 33% at base — the wideners are needed — reaches the nominal 68% at ×1.61, and covers 100% at the ×2.25 the product prescribed. An interval that covers every other reading's value cannot support a scale.

No general independence test is offered, because six targets from one source cohort cannot support one. Any further pair is declared dependent by amendment, with its measurement.

Condition Factor Basis
Small n — fewer than 5 subjects × 1.25 provisional
Instrumentation or cohort mismatch × 1.50 measured. Between-cohort variation with all known variables held constant is about 1.5× (NS-0012 against NS-0026)
Species extrapolation × 1.50 provisional
Preparation extrapolation — in vitro × 2.00 provisional
Proxy substitution — Type C, no perturbation delivered × 2.00 provisional
Parameter departure × 1.25 provisional

A factor marked provisional is a placeholder with no measurement behind it. It is recorded as provisional on the face of the reading, and every provisional factor is listed for review each cycle. This is the honest state of the interval: the wideners are real, their sizes are mostly not yet measured, and saying so is better than either omitting them or implying a precision that does not exist. Replacing a provisional factor with a measured one is a MINOR version change and triggers re-derivation of the affected intervals.

# Question Threshold
5.1 Is the base the between-unit standard deviation rather than the standard error? SD = Yes. SEM = No, and the interval is wrong.
5.2 Is every applied widener named in interval.basis? All named = Yes. A widener applied but unnamed = No.
5.3 Does the interval bracket the value? Yes required. A value outside its own interval is a defect, not a finding.
5.4 Does the interval span the empirical cutoff? Where it does, the reading's own summary line must state that it does not resolve the question it was taken to answer. Being unable to tell is a result and is published like any other.

Domain 6 — Type and modality tier

Both are assigned by decision tree. Neither admits discretion.

# Question Threshold
6.1 Was a perturbation of stated modality and intensity delivered? A stimulation record naming modality and intensity with units = Yes. Either absent = No → the reading is Type C or E and can never share the Type A scale.
6.2 Did the issuer compute the value from the underlying data? A computation record with library, version and commit = Yes. Absent = No → Type D, sharing the Type A scale only where the third party's method and parameters match.
6.3 Is the value a ratio between two issued readings of the same subjects? A derived_from naming two issued Type A or B readings from one deposit = Yes → Type F, dimensionless, never on the Type A scale.
6.4 Is the data the issuer's own? source.access is the issuer's own deposit = Yes → Type A. A third-party deposit under a named open license = No → Type B.
6.5 What is the measurement modality tier? Assigned per METHOD.md §3.4 and displayed on the face of the reading. The tier ceiling is applied without exception.

The tier is not a restatement of the type. The type records where the data came from; the tier records what kind of evidentiary object the reading is. A Type D published value can be M1 or M4 depending on what the third party actually measured, and the difference matters more than the provenance does.


Domain 7 — Caveats

Free text is unbounded discretion, and the caveat field is where an inconvenient fact goes to be worded into harmlessness.

# Question Threshold
7.1 Does the reading carry a caveat for every triggered condition? Each of the following triggers a required caveat: any Probably yes or Probably no answer above; every applied interval widener; every provisional widener factor; any declared parameter departure; small n; species or preparation extrapolation; missing waveform or repetition rate.
7.2 Is each required caveat written from the stem for its condition? The stems are fixed in schema/caveat-stems.json. Free text may be added; it may never replace a required caveat.
7.3 Does the caveat field name what a critic should attack first? Required, and it must name one of: a specific widener, a specific parameter departure, a specific rejected channel or trial class, or a specific source limitation. A caveat list naming none of those = No.

Fixing the stems makes caveat presence fully reproducible between two raters even though caveat wording cannot be. Presence is the part that matters.


Domain 8 — Rating, disagreement and adjudication

# Question Threshold
8.1 Was the reading produced independently by two raters who did not confer? Both ratings recorded in full, at domain level, before adjudication.
8.2 Are the pre-adjudication ratings retained? Permanently. Most bodies adjudicate and discard the disagreement. Retaining it is the only thing that makes the reliability statistics in METHOD.md §9.3 verifiable rather than asserted.
8.3 Is the adjudicator named, with reasoning recorded? Named third party, decision and reasoning in the record.
8.4 Is assessment time recorded? assessment_minutes present for both raters = Yes; absent for either = No. Reported at cycle level.

Conclusion

This instruction is the Nooscope's answer to a failure the register could not previously see. The reproduction gate tests the pipeline against the published world and will catch a pipeline that is broken. It is silent on whether the method is specified tightly enough that two people executing it arrive at the same place, and on the evidence from comparable rubrics the default answer to that is no — unaided expert agreement runs from κ = 0.06 to κ = 0.44, and the one published cross-organization comparison in the adjacent field found near-zero agreement.

Three things follow from adopting it.

It comes before the kill test, not after. The gate is scheduled for month four. This document should be in force first, because the gate will be executed by a person making every judgment above, and a gate whose result depends on who ran it does not discharge its purpose. The cost of writing it first is days; the cost of discovering afterwards that the gate was not reproducible is the gate.

The thresholds here are claims, and most of them are provisional. Six of the seven interval widening factors have no measurement behind them and are marked as such. Several domain thresholds are inherited from an adjacent field rather than derived here. That is the correct state for version 1.0 and the wrong state for version 2.0, and the mechanism for moving between them is the amendment process, not accumulated bench practice.

Low agreement is a defect in this document, not in the raters. Where a domain's agreement falls below κ = 0.40, the response is an amendment to the signalling questions and thresholds above. An instruction to raters to be more careful is not a response, because low agreement is evidence that a judgment is under-specified, and the remedy for under-specification is specification.