Read 2026-07-21. Two color-field works, read as a warm/cool pair:
rothko_no31_yellow-stripe_1958.jpg, 3354×3794.rothko_green-on-maroon_1961_thyssen.jpg, 2991×3384 (5px crop; see PROVENANCE.md).Both raw edge mode (the colored ground is the composition, painted to the edge). Sources and rights: PROVENANCE.md (private structural-research use, © Kate Rothko Prizel & Christopher Rothko / ARS / VEGAP). Pre-registration: PREREGISTRATION.md. Study-local color candidates: color_subset.py, CANDIDATES_local.md.
The counter-pole to Malevich (Study 12). Black Square was all-edge, no interior; a Rothko is all-interior, no committed edge. The study asks whether color carries a Rothko the way tone carried Malevich, and whether the instrument catches what he used. This record is IN PROGRESS: through the first-run trust gate and the color subset, not yet through synthetic controls, the critique, or gather. Findings are labeled by how far they have earned.
Entries 11–13 (2026-07-22) close the study's last open control and record what the spine returned. A source-parity gap the first pass did not check (bit budget, Entry 11) raised the obvious challenge to the colour finding: is B's reading its colour or its compression? Measured both ways, the answer is colour — compression destroys noise-driven CBI and leaves a true chroma boundary untouched (Entry 12), so the split is corroborated from a third direction and B's subsampling if anything suppresses its own value. The same control demoted PM from calibrated to merely discriminating. Entry 13 records the spine outcome (I-25, I-26) and what this study still owes the field_qa guard.
Entry 14 (2026-07-22, on the way to gather) restates two numbers as ranges. A resolution sweep the study never ran shows void_topology_chi is a thin-structure measure (I-10): A 9–30 and B 0–2 across 800px–native, so the headline "30 / 0" is one scale's draw. The contrast holds at every scale and the reading is unchanged; the integers are retired as integers. CBI moves the same way and its A-to-B gap narrows from 3.6× to 1.27× with resolution, which bounds "B ≫ A" to the working scale. tlg is flat across the sweep and is the resolution-robust leg.
Labels: [stable] pipeline-stable · [narrowed] true but scoped · [aid] reading aid · [exploratory] post-hoc / early. [See Critique] Here
Warm/cool, saturated/muted: A is a salmon-over-crimson field on a bright yellow ground that doubles as the stripe between the bands; B is a single muted green rectangle sitting high in a wide grey-violet ground, near-monochrome, its bottom edge dissolving, with two small dark marks in the ground. The confounds are stated up front (PROVENANCE.md): the pair differs on temperature AND form-count AND ground-role, because no two Rothkos differ on one variable.
The one source decision worth recording: the two reproductions are not framed alike. The auction plate (A) is already trimmed to the painted field; the museum plate (B) showed the unpainted tacking-edge canvas as a hard bright spike in the outer ~5px, top and bottom (a reproduction artifact, per-line luminance 158→110 at top vs 78 interior). B is pinned as a 5px crop that removes the spike while keeping the legitimate lighter-ground gradient, bringing it to parity with A so both run raw and identically. A pair's cross-work reads depend on symmetric source treatment. (General note owed to OBSERVATIONS: auction plates trim to the salable image, museum plates show the full object — a Step-0 check for any mixed-source pair.)
The gate stopped the read before any number was trusted, and it caught a confident false verdict. On B, the entire structural mask — 7,813 pixels, 0.077% coverage — sits in one patch at the top-right corner (x 0.75–0.83, y 0.01–0.05): the small dark mark. The green rectangle is invisible to the edge subsystem. So every edge-derived number on B was reading that mark, not the painting: kernel dx +0.289, dy −0.466, μ 1.000, x_p 1.000, and a reading of "coherent_mass · strong connected structural mass" — a confident, wholly false whole-field verdict off a near-empty field. On A, coverage is 1.54% (below the 3–14% healthy band but real), the mask spread across the composition (the band boundaries); its "diffuse_field" reading is weak but honest.
Instrument gap surfaced (owed to the harvest). The shipped field_qa guard (I-23) caught neither: it flags the dead field (100% coverage, robust01 collapse) but has no low-coverage floor. And this is sharper than "add a floor": Rothko-B is a case the shipped gate passes and still gets wrong. B's edge field is NOT degenerate — p98(edge) > 0, because the corner mark is a real high-contrast edge — so I-24's p98 > 0 gate sees nothing wrong; yet 0.077% coverage still produces the false "coherent_mass." So the guard has two independent failure modes with opposite signatures: the dead field (p98 ≈ 0, Malevich's synthetic) and the starved field (p98 > 0 but coverage ≈ 0, Rothko-B). Neither check catches the other's case. This amends I-24's own conclusion, which retired coverage as a "downstream symptom" — Rothko shows coverage IS the gate for a failure p98 cannot see. (Harvest item; both guards needed, stated as a pair.)
Pre-registered: β would split the pair on chroma. It does not. β is substrate-dominated on both (I-12): the bare yellow ground alone reads β 0.60, the bare grey-violet 0.55, as high as or higher than the whole works (A 0.39, B 0.60). And the desaturate/saturate control shows β is chroma-invariant: desaturating A moved β 0.390→0.416, saturating B 0.598→0.620, both flat. β assumes a figure on a neutral substrate; a Rothko's "substrate" IS a compositional color field, so β reads the ground, not the arrangement. This is a class fact, not a Rothko quirk: β does not apply to color-field work. (Candidate note for OBSERVATIONS: I-12 extended to a work class.)
> Read with Entries 11–13. The direction below survives and gains a third, > independent control (Entry 12: compression destroys noise-driven CBI and leaves a > true chroma boundary at 1.000, so B is signal and A is noise). Two re-scopings > apply: PM is a discriminating but not calibrated control — it takes the clean > chromatic pole from 1.000 to 0.000 — so the PM values here are an ordering, not > corrected absolutes; and the plates are not compression-matched (Entry 11), > which bounds the numbers without changing their direction. The 0.20 threshold > referenced below is retired corpus-wide (I-25). > > And read with Entry 14. Every CBI value below is a 1500px read. CBI rises > monotonically with resolution on both plates and the A-to-B gap narrows as it > does: 3.6× at 800px, 1.9× at 1500px, 1.27× at native (A 0.204 / B 0.260). > Entry 12's compression table quotes A's original as 0.2040, which is this same > measure at native, sitting three entries away from the 0.131 below without either > being marked as scaled. The direction never reverses, but resolution is the axis > that compresses it hardest, and it is the one axis Entry 04's three controls > (theorem poles, PM sweep, compression ladder) never varied.
Chromatic Boundary Independence measures the fraction of strong chromatic boundaries that sit where luminance has none: color drawing lines light doesn't. Correction (survey, 2026-07-21): CBI is already a spine measure (vtl.py, output as chromatic_boundary_independence in every read_image — the "Matisse/Bonnard signal"). It was in the first-run gate report; I did not read the vtl block and reimplemented it study-local in color_subset.py. The values agree (spine A 0.131 / B 0.246; port A 0.127 / B 0.241), so the finding stands, but CBI is the instrument's, not this study's coinage.
CBI answers the core question β fumbled: is color doing the structural work. But those RAW values 0.131/0.246 rest on a bare 0.20 threshold with no control, so the control was built (cbi_anchor.py, the I-R computable-expectation lever) — and it found a real problem before the number carried a critique.
The anchor: CBI is correct in principle, catastrophically noise-sensitive in practice. On a clean synthetic, the poles are exactly as documented — an equal-luminance chromatic boundary reads CBI 1.000, a pure value boundary 0.000. But add even σ=0.002 grain and the chromatic pole collapses to 0.055: the noise floods the chroma-edge field (the "strong chromatic boundary" set jumps from 1,000 real-boundary pixels to 37,500 noise pixels), so real-image CBI is deflated by an uncontrolled amount. The measure has no noise-floor control, and every real painting has canvas texture and brushwork. This is the I-23 family inverted: sparse input fooled the edge subsystem high; noisy input fools CBI low. Corpus-wide — CBI has been in all twelve prior studies' output at raw, noise-deflated values.
Noise-controlled (PM-denoised), the finding survives and sharpens. A: raw 0.131 → PM 0.055 (its raw value was mostly noise; denoised it falls to the value-pole floor — genuinely light-drawn). B: raw 0.246 → PM 0.209 (its autonomy survives denoising — genuinely color-drawn). The direction B ≫ A widens to ~4× and A lands at the value pole, so "color carries B, light carries A" is now demonstrated with the control, not inferred from a straddled threshold. The raw values and the 0.20 cutoff are retired here; the PM-denoised split carries the claim.
PM is now a load-bearing control, so it was swept too (diag_pm_sweep.py, the Malevich round-5 discipline — a control is an apparatus with its own free parameter). Across 12 settings (n_iter 5–40 × kappa 0.06–0.24): B stays color-drawn (0.168–0.224), A falls to the value-pole floor (0.021–0.082), and the direction never closes (ratio 2.7×–7.9×). It widens as PM strengthens — A drifts toward zero (its CBI was noise), B holds (its CBI is signal). So B does not drift toward A; the split is Rothko's color, not the denoiser's aggressiveness. The claim is anchored on both ends now: the theorem-object poles (measurement) and the PM sweep (denoiser). The last free parameter is closed before the critique. (Spine harvest: CBI must run on noise-controlled fields; the fix, the mechanism, and the denoiser's stability are all known. The corpus-wide re-check is now done — see CBI_CORPUS_RECHECK.md and the spine-thread brief HANDOFF_spine_CBI.md.)
> Read with Entry 14 (2026-07-22). The direction below is confirmed at every > resolution tested. The integers are not: void_topology_chi uses a 31px local > block and a 256px body floor, both absolute pixel scales, so it is a > thin-structure measure in the I-10 sense and must be swept before its value is > quoted. Swept, A runs 9–30 across 800px–native and B runs 0–2; the "30" > below is the peak of that sweep, at the 1500px working scale this study happened > to use, and the "0 exactly" holds at 800–2600px but reads 2 at native. Restate as > ranges with the scale attached. The finding is unaffected.
The cleanest discriminator in the study is an integer already in the spine, and it was sitting unread: void_topology_chi reads A 30 distinct void bodies, B 0 (both at the study's 1500px working scale; ranges in Entry 14) — the instrument's own count of internal luminance structure, no threshold, no coinage, no control. B has no inside; A has an interior the instrument can count. That is the counter-pole in two integers, and it is the contrast that carries it, not the value of either.
Verified, per the Malevich island-count discipline (the integer is only as good as what the bodies are). First, a computation check: my hand-rolled replication gave B 46 (a 2–98 percentile norm amplified B's sub-threshold noise into phantom bodies); the spine's own min-max _norm01 gives B 0 at this scale (voids_frac 0.000, nothing survives the local threshold; 2 at native, Entry 14). The spine value is right; my first pass would have counted noise — the fragility this check exists to catch, caught on me. Second, the identity check: A's 30 bodies trace the crimson band's boundary against the yellow (segments around the lower rectangle; only 2 on the pink, because pink-on-yellow is a weak value boundary and crimson-on-yellow a strong one). Pointable structure, not artifacts. A_void_bodies.png.
And it converges with CBI from the opposite direction. A's interior is value-defined (a crimson/yellow value boundary), so luminance-based void topology finds it (30) and CBI reads it low (0.13, light-drawn). B's interior is chroma-defined (green/grey near-equal value), so luminance topology finds nothing (0) and CBI reads it high (0.25, color-drawn). Two independent measures, same story: A's structure is a value boundary, B's is a chroma boundary. The void-body contrast is the headline; CBI is the color-specific refinement underneath it. (Both quantities are scale-dependent and both are restated as swept ranges in Entry 14; the contrast is what survives, in both.)
The palette read sits beneath both: A chroma ~0.13, lightness wide (0.41–0.76), a lightness-structured palette whose accent is less saturated than the ground ("muted-key") — No. 31 reads loud but its energy is value plus the yellow's hue, not chroma contrast. B chroma ~0.03, near-monochrome in both dimensions. chroma_mean (0.14 vs 0.03) carries the split β-free; the spine's chroma_variance does NOT (A 0.091 ≈ B 0.087), which is why chroma_mean earns its place.
The upper-right mark is a small dark red-brown hooked gestural stroke at (0.79, 0.05), mild contrast (0.058 below the green). It is the composition's only edge-event; whatever its origin, it functions as the field's single accent. It is read with the accent lens, not the kernel's mass machinery (which mislabeled a 0.08% speck as the painting's structure). Origin not claimed: no museum documentation calls it intent, pentimento, or damage; Rothko's thin unvarnished washes let underlayers and handling marks read through, so it is undecidable from the plate (same discipline as Malevich's under-paintings — measured, not attributed). A second, fainter mark sits bottom-center; the collector reads it as a diffuse stain (tone near the green) versus the upper-right gesture — one belongs to the atmosphere, one punctuates it. Whether they are one family is an open object-level question.
Because the edge subsystem starves, B's structure was mapped on the tone+chroma modulation field (read2_color.py). It concentrates entirely at the boundary diffusion zone — the feathered green-to-grey transition — with hotspots at the bottom-left corner bleed and the right side, while the green center is the lowest-modulation region (calm, uniform). The 3×3 grid puts the minimum dead center (0.161) and maxima top-right (0.242, the mark) and down the left column (0.208). That is the collector's own read, independently measured: a calm reinforced center with two offsetting disturbances, left and right. Credited to his eye; measured as structure; the grid found modulation minima and maxima — whether that is "a calm center with two offsets" is the collector's framing, not the instrument's verdict — and it is bounded by this plate (a stain here could be the painting or the reproduction).
A, run to the same depth (the asymmetry closed). A's modulation map peaks in the middle row (0.29, the pink / yellow-stripe / crimson transition zone) and down the left column, exactly where the two-band structure and the crimson-edge void bodies sit. So both works put their structure at boundaries, but of opposite kind: A's is a horizontal value transition (the band edges, caught by luminance topology and the modulation map); B's is a soft chroma perimeter (caught only by CBI and the modulation map, invisible to luminance topology). The pair is now measured symmetrically, not asserted on one arm.
> Resolved at gather (2026-07-22). This measure duplicated > toolkit/patterns/hue_tension.py, on the shelf since 2026-07-02 from the same > source notebook — the fourth I-28 rediscovery, and the one where the two > versions disagree (B cool 0.9286 here vs 0.9988 floored; tension 0.072 vs 0.0012, > 60×). Tested against known-answer inputs rather than assumed > (warmcool_variant_test.py): this study's version is the correct one. The > shelf's relative chroma floor is framing-dependent and reports "monochrome pole" > for a full warm/cool standoff on full-bleed colour work. The numbers below stand > as published; the shelf form is retired and this one entered the register as C-8.
Warm/cool (color_subset.warm_cool): A is 99.97% warm, B is 92.86% cool, and both have near-zero warm/cool tension (0.0003, 0.072). (Both at the study's 1500px scale; native reads A 0.985 / 0.015 and B 0.930 / 0.071, so B is scale-stable and A moves slightly — Entry 14's lesson applied.) Rothko commits to one temperature per work and modulates within it — A within warm (pink/crimson/ yellow), B within cool (green/grey-violet). He is not working warm/cool complementarity at all; the interplay is within-family. This is the launch point for the color-interplay / hue-interval measure (future).
The tonal survey followed the process lesson — spine first — and found the tonal tools already there: tonal_local_global, tonal_hierarchy, tonal_gestural_offset, none reinvented. And tonal_local_global (tlg) is a third measure telling the counter-pole: A tlg 1.0 ("high tonal plane independence — strong value differentiation across regions"; quadrant zones light-top / dark-bottom, pink 4.7 / crimson 2.6), B tlg 0.126 ("mild variation"; all four quadrants ~2.7–2.97, no value planes). So three independent spine measures converge — void_topology 30/0, tlg 1.0/0.126, CBI(PM) 0.055/0.209 — A is value-structured, B is chroma-structured. The tonal side confirms what void and color already said.
Two caveats. tonal_gestural_offset is corrupt on B (A 0.031, B 0.003): it uses the gestural/edge centroid, which on B is the corner mark (Entry 02), so B's offset is the tonal field vs a speck. Set aside on B, like the kernel. And the spine's tonal tools measure BETWEEN-region planes (tlg's 2×2 grid), not the WITHIN-field atmospheric modulation — Rothko's veiling, the soft tonal transitions inside a field that the collector described and the hand-rolled modulation map (Entry 07) caught. The principled tool for that is Local Contrast Ratio (Parallax Cell 9, "slow tonal transitions without edge detection"), which the spine lacks — so it was ported study-local (lcr_atmosphere.py).
The atmosphere, measured — and it flips. Run LCR on luminance (tonal atmosphere) and on chroma (chromatic atmosphere): A tonal-LCR 0.035 / chromatic-LCR 0.014 (ratio 0.39 — tonal atmosphere, the crimson/yellow value fade); B tonal-LCR 0.007 / chromatic-LCR 0.026 (ratio 3.49 — chromatic atmosphere, the green veiling). A luminance-only LCR read B as near-empty (0.007), which was the same false-absence the edge subsystem gave — B's veiling is color, not value, and on the chroma field it is the strongest atmosphere signal in the pair. So the atmosphere completes the counter-pole in its own dimension: A works in value at every level (structure, boundary, atmosphere), B in chroma at every level. The chr/ton ratio (0.39 vs 3.49, ~9×) is a clean new discriminator — the "atmosphere temperature" of the work. Compression-controlled (Entry 12 iii, 2026-07-22): LCR samples at an 8px radius and JPEG chroma subsampling works on 8×8 blocks, so this ratio had a live scale collision with the plates' unequal bit budget (Entry 11). Degrading A down the ladder moves its ratio 0.39 → 0.47 at quality 12, against an A-to-B gap of 8.9× — ~2% of the difference, and in the direction that would close it. The flip is the paintings', not the compression.
The whole study reduces to void_topology_chi: A 30, B 0 at 1500px, A 9–30, B 0–2 swept across 800px–native (Entry 14) — and it does not stand alone: tlg (1.0 / 0.126) and PM-controlled CBI (0.055 / 0.209) point the same way, from the tonal and color sides. A has an interior (a value boundary the instrument counts); B has none (its boundary is chroma, and luminance topology cannot see it). And it nests with Malevich: Black Square was all-edge with an empty interior (the gradient-void); Rothko-B is all-interior with no countable structure at all (0 void bodies) because its one boundary is refused to luminance. Both artists put structure at the boundary of a uniform field — Malevich's boundary edge-committed (a hard ring), Rothko's edge-refused (a soft chroma diffusion) — which is why Malevich needed the edge subsystem and Rothko needs the color/tone subset. The counter-pole, measured on the first pass rather than asserted, and corroborated by three independent spine measures (void topology, tlg, PM-controlled CBI) pointing the same way — two of them threshold-free. (Entry 12 adds a fourth, independent of all three: under a compression ladder B's CBI behaves like a real chroma boundary and A's like noise. The two threshold-free luminance measures still carry the reading on their own.)
Entry 01 brought the plates to parity on framing (the 5px crop removing B's tacking-edge spike) and PROVENANCE matched them on resolution (both 3000–3800px). A third axis went unchecked, and on a colour measure it is the one that matters: bit budget.
| source | px | file | 8px chroma-block power in the CBI-qualifying set | |
|---|---|---|---|---|
| A No. 31 | Christie's catalogue master | 3354×3794 | 17.2 MB | 0.78% |
| B Green on Maroon | Thyssen museum web viewer | 3000×3393 | 1.2 MB | 26.8% |
Same pixel dimensions, roughly 14× different bytes per pixel. B's chroma is 8×8 subsampled hard enough that 26.8% of the spectral power of its chroma-boundary set sits exactly on the JPEG block period; A's is 0.78%, i.e. absent. (The pinned B reads 27.6% — re-saving after the crop cannot restore discarded chroma.) Measured on the qualifying set's column-density spectrum.
No better B exists in what we hold. RP-Finds/ was searched: the 17.2MB file is A's own source, the Thyssen visor.jpg is B's, and there is no other Green on Maroon plate. Nor is there a source-matched substitute pair — the other low-chroma works (black-on-maroon, tate-1958, T01031_9) are ~1 MP against A's 12.7 MP. So this asymmetry was not a choice that could have been made differently with the sources on hand. It is a disclosed bound, not a correctable error, and it is the reason Entry 12 exists.
Generalises Entry 01's owed note: the mixed-source Step-0 check is not only "do the plates frame alike," it is "do they carry comparable bit budget," because an auction catalogue master and a museum web viewer differ by an order of magnitude on exactly the channel a colour measure reads.
Entry 11's asymmetry raises an obvious challenge to Entry 04: is B's high CBI its colour, or its compression? Asserting either way would repeat the mistake this study keeps catching. It was measured, two ways, and the answer is clean.
(i) Degrade A toward B's regime. Re-encode the clean 17.2MB A at falling quality with 4:2:0 chroma subsampling and re-measure:
| A at quality | MB | 8px block power | CBI |
|---|---|---|---|
| original | 17.2 | 0.78% | 0.2040 |
| 50 | 2.2 | 6.70% | 0.0607 |
| 30 | 1.3 | 12.25% | 0.0653 |
| 20 | 0.8 | 15.71% | 0.0962 |
| 12 | 0.5 | — | 0.0000 |
Compression does not manufacture CBI; it destroys it. Driving A's block signature from 0.78% to 15.7% drove its CBI down 0.204 → 0.096, and at q12 to zero. So B's 0.2596 cannot be an artifact of B's subsampling.
(ii) The theorem side — does compression preserve a REAL chroma boundary? The clean equal-luminance chromatic pole (truth = 1.000), through the same ladder: 1.0000 at lossless, and 1.0000 at every quality down to 12. A genuine chroma boundary is completely compression-robust.
The two together are a discriminator, and they corroborate Entry 04 from a third direction. Real chroma boundaries survive compression; noise-driven CBI dies under it. B survives both PM and compression → its CBI is signal. A dies under both (PM 0.204→0.023; q50 0.204→0.061) → its CBI is noise. That is the same verdict the anchor and the PM sweep reached, now on a control that does not share their machinery. B ≫ A is not weakened by the source gap; if anything B's subsampling suppresses its own value, so the split is conservative.
And it exposes a real problem with PM as the control of record. On the clean chromatic pole, PM takes CBI from 1.000 to 0.000 — the denoiser destroys the exact signal the measure exists to detect, where compression preserves it intact. So PM discriminates on these two images (Entry 04's sweep is real: 12 settings, direction never closes) but it is not calibrated — it does not return truth on a known input, so a PM-controlled value is not "the true CBI," only a noise-suppressed ordering. Entry 04's numbers stand as an ordering; they should not be quoted as corrected absolute values. The compression ladder is the better control here precisely because it leaves the theorem pole untouched.
(iii) The same ladder, run on the atmosphere-temperature ratio (Entry 09) — it holds. That ratio is chromatic-LCR ÷ tonal-LCR, and LCR's local radius is 8px while JPEG chroma subsampling works on 8×8 blocks — the two scales collide, so B's 3.49 could in principle have been reading its own compression. Tested rather than assumed. Baseline reproduces Entry 09 exactly (A 0.0352 / 0.0139 / 0.39; B 0.0073 / 0.0256 / 3.49), and degrading A down the ladder moves its ratio only 0.39 → 0.41 → 0.43 → 0.47 at quality 12, a file ten times smaller than B's actual source. Compression shifts the ratio 1.2× against an A-to-B gap of 8.9× — about 2% of the difference, and in the direction that would close the gap, not open it. The scale collision is real in principle and negligible in fact; the atmosphere finding is compression-controlled and stands.
Also settled: the corpus re-check's causal wording. Across its own 15 images corr(noise, |raw→PM shift|) = −0.064 — the shift does not scale with the noise metric used, which is consistent with PM lowering CBI on any image rather than removing a noise floor. The re-check's universality (all 15 negative) is solid; its attribution of that to noise is not, and the corrected account is that PM depresses the statistic generally.
The CBI work was briefed out (HANDOFF_spine_CBI.md) and has been actioned. What returned, and how it lands on this record:
robust01 amplifies a near-empty field to full scale, so a chroma-free image reads CBI 1.000 where the theorem demands 0.000, and the p98 guard cannot see it (p98 reads a healthy 1.0). Shipped guard: field_qa.chroma_support, flagging at machine-zero only, so low-but-real chroma (B's 0.039) is reported and never suppressed.void_topology. Two consumers, one normalisation.field_qa's starved field, CBI's noise behaviour, and the _norm01 phantom bodies are one thing: a measure that is correct on typical input and returns confident, plausible, wrong numbers on a regime it was never checked against — sparse, near-empty, noisy, or renormalised — and never raises an error. Kin to I-23/I-25/I-26.p98 > 0, coverage 0.077%, a false coherent_mass) is a case the shipped I-23 gate passes and still gets wrong. Coverage was retired there as a "downstream symptom"; Entry 02 shows coverage IS the gate for a failure p98 cannot see. Two guards, two signatures, neither catching the other's case. Logged for the spine; not this study's to fix.How the CBI numbers in this record should be read. They are within-pair orderings on plates of unequal bit budget, anchored on the theorem poles (Entry 04), swept on the denoiser (Entry 04), and controlled against compression (Entry 12). They are not absolute values, not cross-work ranks, and not measured against the retired 0.20 line. The pair's conclusion does not rest on them alone: void_topology_chi (30/0 at 1500px, 9–30 vs 0–2 swept) and tlg (1.0/0.126, the one measure of the three that is resolution-robust — Entry 14) are threshold-free and luminance-based, and carry the reading without the colour measure.
Found by re-running the pinned plates at native resolution instead of the study's 1500px working scale, on the way to gather. Everything stored reproduces to the decimal (mask coverage 0.0154 / 0.0008, μ 0.320 / 1.000, Δy −0.040 / −0.466, β 0.425 / 0.635, τ 0.450 / 0.865, chroma_mean 0.1356 / 0.0278, the atmosphere ratio 0.39 / 3.49). Two things do not, and one of them is the headline.
| px | A chi | B chi | A cbi | B cbi | A tlg | B tlg |
|---|---|---|---|---|---|---|
| 800 | 9 | 0 | 0.060 | 0.217 | 1.000 | 0.128 |
| 1000 | 18 | 0 | 0.086 | 0.228 | 1.000 | 0.127 |
| 1200 | 26 | 0 | 0.105 | 0.238 | 1.000 | 0.125 |
| 1500 (the study's scale) | 30 | 0 | 0.131 | 0.246 | 1.000 | 0.126 |
| 2000 | 26 | 0 | 0.173 | 0.257 | 1.000 | 0.128 |
| 2600 | 28 | 0 | 0.183 | 0.261 | 1.000 | 0.129 |
| native (3794 / 3384) | 24 | 2 | 0.204 | 0.260 | 1.000 | 0.137 |
void_topology_chi is a thin-structure measure and was quoted as if it were a count. It thresholds with threshold_local(block_size=31) and keeps bodies over 256px, both absolute pixel scales, so it is squarely in I-10's resolution-dependent family, and I-10's rule is that such a measure is a reading aid until it has been swept. It never was here. A's count runs 9 to 30 and the study's 30 is the peak of the sweep, landing on the scale the study happened to work at. B runs 0 to 2. The sting is that the study's own pre-registration had set fine-structure measures aside as "not read as findings", and the headline then landed on one.
What survives, and it survives everywhere. A ≥ 9 against B ≤ 2 at every scale tested: A has a countable luminance interior, B has effectively none. The contrast is 12× at its narrowest (800px, 9 vs 0 — unbounded) and never inverts or closes. The reading is untouched. What is retired is the pair of integers as integers, and the phrase "B 0 exactly", which is true at 800–2600px and reads 2 at native.
CBI is the more exposed one, and it is exposed on the axis its three controls did not vary. It rises monotonically with resolution on both plates, and the gap narrows as it rises: B/A is 3.6× at 800px, 1.9× at the study's 1500px, and 1.27× at native. Entry 04's controls (theorem poles, PM sweep, compression ladder) are real and all hold, but all three vary noise and encoding, never scale. And the record already contained the evidence without noticing it: Entry 12's ladder quotes A's original as 0.2040, which is simply this measure at native, three entries away from Entry 04's 0.131. B ≫ A is a 1500px statement. At native it is B > A by a quarter, which is not the same claim. Both entries now say so; neither number is withdrawn, because neither is wrong at the scale it was taken.
tlg is clean — 1.000 against 0.125–0.137 at every scale. Of the three converging measures it is the only resolution-robust one, and it should carry the headline it can support.
Method note, which is this study's own lesson turned on itself: the scorecard says "pre-registration protects against inventing a story; it does not protect against a true story confirmed by a false measurement", and names inspecting the mask as what caught prediction 5. The same discipline applied to chi — asking what the count was computed on, at what scale — would have caught this in-study. It was found by re-running, which is the cheap version of the same question. Producer: sweep_resolution.py.
| # | Prediction | Verdict |
|---|---|---|
| 1 | Colour is the spine; β splits the pair, A high / B low | REFUTED, and inverted — A 0.390, B 0.598. Both substrate-dominated (bare yellow 0.599, bare grey-violet 0.546), and β is chroma-invariant: desaturating A moved it 0.390→0.416, saturating B 0.598→0.620. The prereg named its own consequence ("if β does not separate A from B, the colour subsystem is not reading composition on Rothko, and the study says so") and Entry 03 said it. The thesis survived on a different carrier (CBI, chroma_mean, void topology); the named carrier failed |
| 2 | Substrate check: B's β at risk of substrate domination | HIT, and wider than predicted — both works are substrate-dominated, not just B. The caveat fired before β was quoted, as designed |
| 3 | Both MIDTONE works; low τ, lowest for B; thin shadow/highlight | SPLIT — substance right, metric direction wrong. Midtone concentration hit hard (A mid 0.600, B mid 0.861, B highlight exactly 0.0). But τ went the other way: B reads 0.865, the highest in the pair and above Malevich's 0.699. τ measures tonal concentration, so "close-value veils" predicts high τ, not low. A specification error in the prereg, not a fact about Rothko — and it inverts the "Rothko refuses the poles" framing: B is more tonally concentrated than Black Square, not less |
| 4 | gap HIGH both; mask locks onto band boundaries; field_qa alive but low coverage | SPLIT, with the study's most useful call inside it. The field_qa clause is a clean HIT — alive (p98 > 0) and low coverage, A 1.54% / B 0.077%, predicted before the run. gap is moderate not high (0.346 / 0.321 against Malevich's 0.509). Mask-on-band-boundaries hit for A, failed for B, where it locked onto the corner mark (Entry 02) |
| 5 | Mass/kernel: A distributed, B high and single; B centroid above centre | Numbers matched; mechanism false — the instructive one. A μ 0.320 (distributed ✓). B μ 1.000, Δy −0.466 — both exactly as predicted, and both computed off a mask that is 0.077% of the frame and contains only the corner mark. The prior critique, the pre-registration, and the number all agreed, and the number was still an artifact. Only inspecting the mask caught it (Entry 02). Pre-registration protects against inventing a story; it does not protect against a true story confirmed by a false measurement |
| 6 | DGI/torque/R_spatial: control-gated, probably no purchase | Not tested, correctly. The matched soft-rectangle synthetic was never built, so per the prereg's own order of operations these were not quoted. DGI (A 53.3, B 48.4) sits unread in the gate report and stays there |
| 7 | Address moderate, non-figural | Not run |
| 8 | RCP Not-a-basin on the mass field | Not run |
| 9 | Headline: colour + soft-structure carry both; brightness/edge/gravity weak; the pair separates on colour and tone more than geometry | HIT in direction, with a twist worth keeping. Edge/gravity weak: emphatic (Entry 02). Separation on colour/tone rather than geometry: yes. But the cleanest discriminator turned out to be a luminance measure reading an absence — void_topology_chi, 30 / 0 at the working scale and 9–30 / 0–2 swept (Entry 14) — not a colour measure reading a presence. "Colour carries Rothko" is true of B's boundary; what proved it was light failing to find one |
The pattern in the misses, stated because it is the useful part. The study predicted where the instrument would struggle far better than it predicted what the values would be. Both subsystem-health calls landed (2 and 4's field_qa clause — B's starved edge field was called before the first read). The two real misses are of different kinds and only one is about Rothko: prediction 3 was a metric-definition error (τ's direction), and prediction 5 was confirmation by artifact — the rarest and worst kind, because nothing in the process disagreed with it. Prediction 1 was refuted outright and the study reported it, which is the scorecard's best evidence that it was not fitted to its conclusion.
This study, before the critique (not "cleanup" — half the pair):
fig_edge_void.png). L-M opponency in the luminance void: both compositions reappear in the color channel; B is 100% void to luminance yet fully legible in color. Ported study-local (edge_void.py).New candidates (the genuinely-new, post-survey):
chroma_mean + warm_cool/temperature → toolkit/CANDIDATES.md at gather. These are NOT in the spine and each earns it (chroma_mean discriminates where the spine's chroma_variance does not; temperature is a clean new axis). NOTE: CBI is NOT a candidate — it is already spine (Entry 04); its calibration is a spine task, moved to the harvest.fig_atmosphere.png) — but the promotion target is CORRECTED (2026-07-22). Run on luminance AND chroma; the atmosphere flips: A tonal (chr/ton 0.39), B chromatic (3.49) — the counter-pole completed in its atmosphere dimension. LCR itself is NOT a promotion candidate: toolkit/patterns/local_contrast_ratio.py has existed since 2026-07-02, from the same Parallax Cell 9 source, with the same 8/32 radii. Entry 09's "the spine lacks it" was true of the spine and wrong about the shelf, which was never checked — the third time this study or its predecessor rebuilt an existing tool (after CBI and void_topology_chi), now logged as I-28. What is genuinely new and does earn promotion is the chromatic extension: running LCR on the chroma field, and the chr/ton ratio as an "atmosphere temperature" scalar. That belongs as an extension to the existing shelf pattern, not as a new tool.vtl-color-scheme-dialectic-reader as the adversarial harmony layer.tonal_local_global, tonal_hierarchy, tonal_gestural_offset) are already in the spine — none reinvented, the lesson held. tlg is a third converging measure (A 1.0 / B 0.126). The only gap is within-field atmosphere → Local Contrast Ratio (item 2). tonal_gestural_offset is corrupt on B (uses the edge centroid = the corner mark); set aside on B.Instrument harvest (spine changes — NOT this study, flagged for later):
field_qa guard has TWO independent failure modes, not one. Malevich's dead field: p98(edge) ≈ 0 (I-24's gate catches it). Rothko-B's starved field: p98(edge) > 0 (the corner mark is a real edge, so the field is NOT degenerate) yet 0.077% coverage produces a false "coherent_mass" — a case I-24 passes and still gets wrong. So the fix is not "add a coverage floor"; it is two distinct guards: p98 > 0 for the dead field, a coverage floor for the starved field. Neither catches the other's signature. This amends I-24's own conclusion (which retired coverage as "downstream symptom") — coverage IS the gate for a failure p98 cannot see. Best-evidenced instrument item since I-24; fold both into the ISSUES.md I-Q writeup as a pair.CBI_CORPUS_RECHECK.md + HANDOFF_spine_CBI.md). The anchor found the mechanism: CBI reads its poles (1.0 / 0.0) only on clean synthetics, and σ=0.002 grain collapses the clean chromatic pole to 0.055. The corpus re-check then measured the effect on the real prior images and corrected the sign I first assumed: on real color-field images the raw CBI is INFLATED, not deflated — flat-luminance interiors are full of chroma-noise that counts as "strong chroma + weak luminance," i.e. false pure-chromatic-boundary. All 15 corpus images drop under a PM control (raw→PM −0.019 to −0.091). Deflation is the clean-pole behavior; inflation is the real-image behavior. What the re-check settled: - The domain hypothesis is refuted. AI plates are not smoother than paintings (mean noise 0.013 both; Sora/MJ among the noisiest), so the AI home domain never supplied implicit noise control. This was closer to a flaw all along than a transfer flaw; it stayed invisible because the one founding claim (MJ ≈ 0) is floored and cannot inflate off the floor. - This unifies two existing caveats. The METRICS.md dark-key caveat (Caravaggio, 0.236→0.190 denoised) and I-12 the substrate class (Soejima) are the same mechanism at two texture sources, and the real scope is universal, not those two special conditions. - Pollock is NOT a deflation casualty — its raw CBI is already low (0.048) and stays low under control (0.015); "color follows luminance" holds direction. - The cross-work compare z-tables ARE contaminated (raw CBI across works with different noise floors); Klimt's "0.191 corpus high" ranking claim is the load-bearing exposure. These are the reverse-fix candidates. - Survivors: MJ ≈ 0, and all low-CBI directional reads (Matisse, Pollock, Bruegel), plus this study's B ≫ A. The fix: run CBI on noise-controlled fields (PM works, swept stable) or report a noise-floor beside the raw value. Anchor in hand (cbi_anchor.py), mechanism known, denoiser swept (diag_pm_sweep.py), corpus quantified (corpus_cbi_recheck.py). Not a Rothko task; top of the harvest, brief written.A note for the harvest writeup — one family, three coats. The field_qa starved field (item 1), CBI's noise-sensitivity (item 2), and the _norm01 phantom-void-bodies incident (Entry 05) are the same finding three times: a measure that behaves correctly on clean/typical input and returns confident, plausible, WRONG numbers on an input regime it was never checked against — sparse (field_qa high), noisy (CBI inflated on real images, and sign-flipped on the clean pole), or renormalized (phantom bodies) — and fails silently. The lesson each time is the same: check the measure against its regime extremes (the computable poles, the noise floor, the alternative normalization) before trusting it, because the failure never raises an error. Name the family once when it graduates to OBSERVATIONS; kin to I-23/I-24.
void_topology_chi already in the spine after I reimplemented/missed them. Add to the entry protocol: audit read_image's full output (including the vtl block) and grep the spine for the measure before building a new one.Open object-level questions (need the object / documentation, not the plate):
Research extensions (art-historical, beyond this pair — see CRITIQUE_NOTES.md):
chroma_mean should be read as evidence about his intent (reinforces "not intent," for a documented reason).In progress: through the first-run trust gate, the color/tone subset, the CBI anchor (both ends), the tonal survey, the atmosphere read, and the resolution sweep. Fourteen entries; the counter-pole rests on three converging spine measures plus the atmosphere flip, with two of the three now stated as swept ranges rather than point values (Entry 14) and tlg identified as the resolution-robust leg. The CBI corpus re-check is done and briefed to the spine thread (CBI_CORPUS_RECHECK.md, HANDOFF_spine_CBI.md): the finding re-scoped from "deflated" to "inflated on real images," the AI-domain hypothesis was refuted, and it unified the existing dark-key (METRICS) and substrate (I-12) caveats into one mechanism. Next per the future list: the critique on confirmation (the pre-critique study work is complete). Nothing has touched the spine.