The corpus's most ambitious displaced-subject test: the nominal subjects — King Philip IV and Queen Mariana — appear in their own state portrait only as a dim reflection in a back-wall mirror. Measured (v3 verified boxes — the first box set failed a review audit and was re-pinned; Entry 04): the reflection is painted at full local substance and given 0.61% of the frame — 15.2× less mass-share and 50.2× less saliency-share than the Infanta — ratios that never fall below 13.5× / 43× in 200 jittered re-draws of every box (Entry 12) — dead-last of the named objects on saliency; the dog carries five times their weight; the other lit aperture, the doorway with a chamberlain in it, is the densest local saliency peak in the dark wall and out-pulls the mirror 10.21× (Entry 12; and in every one of 200 jitter draws; the single brightest saliency point is the Infanta's own dress). Around them: a shadow-key at the corpus's pole, statistically tied with the Caravaggio (0.896 vs 0.903, same-pipeline compare; the first-pass "deepest in corpus" is refuted in Entry 06), colour riding the drawing (coupling +0.361, corpus high), the most anti-default composition measured (DGI 18.8), and a central address channel reported as a measurement: it beats a phase-scramble null (p ≈ 0.002 — real figural content, not the room's spectrum) and sits at the 70.4th percentile (95% CI 66.4–74.4, pooled over two independent seeds) of rearrangements of its own cast (Entry 08) — real central bilateral mass, but a composition maximally off the default basin is only mildly better than shuffling its own figures on central address, and conditional on the Infanta being central the other six placements score below random (−1.4σ, n=40 of 300, 7.5th pctile [exploratory]): the frieze dissipates central address rather than concentrating it. Pre-registered (10 predictions, locked before any number); every claim below carries its grade.
Russell Parrish · Parallax Metrology · 2026. Deterministic and reproducible from the scripts named in each entry. Art-historical context cited in Sources. More info: www.parallaxmetrology.com, full critique
Extreme dark key, measured on the first file grabbed (a Prado colour-corrected "FXD" derivative): median L 46, 72% of pixels below L70, 5% above L150. On a surface this dark the I-11 speck check is mandatory, and it fired at the worst tier: PM verdict "meaningful" (0.0178). The cause was the file, not the canvas. The FXD's shadow-lift and sharpening had raised the black floor (a hazy band the eye read as glare) and amplified craquelure. Re-deriving the working image from the raw Prado gigapixel (26065×30000, 10× box-average to 2600px) cut dark-void high-frequency noise from 6.86 to 2.92, deepened the void from L45 to L22, and dropped the verdict to "light" (0.0096).
Las Meninas brought a condition the corpus hadn't met: a large foreground repoussoir object — the canvas-back Velázquez stands behind (x ≈ 0.02–0.16) — whose hard right edge is the strongest edge in the picture. Measured artifacts, each with handling: the canvas-back edge + left dark margin (x < 0.17); two physical canvas seams at x ≈ 0.259 and 0.636 (3.9× and 3.6× the dark-void median vertical-edge energy — the canvas is stitched from three widths); the lit doorjamb at the right edge (x > 0.975); a floor contour-trap along the bottom (y > 0.93). Handling: masked for the accent read (and any future armature read); thin and minor for mass, tone and saliency. Seam 0.259 passes through the Velázquez and Sarmiento probe boxes (noted where those are quoted); the mirror and doorway boxes are clear of both seams — checked, since the mirror's local substance carries the headline.
The default structural accent channel returned 24 accents of which 21 were artifact — a column of fragments down the canvas-back's edge plus the doorjamb (the Degas I-7 edge-bias in its most extreme observed form). The chromatic channel (σ6) was 33% artifact; after removing the four named zones, 12 real accents remain and they land on the figure frieze: Velázquez's face, the búcaro-menina, the Infanta, Maribárbola, Pertusato, the chaperones, the dog. Top leverage 0.83 falling smoothly to 0.49 — distributed punctuation, no climax.
review_checks.json). The
royals carry no punctuation.Is the Santiago cross one of the twelve? Checked (review 6). The
red cross of Santiago on Velázquez's breast is post-1659 paint — added ~3
years after the picture, after the knighthood (per Laing not by the King, which
kills the legend). If it were one of the twelve accents, a punctuation mark of
the study would not belong to the 1656 composition. Rendered and checked
(scout_velazquez_cross.png): the two accents inside his figure box
sit on his eye and the red paint on his palette — both original;
the cross itself carries no accent. The read is clean of the later
addition.
subject_probes.py (I-4 mass-locked reads on the
PM-denoised fields; nine verified boxes). The crux — the mirror against the
Infanta and the doorway:
| channel | Infanta ÷ mirror (v3 boxes) |
|---|---|
| mass, local intensity | 1.54× |
| edge, local | 1.44× |
| mass, share of frame | 15.2× |
| saliency, local | 5.01× |
| saliency, share of frame | 50.2× |
scout_zone_boxes_v3.png), every box zoom-verified
(scout_v3_zooms.png) before the rerun. Isabel then failed a second
way: the I-4 nearest-to-centre mask lock chose a 0.8%-of-box blob near the box
centre over her 27%-of-box dress at distance 0.17 — fixed by pinning the box
centre to the dress-mass centroid, with the chosen mask rendered
(scout_isabel_mask2.png). Pattern-shelf lesson: the I-4 selector
prefers proximity over size; centre loose boxes on the object's mass centroid
and always render the chosen mask. Net effect of all corrections: every
crux ratio strengthened (mass-share 10.4→15.2×, saliency-share
25.7→50.2×), the mirror fell to dead-last on saliency (9/9), and its mass rank
moved 7/9→6/9 — proof the ranks are box-set-dependent, as already scoped; the
shares carry the claim.The marginalization is by extent, not intensity: the reflection is painted with real substance (mass-mean 0.54, edge 0.36 — not a smudge) and allotted 0.61% of the frame. Same signature as the Icarus, and the comparison now carries its numbers (review: it had been asserted without them):
| Icarus (Study 7) | the royal mirror (this study) | |
|---|---|---|
| mass, share of frame | 0.18% | 0.61% |
| subject ÷ biggest actor, mass-share | 25.5× (ploughman) | 15.2× (Infanta) |
| saliency, share of frame | 0.22% | 0.33% |
| subject ÷ biggest actor, saliency-share | 21.6× | 50.2× |
| rank among named objects (analyst-dependent) | 8/11 mass | 6/9 mass · 9/9 saliency |
The dark back wall holds two lit rectangles. They are opposites: the open doorway (José Nieto on the lit stair) is the densest local saliency peak in the back wall — local mean 0.556 — while the mirror is a saliency void at 0.055. The doorway out-pulls the royal mirror 10.1× on saliency, 1.53× on mass [confirmatory, prediction 8; v3 boxes].
And the pull is contrast, not brightness (review 4's sharpening). The Ortega probe puts the two lit planes within 18% of each other in luminance (foreground frieze 0.232, through-door light 0.274) — yet the doorway is the denser local saliency peak by a wide margin. Its optical pull is a local-contrast effect against the dark back wall, not a brightness effect. Which sharpens the study's headline convergence: Snyder & Cohen's vanishing point coincides not with the brightest thing in the picture but with a local-contrast peak — a more specific agreement between the perspectival and optical organizations than brightness alone would give [exploratory sharpening, from measured values].
State mixed_field / offcenter_mass (not dissolution); the island anchor is the Infanta, robust at object identity across both edge modes (not at location: her body edge-aware, her skirt-base raw); entropy 0.79–0.81. Mass pulls hard down (δy +0.214; bottom third 50% of mass vs top 18%) and sits dead-centre horizontally (δx +0.039 — the pre-registered "slightly left" was wrong; scored half-right). μ 0.561: the bright figure-frieze is one connected structure. Colour is a servant of the drawing: β 0.049, coupling +0.361 (corpus high). DGI 18.8, "authored against gravity" — the most extreme off-basin composition measured.
Tone: shadow_mass 0.897 · τ 0.750 · highlight 0.001 — the Infanta's dress, the brightest object in the picture, holds one part in a thousand of the tonal mass. "Darkest in the corpus" is REFUTED by the same-pipeline compare (Entry 11): the Caravaggio reads 0.903 against the Meninas's 0.896 — a 0.007 gap, a statistical tie at the corpus's shadow pole (and the first-pass "Caravaggio ~0.35" was a misquote: that is the Bruegel's number). Against the measured ~0.13 chain delta, the ordering between the two Spanish tenebrists is undecidable at current chain control.
0.15–0.26, subadditive, cut-stable — every half and quadrant carries more torque than the whole; the frieze's local leans cancel into a still composition. Joins Caravaggio, Pollock, Bruegel and the HCB in the cancellation camp [confirmatory, prediction 7 — weak band: 0.1–0.6 spans most of the subadditive range]. Appended to the register.
Central frontality 0.425, axis locked dead-centre (offset 0.004, lock 0.991) — but mid-pack: below the Klimt (0.663) and the Bruegel (0.537). The pre-registered thesis that Las Meninas might be the corpus's figural-address maximum is refuted [confirmatory-refuted, prediction 6; the 0.4–0.65 band was also weak — it covers two-thirds of the corpus spread].
Figural vs spectral. Address is scale-robust here (0.425 @2600 vs 0.431 @1300, a 1.4 percent shift — the stated justification for sweeping at 1300px). The informative null is the phase-scramble (it preserves the power spectrum, so it tests whether the room's banding could produce the signal; patch-shuffle destroys the spectrum and only asks "any structure at all"). At n=100, 0 of 100 draws reached the real value — honest interim p < 0.01 with the point estimate tail-model-dependent (a normal tail gives ≈0.003, but sd/mean = 0.97 is the exponential signature and an exponential tail gives ≈0.027, a 10× spread the n=100 summary cannot discriminate; "p = 0" and "p ≈ 0.003" are both retired — the first as a resolution floor, the second as an extrapolation). Settled empirically at n=1000 (phase-only, quantiles): null median 0.109, q99 0.398, q99.9 0.430, max 0.481 — 1 of 1000 draws ≥ the real 0.431: p ≈ 0.002 by the conventional (r+1)/(n+1) estimator (r/n = 0.001 is the biased form — the p-=-0 logic at its own floor; review 3), 95 percent CI on the exceedance ≤ ~0.006. Note the real value is no longer above the observed null max (0.481 exceeds it): the verdict rests on the empirical percentile, never on clearing a maximum. FIGURAL stands — under a decision rule recorded honestly as adopted AFTER this sweep landed (binding forward, no retroactive authority): p < 0.01 verdict as written; 0.01–0.05 marginal, no "settled"; ≥ 0.05 withdrawn.
null_figure_shuffle.py).
The phase-scramble tests "could the room's spectrum reproduce this?" — it cannot.
The matched null asks the sharper question phase-scramble cannot — is this
arrangement special? — by keeping every figure intact (own paint, size,
height) and permuting only the seven figures' x-positions, room/mirror/
doorway fixed. Reference is the synth-real (identity permutation through the same
paste pipeline): address 0.419, a 2.8 percent synthesis artifact against the true
0.431 — well-powered. Read as a measurement: Velázquez's arrangement scores at
the 73.5th percentile of rearrangements of his own cast — real 0.419 vs null mean
0.348, sd 0.104, an effect of +0.7σ (run A alone). Pooled with the independent
replication (run C, fresh seed, n=300) the estimate is 148/500 ≥ real, the
70.4th percentile, 95% CI [66.4, 74.4]; run A's figure is the upper edge of
that interval and is retired as the quoted number. That is the number, precisely placed.
Its p is a footnote: 53/200 ≥ real, raw p 0.265 (95% CI [0.204, 0.326]); and
more draws cannot move it — the limiting p is ~0.27, so n=1000 returns
~0.27, not zero. Against the pre-registered rule the figural verdict does
not stand, but the informative content is the percentile: a composition sitting
maximally off the default-gravity basin (DGI 18.8) is only +0.7σ better than
shuffling its own cast on central address.What the percentile decomposes into
(shuffle_variance.py). Of the address variance the shuffle
produces, only 36% (η²) is explained by which figure is central; 64% is
the other six positions. Within that 36% the Infanta is genuinely the
best central figure — mean 0.494 when she lands central, vs 0.270–0.365
for every other figure. And the conditional read is the sharpest number in
this entry: given the Infanta central — the real configuration — random
placements of the other six average 0.494 (n=40 of the N=300 shuffles land
her central, ≈ the 300/7 = 43 expected; sd 0.054), while Velázquez's
actual placement scores 0.419: −1.38σ within its own group (run B), the 7.5th
percentile. Only 3 of 40 random arrangements sharing his central figure
score lower. So the +0.7σ decomposes exactly: all of it comes from the
Infanta being central (the strongest bilateral body, in the middle), and the
other six placements give roughly half of it back — they score below
chance, not merely "not maximizing." This is the address-channel version of
Entry 03's accent finding (distributed punctuation, 0.83→0.49, no climax),
arriving independently through a different lens: two channels, same read —
the frieze is arranged to dissipate central address, not concentrate it
[exploratory: post-hoc, n=40 conditional group; the Entry-03
convergence is interpretation, the −1.38σ is measured].
The two nulls bracket; neither answers the deepest question. The
matched null holds central occupancy fixed — someone is always central —
so it tests which figure is central, not whether central occupancy
is itself remarkable. Read together the nulls are bounds: phase-scramble too
permissive (destroys the figures), figure-shuffle too conservative in exactly
this way (it can never vacate the centre). Between them the residue — central
bilateral figural mass exists and is real content — survives because neither
null can test it. Robustness-checked (null_figure_shuffle_robust.py):
not a seed or fill artifact — across 3 seeds × 3 background fills (per-row median,
heavy-blur bleed, flat dark-void) the address p stays 0.21–0.40, even where the
dark-fill drops the synth-real reference to 0.384.
The review caught the study quoting mass_divergence as 0.080 in one
row and 0.0746 in another; a second review caught the first fix mis-attributing
the gap, so the missing cell of the 2×2 was run (review_checks.json):
| mass_divergence | no-PM | PM |
|---|---|---|
| 2600px | 0.0991 | 0.0801 (the spine's 0.080) |
| 1300px | 0.0746 (the null run's real) | 0.0682 |
Scale effect −14.9% within the PM pipeline (−24.7% within no-PM — the first fix quoted the no-PM column as if it explained the spine-to-null gap). The null run is the no-PM @1300 cell: it differs from the spine in two ways whose effects partially offset (net −6.9%), and its label now says so. The conclusions survive at the honest magnitudes: mass_divergence is scale-sensitive (−15% under the spine's own pipeline) where address is not (1.4%); the null test remains internally valid (real and nulls share one pipeline and scale: no-PM @1300); the two numbers are never interchangeable, and the study quotes 0.080 (@2600, PM) as the finding. That null test: real 0.0746 vs phase-scramble 0.0086 and patch-shuffle 0.029, no draw reaching the real value (p < 0.01 at n=100) — the eye-vs-weight displacement is compositional. DGI's null was then run (Entry 11), on exactly the review's rationale that extremes are where metrics break — and its own 2×2 was completed to bridge the 18.0-vs-18.8 pair the first draft left unexplained: DGI reads 28.2 (no-PM) / 18.8 (PM, the spine) @2600, and 18.0 (no-PM, the null run's cell) / 14.9 (PM) @1300 — the null-run value matches the spine only by two offsetting pipeline effects, the same accident as mass_divergence, and the two are never interchangeable.
A deep source pass (meninas_sources.md — graded A–D per
claim, most locators Grade C and unverified; the transmission layer itself
audited and found lossy) lands on the study in both directions.
The headline convergence. The established geometric result contra Foucault (Snyder & Cohen 1980; Steinberg 1981) is that the vanishing point is not central but at the doorway, near Nieto's elbow. Entry 05 found the doorway is the densest local saliency peak in the back wall (0.556; it out-pulls the mirror 10.35× on clean jitter draws, Entry 12 — the single global argmax is the Infanta's dress, so the convergence rests on the doorway being a dominant local peak coincident with the vanishing point, not the global maximum) — a completely independent channel, luminance-contrast attention rather than ruler-work on orthogonals, landing on the same object. The picture's optical organization agrees with its perspectival organization, and the mirror-centred reading loses on a second, unrelated axis. Three more: Steinberg's shifting focal centre and López-Rey's three foci ↔ the distributed punctuation (0.83→0.49, no climax) and moderate-magnitude address; Clark 1960's tone thesis ("the tonal relations are true... and it holds") ↔ β 0.049 / coupling +0.361; White 1969's lower-half observation ↔ δy +0.214. And Garrido 1992's reflectography (no underdrawing; painted direct) is the one technical datum behind DGI 18.8: there was no plan — he found it in paint.
The convergences, condition-split (review 4 — the study's own WTDNP applied to its allies). Clark wrote in 1960 and Ortega in 1953; Brealey's 1984 cleaning removed a yellow veil accumulating since the 19th century — both men described a picture wearing ~150 years of dirt, and this study measures the post-cleaning surface. So the split is: Snyder & Cohen and Steinberg are varnish-invariant (perspective geometry survives any cleaning — the headline convergence is safe from the condition record entirely, a point in its favour); Clark and Ortega are cross-condition-state comparisons — both probably survive (a roughly uniform veil scales tone rather than restructuring it, and lit-dark-lit is robust to scaling), but "probably survives" is now said rather than assumed. And Clark carries a reflexive sting: his thesis is that the picture holds though the colours are drab — he saw drab partly through a yellow veil, this study measures drab partly through faded azurite. Two condition states, agreeing. Might be the painting; might be two states meeting in the middle [condition-split: exploratory].
Two of the literature's decidable questions, run today — with their
scope stated (review 4). Miller 1998 claimed the mirror betrays
itself as a mirror by borrowed light. Measured (miller_probe.json):
the mirror interior reads 2.3–2.7× the two Mazo canvases on the same wall
and 2.0× the bare wall, while still 0.66× the doorway. This confirms Miller's
premise (the mirror is anomalously bright for that wall) and leaves his
inference untouched: the Mazo ratio is content-confounded (their
subjects are dark — any painted rectangle depicting a bright scene would beat
them), and 2.0× bare wall is a bar ordinary paint clears easily (the Infanta
clears it by far more). Miller's actual argument — that the brightness can only
be borrowed, a consistency claim about the room's light — needs an
illumination model, not a ratio. What is genuinely this study's: locally
anomalous AND globally weak, both measured
[premise confirmed; inference not tested].
Attribution caveat (review 6): the two back-wall canvases are the
Miller-probe controls, so their identity is load-bearing — and it is contested.
Kahr reads them as Rubens-derived Minerva punishing Arachne + Apollo
and Marsyas; the National Trust catalogue (Laing) has Minerva and
Arachne + The Judgement of Midas after Jordaens (Mazo's copies in the
Prado, inv. 1551/1712). The probe needs only two dark painted rectangles on that
wall, which both readings grant — but "Mazo canvases" is one identification of
two, flagged. Ortega y Gasset 1953 described a three-part light structure: lit
foreground, darkened intermediate zone of silhouettes, lit background.
Measured (ortega_probe.json): foreground frieze 0.232 /
intermediate 0.074 / through-door light 0.274 — lit-dark-lit holds, the
middle at a third of either lit plane. Prediction 9's hypothesis, pre-registered
73 years early, confirmed. The mirror-vs-door attention conflict is scoped, not
attacked — the received claim could not be attributed to a named scholar and no
eye-tracking study of this painting exists, so the doorway-over-mirror result
(10.35× on clean jitter draws) stands as the
bottom-up model against an unattributed sentence, with the deciding experiment
named.
clark_grid.py, clark_grid_sweep.py). Clark 1960
reads the picture on a grid of quarters and sevenths; the canvas is cut down,
more on the right (López-Rey II:306). Test: do the picture's strong verticals
fit a quarters/sevenths grid better than spacing-matched random peaks, and does
the best fit select a trim? The default configuration said yes, p = 0.021 — but
that was one peak-detection knob. Swept across 18 reasonable thresholds the
fit-p ranges 0.022 → 0.906 (median 0.194); only 4 of 18 are marginal. So
the significance claim is withdrawn — the grid does not reliably beat
chance; it is threshold-dependent, mostly no-signal. What is knob-robust
is the trim direction: every reasonable threshold puts the best
restoration heavier on the right (median right−left +10%), agreeing with
the documented "cut more on the right." (This is knob-robustness, not
replication — the 18 configs are settings on one image, heavily correlated: n=1,
robust to the knob, not 18 confirmations.) And the seam-panel widths
(0.259/0.377/0.364, left ~30% narrower) are not "weak-significance" either —
they carry no p-value at all: they are an inference conditional on an
untested assumption (uniform loom widths; Velázquez may simply have used
a narrower left strip). So the trim evidence is one unfalsified assumption
(the seam widths) plus one direction-only corroboration (Clark's grid) —
both right-heavy, neither significance-grade [exploratory;
Clark Grade C; fit-significance withdrawn, direction retained; seam widths
assumption-conditional].What the instrument cannot touch, said plainly: Foucault's absent centre is structural, not photometric — 0.61% of frame neither confirms nor refutes it; the mirror-reflects-canvas-or-couple debate is geometry and optics outside these metrics; gaze stays beyond the C-4 ceiling (the literature cannot even agree whether three or five figures look out), and the column sweep's lobes-on-brightness result is evidence that address ≠ gaze — a methodological finding about C-4 itself. Alpers 1983 (meaning in the representation, not the plot) is the nearest thing the field has offered to this instrument's manifesto; the study lands where she pointed.
Persistence (full read, PM pipeline): 0.558, fragile_to = low_contrast — compress the tonal range and the structure dies [stable]. That completes a three-way corpus split: Bruegel fragile to blur (fine edges), Klimt to grayscale (colour), Meninas to contrast (tone) — a perfect 1:1 across three works and three perturbations, which is 1-in-6 by chance and a post-hoc pattern [exploratory].
The DGI null (n=100 per family, fields-only, no-PM @1300 for both real and nulls — the same cell throughout; the spine's 18.8 is PM @2600, bridged by the 2×2 in Entry 09): real 18.0 vs phase-scramble [30.2–62.7, median 46] and patch-shuffle [47.4–65.8, median 58]. Direction, stated: lower DGI = further from the centred default-gravity basin, so the real value sitting below every null draw is the confirming side — each re-arrangement scores MORE default-gravity than the painting [stable within its cell]. And a scope the cleanliness demands (review 4): a null with zero overlap in 200 draws may be answering "does this image have any composition?" rather than "is this composition non-default?" — phase-scrambling destroys all arrangement, so any structured image may beat any scramble. The address null was informative precisely because it was close; this one may be free. The matched test is the figure-position shuffle (next).
Compare vs the poles (n=5: + Caravaggio, Bruegel, Klimt, Degas), read as raw values and ranks — z's saturate at this n and carry no magnitude (review 4): the decisive row is tonal — Caravaggio 0.903 vs Meninas 0.896 shadow_mass, a tie at the corpus's shadow pole that refutes prediction 5's superlative (Entry 06). Elsewhere: coupling_index 0.301, rank 1/5 (Caravaggio 0.195, Bruegel 0.192, the moderns ≈0 — colour-describes-form beyond even the other tenebrist); β 0.060, lowest with the Caravaggio's 0.061 (the moderns 0.23–0.38): the two Spanish works pair on colour-tone behaviour across the board; Klimt holds the highlight pole (0.372 vs everyone else ≤0.076), Degas the torque pole. Placement only, n=5 [placement]. The mass_divergence null's cheap-cell caveat stands: it proves non-nullness of the no-PM @1300 cell, transferred to the claim cell across a measured 6.9% gap with an 8.7× margin.
The review-5 worry, stated fairly: four placement errors were found
across v1→v3 and every correction strengthened the headline — unanimity
that either means the errors were all conservative, or the corrector was
steering. Hand-placement cannot settle that; a jitter analysis can
(jitter_probes.py): all nine boxes perturbed independently in
position and extent, uniform ±5% of each box's own dimensions (covering
the requested ±3–5% band), n=200 draws, the exact I-4 selector rerun
per draw on once-computed PM fields. And — the load-bearing addition after
review 5 — each draw is tagged clean or failed: the selector reports
whether the component it chose is the box's largest qualifying one, so
the distribution can be split by whether the I-4 lock held. You can eyeball
nine masks; you cannot eyeball eighteen hundred.
| crux ratio | v3 point | ALL draws: med [q05,q95], floor | CLEAN draws: med [q05,q95], floor |
|---|---|---|---|
| Infanta ÷ mirror, mass-share | 15.2× | 16.1× [14.0, 95.7], 13.6× | 15.5× [14.0, 20.6], 13.6× |
| Infanta ÷ mirror, saliency-share | 50.2× | 53.2× [45.1, 174.9], 43.2× | 51.0× [44.6, 66.0], 43.2× |
| Infanta ÷ mirror, saliency local | 5.01× | 5.02× [2.56, 5.47], 2.49× | 5.11× [4.75, 5.50], 4.64× |
| Infanta ÷ mirror, mass local | 1.54× | 1.54× [1.41, 1.62], 1.37× | 1.55× [1.48, 1.63], 1.43× |
| doorway ÷ mirror, saliency local | 10.1× | 10.21× [5.21, 11.03], 5.11× | 10.35× [9.72, 11.05], 9.53× |
| doorway ÷ mirror, mass local | 1.53× | 1.53× [1.39, 1.62], 1.36× | 1.55× [1.48, 1.62], 1.44× |
Of 200 valid draws, 155 are clean. This run fixes a harness bug AND adopts the largest-component convention for the doorway; the convention × jitter 2×2 (finding 4) shows the bug was worth ~2% and the convention ~167%, so the doorway value here is a convention result, not a bug fix.
Four findings.
(1) The direction never flips — in any of 200 draws, on any crux channel. The Infanta carries at least 13.6× the mirror's mass-share and at least 43× its saliency-share under any box set within ±5% of the pinned one; the clean-draw floors hold at 13.6× / 43.2×. And the doorway out-densities the mirror in every single draw (P = 1.0), floor 5.11× all-draws, 9.53× clean. Box placement is retired as a threat to the finding.
(2) The fat 'all-draws' mass-share tail is a mirror component-flip, not placement — review 5's mechanism, confirmed. ±5% on a box moves an area-based share ~10%, so the ratio should move 15.2 → ~16.8; a q95 of 95.7× requires a box to lose ~84% of its content, which geometry at ±5% cannot do. What can is the selector taking a small off-object blob. Split by lock status the tail collapses: mass-share q95 falls 95.7× → 20.6×. Because a mirror flip makes the mirror catch less, these inflate the ratio and live in the upper tail — so the floor (13.6×) is clean and the headline is safe.
(3) The real diagnostic is the decision margin, not a jitter failure-rate
(lock_margins.py). A jitter flip-rate is a noisy proxy; the
exposure is a static property of each box — how close the runner-up component
sits to the winner. Measured for all nine: doorway margin 0.0015 (a
one-pixel shift decides the read), mirror 0.105, Infanta a single
qualifying component — no decision exists at all, the rest 0.067–0.33. Two
things follow that a flip-rate hid: the Infanta headline never had an
exposure (nothing to flip), and a mirror flip would read the runner-up at
0.107 saliency vs the winner's 0.055 — it would weaken the study's ratio, not
flatter it. The only knife-edge pointing the study's way is the doorway, and
it is handled (finding 4).
(4) The doorway: a harness bug AND a convention change, conflated in one
pass — the 2×2 prices them apart (convention_jitter.py). Two
things moved between last round's 3.75× and this round's 10.2×, and the honest
accounting separates them. (a) A harness bug: the jitter loop computed
the box window with int(round(x·w)) while zone_read
truncates — a one-pixel mismatch, decisive at the doorway's 0.0015 margin.
(b) A convention change: the doorway's selection rule was switched from
nearest-to-centre to largest-component, on the annular-target ground (the
lit aperture surrounds the box centre — Nieto's dark body sits in the middle —
so nearest aims into the hole where a 0.6%-of-box highlight sits; verified in
scout_doorway_components.png). Priced by the 2×2:
fixed-harness + nearest = 3.83× (bug worth ~2%); fixed-harness + largest =
10.23× (convention worth ~167%). So — correcting last round, which credited
the bug — the 3.75× was not a poisoned statistic; it was the correct
doorway number under nearest-to-centre. The 10.2× is a convention result
on an annular target, justified on structural grounds stated before the number,
its worth measured. This is the study's recurring error class (PM × scale;
chain × pipeline; Caravaggio × Bruegel: two things move, effect credited to the
wrong one) — caught here on myself. The convention removes the coin-flip (0
flips by construction under largest); the value is stable under jitter
(10.21× all-draws, 10.35× clean), and every fat-tail mass-share
outlier is a mirror lock-flip (15.4× locked, 92.7× flipped), never
Infanta or doorway.
The maximum claim, settled box-free. Entry 05 had called the doorway "the saliency maximum, the strongest pull of any object measured" — silently switching between a box-free per-pixel argmax and a ranking over nine hand-boxes. Measured with no box: the global saliency argmax is at (0.537, 0.772), inside the Infanta's dress — not the doorway. The picture's two attention poles are the living Infanta (the argmax) and the doorway (the largest single concentration of top-decile saliency mass, ~54%, though its box overlaps the Infanta's by ~10%); the royals are neither. "The saliency maximum is at the doorway" is withdrawn; the doorway is the densest local peak, which is what the Snyder & Cohen convergence needs.
Asymmetric robustness, measured — and it splits by error scale. The review's gift (mirror saliency-share invariant across the v1→v3 re-pin, 0.336‰ → 0.336‰, ratio 0.998, while the Infanta's shares doubled) is real. Fine jitter shows the large lit figures barely move (frame-share cv ≈ 0.03: Infanta, Sarmiento, Isabel). Both are true: the mirror is invariant under gross error (a body-width miss) because nothing bright sits near it to catch by mistake — the buried thing is stably buried — while the small apertures are the ones the selector abandons under fine jitter (finding 3). The crux survives at both scales (gross: the re-pin moved the ratio 10.4→15.2×, same direction; fine, clean: floor 13.8×).
scout_all_masks.png), and the overlay is generated from the
PROBES list rather than a hand-copy — so a reader who has never met me can check
every box against the pixels. That is the answer.The uniform-rule question, closed. Isabel's mask-lock fix had changed
a rule for one object; the review asked whether it was applied to the other
eight. All nine locks were then verified, not asserted
(mask_verification.json; masks rendered in
scout_all_masks.png): 8/9 chose their box's dominant mass
component (7/9 natively — Isabel passes only because her box was treated).
The one remaining failure is Velázquez (chosen component 9.1% of box vs
largest 18.0%). Applying Isabel's treatment uniformly — re-centring his box on
the dominant component's centroid — made him worse, not better
(saliency mean 0.137 → 0.038): his box's dominant mass is wall structure,
not the painter (dark-on-dark, I-16, with the x≈0.259 seam through the
box). So the rule is stated exactly: mass-centroid pinning applies where the
dominant component is the object (Isabel, zoom-verified); where no
placement yields an object-locked dominant component, the object is flagged
unmeasurable-reliable and both readings reported (Velázquez: v3 box
0.387 mass-mean / 0.47% share / 0.137 sal-mean vs centroid-pinned 0.297 /
0.68% / 0.038). He appears in no crux ratio; his rank rows carry the flag
[confirmatory hardening, review 5].
selector_sweep.py) compared the three
at the pinned boxes and found them identical (15.3–15.7× / 49.9–52.2× /
9.9–10.1×) — but at the pin, nearest and largest select the same
components (this entry's own "changes no number at v3"), so the static sweep
restates a known identity and cannot see the doorway's knife-edge, which is
expressed only off the pin. The review caught the claim crediting the wrong
instrument. The missing cells — convention × jitter, the same 2×2 shape as
scale × PM — were then run (convention_jitter.py: the full n=200
jitter under each convention applied uniformly, same seed, identical box sets
across conventions):| convention × jitter (n=200) | Infanta ÷ mirror mass-share: med, floor | saliency-share: med, floor | doorway ÷ mirror sal: med, floor | door flip |
|---|---|---|---|---|
| nearest-to-centre | 16.07×, 13.56× | 53.2×, 43.2× | 3.83×, 1.75× | 52% |
| centroid-pinned | 16.94×, 13.59× | 55.4×, 38.5× | 9.81×, 4.80× | 0% |
| largest-component | 16.07×, 13.56× | 53.2×, 43.2× | 10.23×, 9.53× | 0% |
Two findings, cleanly split. (a) The headline is convention-invariant on AND off the pin: Infanta ÷ mirror floors 13.56/13.59/13.56× across all three conventions under full jitter (saliency-share floor dips to 38.5× under centroid — still ≥38×). No hand-made selection choice moves the burial result, and that claim now rests on the jittered cells, not the static identity. (b) The doorway is the one convention-sensitive number, and only off the pin: under all-nearest it coin-flips exactly as the 0.0015 margin predicted (52% flip, median 3.83×); under largest it is tight (10.23×, floor 9.53×). The honest statement of the doorway choice: the annular-target argument justifies largest-component on structural grounds; the convention × jitter cells measure what the choice is worth (3.83× → 10.23× median); the static sweep never could and is no longer cited for it. Convention is retired for the headline by measurement, and pinned open-eyed for the doorway by stated principle plus its measured effect.
| # | prediction | result |
|---|---|---|
| 1 | royals mass/saliency subordinate, order of magnitude | Confirmed, strong (a real bet; ratios strengthened by the v3 box audit, jitter-hardened in Entry 12): 15.2× / 50.2× on the shares, floors 13.5× / 43× across 200 jitter draws; royals dead-last on saliency |
| 2 | mass centroid down, slightly left | Half-right, now trim-scoped: down ✓ (δy +0.214); the ✗ on "left" may be the 18th-century trim's artifact — the canvas was cut more on the right (Entry 10) |
| 3 | eye & weight roughly coincide, both ignore the mirror | ✓ divergence 0.080 (@2600 PM; the 0.0746 null value is the same quantity @1300 — Entry 09); both low-centre |
| 4 | hierarchy/counterweighted, NOT dissolution | ✓ mixed_field, anchor = Infanta, μ 0.561 |
| 5 | deepest shadow-key in corpus | REFUTED as 'deepest' by the same-pipeline compare: Caravaggio 0.903 vs Meninas 0.896, a tie at the shadow pole (the cited ~0.35 was Bruegel's number); the 2×2 chain work stands — the 0.13 delta swamps the 0.007 gap, ordering undecidable |
| 6 | address 0.4–0.65, figural, may rival Klimt | Number ✓ (weak band); measured, not a verdict; magnitude ✗. Beats the phase null (p ≈ 0.002 — real content); against the matched shuffle it is a 70.4th-percentile arrangement of its own cast, CI [66.4, 74.4] (pooled 148/500 over two independent seeds — the figural verdict does not stand but the percentile is the finding, Entry 08). η²: who-is-central explains 36%; conditional on the Infanta central, the other six placements score −1.32σ below random (pooled) — the frieze dissipates central address [exploratory]. Not the corpus max |
| 7 | R_spatial subadditive 0.1–0.6 | Confirmed, weak band: 0.15–0.26 |
| 8 | apertures fire; doorway beats mirror | Confirmed, jitter-hardened (Entry 12): doorway ÷ mirror saliency 10.21× [5.21, 11.03] all draws, 10.35× [9.72, 11.05] clean, and the doorway out-densities the mirror in every one of 200 draws. The v3 10.1× restored after a harness bug and an annular-target knife-edge were found and fixed (last round's "~4×" was the bug, retracted) |
| 9 | depth: doorway aperture caught, "depth" not (C-1) | Confirmed in its Ortega form (Entries 10–11): lit-dark-lit holds (0.232/0.074/0.274); the aperture is the densest local saliency peak (the global argmax is the Infanta's dress, Entry 12); "depth" as a bound percept stays out of reach (C-1) |
| 10 | nulls: placement compositional, address figural | Split, honest form. mass_divergence composition-dependent (p < 0.01, n=100) — holds. Against the matched shuffle: address is a measured +0.7σ / 74th-percentile arrangement (raw p 0.265, limiting p ~0.27 — figural verdict does not stand, the finding is the percentile, Entry 08); DGI is NOT tested — the 37% synthesis artifact swamps the 11% effect, a false-transfer refused, so no verdict (Entry 11). Composition-dependence: confirmed for mass_divergence, measured-but-modest for address, untestable at this fidelity for DGI |
lock_margins.py): the Infanta headline has no exposure (single component), a mirror flip would weaken the ratio not flatter it (runner-up 0.107 vs winner 0.055), and the doorway's knife-edge (margin 0.0015, an annular target) is handled by a call-site largest-component rule that changes no v3 number. A one-pixel harness bug was fixed, but the convention × jitter 2×2 prices it at only ~2%: the doorway's 3.75×→10.2× lift is a convention change (nearest→largest, annular-target ground), not a bug fix (finding 4). Clean-draw floors 13.6× / 43.2× / 9.53×. Jitter covers fine placement error, not gross mispinning (that is the zoom-verification protocol's job; the v1→v3 episode is what skipping it looks like).meninas_sources.md — the study's graded source list (A–D per claim; the transmission layer audited and found lossy — wrong DOIs, a corrupted title, a phantom page number, and the dimension error this header first inherited). Priority pulls: Garrido 1992 pp. 579–591; López-Rey 1999 II:306; Snyder & Cohen 1980; Clark 1960 pp. 32–40.miller_probe.json / ortega_probe.json.PREREGISTRATION.md (10 predictions, locked pre-run) · labbook.md (markdown twin) · report_pm_nopersist.json / report_nopersist.json · subject_probes.py/.json · accent_clean.py/.json · null_address.json · null_address_1000.json · null_massdiv.json · review_checks.json (the review's measurement batch: massdiv decomposition, shadow chain, column sweep, mirror cap-check, seam overlap) · jitter_probes.py/.json (n=200 box jitter, harness-bug-fixed) · lock_margins.py/.json (I-4 decision margins) · doorway_repin.py/.json + scout_doorway_components.png (the annular-target diagnosis) · null_figure_shuffle.py/.json + _robust.py (matched null + seed/fill robustness — address measured at +0.7σ, DGI not tested) + shuffle_variance.py (η² decomposition) · selector_sweep.py + convention_jitter.py (convention at pin + convention × jitter) · clark_grid.py + clark_grid_sweep.py (the trim estimator and the threshold sweep that withdrew its fit-significance) · mask_verification.json / scout_all_masks.png (all nine I-4 locks) · make_charts.py · instrument findings I-1…I-19, candidates C-3/C-4 (toolkit/CANDIDATES.md, OBSERVATIONS.md — appended).