Study 10 — Midjourney corpus, Little Red Base vs Steered (2×3×8) — lab recordImage Structure Reader · Study 10 · lab record · MJ corpus, 2×3×8

Read 2026-07-13. 48 born-digital PNG frames (Midjourney v8.1; sampling confirmed no-selection, see Entry 01), two prompts (PROVENANCE.md, verbatim), three aspect ratios (1:1, 3:4, 4:3), eight draws per cell. Pre-registration after contact-sheet inspection, before any run: PREREGISTRATION.md (8 corpus-level predictions). The corpus's first within-engine ensemble: draw variance is directly measurable. MJ studied on its own terms (user direction); other engines appear only in the register and same-coordinate placement. Per-frame table: corpus_rows.json (spine, stance, verified red-mask figure Δx, RCP cues). Corpus mode: claims live at cohort level; per-image language is exemplar-only, pixel-checked.

Labels: [stable] pipeline-stable · [narrowed] scoped · [aid] reading aid · [confirmatory] pre-registered · [exploratory] post-hoc.

More info: www.parallaxmetrology.com, full critique


Entry 01 — Selection, design, condition [stable]

Governing frame: this is a composition study, not a prompt case study. The directed prompt is the perturbation used to probe MJ's compositional field, not the object of study — so the question is never "did the engine obey directive X" but "what does the field do when pushed." That stance licenses Entry 11's honest nulls-of-method (no per-directive obedience verdict is owed, because obedience isn't the object), makes "42%→71%" (Entry 05) a statement about the field's steerability rather than a compliance score, and keeps the aspect interaction (Entry 06) central — frame shape conditioning the field is a composition finding. A 2 (cohort) × 3 (aspect) × 8 (draw) factorial from one engine — the first study where "how much is the steering, how much is the draw" is a measured quantity rather than a caveat. Sampling, user-confirmed: prompt in, image out, no selection — two 2×2 grids per size per cohort, all tiles retained; so draws are 2 jobs × 4 tiles per cell (A–D / E–H), complete but nested (job-clustering checked in Entry 09). Conditions: born-digital PNG, full-bleed, no PM; three native sizes read as factors, never pooled blindly (I-10); fine-structure claims off the table (MJ upscaler). One filename-irregularity incident in the Steered set (double underscore, stray dot) crashed the first batch pass at frame 25; the parser was made tolerant (aspect normalized, cohort from directory), no data lost — logged because silent parsing is how corpora rot. Note from provenance: the base prompt itself supplies the architecture (column, arches, tunnel), so architectural presence is not a steering deliverable in this study.

Entry 02 — What the eyes said first (contact sheets, pre-registered) [stable]

Default: graphic-illustrative; small staffage figure in profile, placement roaming; deep one-point arch tunnels; ambient daylight; green/stone with a small red accent. Scene-mode, not the centred-emblematic basin. Steered: the style itself shifts (softer, painterly); figure larger; pose contract visibly delivered in most draws (lantern, look-back, tension); white-hot core frequent — and in many draws halo-adjacent behind a centre-ish figure, despite the "no center halo" refusal. These eyes-first observations set the registered bets, including the risky one (P2) that steering might raise DGI here.

Default cohort contact sheet: rows are 1:1 / 3:4 / 4:3, columns draws A-H. Scene
Default cohort contact sheet: rows are 1:1 / 3:4 / 4:3, columns draws A-H. Scene-mode: small roaming staffage figure, deep arch tunnels, ambient light.
Steered cohort contact sheet, same layout. The style itself shifts; pose contrac
Steered cohort contact sheet, same layout. The style itself shifts; pose contract visible in most draws; white-hot core frequent and often halo-adjacent.

Entry 03 — Headline: draw noise vs steering, measured — and the steering moved style harder than structure [stable, confirmatory core]

The axis table (cohort mean ± sd over 24 draws). [Corrected at review: the first denominator pooled draws across aspects, and Entry 06 shows aspect is a real factor — so it carried ~15% aspect spread. The separation column now divides by the draw-only sd, pooled over the six cohort×aspect cells (canonical_stats.json, the auditor's source of truth); every value rose slightly, nothing flipped.]

The ± is the cohort sd (describes each cohort's spread, aspect included); the draw-sd is the actual separation denominator (pooled within the six cohort×aspect cells, so aspect spread is backed out), and sep = |Δmean| / draw-sd reconciles on every row from the printed draw-sd.

axisDefault (±sd)Steered (±sd)draw-sdsep
coupling0.088 ± 0.1330.265 ± 0.1180.1191.48
β0.155 ± 0.0650.082 ± 0.0380.0491.48
shadow_mass0.592 ± 0.1220.755 ± 0.1130.1121.45
τ0.471 ± 0.0960.575 ± 0.0930.0821.28
gap_fraction0.418 ± 0.0430.453 ± 0.0420.0390.90†
mass Δx−0.040 ± 0.069+0.030 ± 0.0940.0790.89
theta0.043 ± 0.0260.023 ± 0.0210.0230.87
μ0.462 ± 0.1830.337 ± 0.1110.1530.81
figure Δx+0.012 ± 0.124+0.088 ± 0.0880.1060.72
grid_asym0.089 ± 0.0540.123 ± 0.0710.0630.54
torque0.131 ± 0.0820.200 ± 0.1820.1350.51
imbalance0.058 ± 0.0360.075 ± 0.0350.0350.47
DGI50.3 ± 8.047.1 ± 8.78.220.40

† job-clustered axes (Entry 09): effective draw count on coupling and gap_fraction is nearer 4 per cell than 8, so these two separations carry wider uncertainty than the others. Note μ's draw-sd (0.153) slightly exceeds its cohort sd (pooled ~0.151): μ has essentially no aspect effect, so backing aspect out is a wash and the value is sampling noise at n=8.

Prediction 4 (registered: no arrangement axis separates ≥ 2× pooled sd) — hit under both denominators: the maximum separation anywhere is 1.48 on the corrected draw-sd (1.40 on the registered pooled sd), and every arrangement axis sits below 1. Prediction 5 (the style shift dominates: tonal and colour-policy axes lead) — hit: the four top separators are shadow, coupling, β, τ. The steering prompt asked for arrangement (displacement, void discipline, peripheral pull); what it moved most, at cohort level, is tonal key and colour policy — the painterly shift the contact sheets showed. Structure moved too, but inside draw noise on most axes: the DGI gap is 3.2 points against per-cell draw sds of 5.6–10.4 (~0.4 draw-sd). Prediction 2, the registered risky bet (steered DGI ≥ default) — wrong: steering lowered DGI slightly (50.3 → 47.1), and weakly. Both misses are the finding: on this engine and prompt pair, one paragraph of compositional direction re-styled the image reliably and re-structured it only marginally.

Entry 04 — The basin is conditional: this engine+prompt does not produce the emblem [stable]

Default cohort: DGI 50.3 ± 8.0 (range 27.2–63.7), state counterweighted in 21/24, stance standing in 17/24 with zero torquing draws (the C-2 cluster). RCP cues across all 48: no frame reaches ring-fit ≥ 0.5 (max 0.38); density bowls are weak (mean centre-to-rim drop 0.05 Default / 0.11 Steered; only 5/48 frames exceed 0.3) — the deep tunnels put their light at the end of the corridor, not wrapped around a centred subject; several profiles invert. Mass centre-lock passes in 17/24 Default frames but multi-signal agreement never assembles: zero Hard-RCP candidates in either cohort; the modal call is Not-RCP. Prediction 1 — hit on RCP and on "scene-mode," near-miss on the DGI band (predicted 40s, measured 50.3). Read with the program's spatial-priors framing: the centred-emblematic basin is a conditional attractor — engine- and prompt-dependent — not a universal AI default, and this base prompt (walking figure + supplied architecture) lands in a scene attractor instead: recessional, counterweighted, standing.

Entry 05 — Contract delivery is distributional, and placement is real but noisy [stable]

Per-draw counts. Figure right-of-axis (verified red-mask Δx > +0.05): 17/24 Steered vs 10/24 Default — the base rate matters: the direction moved the right-of-axis rate from 42% to 71%, not to certainty (a steerability reading, not a compliance score — the governing frame). Cohort means: figure Δx +0.088 vs +0.012. Draw scatter: Default sd 0.124 vs Steered 0.088 — prediction 3 hit: the staffage roams, the directed pose-and-place is tighter. Torque: 5/24 Steered draws cross the torquing threshold (0/24 Default) — the "torso torqued" item surfaces at field level in about a fifth of draws. The white-hot core (τ/shadow shift) is the most reliably delivered item of all — and it is a style item. Not audited per-frame this tranche: particle ban, cracked stair, moss asymmetry (queued; needs masks).

Entry 06 — The aspect factor: real, and it interacts [stable, exploratory]

Prediction 6 (basin strongest at 1:1) — wrong twice. DGI by aspect: Default 4:3 highest (54.2), 1:1 middle (50.6), 3:4 lowest (46.1); Steered 1:1 highest (49.5), 3:4 lowest (44.7). The cohorts reorder under aspect — an interaction, not a main effect. And the figure is most displaced in Steered 1:1 (+0.154 mean), not 4:3 (+0.057): the square, with the least lateral room, took the displacement direction hardest. Draw noise is also aspect-dependent: 1:1 is the noisiest cell in both cohorts (DGI sd ~10 vs ~6–9). At n=8 per cell these are directions, not laws — quoted as the first measured evidence that frame shape conditions both the attractor and the steerability, a factor the corpus had never varied.

The aspect factor: Default Gravity (left) and figure Δx (right) by cohort×aspect
The aspect factor: Default Gravity (left) and figure Δx (right) by cohort×aspect. The cohorts reorder across aspect (an interaction), and the square displaces the figure hardest.

Entry 07 — The colour policy is not engine-constant: it moved with the style [narrowed]

Within-cohort colour-policy spreads are modest (cbi sd 0.014–0.021, β sd 0.038–0.065) but the cohorts differ hard: coupling 0.088 → 0.265, β 0.155 → 0.082. On this engine, the colour-light relationship is itself steerable — it tracked the painterly shift rather than persisting as an invariant house style. Registered as an open question (P7); answered: not invariant here. Scoped strictly within-MJ: this says nothing about any other engine's policy, and the style shift was requested (the light directives), so "steerable via style language" is the honest phrasing. cbi, by contrast, is flat everywhere (0.027 vs 0.025): neither cohort draws colour boundaries light doesn't draw.

Entry 08 — Exemplar pixel-checks [stable]

Corpus discipline: two exemplars verified against pixels (exemplar_Default_C4x3.png, exemplar_Steered_A1x1.png). The Steered 1:1 exemplar shows the delivered pose (lantern, look-back), anchor adjacent to the figure mass, three resistant corridors (r = 0.63–0.66), entropy 0.94; islands map onto nameable structure (foliage arch, stair treads, figure). Readouts sane; no anomaly requiring escalation. Contact sheets (contact_Default.png, contact_Steered.png) are the corpus's visual record.

Default exemplar (C, 4:3) island readout, pixel-checked.
Default exemplar (C, 4:3) island readout, pixel-checked.
Steered exemplar (A, 1:1) island readout: delivered pose, anchor beside the figu
Steered exemplar (A, 1:1) island readout: delivered pose, anchor beside the figure, three resistant corridors (r = 0.63-0.66).

Entry 09 — The nesting, checked: tiles are near-independent on arrangement, job-clustered on style [stable]

Draws are 2 jobs × 4 tiles per cell, so the eight are not automatically independent — the four tiles of one grid could share compositional DNA. Measured (one-way random-effects ICC by job, pooled over the six cells; canonical_stats.json): ICC(job) ≈ 0 on all axes except two — τ, shadow, β, theta, torque, highlight, cbi all 0.00 (effective n = 8); DGI 0.07, figure Δx 0.06, mass Δx 0.12 (effective n ≈ 6–7). The two clustered axes are both colour/facture: coupling ICC 0.36 and gap_fraction 0.25 — tiles within a grid share a colour-light look — putting their effective n nearer 4 per cell (design effect ≈ 2). (These ICCs are themselves estimated from only 12 job-clusters, 2 per cell × 6 cells, so 0.36 carries real uncertainty and "design effect ≈ 2" is approximate — enough to flag the two axes, not to price them precisely.) Consequence, stated where it bites: the arrangement conclusions (steering inside draw noise) keep full precision; the strongest style separation (coupling, 1.48) is real but less precisely estimated than its row suggests, flagged † in Entry 03. The hidden assumption is now a stated, checked one.

Entry 10 — Cross-study marks, recorded here (the earlier studies stay as time capsules) [narrowed]

Per user direction the Sora study is not edited; the corrections of record live here, in the usual cross-corpus place. (a) The basin is engine-and-prompt-conditional (Entry 04): a second engine on a sibling prompt produced scene-mode, not the emblem — so any earlier "a generative engine produces the basin" framing reads, in hindsight, as scoped to its engine and prompt. (b) The colour-light policy is not an invariant of generative engines as a class: the axis that held fixed in the corpus's other AI study is the axis that moved most here (coupling 0.088 → 0.265, β 0.155 → 0.082) — noting the style shift was requested in this steering prompt. (c) C-2 ("AI clusters at standing") now has split evidence: supported by this Default cohort (17/24 standing, 0 torquing) and broken by the other engine's default (torque 0.572) — stance-at-default reads as engine-conditioned, like the basin and the colour policy. All three marks carry to OBSERVATIONS at study close.

Entry 11 — Contract masks: the reliably-delivered item is the light, the arrangement items are weak or unmeasurable [narrowed]

Per-frame masks for the objectively-checkable steering items, cohort medians over 24 draws (contract_masks.json; distributional, not per-draw verdicts).

The pattern holds and sharpens Entry 05: the item delivered with least ambiguity is the style/light one; the spatial-discipline items (particle ban, breathing lane) are the ones the instrument either can't cleanly see or finds only weakly. Two of four contract masks are honest nulls-of-method, logged as such.

Entry 12 — Candidates as cohort ranges: both cohorts sit in the distributed-order camp, neither is the emblem [stable]

C-3 (R_spatial) and C-4 (address) run on all 48 as the register's first cohort-range rows — a corpus form (per-draw range + verdict distribution), not a single value (candidates_cohort.json; candidates only this pass, nulls owed on exemplars).

R_spatial. Default: 20/24 subadditive, 3 additive, 1 superadditive; median r_min 0.21. Steered: 13/24 additive-cut-sensitive, 9 subadditive, 2 superadditive; median r_min 0.32. Both cohorts live in the subadditive-to-additive band — the distributed-order/cancellation camp the corpus knows from its real deep-space works — and neither cohort produces the superadditive emblem extreme. The steering nudged the distribution toward additive/cut-sensitive (more draws where the whole neither cleanly cancels nor emerges), consistent with a busier, less singly-resolved field; but at cohort level this is a shift in a distribution, not a category change.

Address (raw, no null this pass). Default central frontality median 0.10 (range −0.28 to 0.40), Steered 0.16 (−0.10 to 0.47). Both low. Per the corpus discipline (the Bruegel lesson: raw central frontality without a null can be spectral, not figural), no figural-address claim is made — these are raw values, and the null_sweep that would separate figural from spectral is owed (queued on 2–4 exemplars, not the full 48, for compute). One raw curiosity held at [aid]: the Steered symmetry axis sits more centred (|offset| median 0.07 vs 0.20), which would be worth a null if it survives to exemplars.

Entry 13 — Synthesis [reading]

The MJ corpus measures what a paragraph of compositional direction does to a generative engine when the sampling is complete and unselected. Three things, in order of how well they held. First, the reliably-moved layer is style, not structure: the four largest cohort separations are tonal-and-colour (shadow, coupling, β, τ), the one contract item delivered without ambiguity is the white-hot core, and and the arrangement axes sit at or below the detection floor (Entry 14: MDE d ≈ 0.81; placement marginal, regime underpowered). One paragraph re-styled the image dependably — and, Entry 14 sharpens, moved one latent style dimension, not four — and re-structured it only at the margins. Second, the default here is a scene, not an emblem: DGI ~50, standing, subadditive, zero Hard-RCP across 48 frames — the centred-emblematic basin did not appear, so on this engine and prompt it is a conditional attractor, not a default. Third, what structure did move, moved distributionally: figure displacement went from a 42% to a 71% base rate, not to certainty; torque crossed to torquing in a fifth of draws; the R_spatial distribution shifted toward additive. The honest one-line account: the direction changed the palette reliably (one style dimension), the placement probabilistically (a unimodal rightward shift, at the detection floor), and the compositional regime below what this design could detect — with the caveat that this instrument sees light better than it sees the spatial-discipline items the prompt leaned on, so "barely" is partly a measurement ceiling (Entry 11).

And the study's methodological product: the first generative-variance baseline in the corpus. Because every draw was kept, "steering vs draw noise" is a measured ratio, and the nesting (2 jobs × 4 tiles) is checked, not assumed (Entry 09). That baseline is what lets contract delivery be stated as a rate rather than an anecdote.

What This Does Not Prove

Entry 14 — Ensemble analyses: the tests the single pair could not support [stable, exploratory]

Four questions the Sora pair could not ask, run on the 48 (existing vectors, no new generation).

Left: figure Δx vs non-figure mass Δx — the surround co-moves WITH the figure (S
Left: figure Δx vs non-figure mass Δx — the surround co-moves WITH the figure (Steered r +0.49), it does not counterbalance. Right: figure Δx distribution — a unimodal rightward shift under steering, not a comply/ignore split.

(1) The field's window: mass stays near-centred, and moves as a unit when pushed. The composition question the ensemble can answer: how far does MJ's field let structural mass wander, and what does it do with a displaced figure? Total structural mass stays in a tight central band across all 48 draws — mass Δx +0.030 ± 0.094 Steered, −0.040 ± 0.069 Default (spine values, as in Entry 03), both means within 0.04 of centre and neither cohort's mass leaving roughly ±0.2, even as the figure ranges over half the frame (Steered figure −0.02 to +0.41) and the prompt pushes displacement. That narrowness is the finding: a tight compositional window, measured. And when the field is pushed, it co-moves as a unit rather than counterbalancing — the de-biased test (figure Δx vs non-figure mass Δx, red-mask region removed, so the figure isn't correlated against a mass it belongs to) gives Steered r +0.49, essentially unchanged from the total-mass +0.50, with slope +0.51: the surround architecture shifts with the figure, it does not compensate. (Default shows a weak opposite tendency, r −0.30 — a faint surround compensation when the figure roams undirected.) Caveat that scopes it: MJ resists large figure excursion in the first place, so this reads the field in its low-excursion regime; whether a compensation mechanism would engage under larger forced displacement is not testable on an engine that absorbs the push. One scoping clause on the other AI study, no more: a tight-window field and a single-frame counterweight are different claims, and this study measures the former.

(2) "Style" is one knob, not four. The four top separators intercorrelate hard across the 48: shadow–τ +0.89, shadow–coupling +0.69, coupling–gap +0.79, β anti-correlated with all (−0.54 to −0.74). They are one latent painterly/tenebrist dimension, not four independent properties. Sharpens Entry 03: the steering moved one style axis, and the several metrics that led the separation are that axis seen from different sides.

(3) Placement is a unimodal shift, not comply-or-ignore. Steered figure Δx bins: 0 left / 7 centre / 17 right — steering emptied the left bin and carried the whole distribution rightward; no second mode. So the 42%→71% right-of-axis rate (Entry 05) is the engine nudging every draw, not obeying on a subset and ignoring the rest — the mean±sd framing was hiding a clean one-directional shift, now shown.

(4) Minimum detectable effect — "barely moved" made rigorous, and refined to three tiers. At n = 24/cohort (α .05, power .80) the minimum detectable Cohen's d ≈ 0.81, or ≈ 1.14 under the coupling-level clustering (deff ≈ 2). Against that floor: style axes (d 1.28–1.48) clear it decisively; placement axes sit AT it — mass Δx 0.89, figure Δx 0.72, real but marginal; regime axes sit below it — grid_asym 0.54, torque 0.51, DGI 0.40, genuinely underpowered. So the honest headline is not a flat "structure barely moved" but a powered three-tier statement: steering reliably moved one style dimension, marginally moved figure/mass placement, and moved the compositional regime by less than this design could detect. The last tier is a power statement, not a proof of no effect.

Owed (not run): a formal mixed-effects variance partition (#2) — statsmodels is not in the environment and no dependency was added mid-study; the ICC (Entry 09) and this MDE give its substance by hand, and the fitted model with CIs is queued.

Entry 15 — Exemplar nulls: no figural address anywhere; R_spatial stability confirmed [stable]

The C-3/C-4 nulls owed from Entry 12, on four exemplars (max-address and median per cohort; address null_sweep n=100 both nulls, R_spatial 5-seed).

Address — no figural address in the corpus, even at its most bilateral. Per the Bruegel discipline (raw central frontality can be spectral, not figural), each exemplar's real value is read against its scramble distribution, and none is figural. Three of the four are spectral or near-null: Default median 0.09 and Steered median 0.15 sit below their scramble maxima (0.32 / 0.41 — no distinctive address); Default max 0.40 has a scramble mean of only 0.17 but a max of 0.50, so it too reads spectral-dominated with at most a small margin (the HCB pattern). Steered max 0.47 is genuinely composition-dependent — its scramble mean collapses to 0.01, and only the extreme tail (max 0.50) still reaches the real value — but the frame builds that bilateral structure from the arch/light-column, not a figure owning the central column, so it is the corpus's non-figural composition-dependent category (the Sora-Steered third category), not figural address. Verdict: MJ, even at its most symmetric, builds architectural bilateral symmetry, not figural address — which retroactively validates Entry 12's refusal to make a figural claim from the raw values.

R_spatial — the cohort verdict is not a null artifact. Exemplar real values are cut-stable in the subadditive-to-additive band (0.14–0.21, 0.06–0.38 subadditive; 0.30–3.31, 0.39–1.39 additive-cut-sensitive) while their matched-stats nulls are seed-chaotic (r_max spans 0.29–9.18 across seeds). Same stability-vs-chaos signature as the rest of the corpus: the cohort placement (Entry 12) survives the null.

Entry 16 — Variance partition: one estimate for "how much is steering" [stable — the methodological product]

The unified variance-components model owed since Entry 03/06/09 (variance_components.json; balanced nested ANOVA — for this fully-balanced 2×3×2job×4tile design the ANOVA point estimates equal the REML ones, so this is the mixed-model partition, minus random-effect CIs). Per axis, % of total variance attributable to each source:

Variance partition per axis (balanced nested ANOVA): steering, aspect, cohort×as
Variance partition per axis (balanced nested ANOVA): steering, aspect, cohort×aspect, job (draw-cluster), tile (draw residual). Steering tall on the style axes (left); draw noise dominates the arrangement axes (right).
axissteeringaspectcoh×aspjobtile(draw)
coupling34.04.17.619.035.3
shadow33.23.87.55.050.4
β33.27.06.93.549.4
τ24.213.410.46.545.6
delta_x (mass)15.80.613.714.555.5
θ16.37.90.96.068.9
μ15.02.62.715.564.2
gap_fraction14.97.313.818.046.1
figure_dx11.63.56.313.465.1
grid_asym6.96.73.413.169.8
torque5.912.23.58.370.1
DGI3.87.83.015.470.0

The whole study in one table: steering owns 24–34% of the variance on the style axes and 4–16% on the arrangement axes; draw noise (job+tile) owns 50–85% almost everywhere. Two cross-checks fall out: coupling's job component (19.0%) independently confirms Entry 09's ICC 0.36 (the grid-clustered colour axis), and aspect's largest main-effect shares land on τ (13.4) and torque (12.2), while the cohort×aspect interaction Entry 06 flagged sits in its own column (largest on gap_fraction 13.8 and mass Δx 13.7) — both now quantified. The figure renders it.

Entry 17 — The particle ban: a pre-semantic ceiling, not a tooling gap [narrowed]

Entry 11 could not test the "no particle fill in right quadrant" directive because the bright-blob mask conflated bokeh with the overexposed hot core. The chroma-gated detector (bokeh_detector.py: small, near-circular, LOW-saturation bright discs — the gate excludes the warm hot core) fixes that confound, and its result is a principled ceiling. Right-quadrant count is equal across cohorts (median 54 Default vs 52 Steered) — no measurable effect of the ban — but the overlay (bokeh_check.png) shows why the equality is not the clean finding it looks like: the detector is dominated by backlit leaf-gaps at the tunnel opening, not bokeh. In this scene class, intentional bokeh and light-through-foliage are the same low-level feature (small bright neutral discs); separating them is a semantic distinction above the instrument's pre-semantic ceiling. So the honest close on the particle ban: the item is not measurable by this instrument on this scene, for a principled reason (it requires telling decorative bokeh from scene light, which is recognition), not a tooling gap. The chroma gate was still the right fix — it removed the hot-core confound — it just revealed a second, deeper one.

The chroma-gated bokeh detector on one Steered frame (red = kept right-quadrant
The chroma-gated bokeh detector on one Steered frame (red = kept right-quadrant discs): it catches backlit leaf-gaps at the tunnel opening, not decorative bokeh — the two are the same low-level feature, a pre-semantic ceiling.

Entry 18 — Bootstrap CIs: the three tiers become interval statements, and β is the standout [stable]

Entry 16's partition gave point estimates; this puts 95% CIs on them via a hierarchical bootstrap that respects the nesting (resample the 2 jobs within each cohort×aspect cell, then the 4 tiles within each resampled job, so grid-tile correlation is preserved; B = 2000, mixed_bootstrap.json). Distribution-free — the honest form at 12 job-clusters, and the reason a bootstrap beats REML here: no normality assumed. The steering effect is reported as Cohen's d (draw-sd denominator) with its CI, which answers "which steering effects are robustly non-zero" directly.

Steering effect (Cohen's d) per axis with 95% bootstrap CIs: style axes robustly
Steering effect (Cohen's d) per axis with 95% bootstrap CIs: style axes robustly large (gold, lower bound > 0.72), placement real but size-uncertain (blue), regime hugging zero (red). The MDE ≈ 0.8 line marks the detection floor.
tieraxessteering d [95% CI]
robustly largeβ1.48 [1.12, 2.64]
coupling / shadow / τ1.48 [0.82, 3.02] / 1.45 [0.78, 2.69] / 1.28 [0.72, 2.33]
real, size-uncertainmass Δx / gap / θ / μ / figure Δx0.89 [0.24, 2.09] / 0.90 [0.26, 2.09] / 0.87 [0.29, 1.70] / 0.81 [0.25, 1.73] / 0.72 [0.13, 1.43]
indistinguishable from negligiblegrid_asym / torque / DGI0.54 [0.05, 1.39] / 0.51 [0.04, 1.11] / 0.40 [0.03, 1.22]

The MDE three-tier of Entry 14 is now an interval statement: the four style axes have d-CIs bounded above 0.72 (all robustly at least medium-large); the placement/cohesion axes have CIs that exclude 0 but reach down into small-effect territory (real, size-uncertain); the regime axes (DGI, torque, grid_asym) have lower bounds at 0.03–0.05, hugging zero — consistent with negligible. β is the single most robustly-moved axis — the only one whose d-CI lower bound (1.12) is itself a large effect, and its steering variance-share CI [21.3, 52.2] is the highest-floored of any axis. Colour retreating from the void (β) under the painterly shift is the steering's most certain structural consequence.

Honesty on the shares: the variance-component CIs are wide (coupling steering share 34% [12.7, 56.5]; DGI 3.8% [0, 23.3]) — the faithful reflection of only 12 clusters. So Entry 16's percentages are central tendencies with real uncertainty on the random terms; the fixed steering effect (d, powered by 24 tiles/cohort) is the better-estimated output and its CIs are the ones to trust. One caveat even there: the bootstrap resamples only 2 jobs per cell, so the between-job variation feeding every interval — d-CIs included, not just the share-CIs — rests on a thin cluster resample; read the interval widths as themselves approximate. Either way the ordering and the tiers hold.


Pre-registration scorecard (8 corpus-level predictions)

#PredictionVerdict
1Default is scene-mode: DGI ~40s, most draws Not-RCPMostly hit — Not-RCP decisively (0 Hard in 48, no ring-fit ≥ 0.5); DGI 50.3 overshoots the predicted 40s
2Steered DGI ≥ Default (the risky bet)Wrong — 47.1 < 50.3, weakly (0.40 draw-sd)
3Figure-placement draw scatter: Default > SteeredHit — sd 0.124 vs 0.088
4No arrangement axis separates ≥ 2× sd (registered vs pooled sd)Hit under both denominators — max anywhere 1.48 on draw-sd (1.40 on the registered pooled sd); every arrangement axis < 1
5Style axes (τ/shadow, gap, theta) lead the separationHit — shadow/coupling/β/τ are the top four
6Basin strongest at 1:1; figure most lateral at 4:3Wrong twice — Default peaks at 4:3; displacement peaks at Steered 1:1; a cohort×aspect interaction
7Colour policy low-variance within cohorts; cross-cohort openAnswered — within-cohort modest, cross-cohort large: the policy moved with the style
8Right-of-axis ~half of Steered drawsUnderestimate — 17/24 (71%), against a 42% base rate

Three wrong, one near-miss, and the wrongs carry the study: steering re-styled more than it re-structured; the aspect factor interacts; the scene attractor, not the emblem, is this engine+prompt's default.

Why the ensemble is the unit [methodological note]

A single generated image is an anecdote; the ensemble exposes the generator's operating characteristics. Every quantity this study rests on — variance, steerability, the aspect interaction, response consistency, the field window — is a property of the distribution, not of any one draw, and cannot be estimated from a single exemplar. That is why all 48 draws were kept and none selected: the unit of analysis is the sample of engine behaviour, not the image, and generating more would have refined the same properties rather than revealed different ones. The two cohorts are, in these terms, the engine's native operating regime and a perturbed operating regime, and a directed prompt is read as a perturbation of the distribution (the governing frame, Entry 01), not an instruction an output obeys or fails.

Status

Tranche 3 complete (Entries 01–18 + scorecard + What This Does Not Prove): contract masks (11), cohort candidates (12), ensemble analyses (14), exemplar nulls (15), variance partition (16), bokeh ceiling (17), bootstrap CIs (18). Register rows appended (C-3/C-4 cohort form).

Owed — all tranche-2 analysis items closed this pass (Entries 15–17): address exemplar nulls (Entry 15: no figural address), C-3/C-4 exemplar nulls (Entry 15: R_spatial stability confirmed), variance-decomposition figure + model (Entry 16), chroma-gated bokeh detector (Entry 17: a pre-semantic ceiling). Remaining:

(The mixed-model CIs owed against Entry 16 were delivered in Entry 18 via a nesting-respecting bootstrap — statsmodels/REML not needed, and preferable at 12 clusters.)