Steering authority

Do the steering clauses in a prompt move a picture's composition, or only its tone? Measured on an existing paired Midjourney corpus with the Crit and Field kernels, unmodified.

Paired sitters
106
default vs steered, 4 variants each
Predictions held
1 of 11
written before results
DGI shift after tone control
+0.01
dice; from +2.54 points raw (p = 0.0009)
Best-of-4 selection
1.1–1.6 dice
against 0.0–0.5 for any clause
Human paintings
1.3–2.1× wider
after tone control; more asymmetric

Abstract

Midjourney was given the same 112 portrait prompts twice: once plain, and once with a single steering clause appended naming a compositional goal (place the figure off-centre, build the form from broad value masses, break the symmetry of the field, and so on). Four variants were generated per prompt, so every job carries its own record of what the sampler does when nothing changes. That corpus, 927 images, had never been measured. This study measures all of them with the same two kernels that power The Crit and The Field, pairs each steered job against its default twin by sitter, and asks one question: do the clauses move a picture's composition, or only its tone?

Every effect is reported in dice, the spread among the four variants of a single prompt. One dice is what re-rolling the same prompt already gives you for free, so an effect under 1.0 is weaker than the variation a user sees without changing anything.

Eleven predictions, one or two per clause naming the axis it should move and in which direction, plus a global prediction that the Default Gravity Index should fall if steering escapes the model's habitual composition, were written before any number was computed. One held. Two clauses did appear to move structure, in the wrong direction: steered images measured as more centred and more settled. That turned out to be an artifact. The Crit's material mask is part luminance, so darkening a picture moves every structural reading; these clauses darkened the pictures, and once the tone change is regressed out the entire DGI shift disappears (+2.54 points, p = 0.0009, becomes zero).

The study then re-measured with tone equalised in both arms: each image's luminance remapped onto a common reference by a monotone mapping, which leaves placement, edges and ordering intact and changes only the distribution of values. Under that control the compositional instructions move structure by a quarter to a half of a dice, in directions the words did not ask for, while the tonal instructions move tone by one to two dice. Steering also failed at the aim stated in the original spreadsheet: no structural axis widened, so the model's operating volume did not expand. Finally three levers are placed on one scale, the steering clause, the genre register already present in every prompt, and simply keeping the best of the four variants already generated. Selection wins, the genre word is second, and the explicit compositional instruction is last.

What was done, in order

  1. Decoded all 927 images exactly as The Crit's app does: RGBA, longest side 512 px, bilinear.
  2. Measured each with the product kernels imported unmodified, 99 numeric columns per image.
  3. Grouped the four variants of each prompt into a job, and paired steered jobs to default jobs by sitter (106 pairs).
  4. Set the unit of measurement: the median within-job spread across the four variants, the dice.
  5. Tested the eleven pre-declared predictions, with Benjamini-Hochberg correction, plus a separate 534-test exploratory scan corrected on its own.
  6. Tested the tone confound directly: correlation of every structural change with the tone change, and the DGI shift before and after regressing tone out.
  7. Re-measured with tone equalised, applying the identical remap to both arms, then repeated the structural tests, added equivalence tests with a declared margin of one dice, and measured whether the spread widened or narrowed.
  8. Compared three levers on the dice scale, and measured 125 human paintings as a reference band, tone-matched the same way.
ElementThis study
Corpus927 Midjourney portraits, 112 sitters, 8 registers, generated by the owner before this analysis
UnitA job: one prompt, four variants
TreatmentOne of six steering clauses appended to an otherwise identical prompt
PairingSame sitter, same base prompt, steered against default (106 pairs)
ScaleDice: the median spread among four variants of one prompt
Main controlMonotone luminance remap applied identically to both arms
Reference band116 museum photographs and 9 Cézannes, tone-matched to the AI reference
Pre-registered11 clause-to-axis predictions and 1 global prediction, fixed before results

1. The corpus was already the experiment

112 sitters, each generated twice by Midjourney: once with a base prompt and once with one of six steering clauses appended, four variants per job. The four variants are the sampler's own spread, called the dice here. Every effect below is measured in dice units, so "1.0" means the clause moved the picture as much as re-rolling the same prompt does.

Images decoded exactly as The Crit's app does (RGBA, longest side 512, bilinear) and measured by products/the-crit/js/kernel.js and products/the-field/js/color-kernel.js, imported unmodified (sha256 460b745c235f6c63 and 6eef90e7d1ab8a24, Node v24.13.1). 99 numeric columns per image.

2. One of eleven predictions held

Each clause versus the axis it names. Orange is the raw comparison, blue after tone control. The grey band is the sampler's own spread: anything inside it is weaker than re-rolling the prompt.
ClauseAxis it namesPredictedShift (dice)pVerdict
off-center masscentroid offsetup+0.370.17no effect
environmental asymmetrysector variationup+0.090.77no effect
mass economydispersion (SDI)down+0.100.7no effect
chiaroscuro / value-structurelightness contrastup+0.750.33no effect
color autonomycolour independenceup+0.350.12no effect
tonal oppositionspotlight concentrationup+0.790.00056held

3. The apparent structural win was tone

Raw, the Default Gravity Index rose under the off-center and asymmetry clauses, meaning the steered images sat closer to the default attractor. The Crit's material mask is tone-thresholded, and those clauses darkened the pictures. Mask size tracks lightness at r = −0.68. Regress the tone change out and the DGI shift disappears entirely: +2.54 (p = 0.0009) becomes −0.00 (p = 1.0).

Each point is one sitter: how much the clause changed tone, against how much it changed DGI.

Three examples, with the control

Each steered image's luminance is remapped onto its paired default: a monotone mapping, so placement and edges survive and only the tone distribution changes. The first two rows show clauses that asked for a compositional change, moved the named axis the wrong way, and moved tone several dice instead. The third row is the one prediction that held.

portrait of a French courtier 18th-century French Rococo
Steering attempted: off-center mass
“the figure placed well off-center, its weight answered by a large quiet passage of shadow”
aimed at centroid offset: moved -1.56 dice. Tone moved -4.07 dice.
default prompt
DGI 34 · offset 0.052 · tone 3.30
+ steering clause
DGI 54 · offset 0.026 · tone 1.72
steered, tone matched to the default
the control: same placement, default tone
portrait of a Medici courtier Italian Renaissance oil pai
Steering attempted: chiaroscuro / value-structure
“form built from broad masses of light and deep shadow, the darks modeling the structure”
aimed at lightness contrast: moved +1.15 dice. Tone moved -2.47 dice.
default prompt
DGI 46 · offset 0.021 · tone 1.73
+ steering clause
DGI 60 · offset 0.005 · tone 0.77
steered, tone matched to the default
the control: same placement, default tone
portrait of a New England minister early American colonia
Steering attempted: tonal opposition
“the lit head set against an independent field of muted color”
aimed at spotlight concentration: moved +2.48 dice. Tone moved +0.52 dice.
default prompt
DGI 53 · offset 0.020 · tone 1.49
+ steering clause
DGI 40 · offset 0.052 · tone 1.70
steered, tone matched to the default
the control: same placement, default tone

4. With tone equalised, what survives is small and unasked-for

Both arms now receive identical processing. Pooled across all clauses, every structural axis is statistically equivalent to no change within one dice, and the sample can detect about a quarter of a dice. Four axes are nonetheless non-zero: dispersion +0.51, inner mass fraction −0.33, centroid offset +0.27 and centering +0.22. So the clauses do touch structure, faintly, but not in the direction they asked for: mass economy asked dispersion down and it went up, and centering rose rather than fell. Per clause, on its own named axis, only the two tonal predictions reach significance.

AxisShift (dice)90% CIDetectable at 80%Equivalent within 1 dice
DGI score+0.01[-0.15, +0.18]0.28yes
centroid offset+0.27[+0.11, +0.44]0.28yes
sector variation+0.08[-0.11, +0.26]0.31yes
dispersion (SDI)+0.51[+0.34, +0.67]0.28yes
centering+0.22[+0.05, +0.39]0.29yes
symmetry-0.08[-0.26, +0.11]0.31yes
frame outside mask (rv)-0.06[-0.29, +0.18]0.40yes
inner mass fraction-0.33[-0.49, -0.17]0.28yes

5. Steering did not widen the model

The experiment's stated aim was to expand Midjourney's operating volume. It did not. With identical processing in both arms, no axis widened: within-job ranges sit between ×0.77 and ×1.06, and only the DGI narrowing reaches significance (×0.85, p = 0.017). Spread across sitters is unchanged. An earlier version of this table showed a dramatic narrowing on every axis; that was an artifact of processing only the steered arm, and it is the same lesson as section 3.

AxisWithin-job rangepAcross-sitter spreadp
DGI score×0.850.02×0.970.5
centroid offset×1.060.4×1.130.3
sector variation×0.920.2×0.920.3
dispersion (SDI)×0.910.2×0.900.3
centering×0.880.07×0.880.3
symmetry×0.920.2×0.920.3
frame outside mask (rv)×0.770.07×0.950.7
inner mass fraction×0.930.3×0.990.8

6. Three levers, measured on the same scale

What actually moves structure: the steering clause, the style register (Dutch Golden Age, Rococo and so on, tone-matched), and simply picking the best of the four variants already generated.
AxisSteering clauseStyle registerBest-of-4 selection
DGI score0.010.551.06
centroid offset0.270.471.39
sector variation0.080.621.26
dispersion (SDI)0.510.641.12
centering0.220.601.10
symmetry0.080.621.06
frame outside mask (rv)0.060.921.62
inner mass fraction0.330.601.29

Registers are confounded with their sitters and subjects, so that column is an upper bound on what a style word does. Selection needs no prompt change at all.

7. Where the human paintings sit

All images tone-matched to one reference before measuring, so this is not a tone comparison.

Each dot is one image or one job mean; the black bar is the group mean. Both human sets share a colour because the contrast being drawn is AI against human, not museum against Cézanne.
AxisAI defaultAI steeredMuseumCézanneHuman spreadGap (AI SD)
DGI score50.75950.83445.66445.667×1.48-0.7
centroid offset0.0220.0250.0250.018×1.70+0.2
sector variation0.2780.2820.3730.424×1.73+1.0
dispersion (SDI)0.2690.2720.2650.254×1.77-0.3
centering0.8070.8230.7590.750×1.33-0.4
symmetry0.6520.6470.5370.470×1.65-1.0
frame outside mask (rv)0.2520.2510.3200.403×2.14+0.7
inner mass fraction0.1790.1710.2020.242×1.93+0.8

Before tone control, for contrast

Raw, the gap looks enormous: DGI 50.4 for AI against 35.0 for the museum set and 33.6 for Cézanne. Most of that is tone, and most of it closes once tone is equalised. What survives is the asymmetry and the width.

AxisAI default (raw)Museum (raw)Cézanne (raw)
DGI score50.36634.97433.556
centroid offset0.0230.0440.047
sector variation0.2890.5500.600
dispersion (SDI)0.2680.2560.240
centering0.8000.5930.550
symmetry0.6390.3380.260
frame outside mask (rv)0.2560.5850.626
inner mass fraction0.1800.2170.243
museum painting
DGI 12 · offset 0.041 · sector var 1.01
museum painting
DGI 33 · offset 0.057 · sector var 0.85
AI default
DGI 57 · offset 0.018 · sector var 0.23

The museum set is mixed genre, not portraits only, and these are photographs of paintings, so framing and surface are part of what is measured. Treat it as a reference band, not a matched control.

8. What this says about steering images

  1. Words are a tone lever, and a strong one. Value, shadow mass, palette count and lightness move on command by one to two dice.
  2. Words did not move composition to order. The structural movement that survives tone control is a quarter to a half of a dice, is not aligned with what the clause asked for, and is dwarfed by the sampler's own variation. The large effects that looked structural were tone acting through a tone-thresholded mask.
  3. Steering narrowed the output distribution. More instruction, less variety, and no new territory.
  4. Selection is the cheapest geometry lever available today. Best-of-four gives 1.1 to 1.6 dice with no prompt change, several times the matched clause effect, and the price of stronger selection is the keep-rate arithmetic in the curation study next door.
  5. The style register beats the compositional instruction on every structural axis, which suggests geometry here is carried by genre priors rather than by spatial words.
  6. Geometry likely needs a different injection site: layout conditioning, spatial control input, inpainting, crop or post-edit.
  7. For the instrument: any structural comparison across corpora whose tone differs needs this control first, or the mask polarity will manufacture a structural result.

Limits

One generator, one genre, one prompt family, about 18 pairs per clause and four variants per job. "No effect" means smaller than the dice, not zero: the sample can detect roughly a quarter of a dice pooled and about 0.7 of a dice per clause. The eight tracks may differ in generation date or model version, which is not recorded. Tone matching is monotone per image: it preserves ordering, not local contrast. No human judgement of these images was collected. Every statement here is about measured coordinates.

Definitions

Design terms

TermMeaning here
JobOne prompt submitted once, yielding four variants.
VariantOne of the four images Midjourney returns for a job. Differences among them come from the sampler, not from the prompt.
Dice (dice SD)The median standard deviation among the four variants of a job, per axis. The unit for every effect in this report. 1.0 dice = as much movement as re-rolling the prompt.
ClauseThe steering sentence appended to the base prompt, one of six.
RegisterThe genre phrase already in every prompt: Dutch Golden Age, Spanish Golden Age, 18th-century English, Flemish Baroque, Italian Renaissance, Early American Colonial, French Rococo, Northern Renaissance.
TrackA batch folder in the corpus (Track 1 to 8), 14 sitters each. Generation date and model version are not recorded.
Tone matchingRemapping an image's luminance onto a reference distribution. Monotone per image: the order of pixel values is preserved, so placement and edges survive and only the distribution of values changes.
Best-of-4Keeping whichever of the four already-generated variants scores highest on an axis. A selection lever, not a prompt change.

Structural axes (The Crit kernel)

AxisDefinition
Material maskThe pixels counted as material. Score = 0.65 × normalised Sobel edge magnitude + 0.35 × (1 − normalised blurred luminance), thresholded by Otsu, then morphologically opened and closed. The 0.35 term is why tone leaks into every structural reading.
Mass mapPer-pixel weight inside the mask; the centroid and all moments below are computed on it.
Centroid offset (kernel.dx)Horizontal distance from the mass centroid to the frame centre, divided by frame width. Larger = the weight sits further off-centre.
Dispersion, SDI (kernel.sdi)Mass-weighted mean distance from the centroid, divided by the frame diagonal. Larger = mass spread wider from its own centre.
Sector variation (radial.cvSectors)Coefficient of variation of mass across 8 sectors around the frame centre. Larger = more lopsided distribution of weight around the frame.
Inner mass fractionShare of total mass inside the inner radial band. Larger = more mass gathered at the centre.
Frame outside mask (kernel.rv)1 − (mask pixels ÷ all pixels): the share of the frame that is not material. Larger = more empty field.
DGI, Default Gravity Index100 × (0.35 × centering + 0.25 × lock + 0.25 × central + 0.15 × symmetry). Higher = closer to the default attractor: centred, frame-locked, centrally concentrated, sector-symmetric. The Crit's headline number.
centering1 − clamp(centroid-to-frame-centre distance ÷ diagonal ÷ 0.18). Higher = more centred.
lock1 − clamp((radial compliance about the centroid − about the frame centre + 0.08) ÷ 0.16). Higher = radial structure aligned to the frame rather than to the subject.
centralThe inner mass fraction, clamped to [0,1].
symmetry1 − clamp(sector variation ÷ 0.8). Higher = weight distributed more evenly around the frame.
Tonal zones, mean tonal zonePixels binned into 10 luminance zones (Rec. 601 luma); the mean zone index is the image's overall value. Lower = darker.

Colour axes (The Field kernel)

AxisDefinition
Spotlight concentrationLargest connected component of the light mask, as a fraction. Higher = one dominant lit region rather than scattered highlights.
Colour independence0.5 × |mean colour-edge response − mean luminance-gradient response| + 0.5 × (1 − mass/colour alignment). Higher = colour doing work the light is not doing.
Coupling IndexThe Field's headline: colour bound to structure (align) versus colour running free (bind). Higher = more bound.
Effective palette countInverse Simpson index of palette weights, 1 ÷ Σw². Roughly, how many colours actually carry the image.
Lightness contrast95th minus 5th percentile of OKLab L. Higher = wider value range.
Dark mass fractionShare of pixels in the dark mask.

Statistical terms

TermMeaning here
Paired testEach steered job is compared with the default job of the same sitter, so subject and register cancel out.
pProbability of seeing an effect this large if the clause did nothing. Small = unlikely to be chance alone.
Benjamini-Hochberg (BH)A correction applied when many tests run at once, controlling the share of false positives among the findings called significant. Applied separately to the 11 pre-declared tests and the 534-test scan.
Equivalence testThe reverse of a significance test: it asks whether an effect is small enough to rule out anything bigger than a declared margin, here one dice.
MDEMinimum detectable effect: the smallest effect this sample would catch 80% of the time. Roughly a quarter of a dice pooled, 0.7 for a single clause.
eta-squaredShare of variance in an axis explained by a grouping, used for the register comparison.
rCorrelation, −1 to +1, used for the tone confound.
ExploratoryTests not declared in advance. Reported with their own correction and never used to claim a prediction was confirmed.

Methods in full

Decoding. Every image opened with Pillow, converted to RGBA, scaled so the longest side is 512 px using bilinear resampling, matching The Crit's analysis path (its ANALYSIS_MAX_DIM is 512).

Measurement. Raw RGBA passed to Node, which imports products/the-crit/js/kernel.js (sha256 460b745c235f6c63) and products/the-field/js/color-kernel.js (sha256 6eef90e7d1ab8a24) unmodified, Node v24.13.1. Numeric leaves flattened to 99 columns per image. No kernel parameter was changed for this study.

Clause assignment. Taken from the pasted prompt text, not the spreadsheet's own label: three rows carry a label that disagrees with their prompt. A sitter whose truncated filename matches two steered prompts with different clauses is dropped as ambiguous.

Pairing and scale. Job mean per axis; steered minus default per sitter; divided by the median within-job standard deviation (the dice). 106 pairs, 17 to 19 per clause.

Tone control. Each image's luminance quantiles are remapped onto a reference set of quantiles, and the resulting per-pixel gain is applied to RGB. Three references are used: the paired default job (for steered against default), one corpus-wide reference (for the register comparison, which removes each register's tonal signature), and the AI-default reference (for the human paintings). In every reported comparison both arms receive the same treatment; an earlier version processed only one arm, which produced a false narrowing result, corrected here.

Statistics. Paired t-tests; BH at 0.05 within each family; two one-sided equivalence tests against a one-dice margin; Levene's test for spread; one-way ANOVA with eta-squared for register; Pearson correlation and simple linear regression for the tone confound. All analysis in NumPy and SciPy.

Reproduction. run_measure.py, measure.mjs, analyze.py, tone_control*.py, extra_analysis.py, final_analysis.py, build_report.py, with measurements*.csv, human_*.csv, analysis.json and final_analysis.json beside this page.