Steering authority
Do the steering clauses in a prompt move a picture's composition, or only its tone? Measured on an existing paired Midjourney corpus with the Crit and Field kernels, unmodified.
Abstract
Midjourney was given the same 112 portrait prompts twice: once plain, and once with a single steering clause appended naming a compositional goal (place the figure off-centre, build the form from broad value masses, break the symmetry of the field, and so on). Four variants were generated per prompt, so every job carries its own record of what the sampler does when nothing changes. That corpus, 927 images, had never been measured. This study measures all of them with the same two kernels that power The Crit and The Field, pairs each steered job against its default twin by sitter, and asks one question: do the clauses move a picture's composition, or only its tone?
Every effect is reported in dice, the spread among the four variants of a single prompt. One dice is what re-rolling the same prompt already gives you for free, so an effect under 1.0 is weaker than the variation a user sees without changing anything.
Eleven predictions, one or two per clause naming the axis it should move and in which direction, plus a global prediction that the Default Gravity Index should fall if steering escapes the model's habitual composition, were written before any number was computed. One held. Two clauses did appear to move structure, in the wrong direction: steered images measured as more centred and more settled. That turned out to be an artifact. The Crit's material mask is part luminance, so darkening a picture moves every structural reading; these clauses darkened the pictures, and once the tone change is regressed out the entire DGI shift disappears (+2.54 points, p = 0.0009, becomes zero).
The study then re-measured with tone equalised in both arms: each image's luminance remapped onto a common reference by a monotone mapping, which leaves placement, edges and ordering intact and changes only the distribution of values. Under that control the compositional instructions move structure by a quarter to a half of a dice, in directions the words did not ask for, while the tonal instructions move tone by one to two dice. Steering also failed at the aim stated in the original spreadsheet: no structural axis widened, so the model's operating volume did not expand. Finally three levers are placed on one scale, the steering clause, the genre register already present in every prompt, and simply keeping the best of the four variants already generated. Selection wins, the genre word is second, and the explicit compositional instruction is last.
What was done, in order
- Decoded all 927 images exactly as The Crit's app does: RGBA, longest side 512 px, bilinear.
- Measured each with the product kernels imported unmodified, 99 numeric columns per image.
- Grouped the four variants of each prompt into a job, and paired steered jobs to default jobs by sitter (106 pairs).
- Set the unit of measurement: the median within-job spread across the four variants, the dice.
- Tested the eleven pre-declared predictions, with Benjamini-Hochberg correction, plus a separate 534-test exploratory scan corrected on its own.
- Tested the tone confound directly: correlation of every structural change with the tone change, and the DGI shift before and after regressing tone out.
- Re-measured with tone equalised, applying the identical remap to both arms, then repeated the structural tests, added equivalence tests with a declared margin of one dice, and measured whether the spread widened or narrowed.
- Compared three levers on the dice scale, and measured 125 human paintings as a reference band, tone-matched the same way.
| Element | This study |
|---|---|
| Corpus | 927 Midjourney portraits, 112 sitters, 8 registers, generated by the owner before this analysis |
| Unit | A job: one prompt, four variants |
| Treatment | One of six steering clauses appended to an otherwise identical prompt |
| Pairing | Same sitter, same base prompt, steered against default (106 pairs) |
| Scale | Dice: the median spread among four variants of one prompt |
| Main control | Monotone luminance remap applied identically to both arms |
| Reference band | 116 museum photographs and 9 Cézannes, tone-matched to the AI reference |
| Pre-registered | 11 clause-to-axis predictions and 1 global prediction, fixed before results |
1. The corpus was already the experiment
112 sitters, each generated twice by Midjourney: once with a base prompt and once with one of six steering clauses appended, four variants per job. The four variants are the sampler's own spread, called the dice here. Every effect below is measured in dice units, so "1.0" means the clause moved the picture as much as re-rolling the same prompt does.
Images decoded exactly as The Crit's app does (RGBA, longest side 512, bilinear) and measured by
products/the-crit/js/kernel.js and products/the-field/js/color-kernel.js, imported unmodified
(sha256 460b745c235f6c63 and 6eef90e7d1ab8a24, Node v24.13.1). 99 numeric columns per image.
2. One of eleven predictions held
| Clause | Axis it names | Predicted | Shift (dice) | p | Verdict |
|---|---|---|---|---|---|
| off-center mass | centroid offset | up | +0.37 | 0.17 | no effect |
| environmental asymmetry | sector variation | up | +0.09 | 0.77 | no effect |
| mass economy | dispersion (SDI) | down | +0.10 | 0.7 | no effect |
| chiaroscuro / value-structure | lightness contrast | up | +0.75 | 0.33 | no effect |
| color autonomy | colour independence | up | +0.35 | 0.12 | no effect |
| tonal opposition | spotlight concentration | up | +0.79 | 0.00056 | held |
3. The apparent structural win was tone
Raw, the Default Gravity Index rose under the off-center and asymmetry clauses, meaning the steered images sat closer to the default attractor. The Crit's material mask is tone-thresholded, and those clauses darkened the pictures. Mask size tracks lightness at r = −0.68. Regress the tone change out and the DGI shift disappears entirely: +2.54 (p = 0.0009) becomes −0.00 (p = 1.0).
Three examples, with the control
Each steered image's luminance is remapped onto its paired default: a monotone mapping, so placement and edges survive and only the tone distribution changes. The first two rows show clauses that asked for a compositional change, moved the named axis the wrong way, and moved tone several dice instead. The third row is the one prediction that held.
“the figure placed well off-center, its weight answered by a large quiet passage of shadow”
aimed at centroid offset: moved -1.56 dice. Tone moved -4.07 dice.
DGI 34 · offset 0.052 · tone 3.30
DGI 54 · offset 0.026 · tone 1.72
the control: same placement, default tone
“form built from broad masses of light and deep shadow, the darks modeling the structure”
aimed at lightness contrast: moved +1.15 dice. Tone moved -2.47 dice.
DGI 46 · offset 0.021 · tone 1.73
DGI 60 · offset 0.005 · tone 0.77
the control: same placement, default tone
“the lit head set against an independent field of muted color”
aimed at spotlight concentration: moved +2.48 dice. Tone moved +0.52 dice.
DGI 53 · offset 0.020 · tone 1.49
DGI 40 · offset 0.052 · tone 1.70
the control: same placement, default tone
4. With tone equalised, what survives is small and unasked-for
Both arms now receive identical processing. Pooled across all clauses, every structural axis is statistically equivalent to no change within one dice, and the sample can detect about a quarter of a dice. Four axes are nonetheless non-zero: dispersion +0.51, inner mass fraction −0.33, centroid offset +0.27 and centering +0.22. So the clauses do touch structure, faintly, but not in the direction they asked for: mass economy asked dispersion down and it went up, and centering rose rather than fell. Per clause, on its own named axis, only the two tonal predictions reach significance.
| Axis | Shift (dice) | 90% CI | Detectable at 80% | Equivalent within 1 dice |
|---|---|---|---|---|
| DGI score | +0.01 | [-0.15, +0.18] | 0.28 | yes |
| centroid offset | +0.27 | [+0.11, +0.44] | 0.28 | yes |
| sector variation | +0.08 | [-0.11, +0.26] | 0.31 | yes |
| dispersion (SDI) | +0.51 | [+0.34, +0.67] | 0.28 | yes |
| centering | +0.22 | [+0.05, +0.39] | 0.29 | yes |
| symmetry | -0.08 | [-0.26, +0.11] | 0.31 | yes |
| frame outside mask (rv) | -0.06 | [-0.29, +0.18] | 0.40 | yes |
| inner mass fraction | -0.33 | [-0.49, -0.17] | 0.28 | yes |
5. Steering did not widen the model
The experiment's stated aim was to expand Midjourney's operating volume. It did not. With identical processing in both arms, no axis widened: within-job ranges sit between ×0.77 and ×1.06, and only the DGI narrowing reaches significance (×0.85, p = 0.017). Spread across sitters is unchanged. An earlier version of this table showed a dramatic narrowing on every axis; that was an artifact of processing only the steered arm, and it is the same lesson as section 3.
| Axis | Within-job range | p | Across-sitter spread | p |
|---|---|---|---|---|
| DGI score | ×0.85 | 0.02 | ×0.97 | 0.5 |
| centroid offset | ×1.06 | 0.4 | ×1.13 | 0.3 |
| sector variation | ×0.92 | 0.2 | ×0.92 | 0.3 |
| dispersion (SDI) | ×0.91 | 0.2 | ×0.90 | 0.3 |
| centering | ×0.88 | 0.07 | ×0.88 | 0.3 |
| symmetry | ×0.92 | 0.2 | ×0.92 | 0.3 |
| frame outside mask (rv) | ×0.77 | 0.07 | ×0.95 | 0.7 |
| inner mass fraction | ×0.93 | 0.3 | ×0.99 | 0.8 |
6. Three levers, measured on the same scale
| Axis | Steering clause | Style register | Best-of-4 selection |
|---|---|---|---|
| DGI score | 0.01 | 0.55 | 1.06 |
| centroid offset | 0.27 | 0.47 | 1.39 |
| sector variation | 0.08 | 0.62 | 1.26 |
| dispersion (SDI) | 0.51 | 0.64 | 1.12 |
| centering | 0.22 | 0.60 | 1.10 |
| symmetry | 0.08 | 0.62 | 1.06 |
| frame outside mask (rv) | 0.06 | 0.92 | 1.62 |
| inner mass fraction | 0.33 | 0.60 | 1.29 |
Registers are confounded with their sitters and subjects, so that column is an upper bound on what a style word does. Selection needs no prompt change at all.
7. Where the human paintings sit
All images tone-matched to one reference before measuring, so this is not a tone comparison.
| Axis | AI default | AI steered | Museum | Cézanne | Human spread | Gap (AI SD) |
|---|---|---|---|---|---|---|
| DGI score | 50.759 | 50.834 | 45.664 | 45.667 | ×1.48 | -0.7 |
| centroid offset | 0.022 | 0.025 | 0.025 | 0.018 | ×1.70 | +0.2 |
| sector variation | 0.278 | 0.282 | 0.373 | 0.424 | ×1.73 | +1.0 |
| dispersion (SDI) | 0.269 | 0.272 | 0.265 | 0.254 | ×1.77 | -0.3 |
| centering | 0.807 | 0.823 | 0.759 | 0.750 | ×1.33 | -0.4 |
| symmetry | 0.652 | 0.647 | 0.537 | 0.470 | ×1.65 | -1.0 |
| frame outside mask (rv) | 0.252 | 0.251 | 0.320 | 0.403 | ×2.14 | +0.7 |
| inner mass fraction | 0.179 | 0.171 | 0.202 | 0.242 | ×1.93 | +0.8 |
Before tone control, for contrast
Raw, the gap looks enormous: DGI 50.4 for AI against 35.0 for the museum set and 33.6 for Cézanne. Most of that is tone, and most of it closes once tone is equalised. What survives is the asymmetry and the width.
| Axis | AI default (raw) | Museum (raw) | Cézanne (raw) |
|---|---|---|---|
| DGI score | 50.366 | 34.974 | 33.556 |
| centroid offset | 0.023 | 0.044 | 0.047 |
| sector variation | 0.289 | 0.550 | 0.600 |
| dispersion (SDI) | 0.268 | 0.256 | 0.240 |
| centering | 0.800 | 0.593 | 0.550 |
| symmetry | 0.639 | 0.338 | 0.260 |
| frame outside mask (rv) | 0.256 | 0.585 | 0.626 |
| inner mass fraction | 0.180 | 0.217 | 0.243 |
DGI 12 · offset 0.041 · sector var 1.01
DGI 33 · offset 0.057 · sector var 0.85
DGI 57 · offset 0.018 · sector var 0.23
The museum set is mixed genre, not portraits only, and these are photographs of paintings, so framing and surface are part of what is measured. Treat it as a reference band, not a matched control.
8. What this says about steering images
- Words are a tone lever, and a strong one. Value, shadow mass, palette count and lightness move on command by one to two dice.
- Words did not move composition to order. The structural movement that survives tone control is a quarter to a half of a dice, is not aligned with what the clause asked for, and is dwarfed by the sampler's own variation. The large effects that looked structural were tone acting through a tone-thresholded mask.
- Steering narrowed the output distribution. More instruction, less variety, and no new territory.
- Selection is the cheapest geometry lever available today. Best-of-four gives 1.1 to 1.6 dice with no prompt change, several times the matched clause effect, and the price of stronger selection is the keep-rate arithmetic in the curation study next door.
- The style register beats the compositional instruction on every structural axis, which suggests geometry here is carried by genre priors rather than by spatial words.
- Geometry likely needs a different injection site: layout conditioning, spatial control input, inpainting, crop or post-edit.
- For the instrument: any structural comparison across corpora whose tone differs needs this control first, or the mask polarity will manufacture a structural result.
Limits
One generator, one genre, one prompt family, about 18 pairs per clause and four variants per job. "No effect" means smaller than the dice, not zero: the sample can detect roughly a quarter of a dice pooled and about 0.7 of a dice per clause. The eight tracks may differ in generation date or model version, which is not recorded. Tone matching is monotone per image: it preserves ordering, not local contrast. No human judgement of these images was collected. Every statement here is about measured coordinates.
Definitions
Design terms
| Term | Meaning here |
|---|---|
| Job | One prompt submitted once, yielding four variants. |
| Variant | One of the four images Midjourney returns for a job. Differences among them come from the sampler, not from the prompt. |
| Dice (dice SD) | The median standard deviation among the four variants of a job, per axis. The unit for every effect in this report. 1.0 dice = as much movement as re-rolling the prompt. |
| Clause | The steering sentence appended to the base prompt, one of six. |
| Register | The genre phrase already in every prompt: Dutch Golden Age, Spanish Golden Age, 18th-century English, Flemish Baroque, Italian Renaissance, Early American Colonial, French Rococo, Northern Renaissance. |
| Track | A batch folder in the corpus (Track 1 to 8), 14 sitters each. Generation date and model version are not recorded. |
| Tone matching | Remapping an image's luminance onto a reference distribution. Monotone per image: the order of pixel values is preserved, so placement and edges survive and only the distribution of values changes. |
| Best-of-4 | Keeping whichever of the four already-generated variants scores highest on an axis. A selection lever, not a prompt change. |
Structural axes (The Crit kernel)
| Axis | Definition |
|---|---|
| Material mask | The pixels counted as material. Score = 0.65 × normalised Sobel edge magnitude + 0.35 × (1 − normalised blurred luminance), thresholded by Otsu, then morphologically opened and closed. The 0.35 term is why tone leaks into every structural reading. |
| Mass map | Per-pixel weight inside the mask; the centroid and all moments below are computed on it. |
| Centroid offset (kernel.dx) | Horizontal distance from the mass centroid to the frame centre, divided by frame width. Larger = the weight sits further off-centre. |
| Dispersion, SDI (kernel.sdi) | Mass-weighted mean distance from the centroid, divided by the frame diagonal. Larger = mass spread wider from its own centre. |
| Sector variation (radial.cvSectors) | Coefficient of variation of mass across 8 sectors around the frame centre. Larger = more lopsided distribution of weight around the frame. |
| Inner mass fraction | Share of total mass inside the inner radial band. Larger = more mass gathered at the centre. |
| Frame outside mask (kernel.rv) | 1 − (mask pixels ÷ all pixels): the share of the frame that is not material. Larger = more empty field. |
| DGI, Default Gravity Index | 100 × (0.35 × centering + 0.25 × lock + 0.25 × central + 0.15 × symmetry). Higher = closer to the default attractor: centred, frame-locked, centrally concentrated, sector-symmetric. The Crit's headline number. |
| centering | 1 − clamp(centroid-to-frame-centre distance ÷ diagonal ÷ 0.18). Higher = more centred. |
| lock | 1 − clamp((radial compliance about the centroid − about the frame centre + 0.08) ÷ 0.16). Higher = radial structure aligned to the frame rather than to the subject. |
| central | The inner mass fraction, clamped to [0,1]. |
| symmetry | 1 − clamp(sector variation ÷ 0.8). Higher = weight distributed more evenly around the frame. |
| Tonal zones, mean tonal zone | Pixels binned into 10 luminance zones (Rec. 601 luma); the mean zone index is the image's overall value. Lower = darker. |
Colour axes (The Field kernel)
| Axis | Definition |
|---|---|
| Spotlight concentration | Largest connected component of the light mask, as a fraction. Higher = one dominant lit region rather than scattered highlights. |
| Colour independence | 0.5 × |mean colour-edge response − mean luminance-gradient response| + 0.5 × (1 − mass/colour alignment). Higher = colour doing work the light is not doing. |
| Coupling Index | The Field's headline: colour bound to structure (align) versus colour running free (bind). Higher = more bound. |
| Effective palette count | Inverse Simpson index of palette weights, 1 ÷ Σw². Roughly, how many colours actually carry the image. |
| Lightness contrast | 95th minus 5th percentile of OKLab L. Higher = wider value range. |
| Dark mass fraction | Share of pixels in the dark mask. |
Statistical terms
| Term | Meaning here |
|---|---|
| Paired test | Each steered job is compared with the default job of the same sitter, so subject and register cancel out. |
| p | Probability of seeing an effect this large if the clause did nothing. Small = unlikely to be chance alone. |
| Benjamini-Hochberg (BH) | A correction applied when many tests run at once, controlling the share of false positives among the findings called significant. Applied separately to the 11 pre-declared tests and the 534-test scan. |
| Equivalence test | The reverse of a significance test: it asks whether an effect is small enough to rule out anything bigger than a declared margin, here one dice. |
| MDE | Minimum detectable effect: the smallest effect this sample would catch 80% of the time. Roughly a quarter of a dice pooled, 0.7 for a single clause. |
| eta-squared | Share of variance in an axis explained by a grouping, used for the register comparison. |
| r | Correlation, −1 to +1, used for the tone confound. |
| Exploratory | Tests not declared in advance. Reported with their own correction and never used to claim a prediction was confirmed. |
Methods in full
Decoding. Every image opened with Pillow, converted to RGBA, scaled so the longest side is 512 px using
bilinear resampling, matching The Crit's analysis path (its ANALYSIS_MAX_DIM is 512).
Measurement. Raw RGBA passed to Node, which imports products/the-crit/js/kernel.js (sha256
460b745c235f6c63) and products/the-field/js/color-kernel.js (sha256 6eef90e7d1ab8a24) unmodified, Node
v24.13.1. Numeric leaves flattened to 99 columns per image. No kernel parameter was changed for this study.
Clause assignment. Taken from the pasted prompt text, not the spreadsheet's own label: three rows carry a label that disagrees with their prompt. A sitter whose truncated filename matches two steered prompts with different clauses is dropped as ambiguous.
Pairing and scale. Job mean per axis; steered minus default per sitter; divided by the median within-job standard deviation (the dice). 106 pairs, 17 to 19 per clause.
Tone control. Each image's luminance quantiles are remapped onto a reference set of quantiles, and the resulting per-pixel gain is applied to RGB. Three references are used: the paired default job (for steered against default), one corpus-wide reference (for the register comparison, which removes each register's tonal signature), and the AI-default reference (for the human paintings). In every reported comparison both arms receive the same treatment; an earlier version processed only one arm, which produced a false narrowing result, corrected here.
Statistics. Paired t-tests; BH at 0.05 within each family; two one-sided equivalence tests against a one-dice margin; Levene's test for spread; one-way ANOVA with eta-squared for register; Pearson correlation and simple linear regression for the tone confound. All analysis in NumPy and SciPy.
Reproduction. run_measure.py, measure.mjs, analyze.py,
tone_control*.py, extra_analysis.py, final_analysis.py,
build_report.py, with measurements*.csv, human_*.csv,
analysis.json and final_analysis.json beside this page.