Smaller than the dice
What prompt steering actually moves, and why the instrument agreed with it at first.
A prompt is offered as a control surface. Add a clause about composition and the picture is supposed to change in that way. The advice circulates in exactly that form, as a list of phrases that work, and it is almost never accompanied by the one number that would make it checkable: how much the picture changes when you change nothing at all.
Generators are stochastic. Run the same prompt four times and you get four different pictures. Any honest claim about a prompt has to clear that variation first. This piece proposes that variation as the unit of measurement, calls it the dice, and uses it on a steering experiment that was already sitting on disk. It also documents the part that nearly fooled the measurement: the first, strongest result was an artifact of the instrument, and the instrument had failed this way before.
1. The missing denominator
Here are ten prompts from the corpus. Each has four variants, generated together, from identical text. The black bar is the average of the four.
This is the sampler's own spread. One dice is the median standard deviation among the four variants of a prompt, computed per axis. It is the amount a user already gets for free, by pressing the button again. For the Default Gravity Index, one dice is 7.8 points. For centroid offset it is 0.0169 of the frame width.
There is nothing new in the statistics. This is an effect size, a difference divided by a standard deviation, and the only choice being made is which standard deviation to divide by. Psychology divides by the spread between people. Here the relevant spread is the one the machine produces on its own, because that is what the user is competing with. Naming it and quoting it routinely is the whole proposal.
The rule that follows is uncomfortable and simple: an effect under one dice is weaker than re-rolling. It may still be real. It is not yet a control surface, because a user cannot reach it reliably without generating repeatedly, and if they are generating repeatedly they have a stronger lever available anyway (section 7).
2. A test that was already on disk
The corpus was generated before this analysis, for a different purpose, which makes it a better test than anything designed afterwards. 112 sitters across eight painting registers were each rendered twice by Midjourney: once with a plain prompt, and once with the identical prompt plus one steering clause. Four variants per job, 927 images. The stated aim in the originating spreadsheet was to test whether steering could expand the model's operating volume.
| Clause | Text appended to the prompt |
|---|---|
| off-center mass | “the figure placed well off-center, its weight answered by a large quiet passage of shadow” |
| chiaroscuro / value-structure | “form built from broad masses of light and deep shadow, the darks modeling the structure” |
| color autonomy | “cool shadows set against warm half-lights, the color departing from the light” |
| environmental asymmetry | “asymmetric composition, the sitter turned with open gaze-room to one side” |
| mass economy | “constructed from a few large coherent value shapes rather than even all-over detail” |
| tonal opposition | “the lit head set against an independent field of muted color” |
Each clause names something a painter would recognise, and each maps onto an axis the instruments already measure. Before any number was computed, eleven predictions were written down, each naming an axis and a direction, plus one global prediction: if steering escapes the model's habitual composition, the Default Gravity Index should fall. Measurement used the kernels behind The Crit and The Field, imported unmodified, on images decoded exactly as the apps decode them.
3. First result: one in eleven
One prediction held on the first pass: the clause asking for a lit head against an independent field raised spotlight concentration by +0.72 dice. The compositional predictions did not hold. Two of them appeared to move in the wrong direction, and that is the interesting part, because two clauses asking for off-centre weight and asymmetry produced images that measured as more centred and more settled: the Default Gravity Index rose by 2.54 points (p = 0.0009).
4. The instrument moved with the light
A result that contradicts the instruction is worth more scrutiny than one that confirms it. The Crit's material mask, the thing that decides which pixels count as substance before any geometry is computed, is built like this:
score = 0.65 × normalised edge magnitude + 0.35 × (1 − normalised luminance), thresholded by Otsu, then opened and closed.
Thirty-five per cent of that score is darkness. Change a picture's tone and the mask changes; change the mask and the centroid, the dispersion, the sector balance and the index built on top of them all change with it. The steering clauses did change tone, because most of them mention shadow.
The correlation between the tone change and the index change is -0.49. Across the whole corpus, mask coverage against image lightness is -0.62. Regress the tone change out of the index change and the entire effect disappears: from +2.54 points (p = 0.0009) to -0.00 (p = 1.00). The first strong result was the instrument agreeing with itself.
This program has met this failure before
The interesting thing is not that it happened. It is that it is a recurring species, documented repeatedly in this body of work, with the same shape each time: a number that reads as structure turns out to be carried by light, or by a normalisation that manufactures the signal it reports.
| When | What was claimed | What it turned out to be | Source |
|---|---|---|---|
| Saltimbanques | A persistence-under-resolution heartbeat read composition | Falsified as a composition reading: persistence rank-orders by chiaroscuro | ZTOYBOX/Saltimbanques/CALIBRATION_RESULTS.md |
| The Crit kernel audit | Structural coordinates are comparable across images | The mask carries a polarity assumption, material = dark; light-on-dark inverts it, frame-outside-mask 0.894 to 0.132 | image-steering-accreditation/references/23_kernel_audit.md §7.2 |
| The Synth II | softRatio measured how much gradient fails to resolve into a boundary | Percentile thresholds fixed the counts, so it returned 0.778 for every image: “a constant wearing a measurement's clothes” | products/the-synth-ii/README.md |
| ISR, observation I-23 | A normalised evidence field was healthy | robust01 zeroes a field when support is sparse, and its mirror stretches numerical dust to full scale while reading perfectly healthy | isr-library/OBSERVATIONS.md |
| This study | Two clauses moved composition (the wrong way) | Entirely tone: the shift went from +2.54 DGI points to zero once tone was regressed out | this page |
The lesson is not to distrust the instruments. It is that a structural claim compared across images whose tone differs is not yet a structural claim, and the fix is cheap.
The control
Each steered image's luminance is remapped onto the distribution of its own plain twin. The mapping is monotone, so every pixel keeps its rank: placement, edges and arrangement survive untouched, and only the distribution of values changes. Both arms then receive identical processing, which matters more than it sounds: an earlier version of this analysis remapped only the steered arm, and manufactured a result that looked like steering narrowing the model.
“the figure placed well off-center, its weight answered by a large quiet passage of shadow”
Aimed at centroid offset, which moved -1.56 dice. Tone moved -4.07 dice.
DGI 34 · offset 0.052 · tone 3.30
DGI 54 · offset 0.026 · tone 1.72
the control: same placement, plain tone
“form built from broad masses of light and deep shadow, the darks modeling the structure”
Aimed at lightness contrast, which moved +1.15 dice. Tone moved -2.47 dice.
DGI 46 · offset 0.021 · tone 1.73
DGI 60 · offset 0.005 · tone 0.77
the control: same placement, plain tone
“the lit head set against an independent field of muted color”
Aimed at spotlight concentration, which moved +2.48 dice. Tone moved +0.52 dice.
DGI 53 · offset 0.020 · tone 1.49
DGI 40 · offset 0.052 · tone 1.70
the control: same placement, plain tone
5. What survives the control
With tone equalised in both arms, pooled across all clauses, every structural axis is statistically equivalent to zero within one dice, and the sample can detect about a quarter of a dice.
| Axis | Shift (dice) | 90% interval | Detectable at 80% |
|---|---|---|---|
| DGI | +0.01 | [-0.15, +0.18] | 0.28 |
| centroid offset | +0.27 | [+0.11, +0.44] | 0.28 |
| sector variation | +0.08 | [-0.11, +0.26] | 0.31 |
| dispersion | +0.51 | [+0.34, +0.67] | 0.28 |
| centering | +0.22 | [+0.05, +0.39] | 0.29 |
| symmetry | -0.08 | [-0.26, +0.11] | 0.31 |
| frame outside mask | -0.06 | [-0.29, +0.18] | 0.40 |
| inner mass fraction | -0.33 | [-0.49, -0.17] | 0.28 |
Four of those are nonetheless non-zero, which deserves saying plainly: the clauses do touch structure. They touch it faintly, at a quarter to a half of a dice, and not where they aimed. Mass economy asked for dispersion to fall and it rose. Centering rose rather than fell under clauses asking for off-centre weight. Per clause, on its own named axis, only the two tonal predictions survive: spotlight concentration (+0.79 dice) and lightness contrast (+0.75 dice).
What the words did move, measured in the same units, was tone: dark mass fraction rose 2.22 dice under the chiaroscuro clause; the off-centre clause made pictures 1.70 dice darker, obeying its own phrase about a large quiet passage of shadow while ignoring the placement instruction it was attached to. Those two figures come from the uncontrolled comparison, necessarily: the control equalises tone, so it would define the tonal effect out of existence. Tone is the one thing here that must be measured before the control, and structure the one thing that must be measured after it.
What this does not say
It does not say the model cannot make an off-centre portrait. It plainly can: the four variants of a single prompt already scatter across the placement axis, and section 7 shows that picking among them reaches further than any clause does. The capability is in the model. What failed is the attempt to address it in words. A user who wants a specific composition is not blocked; they are being asked to find it by generating and choosing rather than by instructing.
6. Nothing widened
The original aim was expansion: more territory, not a shifted average. With identical processing in both arms, no axis widened.
| Axis | Within-job range | p | Spread across sitters |
|---|---|---|---|
| DGI | ×0.85 | 0.017 | ×0.97 |
| centroid offset | ×1.06 | 0.43 | ×1.13 |
| sector variation | ×0.92 | 0.2 | ×0.92 |
| dispersion | ×0.91 | 0.17 | ×0.90 |
| centering | ×0.88 | 0.072 | ×0.88 |
| symmetry | ×0.92 | 0.2 | ×0.92 |
| frame outside mask | ×0.77 | 0.066 | ×0.95 |
| inner mass fraction | ×0.93 | 0.28 | ×0.99 |
7. Three levers, one scale
| Axis | Steering clause | Genre register | Best of four |
|---|---|---|---|
| frame outside mask | 0.06 | 0.92 | 1.62 |
| centroid offset | 0.27 | 0.47 | 1.39 |
| inner mass fraction | 0.33 | 0.60 | 1.29 |
| sector variation | 0.08 | 0.62 | 1.26 |
| dispersion | 0.51 | 0.64 | 1.12 |
| centering | 0.22 | 0.60 | 1.10 |
| DGI | 0.01 | 0.55 | 1.06 |
| symmetry | 0.08 | 0.62 | 1.06 |
Two results here are worth sitting with. First, the genre word beats the compositional instruction on every structural axis. "17th-century Dutch Golden Age" moves the geometry more than "the figure placed well off-center" does. Composition in this model appears to ride on genre priors rather than on spatial language. The register comparison is confounded with its sitters and subjects, so treat it as an upper bound, but the direction is consistent across all eight axes.
Second, selection beats writing. Keeping the best of four already-generated variants moves structure by 1.06 to 1.62 dice, several times any clause. It needs no prompt craft and no model access. Its price is throughput, and that price is calculable: the worst-case shift available from keeping a fraction of outputs is the superquantile gap of the measured distribution, which is the arithmetic developed in the companion curation study.
There is a catch, and it is the reason this result is useful rather than deflating. Selecting on an axis requires measuring that axis. Eye-balling four variants for "more off-centre weight" is exactly the judgement the dice comparison shows people are poor at calibrating. A selection lever needs an instrument in the loop, which is what The Crit and The Field are: not verdicts on quality, but the measurement that makes choosing repeatable.
8. A reference band: paintings
To know whether any of these differences are large, it helps to see a population that was not generated: 116 museum photographs and nine Cézannes, measured the same way, and tone-matched to the same reference.
| Axis | AI raw | Museum raw | AI controlled | Museum controlled | Human spread |
|---|---|---|---|---|---|
| DGI | 50.366 | 34.974 | 50.759 | 45.664 | ×1.48 |
| sector variation | 0.289 | 0.550 | 0.278 | 0.373 | ×1.73 |
| symmetry | 0.639 | 0.338 | 0.652 | 0.537 | ×1.65 |
| centroid offset | 0.023 | 0.044 | 0.022 | 0.025 | ×1.70 |
| frame outside mask | 0.256 | 0.585 | 0.252 | 0.320 | ×2.14 |
| inner mass fraction | 0.180 | 0.217 | 0.179 | 0.202 | ×1.93 |
DGI 12 · sector variation 1.01
DGI 56 · sector variation 0.19
Raw, the paintings sit 15 index points away from the generated portraits. Tone-controlled, five points, about 0.7 of an AI standard deviation. What survives the control is that paintings are about one standard deviation more asymmetric, and that they occupy 1.3 to 2.1 times more of every structural axis. The steered images sit on top of the defaults, not between them and the paintings. Note also what the control costs: tone is not a nuisance in a painting, it is a large part of the composition, so the tone-matched comparison answers a narrow question deliberately.
9. What to do with this
If you are testing a prompt trick
- Generate the same prompt several times and measure the spread. That is your denominator.
- Run the variant prompt on the same subjects, paired, and measure the difference on the axis you care about.
- Report the difference in dice. Under 1.0, say so.
- Write the prediction down first. Of eleven predictions here, the one that held was the one whose clause named a tonal outcome, and it would have been easy to find a supporting number afterwards among 99 columns.
If you are building the instrument
- Ask what fraction of your structural measure is luminance. If any, tone-control before comparing corpora.
- Apply identical processing to both arms, including the control itself.
- Report effects against the generator's variance, not in raw units, so the reader can tell a lever from a rounding error.
- Keep the failures in the same document as the findings. Every entry in the table in section 4 was found by the people who built the thing.
If you want to steer composition
Words moved tone reliably and geometry barely. That suggests the lever is elsewhere: layout conditioning or spatial control input, inpainting, crop, post-edit, or selection with a target. Selection is available today at a known price. The rest is an architecture question, not a prompt question.
10. What would overturn this
Stated so that it can be checked rather than argued with:
- Another generator, or another genre, where a compositional clause moves its named axis by more than one dice under the same paired design. That would make this a fact about Midjourney portraits, not about prompting.
- A tone-invariant structural instrument that finds the compositional effects the tone-thresholded one missed. The control used here removes tone from both arms; an instrument that never let tone in would be a stronger test, and the mask could be rebuilt on edges alone to try it.
- A larger sample. About 18 pairs per clause leaves effects below roughly 0.7 dice undetectable per clause. Effects between a quarter and a dice, which is where the surviving structural movement sits, would be resolved by a few hundred pairs.
- Clauses written as spatial constraints rather than painterly description: explicit fractions, sides, margins. These six were written in the language of painting, which may be the reason they landed on tone.
Limits
One generator, one genre, one prompt family. About 18 pairs per clause and 106 pooled; four variants per job. "No effect" here means smaller than roughly a quarter of a dice pooled, or 0.7 for a single clause: it does not mean zero. The eight generation batches may differ in date or model version, which is not recorded. Tone matching preserves rank order, not local contrast. The register comparison is confounded with subject. The museum set is mixed genre and consists of photographs of paintings, so framing and surface are measured too. No human judgement of these images was collected: every statement here is about measured coordinates, and none of it is a claim about whether any picture is good.
Method
Images decoded with Pillow to RGBA, longest side 512 pixels, bilinear, matching The Crit's analysis path.
Measurement in Node v24.13.1 importing products/the-crit/js/kernel.js (sha256 460b745c235f6c63) and
products/the-field/js/color-kernel.js (sha256 6eef90e7d1ab8a24) unmodified, 99 numeric columns per
image. Clause assignment from the pasted prompt text rather than the spreadsheet label, since three rows
disagree with their own prompt; ambiguous sitters dropped. Job means per axis, paired differences per sitter,
divided by the median within-job standard deviation. Tone control remaps luminance quantiles and applies the
resulting gain to RGB, against three references: the paired plain job, one corpus-wide reference for the register
comparison, and the AI reference for the paintings. Paired t-tests; Benjamini-Hochberg at 0.05 within each family
of tests; two one-sided tests against a one-dice margin for equivalence; Levene for spread; one-way ANOVA with
eta-squared for register; Pearson correlation and linear regression for the confound. Data, scripts and the full
result tables are in this folder; the underlying study is in ../steering-authority/. The
measurement pipeline, analysis and this page were built and run with an AI assistant under the author's
direction; every number here is reproducible from the scripts beside it.