Smaller than the dice

What prompt steering actually moves, and why the instrument agreed with it at first.

Parallax Metrology · 12 September 2026 · measurements on 927 generated portraits and 125 paintings

A prompt is offered as a control surface. Add a clause about composition and the picture is supposed to change in that way. The advice circulates in exactly that form, as a list of phrases that work, and it is almost never accompanied by the one number that would make it checkable: how much the picture changes when you change nothing at all.

Generators are stochastic. Run the same prompt four times and you get four different pictures. Any honest claim about a prompt has to clear that variation first. This piece proposes that variation as the unit of measurement, calls it the dice, and uses it on a steering experiment that was already sitting on disk. It also documents the part that nearly fooled the measurement: the first, strongest result was an artifact of the instrument, and the instrument had failed this way before.

The short version. Measured against the dice, six compositional steering clauses moved composition by a quarter to a half of a dice, in directions they did not ask for. The same clauses moved tone by one to two dice. The apparent structural effect in the first pass was the mask reacting to darkness. A genre word already in the prompt moves composition more than a compositional instruction does, and simply keeping the best of four variants moves it more than either.

1. The missing denominator

Here are ten prompts from the corpus. Each has four variants, generated together, from identical text. The black bar is the average of the four.

The same prompt, four times. On the Default Gravity Index, single prompts produce variants from about 30 to 63. On centroid offset, the placement of the picture's mass, the four variants of one prompt routinely differ by more than the average difference between prompts.

This is the sampler's own spread. One dice is the median standard deviation among the four variants of a prompt, computed per axis. It is the amount a user already gets for free, by pressing the button again. For the Default Gravity Index, one dice is 7.8 points. For centroid offset it is 0.0169 of the frame width.

There is nothing new in the statistics. This is an effect size, a difference divided by a standard deviation, and the only choice being made is which standard deviation to divide by. Psychology divides by the spread between people. Here the relevant spread is the one the machine produces on its own, because that is what the user is competing with. Naming it and quoting it routinely is the whole proposal.

The rule that follows is uncomfortable and simple: an effect under one dice is weaker than re-rolling. It may still be real. It is not yet a control surface, because a user cannot reach it reliably without generating repeatedly, and if they are generating repeatedly they have a stronger lever available anyway (section 7).

2. A test that was already on disk

The corpus was generated before this analysis, for a different purpose, which makes it a better test than anything designed afterwards. 112 sitters across eight painting registers were each rendered twice by Midjourney: once with a plain prompt, and once with the identical prompt plus one steering clause. Four variants per job, 927 images. The stated aim in the originating spreadsheet was to test whether steering could expand the model's operating volume.

ClauseText appended to the prompt
off-center mass“the figure placed well off-center, its weight answered by a large quiet passage of shadow”
chiaroscuro / value-structure“form built from broad masses of light and deep shadow, the darks modeling the structure”
color autonomy“cool shadows set against warm half-lights, the color departing from the light”
environmental asymmetry“asymmetric composition, the sitter turned with open gaze-room to one side”
mass economy“constructed from a few large coherent value shapes rather than even all-over detail”
tonal opposition“the lit head set against an independent field of muted color”

Each clause names something a painter would recognise, and each maps onto an axis the instruments already measure. Before any number was computed, eleven predictions were written down, each naming an axis and a direction, plus one global prediction: if steering escapes the model's habitual composition, the Default Gravity Index should fall. Measurement used the kernels behind The Crit and The Field, imported unmodified, on images decoded exactly as the apps decode them.

3. First result: one in eleven

Each clause against the axis it names. Orange is the first pass. Blue is after the control described in the next section. The grey band is one dice wide in each direction: inside it, the clause is weaker than pressing the button again.

One prediction held on the first pass: the clause asking for a lit head against an independent field raised spotlight concentration by +0.72 dice. The compositional predictions did not hold. Two of them appeared to move in the wrong direction, and that is the interesting part, because two clauses asking for off-centre weight and asymmetry produced images that measured as more centred and more settled: the Default Gravity Index rose by 2.54 points (p = 0.0009).

4. The instrument moved with the light

A result that contradicts the instruction is worth more scrutiny than one that confirms it. The Crit's material mask, the thing that decides which pixels count as substance before any geometry is computed, is built like this:

score = 0.65 × normalised edge magnitude + 0.35 × (1 − normalised luminance), thresholded by Otsu, then opened and closed.

Thirty-five per cent of that score is darkness. Change a picture's tone and the mask changes; change the mask and the centroid, the dispersion, the sector balance and the index built on top of them all change with it. The steering clauses did change tone, because most of them mention shadow.

Left: how much each clause changed tone, against how much it changed the Default Gravity Index, one point per sitter. Right: every image in the corpus, its lightness against the share of the frame its mask claims.

The correlation between the tone change and the index change is -0.49. Across the whole corpus, mask coverage against image lightness is -0.62. Regress the tone change out of the index change and the entire effect disappears: from +2.54 points (p = 0.0009) to -0.00 (p = 1.00). The first strong result was the instrument agreeing with itself.

This program has met this failure before

The interesting thing is not that it happened. It is that it is a recurring species, documented repeatedly in this body of work, with the same shape each time: a number that reads as structure turns out to be carried by light, or by a normalisation that manufactures the signal it reports.

WhenWhat was claimedWhat it turned out to beSource
SaltimbanquesA persistence-under-resolution heartbeat read compositionFalsified as a composition reading: persistence rank-orders by chiaroscuroZTOYBOX/Saltimbanques/CALIBRATION_RESULTS.md
The Crit kernel auditStructural coordinates are comparable across imagesThe mask carries a polarity assumption, material = dark; light-on-dark inverts it, frame-outside-mask 0.894 to 0.132image-steering-accreditation/references/23_kernel_audit.md §7.2
The Synth IIsoftRatio measured how much gradient fails to resolve into a boundaryPercentile thresholds fixed the counts, so it returned 0.778 for every image: “a constant wearing a measurement's clothes”products/the-synth-ii/README.md
ISR, observation I-23A normalised evidence field was healthyrobust01 zeroes a field when support is sparse, and its mirror stretches numerical dust to full scale while reading perfectly healthyisr-library/OBSERVATIONS.md
This studyTwo clauses moved composition (the wrong way)Entirely tone: the shift went from +2.54 DGI points to zero once tone was regressed outthis page

The lesson is not to distrust the instruments. It is that a structural claim compared across images whose tone differs is not yet a structural claim, and the fix is cheap.

The control

Each steered image's luminance is remapped onto the distribution of its own plain twin. The mapping is monotone, so every pixel keeps its rank: placement, edges and arrangement survive untouched, and only the distribution of values changes. Both arms then receive identical processing, which matters more than it sounds: an earlier version of this analysis remapped only the steered arm, and manufactured a result that looked like steering narrowing the model.

portrait of a French courtier 18th-century French Rococo
Steering attempted: off-center mass
“the figure placed well off-center, its weight answered by a large quiet passage of shadow”
Aimed at centroid offset, which moved -1.56 dice. Tone moved -4.07 dice.
the plain prompt
DGI 34 · offset 0.052 · tone 3.30
plus the steering clause
DGI 54 · offset 0.026 · tone 1.72
steered, tone matched to the plain one
the control: same placement, plain tone
portrait of a Medici courtier Italian Renaissance oil pai
Steering attempted: chiaroscuro / value-structure
“form built from broad masses of light and deep shadow, the darks modeling the structure”
Aimed at lightness contrast, which moved +1.15 dice. Tone moved -2.47 dice.
the plain prompt
DGI 46 · offset 0.021 · tone 1.73
plus the steering clause
DGI 60 · offset 0.005 · tone 0.77
steered, tone matched to the plain one
the control: same placement, plain tone
portrait of a New England minister early American colonia
Steering attempted: tonal opposition
“the lit head set against an independent field of muted color”
Aimed at spotlight concentration, which moved +2.48 dice. Tone moved +0.52 dice.
the plain prompt
DGI 53 · offset 0.020 · tone 1.49
plus the steering clause
DGI 40 · offset 0.052 · tone 1.70
steered, tone matched to the plain one
the control: same placement, plain tone

5. What survives the control

With tone equalised in both arms, pooled across all clauses, every structural axis is statistically equivalent to zero within one dice, and the sample can detect about a quarter of a dice.

AxisShift (dice)90% intervalDetectable at 80%
DGI+0.01[-0.15, +0.18]0.28
centroid offset+0.27[+0.11, +0.44]0.28
sector variation+0.08[-0.11, +0.26]0.31
dispersion+0.51[+0.34, +0.67]0.28
centering+0.22[+0.05, +0.39]0.29
symmetry-0.08[-0.26, +0.11]0.31
frame outside mask-0.06[-0.29, +0.18]0.40
inner mass fraction-0.33[-0.49, -0.17]0.28

Four of those are nonetheless non-zero, which deserves saying plainly: the clauses do touch structure. They touch it faintly, at a quarter to a half of a dice, and not where they aimed. Mass economy asked for dispersion to fall and it rose. Centering rose rather than fell under clauses asking for off-centre weight. Per clause, on its own named axis, only the two tonal predictions survive: spotlight concentration (+0.79 dice) and lightness contrast (+0.75 dice).

What the words did move, measured in the same units, was tone: dark mass fraction rose 2.22 dice under the chiaroscuro clause; the off-centre clause made pictures 1.70 dice darker, obeying its own phrase about a large quiet passage of shadow while ignoring the placement instruction it was attached to. Those two figures come from the uncontrolled comparison, necessarily: the control equalises tone, so it would define the tonal effect out of existence. Tone is the one thing here that must be measured before the control, and structure the one thing that must be measured after it.

What this does not say

It does not say the model cannot make an off-centre portrait. It plainly can: the four variants of a single prompt already scatter across the placement axis, and section 7 shows that picking among them reaches further than any clause does. The capability is in the model. What failed is the attempt to address it in words. A user who wants a specific composition is not blocked; they are being asked to find it by generating and choosing rather than by instructing.

6. Nothing widened

The original aim was expansion: more territory, not a shifted average. With identical processing in both arms, no axis widened.

AxisWithin-job rangepSpread across sitters
DGI×0.850.017×0.97
centroid offset×1.060.43×1.13
sector variation×0.920.2×0.92
dispersion×0.910.17×0.90
centering×0.880.072×0.88
symmetry×0.920.2×0.92
frame outside mask×0.770.066×0.95
inner mass fraction×0.930.28×0.99

7. Three levers, one scale

What moves structure, all in dice. The genre register is the period phrase already present in every prompt, measured on tone-matched images. Best of four is simply keeping the variant that scores highest, with no prompt change at all.
AxisSteering clauseGenre registerBest of four
frame outside mask0.060.921.62
centroid offset0.270.471.39
inner mass fraction0.330.601.29
sector variation0.080.621.26
dispersion0.510.641.12
centering0.220.601.10
DGI0.010.551.06
symmetry0.080.621.06

Two results here are worth sitting with. First, the genre word beats the compositional instruction on every structural axis. "17th-century Dutch Golden Age" moves the geometry more than "the figure placed well off-center" does. Composition in this model appears to ride on genre priors rather than on spatial language. The register comparison is confounded with its sitters and subjects, so treat it as an upper bound, but the direction is consistent across all eight axes.

Second, selection beats writing. Keeping the best of four already-generated variants moves structure by 1.06 to 1.62 dice, several times any clause. It needs no prompt craft and no model access. Its price is throughput, and that price is calculable: the worst-case shift available from keeping a fraction of outputs is the superquantile gap of the measured distribution, which is the arithmetic developed in the companion curation study.

There is a catch, and it is the reason this result is useful rather than deflating. Selecting on an axis requires measuring that axis. Eye-balling four variants for "more off-centre weight" is exactly the judgement the dice comparison shows people are poor at calibrating. A selection lever needs an instrument in the loop, which is what The Crit and The Field are: not verdicts on quality, but the measurement that makes choosing repeatable.

8. A reference band: paintings

To know whether any of these differences are large, it helps to see a population that was not generated: 116 museum photographs and nine Cézannes, measured the same way, and tone-matched to the same reference.

Default Gravity Index, before and after the control. Raw, the gap between generated portraits and paintings looks enormous. Most of it is tone.
AxisAI rawMuseum rawAI controlledMuseum controlledHuman spread
DGI50.36634.97450.75945.664×1.48
sector variation0.2890.5500.2780.373×1.73
symmetry0.6390.3380.6520.537×1.65
centroid offset0.0230.0440.0220.025×1.70
frame outside mask0.2560.5850.2520.320×2.14
inner mass fraction0.1800.2170.1790.202×1.93
a painting
DGI 12 · sector variation 1.01
a generated portrait
DGI 56 · sector variation 0.19

Raw, the paintings sit 15 index points away from the generated portraits. Tone-controlled, five points, about 0.7 of an AI standard deviation. What survives the control is that paintings are about one standard deviation more asymmetric, and that they occupy 1.3 to 2.1 times more of every structural axis. The steered images sit on top of the defaults, not between them and the paintings. Note also what the control costs: tone is not a nuisance in a painting, it is a large part of the composition, so the tone-matched comparison answers a narrow question deliberately.

9. What to do with this

If you are testing a prompt trick

  1. Generate the same prompt several times and measure the spread. That is your denominator.
  2. Run the variant prompt on the same subjects, paired, and measure the difference on the axis you care about.
  3. Report the difference in dice. Under 1.0, say so.
  4. Write the prediction down first. Of eleven predictions here, the one that held was the one whose clause named a tonal outcome, and it would have been easy to find a supporting number afterwards among 99 columns.

If you are building the instrument

  1. Ask what fraction of your structural measure is luminance. If any, tone-control before comparing corpora.
  2. Apply identical processing to both arms, including the control itself.
  3. Report effects against the generator's variance, not in raw units, so the reader can tell a lever from a rounding error.
  4. Keep the failures in the same document as the findings. Every entry in the table in section 4 was found by the people who built the thing.

If you want to steer composition

Words moved tone reliably and geometry barely. That suggests the lever is elsewhere: layout conditioning or spatial control input, inpainting, crop, post-edit, or selection with a target. Selection is available today at a known price. The rest is an architecture question, not a prompt question.

10. What would overturn this

Stated so that it can be checked rather than argued with:

The protocol, in four lines. Generate the same prompt n times and record the spread of the measure you care about: that is one dice. Run the variant prompt on the same subjects, paired. Report the difference divided by the dice. If it is under 1.0, say so in the sentence that reports it.

Limits

One generator, one genre, one prompt family. About 18 pairs per clause and 106 pooled; four variants per job. "No effect" here means smaller than roughly a quarter of a dice pooled, or 0.7 for a single clause: it does not mean zero. The eight generation batches may differ in date or model version, which is not recorded. Tone matching preserves rank order, not local contrast. The register comparison is confounded with subject. The museum set is mixed genre and consists of photographs of paintings, so framing and surface are measured too. No human judgement of these images was collected: every statement here is about measured coordinates, and none of it is a claim about whether any picture is good.

Method

Images decoded with Pillow to RGBA, longest side 512 pixels, bilinear, matching The Crit's analysis path. Measurement in Node v24.13.1 importing products/the-crit/js/kernel.js (sha256 460b745c235f6c63) and products/the-field/js/color-kernel.js (sha256 6eef90e7d1ab8a24) unmodified, 99 numeric columns per image. Clause assignment from the pasted prompt text rather than the spreadsheet label, since three rows disagree with their own prompt; ambiguous sitters dropped. Job means per axis, paired differences per sitter, divided by the median within-job standard deviation. Tone control remaps luminance quantiles and applies the resulting gain to RGB, against three references: the paired plain job, one corpus-wide reference for the register comparison, and the AI reference for the paintings. Paired t-tests; Benjamini-Hochberg at 0.05 within each family of tests; two one-sided tests against a one-dice margin for equivalence; Levene for spread; one-way ANOVA with eta-squared for register; Pearson correlation and linear regression for the confound. Data, scripts and the full result tables are in this folder; the underlying study is in ../steering-authority/. The measurement pipeline, analysis and this page were built and run with an AI assistant under the author's direction; every number here is reproducible from the scripts beside it.