Give Midjourney a fairy-tale scene and a paragraph of compositional direction, keep every draw it returns, and a pattern comes out of the measurements that is quieter than a single striking image would suggest. The direction reliably repaints the picture: it darkens the key, warms the colour into the light, shifts one painterly dimension of the whole surface, and it does this on almost every draw. It barely rebuilds the picture. Across forty-eight frames the compositional field keeps its shape: the structural mass stays in a tight band near centre no matter where the figure is sent, the field moves as a unit rather than counterbalancing, and the centred emblem this library watches for never forms. What steering owns is the surface; what the engine keeps is the arrangement. This is a composition study that happens to use a directed prompt as its instrument, so the question throughout is not whether the engine obeyed, but what its field does when pushed.
Russell Parrish · Parallax Metrology · 2026. Every claim traces to the study record; grades and confirmatory/exploratory status live in the lab book. No claims about model internals, training, or intent (the RCP guard). "Base" and "Steered" are the maker's prompt labels; the instrument tested what the field does, not whether a directive was met. More info: www.parallaxmetrology.com
The material is a small factorial: one plain prompt ("Little Red Riding Hood walking on a stair path in a forest, stone column and arches along a tunnel"), one directed prompt written in compositional vocabulary (displace the figure, discipline the void, harden the light, torque the pose, five refusal clauses), each rendered at three aspect ratios, eight draws apiece. The eight are two Midjourney grids of four, and every tile was kept: no selection, no edits, no re-rolls held back. That completeness is what turns this from an anecdote into a baseline, because it makes draw-to-draw variation a measured quantity rather than a caveat, and it lets the study separate what the direction does from what the dice do.
The governing frame matters more here than in a single-image read, and it starts from a shift in what a prompt is. A prompt should not be viewed as specifying an image, but as perturbing a distribution from which images are sampled: the engine holds a prior over composition, the prompt nudges it, and what comes out is a draw from the perturbed distribution, not a rendered instruction. So the two cohorts are better named than "Base" and "Steered" (the maker's labels): they are the engine's native operating regime and a perturbed operating regime, and the study measures the difference in the dynamics of a system, not a compliance rate. "The figure went right in seventeen of twenty-four draws" is therefore a fact about the field's steerability, and the detection limits below (directives the pre-semantic instrument cannot resolve) cost nothing, because obedience was never the question. Three scope statements belong up front. The sample is complete but nested (two grids of four, so the eight draws per cell are not eight independent throws, and the study checks that clustering rather than assuming it away). The three aspect ratios are treated as a factor, never pooled blindly. And every effect below carries an interval: at this sample size the study was powered to detect a Cohen's d of about 0.8, and the bootstrap intervals rest on only twelve grid-clusters, so their widths are themselves approximate. Nothing here is a law about Midjourney; it is one prompt family on one version, measured carefully.
The library has a name for the shape a naive prompt often collapses into: the basin, a centred figure under a radial halo whose mass, void and gesture all settle toward one equilibrium. Midjourney, on this prompt, does not go there. The Base cohort reads as a scene: Default Gravity Index averages 50.3, the stance stands rather than torques in seventeen of twenty-four draws, and across all forty-eight frames not one reaches the radial-collapse threshold (no frame's salient structure conforms to concentric arcs past the classifier's floor; the maximum ring-conformity seen is 0.38 against a bar of 0.50). The deep one-point tunnels put their light at the end of a corridor, not wrapped around a subject, so the density bowls that the basin needs never form. The composition that recurs is instead the one this corpus knows from its deep-space paintings: recessional and quiet, a whole that cancels rather than emerges. Read against the program's spatial-priors work, the lesson is specific: the centred-emblematic basin is a conditional attractor, engine and prompt dependent, and this engine on this prompt lands in a scene attractor instead. The emblem is something a generator can produce, not something it must.
Set the two cohorts side by side in one coordinate space and the separation sorts cleanly by kind. The axes that move are tonal and chromatic: the steered cohort is darker (shadow mass 0.59 to 0.76), its colour welded harder to the light (coupling 0.09 to 0.27), its opponent-colour activity in the shadow void falling as the darks deepen (the measure the library calls beta: colour working where drawing is absent). The axes that barely move are the arrangement ones: overall gravity, torque, grid asymmetry, figure and mass placement all shift far less. A variance partition makes the proportions exact. Of the total spread on the colour axes, steering accounts for roughly a third (coupling 34%, shadow and the opponent measure about 33% each); of the spread on the compositional-regime axes it accounts for almost nothing (gravity 3.8%, torque 5.9%), with draw-to-draw noise owning the majority everywhere. And the several style metrics that lead the separation are not independent knobs; they intercorrelate so tightly (up to 0.89) that they are one latent painterly dimension seen from different sides. The direction predominantly affected a single latent painterly dimension. Put differently, it turned one dial, hard.
Because every draw was kept, those proportions come with intervals, and the intervals sharpen the claim into three tiers. The four surface axes are robustly moved: their steering effect sizes sit with lower confidence bounds above 0.7, all at least medium-large. One stands alone: colour retreating from the void has an effect of 1.48 [1.12, 2.64], the only axis whose lower bound is itself a large effect, which makes it the single most certain structural consequence of the direction. The placement and cohesion axes form a middle tier, real but of uncertain size, their intervals excluding zero yet reaching down into small-effect territory. And the compositional-regime axes form a floor: overall gravity moves by 0.4 [0.03, 1.22], an interval that hugs zero and is indistinguishable from negligible at this power. So the honest headline is not "structure barely moved" asserted flatly, but a powered statement: the direction reliably repainted one dimension of the surface, marginally nudged placement, and left the compositional regime below what this design could detect.
The composition question the ensemble can answer, which a single pair never could, is how far the field lets its mass wander. The answer is: not far. Total structural mass stays in a tight central band across all forty-eight draws, its horizontal offset averaging within a few hundredths of centre in both cohorts and rarely leaving a narrow window, even as the figure ranges across half the frame and the prompt pushes it outward. That narrowness is itself the finding, a tight operating window measured rather than asserted. Individual frames read as balanced (each carries a counterweight region, the state the classifier labels "counterweighted"), but that balance is not a response to the figure: when the field is pushed, it does not add a counterweight the way a composed picture would, it co-moves. Correlating the figure's position against the mass of everything except the figure, so that the figure is not being matched against a mass it belongs to, the surround shifts with the figure rather than against it (a positive relationship of about +0.49). The engine absorbs a displacement into a small, unified excursion of the whole field, and keeps the centre. One caveat scopes it precisely: Midjourney resists large figure excursion in the first place, so this reads the field in its low-excursion regime, and whether a compensating mechanism would engage under a harder push is not testable on a field that will not take the harder push.
Because the corpus was rendered at three aspect ratios, it can ask something the library had never varied: does the shape of the rectangle change the composition inside it? It does, and not as a simple main effect. The Base cohort is most basin-like at the wide 4:3 (gravity 54.2) and least at the tall 3:4 (46.1); the steered cohort reorders, peaking at the square. And the square, with the least lateral room, is where the figure is displaced hardest of all (mean offset 0.154, against about 0.05 in the wider frames), as if the direction's push had nowhere to spread and concentrated. Held to an exploratory grade, at eight draws per cell, this is the first measured evidence that frame shape conditions both the attractor a generator falls into and how far it can be steered out of it, which is squarely a composition finding rather than a prompt one.
Three items resolve as limits rather than results, and the study keeps them that way. The contract's spatial-discipline directives (a cleared right-side lane, a ban on particle fill) could not be cleanly measured: a chroma-gated detector built to isolate decorative bokeh from the overexposed light-core ended up counting backlit gaps in the foliage instead, because in a forest tunnel intentional bokeh and light-through-leaves are the same low-level feature, and telling them apart is a semantic distinction above a pre-semantic instrument's ceiling. (The overlay below shows the detector, in red, settling on backlit gaps in the foliage at the tunnel mouth rather than on decorative particles.) The nulls on the most symmetric frames confirm that no draw achieves figural address: the one strongly composition-dependent case builds its bilateral symmetry from an arch and a light-column, not from a figure facing the viewer. And the counterweight question can only be read in the field's low-excursion regime, as noted. None of these is a tooling failure; each is a boundary drawn where the instrument's pre-semantic design actually ends, which is worth more to the record than a forced verdict would be.
The one clear surface change is the overexposed core, present far more strongly in the steered cohort, and it is a tonal change, not a spatial one. Which returns the essay to its measurement: the lever that works is the surface.
Read as a whole, the corpus suggests that this generator holds a strong prior over composition and a weak one over surface, or at least that a paragraph of direction reaches the second far more easily than the first. You can repaint its forest, dependably and on nearly every draw; you can barely rearrange it. The compositional field has a narrow window it returns to, and it answers a shove not by building a counterweight but by shifting a little and keeping its centre. Whether that window is a virtue (a house sense of balance) or a limit (a resistance to real spatial argument) is not a question the instrument can grade, and it is left open. What the instrument can say is that the window is real, it is measured, and it is tight.
A single generated image is an anecdote; an ensemble exposes the operating characteristics of the generator. The quantities this study measured, variance, steerability, the interaction with aspect ratio, the consistency of response, cannot be inferred from a single exemplar, because they are properties of the distribution rather than of any one sample. That is the reason forty-eight draws were kept and none discarded: choosing one good output would have measured a taste, not an engine, and generating more would have refined the same properties rather than revealed different ones. The unit of analysis is not the image but the sample of engine behaviour.
Which is the widest thing the corpus has to say. This study treats a generative model not as a machine that produces individual images, but as a stochastic system with measurable operating characteristics. Individual images are observations; the ensemble is the object of study. Steerability, variance, operating windows, and structural persistence are properties of the distribution, not of any single draw, and a directed prompt is best understood not as an instruction that an output obeys or fails, but as a perturbation whose effect is visible only across the distribution it shifts.
| Not proved | Why |
|---|---|
| Model internals or intent | Geometry read against prompts, one engine version. No claims about training, sampling, or "obedience". "Base" and "Steered" are the maker's labels. |
| A law about the engine | One prompt family, one version (v8.1), forty-eight frames. The variance baseline is real but local to this design; other prompts, versions, or subjects are separate questions. |
| The basin's absence in general | This prompt produced a scene, which scopes the basin as conditional here; it does not measure how often the emblem appears elsewhere. |
| Counterbalancing does not happen | Only that it does not, in the field's low-excursion regime; the engine resists the large displacement under which a counterweight would be tested. |
| Two contract items | The particle ban and breathing lane are unmeasured, not delivered-or-not; the distinction they need is above the pre-semantic ceiling. |
| Precise variance shares | The partition's percentages carry wide intervals (twelve grid-clusters); the fixed steering effects are better estimated, but even their interval widths are approximate. |
| Fine structure | Upscaled PNGs: no claims below upscaler scale. Cohort-level throughout; per-image reads are exemplar-only and pixel-checked. |
The record (Entries 01 to 18) carries the pre-registration with its scored misses, a nested-sampling check, a variance partition with bootstrap intervals, honest nulls of method, and a numeric auditor that reconciles every figure in the lab book against its source. Corrections are part of the record, not cleaned out of it.
PROVENANCE.md. Pre-registration (before any number): PREREGISTRATION.md.studies/AI-red-riding-MJ/: labbook.md + HTML twin (Entries 01–18), per-frame vectors, variance and bootstrap JSONs, contact sheets, exemplar readouts, and audit.py (the numeric reconciler). Register: toolkit/CANDIDATES.md (C-3, C-4 cohort-range rows).