Image Structure Reader · Study 10 · Critique, descriptive mode

The palette moves; the composition holds.A Midjourney corpus: forty-eight unselected draws of Little Red Riding Hood, measured on their own terms

Give Midjourney a fairy-tale scene and a paragraph of compositional direction, keep every draw it returns, and a pattern comes out of the measurements that is quieter than a single striking image would suggest. The direction reliably repaints the picture: it darkens the key, warms the colour into the light, shifts one painterly dimension of the whole surface, and it does this on almost every draw. It barely rebuilds the picture. Across forty-eight frames the compositional field keeps its shape: the structural mass stays in a tight band near centre no matter where the figure is sent, the field moves as a unit rather than counterbalancing, and the centred emblem this library watches for never forms. What steering owns is the surface; what the engine keeps is the arrangement. This is a composition study that happens to use a directed prompt as its instrument, so the question throughout is not whether the engine obeyed, but what its field does when pushed.

READ 2026-07-13/14
SOURCE Midjourney v8.1, user-generated, prompts on record
CORPUS 2 prompts × 3 aspects × 8 draws = 48 PNG, complete and unselected
RECORD lab book Entries 01–18, pre-registered
MODE descriptive critique (isr-critique)

Russell Parrish · Parallax Metrology · 2026. Every claim traces to the study record; grades and confirmatory/exploratory status live in the lab book. No claims about model internals, training, or intent (the RCP guard). "Base" and "Steered" are the maker's prompt labels; the instrument tested what the field does, not whether a directive was met. More info: www.parallaxmetrology.com

  • Instrument lineage: the ISR library: Matisse, Degas, Pollock, Caravaggio, Soejima, Klimt, Bruegel, Cartier-Bresson, Sora Pair, MidJourney Ensemble, Velazquez, and Malevich
  • Orientation, for a cold reader

    The corpus, the frame, and the scope

    The material is a small factorial: one plain prompt ("Little Red Riding Hood walking on a stair path in a forest, stone column and arches along a tunnel"), one directed prompt written in compositional vocabulary (displace the figure, discipline the void, harden the light, torque the pose, five refusal clauses), each rendered at three aspect ratios, eight draws apiece. The eight are two Midjourney grids of four, and every tile was kept: no selection, no edits, no re-rolls held back. That completeness is what turns this from an anecdote into a baseline, because it makes draw-to-draw variation a measured quantity rather than a caveat, and it lets the study separate what the direction does from what the dice do.

    The governing frame matters more here than in a single-image read, and it starts from a shift in what a prompt is. A prompt should not be viewed as specifying an image, but as perturbing a distribution from which images are sampled: the engine holds a prior over composition, the prompt nudges it, and what comes out is a draw from the perturbed distribution, not a rendered instruction. So the two cohorts are better named than "Base" and "Steered" (the maker's labels): they are the engine's native operating regime and a perturbed operating regime, and the study measures the difference in the dynamics of a system, not a compliance rate. "The figure went right in seventeen of twenty-four draws" is therefore a fact about the field's steerability, and the detection limits below (directives the pre-semantic instrument cannot resolve) cost nothing, because obedience was never the question. Three scope statements belong up front. The sample is complete but nested (two grids of four, so the eight draws per cell are not eight independent throws, and the study checks that clustering rather than assuming it away). The three aspect ratios are treated as a factor, never pooled blindly. And every effect below carries an interval: at this sample size the study was powered to detect a Cohen's d of about 0.8, and the bootstrap intervals rest on only twelve grid-clusters, so their widths are themselves approximate. Nothing here is a law about Midjourney; it is one prompt family on one version, measured carefully.

    The Base cohort, all twenty-four draws (rows are the three aspect ratios). Scene-mode: a s
    The Base cohort, all twenty-four draws (rows are the three aspect ratios). Scene-mode: a small figure on a stair, deep arch tunnels receding to a bright opening, ambient light. The figure roams from draw to draw; the architecture is prompt-supplied in both cohorts.
    The Steered cohort, all twenty-four draws, same layout. The style itself shifts to a softe
    The Steered cohort, all twenty-four draws, same layout. The style itself shifts to a softer painterly key; pose cues like the lantern and the backward glance recur across the sheet; the white-hot core is frequent, and often halo-adjacent, a centred glow the direction did not dispel.

    The default is a scene, not an emblem

    The basin this library watches for does not appear [stable]

    The library has a name for the shape a naive prompt often collapses into: the basin, a centred figure under a radial halo whose mass, void and gesture all settle toward one equilibrium. Midjourney, on this prompt, does not go there. The Base cohort reads as a scene: Default Gravity Index averages 50.3, the stance stands rather than torques in seventeen of twenty-four draws, and across all forty-eight frames not one reaches the radial-collapse threshold (no frame's salient structure conforms to concentric arcs past the classifier's floor; the maximum ring-conformity seen is 0.38 against a bar of 0.50). The deep one-point tunnels put their light at the end of a corridor, not wrapped around a subject, so the density bowls that the basin needs never form. The composition that recurs is instead the one this corpus knows from its deep-space paintings: recessional and quiet, a whole that cancels rather than emerges. Read against the program's spatial-priors work, the lesson is specific: the centred-emblematic basin is a conditional attractor, engine and prompt dependent, and this engine on this prompt lands in a scene attractor instead. The emblem is something a generator can produce, not something it must.

    What the direction moves

    One dimension of the surface, reliably; the arrangement, barely [stable]

    Set the two cohorts side by side in one coordinate space and the separation sorts cleanly by kind. The axes that move are tonal and chromatic: the steered cohort is darker (shadow mass 0.59 to 0.76), its colour welded harder to the light (coupling 0.09 to 0.27), its opponent-colour activity in the shadow void falling as the darks deepen (the measure the library calls beta: colour working where drawing is absent). The axes that barely move are the arrangement ones: overall gravity, torque, grid asymmetry, figure and mass placement all shift far less. A variance partition makes the proportions exact. Of the total spread on the colour axes, steering accounts for roughly a third (coupling 34%, shadow and the opponent measure about 33% each); of the spread on the compositional-regime axes it accounts for almost nothing (gravity 3.8%, torque 5.9%), with draw-to-draw noise owning the majority everywhere. And the several style metrics that lead the separation are not independent knobs; they intercorrelate so tightly (up to 0.89) that they are one latent painterly dimension seen from different sides. The direction predominantly affected a single latent painterly dimension. Put differently, it turned one dial, hard.

    The whole study in one figure: the share of each axis's variance owned by steering (gold),
    The whole study in one figure: the share of each axis's variance owned by steering (gold), aspect, their interaction, grid-clustering, and draw residual (grey). Steering is tall on the colour and tonal axes at left and a sliver on the arrangement axes at right, where draw noise takes over.

    Because every draw was kept, those proportions come with intervals, and the intervals sharpen the claim into three tiers. The four surface axes are robustly moved: their steering effect sizes sit with lower confidence bounds above 0.7, all at least medium-large. One stands alone: colour retreating from the void has an effect of 1.48 [1.12, 2.64], the only axis whose lower bound is itself a large effect, which makes it the single most certain structural consequence of the direction. The placement and cohesion axes form a middle tier, real but of uncertain size, their intervals excluding zero yet reaching down into small-effect territory. And the compositional-regime axes form a floor: overall gravity moves by 0.4 [0.03, 1.22], an interval that hugs zero and is indistinguishable from negligible at this power. So the honest headline is not "structure barely moved" asserted flatly, but a powered statement: the direction reliably repainted one dimension of the surface, marginally nudged placement, and left the compositional regime below what this design could detect.

    Steering effect per axis with 95% bootstrap intervals. The four style axes (gold) sit  rob
    Steering effect per axis with 95% bootstrap intervals. The four style axes (gold) sit robustly above the detection floor; the placement axes (blue) are real but reach into small-effect territory; the regime axes (red) hug zero. Colour retreating from the void (beta) is the one axis whose interval floor is itself a large effect.

    The field's window

    The mass stays near centre, and moves as a unit when pushed [stable]

    The composition question the ensemble can answer, which a single pair never could, is how far the field lets its mass wander. The answer is: not far. Total structural mass stays in a tight central band across all forty-eight draws, its horizontal offset averaging within a few hundredths of centre in both cohorts and rarely leaving a narrow window, even as the figure ranges across half the frame and the prompt pushes it outward. That narrowness is itself the finding, a tight operating window measured rather than asserted. Individual frames read as balanced (each carries a counterweight region, the state the classifier labels "counterweighted"), but that balance is not a response to the figure: when the field is pushed, it does not add a counterweight the way a composed picture would, it co-moves. Correlating the figure's position against the mass of everything except the figure, so that the figure is not being matched against a mass it belongs to, the surround shifts with the figure rather than against it (a positive relationship of about +0.49). The engine absorbs a displacement into a small, unified excursion of the whole field, and keeps the centre. One caveat scopes it precisely: Midjourney resists large figure excursion in the first place, so this reads the field in its low-excursion regime, and whether a compensating mechanism would engage under a harder push is not testable on a field that will not take the harder push.

    Left: the figure's position against the mass of everything except the figure; the surround
    Left: the figure's position against the mass of everything except the figure; the surround co-moves with the figure (a positive relationship) rather than counterbalancing it. Right: the figure-position distribution shifts rightward as one mode, not as a comply-or-ignore split.

    The frame shape is part of the composition

    Aspect conditions the field, and interacts with the direction [exploratory]

    Because the corpus was rendered at three aspect ratios, it can ask something the library had never varied: does the shape of the rectangle change the composition inside it? It does, and not as a simple main effect. The Base cohort is most basin-like at the wide 4:3 (gravity 54.2) and least at the tall 3:4 (46.1); the steered cohort reorders, peaking at the square. And the square, with the least lateral room, is where the figure is displaced hardest of all (mean offset 0.154, against about 0.05 in the wider frames), as if the direction's push had nowhere to spread and concentrated. Held to an exploratory grade, at eight draws per cell, this is the first measured evidence that frame shape conditions both the attractor a generator falls into and how far it can be steered out of it, which is squarely a composition finding rather than a prompt one.

    What the instrument cannot see, honestly

    Three ceilings, marked as ceilings [narrowed]

    Three items resolve as limits rather than results, and the study keeps them that way. The contract's spatial-discipline directives (a cleared right-side lane, a ban on particle fill) could not be cleanly measured: a chroma-gated detector built to isolate decorative bokeh from the overexposed light-core ended up counting backlit gaps in the foliage instead, because in a forest tunnel intentional bokeh and light-through-leaves are the same low-level feature, and telling them apart is a semantic distinction above a pre-semantic instrument's ceiling. (The overlay below shows the detector, in red, settling on backlit gaps in the foliage at the tunnel mouth rather than on decorative particles.) The nulls on the most symmetric frames confirm that no draw achieves figural address: the one strongly composition-dependent case builds its bilateral symmetry from an arch and a light-column, not from a figure facing the viewer. And the counterweight question can only be read in the field's low-excursion regime, as noted. None of these is a tooling failure; each is a boundary drawn where the instrument's pre-semantic design actually ends, which is worth more to the record than a forced verdict would be.

    The chroma-gated bokeh detector on one Steered frame, kept right-quadrant discs in red. Th
    The chroma-gated bokeh detector on one Steered frame, kept right-quadrant discs in red. The chroma gate removed the hot-core confound, but the detector then lands on backlit leaf-gaps at the tunnel opening: in a forest, bokeh and light-through-leaves are one low-level feature.

    The one clear surface change is the overexposed core, present far more strongly in the steered cohort, and it is a tonal change, not a spatial one. Which returns the essay to its measurement: the lever that works is the surface.

    Optional interpretation

    Clearly marked, downstream of the numbers [reading]

    Read as a whole, the corpus suggests that this generator holds a strong prior over composition and a weak one over surface, or at least that a paragraph of direction reaches the second far more easily than the first. You can repaint its forest, dependably and on nearly every draw; you can barely rearrange it. The compositional field has a narrow window it returns to, and it answers a shove not by building a counterweight but by shifting a little and keeping its centre. Whether that window is a virtue (a house sense of balance) or a limit (a resistance to real spatial argument) is not a question the instrument can grade, and it is left open. What the instrument can say is that the window is real, it is measured, and it is tight.

    The ensemble is the object

    Why forty-eight draws, and not one

    A single generated image is an anecdote; an ensemble exposes the operating characteristics of the generator. The quantities this study measured, variance, steerability, the interaction with aspect ratio, the consistency of response, cannot be inferred from a single exemplar, because they are properties of the distribution rather than of any one sample. That is the reason forty-eight draws were kept and none discarded: choosing one good output would have measured a taste, not an engine, and generating more would have refined the same properties rather than revealed different ones. The unit of analysis is not the image but the sample of engine behaviour.

    Which is the widest thing the corpus has to say. This study treats a generative model not as a machine that produces individual images, but as a stochastic system with measurable operating characteristics. Individual images are observations; the ensemble is the object of study. Steerability, variance, operating windows, and structural persistence are properties of the distribution, not of any single draw, and a directed prompt is best understood not as an instruction that an output obeys or fails, but as a perturbation whose effect is visible only across the distribution it shifts.

    What this does not prove

    The scope, stated plainly

    Not provedWhy
    Model internals or intentGeometry read against prompts, one engine version. No claims about training, sampling, or "obedience". "Base" and "Steered" are the maker's labels.
    A law about the engineOne prompt family, one version (v8.1), forty-eight frames. The variance baseline is real but local to this design; other prompts, versions, or subjects are separate questions.
    The basin's absence in generalThis prompt produced a scene, which scopes the basin as conditional here; it does not measure how often the emblem appears elsewhere.
    Counterbalancing does not happenOnly that it does not, in the field's low-excursion regime; the engine resists the large displacement under which a counterweight would be tested.
    Two contract itemsThe particle ban and breathing lane are unmeasured, not delivered-or-not; the distinction they need is above the pre-semantic ceiling.
    Precise variance sharesThe partition's percentages carry wide intervals (twelve grid-clusters); the fixed steering effects are better estimated, but even their interval widths are approximate.
    Fine structureUpscaled PNGs: no claims below upscaler scale. Cohort-level throughout; per-image reads are exemplar-only and pixel-checked.

    The record (Entries 01 to 18) carries the pre-registration with its scored misses, a nested-sampling check, a variance partition with bootstrap intervals, honest nulls of method, and a numeric auditor that reconciles every figure in the lab book against its source. Corrections are part of the record, not cleaned out of it.

    Sources & record

    Citations and artifacts