We use spatial instruction compliance as a geometric probe to characterize the compositional attractor basin structure of MidJourney v6 and Recraft V4. A single neutral subject across eight spatial placement prompts. 48 outputs. The prompt defines the subject. The engine defines the world.
We use spatial instruction compliance as a geometric probe to characterize the compositional attractor basin structure of two generative image engines, MidJourney v6 and Recraft V4, using the Visual Thinking Lens (VTL) deterministic measurement framework. A single neutral subject (smooth river stone on wet dark sand) was generated across eight spatial placement prompts targeting cardinal and diagonal positions in a square frame, producing 48 total outputs (32 MidJourney, 16 Recraft). Mass displacement, radial void ratio, angular deviation, mass concentration, and centrality pressure were measured for each output. Results reveal engine-specific attractor basin signatures: both engines exhibit center-gravity pull, Recraft shows higher corner-instruction compliance (4/4 vs 2/4), and both fail a shared left-center instruction consistent with a directional training prior. A structural finding emerges from radial void ratio and angular deviation: the two engines interpret the same prompt as fundamentally different image types, and this prior scene model strongly conditions downstream spatial decisions. Compliance is the probe. Attractor basin geometry is the finding.
A reader looking at an image from this corpus might locate the stone visually and conclude: the stone is in the upper-left, the prompt asked for upper-left, therefore the model complied. The VTL kernel may disagree. The kernel does not detect objects. It measures the distribution of compositional energy across the image field — the aggregate of luminance gradients, edge density, tonal contrast, and spatial weight. The output is a mass centroid: a coordinate representing where the bulk of compositional energy lives. This is a fundamentally different quantity from object location.
Object location answers: where is the named subject? Mass centroid answers: where does the composition’s gravitational center reside? In a simple image with one subject on a plain background, these coincide. In a complex scene with competing elements — a horizon, atmospheric recession, tonal gradients, textural fields — they can diverge substantially.
Artists have prioritized mass placement over object placement for centuries, and the reason is structural. A painting is not a diagram with a labeled subject. It is a field of visual energy that the eye traverses according to forces the artist controls. Object placement tells the viewer what to look at. Mass placement tells the composition where to rest.
Strong compositions are built from decisions about weight, balance, and hierarchy that operate independently of semantic content. Which mass is largest, darkest, or most isolated determines the lead — the visual subject the eye returns to. These decisions happen at the level of tonal distribution and spatial weight, not at the level of named objects.
When a generative model receives a prompt asking for a stone in the upper-left corner, it is being asked to do two things simultaneously: place a named object and distribute compositional mass. A model that places the object correctly but distributes mass toward the center has technically obeyed the instruction while compositionally contradicting it. The eye lands on the stone briefly, then drifts to where the weight actually is. The composition is not about the stone.
The more specific failure mode: in several outputs in this corpus, particularly Recraft’s landscape scenes, the compositional lead has been transferred to something else entirely. The wave ripples in the wet sand carry more tonal variation than the stone. The horizon gradient creates a horizontal mass band the eye reads as primary. When the stone is placed in a corner, the void at center becomes the largest undivided field in the frame — and a large undivided field reads as compositional weight regardless of what the artist intended. The center void becomes the dominant compositional mass field.
The argument is not that any engine is wrong. It is that the tension is currently invisible. Without measurement, neither the model nor the person prompting it can know whether the composition is supporting the intended subject or quietly overriding it. The kernel makes that invisible process legible.
All compliance scores in this paper measure mass centroid direction, not object location. A prompt is scored as compliant when the mean displacement vector of the compositional mass field points in the instructed direction — regardless of where the named subject visually appears. A model that correctly places its subject but fails to align the compositional mass field will score as non-compliant. The kernel measures what the composition does, not what the subject does.
MidJourney’s flat overhead scene model produces images where object location and mass centroid are nearly identical. One subject on a uniform dark surface. Nothing else competes for compositional energy. When the stone moves, the centroid moves with it. The kernel and the eye agree because the scene is simple enough that they are measuring the same thing.
Recraft’s landscape scene model produces images where object location and mass centroid routinely diverge. The horizon introduces a luminance band. The atmospheric gradient distributes tonal weight vertically. A small stone correctly placed in the upper-left contributes less to the centroid than the combined luminance of sky, horizon, and ground plane beneath it.
This tradeoff has an unexpected implication. MidJourney’s reduced scene ontology becomes an advantage for certain positional compliance tasks. By collapsing the world into a flat surface, it removes the competing compositional forces that constrain Recraft. More world simulation produces more realistic outputs but reduces instruction flexibility. Less world simulation produces flatter outputs but more literal positional obedience. Neither is simply better. They are differently constrained.
A smooth river stone on wet dark sand was selected as the evaluation subject: no face, no obvious intrinsic orientation, no culturally encoded placement conventions, sufficient contrast for reliable mass detection. Eight prompts targeting four diagonal corners (P1–P4) and four axis-edge positions (P5–P8). Square format (1:1) specified for all generations. No prompt engineering, no iteration, no style modifiers.
Each output was passed through the VTL kernel extractor. Compliance was assessed by comparing the sign of mean displacement vectors against the expected direction. For corner prompts, both delta_x and delta_y signs must match. For edge prompts, only the relevant axis is scored.
Green labels = MidJourney Gold labels = Recraft — Same prompts, different scene ontologies.
Before reporting displacement data, a qualitative observation is necessary because it conditions all quantitative findings. MidJourney interpreted all eight prompts as a top-down texture study. Recraft interpreted all eight prompts as a landscape scene with a visible horizon line, atmospheric depth, and tonal recession. This is not a stylistic difference. It is a different compositional physics.
The same textual instruction does not define the same compositional task across engines. The geometric outputs are not independent of semantic priors. Semantic interpretation reshapes the compositional field itself.
| Metric | MidJourney | Recraft | Delta | n |
|---|---|---|---|---|
| Mean radial distance | 0.099 | 0.151 | +53% | MJ=32, RC=16 |
| Mean mu (mass concentration) | 0.141 | 0.203 | +44% | MJ=32, RC=16 |
| Mean theta (angular deviation) | 0.006 | 0.066 | +1000% | MJ=32, RC=16 |
| Mean r_v (void ratio) | 0.837 | 0.936 | +12% | MJ=32, RC=16 |
| Centrality pressure | 0.473 | 0.502 | +6% | MJ=32, RC=16 |
| 8-prompt compliance | 4/8 | 6/8 | +25pp | 8 prompts each |
MidJourney theta values are consistently near zero (range 0.0003 to 0.050, mean 0.006) — the signature of a flat overhead view with no scene geometry to introduce diagonal structure. Recraft theta values are substantially higher and more variable (range 0.008 to 0.166, mean 0.066), reflecting the diagonal gradient introduced by the landscape horizon.
Radial void ratio and angular deviation provide a kernel-level signature of scene interpretation without semantic analysis. MidJourney’s near-zero theta and high-uniform r_v are the geometric fingerprint of a top-down flat surface. Recraft’s elevated theta and variable r_v are the fingerprint of a horizon-containing landscape. The scene model is measurable.
Both engines fail P5 (left-center placement). Both comply with P6 (right-center). The asymmetry suggests a directional prior shared across both systems, potentially reflecting statistical regularities in photographic or design training distributions. Whether this reflects training data, architectural bias, or another mechanism is not determined by this corpus. The causal mechanism remains a hypothesis.
MidJourney succeeds at top-center (P7) but fails bottom-center (P8). Recraft succeeds at bottom-center (P8) but fails top-center (P7). The crossover is explained by the scene model. Recraft’s landscape prior anchors subjects to the lower half. MidJourney’s flat overhead prior has no vertical gravity. Failure patterns are mirror opposites.
Spatial instruction compliance depends not only on positional understanding but on the engine’s prior interpretation of what type of image is being generated. A prompt defines the subject. The engine defines the world. These are different operations and both must be measured.
Outlined points show expected positions at normalized radius 0.20. Filled points show actual mean displacement. Y-axis inverted: negative delta_y = upper frame.
MJ P6: stone at right edge, high mu = tight isolation on flat surface. RC P6/P8: correct direction, lower mu = scene complexity distributes mass.
The central finding is that Recraft and MidJourney are not solving the same compositional problem. MidJourney interprets the prompt as a 2D subject placement task. Recraft interprets it as a scene-building task. Comparing spatial instruction compliance between them without accounting for this prior is like comparing two navigators who received the same destination but are driving different road networks.
Standard generative image benchmarks assess whether the output contains the requested subject. This study shows that containment is a necessary but insufficient metric. An engine can correctly generate a stone in every output while systematically failing to place it where instructed. The failure is invisible to semantic evaluation and visible only to structural measurement.
A deeper implication follows: the geometric outputs are not independent of semantic priors. Semantic interpretation reshapes the compositional field itself. The kernels in this study do not detect whether the output contains a stone. They detect the resulting field organization after the engine has decided what kind of image it is making. That decision precedes placement and substantially constrains it.
“These findings suggest that compositional evaluation cannot assume a neutral image manifold shared across engines. The prompt defines the subject, but the engine defines the world.”
n=2 Recraft outputs per prompt (n=16 total). Mean positions computed from two outputs are statistically fragile. Conclusions about Recraft directional behavior should be treated as preliminary. At n=8 outputs per prompt per engine, confidence intervals narrow sufficiently to distinguish genuine compliance from noise at p<0.05 using a one-sample sign test. The stone-on-sand subject introduced a scene interpretation confound not anticipated at study design. Future experiments should test multiple subjects across multiple scene types.
All prompts submitted without style modifiers, seeds, or iteration. Square format (1:1) specified. P1–P4 target diagonal corners; P5–P8 target axis-edge positions.
Horizontal and vertical offset of compositional mass centroid from geometric center. Primary axis for detecting attractor behavior. Negative dy = upper frame.
Proportion of compositional space empty relative to the subject mass. High values indicate flat overhead compositions; variable values indicate scene complexity.
How compressed the marks are. Measures spatial concentration of compositional elements.
Compositional coherence (0 = diffuse, 1 = single point). High mu on flat-surface images; lower mu on scene images with distributed compositional elements.
How hard the edges pull against center. Measures resistance to the central attractor.
Departure of compositional mass from horizontal axis. Near-zero in flat overhead views (MJ: mean 0.006); elevated in horizon-containing landscapes (RC: mean 0.066).
Surface depth and layering. Measures how models build spatial complexity above a base plane.
Fraction of mass within central 50% of frame. Complements displacement: measures how much stays near center regardless of direction.