Image Generation — Attempt 1

Can a generator reproduce a composition it cannot read? Eight AI renders of the Family of Saltimbanques structure (different subjects), scored by the instrument.

Russell Parrish · Parallax Metrology / Saltimbanques · 2026. Renders: GPT image generation and MidJourney, first attempt. Scoring deterministic, reproducible (mjgpt_run.py).

Premise. The project's thesis is that generators cannot render multi-mass composition because they cannot read it. To test it we did the inverse of analysis: we wrote prompts that dictate the structure of Family of Saltimbanques — a clustered group, one small isolated counterweight across a void, psychological disconnection, an active empty ground — in different subject matter, generated eight first-attempt renders, and then scored each with the same instrument against the structural facts of the original: the void (field_void), the eccentric counterweight (balance_dependency), evidence-diversity (EDI), and what the masses are built from (fragile_to). Result: the renders capture the mood and the gist; the original remains alone in combining a genuinely empty field and a strong counterweight. The instrument cleanly separates structural reproduction from surface imitation — including one render that fools the eye and is caught by the measure.

1. Method — prompt the structure, not the scene

Generators default to centering the group, making figures interact and meet eyes, filling the background, and balancing symmetrically — every one of which destroys this composition. So the prompt strategy was to override those defaults with explicit structural directives (a reusable "compositional core" plus a Picasso-free style add-on), then vary only the subject (lunar crew, salt-flat travelers, etc.). The core directives:

The empty ground is the main subject — a vast, pale, featureless expanse filling ~60%. A group of 4–5 figures clustered LEFT and CENTER at varied heights, reading as one mass but not touching, not interacting. ONE smaller figure SEATED ALONE in the lower-RIGHT, separated by a wide gap; the deliberate counterweight — do not move it toward centre. No figure looks at any other or at the viewer; no event, no narrative; frozen, statuesque stillness. Render the group tonally; render the lone figure in a distinct saturated colour note. Dusky, desaturated palette, soft sourceless light, dry matte fresco-like surface. Forbid: centering, interaction, a busy background, symmetry, a storytelling scene.

Eight renders came back (two from GPT image generation, six from MidJourney). None was cherry-picked; these are first-attempt outputs.

original
The target structure. Picasso, Family of Saltimbanques (1905). The standing troupe upper-left; the seated woman small, apart, lower-right, across the void; figures co-present yet disconnected. The instrument reads this as field_void 0.176, EDI 0.217 (highest in our set), and the seated woman as the sole counterweight (balance_dependency +0.043).

2. Result — none reaches the corner

Scoring each render and the original on two axes — how empty the ground actually is (field_void) and how strong the counterweight is (balance_dependency, measured on persistence-stable regions) — the original sits alone at the high-right. The renders trade off: the emptiest has a weak, mis-placed counterweight; those that find a counterweight have busier grounds. No render achieves both at once.

summary scatter
Figure A. Eight renders (square = GPT, circle = MidJourney; cyan = a lower-right counterweight was found) against the original (orange star). The original uniquely combines a genuinely empty field with a strong counterweight. The renders cluster low-left; none reaches the corner.
Table 1. Structural scores. field_void = fraction of flat/empty ground; EDI = evidence-diversity; fragile_to = what the structure is built from; counterweight = max balance_dependency and its position (centred coords); "LR?" = is it a genuine lower-right counterweight. Original in gold; the structural success in cyan.
renderfield_voidEDIfragile_tocounterweightpositionLR?
★ original (Picasso)0.1760.217gamma_down+0.043(+0.42,+0.36)
GPT_0643440.0550.394grayscale+0.022(+0.42,+0.36)
GPT_0643090.0580.000gamma_down+0.015(−0.28,+0.03)
MJ_e63f9f_30.0520.283gamma_down+0.038(+0.20,+0.32)
MJ_e2e947_20.0120.438gamma_down+0.025(+0.36,+0.27)
MJ_f7a04f_00.0990.048contrast_norm+0.029(+0.29,+0.37)
MJ_3c6996_10.0040.362contrast_norm+0.086(+0.17,+0.01)
MJ_995882_10.2140.136contrast_norm+0.016(+0.11,+0.34)
MJ_995882_30.1550.106contrast_norm+0.008(−0.15,+0.21)

4 of 8 reproduced a genuine lower-right counterweight. The emptiest render (MJ_995882_1, field_void 0.214) put its strongest mass near centre — empty ground, but no eccentric balance. The original's combination is not matched.

3. The two that matter

Each render's diagnostic shows the render, the texture-robust Notan mass field it produces, and the balance overlay (cyan arrow = counterweight, red = heavy side; white cross = frame centre).

The structural success — GPT_064344

GPT 064344 diagnostic
Figure B. The counterweight lands at (+0.42, +0.36) — essentially the original woman's coordinates — with high EDI (0.394, competing constructions) and fragile_to=grayscale (the lone figure colour-built, the original's signature). This render got the structure, not just the look.

The instructive failure — GPT_064309

GPT 064309 diagnostic
Figure C. To the eye this is the most faithful render — dark tonal cluster upper-left, a rose seated figure lower-right, apparent emptiness between. The Notan field (centre) confirms the rose figure as a real mass. But the "void" is painterly-busy: field_void 0.058, a third of the original's, because the scrubbed mottled ground carries low-level tonal mass everywhere. So the masses do not cleanly separate, the support field fragments across the ground texture, and the counterweight cannot dominate (EDI 0.000; no clean lower-right counterweight). A surface imitation, not a structural one — the eye is fooled, the reader is not. This single image is the project's thesis in miniature.

4. The MidJourney renders

MidJourney captured the muted grisaille and the averted, disconnected gazes well, but tended to collapse the troupe into a tight vertical cluster with little void — the "group-photo" pull. Two of six found a lower-right counterweight.

MJ e63f9f_3
MJ_e63f9f_3 — counterweight found (+0.20,+0.32); decent diversity (EDI 0.283).
MJ e2e947_2
MJ_e2e947_2 — counterweight found (+0.36,+0.27); highest EDI of all (0.438) but almost no void (0.012).
MJ f7a04f_0
MJ_f7a04f_0 — counterweight found (+0.29,+0.37); more open ground (void 0.099).
MJ 3c6996_1
MJ_3c6996_1 — strong central mass, no eccentric counterweight; near-zero void (0.004) — the group-photo failure.
MJ 995882_1
MJ_995882_1 — the emptiest ground (void 0.214) but the counterweight sits near centre — void without balance.
MJ 995882_3
MJ_995882_3 — open ground (0.155), weak/centre-left mass; no eccentric counterweight.

5. Findings

(i) The gist is easy; the structure is hard. Every render captured the dusky palette, the stillness, the disconnection. None reproduced the original's combination of true emptiness and a strong eccentric counterweight — the thing that makes the composition work.

(ii) The instrument separates imitation from reproduction. GPT_064309 fools a human (it looks the most like the source) but its busy ground and unseparated masses are caught by field_void and the balance reading. GPT_064344, less obviously "Picasso," is the truer structural match (counterweight at the right coordinates, colour-built lone figure). The measure and the eye disagree, and the measure is right about structure.

(iii) The counterweight is the hardest target. 4/8 placed it; the rest drifted the strongest mass toward centre — exactly the predicted default. A genuine eccentric counterweight, small and far, is what generators resist most.

Honest methodology notes. (1) The counterweight was scored on persistence-stable regions, not the default consensus grain — on the default grain, subordinate masses drop below threshold (a documented blind spot, AUDIT §14–15) and the count undercounts (it gave 2/8 before correction, 4/8 after). (2) field_void is a robust, grain-independent measurement and carries the clearest signal here. (3) These are first-attempt renders from a single prompt strategy; this is Attempt 1, not a benchmark. (4) "Counterweight," "void," and the rest are measurements under the instrument's definitions — a render scored low has failed this structural test, which is a claim about reproducible structure, not about artistic worth.

6. Conclusion

Given explicit structural prompting, current generators reproduce the atmosphere of Family of Saltimbanques readily and its structure only partly — and they resist hardest at exactly the load-bearing element, the small eccentric counterweight. The instrument built to read the original turns out to be the right tool to grade the attempts: it sees through a render that fools the eye, and it locates the structural success that doesn't announce itself. The loop closes — read the painting, prompt the structure, render, and measure whether the render heard. Attempt 1: the gist, yes; the structure, not yet.

Per-render diagnostics, scores (mjgpt_scores.csv), and the summary figure: outputs/mjgpt/. Scoring code: mjgpt_run.py, balance.py, persistence_regions.py. Source renders are first-attempt GPT and MidJourney outputs.