Can a generator reproduce a composition it cannot read? Eight AI renders of the Family of Saltimbanques structure (different subjects), scored by the instrument.
field_void), the eccentric counterweight (balance_dependency), evidence-diversity (EDI), and what the masses are built from (fragile_to). Result: the renders capture the mood and the gist; the original remains alone in combining a genuinely empty field and a strong counterweight. The instrument cleanly separates structural reproduction from surface imitation — including one render that fools the eye and is caught by the measure.Generators default to centering the group, making figures interact and meet eyes, filling the background, and balancing symmetrically — every one of which destroys this composition. So the prompt strategy was to override those defaults with explicit structural directives (a reusable "compositional core" plus a Picasso-free style add-on), then vary only the subject (lunar crew, salt-flat travelers, etc.). The core directives:
The empty ground is the main subject — a vast, pale, featureless expanse filling ~60%. A group of 4–5 figures clustered LEFT and CENTER at varied heights, reading as one mass but not touching, not interacting. ONE smaller figure SEATED ALONE in the lower-RIGHT, separated by a wide gap; the deliberate counterweight — do not move it toward centre. No figure looks at any other or at the viewer; no event, no narrative; frozen, statuesque stillness. Render the group tonally; render the lone figure in a distinct saturated colour note. Dusky, desaturated palette, soft sourceless light, dry matte fresco-like surface. Forbid: centering, interaction, a busy background, symmetry, a storytelling scene.
Eight renders came back (two from GPT image generation, six from MidJourney). None was cherry-picked; these are first-attempt outputs.
field_void 0.176, EDI 0.217 (highest in our set), and the seated woman as the sole counterweight (balance_dependency +0.043).Scoring each render and the original on two axes — how empty the ground actually is (field_void) and how strong the counterweight is (balance_dependency, measured on persistence-stable regions) — the original sits alone at the high-right. The renders trade off: the emptiest has a weak, mis-placed counterweight; those that find a counterweight have busier grounds. No render achieves both at once.
| render | field_void | EDI | fragile_to | counterweight | position | LR? |
|---|---|---|---|---|---|---|
| ★ original (Picasso) | 0.176 | 0.217 | gamma_down | +0.043 | (+0.42,+0.36) | ✓ |
| GPT_064344 | 0.055 | 0.394 | grayscale | +0.022 | (+0.42,+0.36) | ✓ |
| GPT_064309 | 0.058 | 0.000 | gamma_down | +0.015 | (−0.28,+0.03) | ✗ |
| MJ_e63f9f_3 | 0.052 | 0.283 | gamma_down | +0.038 | (+0.20,+0.32) | ✓ |
| MJ_e2e947_2 | 0.012 | 0.438 | gamma_down | +0.025 | (+0.36,+0.27) | ✓ |
| MJ_f7a04f_0 | 0.099 | 0.048 | contrast_norm | +0.029 | (+0.29,+0.37) | ✓ |
| MJ_3c6996_1 | 0.004 | 0.362 | contrast_norm | +0.086 | (+0.17,+0.01) | ✗ |
| MJ_995882_1 | 0.214 | 0.136 | contrast_norm | +0.016 | (+0.11,+0.34) | ✗ |
| MJ_995882_3 | 0.155 | 0.106 | contrast_norm | +0.008 | (−0.15,+0.21) | ✗ |
4 of 8 reproduced a genuine lower-right counterweight. The emptiest render (MJ_995882_1, field_void 0.214) put its strongest mass near centre — empty ground, but no eccentric balance. The original's combination is not matched.
Each render's diagnostic shows the render, the texture-robust Notan mass field it produces, and the balance overlay (cyan arrow = counterweight, red = heavy side; white cross = frame centre).
fragile_to=grayscale (the lone figure colour-built, the original's signature). This render got the structure, not just the look.field_void 0.058, a third of the original's, because the scrubbed mottled ground carries low-level tonal mass everywhere. So the masses do not cleanly separate, the support field fragments across the ground texture, and the counterweight cannot dominate (EDI 0.000; no clean lower-right counterweight). A surface imitation, not a structural one — the eye is fooled, the reader is not. This single image is the project's thesis in miniature.MidJourney captured the muted grisaille and the averted, disconnected gazes well, but tended to collapse the troupe into a tight vertical cluster with little void — the "group-photo" pull. Two of six found a lower-right counterweight.
(i) The gist is easy; the structure is hard. Every render captured the dusky palette, the stillness, the disconnection. None reproduced the original's combination of true emptiness and a strong eccentric counterweight — the thing that makes the composition work.
(ii) The instrument separates imitation from reproduction. GPT_064309 fools a human (it looks the most like the source) but its busy ground and unseparated masses are caught by field_void and the balance reading. GPT_064344, less obviously "Picasso," is the truer structural match (counterweight at the right coordinates, colour-built lone figure). The measure and the eye disagree, and the measure is right about structure.
(iii) The counterweight is the hardest target. 4/8 placed it; the rest drifted the strongest mass toward centre — exactly the predicted default. A genuine eccentric counterweight, small and far, is what generators resist most.
field_void is a robust, grain-independent measurement and carries the clearest signal here. (3) These are first-attempt renders from a single prompt strategy; this is Attempt 1, not a benchmark. (4) "Counterweight," "void," and the rest are measurements under the instrument's definitions — a render scored low has failed this structural test, which is a claim about reproducible structure, not about artistic worth.Given explicit structural prompting, current generators reproduce the atmosphere of Family of Saltimbanques readily and its structure only partly — and they resist hardest at exactly the load-bearing element, the small eccentric counterweight. The instrument built to read the original turns out to be the right tool to grade the attempts: it sees through a render that fools the eye, and it locates the structural success that doesn't announce itself. The loop closes — read the painting, prompt the structure, render, and measure whether the render heard. Attempt 1: the gist, yes; the structure, not yet.