A prompt is often described as steering an image. This study measures how much of a picture's geometry a prompt actually sets on one modern generator, what the sampler's draw keeps for itself, and what happens to those answers when the measuring instrument changes. It began as a bet that the draw decides more than the prompt. That bet lost, and the losing produced a usable map.
- Setup: engine, corpus, instruments, protocol
- What the prompt controls
- How the model satisfies a position word
- The default composition
- Who places the mass
- Does the mass lag toward the centre?
- The lighting control run
- What the instruments taught us
- What was measured and not used
- Limits, and what this does not prove
- Files and reproduction
- Appendix: the recipe
1. Setup
The engine and the corpus
Every image comes from Z-Image Turbo (bf16 weights, Qwen3-4B text encoder) running locally in
ComfyUI 0.17.2 on an Apple M1 Max, at 768 × 768, 4 steps, cfg 1, res_multistep, shift 3.
Generation is deterministic: three re-rendered cells came back pixel-identical to their originals.
| Set | Prompts | Seeds | Images | Purpose |
|---|---|---|---|---|
| Round 1, placement | 27 (3 subjects × 9 positions, with "small in the frame") | 6 | 162 | does a position word place the subject? |
| Round 1, scenes | 27 varied scene prompts | 6 | 162 | composition when nothing geometric is named |
| Round 2, position named | 27 (3 new subjects × 9 positions) | 6 | 162 | position without a size word |
| Round 2, size named | 27 (9-rung size ladder) | 6 | 162 | does a size word set size? |
| Round 2, plain | 27 (9 neutral paraphrases) | 6 | 162 | the engine's default |
| Lighting control | 9 (3 subjects × 3 named backdrops) | 6 | 54 | what changes when lighting is named |
The same six seeds run across every prompt in a set, so a seed's own effect can be separated from the prompt's. Round 2 uses subjects and seeds that appear nowhere in round 1.
The instruments
Five readers measure the same pixels in different ways. They are not independent; they differ in computation, not in data.
| Reader | What it finds | Known bias |
|---|---|---|
| Object mask (built here) | the subject itself, against a background model fitted to the frame border | fails when the subject touches the frame edge |
| ISR mask coherence | five mass families (edge, tone, colour, perturbation, blend) and their centroids | whole-frame averages dilute a small subject |
| ISR Notan mass | value mass, both polarities against the image median | counts a dark backdrop as mass |
| ISR soft structure, saliency, radial | soft gradients, spectral saliency, radial compliance about frame and mass centres | no radial eligibility gate in this port |
| The Crit and The Field (product kernels, unmodified) | the app's structural and colour reading | treats dark as material, so it misreads light subjects on dark grounds |
Protocol
frozen Both rounds were pre-registered: the claim, the kill conditions, the estimand, the referee checks and the decision rules were written and hashed before any evaluation image existed. Changes are dated in an amendments file, each made before measurement, each with the evidence that forced it. exploratory marks everything computed after a verdict was read; none of it can change a verdict.
The statistic throughout is a variance share: of the spread in one measurement across prompts and seeds, how much belongs to the prompt. It is an ICC in a crossed prompt-by-seed design, with the McGraw and Wong interval. Round 2 removes each subject's own mean first, so a subject's habits cannot be counted as prompt control.
| Simulated world | survives | inconclusive | killed (partial) | killed (general) | killed (placement) |
|---|---|---|---|---|---|
| law true (reach 0.85, not 0.1) | 0.53 | 0.47 | 0.00 | 0.00 | 0.00 |
| law true, weaker words (reach 0.7, not 0.2) | 0.11 | 0.89 | 0.00 | 0.00 | 0.00 |
| size leaks into Y (G2 Y = 0.6) | 0.00 | 0.78 | 0.00 | 0.00 | 0.00 |
| position leaks into S (G1 S = 0.6) | 0.00 | 0.77 | 0.00 | 0.00 | 0.00 |
| middling world (all cells 0.4) | 0.00 | 0.57 | 0.00 | 0.00 | 0.00 |
The amendment record
Seven changes were made across the two rounds, each before any evaluation number was seen, each dated with the evidence that forced it. Five were referee failures caught on synthetic or simulated data. They are the reason to trust the numbers that follow, so they are listed rather than summarised.
| Round | Amendment | What changed and why |
|---|---|---|
| round 1 | no. 1, 2026-09-13 | R5 runs on the output image, threshold Spearman >= 0.999 |
| round 1 | no. 2, 2026-09-13 | R1 failed; B1 removed from verdict counting; R1 rule for B2 and B3 corrected |
| round 1 | no. 3, 2026-09-13 | E2 dropped by the feasibility rule; S = 6; interval level recalibrated |
| round 1 | no. 4, 2026-09-13 | the crossed bootstrap is replaced by the McGraw-Wong interval (supersedes the interval in Amendment 3) |
| round 2 | no. 1, 2026-09-13 | the instrument needed two fixes before it passed R1 |
| round 2 | no. 2, 2026-09-13 | interval degrees of freedom corrected and level raised to 95% |
| round 2 | no. 3, 2026-09-13 | power under the final rules, stated before measurement |
Two of these removed capability rather than adding it: a second engine was dropped when FLUX.1-dev exhausted the machine's memory, which caps every claim here at one engine, and a product kernel was removed from the round-2 verdict when it failed a known-answer test.
2. What the prompt controls
frozen Round 1 asked whether the draw beats the prompt at setting composition. It does not. Explicit placement words control placement: measured on the object itself, the prompt explains 0.99 of a red apple's horizontal and vertical position, and the subject lands in the instructed third 81% of the time horizontally and 98% vertically, against 33% by chance. The official verdict was inconclusive on a referee failure (a tone correction broke on near-black images); the measured verdict was a partial kill, and every later bearing agreed the claim was false for placement.
Round 2 then asked the sharper question: within a subject, does a prompt control exactly the axes it names? Five of nine cells confirmed, four stayed open, nothing failed and nothing leaked.
| Family | Axis | Predicted | prompt share (95%) | seed share | Result |
|---|---|---|---|---|---|
| position named | horizontal position | reach | 0.84 [0.74, 0.92] | 0.00 | confirms |
| position named | vertical position | reach | 0.50 [0.32, 0.69] | 0.05 | open |
| position named | size | not | 0.63 [0.45, 0.80] | 0.09 | open |
| size named | horizontal position | not | 0.00 [0.00, 0.10] | 0.25 | confirms |
| size named | vertical position | not | 0.26 [0.12, 0.46] | 0.11 | confirms |
| size named | size | reach | 0.55 [0.39, 0.73] | 0.07 | open |
| plain prompts | horizontal position | not | 0.06 [0.00, 0.22] | 0.11 | confirms |
| plain prompts | vertical position | not | 0.17 [0.05, 0.36] | 0.10 | confirms |
| plain prompts | size | not | 0.51 [0.32, 0.70] | 0.16 | open |
What settled round 1
exploratory Round 1's official verdict was capped by a referee failure: a tone correction reordered brightness on 48 of 324 images, worst case 0.75 against a 0.999 floor. Three checks made after the fact all pointed the same way.
| Check | Result |
|---|---|
| Re-measured with the tone correction removed (every referee check then passes) | killed on placement |
| Dropping every prompt with an image below the R5 floor (6 prompts, 288 images left) | unchanged |
| Measuring the apple itself with a colour mask, object-level | prompt share 0.99 horizontal, 0.99 vertical |
| Placement axis, tone correction removed | Notan reader | Border reader | Result |
|---|---|---|---|
| horizontal position | 0.86 [0.79, 0.92] | 0.84 [0.77, 0.90] | fires |
| vertical position | 0.82 [0.73, 0.89] | 0.76 [0.66, 0.85] | fires |
The colour mask also put the apple in the instructed third 81% of the time horizontally and 98% vertically, against 33% by chance. The same mask failed on the third subject, a woman in a yellow coat whose coloured area is 0.1% of the frame, and that failure is reported rather than dropped.
Round 1, for comparison
Round 1 measured five geometric axes with two whole-frame readers, before the object-level instrument existed. Its placement cells stayed open because the readers disagreed, which is what prompted the rebuild.
| Family | Axis | Notan reader | Border reader | Result |
|---|---|---|---|---|
| placement prompts | horizontal position | 0.71 [0.60, 0.82] | 0.47 [0.32, 0.63] | open |
| placement prompts | vertical position | 0.70 [0.58, 0.80] | 0.54 [0.40, 0.68] | open |
| placement prompts | dispersion | 0.70 [0.58, 0.81] | 0.29 [0.17, 0.45] | open |
| placement prompts | inner mass fraction | 0.70 [0.59, 0.81] | 0.30 [0.18, 0.46] | open |
| placement prompts | sector variation | 0.36 [0.23, 0.53] | 0.34 [0.21, 0.51] | open |
| scene prompts | horizontal position | 0.18 [0.08, 0.33] | 0.26 [0.14, 0.42] | holds |
| scene prompts | vertical position | 0.67 [0.55, 0.78] | 0.79 [0.70, 0.87] | fires |
| scene prompts | dispersion | 0.46 [0.32, 0.62] | 0.53 [0.39, 0.67] | open |
| scene prompts | inner mass fraction | 0.51 [0.38, 0.66] | 0.63 [0.51, 0.76] | open |
| scene prompts | sector variation | 0.52 [0.40, 0.68] | 0.70 [0.58, 0.80] | open |
The pattern across both rounds:
- Horizontal position is a clean control. Named, the prompt owns it; unnamed, it belongs to the draw, with no cross-talk from size words.
- Vertical position is half-controlled. Low and middle placements work; high ones mostly do not.
- Size is a switch, not a dial. Naming size buys about as much as a harmless rewording, and position words carry size with them.
| Template | Subject | Corner reach (1.0 = as far as possible) | Edge contact at corners |
|---|---|---|---|
| round 2 | vase | 0.70 | 0.15 |
| round 2 | cat | 0.43 | 0.70 |
| round 2 | man | 0.71 | 0.96 |
| round 1 | apple | 0.77 | 0.17 |
| round 1 | lighthouse | 0.80 | 0.54 |
| round 1 | coat | 0.93 | 0.58 |
Round 1's template said "small in the frame, with plenty of empty plain background"; round 2's did not. With the size words, subjects travel most of the way to a corner and rarely touch the edge. Without them, the engine reaches a corner by pushing a large subject partly out of frame. The two rounds differ in subjects and seeds as well, so this is a strong hint rather than a controlled comparison, and it is the most practical finding here for anyone writing prompts.
Size, in detail
exploratory Size is the axis that behaves least like a control. Naming it moves the subject between two framings rather than along a scale, and position words carry it along.
| Subject | size follows the ladder (rank correlation) | mean size at the four corners | at centre | reaches the top third | middle third | bottom third |
|---|---|---|---|---|---|---|
| vase | 0.85 | 0.15 | 0.50 | 0.12 | 0.94 | 0.75 |
| cat | 0.52 | 0.40 | 0.60 | 0.29 | 0.94 | 0.61 |
| man | 0.76 | 0.45 | 0.64 | 0.56 | 0.89 | 0.94 |
The vase follows the size ladder well (0.85) and the cat barely at all (0.52): asked for "tiny", it still draws a whole cat. Corner prompts cut the vase to a fifth of its centred size. And the vertical columns show the asymmetry directly: every subject lands in the middle third about nine times in ten, while the vase reaches the top third in one image out of eight.
| Subject, plain paraphrases only | smallest wording | largest wording | range |
|---|---|---|---|
| vase | 0.48 | 0.54 | 0.06 |
| cat | 0.57 | 0.67 | 0.10 |
| man | 0.49 | 0.75 | 0.26 |
Rewording alone moves a person's framing by a quarter of the frame and an object's by almost nothing: the vase is the same size under all nine wordings, while the man swings between full-body and close-up.
3. How the model satisfies a position word
exploratory Counting what the engine actually does when told where to put something, by eye, one non-blind reader, on all 162 round-2 position images.
| Instruction | Vase | Cat | Man in a red jacket |
|---|---|---|---|
| corners | cropped against the edge, 23 of 24 | cropped, about 23 of 24 | bottom corners cropped 12 of 12; top corners rotated sideways 7, cropped 5 |
| top centre | ignored, 6 of 6 | mostly ignored | upside down 3, headless crop 1, ignored 2 |
| bottom centre | shrunk, 6 of 6 | shrunk and cropped, 6 of 6 | shrunk to a small full figure, 6 of 6 |
| middle row | normal | normal | normal |
Four moves, used predictably: crop for corners, shrink for bottom centre, ignore for top centre with objects, and rotate when a person is asked to go high. Rotation never happens to an object. This is what "half-controlled vertical" looks like from the inside.
4. The default composition
exploratory Left to itself, the engine centres an isolated object, large and slightly low.
| Prompts | images | in the centre cell | horizontal spread (SD) | median vertical position | median size |
|---|---|---|---|---|---|
| plain object prompts | 162 | 100% | 0.027 | +0.094 | 0.59 |
| size named | 162 | 87% | 0.037 | +0.095 | 0.60 |
| round-1 scene prompts | 145 | 81% | 0.054 | +0.065 | 0.42 |
| position named | 158 | 45% | 0.157 | +0.119 | 0.44 |
Plain object prompts put 100% of subjects in the centre cell of a three-by-three grid, with a horizontal spread about a sixth of what position words produce. Scenes are looser (81%), since a scene carries its own layout. In every family the subject sits slightly below centre, which fits the model's resistance to "top": a picture is built on a ground plane.
5. Who places the mass
exploratory The object is not the composition. A mass reader integrates the whole field: subject, shadow, backdrop gradient, vignette. So the question splits. Who places the object, and who places the mass?
- Name a position and the mass follows the object. Every reading tracks it, prompt shares 0.73 to 0.99, correlations with the object 0.85 to 0.99.
- Leave it unnamed and the mass drifts off the object. The prompt explains almost none of it; 13% to 38% is a seed effect that repeats across prompts, and the rest is a per-image draw. With the object pinned at centre, colour mass and The Crit's mass barely track it at all.
- The offsets are small, 2% to 9% of the frame, but they are systematic.
What the prompt and the seed each own
| Prompt family | background lightness | overall lightness | horizontal position | vertical position | subject size |
|---|---|---|---|---|---|
| position named (round 1) | 0.17 / 0.79 | 0.10 / 0.86 | 0.99 / 0.00 | 0.92 / 0.00 | 0.86 / 0.01 |
| position named (round 2) | 0.19 / 0.61 | 0.12 / 0.62 | 0.84 / 0.00 | 0.50 / 0.05 | 0.63 / 0.09 |
| size named | 0.01 / 0.69 | 0.30 / 0.47 | 0.00 / 0.26 | 0.26 / 0.11 | 0.55 / 0.07 |
| plain prompts | 0.26 / 0.57 | 0.27 / 0.50 | 0.06 / 0.11 | 0.17 / 0.10 | 0.51 / 0.16 |
| scene prompts | 0.70 / 0.11 | 0.79 / 0.08 | 0.02 / 0.10 | 0.36 / 0.01 | 0.65 / 0.01 |
Read the pairs as prompt share / seed share. On plain backgrounds the seed owns lighting and the prompt owns geometry. In scene prompts that inverts: the scene sets its own light, and the seed's share of brightness falls to about a tenth.
6. Does the mass lag toward the centre?
exploratory When a prompt pushes the object off centre, the mass does not always go with it.
- Small subject, open ground: whole-field readings lag badly (0.36 to 0.76) and sit between the object and the centre in 53% to 80% of images. Structure-thresholded readings (ISR kernel, saliency, The Crit) follow at 0.92 to 1.00.
- Large cropped subject: horizontal readings slightly overshoot (1.05 to 1.17) because the cropped subject reaches the frame edge, while vertical readings still lag (0.56 to 0.75).
7. The lighting control run
exploratory If the seed owns lighting only because the prompt leaves it open, naming the backdrop should take it back. Fifty-four images test that: three subjects, three named backdrops, the same six seeds.
| Named backdrop | mean background lightness | seed spread (SD) | seed share of brightness | value mass vs brightness |
|---|---|---|---|---|
| bright white | 0.90 | 0.076 | 0.02 | -0.33 |
| mid-grey | 0.68 | 0.079 | 0.67 | -0.49 |
| dark charcoal | 0.25 | 0.036 | 0.16 | -0.72 |
| lighting unnamed (G3) | — | 0.080 | 0.57 | -0.68 |
Naming the backdrop sets the level: the spread between levels is four times the spread within one. But the seed keeps its wobble, about ±0.07, the same as when lighting was unnamed. Its share falls only because the prompt's effect grew. Shares are relative; the absolute spread is the honest number.
8. What the instruments taught us
Half of what this study produced is about the measuring, not the model. Each of these cost a correction.
| Finding | How it surfaced | Consequence |
|---|---|---|
| The Crit's mask treats dark as material | known-answer synthetic test: correlation 0.97 for dark subjects, 0.10 for light ones | removed from the round-2 verdict; still reported |
| Whole-frame centroids dilute a small subject | object-level mask read 0.99 where frame-level readers read 0.30 to 0.71 | object-level instrument built for round 2 |
| A tone correction can break on near-black images | referee check R5 failed on 48 of 324 images, worst 0.75 | round 1's official verdict capped at inconclusive |
| Bootstrap intervals are mis-specified with few seeds | a planted null could never reach "survives" | replaced with the McGraw and Wong interval, verified against a published example |
| Border-as-background fails on frame-touching subjects | visual audit: 5 of 30 masks clearly wrong, all at the frame edge | the size axis carries the weakness; validation set must include edge cases |
| Radial compliance needs its eligibility gate | the RCA-2 source gates RC_s; this port does not | do not read dRC as a verdict on off-centre work |
The instrument's own validation, after two fixes, reached 99.5% valid masks and correlations of 0.998 or better with known answers on gradients, vignettes, horizons, grey subjects and both polarities. It had never been tested on subjects that touch the frame, which the generator produces constantly.
Radial compliance, and why it says little here
ISR carries the dual-centre radial reading from the RCA-2 notebook: how well mass decays from the frame centre (RC_f), how well it decays from its own centroid (RC_s), and the difference between them.
| Prompt family | RC_f median | RC_s median | dRC median [quartiles] | labels |
|---|---|---|---|---|
| placement (round 1) | 0.86 | 0.89 | +0.012 [+0.001, +0.072] | neutral 111, self-organizing 51 |
| scenes | 0.90 | 0.90 | +0.001 [-0.001, +0.004] | neutral 155, self-organizing 7 |
| position named | 0.90 | 0.92 | +0.015 [-0.007, +0.109] | neutral 93, self-organizing 69 |
| size named | 0.91 | 0.91 | -0.005 [-0.009, -0.001] | neutral 159, self-organizing 3 |
| plain | 0.91 | 0.90 | -0.006 [-0.012, -0.002] | neutral 162 |
On plain and size-named prompts, where the subject sits centred, the two centres coincide and dRC is essentially zero: 162 of 162 plain images read "neutral". Only the placement families, where the subject is pushed off centre, start to separate (round 2: 69 of 162 read "self-organizing"). So this coordinate has something to say about off-centre work and nothing to say about a centred object. Note also that this port lacks the notebook's radial eligibility gate, which exists to stop RC_s being read on masses that are not radially eligible in the first place.
9. What was measured and not used
Honesty about scope. The full instrument stack ran on every image, and this paper uses a fraction of it.
ISR
Twenty-two measurement blocks were captured for all 810 images and are on disk. Six are used here:
- kernel
- mask_coherence
- notan
- radial
- saliency
- soft_structure
Sixteen are not, although they are computed and available:
- color_analysis
- construction
- default_gravity
- dependency
- evidence
- field
- field_qa
- field_regime
- islands
- mass_grid
- reading
- scanpath
- self_similarity
- soft_read
- tonal
- vtl
The Field
The colour kernel ran on all 864 images and produced 37 colour measurements each. None of them appears in this paper. Colour was never part of either frozen question, and the colour mass family used in section 5 comes from ISR, not from The Field. It is listed among the instruments because it was run, not because it contributed.
What that leaves on the table
- The island and dependency layers read a picture as parts and their relations, which is the natural home for the counterweight question in section 6.
- Tonal hierarchy, field regime and the legacy VTL block would say what kind of picture the engine makes by default, not just where the subject sits.
- Scanpath and construction ask where the eye rests and whether a structure is built from tone or colour, both untouched.
- The Field's 37 colour coordinates would answer the colour version of this study's question: does a prompt control palette the way it controls position?
None of this changes a result above. It marks how much of the measurement was spent and how little was spent well, and it is the cheapest available next work, since the numbers already exist.
10. Limits, and what this does not prove
- One engine. Everything here is Z-Image Turbo at 768 px. Nothing says another model behaves this way. A second engine is the next real step.
- Six subjects, twelve seeds, mostly plain backgrounds. Scenes appear only in round 1.
- The bearings are not independent. Every reader sees the same pixels. They differ in computation, which is weaker than differing in data.
- The exploratory sections are post-hoc and were written after the verdicts. They are leads with tests attached, not results.
- No human judgement was collected. Nothing here says whether any composition is good.
- Sharper for isolated objects, by construction. Gradients and multi-object scenes need the rest of the kernel stack, or adjusted kernels, and were not the target here.
11. Files and reproduction
Four folders under image-steering-accreditation-claude/, each with its frozen bet, its dated
amendments, its referee outputs and its data:
| Folder | Contents |
|---|---|
seed-bet/ | round 1: BET.md, AMENDMENTS.md, VERDICT.md, 324 images, measurements, referee files |
named-axes/ | round 2: the object instrument and its validation, BET.md, AMENDMENTS.md, VERDICT.md, 486 images |
insights-zimage/ | the dives: features, full ISR reads of all 810 images, mass comparison, drag, lighting control, MASS.md, INSIGHTS.md |
zimage-paper/ | this paper and its build script |
Measurement is reproducible from the saved images: the object instrument and interval code are hashed in
named-axes/instrument_freeze.sha256, the product kernels are recorded by their own hashes, and
ISR read every image with persistence off, 4.7 s per image, 810 images in 13 minutes at five in parallel.
Appendix: the recipe
A1. Prompt templates
Each set varies one thing and holds the rest fixed. Subjects are named only where they matter.
| Set | Template | What varies |
|---|---|---|
| Round 1, placement | "A photograph of {subject} in the {position} of the frame, small in the frame, with plenty of empty plain background around it." | 3 subjects (a single red apple, a lighthouse, a woman in a yellow coat) × 9 positions |
| Round 1, scenes | the scene alone, no geometric words | 27 scenes, listed below |
| Round 2, position named | "A photograph of {subject} in the {position} of the frame, with a plain background." | 3 subjects (a blue ceramic vase, a black cat, a man in a red jacket) × 9 positions |
| Round 2, size named | "A photograph of {subject}, {size}, with a plain background." | the same 3 subjects × 9 size rungs |
| Round 2, plain | 9 neutral paraphrases, no geometric words | the same 3 subjects × 9 wordings |
| Lighting control | "A photograph of {subject}, with {backdrop}." | the same 3 subjects × 3 backdrops |
The nine positions: top-left corner, top center, top-right corner, middle left, exact center, middle right, bottom-left corner, bottom center, bottom-right corner.
The nine size rungs, smallest to largest:
- tiny in the frame
- very small in the frame
- small in the frame
- fairly small in the frame
- medium-sized in the frame
- fairly large in the frame
- large in the frame
- very large in the frame
- filling most of the frame
The nine neutral paraphrases:
- A photograph of {subject}, with a plain background.
- A studio photograph of {subject} against a plain backdrop.
- {Subject}, photographed on a plain background.
- An image of {subject} on a simple plain background.
- A clean product-style photo of {subject}, plain background.
- A high-quality photograph showing {subject} against a plain background.
- {Subject} on a plain seamless background, photograph.
- A simple photograph of {subject}, nothing else, plain background.
- A professional photo of {subject} set against a plain background.
The three named backdrops:
- L1: a bright white background, brightly lit
- L2: a plain mid-grey background, evenly lit
- L3: a dark charcoal background, dimly lit
The 27 scene prompts: a portrait of an old fisherman; a bowl of lemons on a kitchen table; a snowy mountain range; a busy street market in Marrakech; a cat sleeping on a windowsill; the interior of a gothic cathedral; a racing bicycle; a field of sunflowers; a jazz trio performing in a small club; an abandoned factory; a child flying a kite on a beach; a stack of old books; a tropical rainforest waterfall; a vintage red sports car; a chess game in progress; a lighthouse in a storm; a plate of sushi; a herd of elephants crossing a river; a modern glass skyscraper; a violin resting on a chair; a desert with sand dunes; a crowded subway car; an owl on a branch at night; a wedding cake; a medieval castle on a hill; a pair of hands knitting; a thunderstorm over a wheat field.
A2. Engine and sampling
| Setting | Value |
|---|---|
| Model | Z-Image Turbo, z_image_turbo_bf16.safetensors |
| Text encoder | Qwen3-4B (qwen_3_4b.safetensors, lumina2 loader) |
| VAE | ae.safetensors |
| Sampler / scheduler | res_multistep / simple, 4 steps, cfg 1.0, denoise 1.0 |
| Model sampling | ModelSamplingAuraFlow, shift 3 |
| Negative prompt | zeroed conditioning (ConditioningZeroOut) |
| Resolution | 768 × 768, one image per job |
| Host | ComfyUI 0.17.2 headless, Apple M1 Max 32 GB, about 45 s per image |
| Seeds, rounds 1 | 1000 + 7919k, k = 0..5 |
| Seeds, round 2 and lighting | 50000 + 104729k, k = 0..5 |
Determinism was checked, not assumed: re-rendering fixed cells reproduced them pixel for pixel in both rounds.
A3. What each axis means
| Axis | Definition |
|---|---|
| horizontal / vertical position | centroid of the subject mask, as a fraction of frame width or height from the centre; positive is right and down |
| size | square root of the subject mask's area fraction, so it scales like a length |
| reach | how far the mean centroid moves toward the instructed edge, over the furthest a subject of that size could go while staying in frame |
| edge contact | share of images where the mask meets the outer 4% ring of the frame |
| mass centroid | the centre of gravity of a mass field over the whole frame, subject and background together |
| background lightness | mean OKLab lightness of the border ring outside the subject mask |
| dRC | radial compliance about the mass's own centroid minus about the frame centre |
A4. The statistic
Prompt share is the intraclass correlation ICC(A,1) in a crossed design: prompts are the subjects, seeds are the raters, one image per cell. It answers "of the spread in this measurement, how much belongs to the prompt rather than to the draw". Seed share is the matching quantity for a seed effect that repeats across prompts. Round 2 subtracts each subject's own mean first, so a subject's habits are not counted as prompt control; that costs two degrees of freedom, which the interval accounts for.
Intervals are the McGraw and Wong (1996) F-based construction, checked against a published worked example before use. The level is 95% rather than 90% because simulation at this design size showed the 90% interval covering as little as 0.82 of the time; at 95% the worst case is 0.89. The first method tried, a crossed bootstrap, was discarded: resampling six seeds with replacement duplicates whole columns, which the formulas treat as real levels, and a planted null could never have been confirmed.
A5. Referee checks
| Check | Asks | Outcome |
|---|---|---|
| Known-answer placement | does the instrument find an object whose position is known, across backgrounds, polarities and sizes? | passed after two fixes; The Crit failed and was removed from counting |
| Estimator recovery | does the interval cover a planted truth at the real design size? | passed at 95% after the method and level were corrected |
| Determinism | does re-rendering a cell reproduce it? | passed, pixel-identical |
| Seed crossing | did every prompt get the same seeds? | passed |
| Mask validity | did the instrument find a subject at all? | passed: 4 invalid of 486 in round 2 |
| Visual audit | do the masks look right on a random sample? | passed narrowly: 5 of 30 wrong against a limit of 6 |
| Tone remap monotone | did the brightness correction preserve order? | failed in round 1, on 48 of 324 images |
A6. Hashes and commands
The instrument, the interval code and the product kernels were hashed at the moment of freezing, so a rerun can prove it used the same code.
| File | sha256 (first 16) |
|---|---|
objmask.py | 3e4bc7a3dde1cca5 |
stats_centred.py | 99eb21644a164d3e |
gen2.py | bd03d34c5566d5e7 |
../seed-bet/gen.py | a401f5802e088b3d |
products/the-crit/js/kernel.js | 460b745c235f6c63 |
products/the-field/js/color-kernel.js | 6eef90e7d1ab8a24 |
To reproduce from the saved images:
| Step | Command |
|---|---|
| measure with the object instrument | python3 analyze2.py measure |
| apply the frozen rules | python3 analyze2.py analyze |
| full ISR read of every image | python3 insights-zimage/isr_batch.py |
| product kernels | python3 steering-authority/run_measure.py <image root> <out.csv> |
| object against mass readings | python3 insights-zimage/mass_compare.py |
| rebuild this paper | python3 zimage-paper/build_paper.py |
A7. Terms
| Term | Meaning here |
|---|---|
| prompt share | fraction of a measurement's variance explained by which prompt was used |
| seed share | fraction explained by a seed effect that repeats across different prompts |
| the draw | everything the prompt does not set: the seed's repeatable part plus the per-image remainder |
| object reading | a measurement of the subject itself, found by a mask |
| mass reading | a measurement of the whole field's centre of gravity, subject and background together |
| Notan mass | value mass counting both lighter and darker deviations from the image median |
| frozen | pre-registered before any evaluation image existed |
| exploratory | computed after a verdict was read; cannot change it |