Patch ledger · study 1 run 2026-09-18 · phase 3 run 2026-09-18/19 · updated 2026-09-19 · local file, not published

A generator with an exact credit ledger, and what it says about paying contributors

Built from Kamb & Ganguli (2025), who showed that convolutional diffusion models behave like a closed-form "patch mosaic" machine. Used as the generator itself, that machine makes every output pixel an exact, owner-labelled blend of training pixels, and "the output without contributor X" costs one rerun. With exact counterfactuals in hand, we tested how much cheaper ways of paying contributors get wrong. Phase 3 then trained ordinary diffusion networks on the same images to check that the machine describes them.

~25%
of contributors' money paid to different people by usage-based receipts than by counterfactual credit (median of four outputs)
0 of 5
cheap credit rules consistent with counterfactual credit, including exact leave-one-out
61–779×
stronger influence of an img2img start image than of any single corpus contributor (post hoc)
p = 0.001, 0.0002
trained networks move in the direction the machine's credit predicts for the removed group; replicated with new start weights (single-step test; direction, not magnitude)
r 0.89–0.94
trained networks match the machine from the same noise: Kamb & Ganguli reproduced on a new corpus with independent code
16 / 16
self-tests passed, including agreement with an independent brute-force implementation to ~1e-6

How it works, in plain terms

In one sentence: it is a picture-maker that builds every image out of small pieces of the paintings it was given, and keeps an exact receipt of which painting each piece came from.

Four steps: library, generate, receipt, what if
Library: every training image is cut into small square patches, each tagged with its owner. Generate: starting from static, over 20 rounds each pixel moves toward a weighted blend of the patches its neighbourhood most resembles (large patches first, small last). Receipt: every blend is explicit, so each pixel's weights by owner are known exactly. What if: remove an owner's patches and rerun from the same static, with no retraining.
One pixel's receipt and the whole image's receipt
A real receipt from output U1. One pixel, averaged over the 20 rounds, draws on the commons, a Met work and five different Mont Sainte-Victoire paintings; the whole image draws about half on the commons. The pixel was chosen, for illustration, as a central pixel with a well-mixed receipt.

Why it matters. Real image generators cannot say what an image would have been without a given contributor, so no payment method can be checked against the truth. This one can, which makes it a test bench for how wrong cheaper payment methods are. Phase 3 then checks that ordinary trained networks of a simple kind behave like it.

Plain-language glossary
TermMeaning here
Usage receiptPaying each owner by how much their patches were drawn on while the image was made
CounterfactualThe image made without a given owner's material
Leave-one-outCredit by removing one owner at a time and measuring how much the image changes
Shapley valueCredit averaged over every order in which owners could join; shares credit fairly among owners who substitute for each other
Total variation (TV)The fraction of money two payment rules send to different people
Frozen protocolRules and thresholds fingerprinted before the data existed, so they cannot be adjusted afterwards
Post hocAn analysis chosen after seeing results: useful for explanation, weaker as evidence
p-valueHow often a result this strong would appear by chance with no real effect; p = 0.001 is about 1 in 1,000
img2imgGenerating from a noised copy of an existing image
CommonsMaterial paid by no rule here, like a model pool

Study 1 verdicts under the frozen protocol

ClaimResultVerdict
H1Usage receipts misallocate: median total-variation distance to Shapley ≥ 0.200.254, bootstrap 90% [0.241, 0.269]PASS
H2Generic material over-credited by usage at least 2×median ratio 1.42 (range 0.78–3.04)FAIL
H3Shapley over the corpus ranks a derivative output's source first (≥ 3 of 4)0 of 4FAIL
RefereeAny cheap rule consistent with Shapley (Spearman ≥ 0.7 and source top-1 ≥ 3/4)noneNONE

Protocol written and hashed before the run; all hashes were checked at start; no deviations. The analysis script was written after the freeze but before any result was inspected. Items marked post hoc were computed after the verdicts were read and change none of them.

What the machine makes

Generated 32 by 32 outputs above training images
Top two rows: generated outputs (32×32, enlarged). U1–U8 start from pure noise; D1–D4 start from a noised painting (img2img). Bottom two rows: some of the training images. Outputs are painterly patch mosaics: the corpus's palette and brushwork rearranged, not depictions of anything. This matches the paper's own outputs for convolution-only models.

H1: usage receipts misallocate a quarter of contributors' money

"Usage" pays each contributor by how much posterior weight their patches carried while the image was made. "Shapley" pays by average counterfactual contribution: how much closer the output gets when the contributor joins each possible coalition of the others. The gap between the two, per output:

Bar chart of misallocation per output
Every output except U1 is above the pre-declared 0.20 line. Paying on final-step weight (orange) is worse than paying on the whole trajectory.
How far this generalises. The 90% interval [0.24, 0.27] covers Shapley estimation noise only, not which outputs happened to be generated. The verdict rests on four outputs, one below the line. For these outputs the gap is real and large; its spread across many outputs is not established. A supporting check over all eight unconditional outputs, against exact removal rather than Shapley, gives a median gap of 0.20. "Misallocated" means relative to Shapley with the declared value function, a chosen reference rather than a proven truth.
Credit shares per contributor for four outputs
Left: output. Middle: which contributor dominates each pixel at the final step (the large light-blue area is the always-present commons). Right: shares under usage (blue), exact leave-one-out (orange) and Shapley (green).

No cheap rule tracks counterfactual credit

Rank agreement of each rule with Shapley
Median over eight outputs. Trajectory usage comes closest (0.70, just under the bar). The vendor-style method that looks for training patches in the output has the lowest agreement.
Per-output rank agreement
Post hoc. The black line is Shapley against itself (two independent halves of the permutations): 0.82–0.98, so the reference ranking is stable. On unconditional outputs U2–U4 the cheap rules, including exact leave-one-out, fall to between −0.2 and 0.5. The disagreement is real, not estimator noise. On derivative outputs (D1–D4), where an input anchors the image, everything agrees more.
RuleSpearman with ShapleyMoney misallocatedSource found (D1–D4)Consistent
Trajectory usage0.700.223/4no
Final-step usage0.680.231/4no
Exact leave-one-out0.600.240/4no
Global similarity (vendor-style)0.55n/a4/4no
Patch retrieval (vendor-style)0.450.341/4no

Exact leave-one-out, often treated as the ground truth for removal, does no better than usage. The nine Mont Sainte-Victoire paintings substitute for each other: removing any one is covered by the rest, so each looks unimportant even when the group matters. Measured post hoc, both cheap rules underpay the series by about the same amount: the Cézannes (with the copy) receive 0.78 of their Shapley share under usage and 0.71 under leave-one-out, while the Met works receive 1.18 and 1.29. Choosing between the two rules is not the fix; only a rule that shares credit among substitutes pays a series what the counterfactual says it is worth.

H3 failed, and why: the img2img input is an unledgered channel

Derivative outputs started from a noised copy of a known painting. Shapley over the corpus never ranked that painting first. The post hoc check shows why: the source's influence arrives through the start image, which the corpus counterfactual holds fixed.

Start-image swap effect versus corpus removal effect
Swapping the start image for a different painting moves the output 61–779× more than removing the most influential contributor from the corpus. The two interventions differ in kind (a whole input versus one of 26 contributors), and the replacement paintings were chosen by hand, so the ratios show scale, not a precise quantity.
OutputSourceStart-image swapLargest corpus removalRatioSource's own removal rank
D1Cézanne B0.2600.00095274×3
D2Met DP-12952-0010.1920.00025779×3
D3Cézanne A0.5680.00119476×8
D4Met DP-14201-0010.0440.0007261×3

Corpus credit does not ignore the source (leave-one-out ranks it 3rd of 26 in three of four cases), but it cannot recover its dominant role. Plain similarity finds the source every time because the output resembles its input, which suggests similarity-based attribution largely detects input reuse rather than corpus contribution.

Settlement consequence. A reference or source image supplied at generation time has to be paid as its own party. Corpus attribution, however exact, cannot be expected to find it. Plain global similarity found the source 4 of 4 times precisely because the output still looks like its input.

Other measurements

A copy is not free

Removing Cézanne A or its verbatim copy has an identical, positive effect on every output (U1: 5.2×10⁻⁴ each). The copy doubles the prior weight on that work, which tilts generation toward it: at each step a work with share a rises to 2a/(1+a). Self-test T5 proves the identity exactly: the corpus with a duplicate equals the original at prior weight 2. Every rule credits the two equally, so a work uploaded twice collects close to double. Measured post hoc on U1–U8: with the copy present, the work and its copy are paid 1.93× what the work alone is paid without it, and the work's influence on outputs rises 3.07× in squared units (1.75× in RMS terms). Duplication amplifies a work's pull on generation as well as its pay; whether pay over- or under-states that depends on the unit, but a duplicate unambiguously biases generation toward its work.

H2 failed: misallocation is not concentrated on filler

A contributor of blurred colour fields got 1.4× its Shapley share under usage at the median, not the predicted 2×. It accounts for at most about 1 point of the 18–28 point gap per output; the rest is spread across real contributors. Only one generic contributor was tested.

Cost of exact counterfactuals

On the Mac's GPU, removing each of the 26 contributors from one output took about 2.6 minutes in total; a 10-permutation Shapley estimate took about 24 minutes per output. No retraining anywhere.

Per-output detail

OutputMisallocation (trajectory usage)Misallocation (final-step usage)Usage rank agreementShapley reliabilityNegative Shapley valuesCommons weight
U10.1770.1740.740.8610.52
U20.2440.3150.490.8530.54
U30.2820.3070.180.9310.57
U40.2640.353-0.030.8200.62
D10.2050.2440.700.9400.47
D20.2530.2160.700.9310.59
D30.1930.1690.930.9350.37
D40.0950.0880.950.9810.38

All headline numbers recomputed independently from the raw saved tensors (verify_study1.py): every value matches to four decimals. Contributor shares exclude the always-present commons, which takes 52–62% of usage payment. Shapley: 10 permutations (5 antithetic pairs) per output, 251 trajectories each; efficiency holds to 1e-18; median relative standard error ≈ 10%; negative values clipped to zero in shares, as declared.

Method

Self-tests (16/16)
Contributors (players)
Development look used to set the schedule
Development schedule grid
Development contributors only. Rows alternate between the position-free and position-aware variants at maximum patch sides 7, 11, 15 and 19.
Freeze record
FileSHA-256
PROTOCOL.mddd663afd0f802fea…
run_eval.py53e0368a68c4c03b…
pl/machine.py2ee175aa061da6dc…
pl/corpus.py0dea3aaedcc892e9…

Study 1: corpus images SHA-256 49acf50d76f15834…, frozen 2026-09-18 06:44:34. Phase 3 freezes (hashes in out/phase3/): round 2 protocol, code, groups and predictions (2026-09-18 23:36:36); score test (2026-09-19 05:36:02); replication (2026-09-19 06:35:24). Every frozen file was unchanged at its analysis.

Two checks prompted by an outside review

A reviewer raised two objections to the headline number: Shapley credit is one chosen definition of contribution rather than ground truth, and the corpus was built to contain substitutes, a duplicate and filler, so its magnitudes may be artefacts of that construction. Both were tested.

Does the answer depend on what counts as "contribution"?

Shapley credit is computed against a value function. Four were declared in advance and recomputed from the saved runs: pixels (the original), blur (composition), edges (structure) and colour (palette).

Value functionUsage vs Shapley (money reallocated)Rank agreement
pixel0.250.70
blur0.390.29
edges0.200.66
colour0.360.42
Two readings. Usage diverges from Shapley under every definition (0.20–0.39), so the finding does not depend on the pixel choice. But the definitions disagree with each other by 0.17–0.35: choosing what counts as contribution moves money about as much as choosing usage over Shapley. It is a policy decision, and it belongs in the contract.

Is the number an artefact of the corpus?

The same machinery ran on five corpora, from all-distinct works to the study-1 stress test, under a protocol frozen beforehand. The prediction was that the divergence would rise with redundancy. It did not, but the more important thing the run revealed is that the intended manipulation barely happened.

Redundancy dial
Left: usage versus Shapley per output (dots) and median (diamond). Right: leave-one-out's share of Shapley credit for the Cézanne group. SI: the declared substitutability index.
CorpusSubstitutability indexUsage vs Shapley (median)Leave-one-out ÷ Shapley, Cézannes
L00.100.247n/a
L10.120.2011.19
L20.100.1940.79
L30.100.2190.83
L40.190.2520.71
The manipulation mostly failed, and the useful evidence is a level comparison. The declared redundancy measure barely moved across the first four corpora (0.100, 0.115, 0.105, 0.099); only the verbatim copy raised it (0.188). Nine paintings of one mountain are not nine substitutable patches. So the pre-registered trend test (p = 0.22) ran against a predictor with almost no range: that is low power, not evidence that the divergence is independent of redundancy, and we do not claim it is. What the sweep does support: the two corpora at opposite ends of the range actually achieved, all-distinct works (index 0.10) and the stress test (index 0.19), diverge by the same amount, 0.247 and 0.252. So the headline magnitude is not manufactured by the stress-test construction. Separately, the substitution effect appears as a step rather than a dose-response: leave-one-out pays the Cézannes 1.19 of their Shapley share with three paintings and 0.79 with six, then flattens (0.83 with nine), with the last step (0.71) adding a duplicate rather than more series length.

Phase 3: does the machine describe a trained network?

Everything above uses the machine as the generator. Phase 3 trains ordinary convolution-only diffusion networks on the same 244 images and asks two questions: do they behave like the machine (the paper's claim), and does the machine's credit describe them (our question)?

They behave like the machine: the paper reproduces

Machine and network from the same noise
Pilot check, same starting noise. Top: the machine (uncalibrated pilot schedule). Middle and bottom: the trained network at 20 and 150 sampling steps. Layout and colour carry over; the network paints larger, smoother regions.
Network (half width, 128 channels)r, 20 stepsr, 150 stepsClipping checkNear-copies
half_s0 at 10000 steps0.9370.9380.9930/100
half_s0 at 20000 steps0.9300.9440.9880/100
half_s0 at 30000 steps0.9180.9400.9870/100
half_s1 at 30000 steps0.9240.9380.9890/100
half_s2 at 30000 steps0.8870.9370.9810/100

r is Pearson correlation between network and calibrated machine from the same noise, median over 100 seeds. The paper reports 0.90–0.96 for ResNets; its figures call this r², but its code computes r (CODE_COMPARISON.md).

Calibrated patch sizes, ours vs authors
The machine's patch size at each noise level, fitted to our network with the authors' procedure, against their fit for CIFAR10. Coarse-to-fine in both, a second independent bearing on the paper's finding.

Width. At 10k steps, each network against a machine calibrated on the half-width network / on the full-width network. Half width: r 0.937 / 0.940 (20 steps), 0.938 / 0.937 (150 steps). The paper's width (256 channels): 0.901 / 0.906 (20 steps), 0.939 / 0.941 (150 steps). Half width is never worse by more than 0.004 (well inside the 0.02 margin), so it was used throughout.

Does the machine's credit describe the network? Round 2: underpowered

Seven networks were retrained, each without one group of contributors, plus two control networks with nothing removed. The machine predicted how each network's images should change.

GroupPrediction vs actualvs control 1vs control 2Frozen statisticFrozen class
G1+0.108+0.118-0.007+0.053predicted
G2-0.019-0.032+0.025-0.015undetected
G3-0.073-0.093-0.120+0.034predicted
G4-0.034-0.092-0.012+0.018predicted
G5-0.151-0.158-0.041-0.051undetected
G6-0.032-0.034-0.073+0.021predicted
G7+0.058+0.039-0.048+0.063predicted
The frozen rule said CONFIRMED (5 of 7). I do not stand behind it. The two controls, identical recipes with nothing removed, disagree with each other as much as groups differ from them. Retraining noise (mean squared change 0.008–0.059) is an order of magnitude larger than the predicted group effects (0.0003–0.0019), and the frozen bootstrap did not include network-to-network variation. Honest reading: underpowered. Details in PHASE3_RESULTS.md.
Round 2 specificity matrix
Post hoc, needing no controls: does each network's change match its own group's prediction better than other groups' predictions? Rows dominate (each prediction correlates about equally with every network), and the diagonal does not stand out (p = 0.075).

Phase 3b: single-step score test

Round 2 compared finished images, and 20 sampling steps amplify small differences between networks. The follow-up asks each network one question at a time at identical inputs (what noise do you see here?), where a group's removal acts directly. Its design was chosen after round 2; its primary test was declared and frozen before it ran.

Score test specificity matrix
Single-step noise predictions at 15 noise levels, identical inputs for every network. Bold diagonal: each network against its own group's prediction. Statistic +0.0389; exact permutation p = 0.0010 (5 of 5,040 assignments score as high); own-prediction ranks 1, 1, 6, 1, 7, 1, 3.
Verdict under its frozen protocol: SPECIFIC. Each retrained network moves in its own group's predicted direction more than in the others'. Post hoc robustness: dropping any single group still gives p ≤ 0.007; with row and column effects removed, 6 of 7 groups show a positive group-specific component; agreement is negative in the early, noisy steps and positive from mid-generation on (−0.015 at t = 0.8 to +0.042 at t = 0.2), where patches are small and each patch's owner matters most. Limits: the group-specific correlations are weak (at most about 0.07 for seed 0 and 0.11 for seed 1 once row and column effects are removed), magnitudes do not match, and the design was chosen after round 2. Consistent with Zhao et al. (2025), who validated an attribution score built on the same machine against retrained networks at larger scale; this is an independent, small-scale confirmation with a different test.
Replicated. Seven leave-out networks retrained from new start weights, same frozen test: specificity 0.059, exact p = 0.0002 (the true assignment is the best of all 5,040), 7 of 7 groups positive once row and column effects are removed, and the same pattern by noise level (negative early, positive from mid-generation on; seed 1 peaks mid-generation rather than at the end). Against the average network instead of a single reference, the seed-0 signal roughly doubles (0.076): the strong row effects were one reference network's quirks. Details: PHASE3B_RESULTS.md.

Where the claims stand

Rated by how hard each would be to knock down. "Frozen" means decided by rules written and hashed before the data existed.

ClaimEvidenceWeak point
Paying by usage misallocates about a quarter of contributors' money relative to Shapley creditFrozen; median gap 0.25 against a 0.20 bar; recomputed independently; a corpus of all-distinct works diverges by the same amount (0.247 vs 0.252)Four outputs per corpus; interval covers Shapley noise only; one machine; whether the gap depends on redundancy is untested (the sweep's manipulation failed)
What counts as "contribution" moves money about as much as usage versus ShapleyFour declared value functions: usage gap 0.20–0.39; between value functions 0.17–0.35Post hoc; four value functions on one corpus
No cheap attribution rule tracks counterfactual credit (usage, exact leave-one-out, two vendor-style similarity methods)Frozen; rank agreement 0.45–0.70 against a 0.7 bar; Shapley's own reliability 0.82–0.98One corpus
Both cheap rules underpay a series of similar worksMeasured post hoc: the series receives 0.78 of its Shapley share under usage and 0.71 under leave-one-out (other works 1.18 and 1.29); the shortfall deepens as the series grows (p = 0.002)Substitute-rich corpus (as real catalogues often are); post hoc
A reference image supplied at generation time is an unledgered channelOutweighs any corpus contributor 61–779×; corpus attribution cannot recover itPost hoc; ratios compare different kinds of intervention
Similarity-based attribution mostly detects input reuse, not corpus contributionFinds the source 4/4 on derivative outputs, tracks counterfactual credit poorlyInference from two measurements
A copy is not free: it nearly doubles the work's pay (×1.93) and amplifies its influence on outputs (×3.07 squared, ×1.75 RMS)Exact identity (self-test); measured post hoc on eight outputsOver- or under-payment depends on the unit of influence
Exact counterfactuals are cheapRemoving any contributor: minutes, no retrainingEngineering fact
Kamb & Ganguli reproduces on a new corpus with independently written codeFrozen checks: r 0.89–0.94 (theirs 0.90–0.96); calibration tracks theirs; width irrelevant; no memorisationConfirmation of their work, not a new finding
Small corrections to the paperIts "r²" is r in its own code; released CelebA network is deeper than describedMinor
One retrained network per condition is too noisy as attribution ground truth at this scaleIdentical recipes differ 7-fold; noise swamps removing ~6% of the dataConsistent with the literature's use of many retrained models
The machine's group credit points the right way in a trained network (weak form)Frozen single-step score test: SPECIFIC, p = 0.001; frozen replication with new start weights: p = 0.0002, 7 of 7 groups positive; robust to the reference and to dropping any group; positive from mid-generation onDirection only; correlations weak; magnitudes do not match; design chosen after round 2; one corpus and architecture; consistent with prior work (Zhao et al., 2025), not a first
NoveltyA quick search found closely related work (see Related work below)Not established; no systematic review
Whether it shows up in finished imagesRound 2 on sampled images: underpoweredOpen

What it adds up to for settlement design

  1. Declare what counts as contribution, in the contract. Pixels, composition, structure or palette move money between contributors about as much as the choice of payment rule does.
  2. Pay on counterfactual (Shapley) contribution under that declaration, not usage. Usage receipts moved about a quarter of contributors' money in every corpus tested, and both cheap rules underpay a series (78% and 71% of its Shapley share): only a rule that shares credit among substitutes fixes that.
  3. Receipt reference inputs as their own party. A source image supplied at generation time outweighs any corpus contributor, and corpus attribution does not recover its role.
  4. Consolidate duplicates before paying. Every rule pays a duplicated work close to double.
  5. Do not use similarity as a measure of contribution. It mostly detects input reuse.
  6. Use an exact-ledger generator as a referee. It gives ground truth that attribution methods can be scored against; its behavioural match to trained networks is independently reproduced, and its group credit points the right way in trained networks' single-step behaviour (as Zhao et al., 2025, found at larger scale for a machine-based score). A natural first job: score that score, gradient methods and similarity against exact ground truth.
  7. Evaluate attribution at fixed inputs, not on finished images. Retraining noise swamped group effects on sampled images; single-step comparisons with shared initialisation resolved them.

Related work and sources

From a quick web search on 2026-09-19 (about 20 minutes), not a systematic review. Novelty is not established.

Sources: Zhao et al. 2025 · Lin et al. 2025 · Wang, J. T. et al. 2024 · Briq et al. 2026 · Dai and Gifford 2026 · Kamb and Ganguli 2025 · Niedoba et al. 2024 · Zheng et al. 2024. Full reference list in the working paper.

What this does and does not show

Shows: in a generator whose ledger is exact by construction, usage-based receipts, leave-one-out and similarity methods each disagree materially with counterfactual credit, and inputs supplied at generation time sit outside any corpus ledger. Trained convolution-only diffusion networks behave like that generator (the paper's result, reproduced independently), and its group credit points the right way in their single-step behaviour (replicated across two start-weight seeds).

Does not show: that the machine gets the size of credit right for trained networks, or that its credit shows in their finished images (round 2 was underpowered). Nothing here concerns attention or latent generators such as Midjourney or Flux, human perception, rights or authorship. Results are specific to this machine, this 244-image corpus, four unconditional outputs for the primary test and one choice of Shapley value function; a different value function could change the reference.

Source: Kamb, M. & Ganguli, S. (2025). An analytic theory of creativity in convolutional diffusion models. ICML 2025 / arXiv:2412.20292v2. Code and data: image-steering-accreditation-claude/patch-ledger/.