Exact counterfactual credit in a patch-mosaic diffusion generator
What it means for paying contributors
Working paper. Draft of 2026-09-19. Every number in this document is read programmatically from the saved results.
Reading guide. Each finding is labelled with its evidence type: [F] decided by a test whose rules were written and hash-frozen before the data existed; [R] a frozen replication; [P] post hoc, computed after the frozen verdicts were read (exploratory, changes no verdict); [D] a mathematical identity or derivation, checked by self-tests; [X] independently recomputed. Section 6 lists what can and cannot be claimed. Please keep these labels and limits when summarising.
The precondition. Everything here presupposes a boxed generator: one whose generation consults a declared, owner-labelled memory at run time, so its ledger is an accounting identity. Commercial image generators are parametric: they store what they learned in weights, and their outputs are often unattributable to individual training examples (Dai and Gifford, 2026). So the proposal is not a better attribution method for models already shipped; it is to build generators so that attribution is exact. For parametric models, what carries over is the referee role (section 8) and the general rules on value functions, generation-time inputs, duplicates and similarity.
Abstract
Paying the contributors of a training corpus requires knowing who contributed to each output, but for trained generators the true answer is unknown, so attribution methods cannot be validated. We use the closed-form "equivariant local score" machine of Kamb and Ganguli (2025), which reproduces the behaviour of convolution-only diffusion models, as the generator itself. In this generator every output pixel is an exact, owner-labelled blend of training pixels, and "the output without contributor X" costs one rerun, so exact counterfactual credit under any declared value function is available. This presupposes a boxed generator with a declared memory; it is not an attribution method for the parametric models in commercial use.
We built a stress-test corpus of 244 museum images (26 paid contributors plus an always-present commons) that contains the failure modes a settlement system must handle: a series of nine paintings of one subject, a verbatim duplicate and generic filler. Under pre-registered tests, paying by usage (posterior weight) diverges from Shapley credit under a pixel value function by 0.18–0.28 of contributors' money across four outputs (median 0.25) [F]. The divergence persists under four value functions (0.20–0.39), but those value functions also disagree with one another by 0.17–0.35, so what counts as "contribution" is itself a first-order settlement choice [P]. None of five cheaper rules tracks Shapley credit [F].
Is the divergence an artefact of the stress-test corpus? a corpus of all-distinct works, with no series, no duplicate and no filler, diverges by 0.25, the same as the stress test (0.25), so the divergence is not a product of that construction; the pre-registered trend test across five corpora is uninformative, because the redundancy measure barely moved [F]; The corpus also reproduces known mechanisms in a generative setting: both cheap rules under-credit a series of similar works (a series receives 78% of its Shapley share under usage and 71% under leave-one-out), a duplicate nearly doubles a work's pay, and a source image supplied at generation dominates outputs outside any corpus ledger [P, D]. The leave-one-out and duplicate magnitudes are properties of this corpus, and the leave-one-out shortfall scales with the size of the series as theory predicts; the usage-versus-Shapley divergence does not, holding across the whole corpus family.
We then trained ordinary convolution-only diffusion networks on the same images. They match the machine from the same noise at r = 0.89–0.94, independently reproducing the paper [F]. The machine's predicted change when a group of contributors is removed is specific to that group in the networks' single-step behaviour (exact permutation p = 0.0010; replicated with new start weights, p = 0.0002) [F, R], but weak: correlations of at most about 0.11 (around 1% of variance), magnitudes off by more than ten-fold, and unresolved on finished images. The analytic machine has been used for attribution before (Zhao et al., 2025, validated against retrained networks at larger scale), and Shapley-based credit and royalties for generative models have been proposed (Wang et al., 2024; Lin et al., 2025). What we did not find in a quick search (section 1.1) is the machine used as the generator to supply exact counterfactual ground truth, and the measurements of cheaper receipts that this makes possible.
How it works, in plain terms
In one sentence: it is a picture-maker that builds every image out of small pieces of the paintings it was given, and keeps an exact receipt of which painting each piece came from.
Figure A. Top two rows: images the machine made (32×32 pixels, enlarged). Bottom two rows: some of the paintings it was given. The outputs are painterly mosaics built from the paintings' colours and brushwork; they do not depict anything. U1–U8 start from pure static; D1–D4 start from a noised copy of a known painting.
Figure B. The four steps.
- Library. Each training image is cut, virtually, into every small square patch it contains (thousands per image), and every patch keeps its owner's name.
- Generate. An image starts as random static. Over 20 rounds, every pixel looks at its small neighbourhood, asks which library patches that neighbourhood resembles, and moves toward a weighted blend of the best matches. Early rounds use large patches (they settle the layout); late rounds use small ones (they settle texture).
- Receipt. Because every blend is explicit, we know exactly how much of each pixel came from each owner. Summed over the image, that is the image's receipt. Nothing is estimated after the fact.
- What if? To ask "what would this image be without painting X?", remove X's patches from the library and rerun from the same static. That takes seconds to minutes, and nothing is retrained.
Figure C. A real receipt. Middle: one pixel of output U1, averaged over the 20 rounds; it draws on the commons, a Met work and five different Mont Sainte-Victoire paintings (at the final round it settles mostly on two of them: Cézanne F 45%, Cézanne I 37%). Right: the whole image. The pixel was chosen, for illustration, as a central pixel with a well-mixed receipt.
Why it matters. Real image generators (Midjourney, Flux and the like) cannot say what an image would have been without a given contributor, so no method for paying contributors can be checked against the truth. This generator can say it exactly. That makes it a test bench: we can measure how wrong the cheaper payment methods are.
Why the trained networks in study 2. Kamb and Ganguli showed that a simple, real class of image generators (convolution-only diffusion networks) behaves almost exactly like this machine. Study 2 trains such networks on our paintings to check that, and then asks the harder question: does the machine's receipt also describe them?
Reading this paper by role
| If you are | Read first | The one thing to take away |
|---|---|---|
| Leadership or business | The precondition; key findings; design rules; section 6 | In a stress-test corpus, paying by usage sent 18–28% of contributors' money to different people than Shapley credit did, and the definition of "contribution" moves money as much; both must be decided in the contract, and exact answers need a boxed generator |
| An artist or rights holder | How it works; sections 3.4–3.6 | Works in a series are underpaid by both cheap rules (about 78% and 71% of what counterfactual credit says); a duplicated work is paid about double; an image you supply at generation is a separate thing to be paid for |
| An engineer or ML researcher | Sections 2 and 4; appendices | The ledger is exact; Kamb and Ganguli reproduces independently; compare networks at fixed inputs, not on finished images |
| An investor or partner | Abstract; sections 6 and 7 | Small-scale, carefully controlled evidence for a settlement design; not a claim about production models |
Terms are explained in plain language in Appendix D.
Key findings
| # | Finding | Type | Key number |
|---|---|---|---|
| 1 | In the stress-test corpus, paying by usage diverges from Shapley credit (pixel value function) | F, X | TV 0.18–0.28, median 0.254; 4 outputs; interval [0.241, 0.269] covers Shapley noise only |
| 2 | The divergence persists under four value functions, which also disagree with each other almost as much: the definition of contribution is a first-order choice | P | usage vs Shapley 0.20–0.39; between value functions 0.17–0.35 |
| 3 | The divergence is not an artefact of the stress-test corpus: a corpus of all-distinct works diverges as much as the stress test | F | 0.25 (all-distinct, index 0.10) vs 0.25 (stress test, index 0.19); the five-corpus trend test is uninformative (predictor range too narrow) |
| 4 | No cheap rule tracks Shapley credit | F | rank agreement 0.45–0.70 against a 0.70 bar |
| 5 | Both cheap rules under-credit a series once substitutes exist (a step, not a dose-response) | P, F | series gets 0.78 of its Shapley share under usage, 0.71 under leave-one-out (Met works 1.18 and 1.29); across corpora leave-one-out 1.19 → 0.71, trend p = 0.002 |
| 6 | A generation-time source image is outside any corpus ledger (qualitative; the ratio is not an effect size) | P | corpus attribution never ranks the source first |
| 7 | Similarity attribution mostly detects input reuse | P | finds the img2img source 4/4; rank agreement with Shapley credit 0.55 |
| 8 | A duplicate nearly doubles a work's pay (an identity) and amplifies its pull on outputs | D, P | pay ×1.93; influence ×3.07 (squared units) |
| 9 | Trained convolution-only networks behave like the machine (Kamb and Ganguli reproduced) | F | r 0.89–0.94 (paper, ResNets: 0.90–0.96) |
| 10 | The machine's group credit points the right way in trained networks, single-step; weak (≈1% of variance); consistent with Zhao et al. (2025) | F, R | p = 0.0010 (seed 0), 0.0002 (seed 1) |
| 11 | Retraining one network per condition is too noisy as ground truth at this scale | F, P | noise 0.008–0.059 vs predicted effects 0.0003–0.0019 |
Design rules for settlement
All seven rules assume a boxed generator, except where marked general, meaning they apply to any attribution or settlement scheme.
- Declare what counts as contribution, in the contract. General. (Finding 2.) Whether contribution means pixels, composition, structure or palette moves money between contributors about as much as the choice between usage and Shapley does. It is a policy decision, not a measurement, and it must be made before settlement.
- Allocate each output's payment by Shapley credit under that declaration, not by usage. (Findings 1, 4, 5.) Usage weight diverged from Shapley credit in every setting tested. Both cheap rules under-credit a series by roughly the same amount (a series received 78% of its Shapley share under usage and 71% under leave-one-out), so the choice between them is not the fix: only a rule that shares credit among substitutes, as Shapley does by construction, pays a series what the counterfactual says it is worth. In a boxed generator Shapley credit is computable exactly; for parametric models it has been estimated by approximate retraining (Lin et al., 2025), and Shapley-based royalties have been proposed (Wang et al., 2024).
- Receipt generation-time inputs as their own party. General. (Finding 6.) A reference or source image supplied at generation largely fixes the output, and corpus attribution cannot recover its role. Record it at generation and pay it by contract.
- Consolidate duplicates before settlement. General. (Finding 8.) Every rule pays a duplicated work close to double, and duplication biases generation toward that work.
- Do not use similarity as a measure of contribution. General. (Findings 4, 7.) It mostly measures resemblance to the input; the two similarity methods tracked Shapley credit worse than the usage and removal rules.
- Use an exact-ledger generator as a referee for attribution vendors. General. (Findings 9, 10.) It supplies exact ground truth, under a declared value function, that attribution methods can be scored against; trained convolution-only networks behave like it, and its credit points the right way in them (weakly; also Zhao et al., 2025).
- Evaluate attribution at fixed inputs, not on finished images. General. (Findings 10, 11.) Retraining noise swamped group-removal effects on sampled images; single-step comparisons with shared initialisation resolved them.
These refine, rather than replace, the rules of the earlier memo (2026-09-11): "pay on consumption" decides how much an output earns; rule 2 here decides who shares it. "Commons pays zero" matches our always-present commons, which is excluded from the comparison. "Don't sell receipts as influence" still holds: Shapley credit here is contribution within the declared generator, under a declared value function, not influence through training on some other model.
1. Background
The settlement problem. A royalty scheme for a generative model needs, for each output, a split of credit across the owners of the material the model draws on. Two families of method are common: participation or usage receipts (how much of a contributor's material the model drew on) and similarity attribution (how much the output resembles each contributor's works). For a trained network neither can be validated, because the true counterfactual (the output had a contributor been absent) is unavailable except by retraining.
Declared memory versus training data. Attribution against the training set of a large trained network is not generally identifiable. Attribution against a declared memory that the generator consults at generation time is an accounting identity: the ledger can be exact. This is the premise of a boxed generator.
Kamb and Ganguli (2025). A convolution-only diffusion network sees only a local neighbourhood of each pixel and treats all positions alike (locality and equivariance). The authors show that the best denoiser under exactly those two constraints has a closed form, an equivariant local score (ELS) machine: at each step every pixel moves toward a posterior-weighted blend of the centre pixels of all training patches that could have produced its neighbourhood. After calibrating one time-dependent patch size, the machine predicts trained ResNet and UNet outputs at median correlation 0.85–0.96 (ResNets 0.90–0.96) on MNIST, FashionMNIST, CIFAR10 and CelebA (32×32), and a UNet with attention at 0.77.
Our move. Use the machine as the generator. Its ledger is then exact by construction, and exact counterfactuals become cheap, which yields ground truth for testing cheaper receipts (study 1). Then ask whether that ledger also describes trained networks (study 2).
1.1 Related work, and what appears to be new here
Based on a quick web search on 2026-09-19 (about 20 minutes), not a systematic review. Novelty is therefore not established; the items below are the closest work found.
Counterfactual credit for data is standard. Leave-one-out and Shapley credit for training data are well established (Ghorbani and Zou, 2019). For diffusion models, ground truth is usually obtained by retraining on many subsets and scoring methods by the linear datamodeling score, LDS (Park et al., 2023; Zheng et al., 2024), or by fine-tuning a model on known images (Wang, S.-Y. et al., 2023).
Shapley credit and royalties for generative models have been proposed. Lin, Lu, Kim and Lee (2025) estimate Shapley values for diffusion-model data contributors by approximating retraining with pruning and fine-tuning. Wang, J. T. et al. (2024) propose compensating copyright owners in proportion to Shapley-based contributions to generated content.
The analytic machine has been used for attribution. Zhao et al. (2025, "NDA") build a patch-level attribution score on Kamb and Ganguli's analytic score, requiring no gradients or retraining, and validate it against 64 retrained networks (LDS) on CIFAR-10 and CelebA, where it performs comparably to gradient-based methods. Briq et al. (2026) give closed-form attribution to clusters of training data in flow-matching models, checked by leave-one-cluster-out retraining. Niedoba et al. (2024) developed a related patch-based local score model concurrently with Kamb and Ganguli.
Attribution at scale may be impossible. Dai and Gifford (2026, Nature Communications) show that outputs of diffusion models trained on large datasets are often unattributable to individual training examples, because important features are spread redundantly across the data. This supports the premise of a boxed generator with a declared memory.
What we did not find. (1) The analytic machine used as the generator itself, so that exact removal and Shapley credit cost a rerun rather than retraining and serve as ground truth. (2) The measurements this enables: usage weights diverging about a quarter from counterfactual credit, leave-one-out under-crediting substitutes, the unledgered generation-time input, and duplicates. Study 1's usage result is directly relevant to weight-based scores such as NDA, though NDA's score is not identical to our usage rule. Study 2 is an independent, small-scale confirmation, with a different test design, of what Zhao et al. showed at larger scale: that the machine describes trained networks.
2. The generator and its ledger
At a reverse-diffusion step with noise level ᾱ, for output pixel x with neighbourhood patch φ(x) and each allowed training patch p (owned by contributor c(p)), the posterior weight is a softmax over patches of −‖φ(x) − √ᾱ·p‖² / (2(1 − ᾱ)). The denoised pixel is the weighted mean of patch centres. Grouping by owner:
m(x) = Σ_c W_c(x) · m_c(x), with Σ_c W_c(x) = 1
where W_c is contributor c's total posterior weight at pixel x and m_c their weighted centre. At the final step the output is m, so every output pixel decomposes exactly by contributor. Removing a contributor means masking their patches and rerunning from the same noise; no retraining is involved. Border pixels use only patches with the same zero-padding pattern (the paper's boundary-broken variant).
Implementation. Written from the paper's equations (the authors' code was consulted afterwards and matches: same computation and border handling). Exact: no top-k, no subsampling. Float32 on the Mac GPU; contributor weights sum to 1 within 5×10⁻⁶. Sixteen self-tests pass [D, X]: agreement with an independent float64 brute-force implementation (≤ 4×10⁻⁶), the ledger identity (0 error), the paper's two-image black/white example, memorisation by the ideal-score variant, the duplicate identity (section 3.6), local consistency approached under step refinement (the paper's Theorem B.3), and coalition masking equal to physical removal.
| Parameter | Study 1 | Study 2 |
|---|---|---|
| Image size | 32×32 RGB, pixels in [−1, 1] | same |
| Sampler | DDIM, 20 steps, cosine schedule | same |
| Patch side by step | 15 → 3, declared (not fitted) | fitted to the pilot network (section 4.2) |
| Corpus | 244 training images | same |
Corpus: a stress test by design. The corpus is built to contain the failure modes a settlement system must handle: a series of nine paintings of one subject, a verbatim duplicate under a second owner, and a contributor of generic filler. Real training data is also heavily redundant, but this corpus leans on redundancy deliberately, so every magnitude measured on it has to be checked against a less redundant corpus; section 3.9 does that, and finds an all-distinct corpus diverging as much as the stress test, while the leave-one-out failure appears only once close substitutes exist. Rights-clear: Met Open Access paintings (CC0) and Cézanne's Mont Sainte-Victoire series (public domain). Four 32×32 training images per work (the centred square plus three half-side crops). 26 paid contributors: 9 Cézanne Mont Sainte-Victoire paintings (A–I), 15 Met works, a verbatim COPY of Cézanne A held by a separate owner, and a GENERIC contributor of four blurred colour fields. A commons of 35 further Met works is always present and paid by no rule in the comparison, standing in for a model pool.
3. Study 1: what cheaper receipts get wrong
3.1 Design (frozen 2026-09-18 06:44:34)
Twelve outputs: eight from pure noise (U1–U8) and four img2img derivatives (D1–D4), each starting from a noised view of a known work at t = 0.7. Six credit rules, shares computed over the 26 contributors:
| Rule | Definition |
|---|---|
| R1 final-step usage | contributor's posterior weight at the final step, averaged over pixels |
| R2 trajectory usage | the same, averaged over all steps |
| R3 exact leave-one-out | mean squared pixel change when the contributor is removed (same noise) |
| R4 Shapley (reference) | permutation Shapley of v(S) = −MSE(output of commons ∪ S, full output); 10 permutations, 5 antithetic pairs |
| R5 global similarity | rank by minimum L2 distance from the output to the contributor's images |
| R6 patch retrieval | share of output pixels whose 5×5 patch has its nearest training patch in the contributor's images |
Pre-registered hypotheses: H1 usage misallocates (median TV(R2, R4) over U1–U4 ≥ 0.20 passes, ≤ 0.10 fails); H2 generic material over-credited ≥ 2×; H3 Shapley ranks a derivative's source first on ≥ 3 of 4; referee: a rule is consistent with Shapley if median Spearman ≥ 0.7 and source top-1 ≥ 3 of 4. Run: 12 outputs, 312 exact removal reruns, 2,008 Shapley trajectories, about 3.3 hours.
Shapley is a declared axiom, not ground truth. Shapley credit is the unique allocation satisfying a set of fairness axioms for a given value function v(S); change v and the credit changes (section 3.8). The machine makes Shapley credit exact, not correct. Throughout, "counterfactual credit" and "Shapley credit" mean Shapley credit under the stated value function, and every "misallocation" is relative to that declared reference.
3.2 In the stress test, usage diverges from Shapley credit [F, X]
Figure 1. Total-variation distance between usage shares and Shapley shares per output. TV is the fraction of money that would be paid to different contributors.
Figure 2. Four outputs in detail. Left: the output. Middle: which contributor dominates each pixel at the final step (the large light-blue area is the always-present commons). Right: each contributor's share under usage (blue), exact leave-one-out (orange) and Shapley (green). Where the bars disagree is where usage receipts would pay the wrong people. D1 and D3 are img2img outputs whose named source is discussed in section 3.5.
H1 passes: median TV(R2, R4) = 0.254 over U1–U4 (per output 0.177, 0.244, 0.282, 0.264); bootstrap 90% interval [0.241, 0.269]. Final-step usage is worse: 0.17–0.35 (median 0.31). Caveats that bound the claim: the interval covers Shapley estimation noise only, not the choice of outputs; the verdict rests on four outputs, one below the line; Shapley rests on one declared value function; and the always-present commons, which takes 0.52–0.62 of usage payment, is outside the comparison. A supporting check against exact removal over all eight unconditional outputs gives median TV 0.20. All study 1 numbers were recomputed from the raw saved tensors by independent code and match to four decimals [X].
Read this as an existence proof plus a mechanism rather than an estimate for natural corpora. The corpus is a stress test, so the first question is whether the number is a product of that construction. Section 3.9 answers it with a corpus of all-distinct works, no series and no duplicate: it diverges by 0.25, the same as the stress test.
3.3 No cheap rule tracks Shapley credit [F]
Figure 3. Median Spearman rank agreement of each rule with Shapley over eight outputs.
| Rule | Spearman with Shapley | Money reallocated (TV) | Source found, D1–D4 |
|---|---|---|---|
| trajectory usage | 0.70 | 0.22 | 3/4 |
| final-step usage | 0.68 | 0.23 | 1/4 |
| exact leave-one-out | 0.60 | 0.24 | 0/4 |
| global similarity (vendor-style) | 0.55 | n/a | 4/4 |
| patch retrieval (vendor-style) | 0.45 | 0.34 | 1/4 |
No rule meets the bar. Shapley's own ranking is stable (split-half reliability 0.82–0.98 [P]), so the disagreement is not estimator noise. On unconditional outputs U2–U4 every cheap rule, exact leave-one-out included, has rank agreement between −0.2 and 0.5. This is relative to the pixel value function; section 3.8 shows the disagreement persists under three others.
3.4 Leave-one-out under-credits substitutes: known theory, quantified [P]
Figure 4. Leave-one-out share divided by Shapley share, by group.
With nine paintings of the same mountain, removing any one is covered by the others. Leave-one-out gives the Cézanne works (with the copy) 0.68–0.96 of their Shapley share on U1–U4 and the Met works 1.12–1.37.
Usage under-credits the series too, by about as much. Median over U1–U4, the Cézanne group receives 0.78 of its Shapley share under usage and 0.71 under leave-one-out, while the Met works receive 1.18 and 1.29. Neither cheap rule is the kinder one: both move money from works with close substitutes to works without them, leave-one-out slightly more. The fix is not to choose between them but to use a rule that shares credit among substitutes, which is what Shapley credit does by construction. The leave-one-out half of this is the textbook redundancy failure of removal-based credit; with nine paintings of one mountain it is expected. What is added is its size in a generative setting.
3.5 A generation-time source image is outside any corpus ledger [F for H3; P for the explanation]
H3 fails: Shapley over the corpus ranks the true source first in 0 of 4 derivatives. The post hoc check explains why: the source enters through the start image, which the corpus counterfactual holds fixed.
Figure 5. Swapping the start image for a different painting versus removing the most influential contributor.
| Output | Source | Start-image swap (MSE) | Largest removal (MSE) | Ratio | Source's own removal rank |
|---|---|---|---|---|---|
| D1 | Cézanne B | 0.260 | 0.00095 | 274× | 3 of 26 |
| D2 | Met DP-12952-001 | 0.192 | 0.00025 | 779× | 3 of 26 |
| D3 | Cézanne A | 0.568 | 0.00119 | 476× | 8 of 26 |
| D4 | Met DP-14201-001 | 0.044 | 0.00072 | 61× | 3 of 26 |
The ratios are not effect sizes. Img2img fixes most of the output by construction, the two interventions differ in kind, and the replacement paintings were chosen by hand; the size of the ratio is close to a tautology. The qualitative point is what stands: a generation-time input sits outside any corpus ledger. Global similarity found the source 4 of 4 times because the output still resembles its input, which suggests that similarity attribution largely detects input reuse rather than corpus contribution.
3.6 Duplicates: an identity, plus its effect on outputs [D, P]
At a single step, if a work holds posterior share a, adding an exact copy under another owner gives the pair 2a/(1 + a) (self-test T5: a corpus with a duplicate equals the original at prior weight 2). Measured on U1–U8 [P]:
Figure 6. Pay and influence ratios with the copy present versus absent.
With the copy present, the work and its copy are paid ×1.93 (1.82–1.97) what the work alone is paid without it; the work's influence on outputs (the change when it is removed entirely) rises ×3.07 (2.25–3.87) in squared units, about ×1.75 in RMS terms. Duplication amplifies a work's pull on generation as well as its pay; whether pay over- or under-states that depends on the unit of influence. The pay result follows from the identity; the measurement adds the influence side.
3.7 Generic material (H2 fails) [F]
The blurred colour-field contributor received 1.42× its Shapley share under usage at the median (range 0.78–3.04), short of the pre-declared 2×. It accounts for at most about one point of the per-output gap; the misallocation is spread across real contributors. Only one generic contributor was tested.
3.8 What counts as contribution: four value functions [P]
Prompted by external review. All 2,008 coalition outputs from study 1 are saved, so Shapley and leave-one-out credit can be recomputed under other value functions without new generation. Four were declared before computing (VALUE_FUNCTION_PROTOCOL.md): pixel (the original), blur (coarse composition), edges (structure) and colour (palette, ignoring layout).
| Value function | Median TV(usage, Shapley), U1–U4 | Median Spearman(usage, Shapley), 8 outputs | Median Spearman(leave-one-out, Shapley) |
|---|---|---|---|
| pixel | 0.25 | 0.70 | 0.60 |
| blur | 0.39 | 0.29 | 0.22 |
| edges | 0.20 | 0.66 | 0.62 |
| colour | 0.36 | 0.42 | 0.44 |
Between the value functions themselves, Shapley credit differs by TV 0.17–0.35 (median over U1–U4 for each pair). Two readings follow. First, usage diverges from Shapley under every value function (≥ 0.20), so that finding does not hinge on the pixel choice; the declared threshold for "does not hinge" was 0.15. Second, and more important for settlement, the definition of contribution moves money between contributors about as much as the choice between usage and Shapley. Usage is not uniquely wrong; it is one more implicit definition, and the definition has to be chosen and declared.
3.9 The redundancy dial [F]
The stress-test corpus makes magnitudes corpus-specific. To see how they depend on redundancy, the same machinery ran on five corpora (frozen 2026-09-19 16:31:43; DIAL_PROTOCOL.md), all with 26 contributors and the same commons: L0 all-distinct works; L1–L3 with 3, 6 and 9 Mont Sainte-Victoire paintings; L4 the study-1 corpus (9 paintings, a verbatim copy, generic filler). Six outputs each; usage, exact leave-one-out and Shapley (8 permutations) per output.
Figure 7. Left: usage-versus-Shapley divergence per output (dots) and median (diamond) by corpus; L0 (all-distinct) and L4 (stress test) bracket the redundancy range actually achieved and agree. Right: what each cheap rule pays the Cézanne series, as a fraction of its Shapley share (medians); both over-credit a short series and under-credit a long one. SI: substitutability index, the fraction of each contributor's patches with a close match in another contributor's work (DIAL_PROTOCOL.md).
| Corpus | Substitutability index | Median TV(usage, Shapley) | Median TV(leave-one-out, Shapley) | Leave-one-out ÷ Shapley, Cézanne group |
|---|---|---|---|---|
| L0 | 0.100 | 0.247 | 0.302 | n/a |
| L1 | 0.115 | 0.201 | 0.253 | 1.19 |
| L2 | 0.105 | 0.194 | 0.233 | 0.79 |
| L3 | 0.099 | 0.219 | 0.272 | 0.83 |
| L4 | 0.188 | 0.252 | 0.287 | 0.71 |
The manipulation mostly failed, so the trend test is uninformative. The declared redundancy measure barely moved across L0 to L3 (0.100, 0.115, 0.105, 0.099: a range of 0.017); only the verbatim copy raised it (0.188). Nine paintings of one mountain are not nine substitutable patches at 5×5. The pre-registered trend test therefore regresses against a predictor with almost no range, and its result (p = 0.22, NOT SUPPORTED) is low power, not evidence of independence. We do not claim the divergence is independent of redundancy.
What the sweep does support is a level comparison. Two corpora sit at opposite ends of the redundancy range actually achieved: L0 (index 0.10, all-distinct works, no series, no duplicate, no filler) and L4 (index 0.19, the stress test). Their divergences are the same: 0.247 and 0.252. So the headline magnitude is not manufactured by the stress-test construction, which is what this check was built to find out, and by the declared rule the L0 level reads as "substantial divergence without redundancy". Under the other value functions at L0: blur 0.33, edges 0.25, colour 0.41.
Secondary: the substitution effect appears as a step, then saturates. Leave-one-out pays the Cézanne group 1.19 of its Shapley share with three paintings, 0.79 with six, 0.83 with nine and 0.71 with nine plus a verbatim duplicate (ordered-trend p = 0.002). The move is almost all in the first step, when substitutes first appear; six to nine paintings is flat or slightly up (0.79 to 0.83), and the last drop adds a duplicate rather than more series length, so it confounds the two. Read it as "the shortfall appears as soon as close substitutes exist, then saturates", not as a dose-response in series length. Usage behaves the same way (1.34, 0.75, 0.87, 0.77 for 3, 6, 9 and 9-plus-a-duplicate paintings, against leave-one-out's 1.19, 0.79, 0.83, 0.71): with a short series both rules over-credit it, and both turn to under-crediting once the series is long enough. Close substitutes therefore control how the cheap rules treat a series (finding 5); whether redundancy also moves the overall usage-versus-Shapley gap (finding 1) is untested here, for the reason above. TV(leave-one-out, Shapley) shows no trend (p = 0.53).
Why the manipulation failed, and what would fix it. Post hoc, the index at larger patch sizes does rise with the series (15×15 patches: 0.10, 0.13, 0.13, 0.14, 0.22), the scale the machine uses early in generation and where the leave-one-out effect appears. So subject-level similarity is not the same thing as patch-level substitutability: nine views of one mountain share composition, not texture. A sweep that actually tests the redundancy question needs corpora whose measured index spans a real range at the scale generation uses (graded augmentations or synthesised near-duplicate patches, carrying the 5×5 index from about 0.1 to 0.5). That is cheap in this setup and is the obvious next experiment; it is not this one.
4. Study 2: does the ledger describe trained networks?
4.1 Training
| Item | Setting |
|---|---|
| Architecture | ResNet, 8 zero-padded 3×3 conv layers (receptive field 17×17), residual, no normalisation; the authors' recipe |
| Optimiser | Adam, lr 10⁻⁴, batch 128, decay 0.999965 per step, noise-prediction loss |
| Width | 128 channels for all main runs; one 256-channel run (the paper's width) for a width check |
| Training | full corpus 30,000 steps (round 1); leave-group-out networks 10,000 steps (length chosen by a frozen rule) |
| Groups removed | five groups of four contributors formed by rule from study 1 removal effects, Cézanne B–E, and the COPY alone |
| Seeds | round 2 at start-weight seed 0; replication at seed 1; controls with nothing removed |
| Compute | about 25 minutes per 10k-step network |
4.2 The paper reproduces [F]
Networks and the calibrated machine were run from the same 100 noise seeds:
| Network | r, 20 steps | r, 150 steps | Near-copies of training images |
|---|---|---|---|
| half_s0 at 10000 steps | 0.937 | 0.938 | 0/100 |
| half_s0 at 20000 steps | 0.930 | 0.944 | 0/100 |
| half_s0 at 30000 steps | 0.918 | 0.940 | 0/100 |
| half_s1 at 30000 steps | 0.924 | 0.938 | 0/100 |
| half_s2 at 30000 steps | 0.887 | 0.937 | 0/100 |
Figure 8. Same starting noise, eight examples. Top row: the machine (pilot, before calibration). Middle and bottom: the trained network at 20 and 150 sampling steps. Layout and colour carry over; the network paints larger, smoother regions.
Agreement r = 0.89–0.94 against the paper's ResNet range of 0.90–0.96. The paper's figures label this number r², but its released code computes the Pearson correlation r (checked against github.com/Kambm/convolutional_diffusion, MIT licence), and its tables label the column "Corr."; our r uses the same computation, so the comparison is like for like. Width does not matter here: the 256-channel network is never better by more than 0.004. No memorisation: 0 of 100 outputs near a training image at every checkpoint. Clipping the first sampling step (needed by our noise schedule) changes outputs negligibly (r ≥ 0.98 against the authors' unclipped arrangement).
Figure 9. Machine patch size by noise level, fitted to our network with the authors' procedure, against their CIFAR10 fit: a second independent bearing on the paper's coarse-to-fine finding.
4.3 Credit on finished images: underpowered [F, P]
Round 2 (frozen 2026-09-18 23:36:36) retrained seven networks without their groups plus two controls with nothing removed, and compared the machine's predicted change in finished images with the actual change. The frozen rule returned CONFIRMED (5 of 7 groups). That verdict is not credible and is not relied on: the two controls, identical recipes with nothing removed, disagree with each other as much as the groups differ from them, and the frozen bootstrap did not include network-to-network variation. A control-free specificity test gives p = 0.075 [P]. Honest reading: underpowered.
Figure 10. Why: on finished images, retraining noise (0.008–0.059) dwarfs the predicted group effects (0.0003–0.0019). Single-step comparisons with shared start weights cut the noise to 0.0005–0.0007 (networks with other start weights: 0.0023–0.0024).
4.4 Credit at fixed inputs: specific, and replicated [F, R]
Test (frozen 2026-09-19 05:36:02). Fixed inputs: the calibrated machine's own full-corpus reverse trajectories for 100 noise seeds, at the 15 steps with 0.15 ≤ t ≤ 0.85. At each input, the machine's predicted change in the noise prediction when group h is removed (exact, from the per-contributor ledger terms) is correlated with each retrained network's actual change. M[h, g] averages these correlations over steps. Statistic S = mean diagonal − mean off-diagonal; row and column effects cancel. Exact one-sided permutation test over all 5,040 assignments of predictions to networks. Its design was chosen after round 2 (disclosed); it measures different quantities, so it is a new test.
| Seed 0 (frozen) | Seed 1 (frozen replication, 2026-09-19 06:35:24) | |
|---|---|---|
| Specificity S | 0.0389 | 0.0590 |
| Exact p | 0.0010 | 0.0002 |
| Verdict | SPECIFIC | REPLICATED |
| Own-prediction rank per network (1 = best of 7) | 1, 1, 6, 1, 7, 1, 3 | 1, 2, 2, 1, 7, 1, 4 |
| Group-specific component positive (row/column effects removed) | 6 of 7 (-0.004 to 0.071) | 7 of 7 (0.022 to 0.105) |
Figures 11–12. Prediction-by-network matrices with row and column effects removed (the test statistic is unchanged by this). Bold diagonal: each network against its own group's prediction.
Figure 13. Own-group agreement by noise level, both seeds: negative in the early, noisy steps and positive from mid-generation on, where patches are small and each patch's owner matters most. Seed 0 rises to the end; seed 1 peaks mid-generation.
Robustness [P]. Dropping any one group leaves p between 0.0014 and 0.0069 (0.0014 is the smallest possible with six). Measuring against the mean of the same-initialisation networks instead of one reference network reduces the row effects' share of matrix variance from 85% to 18% and raises S to 0.076 (p = 0.0026): the large row effects were one reference network's quirks. The group-specific patterns of the two seeds correlate at 0.38.
Limits. The group-specific correlations are weak (at most about 0.07 for seed 0 and 0.11 for seed 1); magnitudes do not match (networks change 0.0005–0.0010 against predicted 0.00002–0.00035, and most of each network's change is retraining noise); one corpus, one architecture, 32×32. A correlation of 0.11 means the prediction accounts for about 1.1% of the variance in a network's change: the p-values establish that the direction is real, not that it is useful. Study 2 is directional, replicated, and not yet quantitatively usable; it confirms at small scale, with a different test, what Zhao et al. (2025) showed at larger scale.
4.5 What study 2 establishes
| Question | Answer |
|---|---|
| Do trained convolution-only networks behave like the machine? | Yes: r 0.89–0.94, the paper reproduced independently |
| Does the machine's group credit point the right way in them? | Yes, in weak form, replicated: p = 0.0010 and 0.0002, single-step, about 1% of variance; consistent with Zhao et al. (2025) at larger scale |
| Does it show up in finished images? | Unresolved: underpowered at this scale |
| Does it get the size of credit right? | No evidence |
5. Evidence ledger
| Claim | Type | Evidence | Weak point |
|---|---|---|---|
| Usage diverges from Shapley credit, in the stress test and in an all-distinct corpus | F, X | TV 0.18–0.28, median 0.254; all-distinct corpus 0.247 vs stress test 0.252 | 4–6 outputs per corpus; Shapley-noise interval only; commons excluded; the redundancy trend test had too little range to test dependence |
| The definition of contribution moves money as much as usage vs Shapley | P | 0.17–0.35 between value functions | four value functions; post hoc |
| No cheap rule tracks Shapley credit | F | Spearman 0.45–0.70 < 0.70 | one corpus; pixel value function (others in 3.8) |
| Both cheap rules under-credit a series; the shortfall grows with it | P, F | series receives 0.78 (usage) and 0.71 (leave-one-out) of its Shapley share; across corpora 1.19 → 0.71, p = 0.002 | known theory, quantified; redundancy-dependent by construction |
| Generation-time input is outside any corpus ledger | P | source never ranked first by corpus Shapley | qualitative; the swap ratio is not an effect size |
| Similarity detects input reuse | P | 4/4 sources; Spearman 0.55 | inference from two measurements |
| Duplicates nearly double pay and amplify pull | D, P | ×1.93 pay, ×3.07 influence | pay side is an identity; influence comparison is unit-dependent |
| Exact counterfactuals are cheap | X | removal of 26 contributors ≈ 2.6 min per output; Shapley ≈ 24 min | engineering fact |
| Trained networks behave like the machine | F | r 0.89–0.94 | confirmation of prior work |
| Machine credit points the right way in trained networks | F, R | p 0.0010, 0.0002 | direction only; ≈1% of variance; single-step; confirms Zhao et al. |
| One retrained network per condition is too noisy at this scale | F, P | noise ≫ effect (Figure 10) | small networks, small corpus |
6. What can and cannot be said
| Supported | Not supported |
|---|---|
| Usage receipts diverged from Shapley credit by 18–28% of contributors' money in the stress test, and by the same amount in a corpus of all-distinct works | That the divergence is independent of redundancy: the sweep could not move patch-level redundancy far enough to test that |
| What counts as contribution (the value function) moves money about as much as usage versus Shapley does, so it must be declared | That any one value function is the correct measure of contribution |
| Leave-one-out and similarity attribution disagree materially with Shapley credit here | That any attribution vendor's product is wrong (none was tested) |
| A generation-time source image must be receipted separately | Anything about copyright, fair use or legal attribution |
| Convolution-only diffusion networks trained on this corpus behave like the machine, and its credit points the right way in their single-step behaviour | That the machine gets credit magnitudes right, or that this holds for finished images |
| An exact-ledger generator can serve as a referee for attribution methods | That results transfer to attention or latent generators (Midjourney, Flux, Stable Diffusion XL) |
| That Shapley credit is ground truth, or equals authorship, rights or fairness: it is a declared axiom | |
| That this is a better attribution method for parametric (commercial) generators: it presupposes a boxed generator | |
| That the methods or findings are new: only a quick search was done, and closely related work exists (section 1.1) |
7. Limitations and threats to validity
- Scale. 32×32 images, 244 training images, 26 paid contributors; convolution-only architectures.
- Constructed corpus, and a failed manipulation. The study 1 corpus is a stress test built around substitution, duplication and filler. Section 3.9 shows an all-distinct corpus diverging as much, so the headline is not manufactured by that construction. But the redundancy sweep moved the measured index by only 0.017 across L0–L3, so its trend test is uninformative: we cannot say whether the divergence depends on redundancy, only that it is present without it.
- Reference relativity. Shapley credit depends on the declared value function; four value functions disagree with each other by 0.17–0.35 (section 3.8). Every "misallocation" is relative to a declared axiom.
- Precondition. The settlement proposal presupposes a boxed generator with a declared memory; commercial generators are parametric.
- Few outputs. The primary study 1 test rests on four outputs; its interval covers estimator noise only.
- Post hoc results. Findings 2 and 5–8 include post hoc analysis; they are exploratory and labelled, and change no verdict.
- Design after seeing data. The single-step test was designed after round 2 (frozen before it ran; replicated).
- Machine as generator. Its outputs are painterly patch mosaics, not depictions; product quality is a separate question.
- Retracted verdict. Round 2's frozen CONFIRMED is recorded but not relied on (section 4.3).
- Novelty. A quick search found closely related work (section 1.1); no systematic literature review was done.
8. Next steps
- Referee experiment. Score real attribution methods against the machine's exact ground truth: gradient-based scores (TRAK and relatives), embedding similarity, and the machine-based NDA score of Zhao et al. (2025).
- A redundancy sweep that works. Corpora whose measured substitutability spans a real range at the scale generation uses (graded augmentations, synthesised near-duplicates), so the trend test has power. Cheap here: no retraining, only reruns.
- Literature review. A systematic search of work citing Kamb and Ganguli (2025) and Niedoba et al. (2024) before any novelty claim.
- Magnitudes and finished images. More start-weight seeds per condition, larger removals, or shared-randomness retraining.
- Asset-level receipts. Combine with the earlier finding that receipts must live at the scale of the owned asset.
- Beyond convolution. Attention layers break locality; the paper reports partial correspondence (0.77). Untested here.
Appendix A. Definitions
- Total variation between share vectors s and s′: TV = ½ Σ_c |s_c − s′_c|, the fraction of money paid to different contributors.
- Shapley value of contributor c: the average over orderings of v(S ∪ {c}) − v(S), with v(S) = −MSE(output generated from commons ∪ S, full output) and v(∅) the commons alone. Estimated with 10 permutations (5 antithetic pairs), 251 trajectories per output; efficiency holds to 10⁻¹⁸; negative values (0–5 of 26 per output) clipped to zero in shares.
- Value functions (section 3.8). blur: squared difference after Gaussian blur, σ = 2 px; edges: squared difference of Sobel gradient magnitude of luminance; colour: squared difference of per-channel 8-bin histograms.
- Specificity statistic S = mean(diag M) − mean(off-diagonal M). Double-centring: M − row means − column means + grand mean. Row-effect share: var(row means) ÷ var(M), the fraction of the matrix's variance explained by row effects.
- Duplicate identity: with uniform prior per patch, duplicating a work whose patches hold posterior share a raises the pair's share to 2a ÷ (2a + (1 − a)) = 2a ÷ (1 + a).
Appendix B. Process record
- Study 1 protocol and code hash-frozen 2026-09-18 06:44:34; no deviations. The analysis script was written after the freeze, before any result was inspected.
- Study 2: running three GPU jobs at once caused Metal command-buffer errors; affected outputs were quarantined and rerun one job at a time with automatic abort on GPU error or NaN. The 256-channel network was stopped at 21,500 of 30,000 steps; the width decision uses its 10k checkpoint. Recorded before the round 2 freeze.
- Round 2's frozen CONFIRMED was examined and not relied on (section 4.3). A mistaken interim explanation (unmatched controls) was corrected in the results record.
- All frozen files were unchanged at their analyses (freezes 2026-09-18 06:44:34, 2026-09-18 23:36:36, 2026-09-19 05:36:02, 2026-09-19 06:35:24).
Appendix C. Reproducibility
All code and data in /patch-ledger/ (local). Key entry points:
tests/test_machine.py (self-tests); run_eval.py and analyze.py (study 1); verify_study1.py (independent
recomputation); train.py, phase3_checks.py, phase3_predict.py, phase3_analyze.py (study 2);
score_test.py and score_test_rep.py (single-step test and replication); build_paper.py (this document).
Result records: RESULTS.md, PHASE3_RESULTS.md, PHASE3B_RESULTS.md, CODE_COMPARISON.md.
Appendix D. Plain-language glossary
| Term | Meaning here |
|---|---|
| Diffusion model | An image generator that starts from random static and removes it step by step until an image appears |
| Patch | A small square of pixels (3×3 up to 15×15 here) cut from a training image |
| Usage (participation) receipt | Paying each owner by how much their patches were drawn on while the image was made |
| Counterfactual | What would have happened otherwise; here, the image made without a given owner's material |
| Leave-one-out | Credit by removing one owner at a time and measuring how much the image changes |
| Shapley value | Credit by averaging an owner's added value over every order in which owners could join; it shares credit fairly among owners who substitute for each other |
| Total variation (TV) | The fraction of money that two payment rules send to different people (0 = identical, 1 = completely different) |
| Frozen protocol | Rules and pass/fail thresholds written down and fingerprinted (hashed) before the data existed, so they cannot be adjusted afterwards |
| Post hoc | An analysis chosen after seeing results: useful for explanation, weaker as evidence |
| p-value | How often a result at least this strong would appear by chance if there were no real effect; p = 0.001 means about 1 in 1,000 |
| Seed / start weights | The random starting point of a network before training; different seeds give different networks from the same recipe |
| img2img | Generating from a noised copy of an existing image rather than from pure static |
| Commons | Material everyone may draw on that is paid by no rule here (like a model pool or public domain) |
| Noise prediction (score) | A network's single-step guess of what static is present in an image; the quantity compared in section 4.4 |
| Retraining noise | How much two networks trained by the same recipe differ by chance alone |
References
- Briq, R., Fried, O., Kamp, M. and Kesselheim, S. (2026). Tracing generated samples to training-data clusters in flow-matching models. arXiv:2608.30081.
- Dai, Z. and Gifford, D. K. (2026). Outputs of generative diffusion models are often unattributable. Nature Communications.
- Ghorbani, A. and Zou, J. (2019). Data Shapley: equitable valuation of data for machine learning. ICML 2019.
- Kamb, M. and Ganguli, S. (2025). An analytic theory of creativity in convolutional diffusion models. ICML 2025 (PMLR 267); arXiv:2412.20292. Code: github.com/Kambm/convolutional_diffusion (MIT).
- Lin, C., Lu, M., Kim, C. and Lee, S.-I. (2025). An efficient framework for crediting data contributors of diffusion models. ICLR 2025; arXiv:2407.03153.
- Niedoba, M., Zwartsenberg, B., Murphy, K. and Wood, F. (2024). Towards a mechanistic explanation of diffusion model generalization. arXiv:2411.19339.
- Park, S. M., Georgiev, K., Ilyas, A., Leclerc, G. and Madry, A. (2023). TRAK: attributing model behavior at scale. ICML 2023.
- Wang, J. T., Deng, Z., Chiba-Okabe, H., Barak, B. and Su, W. J. (2024). An economic solution to copyright challenges of generative AI. arXiv:2404.13964.
- Wang, S.-Y., Efros, A. A., Zhu, J.-Y. and Zhang, R. (2023). Evaluating data attribution for text-to-image models. ICCV 2023.
- Zhao, Y., Du, C., Zheng, X., Pang, T. and Lin, M. (2025). Nonparametric data attribution for diffusion models. arXiv:2510.14269.
- Zheng, X., Pang, T., Du, C., Jiang, J. and Lin, M. (2024). Intriguing properties of data attribution on diffusion models. ICLR 2024.
Provenance of references. Checked against the source pages on 2026-09-19: Briq et al., Dai and Gifford, Lin et al., Wang J. T. et al. and Zhao et al. (Zhao et al.'s full text was read for its method and evaluation). Kamb and Ganguli: read in full. Niedoba et al.: as listed in Kamb and Ganguli's references. Ghorbani and Zou, Park et al., Wang S.-Y. et al. and Zheng et al.: standard references cited from memory; Park et al. and Zheng et al. are also cited by Zhao et al.