Why trust it
A deterministic tool gives the same number every time, which is necessary but not sufficient — a reproducible number can still be measuring the wrong thing. The rules below are what keep a reading honest. Each was earned by a documented correction, and several of the sharpest results in the corpus began as caught errors.
Before the first command: look at the image and say what you see in words. Write down what you expect each key metric to say, and why. Name which metrics will have nothing to grip on this particular work. Carry no target.
Then the numbers test a hypothesis instead of seeding a narrative. A surprise you predicted against is a finding. A surprise rationalised afterward is a story. The lab book is built first; the critique does not begin until the reading is solid.
The R_spatial value was pre-registered as subadditive in the 0.2–0.7 band before any command ran. It landed inside the window. Because the prediction came first, the result is evidence rather than a number talked into meaning something.
A whole-image scalar has purchase on some works and floats on others. The same metric can grip one image and read only global statistics on the next. There are two ways to tell:
This is a per-work judgment, never a verdict stamped on a metric. The rule (finding I-15) is not "β is distributional" — it is "recognise, per work, when a number measures the tool instead of the work."
To ask whether a number needs the intact composition, destroy the composition in a controlled way and see if the number survives. The corpus uses three families, and their disagreement is itself informative — they bracket a question rather than settle it.
| Null family | What it keeps, what it destroys |
|---|---|
| Phase-scramble | Keeps the histogram and the power spectrum; destroys the arrangement. Permissive — it also destroys figures. |
| Patch-shuffle | Keeps the histogram and local texture; destroys the composition. |
| Figure-position shuffle | Permutes named figure boxes; asks whether the real arrangement beats rearrangements of its own cast. Conservative — it cannot vacate the centre. |
A value that survives both scrambles is distributional. A value that dies on both needs the intact composition — the strongest "real" evidence. R_spatial passes; a colour-in-void ranking fails (it is invariant under scrambling, so it is a histogram fact, not a composition fact).
Single null draws are seed-chaotic — one draw flips on luck. The real signal is that the measured value is cut-stable while the nulls are seed-chaotic, not a single point-drop. This retired an earlier single-draw claim and replaced it with a 100-seed sweep reporting an empirical percentile. Null tightness is also image-class-dependent: generated images scramble roughly 3× wider than paintings, so percentile thresholds do not transfer across domains without a recheck.
Every measurement is of (object × reproduction chain × scale window). A photograph of a painting is not the painting (finding I-10). Dense-area measures survive resolution sweeps; thin-structure measures do not. Every number is reported with the resolution and physical scale it came from, and a claim that has not survived a resolution sweep is labelled a reading aid, not a result. Crops are different measurements.
A measure invented in a study does not become part of the instrument by being interesting. It enters a candidate register and must clear a promotion bar across the whole corpus — including a null-control check — before it is ever allowed into the spine. Most candidates do not clear it, and the register records why they were declined as carefully as it would record a promotion.
| Candidate | Decision | Reason on the record |
|---|---|---|
| C-3 R_spatial | Not promoted, kept as a pattern | A torque-emergence axis. The signal is the range-width, not the point value, and no operational rule yet discriminates cut-robust from cut-sensitive works. Promotion waits on that rule. |
| C-4 address | Not promoted, kept as a pattern | Null-validated and it discriminates, but the raw scalar misleads alone — it needs an n≥100 null sweep attached — and its fine sub-signals hit the pre-semantic ceiling. |
| C-1 anomaly | Resolved externally | The Degas structural anomaly resolved into a documented figure only through outside scholarship. The instrument pointed; it could not name. |
An empty register would mean the promotion bar is not doing its job. The declines are the evidence that it is.
An early read of the Degas returned zero structural accents. It looked like a clean null — a painting with no punctuation. It was not: the accent lens was operating at the wrong scale for that brushwork. Corrected, the same painting returns a full accent field. The episode became findings I-5 and I-7, and the standing rule: when a result surprises — especially a null — look under it before believing it. Apparent failures are not discarded. Several became the sharpest results in the corpus. The most recent study, Jeong Seon’s Inwangjesaekdo, is a live case: three overclaims were caught before they shipped — a false extremum, a paper-ageing confound mistaken for brushwork, and a phrase wrongly credited to a translated source — each corrected to what the measurement, or the checked source, actually bears.
This is the honest account of what stands behind the numbers, and what does not. An instrument this young has no external gold standard to be scored against, so the question is not "has it been validated" in the certified sense, but "what kind of confidence is available, and where does it come from."
The published corpus is sixteen deep studies. The instrument's calibration was shaped against a far larger informal corpus: well over ten thousand images, across domains, developed and documented over years (see the wider Parallax Metrology site and artistinfluencer.com). That breadth is why the thresholds and defaults sit where they do. It is design provenance, and it is real. It is not external validation, and this page does not treat it as such: the same eye shaped and judged those runs, which is a bubble, and a number produced under those conditions is a considered stipulation, not an independently confirmed fact.
A single scalar is mechanism-blind. Many different arrangements reach the same centred value: a generated image centres because its gradient offsets, a Cézanne because objects are placed on balanced sides, another work because its voids balance. The number is identical; the paths are not. Average that scalar across a hundred works and the mechanisms wash out, because the mean keeps the value and discards the path.
So the coordinate is not the finding. The finding is the per-work signature: which channel, which island, which void produced the value, and how the metrics co-move. This is why the instrument is, by construction and not by immaturity, a single-work or small controlled-corpus instrument. Large-corpus use is legitimate for distributional questions (where do generated images cluster) and a category error for compositional ones (what is this painting doing). Reading a corpus average as if it described a work is the mistake the instrument is built to avoid.
Because averaging destroys the signal it would need to validate, confidence here is built the way observational sciences build it before a gold standard exists: by convergence across independent witnesses on the same case. A reading earns weight when several things that could disagree instead agree:
Three or four independent witnesses converging on one work is a stronger claim than a hundred works averaging, and it is the validation model the instrument actually runs on. This is the established logic of convergent validity and the nomological network (Campbell and Fiske; Cronbach and Meehl), and of consilience: a construct is trusted when independent methods, each fallible alone, point the same way. The discipline that keeps it from becoming "seeing what you want" is that the co-movement is predicted from the mechanism and written down first. A pattern you predicted and then found is confirmation. A pattern you found and then explained is a story.
A metric that fades in a larger corpus, or moves under perturbation, is not thereby useless. Subtle signal is low-amplitude by nature and will always move; the question is whether its structure holds while its value drifts, whether it keeps its rank against comparators and its place in the correlated pattern. A reading that survives in structure but not in amplitude is a subtle real signal, read at n=1 and in correlation, never asserted as a lone number. Amplitude-fragility alone is never taken as refutation.
The instrument is deterministic and open, which means no one has to trust the author: the same image at the same resolution yields the same coordinates, and anyone can re-run a study and challenge a specific number against specific pixels. That is the credential offered in place of external validation, and it is offered deliberately. Question any finding. What would change the reading is stated the same way every time, in each critique's What This Does Not Prove: name the number, show the pixels, and either the geometry supports the reading or it does not.