Orientation, for a cold reader
The scene: Christ, at the far right with Peter in front of him, extends a hand toward a table of five tax collectors at the left: an old man in spectacles, a bearded man in a black hat, a boy in a feathered hat, a youth slumped over the coins, and a youth seated on a box in profile. A shuttered window hangs over the scene; a raking light crosses the upper wall from outside the frame at the right. Which of the five is Matthew is a live scholarly dispute, addressed below with measurements rather than a verdict. One rule governs every sentence that follows: the instrument reads geometry, tone, and colour. It does not recognise people. Every figure name in this essay is a human label attached to a pixel-verified location, and the confidence of each claim is stated where it is made. Claims come in three strengths here: pipeline-stable (survives denoising and threshold sweeps), narrowed (survives with a smaller magnitude than first measured), and reading aids (informative but sensitive to choices we made; never load-bearing).
Field summary
A luminance-dominant, quiet-dissolved field at extreme tonal concentration: 90.4% of tonal mass in shadow, 0.44% in highlight, 81% of all pixels in the two darkest of ten luminance zones (τ 0.71). Only 6.5% of the frame carries structural mask at all; 72% reads as void, though "void" here mostly means form absorbed below the measurement floor, not empty canvas. These numbers are insensitive to the reproduction's noise (they shift in the fourth decimal under denoising) and they are the frame for everything else: this is a painting in which light is scarce, and the instrument's job is to say where the scarce light was spent.
The obvious reading of that first number is a trap. "Ninety percent shadow" invites the reply that Caravaggio painted in the dark, which every viewer already knows; the measurement's contribution is not the observation but the unit of analysis. The instrument carries no figure/ground prior. It reads the whole canvas as one field, and on that reading the dark is not the background behind the figures: it is the dominant field, and the lit passages are localized activations within it. Two readings point that way and one carries it. The single largest coherent mass in the painting is a passage a viewer would call background (the lit wall, next section), though that reading's direction is partly a property of the segmentation method, so it corroborates rather than proves. The regime classifier labels the dark "absorbed edges, not empty space," form modelled below the contrast floor rather than blank canvas. The load-bearing reading depends on no segmentation choice at all: the dark dominates the tonal field (90.4%) while carrying almost none of the structural mass (6.5%), which is the precise sense in which it is a container rather than a weight: it holds the picture without pulling on it. Two limits kept in view: calling the dark a container is interpretation (the tonal/structural split under it is not), and none of this corrects a viewer's perception, which is a different quantity from the instrument's fields.
Structure
Twenty mass islands with a real hierarchy, and the hierarchy's top is not a body. The largest, most cohesive structural mass in the painting is the lit wall beside the window, verified against the pixels directly (the crop below shows what sits at the anchor's centroid: plaster and window frame, nothing else). The magnitude of its dominance needed correcting and the correction is part of the record: the first-pass read gave it 16.1% of the frame at 2.8× the next-largest region, but the reproduction's speck noise was inflating dark-zone mass, and under denoising the anchor holds at 9.7% and a 1.3× margin: still the largest island, still the wall. The identity of the anchor is pipeline-stable; the drama of its margin was partly noise.
One honesty this deserves that the number alone hides: the direction of this result is partly a property of the method. Relation-first segmentation splits a tonal-cohesion field, so chiaroscuro fragments a figure into small islands while a smooth lit plane stays one, by construction; "the largest connected mass is a wall" is close to what this method returns for any strongly-lit-plane-plus-figures picture, not a discovery peculiar to Caravaggio. And it does not correct anyone's figure/ground perception, which is a different quantity from connected-component area. What is left, stated plainly: the instrument carries no figure/ground prior, and on its own terms the heaviest single mass in the frame is a passage a viewer would call background. Whether Caravaggio leans on that device more than his contemporaries is the question worth asking, and one painting cannot answer it.
The composition is measurably off-center and load-bearing rather than symmetric: mass skews right (34.4% vs 28.0% by column, denoised pipeline), the mass centroid sits right of frame center, and the classifier reads the whole as counterweighted. The balance is carried by two load-bearing regions, not one: the seated youth's mass at bottom-center and the money-table group at the left, which swap rank across noise conditions (an earlier single-region attribution to the money table alone is retracted in the lab book). The upper third of the canvas, so often described as empty, is the opposite: it holds 2.5× the bright-pixel share of the rest of the painting, plus the anchor. The action happens under a loaded band, not under a void.
Value and light
The light behaves like a single constructed system, and it can be audited feature by feature. The wall records a falloff: along the clean wall segment right of the window, luminance rises monotonically toward the light source: a 2.32× right-over-left ratio across ≈68 cm of painted wall, identical in both pipelines; across the full width, the wall left of the window is 4.3× darker than the wall right of it. The window is not the source: its brightest pane (0.278) reads dimmer than the wall it is set in (0.329) and far dimmer than the lit wall above the shadow line (0.439); a luminous opening would have to out-read its own wall. The standard oilskin-window reading survives quantification.
Two shadow edges, two different kinds of edge. The crisp diagonal boundary separating the lit upper wall from the darkness the figures stand in runs at 15.1–17.5° above horizontal (the range spans the two pipelines; the systematic difference exceeds the statistical error, so the range is the honest number), with a penumbra of 0.9–1.8 cm that broadens away from the frame edge, which is the direction an off-right occluder predicts. Above it, a much softer triangular darkening (≈6 cm edge) runs roughly parallel. If the source were the sun, that pairing implies a near occluder (window-frame range) for the crisp edge and a far one (roofline range) for the soft one. This is an arithmetic reading aid, stated with its assumption: it takes the painted penumbrae as faithful. What the pairing establishes without any assumption is that the painting renders two distinct edge-hardnesses on one wall, consistently.
Is the light "physically impossible," as the commentary tradition has it? Not internally. Four witnesses are usable: the sill-shadow edge (13.9°), the beam boundary (15–17.5°), and the shading vectors of the two near-frontal faces (both ≈29–31°, agreeing with each other within about a degree in both pipelines). Together they place one coherent source up-right, outside the frame, above the sill line. The two crisp shadow edges agree at roughly the level a published Vermeer study treated as evidence of a consistent source. What is true, and measured, is only the narrower claim: the source is not the visible window. Coherent light from a deliberately off-stage origin is what the numbers support; "impossible" is not. (A fifth witness failed and the failure is kept: shading direction on the posed heads measures hats and pose, not light; it is usable only for near-frontal faces.)
The faces meter the light's economy. The three frontally-lit table faces sit within 1.31× of each other (0.456–0.597) with no correlation between brightness and position (r = −0.08). Where a nearby lamp would dim the far faces, the profile is flat. That flatness is the measured version of a fork posed in the one prior computational study of this painting (Stork & Nagy 2010): either the light was effectively distant, or Caravaggio corrected each figure's brightness by eye. The fork stays open; the premise it needed is now a curve rather than an impression. And an economy the commentary rarely notes: Christ (0.284) and Peter (0.271) carry about half the luminance of the tax collectors' lit faces. The caller stands in his own shadow zone; the light he brings lands on the called.
Color
Colour here does not act as an independent register. The edge–chroma coupling index reads +0.363 on the original and strengthens to +0.481 denoised (wherever colour appears, drawing is already there), while colour-in-the-void (β) is correspondingly low (0.050 → 0.033). One comparison earns its place: across the four paintings this library has measured, those two numbers put Caravaggio at the bound end of a clean spread whose other pole is full decoupling: the opposite construction from a Matisse that weaponises the colour/structure split or a Degas that dissolves through it. The point is not the ranking; it is that in this painting colour is a property of lit form, never a force of its own.
An apparent counter-signal died under scrutiny, and the record keeps the autopsy: the chromatic-boundary-independence metric initially read high (0.236), suggesting colour edges the light doesn't draw, hypothesized as the costume trim. Mapped pixel by pixel, the signal turned out to live in the dark field (49% of qualifying pixels; the costume zone barely enriched at 1.1×), and the denoised value (0.190) falls below the metric's own autonomy threshold. In a 90%-shadow painting, "no luminance edge" is trivially true nearly everywhere dark, so faint warm/cool modulation masquerades as chromatic autonomy. The finding is a caveat about the metric on dark paintings, not a fact about this painting's colour.
Accents and the gap
The accent lens finds real punctuation, and finding it cleanly required two corrections that are part of the result. First, nearly half the raw detections sat on a reproduction-mount band at the bottom edge, pixel-verified as an artifact (uniform, full-width, and achromatic: the colour channel finds nothing there) and excluded. Second, denoising did not move a single surviving accent (13 of 13 re-detected within 0.03 frame units) but revealed six more that the grain had been masking. Among them is one at the painting's most famous location: a bridge-type accent at the fingertip zone of Christ's extended hand, hovering over the seated youth's head. The noisy read had the sleeve and the reacting faces; the fingertip itself appeared only on a clean signal.
Where the punctuation concentrates is itself the reading: the extended hand and the chain of faces turning toward it, plus one still-life exception: a stylus standing upright on the money table, which both the structural and the chromatic channels select independently. Unlike the broken-colour surfaces where this instrument has found two separate accent systems, here the channels agree, consistent with the Color section: where colour and drawing never decouple, there is no second, colour-only system of marks to find.
The gap the hand reaches across is a measured object. From Christ's fingertip to the nearest lit table-group figure: 35–38 cm (about 11% of the canvas's width), stable across lit-mask thresholds and both pipelines. Its anatomy: within each group, neighboring figures sit 19–27 cm apart; between the groups, bypassing the hand, the nearest faces are 109 cm apart; the hand hangs almost exactly midway (31 cm back to Peter, 33 cm forward to the seated youth), slightly nearer its own group: reaching, not arrived. The seated youth's nearest lit neighbor is not any member of his own group; it is the hand. And the bridge accent sits on the crossing: 21 cm from the hand, 16 cm from the group; the sum is the span. The interval is not empty; it carries the painting's punctuation.
For scale, one geometric comparison: measured by the same landmark protocol on the Wikimedia reproduction, the fingertip gap in Michelangelo's Creation of Adam, the image this hand is commonly read as quoting (a connection carried in the scholarship via Lavin, per the Wikipedia article's citations), is 0.44% of its composition's width. Caravaggio's is 11%. He quotes the hand and opens the gap by a factor of about 25: Michelangelo paints the instant before contact; Caravaggio moves the gesture across a room and makes the distance itself the subject. The comparison proves scale, not intent.
Attention
Structural weight and predicted gaze do not rest together. The saliency centroid sits near frame center while the mass centroid is pulled right toward the wall; the divergence (0.086–0.099 across pipelines) exceeds the instrument's calibration case for "the eye goes where the structure isn't." The heaviest thing in the painting, the wall, is not where a viewer is predicted to look; the small lit faces and hands out-pull it. A composition can hold its weight in one place and spend its attention in another; this one measurably does.
Cross-scale and stability
The gradient field holds one logic across the canvas (LG 0.02, globally consistent; the same modelling applied everywhere), while the tonal field sits at the measurement ceiling for plane independence (TLG 1.0): four quadrants, four value registers: near-black over the table, the lit wall, the mixed lower left, the highlight zone of the standing figures. One technique, four lighting worlds. The composition survives degradation well (persistence 0.81–0.84, the highest of this library's four studies): a structure built from a few large lit masses in darkness has little to lose to blur or grayscale. Stance reads "standing" (torque 0.13), meaning the center of gravity holds its heading across scales, with one honest asterisk: an earlier reading of this same painting from a different reproduction gave 0.24 ("mixed"), and until a second source is measured, that discrepancy is logged as a fact about reproductions, not resolved. Self-similarity is reported only as a range (structure-mask D ≈ 1.56–1.61 across a 2.8× resolution sweep); the dark-mask dimension is saturated at this painting's 86% dark fill and carries no information.
Relations: the relay, the debate, the center, and Peter
The Matthew question, measured three ways. The dispute has two poles. Traditional reading: the bearded man is Matthew, pointing at himself; revisionist (Varriano 2006; Magister 2018): he points at the slumped youth, the true Matthew. Under measurement it decomposes into one stable signal and two structured ambiguities. Stable: the light illuminates the bearded man: his face carries 1.9× the luminance of the slumped youth's bowed head, in both pipelines; the "follow the light" argument of the traditional camp has a number. Ambiguous, precisely: the shadow boundary extended toward the group lands between the two candidates and the nearest-candidate verdict flips with the pipeline. It is reported as non-discriminating, a reading aid that refused to become evidence. And the bearded man's own pointing axis passes within 3–7 cm of both his chest and the slumped youth's head. At a 340 cm canvas, both readings of his gesture lie on his finger's line (a 2D measurement of a foreshortened 3D gesture, and the hand's lit mass is only weakly elongated; the axis swung just 3° between pipelines, but the caveat stands). Meanwhile Christ's famous hand measurably targets nothing: its axis passes at least 37 cm above every candidate head; a hovering hand, not a pointing one, which is precisely the quality the Michelangelo-quotation tradition ascribes to it. On these measurements, "deliberately ambiguous" is not a critic's shrug; it is a described geometric property of the picture, coexisting with a photometric vote for the tradition. Which figure is Matthew is a question about gospel text and series context, and nothing here answers it.
The center is not a void; it looks. A commonplace of the didactic literature puts a pocket of dark negative space "in the exact center of the canvas." Measured, the claim fails twice: the deepest pocket between the groups hangs 86 cm from center (pipeline-stable to the millimeter), and the exact center itself sits 1.2 cm from lit structure: on the feathered boy, the brightest face in the painting (and, to a human eye, the one face turned toward Christ; that is observation, not measurement). The void between the groups is real; its address in the commentary is wrong, and what actually occupies the center is more interesting than emptiness.
Peter, tested. The survey literature reports (from X-ray examination; the primary radiograph was not consulted for this study) that Peter was absent from the first state and added in a second campaign. Two standard explanations circulate: a theological insertion, or a formal one ("Christ was too isolated at the right edge and needed anchoring"). Only one narrow slice of the formal reading is testable here, and it is worth being clear that it is narrow: the instrument can weigh left-right balance, and in balance terms Peter fails the anchor role. Masking him to the dark and re-reading improves every balance measure the instrument has (imbalance 0.1015 → 0.0876; horizontal asymmetry halves; the mass centroid moves toward center), and the native per-region estimate on the unmodified painting agrees (his region's balance-dependency is negative). In centroid terms he loads the right, he does not steady it. But centroid balance is almost certainly not what "anchor" means to an art historian: repoussoir, depth staging, the closing of the open right edge, the diagonal of the two entering figures are all spatial or perceptual claims this instrument does not measure, and none is touched here. So the honest statement is scoped: if the argument were that Peter evens out the picture's weight, the numbers contradict it; the readings actually in circulation are untested. Two further limits: the simulation is the extreme case (the true first state had Christ's body where Peter stands, so the real difference is smaller, though the sign carries the point), and ruling out one motive does not establish another.
Optional interpretation
Marked as interpretation, and resting on the measurements above rather than adding to them: this painting's construction is an economy. The regime classifier's pointer (profile_hint: old_masters; a pointer, never an attribution) attaches to the lowest composite perceptual load of the four works this library has measured; almost nothing here is spent on distributed devices. One off-stage source pays for everything: the wall it strikes becomes the largest structure; the faces it selects become the accents; the boundary it casts becomes the upper canvas's single strong diagonal; the gap it leaves dark becomes the subject. Read this way, the measured facts stop being separate findings: a caller dimmer than the called, a center occupied by the brightest face, a hand that hovers rather than points, and a 37 cm interval carrying a single bridge mark are one device seen from four sides. The picture organizes itself around what the light has not yet finished doing.
Public frame
Where this sits against existing readings
This critique keeps several public frames in the background. The didactic tradition around this painting emphasizes the divine light vector, the empty upper half, a central void between Christ and the table, and the paradox of a window that does not light the room. The measurements here confirm the core of that tradition (the source is real, coherent, and deliberately off-stage; the window is measurably not the source) and correct two of its specifics: the upper band is the loaded band, not the empty one, and the void between the groups hangs well right of center, with the brightest face in the painting occupying the center instead.
On the identity of Matthew, this critique joins neither camp. The traditional reading receives a number (the light selects the bearded man photometrically, 1.9×, in both pipelines); the revisionist reading receives the measured fact that the pointing gesture and the light's own geometry are poised between the candidates. The claim this critique adds to the debate is narrow: the ambiguity is a described property of the picture's geometry, not a gap in the scholarship.
Against the one computational precedent (Stork & Nagy 2010), the relationship is complementary: their renderer matched the wall's falloff and the figures' even brightness by eye; this study supplies the measured curves those matches presupposed, adds penumbra widths in physical units, and resolves their "unnoted" upper shadow into two features with two occluder distances. Their local-versus-solar fork remains open; the measurements here constrain it without closing it. And against the X-ray literature, the engagement is one deliberately narrow inference: given the reported later addition of Peter, the instrument can test only whether he steadies the picture's left-right balance, and finds he does not. The senses of "anchor" that art historians actually use (repoussoir, depth staging, edge-closure, the entering diagonal) are perceptual and spatial claims it does not measure and does not address.
Appendix, sources consulted: D. G. Stork and G. Nagy, “Inferring Caravaggio’s studio lighting and praxis in The calling of St. Matthew by computer graphics modeling,” SPIE vol. 7531 (2010), read in full; J. Varriano, Caravaggio: The Art of Realism (2006), and I. Lavin, Past-Present: Essays on Historicism in Art (1993), both via the citations of the Wikipedia article; the Magister–Sullivan exchange on the identity of Matthew (S. Magister, Caravaggio: il vero Matteo, 2018; C. Sullivan’s reply; E. Lev, 2012) via Italian Insider, read at second hand; the X-ray report of Peter’s later addition via Exploring Art and the published radiograph record (after W. Friedlaender, Caravaggio Studies, 1955), primary radiograph not consulted; standard commentary on the oilskin window via caravaggio.org; the Vermeer source-consistency benchmark (M. K. Johnson et al., SPIE 6810, 2008) as cited within Stork & Nagy; Michelangelo comparison reproduction via Wikimedia Commons. Where a source was read at second hand, this essay says so at the point of use. One asymmetry is worth naming plainly: the measurements can defend themselves, but the art-historical scaffolding around them is thinner, several load-bearing anchors (Varriano and Lavin via Wikipedia, the Magister–Sullivan exchange via one Italian Insider piece, the X-ray finding via secondary sources with the radiograph never consulted) are trusted more than verified. A reader inclined to push back should push there first; the primary literature has not been worked as hard as the pixels.
What this does not prove
n = 1 painting, one reproduction, one primary working resolution. The reproduction is the museum's own native capture (8792×8304, ~48 MB), the most authoritative available and not a casual web JPEG, which is the right source to have used. But two different error sources sit on top of it and only one is controlled. Speck noise in the darks is controlled, by the double-pipeline discipline (every headline ran on the original and a denoised copy); one first-pass number, the anchor's 2.8× margin, was already a partial noise artifact and was corrected, which is reason to hold the rest with the stated tolerance rather than tighter. Reproduction choice is not controlled at all, and the stance reading is the warning: the same painting read "standing" here (0.13) and "mixed" (0.24) off a different photograph, and stance is the only quantity this study gave a second source. Every other crisp decimal here (2.32×, 15–17.5°, 37 cm) therefore carries an unmeasured reproduction-variance error bar that its precision does not show: read those as this photograph's values, firm to the tolerance stated but not yet cross-checked against a second capture. Quantifying that variance across two or three reproductions is the obvious next study, and until it is done the confidence on the single-source numbers is lower than the decimals suggest. The instrument has no semantics: it cannot say which figure is Matthew, and does not; every figure name here is a human gloss on a verified location. The gesture measurements are 2D projections of foreshortened 3D acts. The occluder-distance arithmetic assumes the painted penumbrae are faithful; that assumption is the premise of the studio-reconstruction literature, not a finding of this study. The X-ray report behind the Peter counterfactual was corroborated across secondary sources but the radiograph itself was not examined; if the report is wrong, the counterfactual remains a true statement about what Peter's mass does and loses its historical hook. The Michelangelo comparison is geometry between two reproductions, nothing more. The beam-landing extrapolation failed its stability check and is reported as failed. And nothing in this essay is a judgment of quality: the method locates what the image is doing and stops there.