Spatial Kernel · Embodied-Verification Instrument
A learned success judge answers "does this look correct?" The Spatial Kernel answers "is it correct?" deterministically, from pixels, with no model in the loop. It measures the actual centroid of visual mass against the target an instruction named, then adds an independent two-view depth cross-check.
On real photographs it caught ten flat printed decoys a frontier embodied VLM certified as real 3D objects present. Perception said PASS. Geometry said FLAT.
Robots can see and robots can plan. What no standard pipeline instruments is whether the spatial intent of a natural-language instruction became a geometric fact. The Spatial Kernel measures the actual centroid of visual mass and compares it to the target the instruction named, producing a displacement and a compliance verdict without a learned judge. It then adds a second, independent read: a two-view parallax cross-check that separates a real 3D object from a flat printed decoy. Success is treated as the agreement of those independent reads. Across 19 object types photographed on an ordinary phone, the geometric check caught ten flat decoys a frontier embodied VLM certified as real objects present, within a stated envelope reported honestly alongside the cases where it declines.
The companion diagnostic paper, Confidence Without Correctness, shows why the independent read is necessary. A frontier success judge asked whether a manipulation succeeded returned false passes at a flat 0.95 confidence, right or wrong, because its learned representation smooths over the sub-threshold physical change that decides the outcome: on a public corpus an object's learned box moved 2.4% while its true visible extent moved 67.7%. The forward study, an authority-separated auditor specification, states which parts of the response already stand on evidence and pre-registers the single experiment that would settle the rest. Together they frame the Spatial Kernel as one instrument with a diagnosis behind it and a decisive study ahead of it.
Four Modes Of Spatial Non-Compliance
Real objects photographed from two ~10 cm-offset positions with no depth sensor, the second view is the geometry. Across 19 object types the deterministic two-view parallax check separated real 3D objects from flat printed decoys within a stated envelope. A frontier embodied VLM, run on the same scenes, certified ten flat decoys as real objects present. The kernel read every one of them FLAT.
The package shipped as a clean CLI that claimed to answer "did the robot put the object where the task required." Reading the source showed it did not: it scored the whole frame, not the target object. Isolating the object before the kernel runs, without adding a learned judge, turned a silently-confounded metric into a conformant instrument, fixed honestly and adversarially against a real frontier VLM.
Asked to verify whether a manipulation succeeded, a frontier success judge returned false passes at a flat 0.95 confidence, right or wrong. The cause is representational: on Language-Table an object's true visible extent moved 67.7% while its learned box moved only 2.4%. The judge reasons over the box, so the deciding change is invisible to it. An independent pixel measurement recovers exactly the evidence the learned judge discards, and abstains when blind.
Cluster bias, peripheral dispersion, symmetry without occupancy, and partial placement override: four distinct failure modes where a scene reads compliant to perception while the measured centroid of mass says it is not. Each is a case where "looks right" and "is right" come apart, and each is caught by measuring the actual mass against the named target rather than trusting the perceptual impression.
The Spatial Kernel is a deterministic compliance instrument, not a general 3D reconstruction system or a replacement for perception. The two-view parallax check separates real 3D objects from flat print within a stated envelope: it requires two adequately offset views and adequate texture, and it declines rather than guesses where those conditions fail. The ten-decoy result comes from documented real captures, not a statistical sample from a known population; the four non-compliance modes are characterized categories, not a proof of completeness. The favorable partial-visibility rule that would have raised one boundary case from 1/12 to 10/12 is rejected, not adopted, and any such change becomes a new prospective question rather than a post-hoc repair. The central claim, robustness scales with the number of genuinely independent observables, is part true by construction and part empirical, and the forward study is designed to pin down which is which.
A learned judge produces a verdict and a confidence from a representation that has already smoothed away sub-threshold physical change. It is fast and general, and it is the right tool for semantic questions. The problem is that it will return a confident PASS on a scene where the deciding evidence never entered its representation, and its confidence does not move when it is wrong.
The kernel does not judge intent. It measures the centroid of visual mass against the named target and cross-checks depth across two views, deterministically, from pixels. Low disagreement with the judge means the verdict can be trusted. High disagreement is the warning: "looks right" and "is right" have separated, and this episode deserves a human or a second surface before it enters training.
The intended relationship is diagnostic and complementary. Run the Spatial Kernel beside the learned judge that labels your rollouts. Where they agree, proceed. Where they disagree, you have found a candidate poisoned label, the exact episode a data engine should not train on. The natural place a partner comes in is the third surface: telemetry or a second viewpoint for the strata where a single camera is structurally out of envelope, which is the question the pre-registered forward study is built to answer on real data rather than assert.