Parallax Metrology
Lab Book · Relational Robotics Verification

Relational Auditor

Finding the narrow place where spatial-field work might add something real—without mistaking familiar geometry for novelty.
opened 2026-08-14  ·  repo: Documents/Robotics/  ·  model: gemini-robotics-er-1.6-preview
separate from Robotics-ER 1.6 and ISR library  ·  local only, no commit  ·  147 tests + 261 subtests green
scope: task-outcome verification · system-authored frame condition · measurement-validity abstention · learned grounding + independent measurement · ISR gated behind ordinary baselines · Papers: Confidence without Correctness · The Authority-Separated Auditor

This is the backward trail: what prompted the experiment, what was changed, what the controls said, what genuinely survived contact with real data, and where the ISR hypothesis is still only a hypothesis. The standing block is authoritative; dated entries preserve the path.

Standing Where this actually is — 2026-08-17

affirmative core The demonstrated architecture is a visual frame-condition firewall: the task defines a minimal permitted change, learned grounding maps only the commanded symbols into the image, and a separately calibrated monitor audits the unpermitted complement—first catching, four times over, a natural collateral-motion failure a frontier VLM certified as clean at 0.95 confidence, then eliminating nine direct false passes in the prospective paired-contract test. The learned model does not get final authority merely because it can describe the scene; invalid registration, visibility, scale, or depth evidence yields ABSTAIN.

v3.3 correction A hash-group audit finds calibration-to-validation leakage in the Language-Table no-motion packet: NM009 is byte-identical to calibration NM007, and NM016 duplicates calibration NM004. The evidence is corrected from 16/7 rows to 15 unique calibration / 6 unique validation pairs. Thresholds and five task results do not change. The former “coincidental” equal maximum explanation is withdrawn.

Language-Table authority boundary V3.1/v3.2 PASS counts certify only the measured preservation complement, not target-goal success. Their calibration grounding also used a joint two-endpoint call while task grounding used separate window calls. Future evidence must group scenes before splitting, use identical independent grounding transport, and report TARGET_GOAL, COMPLEMENT_INTEGRITY, and MEASUREMENT_VALIDITY separately.

v2.2 prospective safety result On two new layouts with textured and smooth 1 cm collateral moves, direct Gemini again made nine false passes in 20 task-scored judgments, always at confidence 0.95. The frozen gold cascade caught every admitted physical violation, made zero wrong decisions, kept CAL quiet, produced both permission flips, and abstained on every camera case. The natural failure now replicates at smaller motion and across surface types.

v2.2 is partial, not a clean win Gold v2.2 decided 16/20 correctly and abstained on all four LIGHT judgments. Lab alone false-failed those four; flow corroboration correctly refused to call highlight/shadow residuals physical movement but could not verify them stationary. Thus the repair converts false failure to safe deferral—it has not earned illumination-invariant PASS. Extracted execution decided 12/20 correctly with no unsafe error; target-role aliasing and malformed contracts still cost coverage.

post-result audit survives A sealed replay reproduced every source, artifact, registration, denominator, and direct false pass, and localized every admitted collateral detection to the moved object. Two post-hoc LIGHT shortcuts score 20/20 only by weakening the missing-evidence rule; a smooth true movement shares their proposed low-chroma signature. Unchanged ISR creates two false passes. The v2.2 result stands, but no v2.3 LIGHT resolver is admitted.

v2.3 branch closed without a manufactured win The illumination-validity thread repaired its capture apparatus but did not expose a natural sparse-tracking gap. Three increasingly plain household surfaces remained trackable when actually moved; the final blue candidate never supplied a valid moved pair. No LIGHT-to-PASS resolver is admitted, missing physical evidence still means ABSTAIN, and further object shopping is stopped rather than engineering a favorable counterexample.

phase-two position The architecture is the thesis; offline VLA label auditing is the nearest application; online runtime assurance is the later deployment claim. Near-term work should harden contract interoperability, then test a fixed frame-condition auditor on externally adjudicated VLA outcomes. “Rare but poisonous labels” is an economic hypothesis to measure, not a benefit already demonstrated. Sensor-validity and relational-ISR studies are required only if the claim expands into those operating regimes.

v2.4 contract interoperability A system-authorized semantic IR raises sealed-corpus execution readiness from 22/28 to 24/28, admitting only the two contracts with redundant target-only permissions. Exact target/destination binding and an exact non-target write-set remain system-owned; role swaps, mixed permissions, missing authority, and four malformed contracts still fail closed. No v2.2 outcome is rescored. External VLA outcome data is now the next gate.

v2.5 external-data boundary DROID-100 is suitable for a real-trajectory feasibility pilot, but not yet for the advertised poisoned-label test. Its current metadata contains 46 instructed positive-reward episodes and no instructed zero-reward episodes. More importantly, DROID's task-success label does not state that every non-target object must remain fixed. A complement alert is therefore a downstream-policy incompatibility unless it also violates the stated task; the project must not turn an unstated preservation preference into somebody else's label error.

v2.5 visibility result Fixed-role Gemini grounding is locally valid for 11/12 post-screen episodes, but 0/12 has complete TARGET and DESTINATION coverage in both BEFORE and AFTER under the frozen clear-visibility rule. This is not 0/12 correctness: no outcome was admitted. The firewall correctly refused to turn a blind endpoint into success, while the endpoint implementation demonstrated zero useful coverage on this stratum.

v2.5 fail-closed integrity Twelve simple placement episodes produced 48/48 unique sealed frames and 12 unique cached API responses. A second freeze binds the packet, task selection, prompt, response schema, protocols, and executable code. Confidence never gates admission; unavailable roles were nevertheless often reported at 0.9–1.0 confidence. Admitting partial boxes after the result would recover at most two valid responses and still leave 9/11 incomplete, so temporal evidence—not threshold relaxation—is the next gate.

v2.6 three-way real-trajectory result Temporal vision, gripper/Cartesian telemetry, and a post-primary wrist addendum produce 1/12 vision-temporal observable-chain candidate, 2/12 multi-surface candidates, and 9/12 richer-state-required. No task outcome is scored and containment is never called directly verified. The result demonstrates one genuine cross-surface closure without pretending telemetry is an object sensor.

v2.6 strictness and next hinge Exterior grounding has 22/24 valid views and leaves carried-target visibility as the dominant missing link. Wrist grounding validates 12/12 calls and acquisition 12/12, but only 1/12 carriage links pass the frozen clear-target rule. A post-hoc partial-target sensitivity reaches 10/12 and is not admitted; identity-continuous partial tracking on untouched episodes is the next falsifier.

DROID interpretation DROID is a measurement-validity stress test, not the sole efficacy corpus. The controlled photographs show that the monitor works on controlled real captures. DROID shows that wide environment-facing terminal views often withhold the clean scene evidence that implementation assumes. Arm occlusion may be recoverable from a different moment or view; opaque containment is a modality limit. Both require explicit evidence accounting, never a relaxed PASS.

v2.7 external scene pivot The next primary corpus is the real Language-Table release: planar line-of-sight block rearrangement with human-selected behavior endpoints, 2-D end-effector state, and naturally retained collection mistakes. It is materially closer to the present frame-condition instrument. BridgeData V2 is reserved for the later RGB-D, camera, clutter, and obstruction stress. Neither source silently inherits a target-only preservation policy.

v2.7 external surface result One original Language-Table shard supplies 877/877 CRC-valid episodes and a 48-episode instruction-only frozen sample. Registration succeeds 48/48 with median translation 0.11 px and maximum 0.48 px. Localization responses validate 48/48; 41/48 have TARGET clear in both endpoint windows and 35/48 have all eight objects clear in both. This qualifies a useful external measurement surface without pretending the remaining 13 complement-incomplete episodes are unchanged.

v2.3 qualification attempt 1 The corrected 18-image set is useful but does not qualify a resolver. Q03/Q04 lighting changes destabilize the frozen whole-frame camera gate; reversing the same pairs changes admission. The hinged “smooth” case remains easy for flow (193–245 px supported motion) and appears to have moved at inch scale. Q08 nevertheless confirms that a supported physical veto survives concurrent lighting residuals. Attempt 2 needs fixed fiducials, locked exposure, a truly feature-poor object, and measured 10/30 mm spacers.

v2.3 qualification attempt 2 The second 18-image set repairs the main apparatus failure: all eight ordinary pairs remain locked in both directions, both light-only controls remain stationary, and only the declared camera pair crosses the frozen gate. But the green lid is visually plain, not tracking-poor—its rim, tab, lettering, and shading yield 176–215 corners and strong verified motion (226–285 px). The apparatus qualifies; the pivotal missing-evidence challenge does not yet exist.

prospective v2.1 result On two new layouts, direct Gemini made nine false passes in 16 task-scored judgments, zero false fails, and zero abstentions, always at reported confidence 0.95. The same ordinary measurements under the system-authored gold contract made zero false passes, abstained on all four camera-stress judgments, and made three false fails. The permission-complement firewall therefore removed the observed unsafe error, but its measurement surface still pays a specificity cost.

semantic success, interface failure BEFORE-only extraction recovered the target relation, protected scene remainder, and correct selective pink-bowl permission in 28/28 contracts. Yet the frozen compiler accepted only 11/28: all 14 preserve-all contracts used two semantically adequate POSITION and ORIENTATION clauses instead of one canonical ANY_VISUAL_CHANGE clause, and three selective contracts violated the literal remainder schema. This is not model omission; it is brittle normalization at the language-to-compiler boundary.

v2.2 development antecedent Exact split-clause normalization raised replay compiler readiness from 11/28 to 25/28 without repairing the three malformed selective contracts. A trusted permission profile rejected extracted exceptions that were missing or unauthorized, so the learned layer could not widen the write-set. Lab proposals corroborated by pre-existing forward/backward flow removed all three v2.1 false fails while retaining every task failure: gold replay was 16/16, and the earlier v1.6 replay was 9/9. Both datasets informed that repair; the new v2.2 result above, not this replay, is the confirmation evidence.

supported, narrow A learned model is useful for semantic grounding, but its end-to-end judgment can return confident incorrect verdicts even when an exact preservation constraint is stated. On the historical synthetic v0.2 pilot, direct Gemini produced two stable false passes; learned boxes plus deterministic predicates produced no wrong decisions. That pilot contained 28 labeled rows but 27 unique inputs, used only the clean nuisance slice, and supplied oracle target coordinates. Independent depth converted the constructed height cases from guesswork into evidence.

real mechanism The semantic-physicality problem is not confined to synthetic scenes. In the legacy real-frame subset, stored appearance/objectness judgments were 17/27, with 10 flat-print false passes; this endpoint is related to, but not identical with, height-above-plane. Two-view parallax was useful inside its operating envelope.

not the contribution Pairwise relation math, connected components, and object tracking are standard. After correction, a simple renderer-specific component baseline solved all 112 unique synthetic scenes when given depth. There is no case for presenting those predicates as new robotics mathematics.

replicated real mechanism Across four related five-object photo sets, all 32 source JPEG hashes are unique. Direct Gemini was 24/28 and four times false-passed the same explicitly named target-success-plus-collateral-bowl violation at 0.95 confidence, each time reporting the moved bowl unchanged. Replayed under the current measured-camera policy, grounded relations cover 18/28 and decide 18/18 correctly; this is development evidence, not population accuracy, and the trace does not distinguish perceptual miss from later reasoning override.

open-world complement The unlisted-v1.6 target-only path demonstrates the other half of the architecture: registered change exposes all seven camera-admitted collateral manipulations without collateral identities or boxes and produces no alarms on four held-out CAL pairs. Direct Gemini was not run as an outcome judge on those unlisted pairs, so the replicated direct-judge failure and the open-world result remain two related findings—not yet one sealed head-to-head result.

precise claim Runtime monitoring, geometric constraint checking, and abstention are established ideas. The potentially distinctive discipline is narrower: the learned model may ground task symbols but may not author the system invariant, choose the protected set, set an unseen-task tolerance, or convert invalid measurement into PASS. The target envelope is still partly learned in the current implementation and therefore needs deterministic size/path sanity checks before this firewall is fully independent.

threshold audit Two same-state repeat pairs produced 12 registered object residuals: median 1.34 px, p95 26.90 px, maximum 32.19 px. The frozen 70 px tolerance exceeds twice that observed maximum (64.38 px) and remains below the smallest true collateral displacement (156.25 px). This is retrospective support, not independent calibration.

independent calibration grounded Two new four-layout sets provide 32 empty-target photographs. Clean-prompt grounding produced all 192 requested boxes. Locked no-move residuals are tight (median 2.57 px; maximum 9.84 px), but camera-perturbed residuals reach 82.56 px despite correct identities. A universal 70 px preservation rule is therefore rejected.

prospective related-set confirmation Roles and the calibrated camera rule were frozen before photo-v0.4/v0.5 inspection. Direct Gemini was 12/14, false-passing both collateral-bowl cases. Grounded relations covered 9/14, decided 9/9 correctly, and abstained on all five camera-stress frames. Whole-frame delta again rejected every outcome.

adversarial correction A width-versus-longest-dimension scaling bug overstated Claude's registered residuals and is now fixed. Corrected locked residuals have median 2.35 px and maximum 28.02 px; camera perturbation reaches 214.85 px. Twenty pixels remains too tight for three independent-model residuals, while 60 px is only a twice-maximum heuristic—not a discovered constant.

ordinary tracker baseline SIFT features inside semantic reference boxes solve the available locked-camera preservation mechanism: 18/18 decided outcomes correct, zero false passes, and ten explicit abstentions. The 60 px rule lies inside a broad 34.16–149.63 px all-correct interval. This is not an ISR win; it raises the gate. Semantic boxes and placement still come from Gemini.

ordinary controls replicated Sparse optical flow independently reproduces 18/18 correct decisions, with an all-correct 35.83–148.26 px interval. Registered local change reaches 19 decisions but false-fails two nominal successes because a twice-noise changed-fraction cutoff treats small incidental bowl shifts as task failure. It is a useful attention cue, not metric compliance.

first recognition-light residual With every protected-object box withheld, full-frame SIFT motion clustering detects 3/4 placement-correct collateral cases and false-passes photo-v0.5 IMG_3. No-move calibration contains false repeated-texture clusters of six matches; the missed bowl supplies only five, so threshold lowering is not a valid repair. This earns an unlisted-object experiment, not an ISR claim.

annotation limit Claude's interrupted ruler addendum saved two inch-tick points for all 32 earlier calibration images, giving an approximate locked median of 86.47 px/cm. It did not produce the final validation deliverables, and a ruler scalar at the table edge cannot globally calibrate raised objects under perspective. Claude remains a model, not human ground truth.

ISR gate closed The frozen dense methods ran on 32 new, uniquely hashed images. Both detect all eight declared unlisted manipulations and make no alarms on four CAL pairs; ISR alone avoids one ordinary-Lab photometric false failure. The omitted preregistered full-frame SIFT arm was later restored unchanged and detects 7/8, missing A_PLAIN_3CM. Blind masks show Lab covers 70.2% of the moved-object sweep versus ISR's 9.0%; ISR is 4.44× slower than Lab and loses three detections at a +10% score threshold. The current mass-residual implementation does not enter the robotics path. The surviving result is the target-envelope architecture, not this ISR kernel.

Handoff Primary path—freeze, hypothesis, and resume point

direction The primary path is an authority-separated visual auditor for declared VLA data-use contracts. Learned grounding translates task symbols into identities and approximate regions; system policy owns the permission envelope; independent evidence executes it; invalid evidence abstains. The nearest application is offline data and evaluation auditing—not yet runtime robot control.

primary hypothesis On untouched, externally adjudicated robot-scene outcomes, a fixed permission contract executed by independently calibrated pixel evidence will reduce unsafe false passes relative to direct VLM judgment while retaining useful decided coverage. The accompanying validity hypothesis is that measured occlusion, registration loss, or unstable support converts unsafe decisions into ABSTAIN rather than merely moving errors between PASS and FAIL.

candidate frozen V3.2 dense role-local support is the final development candidate. Its masks, 12% expansion, 45-pixel minimum, and decision ordering must not be tuned on P01/P10/P18/P19/P20. Those five are regression cases. V3.3 leaves six unique quiet controls and a 2 complement-PASS / 0 FAIL / 3 ABSTAIN development result, with the joint-versus-independent grounding mismatch still unresolved.

next study elementfrozen requirementreason
selectionuntouched episodes; exact and near-duplicate scenes grouped before splittingprevents the v3.3 leakage from recurring
groundingidentical independent per-frame calls in calibration, validation, and applicationremoves temporal smoothing and transport mismatch
truthindependent human or instrument adjudicationseparates physical outcome from model agreement
armsdirect VLM; global components; frozen v3.2; combined contract executortests architecture and ordinary alternatives
outputsTARGET_GOAL; COMPLEMENT_INTEGRITY; MEASUREMENT_VALIDITY; OVERALL_VERDICTprevents complement-only PASS from becoming task success
reportingrandom prevalence and enriched challenge strata remain separatepreserves both ecological rate and failure-mechanism power

go gate Continue toward an offline VLA audit utility only if untouched evidence shows fewer unsafe passes than direct judgment, non-trivial decided coverage, appropriate nuisance abstention, no authority widening, and consequential alerts at a review cost that matters. “Rare but poisonous” remains an economic hypothesis until this is measured.

stop conditions Narrow or close the path if performance depends on the fixed colored-block vocabulary, most valid outcomes abstain, ordinary components equal the result at lower cost, or alerts are merely unstated downstream preferences rather than task/contract violations. Do not add camera, depth, runtime, or ISR scope to rescue a failed narrow claim.

resume checklist Read v3.3; verify the v3.1/v3.2 freezes; seal a grouped untouched packet and factorized schema; acquire once; run identical grounding once; seal independent adjudication; only then open arm comparisons. Full handoff: PRIMARY_PATH_HANDOFF_V3.4.md.

2026-08-17 V2.7 scene-facing dataset selection—POV becomes part of the instrument

v3.3 post-result correction Exact image-pair grouping reveals two duplicated no-motion pairs, including one crossing the calibration/validation boundary. Corrected unique denominators are 15/6. V2.9 remains 6/6 quiet, v3.0 becomes 5 quiet / 1 false alarm, v3.1 becomes 4 quiet / 1 false alarm / 1 abstain, and v3.2 remains 6/6 quiet. Frozen source artifacts remain intact.

selection The real Language-Table corpus is selected for the next external pilot. Its xArm6 moves in a two-dimensional plane over a smooth board containing eight colored blocks; collection uses a third-person line-of-sight view, and annotators selected the start and end of coherent behaviors. The official release reports 442,226 real episodes. Direct inspection of the public TFDS metadata confirms 500 TFRecord shards, 428.7 GB total, with 360×640 JPEG observations, 2-D end-effector state/action, reward, and instruction bytes.

corpusrolewhyboundary
Language-Tableprimary scene-facing pilotplanar visible blocks; event-selected endpoints; small relational moves; actor statehindsight descriptions are not preservation contracts; camera lock still must be measured
BridgeData V2secondary stressfixed RGB-D view, randomized alternate cameras, wrist view, clutter, 6-DoF tasksactor occlusion and containment recreate DROID's evidence loss
DROID-100measurement-validity recordreal trajectories exposed endpoint, occlusion, and modality limitsnot a useful clean scene-facing efficacy corpus for the current implementation

camera shift A robust affine/homographic registration may admit a near-planar camera shift only when background inliers span the board and a held-out residual gate passes. Thresholds move to board/object-relative coordinates. A transform that aligns the table but leaves parallax, exposes hidden surfaces, or lacks spatial support cannot repair the pair into PASS; it requires calibrated depth/multiview evidence or ABSTAIN.

obstruction The pusher footprint is system-owned and derived conservatively from 2-D end-effector state. Actor-covered pixels are excluded from measurement but never counted unchanged. A visibility ledger may seek the nearest clear frame only under registered identity continuity. A moving target zone is bounded by maximum growth, displacement, and interruption duration; it cannot expand until the residual disappears.

next gate Acquire one original shard (the first is approximately 912 MB), select a 48-episode instruction-only sample before viewing outcomes, retain endpoint windows, and measure registration, all-block visibility, and actor overlap before scoring anything. Only then freeze the external frame-condition execution.

Frozen Language-Table before and after endpoint examples
Frozen external packet, page 1. The scene-facing board remains stable and visible while the compact pusher often overlaps the commanded block. Selection was sealed from instructions and metadata before these endpoint images were extracted.

acquisition result The shard contains 877 CRC-valid episodes. Exact target-only grammar yields 377 target-to-object, 135 target-to-region, and 49 small-move candidates; 316 are rejected. Two pre-image freezes were rejected when adversarial review found multi-target, tap, hand/actor, or unsupported-role leakage. The final deterministic hash selection contains 16 episodes per stratum.

surface result All 48 endpoint pairs register, with 0.11 px median and 0.48 px maximum estimated translation; the maximum scale delta is 0.27%. All 48 learned localization records validate. TARGET is clear in both endpoint windows for 41/48; all eight objects are clear in both for 35/48. Six target-observable episodes still lack complete protected-scene evidence—the exact distinction between checking the commanded relation and auditing its complement.

still a feasibility result These are model-reported visibility states, not independent truth, and no movement or outcome has been judged. The fixed colored-block vocabulary may favor an ordinary classical detector; it must be run as a comparator. Language-Table descriptions are hindsight labels, not preservation contracts.

independent packet sealed A neutral 21-episode packet contains all 13 learned complement-incomplete cases plus eight seeded complete controls, without disclosing which is which. It holds 126 byte-identical images and requires 1,008 per-object visibility annotations. Instructions, source IDs, target bindings, prior outputs, and the external neutral key remain outside the packet; its validator checks completeness, geometry, image dimensions, and hashes.

independent result narrows admission The blind return is structurally valid and complete. Applying the same “at least one clear frame per three-frame window” rule, the independent annotator confirms 12/13 learned-incomplete episodes but only 5/8 learned-complete controls. The two instruments agree on 293/336 object-window decisions (87.2%); 42 disagreements are Gemini-clear/independent-not-clear and only one is reversed. This asymmetry argues for conservative intersection, not promotion of the original 35/48 model estimate to ground truth. Five episodes are jointly complement-complete; disagreement cases remain measurement-loss stress tests.

interpretation discipline Claude used a newly written color-segmentation instrument, then visually assigned same-color shape identities and corrected every overlay. That is useful independent second-model evidence and partially demonstrates the strength of ordinary vision, but it is neither human ground truth nor a clean fully automated baseline. Because the blind packet intentionally oversampled learned failures, its 21-episode rates do not estimate prevalence across the 48-episode freeze.

conservative dry-run freeze Five jointly complement-complete episodes are sealed for registered-change development; the other sixteen adjudicated episodes remain visibility-loss/disagreement stress cases. Freeze SHA-256: d06ef5a378aad48f969d3824249d4d52b6f6d1b0622d0e70dff324fcee35b27e. This is a mechanism-development set, not an outcome sample.

v2.8 registered-change dry run On the five jointly visible episodes, the commanded object has the largest registered earliest-BEFORE to latest-AFTER displacement in 5/5. Nominal direction/relation agrees in 4/5. LT0103 is a natural label/contract mismatch: the yellow hexagon moves 66.06 px rightward but still ends left of the green circle, so “place right of” is not achieved. This separates original hindsight-description agreement from downstream success.

threshold rejected The three-frame windows are trajectories, not certified repeat captures. Within-window registered displacement has median 0.032 px and p95 3.56 px but a 60.77 px maximum; using that maximum as “noise” suppresses every real target move. No percentile is substituted after inspection. Until a matched natural no-motion calibration is frozen, strict target-only preservation is UNRESOLVED, not PASS.

ordinary comparator A fully automated, no-override fixed-HSV component arm obtains exact two-components-per-color coverage in 3/5 episodes and abstains on two. After the sealed learned target box binds the named component, the target is the largest component displacement in 3/3 measured cases. Ordinary geometry is sufficient where it has coverage; same-color shape identity and a valid motion threshold remain outside its authority.

Language-Table P18 six-frame registered-change example
Natural description/contract mismatch, P18. The yellow hexagon moves rightward across the six-frame sequence but remains left of the green circle at the terminal frame. Movement toward a relation is not completion of that relation.

v2.9 prospective validity threshold—corrected A fresh official shard supplies 23 residual-blind no-motion rows but only 21 unique image pairs: 15 calibration and six validation. The calibration maximum remains 1.264 px; unique held-out validation produces 0/6 false alarms with a 0.911 px maximum. The former equal maximum came from the leaked NM007/NM009 duplicate.

validity before displacement Applied to the five known development episodes, the frozen component surface returns 1 measured-motion/complement PASS, 0 FAIL, 4 ABSTAIN; this is not full target-relation certification. P18 would otherwise false-fail on a 7.11 px non-target shift, but that object's visible support grows from 161 to 499 pixels as the robot uncovers it. The prospectively frozen 10% support-stability gate correctly turns the apparent motion into ABSTAIN. Categorical “clear” from two models was not enough to guarantee metric validity.

coverage cost P10 also abstains because a non-target component changes area by 13.7% even though its displacement is below threshold; P19 and P20 lack exact component coverage. The rule is not loosened post-result. The safety ordering now works on natural external pixels, but the fixed-HSV implementation is usable on only one of five known episodes. The next problem is coverage recovery without sacrificing the held-out false-alarm result.

v3.0 learned-box recovery fails metrically Gemini validates all frozen no-motion rows and calls all eight objects clear in every pair. But repeated, quantized boxes freeze a 0.102 px center threshold and a 0% area-change threshold; one of six unique held-out controls false-alarms at 0.108 px. Applied unchanged, the surface returns 0 complement-PASS, 0 FAIL, 5 ABSTAIN. Categorical coverage rises while usable metric coverage falls to zero.

representation mismatch is the result On P18, deterministic visible support grows 67.7% as the robot uncovers a non-target block, while Gemini's inferred box changes only 2.4% and its center remains below the component motion threshold. Learned boxes recover apparent coverage by smoothing over the nuisance evidence. They remain useful for semantic role binding, but cannot certify their own metric support. Independent pixel validity—not learned confidence or box persistence—is what forces ABSTAIN.

v3.1 bidirectional flow preserves validity, not coverage A frozen registered Lucas–Kanade surface freezes a 0.302 px threshold. Corrected unique held-out validation is 4 quiet, 1 false alarm, 1 abstain. Applied unchanged to the five known tasks it returns 1 complement-PASS, 0 FAIL, 4 ABSTAIN. The result does not justify sparse flow as the recovery surface.

P18 remains safely unresolved Requiring coherent BEFORE→AFTER and AFTER→BEFORE evidence prevents newly exposed pixels from masquerading as a stable tracked object. P18 again abstains, independently corroborating the earlier support-area gate. But smooth blocks often lack six stable corners, and NM006 produces a visually unsupported 0.437 px held-out alarm. Bidirectionality protects the evidence chain; it does not manufacture texture or specificity.

three surfaces, three failure signatures Fixed color components preserve visible support but lose coverage under merging; learned boxes recover categorical identity while smoothing nuisance and quantizing geometry; sparse flow remains nuisance-sensitive but loses support on low-texture objects. More instruments are not automatically better. The next gated candidate is dense local color/support conservation inside bounded learned roles, and it must improve coverage without converting P18 into PASS.

v3.2 dense local support clears its development gate—conditionally The protocol and implementation were sealed before measurement. Thresholds remain 1.225 px centroid and 7.06% visible support after deduplication. All six unique held-out stable-support controls are quiet. Applied unchanged, the surface returns 2 complement-PASS, 0 FAIL, 3 ABSTAIN—the first measured coverage improvement that retains P18's abstention.

Language-Table P10 before and after target movement
Recovered complement coverage, P10. The commanded yellow hexagon moves under the pusher while the protected yellow heart remains at lower left. Role-local dense support isolates the stationary partner; the earlier global-component surface abstained on scene-level support change.

validity signal survives In P18, non-target yellow support grows from 161 to 673 pixels (76.08%) as the robot uncovers it, and the same-color role regions do not overlap. Dense support abstains before interpreting its 8.39 px centroid shift. Learned semantics bind the role; independent pixels retain final authority over whether that role is measurable.

conditional win P19 and P20 remain abstentions, but their same-color learned regions overlap substantially, so physical support change and partition-boundary change cannot be separated. The 6/6 unique controls were also preselected into a stable-support regime and are not a corpus-wide false-alarm estimate. Freeze v3.2; do not tune on these known cases. The next evidence must be prospective and stratified by role separation, collision/overlap, occlusion, collateral motion, and illumination.

Decision record: EXTERNAL_SCENE_DATA_SELECTION_V2.7.md · RESULTS_LANGUAGE_TABLE_SURFACE_V2.7.md · LANGUAGE_TABLE_GROUNDING_PROTOCOL_V2.7.md · sources: Language-Table repository, Language-Table paper, BridgeData V2 project, BridgeData V2 paper.

2026-08-17 V2.6 temporal multi-surface pilot—three branches survive

headline The real-trajectory feasibility gate produces all three preregistered branches: 1/12 VISION_TEMPORAL_CLOSES, 2/12 MULTISURFACE_CLOSES, and 9/12 RICHER_STATE_REQUIRED. These are observable-chain candidates, not task verdicts. D005 closes with temporal exterior vision; D010 closes with exterior vision plus telemetry; the disclosed wrist addendum closes D009's exact missing carriage link.

containment boundary Nine tasks are containment-like. Vision cannot see through opaque containers, and gripper position plus Cartesian motion does not sense object contact. Even D010's multi-surface chain establishes acquisition/transport/release-compatible observations—not direct physical containment. Disappearance is never promoted to “inside.”

DROID temporal exterior-camera samples aligned to gripper telemetry
Telemetry-aligned exterior packet. Each episode has two camera rows and five fixed phases: PRE_ACQUIRE, CARRIED, PRE_RELEASE, POST_RELEASE, and TERMINAL. The carried object commonly disappears behind the robot even when the destination remains observable.
surface/gateresultmeaning
gripper cycles11 single-cycle; 1 six-cyclethe plural tissues task cannot be represented by one final transition
exterior response views22/24 validtwo broad masks rejected; valid alternate views retained
primary exterior branches1 vision / 1 multi / 10 richercarried-target visibility is the dominant missing link
wrist acquisition link12/12target and gripper are spatially localized at acquisition
wrist carriage link1/12strict clear-target rule rejects partial/occluded carried objects
integrated branches1 / 2 / 9one wrist result closes exactly one primary missing link
outcome scoring0coverage is not mislabeled task accuracy
DROID wrist-camera samples aligned to gripper telemetry
Post-primary wrist addendum. The camera makes the carried objects visually apparent, but formal coverage remains narrow: Gemini labels ten carried targets partial and two occluded. Clear TARGET + visible GRIPPER + ≤3% box gap passes only D009.

adversarial sensitivity Admitting `partial` TARGET after seeing the result would raise wrist carriage from 1/12 to 10/12. That favorable rule is not adopted. It becomes the next prospective question: can deterministic continuity or independent annotation show that partial wrist observations preserve target identity?

failed-closed apparatus Two sampler defects were caught before learned grounding: insufficient post-release margin and compressed-video seeking across episode boundaries. A ten-image API request and a temporary-client call also failed before inference. After splitting to five-image view calls, a cached replay corrected over-strict whole-episode rejection to per-view rejection. Every incident and superseded artifact is retained.

Result: RESULTS_DROID100_MULTISURFACE_V2.6.md · protocol: DROID100_MULTISURFACE_PROTOCOL_V2.6.md · sources: DROID schema, DROID paper.

2026-08-16 V2.5 external-data gate—real trajectories admitted, label-cleaning claim withheld

selection The LeRobot DROID-100 sample is the first external feasibility corpus. It provides 100 real robot trajectories, three views, and 46 episodes with both a task instruction and positive terminal reward. The metadata payload was inspected directly and hashed before this decision.

DROID-100 intersectionepisodesuse
all episodes100metadata universe
positive terminal reward81claimed successes
zero terminal reward19failures without task text
task text + positive reward46eligible feasibility pool
task text + zero reward0no instructed failure comparison

claim firewall The source label records task success, not a universal no-collateral-motion contract. The fixed complement monitor may audit a declared downstream use policy, but a complement-only alert cannot be called a poisoned source label. Original-label errors, downstream-policy incompatibilities, measurement failures, and abstentions must be reported separately.

DROID-100 tasked episode start and end frames from exterior camera one
Exterior camera 1 screen. Start/end samples for all 46 instructed positive-reward episodes. Backgrounds are often stable, while actuator pose and task visibility vary.
DROID-100 tasked episode start and end frames from exterior camera two
Exterior camera 2 screen. The alternate view recovers some task evidence but introduces different occlusions; view availability must be measured rather than assumed.

adapter gate The existing unmasked complement field would mostly detect the permitted robot. Before scoring, the system needs an authority-fixed robot exclusion, per-view registration and visibility tests, and deterministic multi-view selection or fusion. No label-quality or accuracy number is extracted from this qualitative screen.

one-sided semantics Verified unexpected change outside the actor mask may raise an audit alert. The converse is harder: masking the robot does not prove that its occluded sweep stayed unchanged. No alert becomes PASS only with measured temporal or multiview coverage of the declared protected complement; otherwise it is ABSTAIN. This keeps an offline QA screen from quietly claiming runtime assurance.

Frozen DROID-100 12-episode packet with two views and before-after frames
Frozen development packet. Each row is one opaque episode; columns are V1 BEFORE, V1 AFTER, V2 BEFORE, and V2 AFTER. Robot-pose changes and view-specific occlusions are visible before any learned grounding is run.
packet gateresultmeaning
development episodes12simple single-object placement instructions
extracted frames48/48 uniquetwo views × two phases
source/image hash verificationvalidsource drift and frame tampering fail closed
fixed-role grounding11/12 locally validD010 rejected for a broad robot mask
complete TARGET + DESTINATION coverage0/12no episode reaches outcome-measurement admission
outcome calls0missing visibility remains ABSTAIN, not success

what bound Among the 11 valid responses, BEFORE coverage is TARGET 10/11 and DESTINATION 11/11; AFTER coverage falls to TARGET 3/11 and DESTINATION 9/11. The 25% robot-overlap threshold is not the dominant blocker. The actor is still holding or occluding the task object, or the role leaves the frame. Reported confidence does not cure absent measurement.

next experiment Keep this material development-only and add a small temporal-coverage pilot before any outcome scoring: search for a post-release frame only under independently verified release, role visibility, and robot separation; otherwise model interruption and occlusion explicitly through a temporal window. Multi-view fusion requires calibrated time/geometry validity. Verified unexpected change may ALERT, but incomplete protected-scene coverage must remain ABSTAIN.

alternatives held for their proper jobs ViFailback is stronger for curated temporal failure diagnosis and measurement-loss cases; SAFE's real rollouts are a direct runtime-monitoring comparator; AHA remains a procedural development comparator. None supplies the missing external preservation authority for free.

Decision record: RESULTS_DROID100_GROUNDING_V2.5.md · DROID100_GROUNDING_PROTOCOL_V2.5.md · EXTERNAL_VLA_DATA_SELECTION_V2.5.md · metadata audit: droid-100-metadata-audit.json · sources: DROID paper, DROID-100 metadata, SAFE, ViFailback.

2026-08-16 V2.4 contract interoperability—coverage recovered without authority drift

result The semantic IR compiles 24/28 sealed extracted contracts versus 22/28 under v2.2. It recovers exactly K003 and K021, whose only problem was redundant POSITION/ORIENTATION permission for the already system-authorized target. All 24 IR records validate against the frozen schema.

gateresultinterpretation
execution-ready22/28 → 24/28two target-only interface losses recovered
non-target authority widening0/28learned contracts cannot enlarge the write-set
newly admittedK003, K021trusted exact target binding only
still rejectedK004, K014, K020, K027schema error, missing remainder, or illegal enumeration

not language understanding The compiler does not infer that “white pedestal mug” means target. An external system profile binds the exact entity ID to the target role. Target-only learned MAY clauses are projected into pre-existing system footprints; mixed target/non-target permissions and mismatched roles are rejected.

architectural consequence Interoperability improves through a monotone projection into a system-owned write-set, not through a more forgiving parser. The learned layer can vary its surface form but cannot author away the protected complement.

Scope: retrospective engineering replay only; prior prospective scores remain frozen. Protocol: CONTRACT_INTEROP_PROTOCOL_V2.4.md · result: RESULTS_CONTRACT_INTEROP_V2.4.md.

2026-08-16 V2.3 branch closure—ordinary tracking survives the proposed challenge

headline The proposed realistic smooth-object weakness did not materialize. That is a useful narrowing, not an invitation to search for a stranger object. The project keeps the conservative firewall and ordinary baselines; it does not claim that ISR or a special structural field is needed to handle these movements.

screensurfaceordinary-flow evidencemeaning
attempt 1white hinged case193–245 px supported movementvisually plain but structurally easy
attempt 2green lid176–215 corners · 226–285 px movementrim, tab, embossing, and shading remain sufficient
exploratory v2.3.3pale plate300 corners · 98/116 px movementstrongly trackable under both lighting manipulations
exploratory v2.3.3blue cup289 corners; valid pair measured 0.98 pxobject is trackable, but no valid moved challenge was captured

what is not earned The branch does not establish illumination-invariant PASS. One v2.3.3 blue intensity pair also contains camera stress, and its other blue pair is effectively stationary; neither can be recast as the missing physical-movement counterexample. The exploratory screen is retained only as the stopping audit, not promoted to a sealed result.

what survives A task-authored permission envelope plus an independent deterministic complement monitor remains the durable architecture. Cheap Lab differencing and sparse flow are valid implementations. Invalid camera or missing measurement evidence still yields ABSTAIN. ISR remains outside the admitted path unless a natural deployment condition demonstrates an advantage over those ordinary controls.

decision Stop reducing object texture. Preserve v2.3 attempts 1 and 2 as development evidence, leave the LIGHT cases unresolved, and redirect future work toward naturally occurring deployment failures—occlusion, entry/exit, deformation, registration loss, or temporal interruption—where the current firewall can be challenged without constructing a pathological prop.

Interpretation: the negative result strengthens the surviving claim by removing an unsupported mechanism story. The work is now narrower, cheaper, and harder to dismiss as an engineered ISR demonstration.

2026-08-16 V2.3 qualification attempt 2—apparatus repaired, challenge still absent

capture validity repaired All 18 photographs are unique. Q01/Q02 are quiet; Q03/Q04 light-only changes stay stationary; Q05–Q08 physical moves are detected; and only Q09 crosses the camera gate. Unlike attempt 1, every admission result is the same in both temporal directions.

V2.3.2 lighting intensity before frame
Q04 BEFORE. Normal-light reference. Frozen registration remains inside the locked-camera envelope.
V2.3.2 lighting intensity after frame
Q04 AFTER. Deliberate intensity change and commanded target placement. The complement remains stationary rather than becoming a photometric false failure.
pairforward registrationreverse registrationfrozen disposition
Q03 light direction25.17 px · 0.9914122.00 px · 1.00853locked · stationary
Q04 light intensity8.95 px · 1.004541.36 px · 0.99918locked · stationary
Q08 move + light0.57 px · 1.000170.06 px · 1.00007locked · moved
Q09 camera389.82 px · 0.91633431.53 px · 1.03122camera stress both ways
V2.3.2 smooth motion plus light before frame
Q08 BEFORE. The green lid looks plain at scene scale but retains a rim, tab, embossing, and stable shading.
V2.3.2 smooth motion plus light after frame
Q08 AFTER. Flow verifies the lid movement under changed lighting with 58 selected-cluster points and 285.07 px full-resolution displacement.

pivotal challenge absent Q06 and Q08 produce 176–215 reference corners and 74–85 supported clusters. This is not missing evidence, so it cannot distinguish a safe LIGHT resolver from one that would hide a real but untrackable movement. “Smooth” must be defined operationally by tracking evidence, not appearance.

Verdict: apparatus/registration qualification passed; no resolver admitted. Full record: RESULTS_V2.3_QUALIFICATION_ATTEMPT_2.md.

2026-08-16 V2.3 qualification attempt 1—productive failure

packet and controls All 18 corrected photographs are unique. Both CAL pairs are quiet, the declared CAMERA pair crosses the gate in both temporal directions, and textured/smooth movements are detected. Q08's moved cluster vetoes the frame despite simultaneous unverified lighting evidence.

V2.3 lighting intensity before frame
Q04 BEFORE. Normal illumination, target mug off the blue target.
V2.3 lighting intensity after frame
Q04 AFTER. Target succeeds under reduced illumination. The frozen whole-frame registration gate incorrectly treats the photometric change as camera stress.
pairforward registrationreverse registrationfrozen disposition
Q03 light direction62.03 px · 0.9759029.87 px · 0.99079direction-dependent gate
Q04 light intensity60.23 px · 0.9717439.67 px · 1.01820direction-dependent gate
Q08 move + light15.21 px · 0.9946146.96 px · 0.98194direction-dependent gate
Q09 camera338.32 px · 0.97181302.59 px · 1.04029camera stress both ways

pivotal challenge absent Q06 and Q08 smooth movements are supported at 192.65 and 244.61 full-resolution pixels. The hinged case has usable latches, seams, and corners, and the nearby inch ruler indicates the nominal 1 cm manipulation is closer to inch scale. Thus the capture does not test whether a LIGHT resolver would conceal an unverified physical movement.

Verdict: no candidate admitted. Full record: RESULTS_V2.3_QUALIFICATION_ATTEMPT_1.md.

2026-08-16 v2.2 prospective result—unsafe errors removed, LIGHT remains unresolved

replicated at one centimetre The direct judge false-passed all eight textured/smooth 1 cm judgments and one forbidden pink-bowl judgment, with no abstentions and 0.95 confidence throughout. Its explanations explicitly asserted that the visibly moved objects remained in place. This extends the natural failure beyond the earlier 3 cm selective case.

prospective v2.2 armdecidedcorrect among decidedfalse passfalse failabstain
direct Gemini20 / 2011 / 20900
extracted contract + v2.212 / 2012 / 12008
gold contract + v2.216 / 2016 / 16004
target-only control20 / 2010 / 201000
unconditional complement16 / 2014 / 16024
frozen Lab-only ablation20 / 2016 / 20040
V2.2 textured one-centimetre violation before frame
V2.2 N09 BEFORE. Bare-wood confirmation layout; the target mug begins off the blue target and collateral identities remain outside the task contract.
V2.2 target success with one-centimetre textured collateral movement
One-centimetre open-world failure. The target succeeds and the textured brown object moves. Direct Gemini passes P and S at 0.95; the gold cascade fails both.
V2.2 lighting nuisance before frame
V2.2 N10 BEFORE. White-field layout before target placement and illumination change.
V2.2 target success under shifted side lighting
LIGHT exposes the remaining edge. Lab proposes multiple shadow/highlight regions; flow cannot verify several. Lab-only fails, while v2.2 abstains. The pixels do not support a PASS claim.

authority boundary learned something too BEFORE-only extraction was 25/28 schema-valid and 22/28 execution-ready. Two P contracts added redundant MAY clauses for an entity named WHITE_PEDESTAL_MUG; strict authority rejected them because that alias was not deterministically bound to the system target. This is safe but costly. Target-role binding is a v2.3 interface hypothesis, not a v2.2 rescore.

Verdict: literal safety gates pass, but the scientific specificity claim receives only partial confirmation. V2.2 demonstrates “false failure → abstention,” not illumination invariance. Full record: RESULTS_SPEC_EXECUTION_V2.2.md.

2026-08-16 Post-result adversarial checkpoint—result retained, easy fixes rejected

integrity and localization All sealed hashes, all 28 opaque-to-source image copies, and all 14 registrations reproduce. Each textured/smooth violation has one compact proposal over the moved object with supported motion of 200.40–384.89 full-resolution pixels; CAL and target-only SUCCESS produce none. The nine direct false passes and prospective denominators are unchanged.

post-hoc candidatedecidedcorrectfalse passabstaindisposition
frozen v2.216 / 2016 / 1604retain
stationary track overrides unverified20 / 2020 / 2000reject: unsafe aggregation
illustrative low-chroma fallback20 / 2020 / 2000reject: counterexample
unchanged ISR + frozen flow17 / 2015 / 1723reject

Goldilocks trap exposed Both favorable 20/20 replays merely fit the observed LIGHT pairs. The chroma fallback places the real N14 smooth-object movement (0.279) on the same side of its proposed 0.35 boundary as every lighting residual; flow happens to rescue N14 in this capture. A texture-poor unverified movement could therefore be cleared. The permissive aggregation is worse: any stationary background patch could outvote missing evidence elsewhere.

next gate, not next threshold V2.3 needs a system-owned target-role token and a separately defined illumination-validity measurement tested against a deliberately texture-poor physical move that defeats flow, plus held-out lighting types. Until that evidence exists, LIGHT remains ABSTAIN.

Full record: ADVERSARIAL_CHECKPOINT_V2.2_POSTRESULT.md. This is explicitly post-result analysis and does not rescore v2.2.

2026-08-16 v2.2 repairs both exposed boundaries—post hoc

compiler boundary The exact POSITION+ORIENTATION remainder pair is now normalized to the executor's conservative visual-change proxy with source-clause provenance. Canonical clauses pass through; partial, duplicate, mixed, enumerated, ambiguous, and invalid forms still reject. Replayed acceptance rises from 11/28 to 25/28: P becomes 14/14, S remains 11/14. After blind extraction, an independent authority profile requires the extracted permission set to equal the system-authorized set; extra or missing permissions reject.

measurement boundary Ordinary Lab remains the sensitive proposal field; forward/backward optical flow corroborates physical displacement inside each unexpanded proposal box. Any supported moved cluster fails—even when a larger stationary background cluster is present—any unverified proposal forces ABSTAIN, and only verified-stationary proposals are cleared. The recurring embossed-mug residual measures 1.92–2.20 full-resolution px; true v2.1 collateral proposals measure 143–336 px.

development replaydecidedcorrectfalse passfalse failcamera abstain
v2.1 gold contracts16 / 1616 / 16004 / 4
v2.1 extracted contracts15 / 1615 / 15004 / 4
v1.6 unlisted cross-replay9 / 99 / 9003 / 3

adversarial correction before capture The first freeze was not treated as sacred merely because replay was perfect. The audit found that a dominant stationary cluster could hide a smaller moved cluster, the model still had authority to invent a schema-valid permission, malformed nested values could escape the standalone normalizer, and the reported 30–120 px stability range described v2.1 alone. All four issues were corrected, tested, disclosed, and re-frozen before any v2.2 image existed.

sensitivity, not victory The cross-set threshold sweep is 25/25 correct through 90 px; v1.6 loses one true movement at 120 px, and v2.1 loses two at 150 px. The largest independent no-move flow is 6.65 px. A moved proposal's 89.20 px cluster is not itself an object-level boundary because another supported moved cluster can veto, so no universal safe interval is claimed; 60 px is retained as a development-supported operational policy. The corrected any-moved-cluster rule restores 16/16 at 0%, 15%, and 50% box expansion, but 100% expansion still misses one movement; zero remains frozen. All 80 independent same-state calibration ROIs stayed stationary. A future texture-poor motion may still be unverified; it must ABSTAIN rather than pass.

Prospective test completed: two new layouts added a locked-camera lighting/highlight nuisance, textured and smooth 1 cm moves, the paired selective permission, CAL, SUCCESS, and CAMERA. V2.1 remains unchanged. Repair: DEVELOPMENT_REPAIR_V2.2.md · red-team record: ADVERSARIAL_CHECKPOINT_V2.2.md · protocol: PROSPECTIVE_REPAIR_PROTOCOL_V2.2.md · result: RESULTS_SPEC_EXECUTION_V2.2.md.

2026-08-16 Repositioning against the failure-monitoring field

what this is The durable proposition is a visual frame-condition firewall. The commanded object and destination define a permitted write-set; registered measurement audits the complement for unauthorized scene change. Learned grounding may translate task symbols into coordinates, but it does not own the invariant or the final verdict. This is an implementation of established runtime-assurance logic, not a claim to have invented monitors or abstention.

occupied ground VLM success detection and free-form failure reasoning are already active lines in Vision-Language Models as Success Detectors and AHA. Sentinel combines a statistical action-consistency monitor with VLM task-progress reasoning. SAFE learns failure scores from VLA internal features and calibrates them with conformal prediction; uncertainty-aware policy steering calibrates a learned verifier to decide whether to act, clarify, or request intervention. “Models can be confidently wrong; monitor them and defer under uncertainty” is motivation, not this project's novelty.

closest distinction Code-as-Monitor is the closest comparator. GPT-4o generates its subgoals and constraint set, a VLM-derived segmenter selects constraint-related entities and parts, and GPT-4o generates executable monitor code. Common thresholds may come from an external knowledge base, but unseen-task thresholds may also come from the VLM. Branch-coverage testing can establish executability while leaving specification completeness unresolved: a valid program may monitor the wrong or incomplete contract. The stricter hypothesis here is that a system-authored permission complement prevents that omission channel.

failure-family correction Sentinel's canonical confident wrong-placement case violates the positive task goal. The photographed hard case satisfies the positive target goal while violating a remainder-of-scene invariant. It is therefore better described as a frame-condition failure, not simply another task-progression failure. Conversely, the four 0.95 Gemini misses used explicitly named protected objects; v1.6 proves recognition-light unlisted detection but did not run direct Gemini as an outcome judge. The strongest conjunction remains to be tested prospectively.

different abstention trigger Selective-prediction systems abstain because a learned answer or policy is uncertain; see also the VQA-focused dual-assessment reliability study. Here the monitor abstains when its measurement preconditions are invalid, even if the learned model is confident. This is established metrological/runtime-assurance discipline—see NASA's formal runtime-assurance framework—but its explicit application to a VLM-authored success claim is the useful design stance.

parallax claim ceiling The empirical result is specific: one homography-compensated image-plane displacement tolerance does not safely span locked-camera and viewpoint-changed scenes with depth variation. It does not establish that Code-as-Monitor's fixed multiview RGB-D pipeline fails. A direct comparison would require equivalent RGB-D trajectories, tracking, and camera conditions; an existing-image pilot must be labeled “CaM-style model-authored monitor,” not a reproduction of CaM.

trackquestionnext gate
A · contract authorityDoes a fixed permission complement catch collateral violations omitted by a model-authored monitor?Seal authored DSL specifications before AFTER images; compare omission and execution errors separately.
B · validityDoes a nuisance/observation gate turn unsafe geometric decisions into calibrated abstentions?Manipulate camera/extrinsic, visibility, depth, and lighting conditions without sacrificing locked-camera coverage.
C · ISR invarianceCan a sparse structural channel resolve dense photometric alarms cheaply when invoked locally?New candidate, calibration, and outcomes; no retuning of the rejected v1.7 field.
D · ISR relationsDo temporal islands or corridors expose consequential relation/clearance changes beyond ordinary scene graphs?Choose a genuinely relational task and beat strong tracking, segmentation, and collision-geometry controls.

pilot result Gemini's 16/16 BEFORE-only specifications selected the target goal, full permission complement, safe missing-evidence behavior, and calibration rather than invented thresholds. Yet its separately prompted direct judge false-passed all 8/8 collateral violations at 0.95, false-failed both camera-stress controls at 0.95, and never abstained. The simple contract-omission hypothesis is weakened; the sharper observed failure is specification–execution inconsistency: the learned workflow represented the invariant but its free-form judge did not enforce it.

16 / 16
authored full complement
8 / 8
direct collateral false passes
0
model-invented thresholds
0
direct abstentions

adversarial qualifier Primitive selection is not complete monitor quality. Six specifications unnecessarily permit a transit sweep, M12 binds destination occupancy anomalously, and M05 adds an unstated orientation constraint. Nine of sixteen retain the minimal source/destination permission pattern without the transit expansion or anomalous binding. These are disclosed post-hoc diagnostics, not changes to the preregistered primary classification.

compiler result A strict two-state compiler makes those semantic defects operational. Literal B1 is ready in 0/16 because every value correctly awaits calibration. With the frozen structure-only policies, only 9/16 contracts are valid and 8/16 execution-ready: six transit permissions are rejected as unobservable, one destination binding is rejected, and one otherwise-valid contract requests an unsupported orientation primitive. The compiler abstains or rejects; it never silently deletes a clause.

v2.1 sealed prospective result Twenty-eight uniquely hashed photographs produced 28 BEFORE-only extractions, 20 direct non-CAL judgments, and four measured arms. The gold firewall scored 13/16 task outcomes with zero false passes; the direct judge scored 7/16 with nine false passes. Target-only produced ten false passes, while an unconditional complement produced four false fails and no correct selective label flip. The gold selective monitor achieved the required P→FAIL/S→PASS flip in one of two layouts.

v2.1 falsifier retained The gold executor false-failed the Set A nominal success under both contracts and its selective pink case under S. Those three errors are attributable to the ordinary change surface, not extraction. They prevent the claim that the compiled arm already dominates without an offsetting cost. Seventeen extracted-arm rows also terminate at compilation, so its zero false passes come with unusably low coverage.

v2.1 armtask resultunsafe false passfalse failcamera abstention
direct learned judge7 / 16900 / 4
extracted + compiled6 / 7 decided014 / 4
gold contract + executor13 / 16034 / 4
target-only control6 / 161004 / 4
unconditional complement12 / 16044 / 4
V2.1 layout B selective-permission before frame
V2.1 N11 BEFORE. The pedestal mug begins off target. Contract P protects every non-target object; contract S alone permits the pink bowl to move.
V2.1 layout B target success and selective pink-bowl movement
Same pixels, opposite correct verdicts. The target succeeds and the pink bowl moves. The gold firewall returns P→FAIL and S→PASS; direct Gemini false-passes P at 0.95.
V2.1 layout B unauthorized-change before frame
V2.1 N12 BEFORE. Neither instruction names which protected object will be disturbed.
V2.1 layout B target success with unauthorized collateral movement
Unlisted collateral failure. Target placement succeeds while another object moves. Direct Gemini passes under P and S at 0.95; the permission complement fails both.

Sequencing: Track A has prospective support for the architecture and a newly exposed compiler-normalization problem. Track B now has a concrete residual: remove Set A incidental-change false fails without losing OTHER detection, CAL quietness, or camera abstention. ISR remains separately gated and may not enter as a post-hoc v2.1 repair. V2.1: protocol · result.

2026-08-15 v1.6 unlisted capture reverses the easy retrospective answer

narrow prospective advantage Thirty-two new photographs form 16 distinct pairs across plain-sheet and repeated-wood layouts. Hashes and the v1.7 thresholds were frozen before outcome scoring. Gemini saw only the target mug and target region; collateral identities and boxes were schema-forbidden.

0 / 4
CAL false alarms · both arms
8 / 8
raw collateral alarms · both arms
8 / 9
ordinary decided correct
9 / 9
ISR decided correct
V1.6 unlisted textured-object before frame
V1.6 B_TEXTURED_3CM BEFORE. The learned stage receives only the pedestal mug and blue target; the collateral objects are deliberately unnamed and unboxed.
V1.6 target success with unlisted textured-object displacement
Recognition-light AFTER. The mug reaches the target and the textured pink object moves. Both dense complement fields detect the unlisted change; this tests open-world coverage, not object recognition.
v1.6 methodraw collateral alarmsCAL false alarmsmedian moved-sweep coveragemedian runtime
ordinary Lab8 / 80 / 470.2%39.97 ms
full-frame SIFT7 / 80 / 4not scored165.59 ms
ISR mass residual8 / 80 / 49.0%177.60 ms

Registration found the two intended camera shifts and one accidental shift in A_PLAIN_1CM; all three abstain before verdict. Both methods decide the remaining seven collateral failures correctly, including the texture-poor 1 cm case in layout B. ISR passes both nominal successes. Ordinary Lab false-fails B_SUCCESS because a 6,176-pixel residual forms on a stationary embossed white mug.

The favorable error was challenged post hoc. The residual does not intersect the target envelope, and global registration is clean. Lucas–Kanade flow directly inside it finds 41 coherent points with only 1.78 full-resolution px displacement; expanded-crop SIFT and flow estimates remain below 1.87 px, versus the frozen 60 px object-motion tolerance. The evidence supports photometric/texture sensitivity rather than hidden mug movement.

protocol correction The sealed runner omitted the protocol's frozen v1.5 full-frame SIFT arm. Run unchanged after discovery, SIFT has 0/4 CAL alarms, detects 7/8 raw and 6/7 camera-admitted collateral moves, and false-passes A_PLAIN_3CM. The CAMERA captures also omit their prescribed target move; they remain camera controls but are excluded from task-outcome denominators.

Claude's blind return validated all 16 neutral pairs and supplied 32 inch-ruler intervals (median 96.35 px/cm at the ruler plane), but also claimed five low-confidence shifts of the textured mug from unregistered differences. Registered SIFT and flow disprove all five at 0.12–2.87 px versus the 60 px movement tolerance. This includes N03/B_SUCCESS, so the ordinary error remains a genuine photometric false alarm. Claude also missed the accidental A_PLAIN_1CM camera stress, reinforcing that a second model is annotation evidence, not physical ground truth.

On seven camera-admitted collateral pairs, both dense methods overlap the verified moved-object polygon 7/7. Lab has 70.2% median swept-object coverage and 0.385 IoU; ISR has 9.0% coverage and 0.086 IoU, though its sparse detections have higher precision. Median local runtime is 39.97 ms/pair for Lab, 165.59 ms for SIFT, and 177.60 ms for ISR.

Dialectical result: the ISR field contains a real photometric-invariance clue, but it wins none of the declared admission routes: no movement missed by all ordinary arms, worse localization coverage, equal abstention coverage, and 4.44× Lab cost. It also falls from 8/8 to 5/8 collateral detections at a +10% score threshold while Lab remains 8/8 through +20%. The current candidate is closed without tuning away either its sparse geometry or ordinary's one false alarm.

Surviving direction: target-conditioned recognition-light scene integrity is viable. For deployment, ordinary registered change is the stronger starting control; islands, corridors, or cheaper structural invariance would be separately preregistered hypotheses, not repairs to v1.7.

2026-08-15 v1.5 withholding boxes exposes the real residual

recognition-light miss The full-frame control receives no protected-object identities or boxes. It searches outside only the task-permitted source/destination envelope for SIFT matches that cluster jointly in position and residual-motion vector.

6
largest no-move false cluster
7
frozen required matches
3 / 4
collateral cases detected
1
false pass · v0.5 IMG_3

The missed bowl provides five coherent matches. Lowering the cutoff to five would also admit no-move repeated-texture clusters, so the miss cannot honestly be tuned away. Object-scoped SIFT and flow both detect it because semantic ROIs suppress the background ambiguity.

What this supports: recognition-light discovery is genuinely harder than verifying enumerated objects. What it does not support: that ISR islands are the solution. Dense motion proposals, better background suppression, segmentation, and standard multi-object tracking remain competitors.

Next gate: a newly captured, genuinely unlisted collateral object with frozen false-alarm calibration and identical ordinary/ISR abstention rules. Existing-data box withholding is only a simulation.

2026-08-15 v1.4 optical flow replicates; pixel change overreacts

second ordinary solution Forward/backward pyramidal Lucas–Kanade flow tracks protected-object corners after whole-frame homography compensation. All 80 locked calibration tracks verify; the worst stationary residual is 6.65 full-resolution px.

80 / 80
flow calibration tracks verified
18 / 18
flow outcome decisions correct
17 / 19
local-change decisions correct
36–148 px
flow all-correct interval

Flow reproduces SIFT with zero false passes or false fails and ten abstentions. Registered local change adds one wrong-object decision but false-fails photo-v0.2 IMG_2 and photo-v0.3 IMG_4: both contain small bowl shifts above the no-move noise floor but below the 60 px operational tolerance.

Dialectical result: sensitive change is not automatically useful compliance. A twice-noise threshold answers “did anything measurably change?” while the task requires a declared answer to “how much change counts as disturbance?” Changed fraction lacks physical units, so it cannot settle that policy alone.

ISR gate: SIFT and flow both solve enumerated-object preservation. The next experiment must withhold a semantic box from a collateral object and compare full-frame ordinary motion clustering with ISR islands. Repeating the bowl test would add volume, not information.

2026-08-15 v1.3 conventional tracking clears the known mechanism

ordinary method sufficient A SIFT ROI tracker uses cached semantic reference boxes, stored camera homographies, and coherent residual-vector clusters. Weak tracks are unverified rather than silently stationary. Cached Gemini boxes still supply target placement.

79 / 80
verified calibration tracks
18 / 18
correct outcome decisions
10
explicit abstentions
34–150 px
all-correct preservation interval

The no-move feature-tracking maximum is 6.68 full-resolution px, but sensor noise is not the task tolerance. A twice-noise cutoff falsely rejects two nominal successes containing small bowl residuals of about 34 and 19 px. The independently derived 60 px candidate separates these incidental shifts from deliberate collateral moves without approaching the observed decision boundary.

Implication: the repeated collateral-bowl false pass does not justify ISR. Once identity is initialized, standard local correspondence catches all four instances. ISR's remaining fair test is recognition-light unlisted change, fragmented or texture-poor regions, or efficient attention under parallax—against flow, segmentation, and modern tracking.

Scope: retrospective algorithm design, correlated tabletop images, semantic initialization and placement from Gemini, camera frames abstained, and no open-world object. This is a baseline result, not a robotics novelty claim.

2026-08-15 v1.2 adversarial checkpoint

material bug corrected The Claude audit applied a 900 px homography using longest-dimension scale even though registration resizes to fixed width. EXIF-rotated portrait images made the mismatch material. Corrected locked residuals are 2.35 px median, 9.17 px p95, and 28.02 px maximum; corrected camera residuals are 28.03 px median, 81.05 px p95, and 214.85 px maximum.

24 / 28
direct · four identical-mechanism false passes
18 / 28
harmonized grounded coverage
18 / 18
harmonized decided correctness
8.09–149.13 px
all-correct prospective tolerance interval

Not Goldilocked on preservation: every locked photo-v0.4/v0.5 verdict is unchanged for any tolerance from 8.09 px inclusive to 149.13 px exclusive. The corrected independent-model maximum is 28.02 px, leaving a broad observed interval. A 60 px candidate comes only from doubling that maximum and rounding; future tolerance must be tied to a declared minimum motion and loss function.

Oracle removed: camera abstention previously trusted the manifest's expected camera flag. It now uses measured registration, and a regression test proves that deliberately false manifest flags cannot suppress abstention. The current outputs do not change because measured and declared conditions agree.

Claims narrowed: placement is only target-center-inside-full-mug-box, not base-centering metrology. Every protected object is explicitly enumerated, so open-world collateral change is untested. Camera calibration strongly separates large perturbations but sparsely samples the mild 34–50 px boundary. V0.4 IMG_4's camera label followed the measurement and is not independent detector validation.

Buried direction: the current verifier solves closed-set preservation. The still-open, ISR-relevant problem is recognition-light unexpected change for regions that were never named or grounded. Ordinary flow, ROI tracking, and segmentation must define that residual first.

2026-08-15 v1.0 blind annotation corrects the tolerance

independent disagreement Claude Opus 4.8 received a sealed folder with 32 original images, neutral labels, schema, and validator—but no Gemini boxes, thresholds, or results. It returned 192 validated object records after HSV segmentation, visual verification, and manual corrections.

5.48 px
Gemini–Claude center median · 192
28.02 px
corrected locked maximum · 96
56.04 px
twice corrected maximum
214.85 px
corrected camera maximum · 48

After the v1.2 scale correction, three locked residuals still exceed 20 px: Blue-A's cup at 22.06 px and Blue-C's bowl at 20.05 and 28.02 px. None exceeds 60 px. Thus the rejection of 20 px survives, while 60 px remains explicitly heuristic.

Median cross-model agreement is tight, but six large pedestal-mug discrepancies arise because the arithmetic box center changes depending on whether the handle and full base count toward the visible extent. The correction is conceptual as well as numerical: a box center is an annotation convention, not a unique physical object center. Future metric work should declare a contact point or object-specific anchor.

Protocol correction: the visible scale is labeled in inches. Claude correctly returned all original ruler fields as ambiguous because the packet demanded centimeter ticks. V0.9 stays unchanged as executed; it is not retroactively rescored.

Interrupted addendum: before its token limit, Claude saved 3-inch/8-inch points for all 32 frames but not the requested final files. The 24 reference/locked frames yield 219.64 px/in (86.47 px/cm) median. This contextualizes 28.02 px as roughly 0.32 cm and 60 px as roughly 0.69 cm, but does not create a global scale for raised objects across the perspective field.

2026-08-15 v0.9 prospective related-set confirmation

architecture survives The user confirmed that photo-v0.4 and photo-v0.5 followed the v0.3 file order. Roles and the operational rule—20 px under locked capture, ABSTAIN on detected camera stress—were written before either set was inspected or sent to Gemini.

12 / 14
direct Gemini · two false passes
9 / 9
grounded accuracy among decisions
9 / 14
grounded coverage
5
camera-stress abstentions

Both direct errors reproduced the exact hard mechanism: correct mug placement plus collateral pink-bowl motion. Direct Gemini said the protected objects appeared unchanged; the grounded path kept placement true but measured registered bowl displacement of 149.1 px and 189.3 px and returned FAIL. Across photo-v0.2 through v0.5, direct Gemini is now 24/28, with all four errors on this one mechanism.

Photo-v0.4 IMG_4 unexpectedly crossed the pre-existing camera gate (scale 0.977; full-resolution translation 50.2 px). Its frozen task-success label was retained, but the grounded verdict became ABSTAIN. This is a nuisance-condition correction from declared measurements, not outcome relabeling.

Scope: these are correlated related-set confirmations, not fourteen independent trials. The 20 px rule is preserved as the prospective execution condition; the later blind audit rejects it as a future cross-annotator threshold. Grounded correctness is conditional on 64.3% coverage.

2026-08-15 v0.8 camera parallax splits the calibration regime

identity stable All 32 calibration frames returned all six requested boxes: 192 boxes, no malformed outputs, no abstentions. The similar white mugs remain distinct, and the blue cup does not collapse into the blue target under locked capture.

2.57 px
locked repeat median · 96 residuals
9.84 px
locked repeat maximum
64.63 px
camera-perturbed p95 · 48 residuals
82.56 px
camera-only maximum · exceeds 70 px

Blue-D exceeds the old tolerance on a camera-only frame. Overlay audit confirms correct mug, cup, target, bowl, and apple identities. Multiple raised objects shift differently after whole-frame registration: this is depth-dependent parallax, not a semantic swap or a moved object.

Design correction: do not inflate one global threshold. For locked capture, carry a provisional 20 px tolerance—twice the observed maximum rounded upward. For detected camera stress, return ABSTAIN on preservation until per-object tracking with uncertainty, depth-aware reprojection, stereo, or another measurement surface is available.

Protocol deviation: grounding was run before independent annotations were sealed. Cached boxes must now be withheld from a separate human annotator; annotation by the same agent that audited them would not be blind.

At this stage: 20 px remained a candidate until independent boxes and ruler endpoints could quantify absolute error and physical scale. No photo-v0.4/v0.5 outcomes had been uploaded or read into this calibration; it was subsequently frozen as the provisional operational rule for v0.9.

2026-08-15 v0.7 independent calibration capture admitted

both sets usable The White and Blue rearrange folders each contain four materially different layouts with an empty target, reference, two no-object-move repeats, and one camera-only perturbation. The similar white mug replaces the earlier black distractor. The Blue condition replaces the white case with a blue cup, deliberately testing target-color versus object-identity grounding.

8
independently arranged layouts
32
calibration-only photographs
24 / 24
camera-condition classifications correct
0.627
minimum registration inlier ratio

All 16 nominally locked repeats remain below the established camera-stress gate; all eight frame-4 perturbations exceed it. Every comparison remains registerable. Small natural camera/exposure variation in frames 2/3 is retained as realistic calibration noise rather than sanitized away.

Admission decision: use both sets. The blue cup is a useful hard condition, not a protocol violation. At this stage photo-v0.4 and photo-v0.5 were inventoried as two additional eight-frame outcome sets and deliberately left unevaluated until the role and camera-policy rule was explicit; v0.9 records their later evaluation.

Still pending: independent human boxes/centers and ruler endpoints. Grounding is now complete, but registration plus learned boxes cannot independently prove that every object was physically motionless; immobility currently rests on the capture protocol plus visual review.

2026-08-15 Same-state repeat calibration audit

threshold survives The v0.2 wrong-object repeat and v0.3 success repeat were treated as same-state localization-noise controls. Each first frame was registered to its repeat; clean-prompt Gemini centers were then compared through the homography.

12
same-state object/target residuals
1.34 px
median repeat residual
32.19 px
worst repeat residual · pink bowl
4.85×
collateral minimum / repeat maximum

The worst repeat drift doubled is 64.38 px, just below the frozen 70 px tolerance. The smallest actual collateral drift is 156.25 px. No threshold or label was refit. Nearly all non-bowl repeat residuals are below 4 px, implicating irregular-object box variability more than global registration.

Limit: only two same-state pairs, one capture regime, and model-produced boxes. Pixel tolerances change with range, focal length, crop, and resolution. The next real calibration requires true no-move repeats and independent annotations gathered before outcome scoring.

2026-08-15 v0.3 changed-object replication

exact mechanism repeat A second eight-photo capture replaced the patterned distractor mug with a small black mug, enlarged the target, reshot the layout, and added a third passing outcome. The 70 px preservation tolerance and model settings were unchanged.

4 / 7
registered pixel delta · all three successes rejected
6 / 7
Gemini direct · same collateral false pass
6 / 7
grounded coverage · six decided correctly
13 / 14
combined grounded coverage · 13/13 decided correctly

On v0.3 IMG_3, direct Gemini again returned PASS at 0.95 confidence and said the pink bowl appeared unchanged. Registered grounding measured the bowl 187.3 px from its expected center and failed preservation while passing target placement. The three v0.3 successes topped out at 28.6 px protected-object residual.

Cross-version sensitivity: under the cleaned grounding prompt, the largest success residual is 39.4 px; the smallest true collateral residual is 156.2 px. Every threshold strictly inside that interval yields the same thirteen decided grounded labels. No threshold was adjusted for v0.3.

audit correction The initial grounding prompt carried an irrelevant synthetic-era sentence about “large and small blue squares.” After removing it and versioning the cache, both photo sets were regrounded. v0.2 remained 7/7; v0.3 IMG_6 returned an impossible pink-bowl box and correctly triggered ABSTAIN. The frame was not retried. The earlier 14/14 observation is preserved as prompt-contaminated history, not the standing result.

Photo v0.3 common before frame
v0.3 BEFORE. Larger target and black wrong-object mug; the white pedestal mug remains the commanded object.
Photo v0.3 collateral bowl failure
Repeated hard failure. Correct mug placement plus bowl displacement. Direct judgment misses it again; explicit relation checking retains it.

Limit: this is a stronger mechanism replication, not independent benchmark accuracy. The room, table, task grammar, target object, and several protected objects are shared; labels were inspected; boxes remain model-produced rather than human annotated.

2026-08-15 First registered multi-object photo dry run

mechanism replicated Eight new iPhone photographs formed one common BEFORE frame and seven correlated AFTER outcomes. Roles were frozen before the API call: two successes, a target success with collateral bowl movement, two wrong-object repeats, a wrong object plus collateral movement, and success/failure camera-shift variants.

5 / 7
registered pixel delta · rejected every frame
6 / 7
Gemini direct · one collateral false pass
7 / 7
Gemini boxes + registered relations
7 / 7
camera-stress labels detected
framemechanismtruthpixel deltadirectgrounded
IMG_2successpassfailpasspass
IMG_3success + pink bowl movesfailfailpassfail
IMG_4 / IMG_5wrong floral mug · repeatfailfailfailfail
IMG_8wrong mug + bowl movesfailfailfailfail
IMG_7success + camera shiftpassfailpasspass
IMG_6wrong mug + camera shiftfailfailfailfail

The direct miss was explicit, not ambiguous: on IMG_3 the model said the pink bowl appeared unchanged and returned PASS at 0.95 confidence. Under the cleaned grounding prompt, the grounded path measured a 207.9 px registered bowl displacement and failed the preservation predicate. Positive frames had worst protected-object residuals of 37.4 and 39.4 px; collateral frames were 156.2 and 207.9 px. The provisional 70 px threshold therefore sits in a broad observed gap, although it is not yet independently calibrated.

Common before frame for the photo dry run
Common BEFORE. White pedestal mug begins at rear right; the blue strip marks its destination. Apple, bowl, floral mug, and white case are protected.
Target success with collateral bowl displacement
Hard failure. The requested mug reaches the target, but the pink bowl also moves. Direct judgment passed; explicit registered relations failed.

Limit: one inspected layout, correlated frames, visually distinct mugs, developmental margins, and learned boxes rather than human annotations. This is mechanism evidence, not a 7-trial accuracy estimate. The negative pixel-delta result matters: parallax from small viewpoint changes creates residuals around stationary raised objects, so whole-frame change is an attention cue, not yet an object-compliance measure.

2026-08-14 Adversarial audit and v0.3 correction

claim narrowed A full code-and-artifact audit found that the conceptual direction survived but the v0.2 harness overstated nuisance coverage and experimental independence. Historical artifacts remain untouched; the corrected generator and local artifact are versioned v0.3.

112 / 112
unique corrected task/input bundles
28
unit and regression tests green
100 / 100
classical RGB decided correctly · 12 height abstentions
24 / 24
legacy fusion survives measured-verdict recomputation

Corrected: exact static background replay; clutter in both frames; unique controls; distinct corridor failures; RGB-detected target map instead of oracle coordinates; raw-response-only cache with current predicate reevaluation; strict grounding validation; signed residual convention; measured rather than ground-truth legacy fusion; media hashes and stronger manifest checks.

Still open: no new Gemini calls validate the corrected generator; the original nuisance gate remains unmet; flat decoy is an explicit RGB-equivalent information-limit construct rather than a photorealistic print; box-level depth attribution, real calibration, prompt robustness, and held-out multi-object rollout data remain unresolved.

2026-08-14 Legacy real-frame bridge normalized

new Existing Robotics-ER photographs were not discarded or casually folded into a new claim. They were normalized into a 126-record manifest with explicit provenance, task endpoint, media, ground truth, and limitations. All are marked legacy exploratory · not held out.

113
complete real stereo height scenes
13
calibrated placement scenes
10
stored direct-Gemini flat-print false passes
1
incomplete stereo pair excluded explicitly
reported resultmachine-readable selectioncorrect
parallax overallall surface_height records99 / 113
clean non-emptynon-empty · no detect failure · finite parallax93 / 103
stored direct VLMprior model result exists17 / 27
retrospective appearance + surfaceprior result · clean detection · raised/flat24 / 24
stored placement agreementplacement + prior verdict13 / 13*

* Only one positive placement case. This is not balanced placement accuracy. The 24/24 fusion excludes three detector failures, is retrospective on development data, and does not solve a printed semantic decoy raised on a real box.

A full-page printed cup decoy
Flat print. Direct Gemini passed this real photograph as a cup resting on the table; parallax reads the content at the table plane.
A real orange mug
Real raised control. Appearance and parallax agree. The point is selective disagreement, not geometry disagreeing everywhere.

2026-08-14 v0.2 hard-control pilot

Historical record: the initial synthetic experiment removed printed IDs, added similar distractors, metric language, matched latent initial states, and classical baselines. The API slice itself used only clean index-000 scenes; nuisance diversity was present in the local corpus, not demonstrated for Gemini.

27
unique Gemini pilot inputs · 28 labeled rows
2 FP
direct judge · 25/27 unique
0 FP
Gemini boxes + deterministic predicates
3 / 3
each false-pass mechanism repeated
conditioncoveragedecided accuracyfalse pass
oracle geometry27 / 27100.0%0
classical RGB components24 / 27100.0%0
classical components + depth27 / 27100.0%0
Gemini direct27 / 2792.6%2
Gemini boxes + RGB predicates*24 / 27100.0%0
Gemini boxes + predicates + depth*27 / 27100.0%0

* Historical grounded conditions used an oracle target rectangle for containment. Corrected v0.3 detects the target from RGB.

Synthetic flat-decoy control
Flat-decoy failure. Direct Gemini passed 3/3 and cited the rendered shadow as height evidence. RGB predicates abstained; depth predicates failed.
Synthetic similar wrong-object control
Similar-object / margin failure. Direct Gemini passed 3/3. Twice it estimated 1–2 cm against a 3.5 cm requirement and still concluded “likely met.”
The model can describe evidence inconsistent with the threshold and still narratively round the verdict up to success. The firewall matters because prose reasoning is not a measurement surface.

00 The one-line thesis

Use a learned model to tell us what might matter and where it is; use independent, declared measurement surfaces to decide whether the task's physical and metric predicates are actually satisfied; return ABSTAIN whenever the required surface is absent.

The potential spatial-field contribution is not another left-of function. It is a recognition-light account of what else changed, what relates to what, and where expensive verification should look next—but only if it outperforms ordinary baselines on real rollouts.

01 Why a third project

Robotics-ER had been run as a mostly single-object, height-above-plane instrument. ISR and later VTL work had moved toward islands and relational scene structure because whole-frame mass becomes confounded when several objects or clutter share the image. The question was not whether to merge the repositories. It was whether the broader spatial work had a credible robotics home.

The risk was obvious: robotics already has excellent deterministic geometry, tracking, occupancy, collision checking, SLAM, and scene graphs. An outsider can easily reinvent a weaker version and mistake unfamiliarity for originality. So this project was deliberately split out and built as a sequence of gates.

Take from both if valid; start over if necessary; do not import ISR merely because its vocabulary sounds relational.

02 What was built

03 The task families

familyclaim being auditedhard control
place_insidelarge blue square inside target and raisedflat decoy · similar small-blue object
left_of3.5 cm edge-to-edge marginwrong object · near metric miss
separate9 cm minimum clearancevisually plausible insufficient gap
move_without_disturbingtarget move while named B remains fixednamed collateral movement
move_preserve_scenetarget move while every known tracked object remains fixedunnamed C moves
clear_corridorstraight path clear of every blocker by 1.8 cmintervening island/object
Before scene with five objects
Registered before frame: the instruction names the target but not every protected object.
After scene with target moved and collateral object displaced
After frame: the requested placement succeeds while the yellow object moves. Pairwise target/B verification can miss the scene-level failure.

04 What the evidence changes

candidate claimcurrent readingwhy
New relation geometrydo not claimClassical components + existing geometry solve the synthetic controls.
Learned judge is enoughrejectedStable false passes on height and metric/identity controls; real flat-print errors already exist.
Learned perception is uselessrejectedGemini boxes grounded objects well; the error lived in unconstrained verdict synthesis.
Independent measurement + abstentionsupported narrowlyNo wrong decisions in the historical grounded conditions; the first real multi-object grounded run was 7/7 while direct judgment false-passed collateral movement.
ISR mass-field delta improves motion auditingreopened narrowlyIt lost 0/4 versus 4/4 retrospectively, then matched 8/8 prospective collateral alarms, matched 0/4 CAL false alarms, and avoided one ordinary photometric false failure. Localization and replication are still absent.

05 Where ISR may enter — and the gate

Candidate A · islands as recognition-light multi-track. Segment the field into stable regions, then attach ordinary geometry or parallax per island. Anchor/satellite roles may provide task-conditioned priority without requiring semantic recognition everywhere.

Candidate B · corridors as relational interference. A corridor between islands can express “between” and obstruction, but must be compared with standard occupancy/collision geometry.

Candidate C · layered field delta as attention gate. The first retrospective test was negative: edge/tone/chroma residuals fragmented moved objects while ordinary Lab detected 4/4. The frozen prospective capture changes that reading: both methods alarm on 8/8 declared collateral moves, while ISR alone ignores one stationary object's appearance shift.

gate failed The mass-residual + connected-area candidate does not enter the robotics path. It has one prospective specificity advantage, but no movement missed by all ordinary arms, 9.0% versus 70.2% moved-object coverage, 0.086 versus 0.385 IoU, equal abstention coverage, 4.44× Lab runtime, and marked +10% threshold sensitivity. Photometric invariance may be studied separately, but cannot rescue this candidate under the frozen gate.

06 Honest limits

07 Decisions & dead ends

decisionstatusreason
Merge Robotics-ER and ISR nowrejectedWould entangle hypotheses before contribution is established.
Start with ISR islands/corridorsrejectedOrdinary baselines must define the residual first.
Treat whole-frame mass as object locationrejectedConfounded by clutter and multi-object scenes.
Print object IDs in the benchmarkremovedTurned semantic grounding into OCR/token matching.
Let RGB infer physical height from shadowrejectedThe flat-decoy false pass demonstrates why appearance is not independent height evidence.
Count abstention as an embarrassing missrejectedAbstention is the correct result when the required measurement surface is absent.
Use legacy real data as held outrejectedIt informed prior method development; provenance is explicit.
Keep a third independent projectchosenAllows clean ablation, ordinary controls, and easy abandonment of invalid ideas.
Promote ISR mass-field deltarejectedIt matches 8/8 at the frozen point and avoids one photometric error, but localizes far less, costs 4.44× Lab, and drops to 5/8 at +10% score threshold. No admission route survives.

08 Bookmarked source-project issues

Held for a separate cleanup conversation. They matter if code or claims are lifted into this project, but were not silently changed in Robotics-ER.

bookmarkimplication here
CLI defaults to isolate=noneNever make whole-frame measurement the implicit real-scene path.
μ terminology driftDo not inherit a name whose paper and implementation measure different constructs.
8D / 9D kernel naming driftFreeze dimensional terminology before citing or lifting a kernel.
86/89 selection was prose-onlyThe new 93/103 rule is executable; it does not reconstruct or validate the historical 86/89 subset.
Coverage gaps in CLI/real pipeline/calibration/malformed/stereoThe new runner tests cache, selection, controls, abstention, and manifests; real calibration and stereo regression remain open.

09 Completed registered experiment and next gate

V2.1 is sealed complete. V2.2 is frozen before new capture and tests whether the two post-hoc repairs generalize rather than merely replay their development sets.

  1. Preserve v2.1 unchanged as the prospective record; do not replace its compiled score with a post-hoc repair.
  2. Capture two new seven-pair layouts: CAL, SUCCESS, LIGHT, PINK, TEXTURED_1CM, SMOOTH_1CM, and CAMERA; every pair receives a fresh BEFORE.
  3. Seal BEFORE-only extraction and normalized compilation before exposing AFTER images.
  4. Compare the cascaded candidate with direct judgment, target-only, unconditional-complement, and frozen Lab-only controls.
  5. Require zero false passes, zero SUCCESS/LIGHT false fails, quiet CAL, all camera abstentions, and both paired PINK label flips.
V2.2's perfect replay is a reason to run the next experiment, not the result of that experiment.

10 File trail

relational_auditor/model.py — scene/evidence/verdict types
relational_auditor/predicates.py — deterministic audit predicates
relational_auditor/synthetic.py — matched benchmark generator
relational_auditor/baselines.py — classical RGB/depth controls
relational_auditor/gemini.py — direct and grounded model paths
relational_auditor/run.py — cache, repeats, metrics, runner
relational_auditor/real_manifest.py — real-data import + validation
relational_auditor/legacy_audit.py — explicit legacy selections
relational_auditor/photo_dry_run.py — registration + recognition-light delta control
relational_auditor/photo_gemini.py — direct and grounded photo conditions
relational_auditor/photo_calibration.py — same-state registered grounding-noise audit
relational_auditor/calibration_capture.py — v0.7 registration and camera-condition audit
relational_auditor/calibration_grounding.py — calibration-only object grounding and repeat residuals
relational_auditor/annotation_audit.py — sealed Claude/Gemini localization and repeat comparison
relational_auditor/ruler_scale_audit.py — preliminary inch-scale audit of interrupted Claude return
relational_auditor/adversarial_checkpoint.py — harmonized denominators, sensitivity, and claim audit
relational_auditor/conventional_tracking.py — calibrated SIFT ROI preservation baseline
relational_auditor/ordinary_motion.py — sparse flow and local-change ordinary controls
relational_auditor/full_frame_motion.py — protected-box-withholding motion-cluster baseline
relational_auditor/unlisted_protocol.py — v1.6 capture-manifest validator
relational_auditor/isr_change_fields.py — frozen ordinary Lab and ISR-derived mass-residual comparison
relational_auditor/unlisted_evaluate.py — sealed prospective runner accepting target-only measurements
relational_auditor/unlisted_admission.py — v1.6 hashes, registration, and camera-gate admission
relational_auditor/unlisted_grounding.py — target-only Gemini grounding and sealed measurements
relational_auditor/unlisted_adversarial_audit.py — disclosed post-hoc audit of the sole ordinary false failure
relational_auditor/unlisted_annotation_audit.py — decoded blind changes, registration, tracking, and ruler scale
relational_auditor/unlisted_localization_audit.py — diagnostic moved-sweep overlap comparison
relational_auditor/unlisted_runtime_audit.py — identical-machine frozen-arm benchmark
relational_auditor/unlisted_full_frame_evaluate.py — restored omitted preregistered v1.5 arm
relational_auditor/unlisted_sensitivity_audit.py — post-hoc one-factor threshold grid
relational_auditor/unlisted_adversarial_checkpoint.py — authoritative chronology, corrections, gates, and denominators
relational_auditor/model_authored_monitor.py — restricted-DSL validation and mechanical contract-scope classification
relational_auditor/model_authored_monitor_author.py — blind BEFORE-only authoring runner with source-hash preflight, metadata stripping, opaque caching, and response sealing
relational_auditor/model_authored_monitor_outcomes.py — sealed 12-pair direct-judgment arm and separated error metrics
relational_auditor/model_authored_monitor_report.py — reproducible primary metrics and disclosed post-hoc semantic audit
relational_auditor/monitor_compiler.py — strict schema-to-execution boundary, permission lint, calibration gate, and unsupported-primitive rejection
relational_auditor/model_authored_monitor_compile_audit.py — post-hoc B1/B2 compilability audit of all 16 sealed specifications
relational_auditor/contract_extractor.py — v2.1 declarative-clause validator and deterministic typed write-set compiler
relational_auditor/spec_execution_protocol.py — neutral-packet, paired-ground-truth, balance, and gold-contract validator
relational_auditor/spec_execution_analysis.py — frozen separated denominators and paired permission-flip analysis
relational_auditor/spec_execution_admission.py — v2.1 readability, uniqueness, registration, camera-gate, and image-hash admission without outcome scoring
relational_auditor/verify_spec_execution_freeze.py — v2.1 pre-capture SHA-256 verification
relational_auditor/spec_execution_extraction_packet.py — opaque hash-sorted BEFORE-only extraction packet and private mapping seal
relational_auditor/spec_execution_contract_author.py — raw-first extraction cache, frozen validation, compilation, incident binding, and response seal
relational_auditor/spec_execution_direct_outcomes.py — paired direct BEFORE/AFTER learned-judgment arm
relational_auditor/spec_execution_measure.py — shared grounding, registration, change measurement, and four deterministic outcome arms
relational_auditor/spec_execution_finalize.py — sealed extraction diagnostics and separated final denominators
relational_auditor/contract_normalizer_v22.py — exact executor-equivalent split-clause normalization with provenance and rejection guards
relational_auditor/spec_execution_v22_development.py — post-hoc Lab-proposal/flow-corroboration replay, cross-dataset check, and sensitivity grid
relational_auditor/spec_execution_v22_prepare.py — neutralized capture validation, registration admission, hashes, and BEFORE-only packet seal
relational_auditor/spec_execution_v22_contract_author.py — sealed BEFORE-only extraction through normalization and exact permission authority
relational_auditor/spec_execution_v22_direct.py — sealed paired direct-judge arm on opaque pair IDs
relational_auditor/spec_execution_v22_measure.py — frozen Lab proposal, flow corroboration, grounding, and ablation arms
relational_auditor/spec_execution_v22_finalize.py — separated denominators, LIGHT audit, and final artifact seal
relational_auditor/spec_execution_v22_adversarial.py — post-result integrity replay, localization audit, and rejected v2.3 candidate comparison
relational_auditor/v23_qualification_audit.py — local v2.3 stage, exposure, registration-direction, Lab, and flow qualification audit
relational_auditor/verify_v22_freeze.py — pre-capture v2.2 method-hash verifier
relational_auditor/gemini_schema_probe.py — no-image transport-schema compatibility diagnostic
Claude_review_v1.6/ — neutralized 16-pair blind masks/contact/scale packet
data/unlisted-v1.6/annotation-key.json — sealed neutral-to-source mapping kept outside the packet
tests/test_relational_auditor.py — 109 tests
AUDIT_CORRECTIONS_V0.3.md — adversarial correction record
PREREGISTRATION.md — initial hypothesis and gates
RESULTS_PILOT_V0.2.md — hardened synthetic/model result
RESULTS_LEGACY_REAL_V0.3.md — legacy bridge result
RESULTS_PHOTO_DRY_RUN_V0.4.md — first registered multi-object photo result
RESULTS_PHOTO_REPLICATION_V0.5.md — changed-object, larger-target replication
RESULTS_REPEAT_CALIBRATION_V0.6.md — retrospective repeat-noise threshold audit
RESULTS_INDEPENDENT_CALIBRATION_V0.8.md — eight-layout grounding calibration and regime split
RESULTS_PROSPECTIVE_CONFIRMATION_V0.9.md — frozen-role v0.4/v0.5 outcome confirmation
RESULTS_CLAUDE_ANNOTATION_AUDIT_V1.0.md — blind second-model calibration audit
ADVERSARIAL_CHECKPOINT_V1.2.md — current authoritative red-team checkpoint
RESULTS_CONVENTIONAL_TRACKING_V1.3.md — ordinary ROI-tracking result and raised ISR gate
RESULTS_ORDINARY_MOTION_V1.4.md — optical-flow replication and local-change failure analysis
RESULTS_FULL_FRAME_MOTION_V1.5.md — recognition-light box-withholding residual
UNLISTED_COLLATERAL_PROTOCOL_V1.6.md — 32-image decisive capture protocol
ISR_COMPARISON_PREREGISTRATION_V1.7.md — frozen methods, predictions, exclusions, and boundary correction
RESULTS_ISR_CHANGE_FIELDS_V1.7.md — negative ISR result and ordinary-control mechanism
RESULTS_UNLISTED_V1.6.md — prospective unlisted result and closed ISR admission gate
ADVERSARIAL_LOOKBACK_V1.8.md — authoritative post-unsealing correction record
RESEARCH_AGENDA_V1.9.md — literature positioning and separate contract-authority, validity, and ISR gates
RESEARCH_AGENDA_V2.4.md — phase-one closure, QA application wedge, open-problem requirements, and phase-two sequence
CONTRACT_INTEROP_PROTOCOL_V2.4.md — authority-preserving semantic-IR protocol and rejection gates
RESULTS_CONTRACT_INTEROP_V2.4.md — 22/28 to 24/28 readiness result with zero authority widening
artifacts/contract-interop-v2.4/ — sealed corpus replay, per-contract IR, normalization, and rejections
EXTERNAL_VLA_DATA_SELECTION_V2.5.md — external-corpus comparison, DROID-100 feasibility protocol, and label-policy boundary
DROID100_ADAPTER_PROTOCOL_V2.5.md — frozen one-sided authority, mask, visibility, and outcome vocabulary
DROID100_GROUNDING_PROTOCOL_V2.5.md — fixed-role authority, visibility admission, and no-outcome execution discipline
RESULTS_DROID100_GROUNDING_V2.5.md — 0/12 endpoint coverage result, adversarial sensitivity, and temporal next gate
DROID100_MULTISURFACE_PROTOCOL_V2.6.md — three-way temporal/telemetry protocol and containment boundary
RESULTS_DROID100_MULTISURFACE_V2.6.md — 1/2/9 integrated observable-chain result and partial-target next gate
EXTERNAL_SCENE_DATA_SELECTION_V2.7.md — Language-Table primary pilot, BridgeData stress role, POV correction, and camera/obstruction gates
LANGUAGE_TABLE_GROUNDING_PROTOCOL_V2.7.md — fixed eight-object authority and endpoint-window visibility admission
RESULTS_LANGUAGE_TABLE_SURFACE_V2.7.md — 48/48 registration, 41/48 target coverage, and 35/48 complement coverage result
data/language-table-v2.7/ — original shard, CRC scan, rejected/final freezes, prompt, schema, and grounding freeze
artifacts/language-table-v2.7/ — sealed endpoint packet, contact sheets, registration screen, 96-response cache, and visibility result
Claude_review_language_table_v2.7/ — sealed 21-episode blind visibility packet, exact template, task, and validator
RESULTS_LANGUAGE_TABLE_CHANGE_V2.8.md — registered endpoint geometry, rejected temporal threshold, source mismatch, and fixed-HSV comparator
artifacts/language-table-v2.8/ — continuous registered-change result and no-override component comparator
RESULTS_LANGUAGE_TABLE_VALIDITY_V2.9.md — fresh-shard natural calibration, duplicate correction, frozen support gate, and 0/6 unique held-out false alarms
data/language-table-v2.9/ — official shard 00001, rejected/final no-motion freezes, and cross-shard exclusion record
artifacts/language-table-v2.9/ — natural no-motion packets, rejected threshold, and prospective attempt-4 result
LANGUAGE_TABLE_LEARNED_SURFACE_PROTOCOL_V3.0.md — frozen learned-box authority and metric-validity comparator
RESULTS_LANGUAGE_TABLE_LEARNED_SURFACE_V3.0.md — 23/23 semantic coverage, failed learned metric calibration, and P18 representation mismatch
artifacts/language-table-v3.0/ — sealed Gemini pair cache, localization result, and learned-surface evaluation
LANGUAGE_TABLE_BIDIRECTIONAL_FLOW_PROTOCOL_V3.1.md — frozen two-direction pixel-flow authority, support rule, and calibration split
RESULTS_LANGUAGE_TABLE_BIDIRECTIONAL_FLOW_V3.1.md — corrected 4/1/1 unique held-out result, complement verdict, and low-texture failure analysis
artifacts/language-table-v3.1/ — sealed bidirectional-flow result
data/language-table-v3.1/flow-freeze.json — pre-execution protocol, code, input, and policy hashes
LANGUAGE_TABLE_DENSE_SUPPORT_PROTOCOL_V3.2.md — frozen role-local HSV partition, validity ordering, and prospective gates
RESULTS_LANGUAGE_TABLE_DENSE_SUPPORT_V3.2.md — 6/6 unique quiet controls, 2/0/3 complement result, P10 recovery, and overlap boundary
artifacts/language-table-v3.2/ — sealed dense-support result and P10 visual reference
data/language-table-v3.2/dense-support-freeze.json — pre-measurement protocol, code, input, and policy hashes
ADVERSARIAL_CHECKPOINT_V3.3.md — duplicate leakage correction, grounding-transport mismatch, verdict factorization, and forward gates
artifacts/language-table-v3.3/adversarial-correction.json — reproducible hash-grouped denominators and unchanged task summaries
PRIMARY_PATH_HANDOFF_V3.4.md — frozen candidate, decisive untouched study, decision gates, and exact resume point
ISR_FORK_LAB_BOOK.html — separate relational-ISR research fork; inherits context but no firewall credit
data/droid100-feasibility-v2.5/selection.json — disclosed post-screen 12-episode stratum and source hashes
relational_auditor/droid100_adapter.py — deterministic source validation, frame extraction, sealing, and verification
relational_auditor/droid100_grounding.py — frozen-role Gemini runner, local validator, coverage adjudicator, cache, and seal
artifacts/droid100-feasibility-v2.5/ — 48-image unscored packet, manifest, seal, and contact sheet
artifacts/droid100-temporal-v2.6/ — sealed telemetry-aligned exterior/wrist packets, cached groundings, integration, rejected apparatus, and contact sheets
artifacts/external-vla-selection-v2.5/droid-100-metadata-audit.json — hashed 100-episode task/reward intersection audit
artifacts/external-vla-selection-v2.5/droid-100-visibility-screen.json — hashed two-view start/end screen and adapter admission consequence
artifacts/external-vla-selection-v2.5/droid100-tasked-camera*-contact-sheet.jpg — all 46 tasked start/end pairs in both exterior views
MODEL_AUTHORED_MONITOR_PILOT_V2.0.md — frozen retrospective authoring comparison, arms, metrics, and claim limits
RESULTS_MODEL_AUTHORED_MONITOR_V2.0.md — contract-authoring and direct-judgment results, falsifier, and post-hoc semantic audit
PROSPECTIVE_SPEC_EXECUTION_PROTOCOL_V2.1.md — paired preserve-all/selective-permission prospective protocol
RESULTS_SPEC_EXECUTION_V2.1.md — sealed prospective result, transport deviations, controls, falsifiers, and next gate
DEVELOPMENT_REPAIR_V2.2.md — post-hoc compiler and measurement repairs, mechanism audit, replay, sensitivity, and limits
ADVERSARIAL_CHECKPOINT_V2.2.md — pre-capture red-team corrections to sensitivity, flow aggregation, permission authority, and malformed-input containment
PROSPECTIVE_REPAIR_PROTOCOL_V2.2.md — frozen new-layout LIGHT/texture/selective-permission confirmation protocol
RESULTS_SPEC_EXECUTION_V2.2.md — sealed prospective partial confirmation, gate audit, aliasing limit, and LIGHT mechanism
ADVERSARIAL_CHECKPOINT_V2.2_POSTRESULT.md — authoritative post-result integrity, localization, Goldilocks, authority, and ISR checkpoint
V2.3_ILLUMINATION_VALIDITY_QUALIFICATION.md — pre-capture LIGHT versus texture-poor-motion challenge-set protocol
RESULTS_V2.3_QUALIFICATION_ATTEMPT_1.md — failed-but-informative first qualification result and attempt-two apparatus changes
artifacts/v2.3-qualification/ — sealed local audit, registrations, proposals, flow tracks, EXIF, and source hashes
RESULTS_V2.3_QUALIFICATION_ATTEMPT_2.md — repaired registration/light controls with the pivotal tracking-poor challenge still absent
artifacts/v2.3.2-qualification/ — sealed attempt-two audit and source hashes
data/model-authored-monitor-v2.0/ — prompt, JSON schema, neutral hash manifest, and withheld mapping key
artifacts/model-authored-monitor-v2.0/ — sealed raw authoring/direct responses, hashes, and resumable opaque caches
data/spec-execution-v2.1/ — frozen extraction packet, gold contracts, neutral capture manifest, private role key, policies, checklist, and hashes
artifacts/spec-execution-v2.1/ — capture/extraction/direct/measurement/final outputs, incident ledger, and SHA-256 seals
artifacts/spec-execution-v2.2-development/ — explicitly post-hoc replay, cross-dataset diagnostics, sensitivity, and seal
data/spec-execution-v2.2/freeze.json — frozen candidate/protocol hashes before new capture
data/spec-execution-v2.2/ — neutral capture manifest, private role key, inherited task contracts/schema/policy, and revision-two method freeze
artifacts/spec-execution-v2.2/ — sealed admission, extraction, direct, measurement, final outputs, caches, and SHA-256 seals
CALIBRATION_CAPTURE_PROTOCOL_V0.7.md — independent no-move capture, annotation, and freeze gate
REAL_FRAME_PROTOCOL_V0.3.md — new capture protocol
data/real-v0.3/legacy-manifest.jsonl — 126 normalized records
artifacts/gemini-pilot-v0.2.json — historical 28-row pilot
artifacts/gemini-error-repeat-v0.2.json — targeted stability run
artifacts/legacy-real-audit-v0.3.json — machine-readable denominators
artifacts/local-v0.3.json — corrected 112-input local audit
artifacts/photo-v0.2/local-analysis.json — registration and delta evidence
artifacts/photo-v0.2/gemini-analysis.json — reconciled direct/grounded photo outcomes
artifacts/photo-v0.3/gemini-analysis.json — v0.3 direct/grounded replication outcomes
artifacts/photo-v0.4/gemini-analysis.json — first frozen-role confirmation output
artifacts/photo-v0.5/gemini-analysis.json — second frozen-role confirmation output
artifacts/photo-repeat-calibration-v0.6.json — machine-readable repeat residuals and margins
data/photo-calibration-v0.7/manifest.template.json — four-layout calibration capture template
data/unlisted-v1.6/manifest.template.json — two-layout unlisted-motion capture template
data/photo-calibration-v0.7/manifest.json — admitted eight-layout White + Blue capture manifest
artifacts/photo-calibration-v0.7/registration-audit.json — 24 camera-condition registration checks
artifacts/photo-calibration-v0.7/grounding-analysis.json — 192 cached boxes and 144 repeat residuals
data/photo-calibration-v0.7/claude-annotations.json — sealed 192-record blind annotation return
artifacts/photo-calibration-v0.7/claude-comparison.json — cross-model centers and registered residuals
artifacts/photo-calibration-v0.7/ruler-scale-preliminary.json — 32-image interrupted ruler-addendum audit
artifacts/adversarial-checkpoint-v1.2.json — recomputed denominators and sensitivity
artifacts/conventional-tracking-v1.3.json — calibration records, tracks, verdicts, and sensitivity
artifacts/ordinary-motion-v1.4.json — flow/change calibration records and outcomes
artifacts/full-frame-motion-v1.5.json — full-frame cluster calibration and outcomes
artifacts/isr-change-fields-v1.7.json — calibration, region geometry, verdicts, and failures for both frozen arms
artifacts/unlisted-v1.6/capture-admission.json — 32 frozen hashes and 16 measured registrations
artifacts/unlisted-v1.6/target-grounding.json — cached target-only semantic boxes
artifacts/unlisted-v1.6/frozen-evaluation-v1.7.json — prospective ordinary/ISR outcome rows
artifacts/unlisted-v1.6/adversarial-audit.json — disclosed stationarity check of B_SUCCESS
artifacts/unlisted-v1.6/independent-annotation-audit.json — blind-return adjudication and ruler scale
artifacts/unlisted-v1.6/localization-audit.json — verified-polygon coverage, precision, and IoU
artifacts/unlisted-v1.6/runtime-audit.json — five-repeat per-arm timing
artifacts/unlisted-v1.6/full-frame-sift-evaluation.json — omitted-arm corrected outcome rows
artifacts/unlisted-v1.6/threshold-sensitivity.json — score and area-cutoff robustness
artifacts/unlisted-v1.6/adversarial-checkpoint-v1.8.json — final machine-readable gate

Reproduce the current standing

whatcommand
testspython3 -m unittest discover -s tests
local 112-scene auditpython3 -m relational_auditor.run --per-case 4
import legacy real materialpython3 -m relational_auditor.real_manifest --import-legacy-root "/path/to/Robotics-ER 1.6" --manifest data/real-v0.3/legacy-manifest.jsonl
legacy selection auditpython3 -m relational_auditor.legacy_audit
registered photo baselinepython3 -m relational_auditor.photo_dry_run
cached photo Gemini auditpython3 -m relational_auditor.photo_gemini --env-file /path/to/.env
same-state repeat calibration auditpython3 -m relational_auditor.photo_calibration
independent calibration groundingpython3 -m relational_auditor.calibration_grounding --env-file /path/to/.env
blind annotation auditpython3 -m relational_auditor.annotation_audit
preliminary ruler scalepython3 -m relational_auditor.ruler_scale_audit
adversarial checkpointpython3 -m relational_auditor.adversarial_checkpoint
conventional ROI trackerpython3 -m relational_auditor.conventional_tracking
ordinary flow/change controlspython3 -m relational_auditor.ordinary_motion
full-frame motion clusteringpython3 -m relational_auditor.full_frame_motion
unlisted capture templatepython3 -m relational_auditor.unlisted_protocol
ordinary vs ISR-derived change fieldspython3 -m relational_auditor.isr_change_fields
prospective unlisted evaluationpython3 -m relational_auditor.unlisted_evaluate
post-hoc favorable-result auditpython3 -m relational_auditor.unlisted_adversarial_audit
blind v1.6 annotation adjudicationpython3 -m relational_auditor.unlisted_annotation_audit
verified-polygon localizationpython3 -m relational_auditor.unlisted_localization_audit
frozen-arm runtimepython3 -m relational_auditor.unlisted_runtime_audit
restored full-frame SIFT armpython3 -m relational_auditor.unlisted_full_frame_evaluate
threshold sensitivitypython3 -m relational_auditor.unlisted_sensitivity_audit
authoritative v1.8 checkpointpython3 -m relational_auditor.unlisted_adversarial_checkpoint
validate one model-authored specificationpython3 -m relational_auditor.model_authored_monitor /path/to/spec.json
preflight frozen authoring packetpython3 -m relational_auditor.model_authored_monitor_author --dry-run
sealed BEFORE-only authoring stagepython3 -m relational_auditor.model_authored_monitor_author --env-file /path/to/.env
sealed direct outcome armpython3 -m relational_auditor.model_authored_monitor_outcomes --env-file /path/to/.env
strict v2.0 compiler auditpython3 -m relational_auditor.model_authored_monitor_compile_audit
validate v2.1 prospective packetpython3 -m relational_auditor.spec_execution_protocol
verify v2.1 pre-capture hashespython3 -m relational_auditor.verify_spec_execution_freeze
admit completed v2.1 capturepython3 -m relational_auditor.spec_execution_admission --image-root /path/to/capture
verify opaque v2.1 extraction packetpython3 -m relational_auditor.spec_execution_extraction_packet --verify
finalize sealed v2.1 armspython3 -m relational_auditor.spec_execution_finalize
replay disclosed v2.2 development repairpython3 -m relational_auditor.spec_execution_v22_development --sensitivity
verify v2.2 method freezepython3 -m relational_auditor.verify_v22_freeze
validate v2.2 neutral capture packetpython3 -m relational_auditor.spec_execution_v22_prepare --verify-packet
sealed v2.2 BEFORE-only extractionpython3 -m relational_auditor.spec_execution_v22_contract_author
sealed v2.2 direct outcomespython3 -m relational_auditor.spec_execution_v22_direct
sealed v2.2 measurement armspython3 -m relational_auditor.spec_execution_v22_measure
finalize v2.2 denominatorspython3 -m relational_auditor.spec_execution_v22_finalize
replay post-result v2.2 adversarial checkpointpython3 -m relational_auditor.spec_execution_v22_adversarial
audit v2.3 qualification capturepython3 -m relational_auditor.v23_qualification_audit
audit v2.3.2 qualification capturepython3 -m relational_auditor.v23_qualification_audit --image-root images_rp/v2.3.2 --output-root artifacts/v2.3.2-qualification
audit contract interoperability v2.4python3 -m relational_auditor.contract_interop_v24
verify frozen DROID-100 packetpython3 -m relational_auditor.droid100_adapter --verify-root artifacts/droid100-feasibility-v2.5
preflight frozen DROID-100 groundingpython3 -m relational_auditor.droid100_grounding --dry-run
replay DROID-100 temporal grounding cachepython3 -m relational_auditor.droid100_temporal_grounding
replay DROID-100 wrist grounding cachepython3 -m relational_auditor.droid100_wrist_grounding
integrate DROID-100 surfacespython3 -m relational_auditor.droid100_multisurface_integrate
rebuild this bookpython3 assemble_lab_book.py

11 Reading order

  1. RESULTS_SPEC_EXECUTION_V2.2.md — current sealed result: replicated direct false passes, safe cascade, LIGHT abstention, and target-alias limit.
  2. ADVERSARIAL_CHECKPOINT_V2.2_POSTRESULT.md — authoritative post-result replay and rejection of overfit LIGHT/ISR fixes.
  3. V2.3_ILLUMINATION_VALIDITY_QUALIFICATION.md — exact next capture for qualifying, but not confirming, an illumination resolver.
  4. RESULTS_V2.3_QUALIFICATION_ATTEMPT_1.md — first capture's registration asymmetry, verified smooth motion, and refined apparatus.
  5. PROSPECTIVE_REPAIR_PROTOCOL_V2.2.md — the frozen new-layout experiment that preceded v2.2 capture.
  6. ADVERSARIAL_CHECKPOINT_V2.2.md — the pre-capture attack on the favorable repair result and the corrections it forced.
  7. DEVELOPMENT_REPAIR_V2.2.md — post-hoc repair mechanism, exact replay evidence, sensitivity, and limits.
  8. RESULTS_SPEC_EXECUTION_V2.1.md — preceding prospective result, controls, falsifiers, transport incidents, and repair trigger.
  9. PROSPECTIVE_SPEC_EXECUTION_PROTOCOL_V2.1.md — frozen paired-contract design that preceded v2.1 capture and API outcomes.
  10. RESEARCH_AGENDA_V1.9.md — literature positioning, evidence separation, and research gates entering v2.1.
  11. RESULTS_MODEL_AUTHORED_MONITOR_V2.0.md — retrospective specification–execution split, direct errors, and permission audit.
  12. MODEL_AUTHORED_MONITOR_PILOT_V2.0.md — frozen existing-data contract-authority pilot and restricted DSL.
  13. RESEARCH_AGENDA_V2.4.md — current phase-two position and exact evidence requirements.
  14. RESULTS_CONTRACT_INTEROP_V2.4.md — current semantic-IR interoperability result.
  15. EXTERNAL_VLA_DATA_SELECTION_V2.5.md — current external-data choice, exact DROID-100 pilot, and source-label versus downstream-policy boundary.
  16. DROID100_ADAPTER_PROTOCOL_V2.5.md — frozen development stratum, actor authority, one-sided audit semantics, and grounding admission gate.
  17. RESULTS_DROID100_GROUNDING_V2.5.md — fixed-role grounding feasibility result and temporal-coverage next gate.
  18. RESULTS_DROID100_MULTISURFACE_V2.6.md — current real-trajectory multi-surface coverage result and next prospective hinge.
  19. EXTERNAL_SCENE_DATA_SELECTION_V2.7.md — current external scene-facing choice and bounded acquisition plan.
  20. RESULTS_LANGUAGE_TABLE_SURFACE_V2.7.md — acquired external surface, registration, visibility coverage, and next adjudication gate.
  21. ADVERSARIAL_LOOKBACK_V1.8.md — authoritative post-unsealing corrections and final gate.
  22. RESULTS_UNLISTED_V1.6.md — prospective unlisted result, blind adjudication, and closed ISR gate.
  23. RESULTS_ISR_CHANGE_FIELDS_V1.7.md — negative ISR-derived result and ordinary Lab control.
  24. ISR_COMPARISON_PREREGISTRATION_V1.7.md — exact methods, exclusions, and implementation correction.
  25. UNLISTED_COLLATERAL_PROTOCOL_V1.6.md — frozen capture and final admission test.
  26. RESULTS_FULL_FRAME_MOTION_V1.5.md — recognition-light box-withholding residual.
  27. RESULTS_ORDINARY_MOTION_V1.4.md — optical-flow replication and local-change failure analysis.
  28. RESULTS_CONVENTIONAL_TRACKING_V1.3.md — ordinary feature-tracking baseline and raised ISR gate.
  29. ADVERSARIAL_CHECKPOINT_V1.2.md — current standing before the conventional baseline, material corrections, and gate.
  30. AUDIT_CORRECTIONS_V0.3.md — earlier synthetic and legacy correction ledger.
  31. RESULTS_PILOT_V0.2.md — historical mechanism result with post-audit qualifiers.
  32. RESULTS_LEGACY_REAL_V0.3.md — what the existing real images do and do not establish.
  33. RESULTS_PHOTO_DRY_RUN_V0.4.md — first registered multi-object photo mechanism check.
  34. RESULTS_PROSPECTIVE_CONFIRMATION_V0.9.md — frozen-role, camera-gated confirmation on photo-v0.4/v0.5.
  35. RESULTS_CLAUDE_ANNOTATION_AUDIT_V1.0.md — blind Claude localization comparison and tolerance correction.
  36. RESULTS_PHOTO_REPLICATION_V0.5.md — changed-object replication and cross-version comparison.
  37. RESULTS_REPEAT_CALIBRATION_V0.6.md — retrospective same-state noise and threshold audit.
  38. RESULTS_INDEPENDENT_CALIBRATION_V0.8.md — independent-layout grounding calibration and camera-regime split.
  39. CALIBRATION_CAPTURE_PROTOCOL_V0.7.md — next independent capture and threshold-freeze protocol.
  40. REAL_FRAME_PROTOCOL_V0.3.md — the next genuinely new experiment.
  41. PREREGISTRATION.md — the question as posed before the v0.1 run.
  42. BOOKMARK_ISR_ROBOTICS_BRIDGE.md — source-project cleanup held outside the experiment.