Parallax Markets — Lab Book

Stage: Markets (transfer test) · Russell Parrish / Parallax Metrology · ORCID 0009-0008-9781-7995
Generated 2026-08-20 by build_labbook.py from cached data. Reproducible.
FINAL — the out-of-sample confirmation ran; the result is LAYERED (an adversarial audit corrected a hasty "non-replication"). On strictly-pre-2016 untouched data (~25–31 monthly cross-sections, ~3,300 names each — window capped in time by Sharadar 13F beginning 2013, but full-breadth in the cross-section). Base dispersion is OUT-OF-SAMPLE SUPPORTIVE: total dispersion D is priced on untouched data and survives standard factors (NW(3) t ≈ −2.0 to −2.8, same sign — the iid −3.01 is reconciled below) — and characterization shows it is robust: stable across months (leave-one-month-out t ∈ [−3.1,−2.6]), distributed (leave-one-read-out t ∈ [−2.3,−3.1]; no single read necessary), monotonic (D quintiles Q1 +0.6% → Q5 −1.3%, rank-IC t=−5.6), and composed 51% of within-fundamentals disagreement (price-conflicts a ~16% minority). The novel pre-registered claim FAILS: the conflicts do NOT add beyond D (PF t=-0.65, PO t=-0.66, joint p=0.45; frozen §10 = PARTIAL). Note D does not literally contain PF/PO (PF=(P−F̄)² is ~16% of D, not a term in it), so the vanishing was a genuine finding, not a tautology. Verified + reconciled: independent reimplementation reproduces the frozen primary exactly (PF −0.65, PO −0.66, joint p=0.4512) AND every D number (D-alone −6.19, D+factors −2.74, rank-IC −5.61). The earlier −3.01 was a plain-iid t; under the pre-registered Newey–West(3) inference the identical audit series is −2.27, so base-D with controls is −2.0 to −2.8 (−1.97 on the strict 25-month frozen panel) — real and negative, modestly weaker than the iid number implied, never dependent on any leak. Geometry is stable across eras: discovery (2016+) vs OOS (2014-15) family-pair dispersion shares match within 1.6pp (fund-fund ~50%, fund-own ~17%, fund-price ~13% in both), family correlations and D rank-persistence (~+0.89) near-identical. The smallest defensible reading: same measured geometry, different predictive attribution — the OOS failure of the novel claim is not the instrument reconstructing a different cross-source geometry across eras, only the finer PF/PO structure's return association not replicating (no ontological "state changed" claim). Note on what D is: ~50% of D is fundamentals disagreeing with other fundamentals (price↔fundamentals ~14%), so the object reads less as "market belief disagreement" and more as internal incoherence among heterogeneous descriptions of the same firm — multi-observer state inconsistency, which fits Parallax better than an investor-belief metaphor. The anomaly-hunt is concluded; a methodological thread — the aggregate-hider — reopened 2026-08-20 (see the ROUND 2 box and §13). Positioning: dispersion-is-priced is a known effect (analyst dispersion; RFS 2026 Machine Forecast Disagreement), so this is "a deterministic multi-source instrument reconstructing known consequential structure on untouched data" — real, non-trivial, NOT novel; the novel structural refinement is what did not survive. The frozen primary verdict stands; the earlier "everything collapsed" was an over-deflation the audit caught, and "confirmed" is downgraded to "OOS-supportive" (the D+factors spec was surfaced after the primary, not pre-frozen). In-sample record below = the discovery process.
ROUND 2 — the aggregate-hider turn (2026-08-20; OPEN). The anomaly hunt concluded; a methodological thread reopened, from a VTL premise: an aggregate scalar can hide the mechanism that produces it (same centered mass — different island topology; same total dispersion — different read-space geometry). D is exactly such a scalar. (1) Reconciliation. No code discrepancy — one independent estimator reproduces every reference number (frozen PF −0.65 / PO −0.66; D-alone −6.19; D+factors −2.74; rank-IC −5.61). The old −3.01 for D+factors was a plain-iid t; the pre-registered Newey–West(3) inference on the identical series is −2.27, so base-D with controls is −2.0 to −2.8 (−1.97 on the strict 25-mo frozen panel). The 25/31/47/59-month spread is mechanical (13F gates PF/PO to 25 mo; momentum-history gates D+factors to 47; D-alone spans 59). (2) Geometry stable across eras. Discovery (2016+) vs OOS (2014-15): family-pair dispersion shares match within 1.6pp (fund-fund ~50%, fund-own ~17%, fund-price ~13% both), family correlations and D rank-persistence (~+0.89) near-identical → same measured geometry, different predictive attribution. (3) Shape decomposition (descriptive, no returns). At ~fixed D, concentration C₁ spans ~0.26–0.78 across the central 90% of mid-D deciles — equal aggregate disagreement genuinely conceals radically different internal geometries (twin at D≈0.30: ATKR C₁=.86, one SUE rupture, vs MLI C₁=.17, distributed ± split). So D is LOSSY about mechanism — the "D is a complete descriptor" outcome is ruled out. Concentration and polarity collapse to one axis (corr −0.86). (4) Confound audit. Mid-range concentration is distributed across reads (good), but the extreme tail C₁>0.8 is 43% insider / 15% sales-growth — sparse-read spikes, a confound to guard, not a signal. (5) Consequence test — RAN ONCE (2026-08-20; frozen PREREG_SHAPE.md, sha 9cd845b7…). PRIMARY (all-10 homogeneous observer system, 26 mo): D t=−2.44 (dispersion still priced negative), C₁_resid t=+3.00, two-sided p<0.05 → REJECTS the shape-null; robust across ≥6-read (t=+4.32) and 8-read (t=+3.58) constructions, interaction null. So the aggregate D is a CONSEQUENTIAL hider, not merely a lossy one. SIGN (C₁_resid positive): at equal aggregate D, DISTRIBUTED (multi-observer) disagreement carries MORE of the negative return; concentrated single-read ruptures are less-negatively priced — the most negatively-priced state is high-D + distributed = broad multi-source incoherence (consistent with — not proof of — the internal-incoherence reading; the sign is the opposite of the prior we correctly declined to freeze; it associates a state with returns, it does not establish the economic mechanism). This is the first NEW Parallax structural hypothesis in the markets arc to survive its frozen test (the PF/PO refinement did not). (6) Temporal stability — RAN ONCE (frozen PREREG_SHAPE_TEMPORAL.md, sha 43da5237…). C₁_resid is positive across the WHOLE chronology with no material sign reversals (one 36-mo rolling window trivially negative, t=−0.05 ≈ 0) and no single year necessary (leave-one-year-out t=+1.6..+2.8); it reproduces significantly in the later un-examined era (later-era all-10, 2018–2026, t=+2.12; ≥6 t=+2.77) at ~half the pre-2016 magnitude, near-zero through 2018-2021 → temporally supportive, regime-varying (two non-overlapping strong periods: pre-2016 t=3.00 and 2021-23 t=3.06 — not a one-period accident; not a constant law; the thin-window +3.00 was magnitude-inflated). Plain-language finding: firms with the same total cross-characteristic incoherence can reach it via one extreme signal or many jointly-conflicting signals, and those states are not return-equivalent — more distributed incoherence → lower subsequent returns — a relation that recurs through time though its strength varies by regime. STATUS: candidate finance contribution — meaningful in-domain replication, unresolved novelty, NO pristine prospective confirmation (retrospective / secondary-historical-validation). Novelty (quarantined): own adversarial literature check (§13.7) — general "shape beyond amount" is KNOWN (closest: belief-polarization, Hardouvelis-Karalas-Vayanos, whose kurtosis is confirmed from the source to be over the analyst-forecast distribution, not cross-characteristic; Stambaugh-Yuan = mean/consensus, the opposite object; MFD = amount only), but NO direct equivalent of "concentration of heterogeneous cross-characteristic dispersion conditional on total dispersion" located — plausible but unresolved, not claimed. UNKNOWN (recorded, not an invitation to search): why the C₁* consequence varies materially through time while keeping its sign. (7) Written up + stress-tested (2026-08-20; §13.8). Manuscript drafted (finance-native, zero Parallax). Referee-response CLOSED short-term reversal (survives; C₁⊥STR, corr≈0), but the value-weighted / liquid check plus an adversarial self-audit revealed the return finding is a small-and-illiquid-stock phenomenon: the FM slope itself is insignificant in liquid names (later-era C₁* t=+0.57 vs +2.12 full), robust only in small caps (t 2.1–3.5). A real narrowing from the "general cross-sectional factor" the draft implied; the earlier "association survives, only magnitude is small-cap" split was an over-claim I caught and corrected. Points toward a limits-to-arbitrage mechanism; the measurement contribution is unaffected. True confirmation still needs a genuinely untouched block or the cross-domain transfer with this exact frozen spec. See §13.
In-sample discovery (2016–2026, exploratory — did not confirm). On the survivorship-free universe (7,724 names, 72% delisted, 2016–2026), multi-family Ω is priced; the principal magnitude-vs-disagreement question is now answered by the exact geometric decomposition of read-vector magnitude (‖x‖² = x̄² + D). Entered jointly: dispersion (D = Var = Ω²) carries the pricing (t=−4.6), common-mode magnitude (x̄²) does NOT (t=+0.8), consensus direction is mildly momentum-like (+2.8). This reverses the earlier "it's extremeness" reading, which controlled for MSQ = x̄²+D — a quantity that contains the dispersion, so it absorbed Ω near-tautologically. And family-pair attribution appeared to sharpen it in-sample: price↔fundamentals (−4.3) and price↔ownership (−5.2) seemed to add beyond total D. [SUPERSEDED — the frozen OOS primary specifically REJECTED this beyond-D refinement: out of sample PF/PO add nothing beyond D (PF −0.65, PO −0.66, joint p=0.45). "Structured cross-family conflict priced beyond D" is a discovery-era interpretation that did NOT replicate. What survived OOS is total dispersion D itself.] Literature-located, not novel: cross-observer dispersion is already priced (analyst dispersion; RFS 2026 Machine Forecast Disagreement); the only non-redundant element is heterogeneous data-family observers. Claim status: structured-disagreement-as-priced — SUPPORTED, not confirmed (robustness re-runs on the clean object pending); investor-belief disagreement vs characteristic inconsistency — NOT distinguished (needs external validation vs analyst dispersion / short interest, data-gated). Trade: a factor-neutral construction earns a modest size-tilted ~+6%/yr gross (t=3.9), ~⅔ a small-cap size-interaction (within-size t=1.5), un-costed — measurement evidence, not a strategy.
This was a decomposition, not a series of kills. One noisy object got progressively identified, each step narrowing the explanation — and one step was itself a flawed attack, corrected by a better one: generic disagreement → (price-only) idio-vol → (multi-family, standard zoo) distinct-looking → (MSQ control) "largely extremeness" → (exact geometric decomposition) the MSQ control was collinear with Ω by construction; the priced component is the dispersion (t=−4.6), not the common-mode magnitude (t=+0.8). Withdrawn as over-reaching labels: "validated," "distinct coordinate," "mixture," reflexivity, "sample-bound," naive-sort "inert," and "it's extremeness not disagreement." Current favored reading: dispersion (disagreement) is priced beyond consensus and the factors — supported, not confirmed, robustness re-runs pending on the clean object. Reporting rule adopted: status ledger; "settled/killed/final" reserved for a pre-registered terminal criterion. The lesson cut both ways — the fix for an over-deflation was not more skepticism but the right geometry (magnitude = consensus² + dispersion).

1. Path B (PRICE-ONLY) — CLOSED: rediscovery of idio-vol

Scope: this section is the FAILED first construction — five reads, all transforms of the price surface, violating the instrument's independence requirement. Its verdict is price-only. The live result (multi-family Ω) is §2 onward; do not read §1 as the global verdict.

Fig 1. Price-only Path B on the survivorship-free universe. The disagreement of five price transforms clears the cross-sectional regression by a mile — but the reads are not independent, so the question is what it actually measures.
Fig 2. The decider (price-only). The raw long-short is huge (t=8.6), but its ALPHA after the idio-vol/size/value/momentum factors is zero (t=-0.21), idio-vol β=+0.67. The price-only signal is the idiosyncratic-volatility trade in disguise.
Fig 3. Price-only factor-zoo: a regression residual survives, but Fig 2 shows it carries no tradeable alpha.

Price-only verdict (CLOSED): the five-price-read construction spans to the idiosyncratic-volatility / low-risk anomaly (zero alpha) — a reconstruction of a known effect. This is a real but known effect, and it is NOT the document's finding; it is the failure that motivated the independence correction in §2.

Fig 4. Robustness (price-only): strengthens in liquid names, so the reconstruction is real, not penny-stock distress — the price-only signal is simply the idio-vol anomaly, cleanly measured.
Controls addedDISAGREEMENT t
none-10.56
idio-vol-7.84
+ total-vol-7.32
+ momentum + reversal-6.07
+ size + beta-5.87

2. The dig — disagreement across independent families

The price-only reconstruction failed because its five reads were transforms of one surface (price), violating the instrument's independence requirement. The dig rebuilt disagreement across independent data families with low observed correlation (mean |corr| 0.12 — non-redundancy, not proven informational independence) — price, fundamentals (gross profitability, ROA, earnings momentum, sales growth, accruals, asset growth, value), institutional 13F flow, and insiders. Gross profitability alone survives real controls (FM t=4.77); "everything is idio-vol" is false.

Fig 5. The dose-response. As independent rooms grow 4→10, the disagreement characteristic sharpens (FM |t| 2.1→5.7, green). The red line is the naive quintile-sort alpha (flat) — later shown to be a coarse-sort artifact: a factor-neutral construction recovers ~6%/yr (§8). Caveat: ~60% of the characteristic is extremeness (energy control), so the green climb overstates genuine disagreement.

3. Adversary — attacking the load-bearing claim

The dose-response was, at first, reported as a monotone sequence and believed, with no adversary — a violation of this program's own audit standard (a build fails if the load-bearing claim never survives an attack). Four attacks (adversary.py):

Fig 6. The decisive placebo. With the read→return link shuffled out, the climb vanishes (|t| 0.05→0.28, flat) — so the real climb (2.1→5.7) is genuine signal, not an artifact of adding dimensions to shrink standard errors.
AttackResultReading
Sub-period4→10 climb present in both halves (0.9→2.2; 3.5→5.7)climb replicates; magnitude back-half-heavy
Placebo (shuffle)|t| flat 0.05→0.28not a dimension/SE artifact
Quality in controlsstrengthens −5.74→−6.19not merely the profitability factor
Nonlinear idio-volsurvives (iv+iv²=−5.91, iv+rank=−6.17); but OMEGA–idio-vol corr rises 0.16→0.29MIXTURE: partly idio-vol recomposition, partly orthogonal residual

What the adversary changed: the dose-response is real (survived every attack). The first read of the last test — "idio-vol correlation rises with rooms → mixture, not a coordinate" — was itself over-deflated. Measuring the two spaces separately:

R²(Ω ~ idio-vol)R²(Ω ~ iv+size+value+mom)
Characteristic space (pooled cross-section)0.079 — 92% orthogonal0.182 — 82% orthogonal
Return space (the quintile long-short trade)0.5690.841 — 84% factor-explained

Interim reading (superseded below): at this stage the characteristic looked largely orthogonal to the standard tested factors, and the selected-room dose-response was real relative to its placebo — so it appeared a coordinate candidate. The MSQ and geometry tests below change that interpretation. "Validated," reflexivity, and the "mixture/not-a-coordinate" over-correction were withdrawn here.

The energy-inclusive completion — disagreement vs extremeness

The "distinct at anomaly-grade strength" reading survived a complete accepted model (each characteristic vs all others, calibration twins) — Ω at t=3.7, mid-pack. But that model lacks a multi-signal-extremeness term (MSQ = mean-squared read level). Var = E[x²]−E[x]², so dispersion hides an energy component. Adding MSQ:

characteristiccomplete model+ MSQ (energy)effect
Ω (disagreement)-3.74-1.87collapses
gross profitability+2.24+3.84strengthens
momentum+5.32+6.40strengthens
accruals+5.21+6.33strengthens
ROA+7.56+7.62unchanged
idio-vol-7.88-7.78unchanged

MSQ crushes Ω specifically (t -3.7→-1.9, below significance) while every accepted twin is unchanged or STRENGTHENS. So MSQ is a legitimate priced control, not an over-fit, and Ω's distinctness from known anomalies was largely multivariate extremeness. Algebraic caveat retained (Ω and MSQ mechanically entangled), but the twins-strengthen asymmetry is the real evidence.

Does the dose-response survive extremeness?

roomsraw Ω |t|Ω residual after MSQ |t|
45.20.0
68.53.1
87.40.6
105.52.2

Raw climb is non-monotone over nested subsets (the clean 2.1→5.7 was partly room-selection-specific); the MSQ-residual climb is erratic and mostly weak. So the "more rooms → sharper disagreement" story is not robust as stated.

Exact geometric decomposition — the MSQ deflation was a collinear-control artifact

MSQ = x̄² + D, so the dispersion is inside MSQ: controlling for MSQ absorbs Ω near-tautologically. The clean test uses the exact geometric decomposition of read-vector magnitude: x = x̄·1 + (x−x̄·1), where the common-mode and centered-residual vectors are orthogonal, so ‖x‖² = x̄² + D exactly. Entering the scalar components jointly (they are not themselves independent — x̄² is a transform of x̄) and asking which is priced:

read-vector componentFM |t| (+ iv,size,value,mom)
consensus direction (x̄)+2.82
consensus magnitude (x̄²)+0.81 — null
dispersion (D = Ω²)−4.56 — significant
(radial R² = x̄²+D, alone)−5.7

The priced component is the dispersion (disagreement), not the common-mode magnitude. So "it's extremeness, not disagreement" was substantially an artifact of a control collinear with Ω by construction. Corrected reading: dispersion of the standardized reads is priced beyond consensus and the standard factors. Methodological lesson: don't decompose dispersion with a control that contains it; use the magnitude = consensus² + dispersion split. (This is the correct form of the cross-domain energy test, too.)

Family-pair attribution — structure matters beyond magnitude

D = (1/2n²)Σ(xᵢ−xⱼ)², so dispersion is an aggregate of pairwise disagreements. Grouping by family (price / fundamentals / ownership-13F / insider):

family conflictalone (+factors)beyond total D
price ↔ ownership-5.15-4.55
price ↔ fundamentals-4.31-3.13
fund ↔ ownership-3.17-1.36
price ↔ insider-1.51
fund ↔ insider-1.74
own ↔ insider-0.71

Discovery-era interpretation, subsequently NOT confirmed OOS: in-sample, price↔fundamentals and price↔ownership appeared to add beyond total dispersion D. The frozen out-of-sample primary rejected this — PF/PO add nothing beyond D on untouched pre-2016 data (PF −0.65, PO −0.66, joint p=0.45). So the paragraph that follows describes the discovery hypothesis, not the confirmed conclusion; what replicated OOS is total dispersion D, not the finer cross-family structure. Caveat: positive "flips" for insider pairs conditional on D are suppression effects, not "insider disagreement predicts positively." Literature location: cross-observer dispersion is already priced (analyst forecast dispersion; RFS 2026 "Machine Forecast Disagreement," ML models as observers) — so "dispersion is priced" is not a discovery. The only non-redundant element: observers are heterogeneous data families, and the structure is price-vs-fundamental/ownership. Status: supported as structured cross-family conflict; NOT supported as investor belief disagreement (vs characteristic inconsistency) or as novel. Robustness re-runs and external validation (vs analyst dispersion / short interest, data-gated) pending.

4. What closed, and what it cost to get here

Fig 4. The time-series arena, closed by two powered nulls (v2 vol, v3 drawdown): geometry — the published arsenal (absorption ratio, financial turbulence, real persistent homology) and ours — never beats volatility persistence.

Path C (arrival ordering / the seismic θ_mv coordinate) reached t=1.97 on 230 large-cap survivors, then went to a powered null (reversal gate t=+4.34; SIGNAL t=+0.38) on the real universe. A clean death: thin-data near-misses do not survive breadth. That is the discipline working, and it is why the Path B pass is credible.

5. Chronological trail

origin (variance-of-disagreement latent in the quiver) → v1 rare-event (underpowered) → v2 vol / v3 drawdown (powered nulls, arena closed) → net research (path is the cross-section; dispersion is priced) → Path B/C on survivors (underpowered; C flickered at 1.97) → seismic timing probe (σ-moveout ⊥ vol, but payoff is cross-sectional) → Sharadar survivorship-free data → Path C null, Path B PASS + robustness + factor-zoo. Full detail in RESULTS_AND_LEARNINGS.md.

6. Scoreboard

TestArenaResultPower
v1time-series rare eventfailunderpowered
v2time-series volnullpowered
v3time-series drawdownnullpowered
Path B (survivors)cross-section disagreementnull (correct sign)failed
Path C (survivors)cross-section arrival orderfail t=1.97passed
Path C (real data)cross-section arrival orderpowered nullpassed
Path B — price-onlycross-section, 5 price readsspans to idio-vol; no alphaCLOSED (rediscovery)
Multi-family Ω — characteristic10 reads, independent familiespriced; survives standard zoo t=3.7 (mid-pack vs anomalies)priced characteristic
Q1 — + extremeness (MSQ)add multi-signal energy termΩ 3.7→1.9 — but MSQ = x̄²+D is collinear with Ω (tautological)reading superseded
Geometry decompositionconsensus / consensus-mag / dispersiondispersion (D=Ω²) t=−4.6 priced; common-mode magnitude t=+0.8 nulldisagreement is the priced component (supported)
Observer dose-response4→10 roomsraw non-monotone; residual-after-MSQ erraticnot robust as sharpening
Trade — naive quintile sortfactor spanning84% factor exposure; no alphaartifact of coarse sort
Trade — factor-neutralresidualize then sort+6%/yr, t=3.9, Sharpe 1.3; but ⅔ small-cap, un-costedmodest size-tilted gross edge

7. Instrument & corpus

General instrument: σ(Ω) / variance of disagreement among independent reads — pathology → seismology → biomechanical → markets. Failed cross-sectional implementation: dispersion across five price transforms (momentum horizons, reversal, trend) — insufficient observer independence; rediscovered idio-vol (§1). Current implementation (the operative object): ten reads drawn across independent data families with low observed correlation (mean |corr| 0.12) — price, fundamentals (gross profitability, ROA, earnings momentum, sales growth, accruals, asset growth, value), institutional 13F, insiders. "Independent" here means independent data families, not proven informational independence — accounting reads share a corporate state; the honest hierarchy is source → family → statistical non-redundancy → observer independence, and only the middle two are demonstrated. Real-data corpus: Sharadar SEP 10-year bulk, 7,724 domestic common stocks on major exchanges, 72% delisted (survivorship-free), median ~4,242 names/month, 2016–2026. Incumbents reconstructed: absorption ratio, financial turbulence, persistent homology (ripser), HAR, short-term reversal, idiosyncratic vol.

8. Interpretation boundaries

9. Forward queue

#Open question (in markets — no pivot)What would settle it
1 ★ Robustness re-runRe-run the attacks on the CLEAN objectPlacebo / sub-period / cost were run on the old (Ω-vs-MSQ) framing. Re-run them on dispersion-controlling-for-consensus-magnitude (D | x̄²). This is what turns 'supported' into confirmed-or-not for the disagreement reading.
2 · Economic identityWhat is the size×extremeness/dispersion interaction the factor-neutral trade exploits?Decompose R ~ radial + size + radial×size + D⊥radial (+ interactions). Then a crude cost pass (turnover, spread proxy, exclude smallest decile) — before it leaves the measurement bucket.
3 · Cross-domainDo other Parallax domains' disagreement results survive the SAME geometry test?Guardrail: control for the radial magnitude of the same normalized read vector, not an arbitrary energy proxy. (Pathology is Colab/unorganized; start with the organized-Python domains.) Tests the instrument's identity program-wide.
4 · Family attributionWhich information families make the dispersion measurable?Drop-one-family; does price / fundamentals / 13F / insiders dominate, or is it distributed? The surviving form of the observer-architecture idea.
✓ Confirmation + audit — DONE (2026-08-19)Pre-registered OOS test on strictly-pre-2016, then adversarial auditLAYERED. Base dispersion D REPLICATES OOS (+factors t=−3.01 iid → −2.27 under pre-registered NW(3); see the 2026-08-20 reconciliation); raw conflicts replicate (PF −3.7, PO −4.4). But the frozen NOVEL claim (conflicts beyond D) FAILS (PF −1.15, PO −1.04). Panel not thin (~3,300 names/mo); limit is #months. Frozen primary verdict stands; audit corrected a hasty 'total non-replication.' PREREGISTRATION.md 0299723b…, harness f9d9a08b…, audit in CONFIRMATION_RESULT.txt.
✓ Reconciliation + geometry — DONE (2026-08-20)Explain −2.74 vs −3.01 and the 25/31/47/59-mo spread; then compare disagreement geometry across erasReconciled: no code discrepancy — same estimator reproduces all references. −3.01 was a plain-iid t; pre-registered NW(3) on the identical series = −2.27. Month spread is mechanical (13F gates PF/PO to 25mo, momentum-history gates D+factors to 47mo, D-alone spans 59mo). Base-D with controls is −2.0 to −2.8. Geometry stable: discovery vs OOS family-pair shares match within 1.6pp; D rank-persistence ~+0.89 both eras → same measured state, only the return-consequence was sample-specific. reconcile_D.py, geometry_compare.py.
✓ Aggregate-hider / shape decomposition — DESCRIPTIVE DONE (2026-08-20)Does equal aggregate D conceal distinct internal read-geometries?YES. At ~fixed D, concentration C₁ spans ≈0.26–0.78 (mid-D deciles); geometry-only twins show same-D/opposite-shape firms. D is lossy about mechanism — 'D is complete' ruled out. C₁ & polarity are one axis (−0.86). Confound: extreme tail C₁>0.8 is sparse-read (insider/salesg) — guard, don't chase. See §13.3–13.4. shape_descriptive.py, dominant_read_audit.py.
✓ Consequence of shape — RAN ONCE (2026-08-20)Does the market care about the mechanism D hides?YES (secondary-historical). Frozen PREREG_SHAPE.md (sha 9cd845b7…), all-10 primary: C₁_resid t=+3.00, p<0.05, REJECTS shape-null; robust ≥6-read (+4.32) & 8-read (+3.58); interaction null. Sign: distributed multi-observer incoherence carries the negative return, not single-read ruptures. FIRST novel structural claim in the arc to survive a frozen test. Status: secondary-historical-validation, not pristine confirmation; C₁_resid literature-location still open. See §13.5.
✓ Shape temporal stability — RAN ONCE (2026-08-20)Is +3.00 persistent or one-period?Temporally supportive, regime-varying. Positive across the whole chronology, no material sign reversals (one rolling window t=−0.05≈0), no single year necessary; reproduces in the later un-examined era (later-era all-10 2018-2026 t=+2.12, ≥6 t=+2.77) at ~half the magnitude; near-zero through 2018-2021. Two non-overlapping strong periods. NOT a one-period accident, NOT a constant law. Frozen PREREG_SHAPE_TEMPORAL.md (sha 43da5237…). §13.7.
✓ Literature-location — DONE (own adversarial web check)Is the shape result novel to finance?General 'shape beyond amount' KNOWN (closest: belief-polarization, Hardouvelis-Karalas-Vayanos). Both closest papers now FULL-TEXT confirmed from source PDFs: Hardouvelis kurtosis is over the analyst-forecast distribution (≈ our polarity, not concentration); MFD is SD of cross-model forecasts (amount only, no shape/concentration test in 54pp). Stambaugh-Yuan = mean/consensus (opposite object); Huang-Li-Wang = aggregation for market-timing. NO direct equivalent of concentration-of-heterogeneous-cross-characteristic-dispersion-conditional-on-amount located → novelty PLAUSIBLE BUT UNRESOLVED (a search cannot prove the negative further). §13.7.
✓ Paper + referee-response + adversarial — DONE (2026-08-20)Does the return finding survive reversal and hold in liquid names?Survives short-term reversal (C₁⊥STR). But VALUE-WEIGHT + liquid-subset FM show it is a small/illiquid phenomenon: slope insignificant in liquid names (later t=+0.57 vs +2.12), robust only in small caps (t 2.1–3.5). Caught + corrected my own 'association survives, only magnitude small-cap' over-claim. §13.8; run_referee_response.py, adversary_check.py.
✓ LTA vs artifact test — RAN ONCE (2026-08-20)Real mispricing, or microstructure artifact?MIXED (pre-committed). Implementation-gap SURVIVES + strengthens (later C₁* +2.12→+3.40 skipping 5 days) → NOT a microstructure/bid-ask artifact. But IVOL gradient (t=+0.76) and above-microcap (t=+0.55) both fail → confined to microcaps, mechanism UNRESOLVED, NOT claimed as limits-to-arbitrage. Frozen PREREG_LTA.md (b9e5d34e…). §13.8.
✓ Reducibility kill (linear) — RAN ONCE (2026-08-20)Is C₁* just a linear repackaging of the ten characteristics?DEAD in the primary sample (pre-committed). Adding the ten reads as linear controls absorbs C₁*: later-era t +2.12 → −0.06; survives only in the thin pre-2016 window (+2.87). Frozen PREREG_REDUCE.md (c6ab1b1b…). §13.8.
✓ Nonlinear reducibility kill — RAN ONCE (2026-08-20) — RETURN QUESTION CLOSEDDoes C₁* survive a nonlinear read-expansion?DEAD (total, pre-committed). Controls + reads² + all 45 pairwise products: later-era C₁* t=+0.62 (coef 24% of baseline). Reducible even nonlinearly. Pre-2016 (n26, unbankable) t=+3.07 NOT pursued. The return contribution is closed; paper stands on the descriptive decomposition + measurement object. Frozen PREREG_NONLIN.md (db22bb13…). §13.8.
★ next — true confirmation of the shape claimIs the shape effect real beyond secondary-historical corroboration?Two routes, either genuinely untouched: (a) future accumulating pre-registered data; (b) cross-domain transfer of the EXACT frozen shape construction to a domain where the aggregate-hider question is not finance-known.
secondary — optionalPF-only over a longer untouched era (drop 13F, ~2008+)NON-confirmatory power diagnostic of the price↔fundamentals component alone. Modified spec; does NOT change the verdict above. Only if the PF component specifically is worth probing at higher power.
parkedtrade costing, external validation, cross-domainDownstream of a replication that did not occur. Cross-domain energy/geometry test on organized-Python domains remains the higher-leverage program question.

Corrections logged this session: "validated" withdrawn (overclaim); reflexivity withdrawn (effect intensifies, contradicting decay); "sample-bound" withdrawn as an alibi (no tradeable edge on data held, full stop). No pivot away from markets — the inquiry stays open here.

10. Status: exploratory, and the multiple-testing discount

The whole markets investigation is exploratory / hypothesis-generating, not confirmatory. ~15 decompositions were run on ONE 2016–2026 Sharadar window, each chosen in response to the previous result — the garden of forking paths. Every "supported" in this document (dispersion priced at t=−4.6, the price↔fundamentals/ownership structure, the ~6% factor-neutral gross) carries a nominal t that does not account for that search path; a referee would haircut them, correctly. Nothing here is confirmed. The only thing that can confirm any claim is a pre-registered out-of-sample test on data not yet touched — full-history tier, a different period, or the cross-domain transfer. Read every t-stat as exploratory.

11. Where this lands (uses) — a diagnostic instrument, not a factor

The defensible contribution is a characteristic-construction-and-interrogation method, plus a deterministic cross-source state variable — not a tradable factor (the ~6% gross stays in the measurement bucket: size-concentrated, un-costed). Plausible uses:

12. Field-hygiene status (a finance referee's checklist)

Addressed: point-in-time lagging (fundamentals by datekey, 13F +50d), survivorship-free universe with delisting, microcap screen (price ≥ $5) + winsorization, block/Fama–MacBeth inference, calibration twins, internal-algebra decomposition. Pending (would matter for publication, not for the exploratory claim): NYSE breakpoints, value- vs equal-weighting, transaction costs / turnover, industry-neutralization, Newey–West / persistence-aware SEs, formal multiple-testing correction, and — the big one — a longer history than 2016–2026 (regimes matter; premia are noisy). These separate an interesting lab experiment from something a journal would accept.

13. The aggregate-hider turn — reconciliation, geometry, shape decomposition (2026-08-20)

Why this section exists. After the frozen OOS primary settled (novel beyond-D claim failed; base-D OOS-supportive), a VTL-derived question reopened the work at the methodological level, not the anomaly level: the aggregate is the hider. A scalar summary can be identical across objects generated by completely different mechanisms — two paintings both "centered" in total mass, one via counterweighted islands, one via a single central blob. D is exactly such a scalar: for a 10-read vector it is a radius from consensus, and many internal configurations sit on the same radius (e.g. [-2,-2,-2,+2,+2,+2,…] bimodal vs [-4,0,0,…] single rupture). The question: does equal aggregate D conceal distinct read-space geometries, and — separately — does the market care about the mechanism it hides?

13.1 Reconciliation — the D numbers are not a code artifact

Before trusting any of this, the base-D result was reconciled against the frozen primary with an independent estimator (reconcile_D.py). Result: no implementation discrepancy. The same estimator reproduces every reference number exactly. The earlier −3.01 for D+factors was audit_confirmation.py's plain-iid Fama–MacBeth t; applying the pre-registered Newey–West(3) to the audit's own coefficient series (imported in-memory, byte-identical rows/months) gives −2.27. The −3.01-vs-−2.74 gap is inference (iid vs NW), not the effect. A forward-return boundary leak (a diagnostic building fwd on the extended cache, letting the last month peek past the outcome boundary) was caught and fixed here.

Diagnostic (all NW(3), clean pre-2016)ReadsControlsMonthsWindowt
frozen PF (full model)full 10full frozen252013-11..2015-11−0.65
frozen PO (full model)full 10full frozen252013-11..2015-11−0.66
D alonemask≥6none592011-01..2015-11−6.19
D + factorsmask≥6factors472012-01..2015-11−2.74
D rank-ICmask≥6none592011-01..2015-11−5.61
D + factors (on the frozen 25-mo panel)full 10factors252013-11..2015-11−1.97

The 25/31/47/59-month spread is mechanical: PF/PO require 13F ownership (→ 25 mo from 2013-11), D+factors requires momentum history (→ 47 mo from 2012-01), D-alone/rank-IC need only D (→ 59 mo from 2011-01). Same estimator throughout. Honest characterization of base-D with controls: −2.0 to −2.8 — real and negative, modestly weaker than the iid −3.01 implied, never dependent on any leak.

13.2 Geometry stable across eras — same measured state

Descriptive, no returns (geometry_compare.py): the disagreement geometry is nearly the same object in the discovery era (2016+) and the OOS era (2014-15).

Family pair (share of total pairwise dispersion)OOS 2014-15Discovery 2016+Δ
fundamentals ↔ fundamentals51.2%49.7%−1.6pp
fundamentals ↔ ownership16.3%17.6%+1.3pp
fundamentals ↔ price14.2%12.7%−1.5pp
fundamentals ↔ insider12.6%14.1%+1.5pp
price ↔ ownership1.9%1.6%−0.3pp

Family-centroid correlations are near-identical across eras (price↔own +0.29/+0.26, own↔insid −0.19/−0.18, fundamentals near-orthogonal to all in both); D's month-to-month rank persistence is +0.888 / +0.896. Smallest defensible reading: same measured geometry, different predictive attribution. The OOS failure of the novel claim is not the instrument reconstructing a different cross-source geometry across eras — only the finer PF/PO structure's return association failing to replicate. No ontological "state changed" claim is warranted.

13.3 Shape decomposition — the aggregate IS a hider (descriptive, no returns)

Per firm-month, inside D, two shape axes (shape_descriptive.py, discovery 2016+, all-10-reads, price≥5, 281,538 firm-months): concentration C₁ = maxᵢ dᵢ² / Σ dᵢ² (dᵢ = xᵢ − x̄; high ⇒ one read ruptures) and polarity P = 2·min(E₊,E₋)/(E₊+E₋) (0 = one-sided rupture, 1 = balanced two-sided split); participation ratio PR reported as a descriptive companion to C₁, not a second test.

Fig. Occupied shape state-space (left): C₁ vs P is a single diagonal spectrum, not two free axes (pooled corr −0.86). Residual C₁ spread within each D decile (right): the boxes stay wide — at ~fixed aggregate dispersion the internal concentration varies substantially, i.e. the aggregate is hiding shape.

Findings. (i) C₁ and polarity collapse to one axis (corr −0.86); PR is −0.93 with C₁. So there is one effective free shape dimension — concentrated/one-sided ↔ distributed/two-sided — not five. (ii) Shape is entangled with D itself (corr(D,C₁) = +0.40): higher-D states are more rupture-like on average, so a C₁ main effect co-moves with D and the honest question is shape conditional on D. (iii) Large residual shape variation at fixed D: across mid-D deciles C₁'s p5–p95 spans ≈0.26 to ≈0.78. The geometry-only twin makes it concrete — two firms at D≈0.30: ATKR C₁=0.86, P=0.28 (a single SUE rupture at −1.69, everything else ≈0) vs MLI C₁=0.17, P=0.87 (disagreement spread across many reads in a balanced ± split). Same radius, opposite internal state. Conclusion: D is lossy about mechanism — the "D is a complete descriptor of the state" outcome is ruled out at the measurement level. What remains open is whether the market cares about the hidden mechanism.

13.4 Confound audit — is "rupture" a geometric condition or one noisy read?

Among firm-months, which read owns the maximum squared deviation (dominant_read_audit.py, no returns):

C₁ bandTop dominant readsTop-3 shareReading
all firm-monthsSUE 25%, GPA 19%, value 18%, inst 10%dominance spread across ~5 reads
C₁ > 0.60 (62k)SUE 30%, GPA 15%, insid 14%, inst 12%59%SUE elevated, no majority
C₁ > 0.70 (30k)SUE 25%, insid 23%, GPA 14%, inst 12%61%insider joins
C₁ > 0.80 (10k)insid 43%, salesg 15%, GPA 11%, SUE 11%69%tail is sparse-read driven

Two-part answer. Through moderate concentration (C₁ ≈ 0.6–0.7, the bulk of "concentrated" states), the rupturing read is spread across five or six channels — concentration has the same distributed character D itself has, so it is a genuine geometric condition, not a single-channel pathology. But the extreme tail (C₁ > 0.8) is confounded: insider net-buying owns 43% and sales growth another 15%, both structurally sparse/spiky reads (intermittent filings → occasional large z-scores). So the most extreme "single rupture" states are disproportionately a sparse read firing — mechanism, not signal. The shape variable is real in the mid-range; any consequence test must be robust to the tail rather than chase it.

13.5 Consequence spec (frozen) and RESULT — ran once 2026-08-20

The consequence question (does shape conditional on amount carry return consequence?) was frozen as a minimal pre-registration (PREREG_SHAPE.md, sha256 9cd845b7…; harness run_shape_consequence.py, sha256 381f9497…) and run ONCE. Per GPT's edit before the run, the primary observer system is all-10 reads (every firm-month the same ten-dimensional instrument), with ≥6/10 demoted to Robustness A; the two roles of D are documented (raw D_raw for the residualization map; z_t(D_raw) as the consequence regressor). Frozen design:

ElementFrozen choice
Shape variableC₁* = C₁ − Ê(C₁ | D), the conditional-concentration residual; Ê(C₁|D) a smooth monotone fit frozen on discovery. (Answers 'among firms at the same D, how unusually concentrated is this one' — not raw D×C₁.)
Confound guardprimary spec winsorizes C₁ at 0.80 (caps the sparse-read tail); frozen robustness variant recomputes C₁ dropping insider & sales-growth. Both must agree in sign to count.
Modelr₍t+21₎ ~ D + C₁* + accepted-factor controls; Fama–MacBeth, NW(3); two-sided on C₁*. Interaction D×C₁* reported as secondary, not primary.
Status label (fixed in advance)retrospective / secondary-historical-validation on pre-2016 — NOT a pristine confirmation (see 13.6).
Interpretation (fixed in advance)C₁* negative & significant in BOTH variants → shape conditional on amount carries consequence (aggregate is a consequential hider). C₁* null → D is lossy about mechanism but sufficient for consequence — a real result, not a failure. Either way we learn. A C₁* null means shape-orthogonal-to-D is inert, not that concentration is irrelevant.
Result (ran once, frozen)monthsD tC₁_resid tverdict
PRIMARY (all-10 homogeneous)26−2.44+3.00REJECT shape-null (p<0.05)
Robustness A (≥6/10, per-count iso)48+4.32same sign — supportive
Robustness B (8-read: drop insid+salesg)26+3.58same sign — supportive
Secondary interaction D×C₁_resid26(C₁_resid +3.11)D×C₁_resid t=−0.29 ≈ null

Reading (within the frozen rules, no overreach). Shape conditional on amount carries detectable incremental consequence beyond aggregate D — the aggregate is a consequential hider, robust across three geometries; the interaction is null (additive, does not scale with D). Sign: C₁_resid is positive, so at equal aggregate D the more-concentrated (single-read rupture) states earn higher returns and distributed disagreement earns lower returns. With D negative, the most negatively-priced state is high-D and distributed = broad multi-source incoherence — broadly distributed cross-read inconsistency is the state associated with the stronger negative return relation (an association, not a demonstrated economic mechanism; a latent known characteristic could in principle generate both broad read-space dispersion and low returns). This is the first NEW Parallax structural hypothesis in the markets arc to survive its frozen test (the PF/PO "beyond-D" refinement failed its frozen primary; this passed). Status: retrospective / secondary-historical-validation — meaningful corroboration of a geometry-derived hypothesis under a frozen spec, stronger than in-sample discovery, but NOT pristine confirmation (pre-2016 was already opened for other purposes). One pre-registered primary, so no forking-paths discount on +3.00; but 26 months is thin (hence the ≥6 robustness at 48 mo, +4.32, and the temporal battery in §13.7). Novelty is quarantined, not claimed: a literature audit (§13.7) finds the general "shape beyond amount" principle KNOWN (closest prior: belief-polarization, Hardouvelis-Karalas-Vayanos 2025) but locates no direct equivalent of concentration-of-heterogeneous-cross-characteristic-dispersion-conditional-on-amount. True confirmation still needs a genuinely untouched time block or the cross-domain transfer with this exact frozen spec.

13.6 Governance — pre-2016 is no longer a pristine holdout

The pre-2016 window was untouched for the original PF/PO confirmation. It has since been opened for D regressions, rank-IC, geometry comparison, and sample reconciliation. It is untouched with respect to C₁-vs-return results specifically (nobody has looked at that there), but it has already shaped our thinking about D, the controls, the samples, and the instrument. So a C₁ test there is secondary historical validation, not independent confirmation. A positive there would be taken far more seriously than another in-sample discovery, but true confirmation needs a genuinely untouched block — future accumulating data, or a cross-domain transfer with the shape hypothesis frozen before consequence is examined. Open decision (awaiting input): (a) run the frozen secondary-validation pass on pre-2016 now; (b) edit the freeze first; or (c) hold the frozen spec unrun and carry it straight to an untouched domain so its first contact with returns is the pristine one. What this turn already bought: we asked whether Markets failed because aggregate D was hiding mechanism — and we now know it is hiding mechanism; what we do not yet know is whether the market cares about the mechanism it hides.

13.7 Temporal stability + literature-location (2026-08-20)

Temporal battery (frozen PREREG_SHAPE_TEMPORAL.md, sha 43da5237…; harness run_shape_temporal.py). Instrument frozen; only time segmentation varies. The all-10 instrument needs 13F so the clean battery lives in the later era (2018-2026 effectively) with pre-2016 as the comparison block; the 2016+ C₁→return relationship had not been examined before, so it is an un-inspected reproduction, not a re-use.

Non-overlapping block (all-10)monthsβ C₁_residNW t
pre-2016 (2013-11..2015-11)26+0.00228+3.00
2018-03..2018-119+0.00036+0.48
2018-12..2021-05 (COVID-spanning)30+0.00003+0.04
2021-06..2023-1130+0.00185+3.06
2023-12..2026-0530+0.00086+1.01

Leave-one-calendar-year-out (later era): all 11 folds positive, t=+1.6..+2.8 — no single year is necessary. Rolling 36-mo path (below): near-zero in the earliest 2017-2020 windows, then climbs and holds positive from ~2020 (peak t=+2.87); no material sign reversal — one window is trivially negative (t=−0.05 ≈ 0). Headline reproduction: pre-2016 all-10 t=+3.00 → pooled later-era all-10 (2018–2026) t=+2.12 (≥6 companion +4.32 → +2.77), same sign and still significant, ~half the magnitude. (The all-10 homogeneous panel begins 2018-03-29, when the ten-read system first has 13F throughout — "later-era", not a full ten years of all-10 data.)

Fig. C₁_resid coefficient, 36-month rolling windows stepped 12 months (2016+), NW t annotated. Positive and rising from ~2020; the earliest COVID-spanning windows sit near zero. Overlapping windows — a diagnostic path, not independent tests.

Verdict (against the frozen ladder): temporally supportive, regime-varying. Positive across the whole chronology with no material sign reversals (one rolling window trivially negative, t=−0.05 ≈ 0) and no single-year dependence; two non-overlapping strong periods (pre-2016 t=3.00 and 2021-23 t=3.06); reproduces significantly in the later un-examined era — so not a one-period accident. But magnitude is regime-modulated (near-zero through 2018-2021) and the thin-window +3.00 was magnitude-inflated (durable coefficient ≈ half, t≈2.1-2.8). Not a constant law, not pristine confirmation. Plain-language finding: firms with the same total cross-characteristic incoherence can reach it via one extreme signal or many jointly-conflicting signals; those states are not return-equivalent (more distributed incoherence → lower subsequent returns), and the relation recurs through time though its strength varies by regime. UNKNOWN (recorded, NOT an invitation to search): why does the C₁* consequence vary materially through time while keeping its broad sign? A thousand macro variables (rates, vol, sentiment, liquidity, COVID) offer post-hoc stories; correlating the rolling coefficient against them now would re-open the garden. It sits as an open question. This moves Markets from "instrument-validation with an interesting secondary result" to a candidate finance contribution: meaningful in-domain replication, unresolved novelty, no pristine prospective confirmation.

Literature-location (novelty quarantined). The general principle "the internal shape of disagreement matters beyond its amount" is KNOWN. Ledger:

ClaimStatus
Disagreement predicts returnsknown (analyst dispersion; Johnson 2004)
Shape of a belief distribution matters beyond amountknown — closest prior: Hardouvelis-Karalas-Vayanos 2025 (range vs kurtosis/polarization of analyst beliefs, conditional)
Deterministic model-forecast dispersion predicts returnsknown — Machine Forecast Disagreement (RFS 2026); no conditional-concentration test located
Aggregating disagreement can hide substructureknown-adjacent, recent (LLM-decomposed retail disagreement, 2026)
Concentration of heterogeneous cross-characteristic dispersion, conditional on total dispersion, predicts returnsno direct prior located — the piece to protect

Own forensic web check (2026-08-20), adversarial (trying to FIND the exact prior): could not locate a direct equivalent. Verified differentiators from each neighbor — Hardouvelis-Karalas-Vayanos (closest): range/kurtosis are moments of the analyst-belief distribution about one quantity, ownership dispersion is a Herfindahl of holdings breadth — neither is concentration across heterogeneous characteristics; their coordinates are homogeneous beliefs, ours are heterogeneous firm-descriptions; their shape statistic is distributional kurtosis, C₁ is one-vs-many concentration of a fixed total cross-read dispersion. Stambaugh-Yuan mispricing score = arithmetic average of 11 anomaly ranks = the consensus/mean direction — the mathematical opposite of concentration-of-dispersion. Machine Forecast Disagreement = forecast dispersion amount; no concentration/higher-moment topology test located. Huang-Li-Wang "Are Disagreements Agreeable?" = PLS aggregation of many disagreement measures to time the market — combines, does not decompose; time-series not cross-section. Cross-sectional valuation-dispersion / higher-moment work is across firms or on the return distribution, not within-firm across-reads. So ~half of D being fundamentals-vs-fundamentals, the object is best named cross-characteristic state incoherence, not "belief disagreement." Novelty: plausible but unresolved, NOT claimed. Full-text confirmation (Hardouvelis, from the source PDF): they measure polarization by the kurtosis of, in their words, "the distribution of analysts' forecasts"; range measures intensity of disagreement; the Herfindahl index is their measure of ownership breadth; kurtosis is tested conditional on range. So it is a genuine shape-conditional-on-amount precedent — but over the analyst-forecast distribution, and their kurtosis/polarization maps onto our polarity axis over beliefs, not our primary concentration (C₁) over heterogeneous characteristics. Full-text confirmation (MFD, from the source NBER PDF): MFD is defined as "the standard deviation of [return forecasts] across investors" — dispersion amount only. A term search of the full 54-page text finds no kurtosis/skewness-as-predictor, no Herfindahl/concentration/participation-ratio, no "one rogue model vs distributed," and no conditional-on-total-dispersion shape test; the robustness section varies the ML model (ridge vs random forest) and geography, not the shape of disagreement. So both closest papers are now full-text-checked and neither contains the C₁* object. Remaining honest limit: "no prior located" ≠ "no prior exists" — an exact prior could still sit under other terminology or in a non-finance/ML paper; a literature search cannot prove the negative beyond this. Structural distinction a referee can read directly: Hardouvelis = distribution of homogeneous beliefs; MFD = dispersion across model forecasts; Stambaugh-Yuan = average direction across signals; Parallax = heterogeneous descriptions of one firm → D = total cross-description inconsistency, C₁* = concentration of that inconsistency conditional on its amount.

13.8 Paper build, referee-response, and adversarial self-validation (2026-08-20)

The frozen Stage-1/Stage-2 evidence was written up as a finance manuscript (measurement/representation framing, zero Parallax vocabulary): RESEARCH_NOTE.md (~6pp), MANUSCRIPT.md (~20pp, standalone HTML with embedded figures), title "Cross-Characteristic Dispersion and the Cross-Section of Stock Returns." Table 1 = plain joint D+C₁; C₁* the estimand-clean confirmation; PF/PO the falsification box; economic magnitude via a frozen conditional sort. Two external-review rounds then materially sharpened it.

Referee-response robustness (pre-declared, run_referee_response.py). (a) Short-term reversal: momentum is 12−1 so 1-month reversal was absent from controls. Adding STR leaves the shape coefficient essentially unchanged (C₁* +3.00→+2.97 pre-2016, +2.12→+2.21 later); the effect is not reversal in disguise. (b) Value-weighted magnitude: the ≈45 bps/21d equal-weighted spread attenuates to near zero value-weighted and among liquid names (later era EW −45 → VW +11 t=+0.47; liquid-EW −6). The magnitude is a small-cap phenomenon.

Adversarial self-validation (adversary_check.py) caught an over-claim of mine. I had framed it as "the statistical association survives; only the economic magnitude is small-cap." That split was false. The Fama–MacBeth slope itself collapses in liquid names: later-era C₁* t = +0.57 on the top-half-market-cap subset (vs +2.12 full), coefficient ~¼ of full; pre-2016 liquid t = +1.65 (vs +3.00); in the small-cap half it is stronger (t = +2.14, +3.49). The STR defense, by contrast, strengthened: C₁ and STR are essentially orthogonal (cross-sectional corr ≈ 0.00–0.04), so C₁ cannot be reversal relabeled (though STR is itself weak in a ≥\$5 universe).

Honest current scope of the return finding: a small-and-illiquid-stock phenomenon in both statistical and economic form — robust in small caps (t 2.1–3.5), weak-to-absent among liquid names (t≈0.6 later), survives reversal. This is a real narrowing from the "general cross-sectional relationship" the draft implied that morning. The measurement contribution (the aggregate-hider decomposition, the descriptive-before-returns spine, the PF/PO falsification) is independent of return magnitude and unaffected.

Limits-to-arbitrage vs microstructure-artifact test (pre-committed verdicts frozen BEFORE running — PREREG_LTA.md, run_lta.py). Committed: REAL = implementation-gap survives + an arbitrage-cost support; ARTIFACT = gap collapses; MIXED = gap survives but supports fail. Result: MIXED. The gap test (primary discriminator) SURVIVES and strengthens — skipping the first 5 trading days takes later-era C₁* from t=+2.12 to +3.40 (gap coef 126% of baseline), so the effect is NOT a microstructure / bid-ask / stale-price artifact (those decay within days). But both mechanism supports fail: C₁*×IVOL interaction t=+0.76 (<1.5), and the slope does not survive excluding the bottom market-cap quintile (later t=+0.55). Committed verdict honored: a genuine, non-artifact predictive signal CONFINED TO MICROCAPS, of UNRESOLVED mechanism — NOT claimed as limits-to-arbitrage. This rules out the near-fatal "just microcap microstructure noise" possibility (a real gain for the finding's validity) while leaving it without a mechanism or large-cap presence.

Reducibility kill (cheapest-kill generator; pre-committed verdicts frozen first — PREREG_REDUCE.md, run_reduce.py). Generator: "what single result ends the claim?" → C₁* is a nonlinear repackaging of the ten reads, so adding them as linear controls should absorb it. Committed: REAL = later-era C₁* |t|≥1.5 with the ten reads controlled; DEAD = |t|<1.5 or sign flip. Result: DEAD in the primary sample. Later era (n100): C₁* t +2.12 → −0.06 once the ten reads are added (ratio −0.03). Pre-2016 (n26): +3.00 → +2.87 survives. Committed verdict (primary = later): REDUCIBLE. A relation that holds in 26 months and vanishes in 100 is not robust; honored the pre-commitment, did NOT switch to the surviving sample. So C₁*'s RETURN predictability is the reads' own linear signals repackaged through concentration, not an incremental geometric effect. This is the biggest deflation of the arc. License: the descriptive aggregate-hider (D lossy at fixed D) is UNAFFECTED (a geometric fact, no returns); the result does NOT license "incremental return information beyond the characteristics" or "a new return factor." The paper retreats to the descriptive decomposition + measurement object; the return section honestly states its predictability is not robustly distinct from the characteristics.

Nonlinear reducibility kill (final return-side test; pre-committed — PREREG_NONLIN.md, run_nonlin.py). The single unrun, most lethal residual: does C₁* survive beyond a NONLINEAR expansion of the ten reads (10 squares + all 45 pairwise products) as controls? Committed: RESURRECT = later-era t≥+2.0 AND coef≥0.5× baseline; DEAD = else. Result: DEAD (total). Later era (n99): C₁* coef=+0.000214, t=+0.62, ratio 0.24 — both criteria fail. Pre-2016 (n26, UNBANKABLE): t=+3.07 survives, but n=26 and NOT pursued per the spec (chasing it is the motivated save). Committed verdict: the return contribution is reducible to the reads even nonlinearly. THE RETURN QUESTION IS CLOSED. License: closes the return contribution; does not license a factor, a general effect, or a mechanism; the descriptive aggregate-hider was out of scope and is unaffected. Manuscript reducible linearly AND nonlinearly; re-frozen. STOP — no further return-side tests.

14. Provenance

Data: data_cache/*.parquet (Yahoo ETFs; Sharadar survivorship-free bulk). Gates: LOCKED_CRITERION.md, PRE-REGISTRATION_v2/v3/pathB/pathC.md (unedited post-run). Runners: run_horserace_v2/v3.py, run_pathB/C_sharadar.py, factor-zoo, robustness. Confirmation + audits (2026-08): run_confirmation.py (frozen harness), verify_characterize.py, audit_confirmation.py, reconcile_D.py, geometry_compare.py, shape_descriptive.py, dominant_read_audit.py. Shape consequence + temporal (frozen, hashed): PREREG_SHAPE.md (9cd845b7…) / run_shape_consequence.py (381f9497…); PREREG_SHAPE_TEMPORAL.md (43da5237…) / run_shape_temporal.py; magnitude sort PREREG_PORTFOLIO.md (0280ab8f…) / run_portfolio.py. Referee-response + adversarial (2026-08): run_plain_joint.py, run_referee_response.py, adversary_check.py. Paper: RESEARCH_NOTE.md, MANUSCRIPT.md/.html, figures fig1-3. Result dumps *_RESULT.txt. Log: RESULTS_AND_LEARNINGS.md. This page: build_labbook.py.