Russell Parrish · Parallax Metrology · www.parallaxmetrology.com
Working manuscript: frozen Stage 1 evidence plus preregistered Stage 2 economic-magnitude characterization. Built from the frozen research note (sha 902543b1…) and the frozen empirical record.
We represent each firm as a vector of ten heterogeneous, standardized descriptions spanning price, accounting fundamentals, institutional ownership, and insider activity. The within-firm dispersion across these standardized characteristics, D, is the conventional summary of how much a firm’s descriptions disagree. We show that D is a lossy summary of firm state: among firms with nearly identical D, the disagreement can be concentrated in a single description or distributed across many, and these are not return-equivalent. In a pre-registered, frozen test on strictly pre-2016 data, the concentration of disagreement carries incremental information about subsequent returns conditional on its amount. Firms whose disagreement is more distributed earn lower subsequent returns than firms of equal D whose disagreement is concentrated in one description. The relation is present in the simplest joint specification, is corroborated by two re-parameterizations on the same data (not independent confirmations), survives adding short-term reversal to the controls, and recurs, in two disjoint regimes with a quiet stretch between them, in a later era whose concentration–return relationship had not previously been examined. A pre-specified semantic decomposition of the same dispersion (price-versus-fundamentals and price-versus-ownership conflict) did not survive the frozen test, while the label-light geometric decomposition did. We locate the object against the closest prior work and do not identify a direct equivalent. We make no causal claim and no claim of prospective confirmation; the evidence is retrospective / secondary historical validation of a frozen hypothesis. The pre-registered primary window is short (26 monthly cross-sections, limited by 13F coverage), so the result leans on larger-sample reproductions (a ~100-month later era and ≥6-read constructions) rather than small-sample Newey–West inference alone; and we characterize an economic magnitude (≈45 bps per 21 days equal-weighted in the longer sample). The relation is concentrated in small, illiquid stocks: among liquid (top-half market-cap) names both the Fama–MacBeth slope and the equal-weighted magnitude weaken sharply and lose significance, so this is a small-cap phenomenon rather than a general cross-sectional factor. In the better-powered later sample the return predictability is moreover reducible to the underlying characteristics, linearly and under a nonlinear (squares and pairwise-product) expansion alike (it is absorbed once the ten reads are controlled), so our contribution is the descriptive decomposition and the measurement object, not an incremental return factor, and we do not claim a tradeable one.
Throughout, “disagreement” and “inconsistency” are shorthand for within-firm dispersion across standardized characteristics; the primitive is dispersion, the interpretive gloss is disagreement.
A large literature relates disagreement (dispersion in beliefs, forecasts, or signals about a firm) to the cross-section of expected returns. Almost all of it treats disagreement as a scalar: the standard deviation of analyst forecasts, the dispersion of machine-model predictions, the spread of a signal. This paper asks a question that a scalar cannot answer. Two firms can carry the same amount of disagreement built in completely different ways: in one, a single description ruptures far from the firm’s own consensus while the rest agree; in the other, no description is extreme but the set splits into opposing camps. The amount is identical; the organization is not. Does the organization matter?
We make the question precise with a simple geometric decomposition.
Represent a firm as a vector of n standardized
descriptions, x = (x₁,…,xₙ). Its mean squared
magnitude decomposes exactly into a consensus term and a dispersion
term, MSQ = (1/n)Σxᵢ² = x̄² + D, where x̄ is the
within-firm mean and D = (1/n)Σ(xᵢ−x̄)² (equivalently
‖x‖² = n(x̄² + D)). Prior disagreement measures typically
retain an analogous scalar amount of dispersion; we ask whether the
internal organization of D, specifically how its
concentration is distributed across the descriptions, carries
information beyond D itself. We measure concentration by
C₁ = maxᵢ(xᵢ−x̄)² / Σ(xᵢ−x̄)²: near 1/n when
disagreement is evenly spread, near 1 when one description
carries it all.
We report three results, in an order chosen to make the object credible before any return is examined.
First, D is descriptively lossy. At
approximately fixed D, C₁ varies enormously
across real firms. We exhibit matched twins (firms with nearly equal
D but opposite internal geometry) selected purely on their
description vectors, with no reference to returns. This establishes the
precondition: the aggregate is hiding structure.
Second, the hidden structure carries consequence. In
a specification frozen and hashed before returns in the pre-2016 panel
were examined for this shape hypothesis (the panel itself had previously
been analyzed for other questions), concentration conditional on amount
predicts subsequent returns. Because D enters negatively
and C₁ positively, the most negatively-priced state at a
given amount of disagreement is broadly distributed
disagreement, not a single-description rupture. The result holds in the
simplest joint regression of D and C₁, and
survives a short-term-reversal control, but it is concentrated in small,
illiquid stocks and is weak-to-absent among liquid names (Section
7).
Third, the relation recurs. Under a separately frozen temporal protocol, the relation is positive across the whole chronology with no material sign reversal and no single-year dependence, and reproduces in a later era at about half magnitude with time-varying strength.
Alongside these we report an informative failure. A semantically natural decomposition (is it specifically the price view conflicting with fundamentals, or with ownership?) looked strong in exploratory work and did not survive the frozen test, while the label-light concentration measure did. We read this not as a defeat but as evidence about what kind of decomposition is durable: how disagreement is organized proved more durable, in this sample, than which named components conflict.
Methodological stance. The disagreement/anomaly literature is unusually exposed to specification search. We therefore adopt a pre-registration discipline throughout: the descriptive stage touches no returns; the consequence specification (reads, normalization, concentration definition, cap, controls, inference, pass/null interpretation) is fixed and hashed before returns in the pre-2016 panel are examined for the concentration/shape hypothesis (the panel itself having previously been analyzed for other questions); the temporal protocol is fixed and hashed separately; and each is run once. The claim-status timeline in Appendix A records the full arc, including the abandoned steps, so the reader can see that the surviving specification was not selected on its return performance.
What we do not claim. We do not identify an economic mechanism, we do not construct a tradeable factor, and we do not offer prospective confirmation. The pre-2016 window, though frozen with respect to this hypothesis, had been examined earlier for other questions; the evidence is retrospective / secondary historical validation. Section 9 is explicit about each limit; Section 10 separates the remaining upgrades, mechanism and confirmation, that would strengthen it (economic magnitude having been characterized in Section 7).
The idea that disagreement predicts returns is well established. Diether, Malloy and Scherbina (2002) and Johnson (2004) document the negative relation between analyst forecast dispersion and subsequent returns; a large body of work refines the measure and debates its interpretation (risk versus mispricing, conditioning on sentiment, optimism, or future profitability). This literature is about the amount of disagreement.
Closest to our question is work on the shape of the belief distribution. Hardouvelis, Karalas and Vayanos (2026) decompose the distribution of analysts’ forecasts into an intensity dimension (range) and a polarization dimension (kurtosis) and show both matter for returns and ownership, testing one conditional on the other. This is a genuine “shape beyond amount” precedent. It differs from our object on three axes: its coordinates are homogeneous beliefs (many analysts forecasting the same quantity) where ours are heterogeneous descriptions of one firm; its shape statistic is distributional kurtosis (polarization), nearer to a two-sided split than to one-vs-many concentration; and its ownership dispersion is a Herfindahl of holdings breadth, not of the disagreement itself.
Bali, Kelly, Mörke and Rahman (2026) construct Machine Forecast Disagreement as the standard deviation of return forecasts across machine-learning “investor-models.” This is a deterministic, model-based disagreement measure with excellent coverage, and it strongly negatively predicts returns, but it is dispersion amount; a full-text reading finds no concentration or higher-moment test conditional on that amount.
The composite-anomaly literature (Stambaugh and Yuan, 2017) combines
many anomaly signals into a mispricing score by averaging
ranks, the consensus/mean direction across signals. This is, if
anything, the opposite construction to ours: a stock on which all
descriptions agree has a strong composite score and low D;
a stock whose descriptions split has near-zero average score and high
D. Higher-moment work on the return distribution
(implied skewness, kurtosis) and recent work on the structural
organization of accounting information are conceptually adjacent
examples of “structure beyond a first or second moment,” but neither
constructs concentration of cross-characteristic dispersion.
Our contribution is therefore specific: not “shape matters” (known), but that the concentration of heterogeneous cross-characteristic dispersion, conditional on its total amount, carries incremental return information. We do not identify a direct equivalent in full-text checks of the closest work, and we do not claim its non-existence.
Universe. We use Sharadar’s survivorship-free US
equity database, which retains delisted names. We keep domestic common
stock on NYSE, NASDAQ, and NYSE-MKT, with month-end price ≥ $5, and
require ≥100 valid names per monthly cross-section. All information is
point-in-time: fundamentals are lagged to their reporting
datekey, and 13F holdings are lagged 50 days beyond the
filing date. We study two non-overlapping periods: a strictly
pre-2016 window used for the primary consequence test,
and a later development era (2016–2026, with the homogeneous
all-ten panel available from March 2018 onward) used to develop
the descriptive construction and, separately, to test temporal
recurrence.
The ten descriptions (“reads”). Each is a firm characteristic computed point-in-time, winsorized at 1/99 and standardized cross-sectionally (mean 0, unit variance) each month:
| # | Read | Definition |
|---|---|---|
| 1 | Momentum | log price ratio, t−12 to t−1 |
| 2 | Gross profitability | gross profit / assets |
| 3 | ROA | net income / assets |
| 4 | Earnings surprise (SUE) | Δ4-quarter EPS / rolling std |
| 5 | Sales growth | 4-quarter revenue growth |
| 6 | Accruals | −(net income − free cash flow) / assets |
| 7 | Asset growth | −(4-quarter asset growth) |
| 8 | Value | −log(price/book) |
| 9 | Institutional breadth Δ | change in count of 13F holders (lagged 50d) |
| 10 | Insider net buying | signed trailing-6-month insider purchases |
These are heterogeneous in economic meaning and span several data families, but they are not assumed statistically or source-independent: several fundamentals (reads 2–7) are transforms of overlapping financial-statement information. The observed mean pairwise absolute correlation is ≈ 0.12: low empirical redundancy, not informational independence. This distinction is load-bearing and we return to it in Section 9.
Controls. Idiosyncratic volatility (residual from a 60-day market-model), size (log market cap), value (−log P/B), and momentum, each standardized cross-sectionally.
For a firm with n present reads xᵢ, let
x̄ = (1/n)Σxᵢ and dᵢ = xᵢ − x̄. We use:
D = (1/n)Σ dᵢ², dispersion of
the reads around the firm’s own consensus.C₁ = maxᵢ dᵢ² / Σ dᵢ²,
bounded in [1/n, 1].P = 2·min(E₊,E₋)/(E₊+E₋), where
E₊ = Σ_{dᵢ>0} dᵢ²: 0 for a one-sided rupture, 1 for a
balanced two-sided split.The decomposition MSQ = (1/n)Σxᵢ² = x̄² + D cleanly
separates the firm’s consensus level from its internal dispersion; this
paper concerns the internal structure of D, holding the
consensus term aside.
Concentration conditional on amount. C₁
mechanically co-moves with D (higher-dispersion states tend
to be more rupture-like). Because our question is explicitly conditional
(at the same amount of disagreement, does concentration
matter), we also form a residual,
C₁* = min(C₁,0.80) − Ê(C₁|D), where Ê(C₁|D) is
a monotone (isotonic) fit of expected capped concentration on
D, estimated once on the later-era development sample and
applied unchanged to the test sample. The 0.80 cap is a pre-test guard
motivated in Section 5. We stress that C₁* is a
conditioning device, not a new statistic: the plain joint
regression of raw D and raw C₁ (Section 7)
makes the same point without it.
Homogeneous observer system. For the primary test we
require all ten reads present, so every firm-month is represented by the
identical ten-description system and D, C₁,
and C₁* have the same meaning across observations. This
costs sample (13F availability restricts the pre-2016 primary to ~26
months); we accept it because a homogeneous instrument is preferable to
a larger sample whose observer composition drifts. A ≥6-read
construction, with concentration residualized within read-count strata,
is carried as robustness.
Pre-registration. The consequence specification and the temporal protocol were each written, hashed, and only then run once (Appendix C). This is the paper’s principal defense against the garden of forking paths, and the reason the descriptive stage (Section 5) is reported before any return regression.
Before testing consequence we establish that D genuinely
hides structure. On 281,538 later-era firm-months with all ten reads
present (price ≥ $5), we examine how much C₁ varies at
fixed D.
Within D deciles, the interquartile and 5–95% ranges of
C₁ are wide throughout the middle of the distribution: for
representative mid-D deciles the 5–95% span of
C₁ runs from roughly 0.26 to 0.78. At nearly identical
aggregate disagreement, firms occupy geometries from near-uniform to
near-rupture.
Matched twins. Selecting purely on geometry (no
returns), we can pair firms at D ≈ 0.30 with opposite
concentration. A representative pair:
| Firm | D | C₁ | Read vector (z-scores) |
|---|---|---|---|
| A | 0.30 | 0.86 | one read at −1.7 (earnings surprise), all others ≈ 0 |
| B | 0.30 | 0.17 | many reads at ±0.5–0.8, balanced positive and negative |
Same amount of disagreement; A is a single-description rupture, B is distributed inconsistency. This is the paper’s central picture (Figure 1 schematic; Figure 2 the real pair).
One shape axis. Concentration and polarity are nearly collinear in this sample (corr ≈ −0.86): a concentrated state is almost mechanically one-sided. We therefore carry a single shape axis, concentration, and note that a separate polarity test is not independent information here.
Confound audit. We ask which read owns the maximum
squared deviation in high-C₁ states. Through moderate
concentration the answer is spread across five or six reads; but the
extreme tail (C₁ > 0.8) is dominated by two structurally
sparse, spiky reads: insider net buying (43%) and sales growth (15%),
whose intermittent large values mechanically produce single-read
ruptures. This is a measurement artifact, not signal, and it motivates
two pre-test guards used in Section 7: cap C₁ at 0.80, and,
as robustness, recompute the geometry after dropping the two sparse
reads.
Crucially: we did not select a shape statistic because it predicted returns; we first established that the aggregate is lossy, then asked whether the loss is consequential.
A natural hypothesis is that specific named conflicts drive
any effect, e.g. the price view (read 1) conflicting with the
fundamental block (reads 2–8) or with the ownership block (read 9).
Define PF = (price − fundamental-mean)² and
PO = (price − ownership)². In the later-era development
sample, PF and PO entered strongly. We froze a
test of whether they add beyond total dispersion D and
consensus, and ran it on the pre-2016 sample. They did not: entered
jointly with D, consensus, and controls,
PF t = −0.65, PO t = −0.66, joint Wald
p = 0.45.
A predefined semantic decomposition did not replicate out of sample. This motivated no change to the dispersion measure; instead we asked the more general, label-light question of how a fixed amount of deviation is distributed across coordinates. That named attribution failed while label-light organization survived is itself informative, and makes the surviving result harder to interpret as a simple continuation of the original semantic decomposition.
Specification (frozen and hashed before returns in the pre-2016 panel were examined for this shape hypothesis; the panel itself had previously been analyzed for other questions). Monthly Fama–MacBeth regressions of 21-day forward returns on the shape variable and controls (idiosyncratic volatility, size, value, momentum), Newey–West(3) inference, ≥100 names/month, price ≥ $5, forward returns winsorized 1/99 per month and truncated at the window end so no month looks past the outcome boundary.
Table 1. Plain joint specification (raw D and
raw C₁ in one regression, no interaction).
| Sample | D coef (t) | C₁ coef (t) |
|---|---|---|
| pre-2016 (n=26) | −0.00313 (−2.66) | +0.00188 (+2.79) |
| pre-2016, C₁ capped 0.80 | −0.00309 (−2.64) | +0.00182 (+2.75) |
| later era 2018–2026 (n=100) | −0.00284 (−4.04) | +0.00080 (+1.99, marginal at 5%) |
Amount and concentration enter with opposite signs, and each retains
explanatory power conditional on the other and the controls; the cap is
inert (capped vs raw C₁, t = +2.75 vs +2.79). In the later
era the concentration coefficient is positive but marginal by the
conventional two-sided 5% threshold.
Secondary interaction diagnostic. Adding
D×C₁ leaves the main coefficients in the same neighborhood,
with interaction t = +1.47 (pre-2016) and
t = +2.24 (later era), mixed across periods and not part of
the core claim.
Table 2. Residualized and robustness constructions (pre-2016 primary).
| Specification | C₁ / C₁* (t) |
|---|---|
Residualized C₁\*, all-ten |
+3.00 |
| Robustness: ≥6-read, within-count residualization | +4.32 |
| Robustness: eight-read (drop insider, sales growth), geometry recomputed | +3.58 |
D coefficient (residualized primary) |
−2.44 |
The residualized form conditions on D nonlinearly
(isotonic) rather than linearly and reaches the same conclusion; raw
C₁ and C₁* are the same relationship
differently parameterized, so the agreement is a robustness check, not
independent evidence. Both robustness constructions agree in sign, and
the D×C₁* interaction is insignificant in the primary.
The eight-read construction is the direct test of whether the
effect is an artifact of the two sparse tail-dominating reads:
it drops insider and sales growth entirely and recomputes D
and C₁ on the remaining eight, yet the effect strengthens
(t=+3.58), so it is not driven by those two channels. (The robustness
specifications running stronger than the plain later-era C₁
has a benign source: the ≥6-read construction spans more months, and
C₁* conditions on D nonlinearly; they are not
independent votes.)
Short-term reversal. Because momentum enters as t−12
to t−1, one-month reversal is not among the frozen controls, and a
natural concern is that a concentrated single-read state is really a
recent extreme price print. Adding short-term reversal (STR, the
past-21-day return) as an additional control leaves the shape
coefficient essentially unchanged (pre-2016 C₁* t = +2.97
vs +3.00, plain C₁ t = +2.77 vs +2.79; later-era
C₁* t = +2.21 vs +2.12), while STR itself is insignificant
(t = +0.36 and −0.68). The effect is not short-term reversal in
disguise. The direct evidence is that C₁ and STR are
essentially uncorrelated in the cross-section (mean correlation ≈ 0.00
to 0.04 in both periods), so concentration cannot be reversal relabeled;
this is consistent with C₁ being direction-agnostic (it
encodes concentration, not the sign of any read) and with the
high-C₁ tail being dominated by non-price reads. We note
that short-term reversal is itself weak in this ≥$5 universe (STR-alone
t = +0.44 and −0.52), which limits how much a reversal control can
absorb; the near-zero correlation settles the question regardless.
Sign and interpretation. Because D is
negative and C₁ (concentration) positive, the most
negatively-priced state at a given amount of disagreement is
high-D, low-C₁: broadly distributed disagreement
across many descriptions, rather than a single-description rupture.
Economically, this is an association between a measured state and
subsequent returns; it does not establish that investors perceive
“incoherence,” nor that a latent known characteristic does not generate
both.
Economic magnitude. To express the estimand in
interpretable units we form a pre-specified conditional sort (frozen
separately; Appendix C): within D quintiles we sort firms
into C₁* terciles, pool the bottom (distributed) and top
(concentrated) terciles across D quintiles (holding
aggregate dispersion approximately fixed), and equal-weight each leg’s
21-day forward return. We report the distributed-minus-concentrated
spread (a zero-net-investment return difference by construction; we do
not claim market-beta neutrality).
| Sample | R_distributed | R_concentrated | distributed − concentrated (t) |
|---|---|---|---|
| pre-2016 (26 mo) | −137 bps | −41 bps | −96 bps / 21d (−3.94) |
| later era 2018–2026 (99 mo) | −74 bps | −29 bps | −45 bps / 21d (−3.56) |
Holding aggregate disagreement roughly fixed, distributed states earn on the order of 45 bps per 21-day period less than concentrated states in the longer later sample (≈96 bps in the thin pre-2016 window, consistent with the larger Fama–MacBeth slope there). This is a characterization of the Stage-1 estimand’s economic size, deliberately not a trading claim: the legs are equal-weighted and un-costed, and by pre-registration this sort neither confirms nor disconfirms the frozen concentration coefficient. A coarse portsort and a Fama–MacBeth slope are two views of the same association, on the same data.
Weighting and liquidity. The magnitude is an
equal-weighted, small-name phenomenon. Value-weighting the legs by
market cap, or restricting to the top half by market cap, attenuates the
spread to near zero and insignificance (later era: value-weighted +11
bps, t = +0.47; liquid equal-weighted −6 bps, t = −0.58; liquid
value-weighted +17 bps, t = +0.71; pre-2016 value-weighted −32 bps, t =
−1.07). The economic footprint therefore lives disproportionately in
small, illiquid stocks, and the equal-weighted ≈45 bps is not
representative of large or liquid names. This concentration is
not only economic: the Fama–MacBeth slope itself weakens sharply in
liquid names. Restricting the cross-sectional regression to the
top half by market cap, the C₁* coefficient falls to
roughly a quarter of its full-sample size and loses significance in the
better-powered later era (later-era liquid t = +0.57 vs
+2.12 full; pre-2016 liquid t = +1.65 vs +3.00,
n = 26), while in the small-cap half it is if anything
stronger (t = +2.14 and +3.49). The finding is therefore a
small-and-illiquid-stock phenomenon in both its statistical and its
economic form, not a general cross-sectional relationship whose
magnitude merely happens to be larger in small names.
Microstructure artifact or mispricing? Because the
effect lives in small, illiquid stocks, we pre-registered a test
(frozen; Appendix C) to distinguish a genuine slow mispricing from a
microstructure artifact such as bid-ask bounce or stale prices.
It is not a microstructure artifact: the
C₁* slope survives, and in fact strengthens, when the
outcome skips the first five trading days (later-era
t = +3.40 on the t+5→t+26 return vs +2.12
contemporaneously, gap coefficient 126% of baseline), whereas bid-ask
and stale-price effects decay within days. But we cannot attribute it to
limits to arbitrage either: the
C₁* × idiosyncratic-volatility interaction (IVOL being the
standard arbitrage-cost proxy) is insignificant (t = +0.76), and the
slope does not survive excluding the bottom market-cap quintile
(later-era t = +0.55). By the pre-registered rule the
result is mixed: a genuine, non-artifact predictive signal
confined to microcap stocks, of unresolved economic mechanism. We do not
claim a limits-to-arbitrage channel.
Reducibility to the characteristics (the decisive
limitation). The frozen test controls for amount D
and four standard factors. A sharper question is whether concentration
adds beyond the ten underlying reads themselves, since
C₁ is built from them. Adding all ten reads as linear
controls, C₁*‘s return predictability is essentially fully
absorbed in the primary later-era sample (t falls from
+2.12 to −0.06), surviving only in the thin pre-2016
window (t = +2.87 vs +3.00). By the pre-registered rule
(primary = later era; Appendix C) this is a reducible result:
C₁*’s return signal is the reads’ own linear return signals
repackaged through the concentration statistic, not an incremental
geometric effect that a linear characteristic model would miss. A
relationship that holds in 26 months and vanishes in 100 is not robust.
A pre-registered nonlinear extension of this kill reaches the same
verdict: adding the ten reads’ squares and all 45 pairwise products as
controls, C₁* in the primary later sample is
t = +0.62 (coefficient 24% of the linear baseline), so its
return signal is reducible to the reads even nonlinearly. (The pre-2016
window survives even this, t = +3.07, but at n = 26 it is
underpowered and we do not bank it.) We therefore do not claim
that concentration carries incremental return information beyond the
underlying characteristics, linearly or nonlinearly; on this evidence
the return question is closed. The concentration geometry
remains a robust descriptive fact (Section 5); its
return predictability is not distinct from the underlying
characteristic signals.
We fixed a temporal protocol (instrument unchanged, only time segmentation varied), hashed it, and ran it once (Appendix C).
Non-overlapping blocks (all-ten). Positive throughout, no material sign reversal:
| Block | months | C₁* (t) |
|---|---|---|
| pre-2016 | 26 | +3.00 |
| 2018-03..2018-11 | 9 | +0.48 |
| 2018-12..2021-05 (COVID-spanning) | 30 | +0.04 |
| 2021-06..2023-11 | 30 | +3.06 |
| 2023-12..2026-05 | 30 | +1.01 |
Leave-one-calendar-year-out (later era). All eleven
folds positive, t = +1.6…+2.8; no single year is
individually necessary. We flag that this reassurance is limited:
because the later-era signal concentrates in the 2021–2023 regime,
dropping any single year leaves that regime largely intact, so
“all folds positive” is closer to mechanical than to independent
corroboration.
Rolling 36-month windows (stepped 12 months). Near
zero over 2017–2020, then rising and remaining positive from ~2020 (peak
t = +2.87); one window is trivially negative
(t = −0.05). Overlapping windows are a diagnostic path, not
independent tests (Figure 3).
Later-era reproduction. Pooled later-era all-ten
C₁* t = +2.12 (≥6-read companion +2.77; plain
C₁ +1.99), same sign at roughly half the pre-2016
magnitude.
The effect recurs but is not constant. Two non-overlapping strong periods (pre-2016 and 2021–23), no single-year dependence, and reproduction in a later un-examined era place it well beyond a one-period accident; the near-zero 2018–2021 stretch and the magnitude attenuation place it well short of a constant law. We record the regime dependence (why the consequence of the same measured geometry varies materially over time while keeping its sign) as an open question, not something we model, to avoid post-hoc correlation against a thousand available macro variables. One specific caution belongs here: the strong 2021–23 window overlaps a high-dispersion macro regime, and while the specification controls for idiosyncratic volatility cross-sectionally and the leave-one-year-out result survives dropping any single year (2020 included), the cross-sectional control does not absorb time-series volatility regimes. We therefore cannot exclude that the relation strengthens in high-dispersion periods, and we do not treat the later-era reproduction as fully independent of that regime.
Established. (i) Aggregate cross-characteristic
dispersion is descriptively lossy: at fixed D, real firms
occupy sharply different concentration geometries. (ii) Conditional on
amount, concentration carries incremental information about subsequent
returns in a frozen pre-2016 test, present in the simplest joint
specification and robust to two re-parameterizations and to a
short-term-reversal control, but concentrated in small, illiquid stocks,
weak-to-absent among liquid names, and (in the better-powered later
sample) reducible to the underlying characteristics when they are added
as linear controls (Section 7). (iii) The relation recurs, weaker and
regime-varying, in a later era whose concentration–return relationship
had not been examined. (iv) Within its microcap scope the effect is not
a microstructure artifact: it survives, and strengthens under, a
five-day implementation gap (Section 7).
Not established. A causal or behavioral mechanism (a
pre-registered test found the effect is not a microstructure artifact
but is confined to microcaps with no significant arbitrage-cost
gradient, so limits to arbitrage is not established either); that its
return predictability is incremental to the underlying characteristics
(in the better-powered later sample it is absorbed once the ten reads
are added as controls, linearly and under a nonlinear
squares-and-products expansion alike, Section 7); that investors
perceive “disagreement” or “incoherence”; that concentration is a
distinct priced risk factor; that any single description is an
informationally independent observer (the reads are low-redundancy,
|corr| ≈ 0.12, not independent); generality across the size
distribution (Section 7: among liquid names both the Fama–MacBeth slope
and the equal-weighted magnitude weaken sharply and lose significance,
so the finding is a small-and-illiquid-stock phenomenon, its presence in
large caps and its tradability both unestablished); and, most
importantly, prospective confirmation. The pre-2016
window, though frozen with respect to this hypothesis, had been examined
earlier for other questions, so the evidence is retrospective /
secondary historical validation rather than a clean out-of-sample
experiment. An omitted known characteristic that generates both
distributed disagreement and low returns remains possible.
Economic relevance has now been characterized through the preregistered conditional sort in Section 7; this does not alter confirmation status. Two live frontiers remain, of different kinds, not points on one confirmation ladder:
Each requires its own pre-registration. We report the economic-magnitude characterization only; mechanism and confirmation remain deliberately untested here.
Conventional dispersion measures the amount of cross-characteristic disagreement but not its internal organization. That distinction is empirically material: firms with nearly identical aggregate dispersion can have sharply different characteristic geometries, and a concentration measure defined independently of returns carries incremental information conditional on dispersion in a frozen historical test, recurring at smaller magnitude in a later period. More distributed disagreement is associated with lower subsequent returns. A semantic decomposition into specific price–fundamental and price–ownership conflicts failed out of sample, while the label-light geometric decomposition survived, suggesting that, in this sample, the organization of dispersion was more durable than attribution to particular characteristic families. Why these states differ economically remains unresolved.
To make specification search auditable, the full sequence, including abandoned steps:
D →
out-of-sample supportive, but known. D across the
ten reads is priced out of sample under Newey–West(3) inference (t ≈ −2
to −3 with controls), but dispersion-is-priced is an established effect,
not a discovery.D out of sample (Section 6).C₁*) →
frozen secondary validation. The label-light geometric
decomposition survived its frozen test (Section 7).The surviving specification (step 4) was frozen and hashed before its pre-2016 concentration–return relationship was examined; steps 1–3 are the path that produced the hypothesis, not evidence selected on its return performance.
Reads standardized cross-sectionally each month after 1/99
winsorization. D, C₁, P as in
Section 4. Ê(C₁|D) fit by isotonic (monotone) regression on
the later-era development sample, per read-count stratum for the ≥6-read
robustness, with endpoint-clamp extrapolation; frozen and applied
unchanged. Forward returns are 21-trading-day log returns, truncated at
the sample boundary. Inference is Fama–MacBeth with Newey–West(3) on the
monthly coefficient series.
| Artifact | Role | SHA-256 (prefix) |
|---|---|---|
PREREGISTRATION.md |
original PF/PO confirmation | 0299723b… |
run_confirmation.py |
PF/PO harness | f9d9a08b… |
PREREG_SHAPE.md |
frozen consequence spec | 9cd845b7… |
run_shape_consequence.py |
consequence harness | 381f9497… |
PREREG_SHAPE_TEMPORAL.md |
frozen temporal protocol | 43da5237… |
run_plain_joint.py |
Table 1 transparency spec | f54c67a6… |
PREREG_PORTFOLIO.md |
frozen Stage-2 magnitude sort | 0280ab8f… |
run_portfolio.py |
economic-magnitude harness | 38a499ae… |
PREREG_LTA.md |
frozen artifact-vs-mechanism test | b9e5d34e… |
run_lta.py |
limits-to-arbitrage / gap harness | 15f95ee8… |
PREREG_REDUCE.md |
frozen reducibility kill | c6ab1b1b… |
run_reduce.py |
reducibility harness (linear) | 4cc153ae… |
PREREG_NONLIN.md |
frozen nonlinear reducibility kill | db22bb13… |
run_nonlin.py |
nonlinear reducibility harness | e55eb5cf… |
RESEARCH_NOTE.md |
frozen Stage-1 note | 902543b1… |
Descriptive and audit code (shape_descriptive.py,
dominant_read_audit.py, reconcile_D.py,
geometry_compare.py) and figures are in the project
record.