Cross-Characteristic Dispersion and the Cross-Section of Stock Returns

Cross-Characteristic Dispersion and the Cross-Section of Stock Returns

Concentration beyond amount

Russell Parrish · Parallax Metrology · www.parallaxmetrology.com

Working manuscript: frozen Stage 1 evidence plus preregistered Stage 2 economic-magnitude characterization. Built from the frozen research note (sha 902543b1…) and the frozen empirical record.


Abstract

We represent each firm as a vector of ten heterogeneous, standardized descriptions spanning price, accounting fundamentals, institutional ownership, and insider activity. The within-firm dispersion across these standardized characteristics, D, is the conventional summary of how much a firm’s descriptions disagree. We show that D is a lossy summary of firm state: among firms with nearly identical D, the disagreement can be concentrated in a single description or distributed across many, and these are not return-equivalent. In a pre-registered, frozen test on strictly pre-2016 data, the concentration of disagreement carries incremental information about subsequent returns conditional on its amount. Firms whose disagreement is more distributed earn lower subsequent returns than firms of equal D whose disagreement is concentrated in one description. The relation is present in the simplest joint specification, is corroborated by two re-parameterizations on the same data (not independent confirmations), survives adding short-term reversal to the controls, and recurs, in two disjoint regimes with a quiet stretch between them, in a later era whose concentration–return relationship had not previously been examined. A pre-specified semantic decomposition of the same dispersion (price-versus-fundamentals and price-versus-ownership conflict) did not survive the frozen test, while the label-light geometric decomposition did. We locate the object against the closest prior work and do not identify a direct equivalent. We make no causal claim and no claim of prospective confirmation; the evidence is retrospective / secondary historical validation of a frozen hypothesis. The pre-registered primary window is short (26 monthly cross-sections, limited by 13F coverage), so the result leans on larger-sample reproductions (a ~100-month later era and ≥6-read constructions) rather than small-sample Newey–West inference alone; and we characterize an economic magnitude (≈45 bps per 21 days equal-weighted in the longer sample). The relation is concentrated in small, illiquid stocks: among liquid (top-half market-cap) names both the Fama–MacBeth slope and the equal-weighted magnitude weaken sharply and lose significance, so this is a small-cap phenomenon rather than a general cross-sectional factor. In the better-powered later sample the return predictability is moreover reducible to the underlying characteristics, linearly and under a nonlinear (squares and pairwise-product) expansion alike (it is absorbed once the ten reads are controlled), so our contribution is the descriptive decomposition and the measurement object, not an incremental return factor, and we do not claim a tradeable one.

Throughout, “disagreement” and “inconsistency” are shorthand for within-firm dispersion across standardized characteristics; the primitive is dispersion, the interpretive gloss is disagreement.


1. Introduction

A large literature relates disagreement (dispersion in beliefs, forecasts, or signals about a firm) to the cross-section of expected returns. Almost all of it treats disagreement as a scalar: the standard deviation of analyst forecasts, the dispersion of machine-model predictions, the spread of a signal. This paper asks a question that a scalar cannot answer. Two firms can carry the same amount of disagreement built in completely different ways: in one, a single description ruptures far from the firm’s own consensus while the rest agree; in the other, no description is extreme but the set splits into opposing camps. The amount is identical; the organization is not. Does the organization matter?

We make the question precise with a simple geometric decomposition. Represent a firm as a vector of n standardized descriptions, x = (x₁,…,xₙ). Its mean squared magnitude decomposes exactly into a consensus term and a dispersion term, MSQ = (1/n)Σxᵢ² = x̄² + D, where is the within-firm mean and D = (1/n)Σ(xᵢ−x̄)² (equivalently ‖x‖² = n(x̄² + D)). Prior disagreement measures typically retain an analogous scalar amount of dispersion; we ask whether the internal organization of D, specifically how its concentration is distributed across the descriptions, carries information beyond D itself. We measure concentration by C₁ = maxᵢ(xᵢ−x̄)² / Σ(xᵢ−x̄)²: near 1/n when disagreement is evenly spread, near 1 when one description carries it all.

We report three results, in an order chosen to make the object credible before any return is examined.

First, D is descriptively lossy. At approximately fixed D, C₁ varies enormously across real firms. We exhibit matched twins (firms with nearly equal D but opposite internal geometry) selected purely on their description vectors, with no reference to returns. This establishes the precondition: the aggregate is hiding structure.

Second, the hidden structure carries consequence. In a specification frozen and hashed before returns in the pre-2016 panel were examined for this shape hypothesis (the panel itself had previously been analyzed for other questions), concentration conditional on amount predicts subsequent returns. Because D enters negatively and C₁ positively, the most negatively-priced state at a given amount of disagreement is broadly distributed disagreement, not a single-description rupture. The result holds in the simplest joint regression of D and C₁, and survives a short-term-reversal control, but it is concentrated in small, illiquid stocks and is weak-to-absent among liquid names (Section 7).

Third, the relation recurs. Under a separately frozen temporal protocol, the relation is positive across the whole chronology with no material sign reversal and no single-year dependence, and reproduces in a later era at about half magnitude with time-varying strength.

Alongside these we report an informative failure. A semantically natural decomposition (is it specifically the price view conflicting with fundamentals, or with ownership?) looked strong in exploratory work and did not survive the frozen test, while the label-light concentration measure did. We read this not as a defeat but as evidence about what kind of decomposition is durable: how disagreement is organized proved more durable, in this sample, than which named components conflict.

Methodological stance. The disagreement/anomaly literature is unusually exposed to specification search. We therefore adopt a pre-registration discipline throughout: the descriptive stage touches no returns; the consequence specification (reads, normalization, concentration definition, cap, controls, inference, pass/null interpretation) is fixed and hashed before returns in the pre-2016 panel are examined for the concentration/shape hypothesis (the panel itself having previously been analyzed for other questions); the temporal protocol is fixed and hashed separately; and each is run once. The claim-status timeline in Appendix A records the full arc, including the abandoned steps, so the reader can see that the surviving specification was not selected on its return performance.

What we do not claim. We do not identify an economic mechanism, we do not construct a tradeable factor, and we do not offer prospective confirmation. The pre-2016 window, though frozen with respect to this hypothesis, had been examined earlier for other questions; the evidence is retrospective / secondary historical validation. Section 9 is explicit about each limit; Section 10 separates the remaining upgrades, mechanism and confirmation, that would strengthen it (economic magnitude having been characterized in Section 7).

The idea that disagreement predicts returns is well established. Diether, Malloy and Scherbina (2002) and Johnson (2004) document the negative relation between analyst forecast dispersion and subsequent returns; a large body of work refines the measure and debates its interpretation (risk versus mispricing, conditioning on sentiment, optimism, or future profitability). This literature is about the amount of disagreement.

Closest to our question is work on the shape of the belief distribution. Hardouvelis, Karalas and Vayanos (2026) decompose the distribution of analysts’ forecasts into an intensity dimension (range) and a polarization dimension (kurtosis) and show both matter for returns and ownership, testing one conditional on the other. This is a genuine “shape beyond amount” precedent. It differs from our object on three axes: its coordinates are homogeneous beliefs (many analysts forecasting the same quantity) where ours are heterogeneous descriptions of one firm; its shape statistic is distributional kurtosis (polarization), nearer to a two-sided split than to one-vs-many concentration; and its ownership dispersion is a Herfindahl of holdings breadth, not of the disagreement itself.

Bali, Kelly, Mörke and Rahman (2026) construct Machine Forecast Disagreement as the standard deviation of return forecasts across machine-learning “investor-models.” This is a deterministic, model-based disagreement measure with excellent coverage, and it strongly negatively predicts returns, but it is dispersion amount; a full-text reading finds no concentration or higher-moment test conditional on that amount.

The composite-anomaly literature (Stambaugh and Yuan, 2017) combines many anomaly signals into a mispricing score by averaging ranks, the consensus/mean direction across signals. This is, if anything, the opposite construction to ours: a stock on which all descriptions agree has a strong composite score and low D; a stock whose descriptions split has near-zero average score and high D. Higher-moment work on the return distribution (implied skewness, kurtosis) and recent work on the structural organization of accounting information are conceptually adjacent examples of “structure beyond a first or second moment,” but neither constructs concentration of cross-characteristic dispersion.

Our contribution is therefore specific: not “shape matters” (known), but that the concentration of heterogeneous cross-characteristic dispersion, conditional on its total amount, carries incremental return information. We do not identify a direct equivalent in full-text checks of the closest work, and we do not claim its non-existence.

3. Data

Universe. We use Sharadar’s survivorship-free US equity database, which retains delisted names. We keep domestic common stock on NYSE, NASDAQ, and NYSE-MKT, with month-end price ≥ $5, and require ≥100 valid names per monthly cross-section. All information is point-in-time: fundamentals are lagged to their reporting datekey, and 13F holdings are lagged 50 days beyond the filing date. We study two non-overlapping periods: a strictly pre-2016 window used for the primary consequence test, and a later development era (2016–2026, with the homogeneous all-ten panel available from March 2018 onward) used to develop the descriptive construction and, separately, to test temporal recurrence.

The ten descriptions (“reads”). Each is a firm characteristic computed point-in-time, winsorized at 1/99 and standardized cross-sectionally (mean 0, unit variance) each month:

# Read Definition
1 Momentum log price ratio, t−12 to t−1
2 Gross profitability gross profit / assets
3 ROA net income / assets
4 Earnings surprise (SUE) Δ4-quarter EPS / rolling std
5 Sales growth 4-quarter revenue growth
6 Accruals −(net income − free cash flow) / assets
7 Asset growth −(4-quarter asset growth)
8 Value −log(price/book)
9 Institutional breadth Δ change in count of 13F holders (lagged 50d)
10 Insider net buying signed trailing-6-month insider purchases

These are heterogeneous in economic meaning and span several data families, but they are not assumed statistically or source-independent: several fundamentals (reads 2–7) are transforms of overlapping financial-statement information. The observed mean pairwise absolute correlation is ≈ 0.12: low empirical redundancy, not informational independence. This distinction is load-bearing and we return to it in Section 9.

Controls. Idiosyncratic volatility (residual from a 60-day market-model), size (log market cap), value (−log P/B), and momentum, each standardized cross-sectionally.

4. Measurement and the geometry of disagreement

For a firm with n present reads xᵢ, let x̄ = (1/n)Σxᵢ and dᵢ = xᵢ − x̄. We use:

The decomposition MSQ = (1/n)Σxᵢ² = x̄² + D cleanly separates the firm’s consensus level from its internal dispersion; this paper concerns the internal structure of D, holding the consensus term aside.

Concentration conditional on amount. C₁ mechanically co-moves with D (higher-dispersion states tend to be more rupture-like). Because our question is explicitly conditional (at the same amount of disagreement, does concentration matter), we also form a residual, C₁* = min(C₁,0.80) − Ê(C₁|D), where Ê(C₁|D) is a monotone (isotonic) fit of expected capped concentration on D, estimated once on the later-era development sample and applied unchanged to the test sample. The 0.80 cap is a pre-test guard motivated in Section 5. We stress that C₁* is a conditioning device, not a new statistic: the plain joint regression of raw D and raw C₁ (Section 7) makes the same point without it.

Homogeneous observer system. For the primary test we require all ten reads present, so every firm-month is represented by the identical ten-description system and D, C₁, and C₁* have the same meaning across observations. This costs sample (13F availability restricts the pre-2016 primary to ~26 months); we accept it because a homogeneous instrument is preferable to a larger sample whose observer composition drifts. A ≥6-read construction, with concentration residualized within read-count strata, is carried as robustness.

Pre-registration. The consequence specification and the temporal protocol were each written, hashed, and only then run once (Appendix C). This is the paper’s principal defense against the garden of forking paths, and the reason the descriptive stage (Section 5) is reported before any return regression.

5. Is the aggregate lossy? (no returns)

Before testing consequence we establish that D genuinely hides structure. On 281,538 later-era firm-months with all ten reads present (price ≥ $5), we examine how much C₁ varies at fixed D.

Within D deciles, the interquartile and 5–95% ranges of C₁ are wide throughout the middle of the distribution: for representative mid-D deciles the 5–95% span of C₁ runs from roughly 0.26 to 0.78. At nearly identical aggregate disagreement, firms occupy geometries from near-uniform to near-rupture.

Matched twins. Selecting purely on geometry (no returns), we can pair firms at D ≈ 0.30 with opposite concentration. A representative pair:

Firm D C₁ Read vector (z-scores)
A 0.30 0.86 one read at −1.7 (earnings surprise), all others ≈ 0
B 0.30 0.17 many reads at ±0.5–0.8, balanced positive and negative

Same amount of disagreement; A is a single-description rupture, B is distributed inconsistency. This is the paper’s central picture (Figure 1 schematic; Figure 2 the real pair).

Figure 1
Figure 2

One shape axis. Concentration and polarity are nearly collinear in this sample (corr ≈ −0.86): a concentrated state is almost mechanically one-sided. We therefore carry a single shape axis, concentration, and note that a separate polarity test is not independent information here.

Confound audit. We ask which read owns the maximum squared deviation in high-C₁ states. Through moderate concentration the answer is spread across five or six reads; but the extreme tail (C₁ > 0.8) is dominated by two structurally sparse, spiky reads: insider net buying (43%) and sales growth (15%), whose intermittent large values mechanically produce single-read ruptures. This is a measurement artifact, not signal, and it motivates two pre-test guards used in Section 7: cap C₁ at 0.80, and, as robustness, recompute the geometry after dropping the two sparse reads.

Crucially: we did not select a shape statistic because it predicted returns; we first established that the aggregate is lossy, then asked whether the loss is consequential.

6. A semantic decomposition and its failure

A natural hypothesis is that specific named conflicts drive any effect, e.g. the price view (read 1) conflicting with the fundamental block (reads 2–8) or with the ownership block (read 9). Define PF = (price − fundamental-mean)² and PO = (price − ownership)². In the later-era development sample, PF and PO entered strongly. We froze a test of whether they add beyond total dispersion D and consensus, and ran it on the pre-2016 sample. They did not: entered jointly with D, consensus, and controls, PF t = −0.65, PO t = −0.66, joint Wald p = 0.45.

A predefined semantic decomposition did not replicate out of sample. This motivated no change to the dispersion measure; instead we asked the more general, label-light question of how a fixed amount of deviation is distributed across coordinates. That named attribution failed while label-light organization survived is itself informative, and makes the surviving result harder to interpret as a simple continuation of the original semantic decomposition.

7. Does the hidden shape carry consequence? (frozen test)

Specification (frozen and hashed before returns in the pre-2016 panel were examined for this shape hypothesis; the panel itself had previously been analyzed for other questions). Monthly Fama–MacBeth regressions of 21-day forward returns on the shape variable and controls (idiosyncratic volatility, size, value, momentum), Newey–West(3) inference, ≥100 names/month, price ≥ $5, forward returns winsorized 1/99 per month and truncated at the window end so no month looks past the outcome boundary.

Table 1. Plain joint specification (raw D and raw C₁ in one regression, no interaction).

Sample D coef (t) C₁ coef (t)
pre-2016 (n=26) −0.00313 (−2.66) +0.00188 (+2.79)
pre-2016, C₁ capped 0.80 −0.00309 (−2.64) +0.00182 (+2.75)
later era 2018–2026 (n=100) −0.00284 (−4.04) +0.00080 (+1.99, marginal at 5%)

Amount and concentration enter with opposite signs, and each retains explanatory power conditional on the other and the controls; the cap is inert (capped vs raw C₁, t = +2.75 vs +2.79). In the later era the concentration coefficient is positive but marginal by the conventional two-sided 5% threshold.

Secondary interaction diagnostic. Adding D×C₁ leaves the main coefficients in the same neighborhood, with interaction t = +1.47 (pre-2016) and t = +2.24 (later era), mixed across periods and not part of the core claim.

Table 2. Residualized and robustness constructions (pre-2016 primary).

Specification C₁ / C₁* (t)
Residualized C₁\*, all-ten +3.00
Robustness: ≥6-read, within-count residualization +4.32
Robustness: eight-read (drop insider, sales growth), geometry recomputed +3.58
D coefficient (residualized primary) −2.44

The residualized form conditions on D nonlinearly (isotonic) rather than linearly and reaches the same conclusion; raw C₁ and C₁* are the same relationship differently parameterized, so the agreement is a robustness check, not independent evidence. Both robustness constructions agree in sign, and the D×C₁* interaction is insignificant in the primary. The eight-read construction is the direct test of whether the effect is an artifact of the two sparse tail-dominating reads: it drops insider and sales growth entirely and recomputes D and C₁ on the remaining eight, yet the effect strengthens (t=+3.58), so it is not driven by those two channels. (The robustness specifications running stronger than the plain later-era C₁ has a benign source: the ≥6-read construction spans more months, and C₁* conditions on D nonlinearly; they are not independent votes.)

Short-term reversal. Because momentum enters as t−12 to t−1, one-month reversal is not among the frozen controls, and a natural concern is that a concentrated single-read state is really a recent extreme price print. Adding short-term reversal (STR, the past-21-day return) as an additional control leaves the shape coefficient essentially unchanged (pre-2016 C₁* t = +2.97 vs +3.00, plain C₁ t = +2.77 vs +2.79; later-era C₁* t = +2.21 vs +2.12), while STR itself is insignificant (t = +0.36 and −0.68). The effect is not short-term reversal in disguise. The direct evidence is that C₁ and STR are essentially uncorrelated in the cross-section (mean correlation ≈ 0.00 to 0.04 in both periods), so concentration cannot be reversal relabeled; this is consistent with C₁ being direction-agnostic (it encodes concentration, not the sign of any read) and with the high-C₁ tail being dominated by non-price reads. We note that short-term reversal is itself weak in this ≥$5 universe (STR-alone t = +0.44 and −0.52), which limits how much a reversal control can absorb; the near-zero correlation settles the question regardless.

Sign and interpretation. Because D is negative and C₁ (concentration) positive, the most negatively-priced state at a given amount of disagreement is high-D, low-C₁: broadly distributed disagreement across many descriptions, rather than a single-description rupture. Economically, this is an association between a measured state and subsequent returns; it does not establish that investors perceive “incoherence,” nor that a latent known characteristic does not generate both.

Economic magnitude. To express the estimand in interpretable units we form a pre-specified conditional sort (frozen separately; Appendix C): within D quintiles we sort firms into C₁* terciles, pool the bottom (distributed) and top (concentrated) terciles across D quintiles (holding aggregate dispersion approximately fixed), and equal-weight each leg’s 21-day forward return. We report the distributed-minus-concentrated spread (a zero-net-investment return difference by construction; we do not claim market-beta neutrality).

Sample R_distributed R_concentrated distributed − concentrated (t)
pre-2016 (26 mo) −137 bps −41 bps −96 bps / 21d (−3.94)
later era 2018–2026 (99 mo) −74 bps −29 bps −45 bps / 21d (−3.56)

Holding aggregate disagreement roughly fixed, distributed states earn on the order of 45 bps per 21-day period less than concentrated states in the longer later sample (≈96 bps in the thin pre-2016 window, consistent with the larger Fama–MacBeth slope there). This is a characterization of the Stage-1 estimand’s economic size, deliberately not a trading claim: the legs are equal-weighted and un-costed, and by pre-registration this sort neither confirms nor disconfirms the frozen concentration coefficient. A coarse portsort and a Fama–MacBeth slope are two views of the same association, on the same data.

Weighting and liquidity. The magnitude is an equal-weighted, small-name phenomenon. Value-weighting the legs by market cap, or restricting to the top half by market cap, attenuates the spread to near zero and insignificance (later era: value-weighted +11 bps, t = +0.47; liquid equal-weighted −6 bps, t = −0.58; liquid value-weighted +17 bps, t = +0.71; pre-2016 value-weighted −32 bps, t = −1.07). The economic footprint therefore lives disproportionately in small, illiquid stocks, and the equal-weighted ≈45 bps is not representative of large or liquid names. This concentration is not only economic: the Fama–MacBeth slope itself weakens sharply in liquid names. Restricting the cross-sectional regression to the top half by market cap, the C₁* coefficient falls to roughly a quarter of its full-sample size and loses significance in the better-powered later era (later-era liquid t = +0.57 vs +2.12 full; pre-2016 liquid t = +1.65 vs +3.00, n = 26), while in the small-cap half it is if anything stronger (t = +2.14 and +3.49). The finding is therefore a small-and-illiquid-stock phenomenon in both its statistical and its economic form, not a general cross-sectional relationship whose magnitude merely happens to be larger in small names.

Microstructure artifact or mispricing? Because the effect lives in small, illiquid stocks, we pre-registered a test (frozen; Appendix C) to distinguish a genuine slow mispricing from a microstructure artifact such as bid-ask bounce or stale prices. It is not a microstructure artifact: the C₁* slope survives, and in fact strengthens, when the outcome skips the first five trading days (later-era t = +3.40 on the t+5→t+26 return vs +2.12 contemporaneously, gap coefficient 126% of baseline), whereas bid-ask and stale-price effects decay within days. But we cannot attribute it to limits to arbitrage either: the C₁* × idiosyncratic-volatility interaction (IVOL being the standard arbitrage-cost proxy) is insignificant (t = +0.76), and the slope does not survive excluding the bottom market-cap quintile (later-era t = +0.55). By the pre-registered rule the result is mixed: a genuine, non-artifact predictive signal confined to microcap stocks, of unresolved economic mechanism. We do not claim a limits-to-arbitrage channel.

Reducibility to the characteristics (the decisive limitation). The frozen test controls for amount D and four standard factors. A sharper question is whether concentration adds beyond the ten underlying reads themselves, since C₁ is built from them. Adding all ten reads as linear controls, C₁*‘s return predictability is essentially fully absorbed in the primary later-era sample (t falls from +2.12 to −0.06), surviving only in the thin pre-2016 window (t = +2.87 vs +3.00). By the pre-registered rule (primary = later era; Appendix C) this is a reducible result: C₁*’s return signal is the reads’ own linear return signals repackaged through the concentration statistic, not an incremental geometric effect that a linear characteristic model would miss. A relationship that holds in 26 months and vanishes in 100 is not robust. A pre-registered nonlinear extension of this kill reaches the same verdict: adding the ten reads’ squares and all 45 pairwise products as controls, C₁* in the primary later sample is t = +0.62 (coefficient 24% of the linear baseline), so its return signal is reducible to the reads even nonlinearly. (The pre-2016 window survives even this, t = +3.07, but at n = 26 it is underpowered and we do not bank it.) We therefore do not claim that concentration carries incremental return information beyond the underlying characteristics, linearly or nonlinearly; on this evidence the return question is closed. The concentration geometry remains a robust descriptive fact (Section 5); its return predictability is not distinct from the underlying characteristic signals.

8. Temporal recurrence

We fixed a temporal protocol (instrument unchanged, only time segmentation varied), hashed it, and ran it once (Appendix C).

Non-overlapping blocks (all-ten). Positive throughout, no material sign reversal:

Block months C₁* (t)
pre-2016 26 +3.00
2018-03..2018-11 9 +0.48
2018-12..2021-05 (COVID-spanning) 30 +0.04
2021-06..2023-11 30 +3.06
2023-12..2026-05 30 +1.01

Leave-one-calendar-year-out (later era). All eleven folds positive, t = +1.6…+2.8; no single year is individually necessary. We flag that this reassurance is limited: because the later-era signal concentrates in the 2021–2023 regime, dropping any single year leaves that regime largely intact, so “all folds positive” is closer to mechanical than to independent corroboration.

Rolling 36-month windows (stepped 12 months). Near zero over 2017–2020, then rising and remaining positive from ~2020 (peak t = +2.87); one window is trivially negative (t = −0.05). Overlapping windows are a diagnostic path, not independent tests (Figure 3).

Figure 3

Later-era reproduction. Pooled later-era all-ten C₁* t = +2.12 (≥6-read companion +2.77; plain C₁ +1.99), same sign at roughly half the pre-2016 magnitude.

The effect recurs but is not constant. Two non-overlapping strong periods (pre-2016 and 2021–23), no single-year dependence, and reproduction in a later un-examined era place it well beyond a one-period accident; the near-zero 2018–2021 stretch and the magnitude attenuation place it well short of a constant law. We record the regime dependence (why the consequence of the same measured geometry varies materially over time while keeping its sign) as an open question, not something we model, to avoid post-hoc correlation against a thousand available macro variables. One specific caution belongs here: the strong 2021–23 window overlaps a high-dispersion macro regime, and while the specification controls for idiosyncratic volatility cross-sectionally and the leave-one-year-out result survives dropping any single year (2020 included), the cross-sectional control does not absorb time-series volatility regimes. We therefore cannot exclude that the relation strengthens in high-dispersion periods, and we do not treat the later-era reproduction as fully independent of that regime.

9. What is and is not established

Established. (i) Aggregate cross-characteristic dispersion is descriptively lossy: at fixed D, real firms occupy sharply different concentration geometries. (ii) Conditional on amount, concentration carries incremental information about subsequent returns in a frozen pre-2016 test, present in the simplest joint specification and robust to two re-parameterizations and to a short-term-reversal control, but concentrated in small, illiquid stocks, weak-to-absent among liquid names, and (in the better-powered later sample) reducible to the underlying characteristics when they are added as linear controls (Section 7). (iii) The relation recurs, weaker and regime-varying, in a later era whose concentration–return relationship had not been examined. (iv) Within its microcap scope the effect is not a microstructure artifact: it survives, and strengthens under, a five-day implementation gap (Section 7).

Not established. A causal or behavioral mechanism (a pre-registered test found the effect is not a microstructure artifact but is confined to microcaps with no significant arbitrage-cost gradient, so limits to arbitrage is not established either); that its return predictability is incremental to the underlying characteristics (in the better-powered later sample it is absorbed once the ten reads are added as controls, linearly and under a nonlinear squares-and-products expansion alike, Section 7); that investors perceive “disagreement” or “incoherence”; that concentration is a distinct priced risk factor; that any single description is an informationally independent observer (the reads are low-redundancy, |corr| ≈ 0.12, not independent); generality across the size distribution (Section 7: among liquid names both the Fama–MacBeth slope and the equal-weighted magnitude weaken sharply and lose significance, so the finding is a small-and-illiquid-stock phenomenon, its presence in large caps and its tradability both unestablished); and, most importantly, prospective confirmation. The pre-2016 window, though frozen with respect to this hypothesis, had been examined earlier for other questions, so the evidence is retrospective / secondary historical validation rather than a clean out-of-sample experiment. An omitted known characteristic that generates both distributed disagreement and low returns remains possible.

10. What would strengthen the evidence

Economic relevance has now been characterized through the preregistered conditional sort in Section 7; this does not alter confirmation status. Two live frontiers remain, of different kinds, not points on one confirmation ladder:

Each requires its own pre-registration. We report the economic-magnitude characterization only; mechanism and confirmation remain deliberately untested here.

11. Conclusion

Conventional dispersion measures the amount of cross-characteristic disagreement but not its internal organization. That distinction is empirically material: firms with nearly identical aggregate dispersion can have sharply different characteristic geometries, and a concentration measure defined independently of returns carries incremental information conditional on dispersion in a frozen historical test, recurring at smaller magnitude in a later period. More distributed disagreement is associated with lower subsequent returns. A semantic decomposition into specific price–fundamental and price–ownership conflicts failed out of sample, while the label-light geometric decomposition survived, suggesting that, in this sample, the organization of dispersion was more durable than attribution to particular characteristic families. Why these states differ economically remains unresolved.


Appendix A: Claim-status timeline (the discovery arc)

To make specification search auditable, the full sequence, including abandoned steps:

  1. Price-only disagreement → known idiosyncratic volatility. A first construction using five price transforms reconstructed the low-volatility anomaly (zero alpha after standard factors). Closed: the reads were one surface, violating heterogeneity.
  2. Total heterogeneous dispersion D → out-of-sample supportive, but known. D across the ten reads is priced out of sample under Newey–West(3) inference (t ≈ −2 to −3 with controls), but dispersion-is-priced is an established effect, not a discovery.
  3. Named semantic structure (PF/PO) → frozen failure. Price-versus-fundamental/ownership conflict looked strong in development and did not add beyond D out of sample (Section 6).
  4. Concentration conditional on dispersion (C₁*) → frozen secondary validation. The label-light geometric decomposition survived its frozen test (Section 7).
  5. Temporal recurrence → supportive, regime-varying (Section 8).

The surviving specification (step 4) was frozen and hashed before its pre-2016 concentration–return relationship was examined; steps 1–3 are the path that produced the hypothesis, not evidence selected on its return performance.

Appendix B: Variable construction details

Reads standardized cross-sectionally each month after 1/99 winsorization. D, C₁, P as in Section 4. Ê(C₁|D) fit by isotonic (monotone) regression on the later-era development sample, per read-count stratum for the ≥6-read robustness, with endpoint-clamp extrapolation; frozen and applied unchanged. Forward returns are 21-trading-day log returns, truncated at the sample boundary. Inference is Fama–MacBeth with Newey–West(3) on the monthly coefficient series.

Appendix C: Frozen specifications

Artifact Role SHA-256 (prefix)
PREREGISTRATION.md original PF/PO confirmation 0299723b…
run_confirmation.py PF/PO harness f9d9a08b…
PREREG_SHAPE.md frozen consequence spec 9cd845b7…
run_shape_consequence.py consequence harness 381f9497…
PREREG_SHAPE_TEMPORAL.md frozen temporal protocol 43da5237…
run_plain_joint.py Table 1 transparency spec f54c67a6…
PREREG_PORTFOLIO.md frozen Stage-2 magnitude sort 0280ab8f…
run_portfolio.py economic-magnitude harness 38a499ae…
PREREG_LTA.md frozen artifact-vs-mechanism test b9e5d34e…
run_lta.py limits-to-arbitrage / gap harness 15f95ee8…
PREREG_REDUCE.md frozen reducibility kill c6ab1b1b…
run_reduce.py reducibility harness (linear) 4cc153ae…
PREREG_NONLIN.md frozen nonlinear reducibility kill db22bb13…
run_nonlin.py nonlinear reducibility harness e55eb5cf…
RESEARCH_NOTE.md frozen Stage-1 note 902543b1…

Descriptive and audit code (shape_descriptive.py, dominant_read_audit.py, reconcile_D.py, geometry_compare.py) and figures are in the project record.