Working paper — theory. Draft for circulation.
Abstract
Generative models have collapsed the cost of producing candidate reasoning while leaving the cost of establishing warrant largely intact. The standard response, that attention is now the scarce resource, is true and insufficient: with costless discard and any informative pre-screen, a larger candidate pool weakly improves outcomes. We locate the binding constraint elsewhere. What abundance consumes is not attention in general but the cost of disposal: the resources required to establish that a candidate does not warrant reliance. Machine generation raises this cost directly, because fluency, structure, apparent citation, and calibrated-sounding hedging are precisely the surface features that once served as cheap disposal cues.
We then identify a mechanism specific to machine-generated candidates and independent of any human cognitive limitation. Candidates drawn from correlated generative lineages carry bounded joint evidential content: with intraclass correlation ρ, effective independent support satisfies n_eff ≤ 1/ρ regardless of volume, while disposal cost grows linearly in volume. Evidential value therefore saturates while cost does not, which produces a finite optimal candidate load with no appeal to overload, fatigue, or bounded rationality. The ceiling is a property of market structure rather than of psychology: adoption raises the number of candidates without raising the number of lineages.
Four further results follow. An identity, coverage × depth = V_eff/C, shows that institutions cannot hold both breadth and thoroughness fixed under rising load and predicts two distinguishable degradation signatures. In a loss system rather than a queue, saturation produces a regime change in which institutional discriminability converges to the discriminability of the triage function, independent of validation quality; this yields enclosure of collective reasoning without ownership of any record, and it is worsened when triage and generation share generative lineage. Where contributors and validators are drawn from the same population, a single technology shock raises candidate supply and lowers validation supply simultaneously, so admission friction and validation credit are complements whose relative necessity depends on whether contributors validate. Finally, greedy value-maximising triage makes the set of never-inspected sources absorbing, so distributional exclusion is a consequence of optimality rather than of bias, and a reserved exploration budget is the standard consistency fix rather than an equity concession.
We argue the object at risk is not a common-pool resource. It is a non-rival good whose use requires a rival complement, and the externality runs through that complementarity. We give ten numbered predictions with refutation conditions. None is yet tested.
1. Introduction
1.1 The problem, stated carefully
Institutions that produce knowledge run two distinct processes. One generates candidate propositions, explanations, objections, and inferences. The other establishes which of them deserve to carry consequential weight. Until recently these processes were coupled by a shared input: both required competent human labour, so the arrival rate of candidates was bounded by roughly the same resource that bounded the capacity to assess them.
Generative models decouple them. The marginal cost of producing a fluent, structured, plausible candidate has fallen by orders of magnitude. The marginal cost of establishing warrant has fallen much less, and in some domains not at all, because it terminates in experiment, measurement, proof, replication, or contact with an outcome that arrives on its own schedule.
The obvious inference, that this is an attention-scarcity problem, is the one we want to resist. Attention scarcity is real, was described by Simon (1971), and was formalised by Sims (2003). It is also insufficient as a theory, for a reason developed in Section 2: under costless discard, more candidates cannot hurt. Any account of harm from abundance must therefore locate a cost that abundance imposes before the decision to discard, or a respect in which the discarded material was never separable from the retained material at low cost. This paper is about that cost.
1.2 What we claim
The paper makes four theoretical claims and one measurement claim.
C1. The binding constraint is the cost of disposal. Define disposal cost as the resources required to establish that a candidate does not merit reliance and to remove it from further consideration without residue. Machine generation raises disposal cost, because the features that historically permitted cheap disposal — incoherence, malformed structure, missing citations, register errors, absent hedging — are the features generative models eliminate first. AI lowers the cost of production and raises the cost of disposal. That conjunction, not abundance alone, is the shock.
C2. Machine corroboration has a diversity ceiling. Candidates drawn from correlated generative processes contribute bounded joint evidential content. With intraclass correlation ρ induced by shared training corpora, retrieval corpora, prompts, or fine-tuning ancestry, effective independent support is bounded by 1/ρ no matter how many candidates arrive. Disposal cost grows in the count; evidential value saturates in the effective count. This asymmetry alone yields a finite optimum, without invoking human limitation, and it identifies generative diversity as an epistemic public good whose supply is determined by foundation-model market structure.
C3. At saturation, epistemic authority relocates to triage. Validation systems are loss systems, not queues: unexamined candidates are dropped, not delayed. Consequently delay does not diverge at capacity; coverage collapses, and institutional discriminability converges to the discriminability of whatever function performs the dropping. The marginal return to improving validation quality falls toward zero while the marginal return to improving triage rises toward one. Whoever supplies triage therefore governs reliance, without owning any record. When triage and generation share generative lineage, generator errors become systematically invisible rather than randomly missed.
C4. The failure is two-sided where contributors validate. If generation and validation draw on the same time budget, and validation is credited at less than its social value, a technology shock that raises the productivity of generation relative to credited validation raises C and lowers V simultaneously. Friction and credit are then complements. Which one an institution needs depends on whether its contributors are drawn from its validating population, which supplies a principled distinction between human contributors and autonomous agents.
M1. Congestion is identifiable without knowing an institution's preferences. Rather than a scalar "warranted reliance," we recommend the achievable frontier between misreliance and forgone reliance, recovered from a proper scoring rule and its calibration/resolution decomposition. Whether the frontier shifts outward is a preference-free empirical question. Whether a scalar R falls is not, because it depends on where on the frontier the institution chose to sit.
1.3 What we do not claim
We do not claim that AI necessarily degrades collective epistemics. The theory is conditional, and it identifies the conditions. Automated triage, retrieval, provenance extraction, and review assistance can raise effective validation capacity and lower disposal cost; there is direct evidence that machine feedback improves human review rather than replacing it. The theory predicts deterioration only where the growth of disposal demand outruns the growth of disposal capacity, and it says which institutional properties determine that race.
We do not claim to have established any of the results empirically. Every proposition below is a theorem about a model or a comparative static within one. Section 13 states what would refute each, and Section 14 states what we think is most likely to be wrong.
We do not claim that abundance is bad. Sections 6 and 8 both contain regions in which additional candidates strictly increase warranted reliance. The interesting quantity is the threshold, and the interesting design question is how to move it.
1.4 Novelty, stated narrowly
The general observation that AI generates plausible artifacts faster than institutions can verify them is no longer novel; by 2026 it has been stated directly, including with the term "epistemic pollution" and with structured claim representations proposed as the design response (Ma 2026; see also Levy 2018 on polluted epistemic environments). The observation that senders impose congestion externalities on scarce receiver attention is older still and already formalised in economics, with results on triage, message pricing, and gatekeeping (Van Zandt 2004; Anderson and de Palma 2009). The observation that the scholarly literature outgrew its appraisal capacity predates generative models by fifteen years (Bastian, Glasziou and Chalmers 2010).
What we believe is new is the specific set: disposal cost as the binding quantity and AI's distinctive effect on it (C1); the diversity ceiling and its market-structure reading (C2); triage dominance under saturation as a mechanism of enclosure, and lineage-sharing as a mechanism of systematic invisibility (C3); the two-sided wedge with the contributor-population distinction (C4); and the frontier-shift identification strategy (M1). Section 11 states the relation to existing formal traditions in detail, including places where existing results cut against our institutional preferences.
2. Why abundance should help, and what has to be true for it not to
2.1 The free-disposal baseline
Consider an institution choosing among N candidate answers to a question. It may inspect them in any order and stop at any time, and it retains the best it has found. This is Weitzman's (1979) optimal search problem, and its central property is monotonicity: the value of the search problem is weakly increasing in N. Adding a candidate can never hurt, because the institution can ignore it.
The same monotonicity holds under weaker conditions. If the institution has any pre-screen whose signal is even slightly informative about candidate quality, and if applying the pre-screen is free, then a larger pool improves the quality of the item selected at any fixed inspection budget, because the pre-screen's top-ranked item is drawn from a better order statistic.
This is not a marginal objection. It is the reason "there is more information now" cannot by itself be a theory of epistemic harm, and it is why decades of information-overload results have remained contested: the effect is context-dependent because in many contexts the free-disposal logic simply dominates.
2.2 The three assumptions that must fail
Free-disposal monotonicity requires three things. Epistemic congestion is exactly the failure of one or more of them.
A1. Disposal is free. Ignoring a candidate must cost nothing. In practice it costs something: the candidate must be received, recognised as a candidate, checked for duplication against the existing record, attributed, and either routed or dismissed. Where the candidate is a proposition rather than a product, dismissal often requires reading enough to determine that it is not novel, not contradictory, and not consequential — which is a substantial fraction of the cost of evaluating it.
A2. The pre-screen is free and its precision is exogenous. Ranking must not consume the capacity that inspection needs, and ranking accuracy must not depend on the volume or composition of the pool. Both fail when candidates are optimised, deliberately or as a by-product of training, for the features the pre-screen uses.
A3. The pool's right tail improves with size. Adding draws helps because the maximum of a larger sample is larger. This requires the draws to be informative about distinct things. It fails when additional candidates are near-copies in the evidential sense even when they are distinct in the textual sense.
The rest of this paper is organised around these failures. Section 6 develops A3, which we take to be the most important and the most specific to machine generation. Section 7 develops A1 and the coverage/depth trade-off it forces. Section 8 develops A2 and its institutional consequence.
2.3 The cost of disposal, and why AI raises it
Definition 1 (Disposal cost). For candidate j, the disposal cost c_j is the expected resource expenditure required to establish that j does not warrant reliance and to remove it from further consideration, including duplication checking, provenance attribution, and the recording required to prevent re-litigation.
Two properties matter. First, c_j is bounded below by a strictly positive constant across a wide range of realistic institutions, because the minimum operations are irreducible. Second, c_j depends on the candidate's surface features, because disposal in practice proceeds by cheap cue before it proceeds by evidence.
The second property is where generative models act. Historically, a large fraction of candidate reasoning could be disposed of by cues that correlated with unreliability at near-zero inspection cost: incoherence, absent structure, missing or malformed references, register mismatch, absence of hedging, failure to anticipate the obvious objection. These cues were serviceable because producing their absence was expensive. Language models produce their absence at negligible cost, and they do so first, because surface fluency is what the training objective most directly rewards.
This gives the paper's central asymmetry in one line:
Generative models reduce the cost of producing a candidate and increase the cost of disposing of one.
The companion literature's framing — that AI makes generation cheap while verification stays costly — is correct but under-specifies the mechanism. Verification of the retained material was always costly. What changed is the price of getting rid of the rest.
In signal-detection terms, developed in Section 7, this is an increase in the evidence-signal mean of invalid candidates, μ_F, holding the mean for valid candidates fixed. Discriminability at any given inspection depth falls, and reaching the features that still discriminate requires deeper inspection.
3. Primitives
3.1 Candidates and effective load
Definition 2 (Candidate reasoning object). A candidate is a proposition or inferential relation that could alter a community's beliefs or actions and whose warranted status is unsettled in that community's record. Factual claims, causal explanations, objections, model outputs, forecasts, proposed inferences, and challenges to existing claims all qualify. Restatements of settled material do not.
Raw count N is not the theory's independent variable. Define per-candidate load as
ℓ_j = c_j^{base} + s_j · k_j · (1 − d_j) · (1 + δ_j)
where c_j^{base} > 0 is the irreducible handling and disposal floor, s_j is stake, k_j is direct verification difficulty, d_j ∈ [0,1] is cheaply detectable duplication, and δ_j is dependency complexity, the number of downstream commitments that change if j is accepted or rejected. Effective load is C = Σ_j ℓ_j.
Two departures from the multiplicative forms used elsewhere are deliberate. The additive floor is required because a zero-stake, unattributed candidate still consumes triage; a purely multiplicative form implies it is free. And correlation does not appear here. Correlated candidates cost roughly what independent ones cost to handle; what they lack is evidential contribution. Folding correlation into a cost multiplier conceals the asymmetry that Section 6 turns into the paper's main result. Correlation belongs on the benefit side.
The functional form is a modelling proposal, not a finding. Its only load-bearing features are the strictly positive floor and the exclusion of correlation.
3.2 The validation pipeline
An institution receives C and produces reliance decisions through three stages.
- Handling. Every candidate incurs
c^{base}. This expenditure occurs before any decision and cannot be avoided by declining to inspect. - Triage. A function
Tassigns each candidate a priority using a signal of precisionτ. Coverageγ ∈ [0,1]is the fraction routed to inspection. - Inspection. Inspected candidates receive effort
eand yield an evidence signal whose discriminability isd'(e).
Gross validation capacity V is resource per unit time available for stages 2 and 3. Effective capacity is
V_eff(C) = [ V − a·C − b·C^η ]₊ , a > 0, b ≥ 0, η > 1
where a·C is the handling floor aggregated over arrivals and b·C^η collects superlinear overhead: pairwise comparison and reconciliation, duplicate detection across a growing record, dependency reconstruction, coordination among validators, and context switching. Whether b > 0 and what η is are empirical questions, and we flag now that no study we know of has measured them. The theory's core results in Sections 6 and 8 do not require b > 0. They require only a > 0, which is Definition 1.
Remark. The a·C term is what breaks free disposal. It is the formal content of assumption A1's failure, and it is the reason a purely bounded specification like V_eff = V/(1+(C/K)^η) should be avoided: that form has no interpretable handling floor and its implied g is eventually concave, which contradicts the superlinear-overhead story it is usually introduced to represent.
3.3 Reliance and answerability
Definition 3 (Reliance). a_j ∈ [0,1] is the consequential weight the institution places on candidate j: the extent to which decisions and further inference are permitted to depend on it.
Definition 4 (Warrant). w_j ∈ [0,1] is the weight j would deserve given full access to the evidence, provenance, independence structure, and outstanding objections. w_j is not directly observable. On resolvable claims it is estimated from outcomes; Section 5 treats the general case.
Definition 5 (Bearer). For each consequential reliance act, β(j) is the agent answerable for it: the party who can be questioned, credited, or held responsible on its account.
Definition 5 is not decoration. Reliance weights are only well-posed if there is something whose calibration record can be scored, and calibration records require persistent identity. Section 12.3 argues that this requirement is a substantive constraint on machine participation, because model identity is unstable in a way human and methodological identity is not.
4. What is actually degraded
The literature offers three candidate answers to "what does abundance degrade": information quality, attention, and discriminability. We take the third, with a decomposition, because the first is false in general (abundance need not lower average quality, and the theory should not depend on it) and the second is too coarse to distinguish interventions.
Definition 6. An institution's epistemic performance decomposes into three components.
- Per-item discrimination
d': given inspection, the ability to distinguish candidates that merit reliance from those that do not. - Coverage
γ: the fraction of arriving candidates that receive inspection at all. - Selection quality
τ: the informativeness of the triage function that determines which candidates are covered.
These come apart, they degrade under different conditions, and they respond to different instruments. An institution with high d', coverage of two per cent, and excellent triage may be near-optimal. An institution with full coverage, perfect triage, and collapsed d' is failing. Aggregating them into a single scalar makes the governance section of any such theory incoherent, because provenance and deduplication act on τ and γ while preregistration and adversarial review act on d'.
Composed discriminability, for a binary accept/rely decision, is approximately
D = γ · d'_inspect + (1 − γ) · d'_triage
with d'_triage < d'_inspect by construction, since triage uses a strict subset of the available evidence. This expression is the engine of Section 8.
5. Measurement: reliance as a scoring problem
5.1 Against a scalar objective
A natural aggregate is R = Σ_j a_j w_j. It should be rejected. It rises with the sheer volume of reliance placed on warranted claims, so a community relying heavily on ten thousand true propositions scores above one relying appropriately on fifty; it is not comparable across institutions of different size; and it registers only one of the two errors that matter.
5.2 Two errors and a frontier
Definition 7.
Misreliance M = Σ_j a_j (1 − w_j)
Forgone reliance F = Σ_j (1 − a_j) w_j
M is weight placed where it is not deserved. F is warranted material the institution failed to use, which is the cost of over-filtering and the formal home of the "suppressed novelty" concern.
Let 𝔉(C) ⊂ ℝ²₊ be the set of (M, F) pairs achievable at load C given the institution's capacity and technology. The institution's choice of point within 𝔉(C) reflects its loss weights, which vary and are usually unobservable.
Definition 8 (Congestion threshold). C* is the supremum of loads such that 𝔉 is weakly increasing in set inclusion up to C*:
C* = sup { C : 𝔉(C′) ⊆ 𝔉(C) for all C′ ≤ C }
Congestion begins where the achievable frontier stops expanding and starts contracting.
Why this matters for identification. Whether a scalar R rises or falls in C depends on the institution's position on the frontier, hence on its preferences. Whether the frontier itself shifts inward does not. The inverted-U-in-R prediction that this literature usually advances is therefore not preference-free and is weak evidence; the frontier-contraction prediction is preference-free and is strong evidence. This is the paper's methodological recommendation and it is available at no theoretical cost.
5.3 One instrument
For claims whose outcomes eventually resolve, the institution's reliance weights a_j are probabilistic forecasts and can be scored. Under the Brier score, Murphy's (1973) decomposition gives
BS = Reliability − Resolution + Uncertainty
where Reliability is calibration error and Resolution is discrimination. This single instrument yields d' (Resolution) and the quality of reliance (Reliability) without a separate construct for each, aggregates across heterogeneous claim types, and comes with an existing inferential apparatus including corrections for selective sampling. We recommend it as the measurement spine in place of paired d' and R constructs.
5.4 The ground-truth restriction, stated as a restriction
w_j is estimable from outcomes only for resolvable claims. For novel causal hypotheses, contested inferential relations, and objections — the objects this literature says matter most — resolution may be unavailable in principle or arrive on decade timescales. This is a real limit and should not be finessed.
We therefore restrict the empirical scope of the discrimination results to three settings where ground truth exists or can be manufactured: (i) resolvable domains, including forecasting, replication outcomes, static-analysis true and false positives, post-merge defect attribution, and preregistered trial outcomes; (ii) seeded-error designs, in which candidate pools contain planted defects of known type and location, giving exact ground truth and permitting independent manipulation of volume, quality, correlation, and verification difficulty; (iii) surrogate criteria with stated assumptions, including downstream conclusion stability, inter-validator agreement, and revision rates. Claims about non-resolvable inferential objects are, in this framework, conjectures by analogy.
6. Result 1: The diversity ceiling
This is the paper's central result. It requires no assumption about human cognition.
6.1 Correlated generation
Let θ be the unknown status of a proposition. Candidate i yields an evidence signal
x_i = m(θ) + β + ε_i , β ~ N(0, σ_β²), ε_i ~ N(0, σ_ε²) i.i.d.
β is a shared generative bias: error induced by common training corpora, common retrieval sources, common prompt framing, shared fine-tuning ancestry, or a shared upstream factual error. ε_i is idiosyncratic. Write σ² = σ_β² + σ_ε² and let
ρ = σ_β² / (σ_β² + σ_ε²)
be the intraclass correlation of candidate errors.
Proposition 1 (Evidential ceiling). The posterior precision about θ obtainable from n such candidates is
I(n) = 1 / ( σ_β² + σ_ε²/n ) → 1/σ_β² as n → ∞
Equivalently, effective independent support is
n_eff = n / (1 + (n − 1)ρ) ≤ 1/ρ
for all n. Evidential content is bounded above by a constant determined by ρ alone, independent of volume.
Proof. Var(x̄ | θ) = σ_β² + σ_ε²/n, whose reciprocal is I(n); the design-effect identity n_eff = n σ²/(σ_ε² + n σ_β²) follows by substitution, and is increasing in n with limit σ²/σ_β² = 1/ρ. ∎
The result is elementary. Its significance is what it is applied to. The classical wisdom-of-crowds and Condorcet results require independence, and their failure under correlated votes is known (Ladha 1992); Hong and Page's diversity decomposition makes the loss exact. What is new is the observation that machine candidate generation is a mechanism that produces high ρ by construction, at scale, while presenting each candidate as textually distinct.
6.2 Cost grows in the count, value grows in the effective count
Let the value of precision be v(I), increasing and weakly concave, and let disposal-plus-handling cost be linear at rate c > 0 per candidate (Definition 1). Net epistemic value of a candidate stream is
W(n) = v( I(n) ) − c·n
Proposition 2 (Interior optimum from correlation alone). For ρ > 0 and c > 0, W has a unique interior maximiser. Taking v locally linear with slope v',
n* = ( σ_ε / σ_β² ) · ( √(v'/c) − σ_ε )
and n* → ∞ as ρ → 0.
Proof. I'(n) = σ_ε² / (n σ_β² + σ_ε²)², positive and Θ(n^{-2}). Marginal value is therefore strictly decreasing to zero while marginal cost is constant, giving a unique crossing. Setting v' I'(n) = c and solving yields the expression. With σ_β² = 0, I(n) = n/σ_ε² is linear and no interior optimum exists. ∎
Three consequences deserve emphasis.
(a) The inverted U does not require human limitation. No fatigue, no bounded rationality, no attention constraint, no queueing. A perfectly rational institution with unlimited stamina and a positive disposal cost still faces a finite optimal candidate load, because correlated candidates stop carrying information before they stop carrying cost. This matters because it makes the result robust to the standard rejoinder that better tools and better trained readers will absorb the volume.
(b) Correlation, not volume, is the policy-relevant variable. ∂n*/∂ρ < 0. Two candidate streams of identical size and identical average quality have different optimal loads and different congestion thresholds if their generative diversity differs. This is directly testable and is the sharpest prediction in the paper.
(c) Tooling improvements lose the race. n* ∝ c^{−1/2}. A k-fold reduction in per-candidate disposal cost raises the absorbable volume by only √k. Against candidate growth that is plausibly geometric, investment in disposal efficiency is a losing race at the exponent level. Capacity-side interventions cannot be the whole answer; diversity-side and admission-side interventions must carry part of the load.
6.3 Naive aggregation converts abundance into overconfidence
Proposition 3 (Double counting). An institution whose aggregation rule treats n concurring candidates as n independent observations overstates posterior precision by the factor
I_naive(n) / I(n) = n / n_eff = 1 + (n − 1)ρ
which grows linearly in n with slope ρ.
Proof. Immediate from Proposition 1. ∎
Miscalibration therefore grows without bound in volume even when discriminability is unchanged, and it grows fastest exactly where generative diversity is lowest. An institution can be getting no worse at judging individual candidates and simultaneously getting linearly more overconfident about its conclusions. This is a distinct failure mode from anything in the overload literature, and it is invisible to metrics based on acceptance rates or reviewer accuracy.
6.4 The ceiling is a market-structure variable
Suppose deployed agents draw on K distinct generative lineages, with within-lineage error correlation near one and between-lineage correlation ρ_K. Then
n_eff ≤ K / (1 + (K − 1) ρ_K) ≤ K
Corollary 4 (Concentration bound). The evidential value of machine corroboration is bounded by the number of independent generative lineages, not by the number of deployments, agents, or candidates. Growth in adoption raises n without raising K.
This has two implications that no purely psychological theory of overload can reach.
First, generative diversity is an epistemic public good. Lineage diversity determines the ceiling on how much machine agreement can ever be worth, no participant is compensated for maintaining it, and it is depleted by ordinary competitive dynamics in a market with strong returns to scale in pretraining. Distillation, synthetic-data training, and reliance on a small number of base models all lower K while raising apparent capability.
Second, this supplies the missing formal bridge between the epistemic argument and the political one. Arguments that a small number of providers should not come to own humanity's accumulated reasoning usually rest on data possession, which is a weak theory of capture because records are copyable. Corollary 4 gives a stronger one: concentration in generative lineage caps the achievable evidential content of machine-assisted inquiry regardless of who holds the records. Section 8.4 gives a second, sharper mechanism.
7. Result 2: Effort rationing and the inspection identity
Section 6 required nothing about human limits. This section adds them, and derives rather than assumes the relation between load and per-item discriminability.
7.1 Discriminability from inspection effort
For inspected candidates, let valid and invalid candidates generate evidence signals with means μ_T, μ_F and common noise σ. Per-item discriminability is d' = (μ_T − μ_F)/σ. Inspection effort e reveals diagnostic features with diminishing returns:
μ_T(e) − μ_F(e) = Δ_max ( 1 − e^{−β e} ) − ψ φ
where β > 0 indexes how quickly inspection reaches diagnostic features, φ is the share of the candidate stream that is machine-generated, and ψ > 0 is the surface-plausibility inflation of Section 2.3: machine generation raises the apparent credibility of invalid candidates at low inspection depth. Heterogeneity of the stream raises σ.
This is the correct place for the AI-specific mechanism, and note what it implies: the effect of φ is not a shift in the level of d' alone but an increase in the effort required to reach any given d'. Formally, ∂e/∂φ > 0 at fixed target d'. Fluency does not merely fool readers; it raises the price of not being fooled.
7.2 The inspection identity
Coverage is γ = m/C where m is the number inspected; effort per inspected item is e = V_eff/m. Therefore, identically,
γ · e = V_eff(C) / C
Proposition 5 (Inspection identity and the breadth–depth trade-off). The product of coverage and per-item effort equals effective capacity per candidate. Since V_eff(C)/C is strictly decreasing in C whenever a > 0, no institution can hold both coverage and inspection depth fixed as load rises. Under the specification of 7.1, d' is increasing in e, so an institution faces a strict choice between coverage loss and discrimination loss.
The identity is definitional; its content is what it forbids. Any claim that an institution "kept up" with rising candidate volume without additional capacity implies that either coverage or depth fell, and the theory says which observable should have moved.
Corollary 6 (Two degradation signatures). Institutions with norms that fix depth — every submission gets three full reviews, every alert gets triaged to resolution — degrade in coverage. Institutions with norms that fix coverage — everything gets looked at — degrade in d'. The signatures are distinguishable in data, and they call for different remedies. This is a prediction that no aggregate overload measure can make.
8. Result 3: Saturation, triage dominance, and enclosure
8.1 Validation is a loss system
Modelling validation as an M/M/1 queue and concluding that delay diverges as arrivals approach service capacity is a category error, because it assumes all arrivals are eventually served. Epistemic institutions drop: desk rejection, unreviewed closure, stale issues, unread submissions, ignored alerts. The correct model is a loss system with abandonment.
In a loss system, delay does not diverge at capacity. Throughput saturates at μ and the loss probability rises. Coverage becomes
γ → min{ 1, μ/λ }
Proposition 7 (Regime change). As λ/μ crosses one, the institution transitions from a regime in which coverage is near-complete and performance is governed by inspection quality, to a regime in which coverage is μ/λ and performance is increasingly governed by the accuracy of the dropping rule. The transition is a change in which parameter is binding, not a smooth degradation of a single parameter.
This is a better nonlinearity claim than delay divergence: it is true of real institutions, it does not depend on a queue discipline nobody uses, and it identifies a regime rather than a curve.
8.2 Triage dominance
From Section 4, D = γ d'_inspect + (1−γ) d'_triage. Substituting γ = μ/λ:
Proposition 8 (Triage dominance). As λ/μ → ∞,
D → d'_triage , ∂D/∂d'_inspect → 0 , ∂D/∂d'_triage → 1
Institutional discriminability converges to the discriminability of the triage function, and the marginal return to improving the validation layer converges to zero.
Proof. Immediate from γ → 0. ∎
The result is trivial algebraically and consequential institutionally. It says that a saturated institution's epistemic quality is a property of its filter, not of its experts, and that ordinary quality-improvement interventions aimed at the validation layer will show shrinking effect sizes as saturation increases. It also predicts something counterintuitive and checkable: interventions that improve review quality should have measurably smaller effects in more saturated venues, holding the intervention fixed.
8.3 Lineage sharing converts random error into invisible error
Let the triage function T and the generator G have error correlation ρ_{TG} arising from shared lineage: same base model, distilled descendant, shared retrieval corpus, shared prompt scaffolding.
Proposition 9 (Correlated blind spots). The probability that an error made by G is flagged by T is decreasing in ρ_{TG}. In the limit ρ_{TG} → 1, T flags no error that G systematically makes.
Proof sketch. T's discriminating signal is a function of the same latent representation that produced G's error; where the representation is wrong in the same direction, the error lies inside T's accepting region by construction. Formally, with T and G errors jointly normal with correlation ρ_{TG}, the conditional probability that T's signal crosses its rejection threshold given a G error decreases monotonically in ρ_{TG}. ∎
The implication is that a triage layer built from the same lineage as the generation layer does not merely miss errors at some rate. It misses the specific errors that matter, which are the systematic ones, while catching idiosyncratic noise. Aggregate triage accuracy statistics will look acceptable while the residual error becomes structured and undetectable.
8.4 Enclosure without ownership
Combining Propositions 8 and 9 yields the paper's institutional punchline.
Corollary 10 (Triage capture). Under saturation, an actor controlling the triage function determines institutional reliance to first order. Control of triage therefore constitutes control of the community's accumulated reasoning without any ownership of, or exclusive access to, the record.
Two consequences.
First, this is a stronger theory of the enclosure of collective reasoning than data possession. Portability mandates, open formats, and interoperable provenance address ownership of the record. They do nothing about ownership of the filter. If the filter is supplied by the same firms supplying generation, the accumulated reasoning of an institution is governed by a vendor even when every byte of it is exportable.
Second, it identifies a structural contradiction in the design of "judging institutions" that both (i) separate generation from ratification and (ii) ration scarce human attention by surfacing only what needs a person. Rationing is selection, and selection performed by a generator-adjacent system is authorship of the ratification agenda. Under Proposition 8 the selection layer dominates outcomes; under Proposition 9 shared lineage makes its failures systematic. The separation of generation from governance therefore requires separation at the triage layer specifically, which is the layer such designs typically automate first.
Structural remedies follow directly and are stated in Section 12.2: independent triage provision, mandated lineage diversity between triage and generation, published and pre-committed triage policy, adversarial triage audit with seeded errors, and a reserved exploration budget.
9. Result 4: Endogenous validation and the two-sided wedge
9.1 Contributors and validators are often the same people
In peer review, code review, clinical review, and internal organisational reasoning, the population that produces candidates is the population that assesses them, drawing on one time budget. Treating V as exogenous therefore removes the most important channel.
Let agent i allocate time t_i = g_i + v_i between generation and validation. Generation yields private return A_g B(g_i) with B increasing and concave, where A_g is generation productivity. Validation produces social value S(v_i) of which the agent privately captures κ · A_v S(v_i), with κ < 1 because validation is a public good within the institution and is credited weakly. Contribution also imposes congestion G(C) of which agent i internalises α_i < 1.
Private optimum:
A_g B'(g_i) − α_i G'(C) = κ A_v S'(v_i)
Social optimum:
A_g B'(g_i) − G'(C) = A_v S'(v_i)
Proposition 11 (Two-sided wedge). With α_i < 1 and κ < 1, the decentralised allocation has both C > C^{FB} and V < V^{FB}. Moreover a technology shock that raises A_g relative to κ A_v moves both away from first best simultaneously: candidate supply rises and validation supply falls from the same cause.
Proof. Comparison of first-order conditions; the wedge on the generation margin is (1−α_i)G'(C) > 0 and on the validation margin is (1−κ)A_v S'(v_i) > 0. Comparative statics in A_g at fixed κ A_v shift g_i up and, since t_i is fixed, v_i down. ∎
This is a stronger claim than congestion alone, and it is the correct diagnosis of what AI does to institutions whose reviewers are its authors. It is also the reason the standard framing — "generation got cheaper, verification did not" — understates the problem. Verification did not merely fail to get cheaper. Its supply fell, because the same agents reallocated.
9.2 Friction and credit are complements, and the choice depends on who contributes
Corollary 12 (Instrument assignment). Raising the private cost of contribution (friction) closes the generation wedge but does nothing to the validation wedge and falls hardest on contributors with the fewest resources. Raising κ (credit for validation) closes the validation wedge and, because the time budget is fixed, also reduces C on the internal margin. Therefore:
- For contributors drawn from the validating population, credit dominates friction: it closes both wedges and is not exclusionary.
- For contributors not drawn from the validating population — external submitters, and above all autonomous agents, which have no time budget and cannot be induced to validate — credit has no effect on
C, and admission instruments are the only available lever.
This yields a principled, non-arbitrary distinction that current governance discussions lack: human contributors should be governed by validation credit; autonomous agents should be governed by submission budgets or posted stakes. The distinction does not rest on suspicion of machine contributions. It rests on the absence of a reciprocity margin.
The natural instrument on the agent margin is a refundable stake rather than a fee: a bond posted with a contribution and returned unless the recipient marks the contribution as waste. This imposes cost on low-value contribution while leaving high-value contribution costless, which addresses the distributional objection to friction. Anderson and de Palma's (2009) result that costlier transmission can raise average message benefit, so that more messages are examined, supplies the formal warrant: friction that raises average quality can increase, not decrease, the amount of material that receives attention.
10. Result 5: The confirmatory triage trap
Triage under scarcity must allocate inspection by expected value, and expected value is estimated from priors, which are estimated from track record.
Proposition 13 (Absorbing exclusion). Under greedy value-maximising triage with a threshold inspection rule, and with source records updating only on inspection, the set of never-inspected sources is absorbing: a source whose prior lies below the inspection threshold never receives inspection, never generates a signal, never updates its posterior, and therefore remains below the threshold in perpetuity, independent of its true quality.
Proof. Threshold rule plus no signal implies posterior equals prior; the prior is below threshold by hypothesis; induct. ∎
The point is not that triage systems are biased. It is that exclusion is a consequence of optimality, and therefore cannot be corrected by removing bias. Greedy allocation in a bandit problem is a known consistency failure, and this is that failure wearing institutional clothes.
Two remarks strengthen it.
Cross-validation from mechanism design. The optimal mechanism for allocating a good under costly verification is prior-favouring: Ben-Porath, Dekel and Lipman (2014) show that the optimal scheme designates a favoured agent who receives the good unverified while others must be verified. Applied here — reliance is the good, claims are the agents, validation is verification — the optimal institution deliberately extends unverified reliance to high-prior sources and demands verification from the rest. Proposition 13 is therefore not an artifact of a crude triage rule. It is the shape of the optimal rule.
Interaction with automated validators. Where automated triage or fact-checking performs unevenly across languages, regions, and domains, priors are worse estimated for under-represented material, so the threshold excludes it at higher rates. Uneven validator competence and absorbing exclusion compound.
Corollary 14 (Exploration budget). Reserving a fraction ε of validation capacity for uniformly sampled below-threshold candidates makes the exclusion set non-absorbing and, by standard results on ε-greedy allocation, yields consistent long-run source rankings at a short-run precision cost of order ε.
The equity constraint is thus not a moral supplement to the theory. It is the consistency fix for a known failure of greedy allocation, and its cost is quantifiable. An institution that declines to reserve an exploration budget is not making a hard-headed efficiency choice; it is choosing an estimator that is inconsistent.
11. What kind of good is this?
11.1 Not a common-pool resource
Ostrom's common-pool resources are defined by difficulty of exclusion and subtractability: one user's appropriation leaves fewer resource units for another. Informational goods lack physical subtractability, which is why Hess and Ostrom (2007) developed a distinct knowledge-commons framework rather than extending fishery models directly.
The harm here does not arise from withdrawal. It arises from addition. Nothing is depleted from the reasoning record; the record grows monotonically. What degrades is the capacity to tell which of its contents merit reliance.
11.2 The correct positive classification
The good is non-rival. Its use requires a rival complement, namely validation capacity. Additions to the non-rival good increase the required input of the rival complement. The externality runs through the complementarity.
This places the object in the theory of impure public goods and congestible public goods (Buchanan 1965; Cornes and Sandler), not in the theory of common-pool resources. We propose the term complementarity congestion for the specific structure: a non-rival good whose consumption requires a rival complementary input, where contribution to the good imposes demand on the complement.
| Classic CPR | Pure public good | Club good | This case | |
|---|---|---|---|---|
| Rival in consumption | yes | no | no (within club) | no |
| Exclusion feasible | difficult | difficult | yes | partial (boundaries) |
| Harmful act | over-withdrawal | free riding | over-crowding of facility | over-contribution |
| Degraded object | resource stock | provision level | facility quality per user | discriminability of the record |
| Rival element | the good | none | the facility | the complementary input |
| Governance target | appropriation | contribution | membership and tolls | reliance, independence, admission |
Naming the category correctly does three things: it prevents the reinvention of standard results, it grants access to existing theory on tolls versus quotas and optimal membership, and it removes the appearance of a bespoke coinage. Ostromian institutional analysis remains applicable to the governance problem; it is the resource dynamics that differ.
11.3 Relation to existing formal traditions
| Tradition | What it supplies | Where it falls short for this problem | How this paper relates |
|---|---|---|---|
| Attention scarcity, rational inattention (Simon 1971; Sims 2003) | bounded processing as a primitive | single decision-maker; no institutional externality | supplies the capacity constraint in V |
| Information congestion (Van Zandt 2004; Anderson and de Palma 2009) | formal senders-congest-receivers model, with triage, message pricing, gatekeeping results | senders are advertisers with profit motives; no notion of warrant, evidence, or correlated content | closest formal antecedent; we add warrant, correlation, and endogenous validation supply, and we adopt their friction result |
| Search and screening (Weitzman 1979) | free-disposal monotonicity | assumes costless discard | we treat the failure of that assumption as the object of study |
| Information overload experiments | performance decline under excess input | context-dependent; volume is the wrong variable | motivates the load decomposition in 3.1 |
| Queueing and congestion | threshold behaviour near capacity | work-conserving service assumption is false here | replaced with a loss system in Section 8 |
| Signal detection | operational discriminability | requires ground truth | adopted, with the scope restriction in 5.4 |
| Information cascades (Bikhchandani et al. 1992) | agreement need not carry independent information | a sequential-social mechanism, not a volume mechanism | Section 6 supplies a generative-correlation mechanism instead |
| Diversity and jury theorems (Hong and Page; Ladha 1992) | exact loss from correlated judgement | not about volume or cost | Proposition 1 is the machine-generation application |
| Costly verification mechanism design (Townsend 1979; Glazer and Rubinstein 2006; Ben-Porath, Dekel and Lipman 2014) | optimal allocation when checking is expensive; value of commitment; prior-favouring optimal rules | not about volume growth or congestion | supplies theorems for Section 12 and cross-validates Proposition 13 |
| Knowledge commons (Hess and Ostrom 2007; Governing Knowledge Commons) | institutional analysis of non-rival resources | resource is not congestible in their treatment | we add the congestible complement |
| Epistemic pollution (Levy 2018) and generation/verification imbalance (Ma 2026) | the closest conceptual statements of the problem | largely qualitative; imbalance asserted rather than modelled | we supply mechanism, nonlinearity, and identification |
| Model collapse (Shumailov et al. 2024) | machine-side analogue of correlated degraded input | about training distributions, not institutional validation | a parallel consequence of the same low-K structure |
Two places where existing results cut against common institutional preferences in this literature, and which we therefore flag rather than bury. Anderson and de Palma find conditions under which a monopoly gatekeeper outperforms decentralised access pricing when nuisance costs are moderate, which is in tension with an unconditional preference for polycentric validation. And Glazer and Rubinstein find that commitment to a verification rule beats discretion, which implies that triage policy should be published and pre-committed rather than left to expert judgement. Both deserve engagement rather than assertion in the opposite direction.
12. Governance implications that follow from the results
Each recommendation below is tied to a specific proposition. Recommendations not so tied have been removed.
12.1 Instruments and their targets
| Instrument | Acts on | Warranted by |
|---|---|---|
Independence accounting: reliance weighted by estimated n_eff, not source count |
the double-counting failure | Propositions 1, 3 |
| Lineage-diversity requirements for consequential machine corroboration | the ceiling 1/ρ |
Corollary 4 |
| Provenance sufficient to identify generative ancestry, not merely authenticity | estimability of ρ |
Propositions 1, 9 |
| Deduplication and indexing | a in V_eff, hence C* |
Section 3.2, Proposition 5 |
| Validation credit and calibration-based status | κ |
Proposition 11 |
| Posted refundable stakes or submission budgets for autonomous agents | C on the non-reciprocal margin |
Corollary 12 |
| Published, pre-committed triage policy | discretion loss in selection | Glazer and Rubinstein; Proposition 8 |
| Triage–generation lineage separation and adversarial triage audit | ρ_{TG} |
Propositions 9, Corollary 10 |
Reserved exploration budget ε |
absorbing exclusion | Proposition 13, Corollary 14 |
| Prospective registration with machine-checkable outcome links | estimability of w_j; calibration records |
Section 5.3 |
| Dependency records and correction propagation | persistence of invalidated reliance | Section 3.1 (δ_j) |
| Answerability: a named bearer per consequential reliance | well-posedness of reliance weights | Definition 5 |
12.2 Mechanisms compose in a fixed order
These are not a menu. Several are impossible without others.
provenance (including generative ancestry)
├──> stable bearer identity ──> calibration record ──> reliance weighting
└──> independence accounting (n_eff) ──> aggregation rules
│
triage policy (published, lineage-separated)
│
coverage allocation (+ ε exploration)
│
inspection ──> prospective registration ──> outcome scoring
│
dependency records ──> correction propagation
Reliance weighting without stable identity is unimplementable. Independence accounting without ancestry provenance is unimplementable. Correction propagation without dependency records is unimplementable. Presenting these as parallel options is the most common error in this design space.
Note also what provenance does not do. Content-provenance standards deliberately certify history rather than truth. Provenance lowers a and makes ρ estimable. It does not raise d'. Conflating the two produces credibility badges that are informative about origin and uninformative about warrant.
12.3 Model identity is unstable, which constrains machine reliance weighting
Reputation requires something persistent to attach to. Model identity is not persistent: versions, fine-tunes, distillations, quantisations, system-prompt variants, routing layers, and multi-agent wrappers all produce entities whose track records are not transferable on any principled basis. A reliance weight earned by one checkpoint says little about its successor and nothing about a distillation deployed under another name.
Two consequences. First, machine reliance weighting requires provenance rich enough to identify the relevant lineage, which is a stronger requirement than authenticity provenance. Second, where lineage cannot be established, reliance weights should attach to methods and evidence chains rather than to sources, and to human bearers rather than to models. This is not a preference for human judgement over machine judgement. It is a consequence of Definition 5: only a persistent identity can accumulate a calibration record, and only a calibration record makes a reliance weight mean anything.
12.4 The objective is not minimal load
The theory does not recommend minimising C. Below C* the achievable frontier is expanding and additional candidates strictly improve the institution. The objective is
choose institutional design I to maximise C*(I) and expand 𝔉(C; I)
subject to the exploration constraint of Corollary 14. Institutional quality is measured by how much candidate reasoning the institution can absorb before its frontier contracts, not by how little it admits.
13. Predictions, thresholds, and refutation conditions
Each prediction states the manipulation, the outcome, and what result would count against the theory. All require seeded-error or resolvable-outcome designs (Section 5.4). A concrete design for the first of these — a controlled AI-assisted code-review study that varies candidate load and lineage correlation while holding the seeded defects constant — is set out in the companion proposed seeded-error experiment. It is a design only; it has not been run.
P1 — Lineage-diversity effect (primary test). Holding candidate count, surface quality, verification difficulty, and validator hours fixed, validator discrimination and calibration are worse for candidate sets generated from a single model lineage than from four or more distinct lineages. Refuted if no significant difference emerges at n ≥ 50 candidates with adequate power. This is the paper's most distinctive prediction and its most direct falsifier.
P2 — Double-counting slope. Institutional overconfidence under naive aggregation grows approximately linearly in the number of concurring machine candidates, with slope increasing in measured lineage correlation. Refuted if overconfidence is flat in n, or if its slope is unrelated to ρ.
P3 — Inspection identity. Within an institution, coverage and per-item inspection depth are negatively related across load levels, with their product tracking V_eff/C. Refuted if both are maintained under rising load without capacity growth.
P4 — Divergent degradation signatures. Depth-preserving institutions degrade in coverage; breadth-preserving institutions degrade in d'. Refuted if degradation is uniform across institutional norms.
P5 — Threshold movement under provenance and deduplication. C* is higher under complete provenance and reliable deduplication than under baseline, holding validator hours fixed: C*_provenance > C*_baseline. Refuted if provenance improves individual judgement accuracy but leaves the threshold unchanged, which would indicate that a is not the binding term.
P6 — Triage dominance. The effect size of interventions that improve inspection quality declines with the saturation ratio λ_C/μ_V, approaching zero; the effect size of interventions that improve triage rises. Refuted if review-quality interventions have load-invariant effects. Existing randomised review-improvement studies can be re-analysed by venue saturation to test this cheaply.
P7 — Correlated blind spots. With seeded errors, the probability that a triage layer flags an error made by a generator declines in the lineage correlation between them, and residual errors become more systematic. Refuted if detection rates are independent of shared lineage.
P8 — Two-sided shift. After AI adoption in a population that both generates and validates, candidate output per member rises and validation hours per member fall, with the effect concentrated where validation is least credited. Refuted if validation hours are stable, or if the fall is uniform across credit regimes.
P9 — Absorbing exclusion and its fix. Under greedy triage, the never-inspected source set is stable over time; introducing an ε-exploration budget measurably improves long-run source-ranking accuracy at a short-run precision cost of order ε. Refuted if exclusion sets churn substantially without intervention.
P10 — Square-root tooling. Absorbable volume scales approximately as c^{−1/2}: a k-fold reduction in per-candidate disposal cost raises tolerable load by roughly √k, not k. Refuted if tooling gains translate proportionally or better into absorbable volume.
Two structural falsifiers deserve separate statement, because they would undermine the theory rather than a single proposition.
F1. If measured error correlation across deployed models is near zero in practice, the diversity ceiling is not binding and Section 6 loses its force. This is an empirical question about the current model ecosystem and it is answerable now.
F2. If per-candidate disposal cost can be driven to approximately zero by automation with reliable ground truth, then a → 0, free-disposal monotonicity is restored, and the theory dissolves. We regard this as unlikely in domains where warrant terminates in experiment, but it may hold in domains with cheap automated oracles, and the theory should be expected to fail there.
14. Limitations, and what we think is most likely to be wrong
The theory is comparative-static. It says where thresholds are and what moves them. It does not model the adaptation race between candidate-generation technology and validation technology, which is the question that determines outcomes over a decade. A dynamic extension in which institutional capacity responds to observed congestion is the obvious next step and is not attempted here.
The correlation model is stylised. A single scalar ρ with a shared additive bias is the simplest structure that produces the ceiling. Real generative correlation is hierarchical, domain-dependent, and prompt-dependent. The direction of Proposition 1 is robust to hierarchy; the closed form in Proposition 2 is not.
ρ may be difficult to estimate in practice. The result is only actionable to the extent that generative ancestry is recordable. Provenance standards currently certify authenticity, not lineage. This is a gap between the theory's requirements and existing infrastructure, and it is the most likely reason the independence-accounting recommendation would fail in deployment.
The disposal-cost premise is asserted for candidate reasoning and demonstrated only by analogy. The claim that machine fluency raises disposal cost is plausible and consistent with what is known about human detection of synthetic material, but ψ in Section 7.1 has not been measured for anything. If ψ ≈ 0, Section 2.3 is wrong and the paper's framing loses its AI-specificity while its Section 6 result survives.
No adversarial model. Contributors here are sincere and self-interested. Deliberate flooding to exhaust a validation layer is a real strategy, it changes the optimal instruments toward bonds and stakes, and it is not modelled.
Single institution, no spillovers. We do not model competition among validators, cross-institutional reliance, or the possibility that one institution's discriminability is a public good for others.
The most likely thing to be wrong is the relative importance we assign to the mechanisms. We have put correlation first because it is the most AI-specific and the most robust to human-capability improvements. It is entirely possible that in practice the a·C handling term dominates everything and the correct theory is much duller: institutions are simply doing more clerical work per unit of insight, and better indexing solves most of it. P5 and P10 are designed to distinguish these.
15. Conclusion
The familiar statement of the problem — AI generates plausible reasoning faster than institutions can verify it — is true, is no longer new, and is not yet a theory, because it does not explain why abundance should hurt when unwanted material can be discarded.
The answer proposed here is that discarding is not free, and that generative models raise its price precisely by producing the surface features that once made it cheap. On top of that, machine candidates carry bounded joint evidential content: correlated generative lineages cap effective independent support at 1/ρ regardless of volume, so evidential value saturates while disposal cost does not. That asymmetry produces a finite optimal candidate load without any appeal to human frailty, makes generative diversity an epistemic public good, and ties the epistemic problem to the market structure of foundation models.
Beyond that, the results are about where authority goes when capacity binds. Coverage and depth cannot both be held; at saturation, institutional discriminability converges to the discriminability of the triage function, so the filter rather than the record becomes the site of capture; shared lineage between filter and generator turns random error into systematically invisible error; where contributors also validate, one shock moves both margins the wrong way; and greedy triage makes exclusion absorbing, so distributional harm follows from optimality and requires a reserved exploration budget rather than good intentions.
None of this recommends less reasoning. Below the threshold, more candidates strictly improve an institution, and the design objective is to raise the threshold rather than to lower the load. What the theory recommends governing is reliance, independence, admission on the non-reciprocal margin, and answerability. What it recommends measuring is not how much an institution produces but how much it can absorb before its frontier begins to contract.
References
Anderson, S. P., and de Palma, A. (2009). Information congestion. RAND Journal of Economics 40(4), 688–709.
Anderson, S. P., and de Palma, A. (2012). Competition for attention in the information (overload) age. RAND Journal of Economics 43(1), 1–25.
Bastian, H., Glasziou, P., and Chalmers, I. (2010). Seventy-five trials and eleven systematic reviews a day: how will we ever keep up? PLoS Medicine 7(9).
Ben-Porath, E., Dekel, E., and Lipman, B. L. (2014). Optimal allocation with costly verification. American Economic Review 104(12), 3779–3813.
Bikhchandani, S., Hirshleifer, D., and Welch, I. (1992). A theory of fads, fashion, custom, and cultural change as informational cascades. Journal of Political Economy 100(5), 992–1026.
Buchanan, J. M. (1965). An economic theory of clubs. Economica 32(125), 1–14.
Chetty, R., Saez, E., and Sándor, L. (2014). What policies increase prosocial behavior? An experiment with referees at the Journal of Public Economics. Journal of Economic Perspectives 28(3), 169–188.
Clark, T., Ciccarese, P. N., and Goble, C. A. (2014). Micropublications: a semantic model for claims, evidence, arguments and annotations in biomedical communications. Journal of Biomedical Semantics 5(28).
Cornes, R., and Sandler, T. (1996). The Theory of Externalities, Public Goods, and Club Goods. 2nd ed. Cambridge University Press.
Dung, P. M. (1995). On the acceptability of arguments and its fundamental role in nonmonotonic reasoning, logic programming and n-person games. Artificial Intelligence 77(2), 321–357.
Glazer, J., and Rubinstein, A. (2006). A study in the pragmatics of persuasion: a game theoretical approach. Theoretical Economics 1(4), 395–410.
Goldberg, S. C. (2010). Relying on Others: An Essay in Epistemology. Oxford University Press.
Groth, P., Gibson, A., and Velterop, J. (2010). The anatomy of a nanopublication. Information Services and Use 30(1–2), 51–56.
Hardwig, J. (1985). Epistemic dependence. Journal of Philosophy 82(7), 335–349.
Hess, C., and Ostrom, E., eds. (2007). Understanding Knowledge as a Commons: From Theory to Practice. MIT Press.
Hong, L., and Page, S. E. (2004). Groups of diverse problem solvers can outperform groups of high-ability problem solvers. PNAS 101(46), 16385–16389.
Kish, L. (1965). Survey Sampling. Wiley. [design effect and effective sample size]
Kunz, W., and Rittel, H. W. J. (1970). Issues as elements of information systems. Working Paper 131, Institute of Urban and Regional Development, University of California, Berkeley.
Ladha, K. K. (1992). The Condorcet jury theorem, free speech, and correlated votes. American Journal of Political Science 36(3), 617–634.
Levy, N. (2018). Epistemic pollution and the epistemic environment. [See Levy's work on polluted epistemic environments; verify edition and venue before citing.]
Loder, T., Van Alstyne, M., and Wash, R. (2006). An economic response to unsolicited communication. Advances in Economic Analysis and Policy 6(1).
Lorenz, J., Rauhut, H., Schweitzer, F., and Helbing, D. (2011). How social influence can undermine the wisdom of crowd effect. PNAS 108(22), 9020–9025.
Ma, J. W. (2026). Toward an engineering of science: rebalancing generation and verification in the age of AI. arXiv:2605.10425.
Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology 12(4), 595–600.
Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press.
Resnick, P., and Zeckhauser, R. (2002). Trust among strangers in internet transactions: empirical analysis of eBay's reputation system. In The Economics of the Internet and E-Commerce, Advances in Applied Microeconomics 11.
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755–759.
Simon, H. A. (1971). Designing organizations for an information-rich world. In M. Greenberger, ed., Computers, Communications, and the Public Interest. Johns Hopkins Press.
Sims, C. A. (2003). Implications of rational inattention. Journal of Monetary Economics 50(3), 665–690.
Townsend, R. M. (1979). Optimal contracts and competitive markets with costly state verification. Journal of Economic Theory 21(2), 265–293.
Van Zandt, T. (2004). Information overload in a network of targeted communication. RAND Journal of Economics 35(3), 542–560.
Weitzman, M. L. (1979). Optimal search for the best alternative. Econometrica 47(3), 641–654.
Note on verification status. This draft is a theory paper and cites empirical literature illustratively rather than exhaustively; the empirical review is the subject of the companion paper. Sources dated 2026 and any figure carried over from the companion review are pending independent verification and are marked as such in the verification appendix. Given that this paper's subject is the cost of establishing warrant, we regard an explicit and auditable verification record as a requirement rather than a courtesy, and recommend that any published version ship one.
Appendix A — Notation
| Symbol | Meaning |
|---|---|
N, C |
raw candidate count; effective candidate load Σ ℓ_j |
ℓ_j |
per-candidate load: floor plus stake × difficulty × novelty × dependency |
c_j, c^{base} |
disposal cost; irreducible handling and disposal floor |
V, V_eff |
gross and effective validation capacity |
a, b, η |
linear handling coefficient; superlinear overhead coefficient and exponent |
γ |
coverage: fraction of candidates inspected |
e |
inspection effort per inspected candidate |
τ, d'_triage |
triage signal precision; discriminability of triage alone |
d', D |
per-item and composed discriminability |
ρ |
intraclass correlation of candidate errors (generative correlation) |
n_eff |
effective independent support, n/(1+(n−1)ρ) |
K |
number of distinct generative lineages |
ρ_{TG} |
error correlation between triage layer and generation layer |
a_j, w_j, β(j) |
reliance weight; warranted weight; answerable bearer |
M, F, 𝔉(C) |
misreliance; forgone reliance; achievable frontier |
C* |
congestion threshold (Definition 8) |
κ, α_i |
share of validation's social value privately captured; share of congestion internalised |
φ, ψ |
machine-generated share of stream; surface-plausibility inflation of μ_F |
ε |
reserved exploration fraction of validation capacity |
Appendix B — Which results depend on which assumptions
| Result | Requires | Does not require |
|---|---|---|
| Prop. 1–4 (diversity ceiling) | ρ > 0, c > 0 |
human limitation, overload, b > 0, queueing |
| Prop. 5–6 (inspection identity) | a > 0, d' increasing in e |
b > 0, correlation |
| Prop. 7–10 (triage dominance, capture) | loss system, d'_triage < d'_inspect |
correlation, superlinear g |
| Prop. 9 (correlated blind spots) | ρ_{TG} > 0 |
saturation |
| Prop. 11–12 (two-sided wedge) | shared time budget, κ < 1, α < 1 |
correlation, saturation |
| Prop. 13–14 (absorbing exclusion) | threshold triage, record updates on inspection only | AI at all — this is a general result |
The table is included because the theory should not stand or fall as a unit. Propositions 13 and 14 hold for any scarce-verification institution. Propositions 5 through 12 hold for any cheap-generation shock. Only Propositions 1 through 4 and 9 are specific to machine generation, and those are the ones we would defend hardest.