Small contracts, measurable limits: mathematical foundations for the MCRP seed release
Research prototype, 25 September 2026. Conditional propositions and synthetic experiments; not evidence of field efficacy. This manuscript advances the counterexamples in dossier 06 into small executable modeling interfaces. The results use elementary probability, constrained incentives and positive-system bounds. The contribution is the alignment of those tools with protocol records, not a claim to invent the underlying mathematics.
Abstract
A protocol intended for a solo researcher, a large collaboration, a single agent and a team of agents should not charge its participants for understanding the internal complexity of everyone else. It should expose what can be checked at the boundary: exact work, declared responsibility, available attention, performed checks, and the conditions under which a reliance decision needs renewal. Four models identify what this interface can and cannot guarantee. Sampling accountable groups before representatives removes one representation-multiplicity incentive under known control. Selective completion can nevertheless reverse offered-panel bounds, while retries spend scarce attention without repairing the selection bias. A common weighted repair envelope controls expected cascades under changing regimes, but neither stable snapshots nor a small expected workload certify a safe queue. Finally, an audit must fit effort incentives, participation and its own funding constraint simultaneously. The package contains exact calculations, independent finite enumerations, reproducible stochastic checks and deliberate negative cases. Every positive result is paired with a premise whose failure is observable in a toy experiment or requires a real-world investigation.
1. The interface as a modeling boundary
The participant-facing actions remain offer → check → rely → amend. They are composable actions, not an exclusive sequence of statuses. The mathematics adds no fifth user action. It specifies what a scheduler or experiment must record behind those actions.
An offer fixes a version, bounded claim, requested check and responsible boundary. A check records performed work and its limits; it is not a vote that converts into general truth. Rely is a separate actor’s decision for a stated use, with observed evidence and renewal conditions. Amend identifies a material change and affected uses. A large team may automate hundreds of internal checks; that does not create hundreds of independently controlled reviewers.
The minimum experiment state therefore includes an immutable target identifier; declared control groups and qualifications; a panel-generation policy; invitation, refusal, completion and reliance events; actual resource debits; and a declared dependency graph. These are inputs and records, not facts made true by a schema. In particular, the mathematics below does not authenticate identity, discover undeclared dependencies or grant scientific authority to a software process.
2. Representation invariance belongs to a distribution
Let be a fixed finite set of eligible accountable groups. Group has interchangeable representatives. Eligibility, capacity and conflict conditions have already been resolved at group level. Let be a nonempty set of feasible -group panels. We compare two policies.
Representative-first policy. Enumerate every feasible representative panel with at most one member from each group, then choose uniformly. Its induced probability of a group panel is
It satisfies the one-seat-per-group rule while still rewarding label multiplicity. With , and ,
One hundred representatives move the probability from to without adding a group or changing the claim’s scientific needs.
Group-first policy. Fix a distribution on using only clone-invariant group attributes. Draw . Conditional on , choose one representative per group by any normalized kernel supported on that group panel.
Proposition 1 (pushforward invariance). If only representative multiplicities change, while , and remain fixed, the distribution of selected group panels remains .
Proof. For any , the probability of all representative realizations mapping to is . No summand from another group panel maps to . Consequently every statistic depending only on the group panel is invariant.
This is a statement about an entire pushforward distribution, not a claim that a cap on individual graph scores suffices. The executable baseline takes all -subsets and uniform ; it does not solve general conflict-aware scheduling. A production implementation must construct feasible panels using actual skill, conflict, independence and capacity constraints before drawing, and define what happens when the feasible set is empty. It must not silently relax coverage.
Falsifier. Declare two secretly controlled accounts as different groups. With and uniform pairs, the true controller behind now receives at least one seat with probability , and both seats with probability . Group-first sampling has not solved Sybil resistance. Douceur’s original analysis establishes why multiple identities undermine redundancy and why identity certification is a substantive assumption; our toy group map supplies that assumption rather than implementing it. The Sybil Attack
Meaning for participants. An individual and a collaboration may expose the same accountable boundary. Splitting a team into more named software processes must not automatically buy the team more external allocation probability. Real new expertise or independently controlled capacity may legitimately change the feasible set; that is not the cloning transformation in Proposition 1.
3. Completion is a second selection mechanism
Label a panel if it belongs to a synthetic risk category. This is a model label, not a diagnosis of any real group. Let among offers, let be its completion probability, and let be the other category’s completion probability. Then
A offered share becomes of completed panels when and . An offered-panel cap is therefore not a completed-panel cap.
Proposition 2 (completion envelope). If , , and , then
Proof. The conditional fraction increases with and and decreases with whenever its denominator is positive. Substitute the respective extrema. The case has zero risky completions.
The missing premise is often the important one: humans cannot be compelled to supply completion. Refusals may be the appropriate response to conflicts, insufficient time, confidential material or inadequate expertise. Treating them as misconduct would change the protocol’s purpose.
Finite retries cost attention but do not remove IID selection bias
Suppose each new attempt independently draws from the same offer distribution, with unchanged completion probabilities. Stop at the first completion or after attempts. Write for attempts actually made. For ,
The probability of risky completion is . Dividing by total completion probability cancels the same sum: the risky share among completions remains for every . For , every request uses attempts and stays unresolved.
Let be invitation/triage cost per attempt and be work cost charged on completion. Then
Expected refusals equal ; expected retries equal . These are different quantities. Standby reservations, partial work, late cancellations and case intake require additional debits in a real ledger. The toy costs are chosen constants, not estimated labor.
For , one attempt costs model units on average, with unresolved. Ten attempts cost about , and fifty about . The completed-panel risk share remains . The growth in completed work is real in this model, but presenting only the completion count conceals the resource cost and composition.
The seeded demonstration runs 20,000 requests per case, 80,000 in total. For the ten-attempt case, 13,810 complete, the observed risky share is , and its pointwise nominal 95% Wilson interval is approximately . This interval describes Monte Carlo variation under the supplied Bernoulli model. It says nothing about uncertainty in real refusal behavior. Adaptive rerolls, learning completion rates, collusion and repeated contacts violate the IID model; they need a new model, not reuse of these error bars.
The denominator is part of the result
Publish eligible, invited, accepted, completed and relied-upon counts separately, plus offered requests, excluded requests, refusal attempts, retries, outstanding work and resource costs. The toy retry model collapses acceptance and completion into one Bernoulli event; the protocol simulator should keep the stages distinct. The mathematical results are intentionally not a substitute for that event log.
Inverse probability weighting can diagnose observed selection when probabilities are known and positive, but it does not produce missing scientific checks or restore a physical panel guarantee. The classical unequal-probability estimator is standard statistical machinery, not a novel trust mechanism. Horvitz and Thompson, 1952
Red-team extension: adapted completion shares
The red mathematical reviewer supplied a stronger envelope and an
independent finite-tree verification in
reviews/red-math/adaptive-completion.md. At every reached
pre-attempt history
,
require
,
,
and
.
Stopping must occur before inspecting the current draw. The risky and
other terminal-completion masses at that history obey
,
where
.
Weight by the probability of reaching each history and sum.
First-completion events are disjoint, so
and the same completion-share envelope follows. This permits adapted
policies and history-dependent completion, but does not
generalize IID retry costs. Cherry-picking an uncounted draw
violates the effective offer bound. Selecting only risky completed
panels for reliance is yet another selection stage; the completion bound
says nothing about that relied-upon distribution. Credit for this
extension and its independent probes belongs to the red mathematical
lane.
4. Repair must survive changing regimes
Let be a nonnegative row vector of outstanding repair items by type in generation . Types can mean data calibration, statistical interpretation, software execution or a rights-handling obligation. The branching abstraction counts work items, not unique scientific truths. Duplicate notices and reusable checks need deduplication before interpretation as labor.
Conditional on history , suppose
componentwise, where belongs to a declared family and can be chosen based on history. This premise is stronger than estimating unconditional average offspring in a convenient sample. Suppose there is a vector and with
Proposition 3 (common weighted repair envelope). For deterministic initial , the expected cumulative weighted work satisfies
Proof. Multiplying the conditional bound by positive gives . Iterated expectation yields . Sum the finite geometric bound and apply monotone convergence to nonnegative partial sums. Independence between generations and a fixed switching sequence are unnecessary under the conditional premise.
If bounds immediate work hours per type- item, the right side also bounds expected cumulative hours. If is only a mathematical witness, its units are abstract weighted work and must not be relabeled as human hours. A failure to find this witness does not prove instability: even a supplied can fail where another positive succeeds. Common Lyapunov methods are established stability tools; the author’s accessible survey distinguishes stable subsystems from stable switching. Lin and Antsaklis, author manuscript
Counterexample: snapshots pass while switching explodes
Take
Each matrix is nilpotent, so repeating either regime alone eventually clears all work in this linear model. Alternating from doubles work every generation and produces 1,024 items in generation 10. No common positive envelope exists: and imply . Thus checking each snapshot’s spectral radius is insufficient even before introducing stochastic human behavior.
For the positive example, use
Both satisfy . Starting with one first-type item gives an expected cumulative weighted bound . Exhaustive enumeration of every length-eight switching sequence checks the finite inequalities; a longer alternating trajectory is included in the CSV output. These tests supplement the proof; they do not estimate a real matrix family.
Finite mean is not a service guarantee
A branching item that produces 50 descendants with probability and none otherwise has mean offspring . It is subcritical in first moment, yet its first generation alone exceeds an eight-item capacity with probability . For nonnegative total work , the preceding expectation bound gives at most the conservative Markov bound , capped at one. It does not establish deadline compliance or acceptable tails.
The capacity interface instead records commitments by person and epoch across all lanes. Three hours of checking, two of audit and four of repair consume nine hours of the same person’s eight-hour day. Three separately feasible lane budgets do not make the joint plan feasible. Skills, simultaneous appointments, legal priority, independence and deadlines impose additional constraints. Our checker only verifies the shared declared hour totals. Unrecorded obligations remain a failure mode; expected future repair bounds cannot be booked as realized work.
5. Incentives must fit the audit and participation budget
Consider a one-shot invitation. Honest work yields reward , costs and is incorrectly sanctioned on an audit with probability . Shirking saves and is detected on an audit with probability . Audit probability is ; the modeled enforceable loss is . Abstaining yields outside utility . Utilities are
For , , , honest work is a weak best response and the expected audit expenditure for invitations at unit cost fits only if
Proposition 4 (one-shot feasibility). Under the stated assumptions, this interval is necessary and sufficient for honest work to weakly dominate both shirking and abstention while satisfying the expected invited audit expenditure constraint. This is not a hard reservation or a pathwise spending guarantee.
Proof. Rearrange , , and . All constraints are affine in . Their intersection is exactly the interval. The code handles zero audit cost, zero effort cost and zero discrimination without dividing by zero. Negative discrimination with zero effort permits only .
With , the interval is approximately . Reducing to 1 makes it empty. Keeping ample funding but lowering reward to and increasing false sanctions to can also make it empty: is needed for effort while is needed to keep honest participation worthwhile. Increasing policing without controlling false positives can exclude contributors.
The sweep uses 102 artificial agents with two reward levels and 51 effort costs. Each chooses the best of honest, shirk and abstain. Ties prefer abstain, then honest, then shirk. It reports both honest share among participants and counts among all invitations, along with expected audit work among participants and expected expenditure if every invited agent participates. The sweep is a mechanism diagnostic: its utilities are neither survey responses nor a fit to scientists, lawyers or software agents.
Red-team correction: fund a hard audit cap
The first draft incorrectly called a worst-case reservation. Red finding RM-1 supplies the counterexample: independent Bernoulli audits at exceed with probability about . The algebra above has been relabeled expected expenditure; the original finding is retained.
For fixed and identical actual audit cost , define . Uniformly sample a concealed subset of invitation indices before behavior, with . Every index has marginal , and every realized audit cost is at most . Noninteger can be implemented by mixing its adjacent integer counts, provided both are at most . Therefore the hard-cap upper bound is
For
,
the expected constraint allows
but the hard cap allows at most
.
hard_audit_interval intersects this bound with effort and
participation; blinded_audit_sample implements the subset
lottery. Concealment is a caller obligation, not a cryptographic
property of a returned Python list. If actors learn their selection
before effort, unaudited actors may shirk and the
common-
incentive calculation no longer applies to their information sets.
Changing invitation populations, variable audit cost, failed delivery
and funding of the reward
are separate obligations. Fixed sample size does not create credible
commitment or independent audit authority.
This is deliberately not a whole-network equilibrium. Audit credibility, collusion, appeal reversal, reviewer judgment, wealth constraints, repeated identity resets and the legitimacy of imposing are outside the model. The simpler prototype may rely on bounded recognition and loss of future assignments, not monetary penalties. Such a loss still needs a justified valuation and a responsible institution; a symbolic variable cannot supply either.
6. Four deep adaptations, one common experiment boundary
Physics and astronomy. Treat calibration and inference as separate repair types. A new instrument calibration can invalidate a derived estimate without invalidating the raw observations. A collaboration’s 100 analysis processes remain one control boundary for panel sampling. Ask whether repair notices reach the actual downstream uses and whether the limited calibration specialists are double-booked. The switched-matrix counterexample represents changing dependency patterns, not a fitted astrophysical pipeline.
Biology. A data-use restriction can remove evidence without proving a biological claim false. Give evidence availability, statistical checks and authorized reuse separate records. Let wet-lab replication complete more slowly than a computational check; the completion model demonstrates why counting only finished work can favor an easy-to-complete category. A wet-lab refusal is not a negative scientific verdict. Consent and restricted data do not enter a public toy fixture.
Economics and social science. Use offered/completed denominators to teach how selection can create apparent institutional improvement. Vary completion rates, setup costs and outside options; report unresolved projects and labor as outcomes. The audit experiment shows a mechanism whose apparent integrity among remaining participants improves while scarce contributors can leave. A substantive causal claim requires a real identification strategy and external data, not this queue.
Law and governance. Treat rights response, scientific disagreement and appeal as different authorities drawing on some of the same people’s time. The capacity checker can reject nine hours booked into eight without deciding legal priority. The audit model’s sanction must not be read as a legally enforceable fine. A timely record, lawful recipient-specific disclosure and authorized reversal require operator processes outside these mathematical models.
7. Claim ledger and reproducibility
| ID | Status | Claim and evidence | Limitation |
|---|---|---|---|
| M1 | T | Group-first pushforward invariance; proof and exact probability tests | Fixed true/declared group map, feasibility and weights |
| M2 | T | Completion conditioning and finite IID retry accounting; proof and independent path enumeration | Retry costs use IID assumptions; the adapted completion-share extension needs bounds at every reached history |
| M3 | T | Common positive envelope bounds conditional expected cumulative repair; proof | Envelope validity is an empirical/operational assumption |
| M4 | T | One-shot audit/participation/funding interval; affine inequalities | Audit and enforceable loss exogenous, no equilibrium selection |
| M5 | I | Executable synthetic implementation reproduces positive and negative cases | Floating-point toy range, no production scheduler |
| M6 | I | 80,000 seeded simulated requests with pointwise uncertainty | Simulation outcomes only, no human evidence |
| M7 | D | Keep simple participant actions while instrumenting full denominators and shared budgets | Usability and effectiveness untested |
Here T denotes a conditional theoretical result, I implementation evidence including explicitly synthetic experiments, and D a design proposal. These labels follow the repository claim vocabulary and do not confer scientific acceptance. Twenty-one unit tests include an independent representative panel enumeration, an independent finite retry-path tree, all switching sequences and utility comparisons on both sides of a feasible audit interval. Three negative results are required demonstrations, not optional caveats: hidden control, refusal-driven completion bias and unstable switching.
From the repository root:
python3 -m unittest discover -s papers/07-release-packet/models/math -p 'test_*.py' -v
python3 papers/07-release-packet/models/math/run_experiments.pyThe second command writes 115 parameter-sweep rows in four CSV files
and the 80,000-request Monte Carlo summary under
models/math/results/. The manifest records exact
source/table SHA-256 digests, seeds, Python version and interval
interpretation. Standard-library Python is sufficient. No service,
credential, network access or actual author data are used. Input
validation catches malformed shapes, negative work and non-finite scalar
inputs. The numerical routines are small research references; they are
not certified against overflow for every finite IEEE-754 value or
combinatorial explosion on large group populations.
The decisive next question is not whether these plots look plausible. It is whether the full event trace preserves the stated boundary when one toy actor refuses, splits its identity, hides a dependency or overloads a shared reviewer. A human pilot should then compare the four-action interface with an ordinary structured referee template with the same information and comparable declared support, measuring actual total labor. If the template works as well with less effort, retain the portable evidence records and simplify the surrounding machinery.