MCRP / Research prototype

Independent red review: mathematics, incentives and allocation

Review scope: manuscript math-foundations.md, models/math/, sampler and completion demonstrations in toy_agents/, and inherited dossier-06 selection and repair objections. This is an adversarial analytical review, not a security or legal certification. Original inspected inputs are preserved as original_*; re-review of author repairs must be explicit.

Recommendation

Suitable for a public research seed after two concrete corrections: distinguish expected audit spending from a hard capacity guarantee, and repair the toy completion denominator. The central group-distribution, completion-conditioning and common-envelope propositions are correct under their stated premises. The current results do not establish practical identity independence, human completion floors, review quality, feasible real schedules or a deployable equilibrium. Those limits are substantial but compatible with an honest prototype release.

Release corrections

RM-1: Expected audit cost is called worst-case reservation (high)

The original Proposition 4 constrains Naq<=B, then calls this a worst-case invited reservation. With independent Bernoulli audit decisions, Naq is the expected cost. At N=100, q=.2, a=.1, B=2 the exact probability of exceeding the budget is 0.4405384151266024, not zero. This cannot certify the shared attention constraint that the same manuscript rightly emphasizes.

The utilities and expected-budget feasibility interval are correct. Rename their budget meaning. For a hard budget with identical audit cost and a fixed invitation population, uniformly select exactly k<=floor(B/a) invitation indices (or mix bounded k values) before seeing behavior, conceal assignment until it can no longer alter effort, and deliver funded audits. Each invitation can have marginal q=k/N despite negative dependence. A hard-cap marginal requires q<=floor(B/a)/N, not merely B/(Na): N=3,a=1,B=1.5 permits at most q=1/3 equally, whereas the expected inequality allows q=1/2. If invitations, actual costs, audit eligibility or strategic participation change, restate the design. Fixed-size sampling is not itself credible funding, commitment, independence of auditors or legal authority to impose loss.

Evidence: independent exact binomial enumeration and fixed-size sample enumeration in probes.py; original wording/code snapshots retained. No attack on the algebra of one-shot expected utilities is intended.

RM-2: Single-offer refusals disappear from unresolved requests (high)

Original completion_bias() reported 10,000 offers, 1,044 completions, 8,956 refusals, zero unresolved, no retries and no reliance. Without a distinct terminal rejection/withdrawal state, the remaining 8,956 requests are unresolved. The published denominator vector is internally misleading even though the completed-risk fraction is calculated correctly. Repair the counter and add an accounting identity test covering the no-retry scenario. If a later model distinguishes withdrawn, rejected and unresolved requests, count these terminal outcomes explicitly; refusal attempts and unresolved requests must not be conflated once retries exist.

Evidence: original_completion.json, generated by executing the original function, and original_scenarios.txt. The independent probe asserts the preserved defect so an author repair cannot erase history.

Theorem boundary audit

Sampling. Proposition 1 is the elementary pushforward identity and is correct. It depends on fixed G,F,Q, a normalized representative kernel and accurate control mapping. The executable sampler applies one skill and static group exclusions, then leaves capacity reservation to the caller. It does not solve arbitrary feasible panels. Correctly disclosed. Adding representatives with different capacity/qualification is not the stated cloning transformation. Capacity-conditioned retry or operator-controlled random seeds can change the realized panel distribution; these require fresh end-to-end experiments, not a claim that the fixed-set theorem failed.

Completion. Proposition 2 is correct; all four seeded Monte Carlo cases reproduce bit-for-bit. Independently solving the Wilson score quadratic reproduces every reported completed-risk interval to <1e-12. These are pointwise simulation intervals conditional on supplied behavior, not error bars for scientific accuracy. IID retries do not cure completion bias. The envelope can be strengthened to adapted policies if bounds hold on every reached conditional history; see adaptive-completion.md for proof and exact finite-tree probes. The IID expected-cost formula must not be carried over.

Reliance selection. Completion is not the final selection stage. A .1 risky share with a=b=1 stays .1 on completion, but an actor that relies only on risky outcomes produces share one among reliance decisions. This is not a protocol violation: reliance is a discretionary scoped act. It is a warning against reusing completed-panel bounds as properties of downstream reliance. The public demo should show the difference and its denominators.

Repair. Proposition 3’s conditional bound, adaptive matrix choice and monotone-convergence proof are correct. Positive weights are abstract unless they dominate actual per-item hours. Snapshot spectral radii and unconditional means are inadequate. Additional counterexample: choose a permanent hidden high-reproduction regime with probability .25 (each item produces two) and extinction regime otherwise. First-generation mean is .5, but E[Z_10]=256. Conditional on observing offspring, the next reproduction multiplier is two. This violates the conditional envelope premise, not the theorem. Estimating a pooled mean matrix and passing it to repair_bound would therefore be an invalid operational bridge.

Audit behavior. Honest work weakly dominating alternatives does not predict universal honest participation. The stated tie-breaking convention prefers abstention; at a participation boundary a feasible weak interval may coexist with zero participation. The manuscript already labels the result weak and the sweep synthetic. Retain that distinction. Reward R and enforceable F are exogenous; funding only audits does not fund rewards, establish collections or make sanctions legitimate. Public claims should say exactly which budget is constrained.

Validation performed

These checks validate bounded calculations, not the whole public release or live deployment. statistical-verification.json records interval comparisons; probe-results.json records counterexamples. probes.py is standard-library-only and writes only its own results file.

Strongest simple next experiments

  1. One coupled queue, one person budget. Draw group panels, attempt whole-panel capacity reservation, record refusals and partial effort, retry under a fixed policy, and let reliance select completed work. Compare initial, reserved, completed and relied distributions. Vary only clone count and refusal behavior first. Expose capacities by true person/control ID, with a deliberately incorrect declaration negative control.
  2. Hard-funded audit game. Compare Bernoulli expected funding with uniform fixed-size concealed assignment, then let agents observe premature audit disclosure or the exhausted budget. Count unmet promised audits and deviations from honest best responses. Add outside options and reward funding separately.
  3. Regime-shift repair. Give the scheduler a pooled empirical estimate while a hidden persistent regime or adaptive attacker changes offspring. Measure conditional-envelope violations, deadline misses, tail workload and reserved/actual hours. A detector should pause new commitments, not pretend that observation repairs the bound.
  4. Human template comparison. Compare the four-action boundary to an ordinary structured referee form under equal attention budgets. Primary outcomes: correctly bounded reliance, update recognition, unresolved work and total participant time. Protocol complexity must earn its cost.

Collective position

The research seed is publishable with clear prototype labeling and corrected accounting. A persuasive launch should lead with what a reader can falsify in ten minutes: representative multiplicity, completion selection, stale dependency awareness and shared capacity. It should not lead with generic trust claims. The most useful next contribution is a small coupled simulator that crosses these boundaries, not a larger menu of disconnected models.