MCRP / Research prototype

Adversarial review: domain meaning, legal operations, and public release

Review date: 25 September 2026. This agent independently inspected the four domain cases, numerical fixtures, actual receipt replay, publication documents, both onboarding paths, and the first dossier’s collective assessment. It is an internal adversarial lane sharing the same operator context, not external peer review or legal clearance.

Judgment

Release a corrected, curated research seed. Do not represent it as an operating review or dispute service. The distinction is substantive: the domain cases invite bounded synthetic exercises; they do not ask a patient, litigant, editor, or laboratory to transfer a consequential decision to the toy engine. The small interface remains useful even when its honest answer is unknown or unavailable.

The strongest material is the juxtaposition of a useful narrow check and a tempting but unjustified broader use: code reconstruction versus calibration, assay arithmetic versus identification, a trial mixture versus a transported policy effect, and citation accuracy versus institutional authority. The release should keep these failures in the first encounter, not bury them in legal text.

Two scientific wording/model boundaries and one legal scope ambiguity should be corrected before export. None requires delaying public discussion until a full institution exists. Attribution, rights, exact public bytes and destination still need the responsible human’s concrete release decision; the agent cannot supply those facts by consensus.

Independent checks

All 15 domain tests passed. Both domain_models.py and replay_receipts.py ran independently. The actual replay makes the initial use current under its toy policy, changes it to reconsideration pending after the declared amendment, and admits a newly checked version. Biology rejects an unsupported broad scope; law rejects unsupported legal_operational authority. The hidden-dependency control correctly records a known miss rather than announcing full correction recall.

These are implementation observations from the inspected version. They do not prove that declarations reflect real facts, that an actual authority is appointed, or that notices reach people. The new counterexamples.py beside this report adds an independent conditional-mean counterexample, a causal countermodel, an interaction probe, and a final-selection counterexample.

Findings

D1 — State conditional exogeneity before deriving the biological conditional mean

Correction required; mathematical precondition. The original biology case stipulates errors with mean zero, then writes the conditional mean as mu + (beta + gamma) T when T=B. Unconditional mean zero is insufficient. Let T=B be Bernoulli one half and epsilon=T-1/2. Then E[epsilon]=0, but the conditional errors are minus one half and plus one half. The displayed conditional-mean expression has omitted a term.

Specify E[epsilon | T,B]=0 for the conditional-mean model. The duplicate-column nonidentifiability result itself remains valid: only a sum of coefficients enters the fitted values. This is a narrow repair to an otherwise correct teaching case.

D2 — Crossing batch and treatment does not by itself establish causation or exact additivity

Correction required in extensible oracle labeling; preserve a named boundary in the teaching prose. The supplied noiseless four cells have equal within-batch contrasts. That supports the stated additive arithmetic illustration. Yet the same table also arises from Y=10+2B+U with observed T=U: intervening on T does not change U, so the causal effect is zero. A full-rank observed design does not supply randomization, exchangeability, consistency or absence of interference. The case mostly observes this distinction, but an implementation must keep additive_model_identification from acquiring a general causal meaning.

The original numerical oracle accepts crossed outcomes 10, 11, 12, 15. It reports within-batch contrasts 1 and 3 and their average 2 as crossed_additive_effect. There is no single coefficient fitting these noiseless cells in the declared additive model. The receipt replay already checks equal contrasts before its amended support. Give the numerical oracle the same semantic honesty: report the average as an average and expose failure of exact additivity, or reject this input outside the fixture’s stated domain. If fitting a projection instead, name its weights and lack of fit. Do not suppress a legitimate interaction as mere error.

D3 — Preserve the federal jurisdiction qualifier in the evidence-rules example

Correction required; low-cost legal scope repair. The original law case says the Federal Rules of Evidence govern most U.S. court proceedings. The cited official page refers to proceedings in the “United States courts,” a federal institutional term. The paraphrase can be read as including state proceedings. Use “covered federal proceedings, subject to the rules’ applicability and exceptions.” No universal U.S. or transnational evidence rule follows.

The official Forms and Rules overview expressly places the federal rules in federal litigation. The official current evidence-rules page was independently fetched. A House Rule 1101 search result corroborated the specified courts, but its direct retrieval failed; this review does not claim full inspection of that page or a current litigation opinion.

D4 — Every decision stage can concentrate influence again

Research boundary; not a release blocker if preserved. The economics case’s completion-selection identity is correct. The later reliance decision can select again. Ten adverse and ninety other completed reviews give an adverse completion share of one tenth. If the editor relies only on the adverse ten, the relied share is one. A correct invitation algorithm and complete completion ledger do not constrain this choice. The first dossier already calls for reliance denominators; carry them into the toy experiment when evaluating capture.

This does not imply every institution must use reviewers proportionally. It means the project must say whether it is measuring representation, expertise, decision quality, or a restriction on discretionary authority. These objectives can conflict. The suggested group-first baseline solves only its specified label invariance problem, not undeclared common control or legitimate heterogeneous expertise. Sanctions and audit rewards remain assumed institutional inputs.

D5 — A source can support a narrow claim without certifying a dissemination route

Satisfactory after the packet’s existing exclusions. Independently fetched OSF’s policy states that mostly or completely AI-generated material is inappropriate for OSF Preprints. The packet correctly excludes that immediate route, including its hosted community servers. This restriction is not a universal ban on public prototype pages or evidence that a cosmetic human rewrite creates eligibility. Do not hide the AI contribution record to gain a submission endpoint.

Independently fetched AEA registration guidance supports the economics case’s distinction between registration and quality or ethics review. It also limits the registry to social-science randomized trials; it is not a general protocol archive or clinical registry. The packet does not propose using it as one. arXiv, bioRxiv, SSRN, JOSS and community discussion remain conditional routes with different eligibility questions, not promised acceptance. This review did not re-fetch every such venue or independently verify a future submission’s eligibility.

D6 — A public research artifact needs an honest responsible party, not a fictional institution

Concrete publication decision, not missing scientific validation. The release assessment correctly recommends a fresh allowlisted export instead of making private repository history public. A source digest identifies bytes; it does not grant rights. AI drafting disclosure is not a human author appointment, and a public Git URL is not an open-source license. The draft explicitly leaves rights-holder approval, license scope, responsible attribution, destination and final artifact approval pending. Keep those states visible until actually done.

An identified maintainer and a modest correction route are enough to start a research conversation. Do not require the entire proposed appeal institution before permitting a static document to be read. Conversely, a public issue tracker will receive unexpected personal allegations unless its purpose and a responsible contact are clear. Publication is not an assurance that no legal duties can arise; real intake and contested-service operation must be reassessed on their facts. Do not promise blanket confidentiality for a counter-notice. The independently fetched Copyright Office statutory text supports conditional forwarding and timing obligations described in the case. No real notice, statutory calendar, jurisdiction decision or safe harbor was tested.

D7 — Human adoption should be judged against a genuinely easy alternative

Next experiment; no release blocker. The six-question card is accessible and the paired exercise permits unknowns and withdrawal. That is a credible entry path, but its ten- and forty-five-minute budgets are proposals, not measured completion times. Three review formats within one domain clinic can create large learning and facilitation effects. Use matched tasks, counterbalanced order, total author/checker/facilitator labor and refusal denominators as proposed.

The first public invitation should show a completed card before asking for a blank one, and let a novice contribute one well-scoped sentence without adopting the vocabulary or a steward role. This is an engagement recommendation, not a reason to add mandatory fields. The protocol loses if a plain claim/evidence/ limitation paragraph yields equal comprehension at lower total effort.

What the public message can defend

The blog can say: this prototype makes a claim’s exact version, check scope, intended use and correction boundary explicit; the code demonstrates how such records behave under specified toy assumptions; the cases expose failure modes that a generic positive badge hides. It can invite people to break these examples.

It cannot yet say: reviewers are independent, truth is certified, legal rights are resolved, corrections reliably reach recipients, human workload falls, or the four domains have adopted a shared standard. The current blog avoids these claims. Different expertise remains necessary even though the external record vocabulary is small.

Provisional collective position

Assent to a corrected curated public prototype, conditional on D1–D3 repairs and the exact export checks. Mathematical/runtime reviewers’ unresolved concrete failures must also be reflected before release. No domain or legal objection here requires suppressing the protocol seed until all operational questions are solved. No assent substitutes for the human attribution, rights and destination decision.