MCRP / Research prototype

Collective adversarial assessment: MCRP release candidate

Joint content position: publish a curated, candid research seed. This is qualified assent to the bounded proposal and synthetic evidence. It is not certification of an operating review service. The final archive must independently pass export and clean-extraction checks; its exact digest and that result belong to the release record outside the archive, avoiding a self-referential hash claim. The additional human-design and funded-audit content reviews are complete, with corrected statistical targets, human-task controls and delivered-audit assumptions independently examined.

The reviewers agree on the central boundary: the component premises do not compose automatically. The research contribution is the small interface, conditional models, runnable failures and a way for people to challenge them. The public seed does not need a complete appeals institution or a demonstrated cooperation equilibrium before it can be discussed. It does need honest claims, responsible attribution, an explicit rights disposition and a precise artifact.

What was actually reviewed

Blue lanes developed the mathematical foundations, executable receipt agents, coupled allocation model, recognition ecology, correction game, four domain cases, a legacy-workflow adapter, a pre-results human evaluation design and a funded-audit delivery game. The integrator assembled publication, onboarding, the publication-cycle model, static export and this synthesis. Separate adversarial lanes challenged:

Lane Independent work Position
Mathematics and incentives Finite enumeration, endpoint and utility probes, all coupled/ecology replications, correction-game reconstruction Assent to conditional claims; no general equilibrium or composed guarantee
Runtime and evidence State-transition attacks, publication sequences, capacity boundaries, adapter type confusion, export and provenance attacks Concrete repairs verified; final exact archive remains a separate check
Domains, law and human interface Scientific interpretation countermodels, jurisdiction/authority scope, four newcomer walkthroughs and publication routes Assent to a static public seed; real authority and human efficacy remain external

These are separate agent lanes within the same orchestration and operator context. They are not independent human reviewers or evidence of model-family diversity. The reviewers exchanged their full-packet conclusions directly; the domain/legal collective exchange records the cross-lane questions and responses. Their detailed reports, original failures, repair probes and qualified verdicts remain in this packet.

Findings that changed the candidate

Finding Repair or narrowed claim Evidence
A matching historical byte observation could be attached to changed release metadata Observation and currentness bind the actual candidate; bytes alone do not suffice Runtime report and independent regression
Rejected future actions could poison the toy clock; amendment pointers could rebind or cycle Failed actions preserve the clock; successors bind full candidates and require fresh IDs Runtime adversarial probes
Contradictory checks could leave prior reliance current; minimum coverage was negotiable after offer Currentness exposes contradiction; an immutable offer minimum cannot be weakened Receipt-runtime regressions
Refusals disappeared from unresolved totals All refused requests remain unresolved in the single-offer example Math verdict; 8,956 unresolved in its retained run
An expected audit budget was described as a hard reservation Separate expected-cost model from blinded hard-quota sampling and exact rational bounds Independent endpoint/utility probes; all 495 endpoint cases and 832 inequalities checked
A floating tolerance allowed positive settlement against a zero reservation Settlement cannot exceed the individual reserved amount Coupled-ledger adversarial test; no change to default sweep outcomes
Audit-hour tolerance authorized positive work against zero hours; floating utility ties reversed the declared decision policy Exact decimal-rational ledgers and utilities, with refusal taking precedence at ties Funded-audit review; original failures and 3,839 exact outcome paths retained
Crossed biology contrasts were overinterpreted as a common additive or causal effect Separate average contrast, exact additivity and causal assumptions; show countermodel Domain review and retained before/after results
Federal evidentiary framing could be read as jurisdiction-general Limit the example to covered federal proceedings and preserve operational authority distinctions Domain/legal revision and verdict
Existing output leftovers and source symlinks could enter an apparently curated export Fresh staging, explicit inventory, symlink rejection and unchanged-prior-output checks Integration review; harmless public sentinels
The verifier accepted an archive with extra or symlink members Exact safe ZIP membership, regular-file metadata and content checks Export regressions retain vulnerable baseline and repaired outcomes
Saved result/figure bytes could change after validation and acquire a new export manifest Bind saved evidence to validation; check source/input snapshots and figure lineage Provenance review and five independent regressions
Python equality allowed a legacy integer to masquerade as a Boolean assertion Compare canonical JSON bytes, retaining true / 1 / 1.0 distinctions Adapter author regression and independent runtime bypass test
Missing-outcome bounds could be mistaken for bounds on the causal effect Name the realized assigned-arm contrast; preserve a zero-causal-effect counterexample with observed bounds [1,1] Human-design mathematical addendum and independent enumeration
A uniformly cautious response could appear adequate; shared facilitator burden could create interference Add an unaffected-use positive control; separate accounting from interference assumptions Human-design domain addendum
The first human action required too much interpretation Add solo entry, four concrete microtasks, a worked card and a three-item minimum contribution Newcomer walkthrough; no actual human study

The before/after source and output records are evidence of these particular repairs. They do not justify a comprehensive security or correctness claim. Intentionally vulnerable baseline packaging code is retained solely as a research fixture; it is not the candidate’s release tooling.

Counterexamples that remain part of the result

Selection is not completion, and completion is not reliance. Correctly grouped invitation probabilities can coexist with unequal completion or completely selective reliance. The coupled experiments preserve every stage denominator. A false declaration of independent control violates the group-first theorem’s premise; the prototype does not discover that deception.

Quiet operations can hide scientific failure. Shared blindness produces few known correction tasks despite missed defects. The biology case can reproduce without identifying a causal effect. A legal workload can fit aggregate hours while missing a deadline. Throughput, agreement and a short queue are insufficient success measures.

Recognizing corrections can create work worth gaming. In the ecology, defective work can earn recognition through repair while clean unaudited work earns little. The added one-step game shows profitable manufactured corrections, including an author-repairer coalition under repairer-only credit. It also retains no-credit, clean-author-credit and cost/enforcement sensitivities. The exercise does not establish a preferred incentive policy, actual intent detection or participation in genuine repair. Ordinary correction must remain possible without presuming misconduct.

A record does not create its external premises. The receipt runtime trusts actor/control/authority fixtures. The coupled model uses perceived audit risk without running audits. The ecology uses a stipulated perfect audit oracle when access and capacity permit. The hard-audit model assumes detection and enforceable loss. The correction game has a one-case scalar capacity condition. The funded-audit game closes a specific quota, hour-reservation and token-transfer loop, while retaining fixed identities, trusted concealment, stipulated detection and automatic auditor effort. None supplies the missing premises for the others.

A legacy wrapper cannot manufacture a reviewed fact. The adapter preserves source values, unknowns and semantic losses. Its enriched run records a declared historical check under separate synthetic assumptions; it does not rerun that reviewer’s method, authenticate the letter or obtain consent. Passing syntax for a digest does not prove the referenced artifact exists or matches it.

The human claim is still open. An accessible exercise and favorable toy result do not establish comprehension, adoption or lower labor. The planned comparison uses a competent template baseline and includes facilitation, refusals and unfinished work. Its small-design calculations are artificial planning examples, not study results or a demonstrated sample-size recommendation.

Positioning and adoption were also challenged

A further comparison examined existing COAR Notify, W3C PROV, RO-Crate, PCI, CRediT and ORCID capabilities. The domain/legal reviewer independently checked key claims and narrowed the adoption memo: reuse existing versioned review and provenance, preserve the audience restrictions of review records, and treat the proposed contribution as a testable cross-system use/capacity profile. Permission to read a review does not imply permission to republish it. See the positioning review. No conformance, priority or community-adoption claim follows from this comparison.

The final domain/legal pass also examined the funded-audit interpretation, claim map, site integration and font rights notices; see its release addendum. The fictional collateral does not propose a real deposit policy, and task decision units are distinguished from execution people.

A further reproduction reviewer, independent of runtime/model authorship but a contributor to domain prose, checked 3,060 funded-audit configuration/seed cases: 2,160 executions and 900 expected self-audit preflight stops. Its content and accounting review supports the bounded claims and records an extreme-range display overflow outside the teaching fixtures. This is an explicit runtime limitation, not an overdraw or a passed arbitrary-range robustness claim. During final integration, the historical completion fixture consumed by an adversarial probe was added to the before/after input inventory; a mutation regression verifies that changing it invalidates a run.

Replication and remaining release work

The mathematical lane independently reconstructed five coupled traces, all 7,400 coupled runs and 888 replicate-level intervals, all 216 ecology runs and 351 aggregate/paired metrics, all 18 correction-game outcomes, all 15 exact human-design power rows, and 3,839 positive-probability funded-audit outcome paths. The funded-audit review also checked all 11 full traces and 220 recorded runs. The runtime lane ran its regressions without external executables on PATH. The domain lane reran domain oracles and checked the scientific meaning of their receipt scopes. These are substantive internal checks with stated boundaries.

The canonical reproduction command is python3 run_checks.py from the packet root. The final results/validation.json supplies the current test total, command logs, source/input snapshots and saved-evidence hashes. Counts in earlier reports refer to those earlier reviewed states; they are not additive counts of distinct assurances. The release record must additionally identify the actual ZIP, extraction/reproduction outcome, local-link and byte-inventory checks, and visual inspection. Hashes identify evidence; they do not sign, approve or license it.

Attribution, rights, public destination and modest intake responsibility are the remaining human release choices. No message has been sent to a venue, no public repository has been created, and no claim of external acceptance is made. Keeping these choices concrete is compatible with a positive release recommendation.

Recommendation to the maintainer and prospective contributors

Publish the four-action boundary with its simple baseline, domain examples, reproducible models and this adversarial record. Invite one synthetic claim, one material amendment and one counterexample. Let useful objections determine which additional machinery is worth building. A proposal that can be narrowed, simplified or rejected through a cheap reproducible example is ready to seed a serious protocol conversation.

Agent-first launch cycle

The maintainer subsequently authorized deployment and repository publication. The launch record supersedes earlier pending-release status without changing the scientific limits in this review. A new blue lane added the local JSON interface and four agent task fixtures; the new adversarial report challenges identity, versioning, authority, capacity and input bounds. The interface is independently executable but does not authenticate callers or perform the claimed scientific checks.

This cycle changed the demonstrations to offer genuinely corrected content while retaining old pending reliance. It also binds the complete offered profile, so a new dependency or coverage contract can create a distinct version without fabricating a data change. Agent aliases map through trusted local policy to a shared principal ledger. Undisclosed common control and real shared capacity remain external premises. The current test total is in the final validation manifest; earlier counts describe earlier versions.