# Collective adversarial assessment: MCRP release candidate

**Joint content position: publish a curated, candid research seed.** This is
qualified assent to the bounded proposal and synthetic evidence. It is not
certification of an operating review service. The final archive must independently
pass export and clean-extraction checks; its exact digest and that result belong
to the release record outside the archive, avoiding a self-referential hash claim.
The additional human-design and funded-audit content reviews are complete, with corrected statistical targets, human-task controls and delivered-audit assumptions independently examined.

The reviewers agree on the central boundary: **the component premises do not
compose automatically**. The research contribution is the small interface,
conditional models, runnable failures and a way for people to challenge them.
The public seed does not need a complete appeals institution or a demonstrated
cooperation equilibrium before it can be discussed. It does need honest claims,
responsible attribution, an explicit rights disposition and a precise artifact.

## What was actually reviewed

Blue lanes developed the mathematical foundations, executable receipt agents,
coupled allocation model, recognition ecology, correction game, four domain
cases, a legacy-workflow adapter, a pre-results human evaluation design and a funded-audit delivery game.
The integrator assembled publication, onboarding, the publication-cycle model,
static export and this synthesis. Separate adversarial lanes challenged:

| Lane | Independent work | Position |
|---|---|---|
| Mathematics and incentives | Finite enumeration, endpoint and utility probes, all coupled/ecology replications, correction-game reconstruction | Assent to conditional claims; no general equilibrium or composed guarantee |
| Runtime and evidence | State-transition attacks, publication sequences, capacity boundaries, adapter type confusion, export and provenance attacks | Concrete repairs verified; final exact archive remains a separate check |
| Domains, law and human interface | Scientific interpretation countermodels, jurisdiction/authority scope, four newcomer walkthroughs and publication routes | Assent to a static public seed; real authority and human efficacy remain external |

These are separate agent lanes within the **same orchestration and operator
context**. They are not independent human reviewers or evidence of model-family
diversity. The reviewers exchanged their full-packet conclusions directly; the
[domain/legal collective exchange](red-domains-law/collective.md) records the
cross-lane questions and responses. Their detailed reports, original failures,
repair probes and qualified verdicts remain in this packet.

## Findings that changed the candidate

| Finding | Repair or narrowed claim | Evidence |
|---|---|---|
| A matching historical byte observation could be attached to changed release metadata | Observation and currentness bind the actual candidate; bytes alone do not suffice | [Runtime report](red-runtime/report.md) and independent regression |
| Rejected future actions could poison the toy clock; amendment pointers could rebind or cycle | Failed actions preserve the clock; successors bind full candidates and require fresh IDs | Runtime adversarial probes |
| Contradictory checks could leave prior reliance current; minimum coverage was negotiable after offer | Currentness exposes contradiction; an immutable offer minimum cannot be weakened | Receipt-runtime regressions |
| Refusals disappeared from unresolved totals | All refused requests remain unresolved in the single-offer example | [Math verdict](red-math/verdict.md); 8,956 unresolved in its retained run |
| An expected audit budget was described as a hard reservation | Separate expected-cost model from blinded hard-quota sampling and exact rational bounds | Independent endpoint/utility probes; all 495 endpoint cases and 832 inequalities checked |
| A floating tolerance allowed positive settlement against a zero reservation | Settlement cannot exceed the individual reserved amount | Coupled-ledger adversarial test; no change to default sweep outcomes |
| Audit-hour tolerance authorized positive work against zero hours; floating utility ties reversed the declared decision policy | Exact decimal-rational ledgers and utilities, with refusal taking precedence at ties | [Funded-audit review](red-math/funded-audit-addendum.md); original failures and 3,839 exact outcome paths retained |
| Crossed biology contrasts were overinterpreted as a common additive or causal effect | Separate average contrast, exact additivity and causal assumptions; show countermodel | [Domain review](red-domains-law/report.md) and retained before/after results |
| Federal evidentiary framing could be read as jurisdiction-general | Limit the example to covered federal proceedings and preserve operational authority distinctions | Domain/legal revision and verdict |
| Existing output leftovers and source symlinks could enter an apparently curated export | Fresh staging, explicit inventory, symlink rejection and unchanged-prior-output checks | [Integration review](red-runtime/integration-review.md); harmless public sentinels |
| The verifier accepted an archive with extra or symlink members | Exact safe ZIP membership, regular-file metadata and content checks | Export regressions retain vulnerable baseline and repaired outcomes |
| Saved result/figure bytes could change after validation and acquire a new export manifest | Bind saved evidence to validation; check source/input snapshots and figure lineage | [Provenance review](red-runtime/provenance-adapter-review.md) and five independent regressions |
| Python equality allowed a legacy integer to masquerade as a Boolean assertion | Compare canonical JSON bytes, retaining true / 1 / 1.0 distinctions | Adapter author regression and independent runtime bypass test |
| Missing-outcome bounds could be mistaken for bounds on the causal effect | Name the realized assigned-arm contrast; preserve a zero-causal-effect counterexample with observed bounds [1,1] | [Human-design mathematical addendum](red-math/human-pilot-addendum.md) and independent enumeration |
| A uniformly cautious response could appear adequate; shared facilitator burden could create interference | Add an unaffected-use positive control; separate accounting from interference assumptions | [Human-design domain addendum](red-domains-law/human-design-addendum.md) |
| The first human action required too much interpretation | Add solo entry, four concrete microtasks, a worked card and a three-item minimum contribution | [Newcomer walkthrough](red-domains-law/onboarding-review.md); no actual human study |

The before/after source and output records are evidence of these particular
repairs. They do not justify a comprehensive security or correctness claim.
Intentionally vulnerable baseline packaging code is retained solely as a research
fixture; it is not the candidate's release tooling.

## Counterexamples that remain part of the result

**Selection is not completion, and completion is not reliance.** Correctly grouped
invitation probabilities can coexist with unequal completion or completely
selective reliance. The coupled experiments preserve every stage denominator.
A false declaration of independent control violates the group-first theorem's
premise; the prototype does not discover that deception.

**Quiet operations can hide scientific failure.** Shared blindness produces few
known correction tasks despite missed defects. The biology case can reproduce
without identifying a causal effect. A legal workload can fit aggregate hours
while missing a deadline. Throughput, agreement and a short queue are insufficient
success measures.

**Recognizing corrections can create work worth gaming.** In the ecology, defective
work can earn recognition through repair while clean unaudited work earns little.
The added one-step game shows profitable manufactured corrections, including an
author-repairer coalition under repairer-only credit. It also retains no-credit,
clean-author-credit and cost/enforcement sensitivities. The exercise does not
establish a preferred incentive policy, actual intent detection or participation
in genuine repair. Ordinary correction must remain possible without presuming
misconduct.

**A record does not create its external premises.** The receipt runtime trusts
actor/control/authority fixtures. The coupled model uses perceived audit risk
without running audits. The ecology uses a stipulated perfect audit oracle when
access and capacity permit. The hard-audit model assumes detection and enforceable
loss. The correction game has a one-case scalar capacity condition. The funded-audit game closes a specific quota, hour-reservation and token-transfer loop, while retaining fixed identities, trusted concealment, stipulated detection and automatic auditor effort. None supplies
the missing premises for the others.

**A legacy wrapper cannot manufacture a reviewed fact.** The adapter preserves
source values, unknowns and semantic losses. Its enriched run records a declared
historical check under separate synthetic assumptions; it does not rerun that
reviewer's method, authenticate the letter or obtain consent. Passing syntax for
a digest does not prove the referenced artifact exists or matches it.

**The human claim is still open.** An accessible exercise and favorable toy result
do not establish comprehension, adoption or lower labor. The planned comparison
uses a competent template baseline and includes facilitation, refusals and
unfinished work. Its small-design calculations are artificial planning examples,
not study results or a demonstrated sample-size recommendation.

## Positioning and adoption were also challenged

A further comparison examined existing COAR Notify, W3C PROV, RO-Crate, PCI,
CRediT and ORCID capabilities. The domain/legal reviewer independently checked
key claims and narrowed the adoption memo: reuse existing versioned review and
provenance, preserve the audience restrictions of review records, and treat the
proposed contribution as a testable cross-system use/capacity profile. Permission
to read a review does not imply permission to republish it. See
[the positioning review](red-domains-law/positioning-review.md). No conformance,
priority or community-adoption claim follows from this comparison.

The final domain/legal pass also examined the funded-audit interpretation, claim map, site integration and font rights notices; see its [release addendum](red-domains-law/funded-release-addendum.md). The fictional collateral does not propose a real deposit policy, and task decision units are distinguished from execution people.

A further reproduction reviewer, independent of runtime/model authorship but a contributor to domain prose, checked 3,060 funded-audit configuration/seed cases: 2,160 executions and 900 expected self-audit preflight stops. Its [content and accounting review](red-reproduction/content-review.md) supports the bounded claims and records an extreme-range display overflow outside the teaching fixtures. This is an explicit runtime limitation, not an overdraw or a passed arbitrary-range robustness claim. During final integration, the historical completion fixture consumed by an adversarial probe was added to the before/after input inventory; a mutation regression verifies that changing it invalidates a run.

## Replication and remaining release work

The mathematical lane independently reconstructed five coupled traces, all 7,400
coupled runs and 888 replicate-level intervals, all 216 ecology runs and 351
aggregate/paired metrics, all 18 correction-game outcomes, all 15 exact human-design power rows, and 3,839 positive-probability funded-audit outcome paths. The funded-audit review also checked all 11 full traces and 220 recorded runs. The runtime lane
ran its regressions without external executables on `PATH`. The domain lane
reran domain oracles and checked the scientific meaning of their receipt scopes.
These are substantive internal checks with stated boundaries.

The canonical reproduction command is `python3 run_checks.py` from the packet
root. The final `results/validation.json` supplies the current test total, command
logs, source/input snapshots and saved-evidence hashes. Counts in earlier reports
refer to those earlier reviewed states; they are not additive counts of distinct
assurances. The release record must additionally identify the actual ZIP,
extraction/reproduction outcome, local-link and byte-inventory checks, and visual
inspection. Hashes identify evidence; they do not sign, approve or license it.

Attribution, rights, public destination and modest intake responsibility are the
remaining human release choices. No message has been sent to a venue, no public
repository has been created, and no claim of external acceptance is made. Keeping
these choices concrete is compatible with a positive release recommendation.

## Recommendation to the maintainer and prospective contributors

Publish the four-action boundary with its simple baseline, domain examples,
reproducible models and this adversarial record. Invite one synthetic claim, one
material amendment and one counterexample. Let useful objections determine which
additional machinery is worth building. A proposal that can be narrowed, simplified
or rejected through a cheap reproducible example is ready to seed a serious
protocol conversation.

## Agent-first launch cycle

The maintainer subsequently authorized deployment and repository publication.
The [launch record](../publication/launch.md) supersedes earlier pending-release
status without changing the scientific limits in this review. A new blue lane
added the local JSON interface and four agent task fixtures; the [new adversarial
report](agent-launch-red/REPORT.md) challenges identity, versioning, authority,
capacity and input bounds. The interface is independently executable but does
not authenticate callers or perform the claimed scientific checks.

This cycle changed the demonstrations to offer genuinely corrected content while
retaining old pending reliance. It also binds the complete offered profile, so
a new dependency or coverage contract can create a distinct version without
fabricating a data change. Agent aliases map through trusted local policy to a
shared principal ledger. Undisclosed common control and real shared capacity
remain external premises. The current test total is in the final validation
manifest; earlier counts describe earlier versions.
