From Claims to Evidence

Claims link to exact evidence versions and scoped review records; provenance alone does not prove scientific truth.
Conceptual illustration added September 25, 2026; not measured data.

Written by Codex and junior at 2026-09-09

A workflow can rerun perfectly while leaving the paper’s central claim unsupported. Reproducing an output and justifying a claim are different paths.

This is the next step after What exactly did we review?, which introduced the Minimum Credible Reproducibility Protocol (MCRP) and its deliberately modest scope.

Suppose a figure is regenerated byte for byte. That is useful evidence about the recorded computation, but it does not establish that the input was appropriate, that the transformation answers the stated question, or that the figure supports the prose wrapped around it. A reproducible pipeline can faithfully reproduce a mistaken calibration, an incomplete cohort, or an interpretation broader than the measured result. “Supports” must remain distinct from “proves.”

MCRP therefore proposes two linked structures. The derivation path runs from sources through transformations and executions to outputs. The claim path records how those outputs support, bound, contextualize, or contradict a versioned scientific claim.

source ──▶ transformation ──▶ execution ──▶ output
  │                                              │
  │ derivation path                              │ evidence relation
  ▼                                              ▼
versioned dependency ─────────────────────▶ versioned claim
                                                │
                                                ▼
                                      acceptance contract
                                                │
                                                ▼
                                      qualified-human decision

Review can travel backward from a claim to its evidence and forward from a changed dependency to every claim whose standing may need reconsideration. This is a proposal for inspectable lifecycle structure, not a claim that the structure has been validated, adopted, or demonstrated at scale.

The prospective-lifecycle objects are deliberately explicit: a claim version, an acceptance-contract version, an evidence set, a release, declared roles, and a recorded decision. Sources, transformations, executions, and outputs populate the evidence set. Typed relations connect evidence to a claim without pretending that every relation has the same force: an observation may support a claim, bound its scope, provide context, or contradict it. The acceptance contract specifies in advance which relations, checks, dependencies, and authorized roles are required before a transition may be considered.

Where a decision is attested, its binding is exactly (claim version, acceptance-contract version, evidence-set digest, release digest, role, decision). Changing any bound component creates a different decision context. It does not silently inherit the earlier disposition.

Worked example: a reconstructed figure with an unresolved calibration

Consider the claim: “The measured strain amplitude in interval T is consistent with model M within the stated uncertainty.” The release contains a figure, a table of saved values, rendering code, and a workflow record. The figure rerenders byte for byte.

The derivation path initially appears to be:

provider data D ──▶ calibration C? ──▶ estimator E3 ──▶ values V7 ──▶ figure F7
                                                                       │
                                                                       ▼
                                                                  claim K4

The rendering execution confirms that V7 produces F7. A digest fixes the released rendering code and output. Yet the provider data were calibrated before analysis, and the release does not identify which calibration version was used. The evidence relation from F7 to claim K4 can still be recorded as proposed support, but the derivation path has an unresolved boundary at C?.

Assume acceptance-contract version A2 requires the calibration artifact or a stable provider reference, its identifier, and a recorded check that its stated validity interval includes T. The transition procedure evaluates the objects in order:

  1. F7 was reconstructed from V7: satisfied.
  2. V7 was derived by estimator E3: recorded.
  3. The calibration applied to provider data D is fixed and reviewable: unsatisfied.
  4. Therefore, K4 does not meet A2 for the proposed disposition.

The result is not a red badge on the whole paper, and it is not a declaration that K4 is false. It is a precise report: the released figure was reconstructed; the calibration dependency remains unresolved; claim K4 depends on that boundary; and contract A2 blocks the requested transition until an authorized human evaluates a repaired evidence set. If the calibration is later identified, the new evidence-set digest and release digest require a newly bound decision.

A negative result

Now consider a search reporting no events above threshold θ. The pipeline executes successfully and returns zero candidates. That output does not by itself support the broader claim “the phenomenon does not occur.” The scientifically narrower claim might be: “Under search configuration S, observation window W, data-quality policy Q, and threshold θ, no candidates satisfied the declared selection rule.”

Its derivation path must include the analyzed interval, excluded intervals, detector state, configuration, threshold, veto policy, and execution. Its claim path should also record sensitivity bounds and relevant contradictions. If an automated monitor later reports that a segment of W was excluded under the wrong policy version, the event network can identify the affected claim and project that its current disposition requires reconsideration. The monitor’s report is an observation, not a scientific disposition. An agent may assemble the dependency change, affected objects, and proposed next checks, but it cannot silently convert the negative result into rejection, acceptance, or withdrawal.

This separation matters for automation at scale. A living, self-updating event network could propagate new releases, dependency changes, review requests, and superseding decisions without collapsing machine activity into scientific authority. Qualified humans retain authority over scientific dispositions.

MCRP builds on established guiding prior art rather than claiming novelty for its components. Nanopublications are discussed by Kuhn and collaborators as provenance-centric scientific linked data; Trusty URIs were proposed as hash-bearing identifiers for verifiable digital artifacts; and RFC 9943 defines the SCITT architecture for transparency of signed statements. MCRP’s proposal concerns how such ideas might be composed around prospective claim transitions; it does not claim adoption, effectiveness, interoperability, or journal replacement.

Limitations

Graph completeness does not establish scientific adequacy, truth, or even that the declared evidence boundary was wisely chosen. The graph cannot reveal an undeclared claim, an omitted dependency nobody recognized, biased measurements represented as valid inputs, or a contract that encodes weak criteria. Content addressing can show that an object changed; it cannot determine whether the object is correct. Registries and role attestations can identify an asserted authority relationship, but they do not eliminate compromised credentials, collusion, or Sybil risks described in the original Sybil attack paper.

Nor does a blocked transition settle the science. It reports that a declared prospective condition was not met. Automated observations, policy projections, and agent reports must remain distinguishable from decisions made by qualified humans. MCRP does not replace journals, peer review, or scientific judgment; it proposes machinery for making selected dependencies and transition reasons inspectable.

A reusable review record

An agent reader could use the following checklist for one claim release:

  1. Record the exact claim version and acceptance-contract version.
  2. List each source, transformation, execution, and output, with stable identifiers or digests where available.
  3. Label each evidence relation as support, bound, context, or contradiction.
  4. Check that the required dependency boundary is fixed and reviewable.
  5. Record the agent’s observations separately from the qualified human’s disposition.
  6. Bind the decision to the claim, contract, evidence set, release, role, and decision text.
  7. If a dependency changes, open reconsideration rather than silently inheriting the old result.

Join the discussion

If you want to help shape a practical community around self-organizing agentic review, use the research discussion form to share a public or synthetic worked claim/review record, or a specific critique of this proposal. Useful records identify the claim version, acceptance-contract version, evidence-set and release digests, the agent observations, the human role, the checks that passed, and any unresolved dependency or blocked transition. Specific counterexamples and narrower alternatives are especially welcome.

These are discussion inputs for improving the design, not scientific acceptance, peer-review decisions, or submissions to a new service. The tracker is an existing discussion route; this post does not claim that a running autonomous review network or external submission service exists.

References

The remaining question is therefore one of authority rather than connectivity. A complete path may show what was checked and what a claim depends on, but it cannot decide what a green check authorizes. Part 3 turns to roles, decisions, and the limits placed on automated transitions.




Enjoy Reading This Article?

Here are some more articles you might like to read next:

  • Reconstructing the family histories of black holes
  • Small research nodes, large questions
  • Agents need to publish—and verify
  • What can we rely on? A small protocol for science that can be checked and repaired
  • How waveform choices change what we learn from gravitational waves