Agents need to publish—and verify
AI contribution: Codex drafted this proposal and its diagram at Richard O’Shaughnessy’s request. It describes a proposed process, not an operating certification service. This byline discloses authorship; it is not a cryptographic signature. A human scientific review of this post is not recorded here.
An agent that can write a page but cannot release it leaves a human holding every routine update. An agent that can publish anything it writes creates a different problem: nobody can tell which checks mattered, who authorized the release, or whether the public page is the one that was reviewed.
Our next objective for the Minimum Credible Reproducibility Protocol (MCRP) is to make bounded autonomous publication and independent verification work together. We need a usable process soon, beginning with a small demonstrable release cycle. The proposed publication policy defines that target and the tests required before enabling it.
Permission to publish is a bounded contract
The first useful delegation could be small: publish a public-source bibliography update, at an allowed destination, within an agreed frequency and an expiry date. It should say which sources are allowed, which checks are mandatory, and who can revoke the permission. The author cannot quietly expand its own contract.
A new scientific interpretation belongs in a different class. Agents can prepare the evidence and prose, but explicit editorial approval of the release—and qualified human judgment of scientific claims—remain distinct requirements. This continues the separation developed in From Claims to Evidence. Autonomy is useful precisely when its scope is understandable.
A second agent is not automatically an independent reviewer
A verifier needs the actual artifacts, the sources, and a fixed checklist. It should retrieve evidence and rerun checks rather than accept the author’s account of success. Its approval credentials must be separate from the publisher’s authority. Shared models, infrastructure, and operators still create shared failure modes; a second model’s agreement is not experimental replication.
Consider an agent-generated figure. The verifier checks the plotted values against the saved data, reruns the declared transformation, and asks whether the caption stays within the measured result. “The code ran” and “the figure supports this claim” remain different findings. Unperformed checks must be visible.
A signature says who attested to which bytes
A signed release can make later changes detectable and connect a statement to an allowed identity. It cannot make a mistaken claim true. Readers need separate labels for provenance verification, passed release checks, and scientific review.
We can build on existing tools: SLSA provenance describes where and how software artifacts were produced, while Sigstore verification checks signatures with explicit identity and issuer constraints. Adopting either requires implementation and a clear trust policy; mentioning the standard is not conformance.
The last check happens outside the build
After deployment, an observer fetches the public page and its assets and compares them with the approved release. A green build does not establish that readers received the right page. The observer records a dated receipt or reports failure.
Corrections then become new, linked releases. A changed figure does not inherit the previous check. A revoked publisher or invalidated dependency triggers reconsideration of affected claims. Public history should explain what changed, while private data remains private.
The immediate milestone is concrete: one synthetic staging release, deliberate failures that prove the gates block bad candidates, and then one narrowly delegated operational note with an independently checked public receipt. The goal is an agent that can finish useful work while leaving enough evidence for someone else to decide what to trust.