Final addendum: prospective human evaluation and full synthesis
25 September 2026. Independently reviewed
onboarding/evaluation-plan.md, the human-pilot code
documentation and tests, manuscripts/synthesis.md, the
revised blog and the assembled collective assessment. Seven human-design
test methods passed in this lane’s run. Mathematical randomization and
power verification was also assigned to the separate mathematics
reviewer; this addendum concerns the meaning of the design and its
public claims. No people were studied.
What the plan gets right
The competent structured-template baseline receives the same source information and amendment notices. That avoids comparing a richly supported protocol against an artificially impoverished control. Frozen first-use parallel allocation avoids claiming that participants can unlearn concepts before crossover. The accountable boundary is the unit: a collaboration’s many representatives do not inflate the sample size. Pre-treatment blocks, fixed two-of-four assignment and concealed allocation are stated separately from the published reproduction seed.
The primary outcome concerns adequate responses to both specified changes, not receipt count or self-reported confidence. Total person-minutes include help, assessment, adjudication and unresolved follow-up, with operational and research measurement costs separated. The horizon limits the estimand rather than silently declaring later work free. Withdrawals preventing retained data use remain missing; pre-allocation refusal is not a randomized failure. These are appropriate design choices, conditional on a future human-led implementation.
The artificial power family is explicitly optimistic, unfitted and incapable of justifying separate effects across four domains. Neither failure to reject nor visual similarity is called equivalence. No institutional approval, recruitment, registration, ethics waiver or human efficacy is implied by generating the planning files. A teaching clinic can change examples and welcome coaching; that activity is correctly distinguished from a frozen comparative study.
Three concrete repairs made to the evaluation prose
1. A cautious script must not pass by refusing every possible use
Both primary cases originally involved a change that defeats or leaves unresolved some current use. A participant could mechanically refuse continuation, request another check and preserve a historical label without showing that they can distinguish affected from unaffected statements. That would reward a generic caution strategy instead of calibrated scope interpretation.
Case N now includes a positive control in the same supplied
information: changing the offset from 8 to 12 changes
10 - offset, but the unaltered supplied raw scalar 10
remains positive. An adequate answer should explain that distinction; it
must not declare every old statement false or unsupported. Recognition
of this control is also a secondary outcome. The participant need not
undertake a real action, and declining the exercise remains allowed.
Both arms receive the same control. This is a proposed rubric repair,
not evidence that the new task is understood or easier. Domain experts
still need to check it before freezing a study.
2. Dividing facilitator time is not the same as eliminating interference
The original plan had a useful overhead allocation rule and prevented cross-arm coaching. Yet a shared facilitator can be fatigued or occupied by one arm’s questions, changing the help available to another boundary. Then an outcome depends on the assignment vector, not only that boundary’s format. Equal division of overhead is an accounting choice; it does not create an individual potential outcome under no interference.
The plan now explicitly requires prospective support availability, session order and overhead rules, plus congestion/deviation records. If material spillovers cannot be prevented, the design must change its randomization unit or define a service-regime estimand before recruitment. This does not demand unnecessary institutional machinery for a casual teaching conversation; it protects the causal interpretation of a future comparative study.
3. Missing-outcome sensitivity bounds are not finite-population causal-effect bounds
The function computes possible realized arm-rate contrasts under binary completions of unavailable outcomes. Its arithmetic is appropriate for that quantity. It does not bound the finite-population average causal effect merely because no outcomes are missing. With identical potential outcomes in both arms, an assignment can put all always-successful boundaries in one arm and all always-unsuccessful boundaries in the other: the arm contrast is one and the causal effect is zero.
This lane and the mathematical reviewer independently flagged the same ambiguity and directly agreed on the repair. The plan now labels the output as sensitivity bounds on the realized complete-data arm contrast, distinguishes randomization uncertainty and unobserved counterfactuals, and expressly denies causal-effect identification or confidence coverage. The mathematics/author lanes own the matching code documentation, result labels and executable counterexample; no function arithmetic was edited in this lane.
After the author lane added the causal-contrast distinction regression and regenerated its documented results, this lane independently reran all eight human-design tests successfully. The evaluation plan’s test count was updated.
Full synthesis and public framing
The synthesis now explicitly tabulates the incompatible assumptions of the component models. It does not infer authenticated authority from a fixture, delivered auditing from a belief, real sanctions from a utility variable, or human cooperation from scripted completion. The accessible blog links to the concrete human and agent paths and the adversarial record. Its disclosure still marks responsible human attribution and publication decisions pending.
The synthesis’s proposed comparison now refers to the same information and comparable declared support while measuring actual total labor, which is consistent with estimating a labor difference. Holding actual total labor fixed and then claiming to estimate savings would have been a different design.
A further editorial refinement may eventually qualify the blog’s opening phrase that affected uses become visible with “declared affected uses”; its immediately following calibration discussion already identifies hidden dependencies. I do not regard this as a current release blocker or introduce a broader guarantee.
Assent to the assembled collective assessment
I read reviews/COLLECTIVE.md as assembled after the
runtime, mathematics and domain passes. I assent to its
positive, bounded content recommendation. The additional
human-design review is now complete in this lane. Its three repairs
should be included in the final record and exact export. The content
does not need a completed efficacy trial or a functioning appeals
institution before public discussion. Those requirements would confuse a
research proposal with an operating service.
The final archive still needs its own frozen-byte verification and clean reproduction, and the actual responsible publisher must settle attribution, rights and destination. This addendum does not certify a changing intermediate archive or authorize contact with participants or venues. The best first public ask remains a small synthetic exercise and one concrete objection. A later human study can test whether that invitation and record actually help.