Executable boundary model: states, scarce work, and repair
Research question: Can four participant actions expose composition failures without a universal reputation score or a large mandatory rule table?
Bounded answer: A small deterministic model can reject several precisely specified mismatches and exhibit counterexamples to broader promises. It cannot establish human usability, scientific truth, actual independence or live authority.
Contribution and claim ledger
| Label | Claim | Evidence / limit |
|---|---|---|
| D | Offer/check/rely/amend forms a small participant API | toy_agents/protocol.py; proposed semantics, not adopted
standard |
| T | Joint reservation enforces specified person and skill capacity | Finite-sum argument below; assumes complete correct capacity keys |
| I | Atomic panels, version mismatches, scope gaps, unknown checks, unauthorized decisions and stale renewal are handled | toy_agents/tests/test_protocol.py; trusted-process
synthetic cases |
| T | Canonical group-first sampling is invariant to redundant labels within a group | Elementary pushforward proof below; static eligibility only |
| I | Hidden dependencies escape amendment propagation | Deliberate negative control hidden_dependency() |
| H | Four actions lower total human labor or improve defect detection | Unestablished; compare structured-template baseline in a human pilot |
The API is a synthesis and executable illustration of the preceding dossier, not a claim that versioned records, graph reachability, resource reservations, randomized panels or state machines are newly invented. Closest technical basis is the prior bounded-reliance work summarized in the self-contained background. No external interoperability, external peer review or empirical evaluation is claimed.
Target and authority composition
Define exact target with claim ID, version and evidence identifier. The simulator treats as an opaque fixture identifier; a cryptographic binding would require a separately authenticated content layer. An offer gives scope and declared prerequisite targets . A check gives scope , reviewer , method, limitations, outcome and dependency-generation snapshot.
For requested scope , selected supporting checks , and minimum independent groups , the implemented coverage condition is
This is stronger than counting distinct groups across the entire panel: one group checking calibration and another checking statistics do not provide two checks of either. All selected checks must refer to the same offer, support the claimed scope, and match the current declared dependency snapshot. Current overlapping contradictory checks also block a new reliance even if omitted from . This is a conservative model choice, not a replacement for reasoned disagreement or domain adjudication.
Typed authority is a predicate for decision kind . The trusted fixture registry may give scientific or publication authority separately. The engine never infers from a support check, team size or a favorable numerical result. It does not validate whether any real institution granted the authority.
Reservation invariant and exact units
Let be effort reserved or spent by accountable capacity key , skill and task . A valid schedule satisfies
For a candidate batch , the ledger adds its nonnegative integer entries to the full committed ledger, validates both inequalities, and commits the entire batch only if every inequality holds. Therefore, by induction from the empty ledger, every accepted sequence satisfies both bounds. A rejected batch changes nothing. Consumption preserves its allocation; cancellation removes only unspent work. This is a finite accounting invariant, not proof of deadlines, productivity or honest reservations. Independent tests compare randomized batch acceptance against a separate schedule-summing oracle, including shared skills and partially fitting batches. A single person represented by two unknown keys breaks the premise.
check_panel uses a detached serial state to expose
all-or-none behavior. It is not scalable storage and supplies no
concurrency, crash-recovery or distributed atomicity guarantee. Units
are synthetic effort tokens per scope skill. A real application must
price setup, supervision, audit, triage, appeals and repair in measured
compatible units and preserve commitments across epochs.
Canonical panels and the label-multiplicity trap
Let be the static qualified, nonconflicted group set. Sample uniformly from its -subsets, then choose an eligible representative inside each chosen group. For group panel , the marginal probability is
Splitting a representative into redundant labels inside a fixed group changes only , whose sum remains one. Thus the pushforward onto group panels is unchanged. This is conditional on correct fixed group labels and unchanged eligibility, not Sybil resistance or a dynamic completion theorem. The executable sampler enumerates subsets, so its cost grows combinatorially; small toy pools only. No efficient general constrained sampler is claimed.
Negative control: for groups A, B, C and two seats, give A representatives, B and C one each. Uniform sampling of feasible representative pairs has panels; A occupies . Its probability is despite a strict one-seat per-group rule. Group-first sampling instead gives . The checked-in run uses and 3,000 draws; exact formulas are the claim, frequencies illustrations.
Completion is a different conditioning event. If risky invitations have fraction and complete with probability , versus for others, then
The demo uses , giving . It records all stage denominators and charged intake. No offered-panel probability bound is carried across capacity filtering, retries or selective completion.
Amendments and bounded currentness
Each offered target has generation . A reliance stores generations of its transitive declared prerequisite closure, including itself. Material amendment increments the changed target generation. A downstream reliance is pending iff its snapshot differs from the current closure snapshot. A new checked and authorized reliance can capture the new snapshot without modifying the old one. Offer ordering requires prerequisites already exist, so the toy dependency graph is a finite DAG; graph-cycle negotiation and undeclared future dependencies are outside its scope.
The currentness view also checks expiry, evidence availability and a
stated maximum observation age. This is not a statement about scientific
correctness. An undisclosed dependency is absent from the closure, so
its amendment cannot change that snapshot. The negative fixture
intentionally produces current_under_toy_policy when an
omniscient observer knows the claim needs reconsideration. The failure
is retained in release evidence.
Evaluation and stopping rule
Run the documented tests and deterministic scenario command before packaging. Do not suppress a failing negative control to improve the number of green tests. Stop any general claim at its first violated premise: missing group identity, unknown capacity, incomplete evidence graph, absent authority or stale observation. If a structured record plus ordinary review performs as well with lower total human labor, retain that simpler interface. No human pilot, calibrated error model or consequential external action occurred in these experiments.
AI assistance: this framework was drafted and checked by AI agents within the same orchestration. The blue/red record identifies this as internal adversarial review; human author approval and external review are separate release decisions.