Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Controller delivery

Provider deliveries are repeated, workers overlap, and a process can stop at any point around publication. The controller therefore needs a small durable record that answers two questions: who may evaluate this delivery now, and which exact result must a retry publish?

In code, DeliveryLedger is the coordination interface for that record. It is a behavior contract rather than a required storage format. FileLedger is its first durable implementation; it uses ordinary files, not SQL or a database. This record is separate from the scan ledger, which records project research, and from the repository-owned review memory rejected in Provenance. The scanner itself remains offline and stateless.

Before the record

The contract takes each route from controller-owned configuration, never from the request body. It fixes the provider instance, the accepted trust-anchor set, and its signed-time rule. Before an adapter sees a delivery, IngressPolicy caps the exact body, header count, and total header bytes and checks the controller-recorded receipt time against a short queue window. Only an IngressCheck that passed those checks can enter ProviderAdapter::authenticate.

Each verifier consumes the IngressCheck itself. Authentication returns the provider facts and a small proof: which configured anchor matched, which trust set it belonged to, an optional signed issue time, and the replay identity. That proof also binds the controller-selected route, receipt time, exact header sequence, and exact body. Its fields are private; an adapter can join decoded provider facts to a successful proof, but cannot relabel its trust set, remove its signed time, or move it to another request. The controller checks that binding, applies the route’s signed-time rule, chooses the replay lifetime, and creates the delivery key. None of those steps trusts a decoded body field before the signature succeeds.

The replay key depends on what the provider actually signs:

Provider inputWhat is authenticatedReplay keyTime rule
GitHub X-Hub-Signature-256HMAC-SHA256 over the exact bodyDomain-separated digest of the exact bodyReplay-only; there is no signed delivery-attempt timestamp header.
Gitea-family X-Gitea-Signature or X-Forgejo-SignatureHMAC-SHA256 over the exact bodyDomain-separated digest of the exact bodyReplay-only; there is no signed delivery-attempt timestamp header.
GitLab Standard WebhooksHMAC-SHA256 over webhook-id.webhook-timestamp.bodySigned webhook-idSignedTimePolicy::Required(max_age); replay-only is not a valid GitLab route.
GitLab policy-job OIDCRS256 token with exact issuer, audience, policy origin, project, job, pipeline, runner, train commit, issue time, and jtiDomain-separated runner ID and jti digestRequired signed age plus token nbf and exp; replay-only is invalid.

This follows the providers’ published contracts: GitHub signs the payload body, Gitea signs the raw body, and GitLab’s signing token follows Standard Webhooks. The matching library code is split into the GitHub, Gitea-family, and GitLab verifiers. GitLab’s legacy plaintext X-Gitlab-Token is deliberately unsupported. GitLabWebhook authenticates the timestamp but does not choose the route policy. Ingress rejects that proof under a replay-only route, so the signed timestamp cannot be silently discarded. A GitHub or Gitea delivery header is useful for logs, but using it as the durable key would let a captured signed body bypass replay protection by changing an unsigned header.

The supported GitLab lane instead uses GitLab’s OIDC ID-token contract through the GitLabOidc verifier. It pins reviewed RSA keys and binds the small merge-request hint only after the token claims authenticate it. The Standard Webhook verifier remains available as an independent library surface, not as this lane’s request path.

WebhookKeyring holds one through eight HMAC keys in zeroizing, redacted memory. Anchor IDs and secret bytes must be unique. The ring owns the trust-set ID carried into its proof, so an adapter does not relabel a successful match by hand. Each key has an inclusive start and exclusive end time selected against controller-owned receipt time. Overlapping windows permit rotation; removing an anchor revokes it. GitLab whsec_ tokens have a strict, secret-safe constructor. Exact-body replay IDs do not include the matching anchor, so rotating a key cannot turn the same signed delivery into new work.

The controller also fixes one ReplayWindow: the largest signed age any route may accept and the largest ingress queue age. A route may require a shorter signed age but cannot exceed that fixed ceiling. An authenticated message ID and signed issue time receive an inclusive replay end computed from the issue time plus both fixed ceilings. Exact-body and other replay-only requests are marked permanent because they carry no authenticated time from which safe deletion can be derived. This choice reaches the ledger as part of AcceptedDelivery rather than being derived from payload fields, and the file record rejects a bounded delivery from a different replay window.

These verifiers establish webhook origin and integrity, not current authorization or an exact repository snapshot. The GitHub and Gitea-family lanes add signed pull-request decoding, controller-owned refresh, merge-rule checks, acquisition, and provider publication. The GitLab lane uses a separately verified OIDC token from its protected policy job instead of the Standard Webhook verifier. Its authenticated claims enter the same delivery and replay contract. Deployment and provider-specific trust rules live in Provider-verified controls.

The full flow

The controller owns the provider route. The selected adapter authenticates the untouched headers and body before any body field is trusted. Only the resulting authenticated delivery reaches the durable record.

controller_deliveryrawraw deliverygatebounds + receipt timeraw->gatefirstfirst refresh+ renewauthauthenticate+ replay identitygate->authclaimclaimauth->claimclaim->firstpublishpublishclaim->publishsaved retrycheck bindingrunrun if active+ heartbeat+ renewfirst->runstagesave exact resultexternalretain externalchain + renewrun->externalsecondfinal refresh+ renewexternal->secondsecond->stagestage->publishcompletemark donepublish->complete

The first refresh resolves the event-bound provider run, not the change’s latest head. It supplies the exact repository, URL dialect, refs, commits, and trees given to the runner, plus the provider gate revision to which publication is bound. GitHub and GitLab enforce their result on that commit. A Gitea-family lane publishes an exact-commit review, while the provider owns how that review affects merging. When the plan enables external verification, the controller retains its exact plan, provider evidence, and assessment after the run. The second refresh then checks the same identity, gate revision, and current authorization before the result is staged. If the change was closed, revoked, or superseded, the controller may publish that fail-closed status; it never publishes an old pass or block as if it were still current. A provider adapter may complete a stale publication without an external write only after independently proving that its staged provider gate is no longer current.

The logical record

This is the logical schema required by the contract. It is not a file format, wire schema, or prescription for how bytes are stored.

PartLogical valueRule
Delivery keyProvider namespace, provider instance, integration ID, delivery IDNames one authenticated provider delivery.
Fixed bindingRepository, change, provider run ID and attempt, object format, event candidate commitReusing the key with a different binding fails before refresh, run, or publication.
Replay lifetimePermanent, or an inclusive replay end based on authenticated timeDecided by trusted ingress and stored with the fixed binding. Only an ended bounded lifetime can permit deletion.
Evaluation IDOpaque controller-created ID with fresh random bytesCreated on the first claim and kept through retries and reclaims. A later row cannot reuse it.
Temporary ownershipEvaluation ID, lease deadline, fenceGrants permission to evaluate; the record, not a worker’s clock, decides whether it is still live.
Saved resultEvaluation ID, check-plan binding, fence, provider run, full run identity, provider gate commit, conclusion, optional report and retained-artifact referenceFrozen as one exact value before provider I/O.
StateNew, running, result saved, doneEach change happens atomically: fully or not at all.

Here, “delivery ID” means the replay identity accepted by ingress. It is the signed message ID for a GitLab Standard Webhook, a domain-separated runner-and-jti identity for the supported GitLab OIDC job, or the controller’s digest of the exact signed body for GitHub and Gitea-family requests. It is never an unsigned convenience header.

How FileLedger lays this record onto disk, its fixed lock set, checksummed frames, and cleanup rules, is on The file ledger. The separate bounded store behind a saved artifact reference is described in Retained provider artifacts; it keeps evidence bytes but never decides this record’s state.

The public Rust boundary has four operations. This abridged excerpt omits documentation and type bounds; the ledger module is authoritative.

#![allow(unused)]
fn main() {
pub trait DeliveryLedger {
    type Error;

    fn claim(&mut self, delivery: &AcceptedDelivery)
        -> Result<DeliveryClaim, Self::Error>;
    fn renew(&mut self, delivery: &AcceptedDelivery, lease: &DeliveryLease)
        -> Result<LeaseRenewal, Self::Error>;
    fn stage(
        &mut self,
        delivery: &AcceptedDelivery,
        lease: &DeliveryLease,
        publication: &Publication,
    ) -> Result<StageOutcome, Self::Error>;
    fn complete(&mut self, delivery: &AcceptedDelivery, staged: &StagedPublication)
        -> Result<LeaseCompletion, Self::Error>;
}
}

claim is the one entry point for new work and retries:

ResultPlain meaningController action
ExecuteThis caller has a live lease.Refresh and evaluate.
PublishAn exact result was already saved.Check its binding, then publish it without refreshing or running again.
BusyAnother live claim currently owns the work.Return the evaluation ID and retry time; do no provider or runner work.
DuplicateThe saved result was published and marked done.Do nothing.
BindingConflictThe same delivery key was reused for different authenticated work.Reject it before any provider refresh, run, or publication.

FileLedger can also reject a new identity with Full, reject an already ended bounded delivery with Expired, or reject a root whose saved lease duration, cap, or replay window differs from the configured one. These are fail-closed admission results, not reasons to evict live, saved, or permanent replay rows.

Four states

The record has four logical states. A lease is temporary permission to run. Its fence is an always-increasing generation number: reclaiming expired work keeps the first evaluation ID but uses a higher fence.

delivery_statesfreshnewrunningrunninglease + fencefresh->runningclaim: Executerunning->runningrenew: same fenceexpired reclaim: higher fencesavedresult savedexact valuerunning->savedstagesaved->savedclaim: Publishpublish error: retrydonedonesaved->donepublish OKthen completedone->doneclaim: Duplicate

A claim against live work may return Busy without changing state. A different authenticated binding returns BindingConflict. A stale renewal or stage returns Lost; the controller does not turn uncertainty into ownership.

The stored deadline is a scheduling hint, not proof. Renewal must preserve the evaluation ID and fence and must not move that deadline backward. After each renewal, the controller subtracts its own current time and returns only the positive time left. The concrete runner renews before launch and halfway through each returned window, capped at five seconds between checks. A zero window or a lost, malformed, or uncheckable renewal returns Stop; the runner then cancels its ProcessKit tree and discards the output. The heartbeat boundary is cooperative, while this runner turns its refusal into process cancellation.

Provider refresh calls have no heartbeat. A concrete adapter must therefore give each refresh a timeout comfortably shorter than the lease window. The controller also renews after the runner returns; the final atomic stage remains the decisive stale-owner check.

Races and retries

Suppose one worker holds fence 7 and another tries to reclaim the expired work:

  • If reclaim wins, the record moves to fence 8. The first worker can no longer renew or save a result, so it makes no publication call.
  • If saving wins, the exact result is frozen under fence 7. Reclaim no longer grants an execution lease; every claim receives Publish until that value is marked done.

Saving happens before external provider I/O because the record and provider cannot share one transaction. Publication may therefore be attempted more than once after an error or ambiguous acknowledgement. The adapter must make publishing the same result again have the same effect, using the authenticated delivery and evaluation ID as its repeat-safe key. A different result under that key must fail closed.

That is the controller contract. GitHub’s create-Check-Run API does not offer an atomic transaction or caller idempotency key: an accepted create with a lost reply can be retried before the first run is visible and leave a duplicate. The concrete adapter reconciles one exact visible run and rejects visible duplicates. The provider-verified-controls phase record states that limit rather than implying exactly-once.

Stop pointWhat the next claim seesSafe next action
During a live runBusy, or Execute after expiryWait, or evaluate again with the same evaluation ID and a higher fence.
After saving, before publicationPublishPublish the exact saved value.
Provider accepted, but its reply was lostPublishRepeat the same provider update.
After publication, while completion is unclearPublish or DuplicateRepeat the same update if needed, then complete the exact saved value.
After completionDuplicateDo nothing.

complete accepts only the exact saved value and is repeatable for that value while its done marker exists. A completion error after the provider accepted an update is kept distinct from an error before publication. On retry, the record must expose either the saved value or the done state, never a new execution lease. Once cleanup safely removes an ended bounded marker, later completion is honestly Lost rather than guessed from missing evidence.

Failure behavior

ConditionRequired behavior
Raw ceilings, receipt time, signature, anchor window, route, trust set, or required signed request time failsReject the delivery before claiming it.
The key has a different authenticated bindingReject it before provider or runner work.
A saved result does not match the authenticated deliveryReject it before provider I/O.
Another claim is liveReport in progress and do no work.
Lease renewal cannot prove ownershipStop, discard runner output, and do not publish it.
A valid refresh for the same delivery reports closure or revocationSave and publish the matching unavailable result, never the old pass or block.
A valid refresh for the same delivery reports supersession or changes its URL dialect, refs, base commit, trees, or provider gate commitSave Superseded and invoke publication, which may prove the staged gate stale and write nothing; never publish the old pass or block.
A refresh returns another repository, change, object format, or event candidateReject it without saving or publishing a result.
Runner output is missing, timed out, too large, tampered with, or bound to the wrong identity or treeSave and publish the matching unavailable result without a report.
Atomic stage loses the fence raceMake no publication call.
Publication fails or its acknowledgement is unclearKeep the exact saved result for another publication attempt.
Completion cannot be confirmedReport a completion error; retry only the exact saved result.

The trusted runner promises that a completed engine result already passed its engine, exit-class, and request checks. The controller independently checks the returned identity, nonempty output, and size. It does not authenticate the engine report itself.

Supervised bootstrap run

run_bootstrap implements the provider-neutral execution step once exact repository and action trees have been acquired. It first reopens both repositories and verifies the requested commits and trees. It derives the sealed bootstrap job from the RunRequest and trusted instants rather than accepting a separately assembled job, so a caller cannot pair one run with another run’s control files. Trusted acquisition may return candidate-independent semantic-template bytes with the roots. Each result names one plan-frozen acquisition identity; job construction matches its producer and context, binds the exact candidate itself, combines its limits and ordering with plan templates, and fails closed on any defect. The job also carries one bounded canonical audit value with the exact acquired bytes, derived envelope bytes, acquisition identities, and digests for the artifact-retention layer. It also reads the selected bootstrap under a fixed byte ceiling and matches its digest to the frozen execution plan before copying it into a fresh private scratch directory.

The directory holds only the copied bootstrap, canonical request files, report, and final result record. The controller creates both output files and retains their open handles. The bootstrap may write through the fixed path names, but the controller never reopens those names; replacing a name therefore cannot replace the object that will be read. The child starts with a cleared environment and closed stdin, stdout, and stderr. The report is bounded by MACHINE_JSON_BYTES; the small result record is written last, so a missing record cannot be confused with a completed report. Exit status and that record are checked together. Missing output, oversized reports, timeout, and runtime tampering retain distinct fail-closed outcomes. Signals, heartbeat loss, and spawn failure all fail closed as Unavailable.

Supervision uses pinned ProcessKit 2.2.5 with Tokio through one cross-platform path. ProcessKit selects the host’s process-tree boundary. The controller enforces a positive wall limit no greater than 120 seconds and renews the ledger before launch and halfway through each newly proven lease window, capped at five seconds between checks. After every terminal path it hard-kills the group and waits up to two seconds for ProcessKit to report it empty before reading either output. Failure to prove that drain is unavailable, not completion. This also covers a clean leader exit, so a surviving descendant cannot escape merely by letting its parent finish. A heartbeat refusal cancels the same tree and discards its output. These bounds cover ordinary supervised execution; they do not promise recovery from a host kernel call that itself never returns.

Focused tests cover wrong commits and trees, a wrong bootstrap digest, the cleared child environment, timeout and heartbeat races, descendants that remain alive after their leader exits, path replacement, and missing, malformed, and oversized output. Every provider service calls this runner only after its adapter has refreshed provider state and acquired the exact roots.

What exists now

The controller crates contain the provider-neutral identities, bounded ingress gate, rotating key ring, signature verifiers, durable raw inbox, DeliveryLedger, FileLedger, bounded artifact store and retrieval route, worker, orchestrator, acquisition boundary, bounded Intersphinx and mdBook semantic-evidence producers, and supervised bootstrap runner. Focused tests cover ingress limits and tampering, replay, rotation and revocation, file corruption, cross-process ownership, reclaim, exact publication retry across restart, full roots, artifact expiry, clock rollback, runner timeout, process descendants, and output replacement. The acquisition result can also carry candidate-independent pre-scan semantic templates into the sealed controls request; candidate binding remains inside job construction. A trusted caller can derive source-bound routes, decoded anchors, and rendered navigation reachability from a completed mdBook build without executing mdBook in the controller; the built-in provider acquisitions currently leave that set empty.

Three merge-gate shapes join those pieces. GitHub uses a signed pull-request event, App refresh, strict App-bound ruleset, authoritative test merge, and App-owned Check Run. GitLab uses policy job OIDC, an enforced merge train, independently owned pipeline execution policy, and the policy job’s own result. Gitea and Forgejo use a signed pull-request event, effective protected-branch rule, dedicated reviewer identity, and that account’s approval or request for changes. All three acquire exact SHA-1 repository and action objects through the same fixed-budget protocol-v2 path, run the bootstrap, and refresh the provider gate again.

Provider-verified controls compares the lanes and links each deployment reference. The GitLab Standard Webhook verifier remains an independent library surface; the supported GitLab lane authenticates OIDC instead. The engine’s forge field still chooses only a URL dialect and never authenticates a provider.

Last change: , commit: 6cafb111