Fixed
- Record
tree_fingerprint_headbeside everytree_fingerprintin verify and run receipts; the post-commit linkage statement uses it as areceipt-headnormalization base, provingLINKED-NORMALIZEDeven when the linked commit’s first parent is not the run baseline. (#1454) - Post-commit linkage now always records a
baselinepredicate withbaselineRelationset tounknown(nevernull) whenrun.jsonhas nobaseline_commit, andverify-commit-linkageconfirms the comparison asconfirmedwhen the statement and recomputed relation both sayunknown. (#1454) - Control crosswalk EC-12 rows now reflect the landed evidence package: relationships are
supportsorpartially-supports, the state rule isevidence-package-manifest, andbrigade evidence controlschecks.brigade/evidence-packages/*/manifest.jsonfor schemabrigade.evidence_package.v1and recomputedentries_sha256. (#1454) - Control crosswalk review fixes on
docs/1406-control-crosswalk: corrected ISO 42001 clause/Annex control identifiers, ISO 27001 A.5.9 and A.8.12 mappings, SSDF practice IDs, SP 800-53 AU-11 and SOC 2 CC8.1/CC7.2 mappings, EU AI Act applicability/bearer metadata, removed Art. 50 and added an explicit Art. 9 no-relationship row, removed unsourced NIST AI 600-1 and CSA AICM mappings withidentifiers-not-sourcedstatus, and corrected framework publication dates/editions. Strengthenedbrigade evidence controlsstate rules for EC-02 through EC-11 to use existing verifiers and journal readers, and no-relationship rows now rendernot_applicable. (#1406) - Test Result attestation export now rederives verify-receipt SHA-256 values
from the writer’s compact sorted-key JSON contract, including receipt
path, and refuses stale or malformed stored digest evidence. Legacy digestless exports remain supported. Export receipt lookup is strict and bounded. Approval collector integration remains deferred. (#1404) - Approval v1 and v2 collectors now require a stored receipt digest on every
matching verify receipt, validate the digest before any tree fingerprint
filter, and re-derive the Test Result attestation from the same snapshot so
receipt.jsonis read exactly once. (#1404) - Agent-change evidence index hardening: run ids are validated against
^[A-Za-z0-9._-]+$, run directories and reference locators are resolved and contained under<target>, symlinked components are refused, Test Result references passrequire_receipt=Truewith the snapshot receipt, therequestpredicate field is resolved from the journal’srequest.signedevent,missingentries are only emitted for kinds required by the policy, and the verifier checks allowed profiles, no-approval cycles, and bounded verify-run scans. (#1404) - Attestation, cosign bundle, and approval verification now reject oversized, duplicate-name, non-finite, malformed-Unicode, over-nested, and cyclic JSON inputs before signature processing. DSSE accepts both standard and URL-safe base64 alphabets under bounded decoded payload and signature sizes, and file reads use no-follow regular-file descriptors. Mapping inputs are copied into bounded plain JSON snapshots before use, and dot-only fallback attestation run IDs are discarded.
brigade center report buildrotates old operator report directories automatically under.brigade/center/reports/intoreports-archive/according to retention (default keep newest 20, configurable viaoperator_report_retentionin.brigade/daily.tomlalongsideallow_operator_report_buildor CLI flag--keep N), capsreports-archive/to the same retention count by deleting the oldest archived directories, never rotates unclosed reports, and supports--dry-runto preview rotation without deleting or moving bundles. Fixes #1415.- Fixed timing flakes in tests
test_release_with_renew_in_flight_never_resurrectsandtest_concurrent_edge_writes_do_not_lose_dependency_edges. (#1396) - Antigravity seats now pass
--print-timeoutderived from the seat timeout toagywhen supported byagy --help, preventing worker turns longer than five minutes from timing out prematurely. (#1431) - Grok Bot hub jobs now count
timeout_secondsas an execution budget from claim time rather than enqueue time, avoiding premature expiry of jobs that waited in queue, and add an independentqueue_ttl_secondsto expire unclaimed jobs. (#1409) brigade run cloud grokbot scout-feedunder Fleet Hub authority now re-selects issues whose previous hub jobs expired, failed, or were canceled, using a fresh idempotency revision instead of permanently marking them as already known. (#1364)- Fleet Hub grants feed and control actors queue-scoped
listauthorization soscout-feed --applycan check queue limits and existing jobs under hub authority, and adds afeed-authoritycheck todoctor. (#1350) - Grok agent dispatch now recognizes snake_case
end_turnstop reasons in read-only mode and passes--output-format jsonin write mode, enablingcli = "grok"seats to complete turns on grok 1.0.13. (#1349) - Direct read-only Grok dispatches now bind session continuation to a launcher UUID and reject external session controls while bound, failing unbound output as an output contract violation. (#1333)
- Claude hook timeout handling retries locked state updates and records pending increments in a sidecar when the timeout journal is full, preventing dropped timeout latches under high concurrency. (#1304)
brigade runon Windows no longer fails withPermissionErrorwhen enforcing journal file modes on read-only descriptors and handles closed pipes cleanly during process termination. (#1098)- Seat process timeouts now properly terminate child processes and report timeout status instead of failing with an
AttributeErroron_SeatProcessRegistry. (#936) brigade skills sync --writenow routes mutations through the shared projection transaction kernel, executing bundle updates, receipts, and install history as an all-or-restored operation. (#924)brigade releasesecurity summaries now exposeprojection_truncated_countwhen nested security checks or top findings undergo truncation. (#678)brigade releasecandidate auditing redacts unsafe Unreleased changelog entries from release candidate artifacts and reports a counted redaction summary. (#674)brigade outcome capturediscovers skills in linked git worktrees by checking the target worktree before falling back to the parent checkout, keeping skill discovery and card fingerprinting repository-bounded. (#874)
Added
- Agent-change evidence index emitter (
brigade receipts export agent-change), verifier (brigade receipts verify-agent-change), and policy initializer (brigade receipts agent-change-policy init), with external policy schema, in-toto predicate, and machine-readable verification output. (#1404) - Post-commit linkage statement (
brigade receipts export commit-linkage) and verifier (brigade receipts verify-commit-linkage) that bind a run’s attested tree to a specific git commit after the commit exists, with exact/normalized tree equivalence, isolated git environment, and machine-readable verification axes. (#1404) brigade evidence controlsemits a versioned, machine-readable control crosswalk (brigade.control_crosswalk.v1/brigade.evidence_controls.v1) mapping Brigade evidence claims to ISO/IEC 42001:2023, NIST AI RMF 1.0, NIST AI 600-1, NIST SP 800-218/SSDF, NIST SP 800-53 Rev. 5.1, EU AI Act, ISO/IEC 27001:2022, AICPA SOC 2, OWASP Agentic Top 10, and CSA AICM identifiers. The normal command and--jsonoutput compute an evidentiary state per row (evidenced_passed,evidenced_failed,untested,not_applicable) from configured workspace artifacts.--render-docregeneratesdocs/control-crosswalk.mdas a static mapping (no state column, no scope, noevaluated_atline) that is a pure function of the bundled template; evidence states are computed at query time and are not part of the document. The header states that this is an evidence index and design documentation, not evidence that a control operated, and the template does not reproduce licensed ISO or AICPA clause text. Fixes #1406.brigade governance inventorynow exports one canonical workspace-scoped inventory with bounded observed worker use, node-local Fleet policy facts, and optional CycloneDX machine-learning-model components. Remote MCP endpoints retain only scheme and authority; user-home rosters are never substituted. Fixes #1408.- Fleet hub roster page at
/deck/roster: roles (impl, review, chef, research, security, scout), seat and cloud lane toggles, consumer defaults, and notes saved in one revisioned transaction;brigade work briefprints thefleet_routingblock;fleet preference setgains--research,--security,--scout; hub schema v19. brigade receipts trailer --run <run-id>printsBrigade-RunandBrigade-Receipttrailer lines for a local run receipt so a conductor can pass them togit commit --trailer.brigade receipts verify --commit <sha>reads the trailers from that commit’s message, resolves the run receipt, and recomputes the digest to check for mismatches. (#1407)brigade receipts export packageexports one verify run as a portable, staged evidence package with a detachedbrigade.evidence_package.v1manifest, andbrigade receipts verify-packageindependently verifies content integrity of the copied files. The package is integrity-only and does not include command logs, graph databases, or other run files. (#1407)brigade run cloudis a native parser branch;brigade run-cloudstill works and prints a one-line deprecation notice.brigade harness fragments --harness {openclaw,hermes}fronts the shared fragment backend, withopenclaw fragmentsandhermes fragmentskept as aliases.tests/test_cli_inventory_contract.pyasserts the parser,docs/command-inventory.md, and the roadmap missing-command check agree after alias normalization, and reports (without failing) parser paths that have no direct test invocation (#1229).- Obsidian Operator Phase 2 atomic structured-file adapter: checked-in
obsidian-plugin/grokbot-operator-adapterplus operator-onlyops/install-grokbot-operator-adapter.sh. Public tools, bind, and proposal lifecycle stay unchanged. The plugin now ships a build-time CommonJS bundle of Zod 3.25.76 (pinned URL plus npm integrity, MIT license and NOTICE in-tree) soaddMcpToolcan register the private CAS tools against the published Local REST API v2 host contract, which wraps aRecord<string, z.ZodTypeAny>shape withz.object(shape).scripts/bundle-zod.js --checkreconstructsmain.jsandSHA256SUMSfrom the checked-in vendor CJS plussrc/adapter.jswith no network and no writes, and exits nonzero on drift. The three tools use the same exact shape{path, expected_sha256, replacement_utf8};parseInputstill rejects extra keys, malformed hashes, and oversize replacements. Brigade’s Python runtime does not depend on Zod, and the vault install still copies onlymain.jsandmanifest.json.obsidian_capabilitiesstays Phase 1 until a matching live fingerprint (versioned tool descriptions plus input keys) is observed; this slice does not enable, reload, or cut over live services. Outbound Streamable JSON-RPC now covers the maximum allowed replacement plus a bounded envelope, and an unsendable replacement failsinvalid_requestat propose time before a single-use approval can be claimed.update_excalidrawbinds the embed note path and revision separately from the scene. A Local REST host object change unregisters the old handle and reacquires once. Refs #1255. - First-party Grok Bot Obsidian Operator Phase 1 connector pack
obsidian-operatoron127.0.0.1:8773with public route/mcpand the closed six-tool inventoryobsidian_capabilities,obsidian_search,obsidian_read,obsidian_action_status,obsidian_propose_action, andobsidian_execute_action. Public capability output stays generic. Preview-first pack lifecycle writes only local config and an explicit unit file; canary is non-mutating, calls authenticatedobsidian_capabilitiesthrough the public listener, and this pack does not cut over live services. Setup stores a validated loopback HTTPS upstream URL and an upstream-key env/file reference, never a raw key. The installed listener builds a fixed Streamable-HTTP native MCP client, verifies CA plus hostname/IP plus SPKI pin before sending Authorization, and maps Local REST search without an unsupported limit, snippets from bounded matches, targeted read/patch, and copy/movepath/destinationwithallowOverwrite: false. Native search, read, and existing mutations retrieve vault tags and deny configured sensitive tags before exposure or execute. The validated Excalidraw helper is launched from a no-follow executable descriptor via/proc/self/fd/<fd>(absolute helper, empty argv, staging cwd, allowlisted environment, 45s deadline, 256 KiB output cap, initialize plus four tools). The action store bounds active proposals and prunes expired proposal families. Private runtime files are read through a no-follow descriptor that requires the current UID and mode 0600. - First-party Grok Bot Backup Steward connector pack
backup-stewardon127.0.0.1:8772with the closed six-tool inventorybackup_overview,backup_target_status,backup_restore_readiness,backup_operation_status,backup_propose_action, andbackup_execute_action. Preview-first pack lifecycle writes only local config and an explicit unit file; canary is non-mutating and this pack does not cut over live services. brigade run cloud launchstarts Cursor Cloud or Jules work from a bounded private prompt file. Provider keys stay in the existing environment resolution, launch JSON never includes the prompt or the private lease holder, and a missing key or bad prompt file makes no provider call. Successful bind registers provider IDs, the prompt hash, the expected artifact, and the holdersyncneeds.- Grok Bot
feed,scout-feed, andbuild-feedcommands can now wake a webhook routine after enqueuing work with--apply, sending a bounded notification payload using sender credentials from a local private configuration file. (#1391) wake.jsonmay declare per-role webhook targets underwebhooks, so a scout enqueue wakes the Scout routine and a build enqueue wakes the Builder instead of every enqueue hitting the default webhook. Secrets never appear in logs. (#1449)- Added an opt-in Worklore fleet backlog ledger to Fleet Hub with
brigade fleet workcommands for managing shared backlog items, tracking lifecycle events, and synchronizing issues and task ledgers. (#1361) brigade care installsupports repeatable--scheduleoverrides per entry, preserves custom schedules across reinstalls, parses structured systemd calendar records, and warns when evidence-ledger writers share a timer calendar. (#1326)- Added the bundled
operations-relayGrok Bot connector pack with loopback bearer isolation to report operational incidents and propose remediation actions for operator review. (#1330) - Added approval-gated fleet remediation allowing Fleet Steward to execute catalogued service recovery actions bound to active Wazuh findings under single-use operator approvals and fenced Fleet claims. (#1310)
- Added the first-party
wazuh-triageconnector pack with six closed tools to ingest, normalize, and classify Wazuh security alerts into sanitized incident bundles and non-executable remediation proposals. (#1306) - Added the first-party
cerebroGrok Bot connector pack providing five tools for Cerebro memory integration with preview-first setup, loopback bearer security, and read-only canary verification. (#1263) - Added
brigade run cloud grokbot scout-feedto enqueue Repository Scout jobs for approved GitHub issues under queue locking, daily caps, and issue idempotency. (#1224) - Added
brigade run cloud grokbot feedto validate private manifests and safely enqueue approved Grok Bot tasks under the queue lock with--apply. (#1176) - Grok Bot queue jobs are now projected into Brigade’s cloud tracker with
grokbot-cloudstatus across Center, work brief, daily status, and operator health. (#1148) - Added a bounded local Grok Bot job queue (
brigade run cloud grokbot status|claim|cancel|expire) with leased claims, safe status projections, and artifact references. (#1143) - Center Agent Activity displays inline provider brand marks, flags lockless or unresponsive runs as stale, defaults machine views to the last 12 hours, and polls every 15 seconds. (#919)
- Added
brigade code export --json(brigade.code-graph-export.v1) and the Center Code Graph view (/view/code) featuring module dependency maps, symbol impact drill-down, and branch change overlays. (#887) - Added
brigade memory topology --jsonandbrigade memory inventory --jsoncontracts and replaced Center’s Cards search with the Memory Operations topology and inventory dashboard. (#875) brigade updatepassesGITHUB_TOKENorGH_TOKENviaAuthorization: Bearerfor GitHub API requests to respect rate limits, retrying unauthenticated on credential errors without leaking token values. (#751)- Added
brigade evidence memory auditto report canonical memory card identity coverage, collisions, legacy aliases, and proposed mappings, surfacing body-free projection health states for dashboard consumers. (#844) - The evidence ledger now enforces a single default-live version per external identifier by tombstoning superseded records, filtering them from search queries, and concealing unreviewed injection-pending bodies in
show --json. (#843) brigade contexttracks context-pack freshness against underlying files and dependent receipts, rejecting control characters in persisted snapshot paths so doctor reports generic unsafe-path findings. (#492)
Deprecated
brigade run-cloudremains a working alias forbrigade run cloudfor at least one stable release and prints its existing one-line replacement notice. The generated inventory marks its command paths deprecated. Usage evidence:brigade evidence search "run-cloud"found historical receipt148c30c8534922236df439cb. (#1230)
Fixed
brigade work verify run --command <script>timeout cancellation now terminates all descendants of the verify child, including processes that moved to a new process group (e.g.timeout <secs> <cmd>without--foreground). Descendants are collected via/proc/<pid>/statbefore and after the group kill and terminated with the same shared grace windows. Fixes #1413.- Grok Bot hub jobs now expire on their own
timeout_secondsinstead of waiting for an operator to callexpireby hand. A job’s deadline isqueued_atplustimeout_seconds, and the hub enforces it in three places:statusandreport-metadatasweep before they answer (listand the mutating actions already did), a claim of a job past its deadline is refused with the bounded reasonjob-expiredrather than therevision-conflictthe sweep’s revision bump used to produce, andfleet_hub_grokbot.start_expiry_sweeperruns the sweep on a 60-second timer for the life ofbrigade fleet hub, so a queue nobody polls still expires (seven of the ten jobs in the 2026-08-31 diagnosis sat queued past their timeout, one for 13h21m against a declared 3600s). Each automatic expiry writes anexpirerow togrokbot_operationsunder anexpire:deadline:<job_id>:<revision>operation id with a NULLactor_node_id, which cannot collide with an operator’s own id. The local (no-hub) queue follows the same rule:grokbot_jobs.statusandget_jobterminalize an elapsed job before projecting it (both take an optionalnow), and a claim past the deadline expires the job and failsjob-expired. Timestamp helpers moved togrokbot_job_clock.pyto keepgrokbot_jobs.pyunder the module-size ceiling. Deadline expiry composes with the lease reason added in #1386 rather than pre-empting it: a read keys on the job’s own deadline only, so a claimed job whose lease alone lapsed staysclaimedand its holder is still refused withlease-expired; a lease-holder call on a job past its own deadline expires the row, records the operation, and still answerslease-expired, since a granted lease is always clamped to the deadline.job-expiredstays the claim-side answer. Fixes #1353. brigade run cloud grokbot reconcile-reportslists completed Repository Scout jobs from the hub when the queue is hub-authoritative, pairs each with the local artifact and private snapshot bytask_hash, and drafts unmarked reports. Leftover pre-cutoverjobs/rows are not scanned.unavailablenow distinguishes a hub job with no local artifact (artifact-missing) from a local job with no stored report (report-missing). Fixes #1367.brigade runno longer loses a finished worker’s result when a long session crosses the 1 MiB capture cap. The cap was charged against every byte read off the transport, including reasoning deltas, tool-call notifications, and command output that Brigade routes and discards, so ordinary session chatter spent the budget, killed the app-server child beforeturn/completed, and left a completed worker’s edits stranded in the worktree. Reading is now bounded separately from retention:proc.MAX_STREAM_BYTEScaps the untrusted stream and is the only thing that terminates the child, whileproc.MAX_CAPTURE_BYTEScaps retained text and is charged only for output Brigade keeps. Retained text keeps its head and tail so a truncated capture still carries the final answer, a turn that completed with truncated text staysok, andoutput-limitclassifies as an output-contract violation rather than a transport failure so truncation never poisons seat health. Worker receipts record the observed byte count and the configured cap. Fixes #1144.- Fleet hub
init_dbnow runs every schema-creating statement (including Grok Bot and run-preferenceensure_schema) inside one lock-guardedBEGIN IMMEDIATEwith the existing locked-database backoff, and the migration connection uses a 30s busy timeout. Concurrent first-touch init on a loaded CI shard no longer leaksOperationalError: database is lockedfrom DDL that used to sit outside the claims-only retry loop.fleet_hub_grokbot.ensure_schemauses per-statementexecuteso it cannotCOMMITa caller-held write transaction. Fixes #1272. - The Grok Bot MCP request gate now replays buffered ASGI receive messages and then calls the original receive callable, instead of inventing a disconnect after the buffered request.
- The isolation posture marker is now tri-state, identity-bound, and fail-closed on persistence (#881 final round). Marker reads previously collapsed every error into “marker absent”: a same-uid writer who replaced the user-level
authority-isolationdirectory with an unreadable stand-in (or corrupted any marker entry) after flipping the repo-writable isolation flag tooffand removing the signed store could admit the unsigned dedupe fallback, letting forged sidecars suppress a canonical import. Reads now return present / confirmed-absent / unknown - only a positively read healthy marker store naming no governing marker reports absence; unreadable roots or entries, malformed bytes, non-regular entries, and incomparable identity bindings report unknown, which the fallback treats like present and refuses. Observing genuineexternal-keyisolation no longer swallows posture-persistence failures: it binds the marker to the stable workspace root identity (device/inode, the same identity the signed authority record carries) and refuses to proceed when the identity cannot be read or the marker cannot be persisted;brigade security doctorreports that as FAIL. Posture markers were keyed by path fingerprint only, so renaming a governed workspace orphaned its marker and let an attacker flip the renamed workspace’s config off before its first post-rename observation: the authenticated reanchor now transfers and re-verifies the posture onto the new path (ahead of store adoption, before any unsigned evaluation), identity-scoped reads keep the moved workspace governed even without a transfer, and an unrelated workspace that recycles the old path is never blocked by the stale path-keyed marker. An import append also no longer reopens the authority state through the legacy helper when its own snapshot acquisition legitimately found nothing: append-path validators receive an explicit absent-snapshot sentinel, so one append consumes exactly one acquisition (spy-tested). New regression tests cover the unreadable marker root, persistence-failure observation, rename carry-over without pre-opening, stale-marker non-governance at the recycled path, and single-acquisition appends over intact and removed stores. Fixes #881. brigade runno longer loses a worker’s result or poisons the seat when the combined app-server output stream exceeds the 1 MiB capture cap. An over-cap codex app-server turn (failure_kind=output-limit) with salvageable final-answer text is now delivered bounded through the normal envelope gate as a successful worker result, with the truncation recorded viaoutput_truncated: truein worker-results.json and the original “combined output exceeded … byte limit” detail kept; an over-cap turn with no salvageable text stays a typed failure but no longer triggers same-seat retry or quarantine. Direct CLI overflow keeps its existing #1012 contract. Fixes #1144.- Fleet spool events are no longer lost silently on platforms without descriptor-relative file APIs (Windows; #1157 round 3 regression).
_ensure_private_dirdocumented returningNonethere but always returned a directory descriptor, sending_spool_lockand_write_spool_atomicinto thedir_fd=/src_dir_fd=/dst_dir_fd=calls CPython documents as unsupported on Windows; the resulting failure surfaced only asreport_event’s silentreturn False, so undelivered events vanished without a trace. Descriptor support is now detected once per call (sys.platformplusos.supports_dir_fdmembership ofos.open,os.rename- which registersos.replace’s dir_fd support,os.unlink, andos.stat): unsupported platforms take the documented path-based fallback (Nonereturned, a symlinked spool directory still refused, best-effort 0700 path chmod,tempfile.mkstemp+os.replacerewrite), and any spool write that fails anyway is logged at WARNING with the reason instead of being swallowed. Regression tests simulate the unsupported platform via monkeypatchedos.supports_dir_fd; POSIX descriptor-scoped behavior is unchanged. Fixes #1157. - Fleet abort hardening residuals (#1157 round 3).
brigade run’s credential-failure callback now appends its classification marker before printing the diagnostic and fires_thread.interrupt_main()from afinally, so a failed stderr write can no longer swallow the callback and let the run continue after the hub rejected this machine’s token (receipt kindfleet-credentials-rejectedeither way); the claim-lost callback records its marker before its fallible print, so a dead stderr yieldsfleet-claim-lost, not a user cancel. Fleet spool atomic rewrites are now descriptor-scoped on POSIX: the exclusive unpredictable temp file (os.open(O_CREAT|O_EXCL|O_NOFOLLOW, 0600)) and the finalos.replace(src, dst, src_dir_fd=…, dst_dir_fd=…)address names relative to the validated spool-directory descriptor instead of re-resolvingpath.parent, so a concurrent directory swap cannot redirect the rewrite (tempfile.mkstempremains only for the no-descriptor fallback). The claim-loss abort’s watchdog grace fire and the callback-finally fire share one lock-guarded check-and-set on the one-shot fired flag, so a callback ending exactly at the grace boundary interrupts exactly once. Fixes #1157. - Fleet claim-loss abort no longer depends on the
on_claim_lostcallback behaving (#1157 round 2). The interrupt now fires independently of callback execution: a callback that blocks forever is cut off after a bounded grace period (fleet_client.CLAIM_LOST_CALLBACK_GRACE_SECONDS) and a callback raisingSystemExit/KeyboardInterrupt(anyBaseException, not justException) is logged - the guarded block is interrupted either way, after the callback has had its chance to record state. A lost-claim abort entered from a non-main thread no longer calls_thread.interrupt_main()(which targets the process main thread and would stab an unrelated thread while the owner kept working): the yieldedClaimDecision.cancel_eventis set instead, and guarded work off the main thread must observe it to unwind (documented indocs/fleet-sync.md). The spool directory’s 0700 privacy is now enforced through a descriptor rather than check-then-chmod on the path: it is openedO_DIRECTORY|O_NOFOLLOW, forced viafchmod, re-verified withfstat, and lock/spool files are openeddir_fdagainst it; when privacy cannot be applied or verified the write is refused instead of warned about. Hub redirect same-origin checks normalize default ports, sohttp://hub→http://hub:80is followed while cross-origin hops stay refused.brigade runclassifies and records a fleet claim-loss abort on the main thread from the heartbeat’s marker, so a heartbeat-thread receipt-write failure can no longer record the abort as “canceled by user”; regression tests assertrun.json’sfailure.kind, not stderr. Fixes #1157. - The canonical-dedupe isolation selector is no longer attacker-writable (#881 follow-up).
_unsigned_dedupe_proof_for_non_isolated_workspacedecided isolation purely from the repo-writable.brigade/security.toml, so a same-uid scanner could setauthority_store.isolationback to"off"or malform the file (invalid config normalizes to off) and, while the signed snapshot was missing, malformed, or mid-rotation, select the unsigned dedupe path where attacker-controlled sidecar records suppress a canonical import. Observing genuineexternal-keyisolation now persists a per-target posture marker outside the workspace at~/.brigade/authority-isolation/<target-fingerprint>(overrideBRIGADE_USER_DIR, beside the existingauthority-signedsticky state), and the unsigned fallback is refused whenever that marker exists or the workspace configuration exists but cannot be parsed; only a healthy config that genuinely readsoffwithout a marker admits it (with the same one-time downgrade warning). Clearing requires the existing explicitbrigade security authority downgrade --target <workspace> --confirmpath, which now removes both sticky markers under the same audit-and-restore discipline; no new CLI verb was added. The isolated-refusal regression test is parameterized over intact, missing, malformed, and rotating store states, asserts the recordless persisted-proof validator is never consulted, and asserts stderr stays free of the downgrade warning; new tests pin marker creation, both refusal attacks, the allowed-with-warning non-isolated case, and downgrade clearing. The fallback’s docstring and warning wording no longer claim receipt validation - the unsigned path validates only the persisted sidecar against the workspace’s own authority record, asdocs/security.mdalready described. - Import authentication now follows a deliberate two-posture split (#881 operator decision). Legacy identity grants stay signed-only in every workspace:
_legacy_import_source_content_identityand the existing-migration-proof path keep requiring one verified authority snapshot, so forged receipt or proof bytes never grant legacy source/content identity anywhere. Canonical dedupe suppression, by contrast, follows the workspace posture: withauthority_store.isolation = "external-key"it keeps requiring that verified snapshot plus its bound persisted sidecar, but in a workspace without external-key isolation no signed store can exist (and a same-uid writer there can already rewrite the import inbox directly; issue #1093 tracks that boundary), so suppression returns to the pre-#881 evidence grade - a persisted sidecar plus reproducible receipt validated against the workspace’s own authority record - via the explicit_unsigned_dedupe_proof_for_non_isolated_workspacepredicate, which logs one warning per process naming the downgrade. This restores duplicate suppression (brigade work import ingestno longer re-imports the same canonical row) for every non-isolated workspace while keeping isolated workspaces on the signed path; the round-1 fixture that asserted unsigned-dedupe rejection without isolation was converted to external-key isolation rather than weakening its assertions, and new tests pin both postures. Documented under “Canonical import dedupe and the isolation posture” indocs/security.md. Fixes #881. - Legacy import authentication and canonical dedupe now anchor to exactly one verified authority snapshot per decision, closing two remaining #881 defects. Receipt, directory, and sidecar validators previously re-read the external authority store independently and discarded the verified snapshot’s record, so a concurrent same-uid writer alternating a forged unsigned binding with a saved valid signed envelope could combine evidence that never existed in one authenticated snapshot (granting a legacy identity or suppressing an incoming import);
_authenticated_legacy_import_proofnow reanchors first, reads one signed record via_read_verified_authority_snapshot, and passes that record into_has_locally_stamped_import_proof,_has_persisted_import_proof, and every directory/file validator, none of which reopen the store. Workspace relocation also no longer mis-orders trust: signedness was evaluated before the proof-directory open triggered authority reanchoring after a rename, soauthority_is_signedstayed false on the not-yet-reanchored store and the first matching canonical import was duplicated instead of suppressed;_acquire_authenticated_authority_snapshotnow runs the reanchor pass before its single verified read and is the one acquisition point used by_append_import_records. Regression tests cover both: a store swap between artifact validation and the final predicate is rejected (test_store_swap_between_validators_cannot_combine_snapshots), and end-to-end canonical dedupe after renaming a workspace without pre-opening authority directories suppresses the duplicate (test_canonical_dedupe_after_rename_without_preopening_authority). Fixes #881. - Concurrent trusted task-ledger writers serialize again instead of failing (PR #1207 follow-up). The round-4 sendback removed the bounded wait around
_acquire_task_ledger_locktogether with the inherited-lock inversion, but that made every overlapping writer of.brigade/work/tasks.jsonfail closed immediately withRunLockError: another brigade run appears active- breaking concurrent edge writes and the work-store characterization’s claim race, where the loser must exit 13 (already claimed) rather than 2. Task-ledger acquisition is deadline-bounded again with a nonblocking retry loop in theDEFAULT_LOCK_DEADLINE_SECONDSstyle: outside writers wait patiently up to 30 s and raise a typedTaskLedgerLockTimeouton expiry; inside a scanner run the bound stays short (5 s) and fails closed withRunLockError, so a marked child already holding the writer lock still cannot pin the task-ledger lock for a run window. Lock validation is unchanged (every attempt goes throughrunguard._acquire_lock’s stale reconciliation), and the documented lock order task ledger → run → writer is untouched. - Security sendback hardening for the import inbox, round 6 (PR #1207). Scanner launches are now supervised: once the direct scanner child exits,
proc.run(supervise_group=True)reaps the owned process group regardless of pipe state and recordsdescendants_reapedon the run receipt, so a forked descendant that closed every capture pipe while holdinginbox.jsonl.writer.lockcan no longer outlive the run and force every later scanner and outside writer intoInboxLockTimeout. A writer-lock timeout during output ingestion now marks the run failed (previously onlyrun["error"]was set), so the command returns nonzero instead of reporting rc 0 with every import rejected asinbox_lock_busy. The remaining canonical writers (import add,import context, single handoff promotion, single promotion, both dismiss modes, inbox archive) hold the canonical writer locks from their initial read through publication - promotion takes the task-ledger lock first, matching the documented order - so a scanner commit landing between a pre-lock read and the locked write survives instead of being deleted by stale state. Ingestion now publishes its run receipt inside the writer-lock window that spans the append, so a receipt failure’s snapshot rollback can never restore pre-append bytes over a trusted commit that landed in the release gap. The per-process inbox-lock registry keeps ownership and reentrancy depth in one entry with serialized initialization: an inner thread context stays protected when an outer context exits first, the descriptor closes exactly once at outermost exit, and concurrent first acquisitions wait for the owner instead of observing a half-initialized entry. - Security sendback hardening for the import inbox, round 5 (PR #1207). The in-process lock registry no longer keeps a stale entry after a lock is released:
_held_inbox_locknow drops the_ACTIVE_LOCKSentry atomically under the module lock before closing the descriptor on the outermost exit,verify_inbox_lock/verify_inbox_writer_locklook the entry up under that same lock andfstatthe live descriptor (regular file, single link, dev/ino equal to the lock path), so an acquire/release cycle can no longer leaveverify_*- and the protected writes gated on it, like a marked child’s_append_import_records_locked- succeeding while no OS lock is held; reentrant depth accounting still keeps the entry until the outermost exit. Scanner stamping and output ingestion now take the writer lock before each canonical snapshot and hold it through publication or rollback-state capture: stamping’s live inbox read moved inside a writer-lock span covering its publication (_scanner_stamp_new_importswraps the locked body), so a marked child committing between the launcher’s read and publication has its row appended after instead of overwritten, and ingestion captures its dedup input and rollback snapshot under the same lock that spans_append_import_records, so a receipt failure after a concurrent commit restores a snapshot that includes it instead of deleting it. Inbox lock acquisition is deadline-bounded on both platforms: POSIXflockretries nonblocking and Windows retriesmsvcrt.locking(LK_NBLCK)against a module default deadline (DEFAULT_LOCK_DEADLINE_SECONDS, 30 s, per-call overridable via the context managers’deadline_seconds), raising a typedInboxLockTimeout(aTimeoutError/OSError) on expiry - an escaped descendant holdingwriter.locknow fails launcher stamping with a recorded run error that proceeds through the existing bounded cleanup/rollback instead of hanging the launcher (and behind its run lock every outside writer), outside writers get the same typed failure, and the ingestion path records aninbox_lock_busyrejection. - Security sendback hardening for the import inbox, round 4 (PR #1207). The round-3 design passed the launcher’s held run-lock descriptor to self-importing scanner children (
BRIGADE_INBOX_LOCK_FDwithproc.run(pass_fds=...)); review found that unsound - any child or descendant sharing that open file description couldflock(LOCK_UN)the run-wide lock open to outside writers mid-run, sibling children adopting the same description got no mutual exclusion between their self-imports, a child holding ambient inbox ownership could invert the documented task-then-inbox lock order and strand the task-ledger lock for the run window, and on Windows the msvcrt adoption branch never actually acquired anything. Inherited-lock adoption is deleted entirely (no descriptor is ever passed to a scanner child, andproc.runno longer acceptspass_fds). Exclusion now uses two locks opened and validated identically (held no-follow dirfd walk,O_NOFOLLOW|O_NONBLOCK, regular file, single link, dev/ino re-check before protected writes): the existinginbox.jsonl.lockrun lock, which only the launcher holds for its whole run window to exclude outside writers, and a newinbox.jsonl.writer.lockthat serializes each canonical read-modify-write (imports, backfill, promote, dismiss, rollback, scanner stamping). Outside writers take run then writer; a self-importing child is recognized by a new capability-free environment marker (BRIGADE_SCANNER_RUN_ID, set by the launcher) and takes only the writer lock it opens itself - so siblings serialize on their own open file descriptions and no child can release the launcher’s run lock; forging or stripping the marker only demotes a process to writer-only or outsider behavior, never breaks exclusion. The global lock order everywhere is task-ledger → run → writer, a test asserts every caller follows it, and the task-ledger acquisition’s five-second wait/failure path that bounded the old inversion is removed in favor of immediate fail-closed behavior; on Windows both locks use the same validated msvcrt path and the marker only selects which locks are taken. - Security sendback hardening for the import inbox, round 3 (PR #1207). The canonical inbox writers (
ledger._write_imports,_append_import_records,_backfill_import_provenance, and_promote_matching_imports) now hold the same run-wide writer lock as scanner runs - moved into a sharedbrigade/work_cmd/inbox_lock.pymodule, reentrant per process - so a concurrentbrigade import, promote, dismiss, or provenance backfill can no longer interleave with scanner stamping or rollback and lose or misattribute trusted rows. Self-importing scanner children join their own run’s exclusion window through an inherited lock descriptor (BRIGADE_INBOX_LOCK_FD, passed withproc.run(pass_fds=...)), so locking writers no longer deadlocks them against their parent run; on Windows (no descriptor passing) that child-side serialization remains a documented residual. The lock file itself is hardened: on POSIX it is opened through a held no-follow directory descriptor withO_NOFOLLOW|O_NONBLOCK, validated byfstatas a single-link regular file (planted symlinks, FIFOs, and hard links are refused), and every protected write re-checks that the path still names the locked device/inode, failing closed when the lock was replaced mid-run. The reviewed-inbox read (_read_import_inbox_raw) now enforces the same 4 MiB snapshot cap and full before/after metadata comparison (size, mtime_ns, bytes-read length) as the publication snapshot, so mid-read growth and in-place mutation fail closed. On Windows, a suspended child whoseResumeThreadcall fails is now reported as a resume failure and terminated instead of being left frozen on its job assignment until timeout. - Runbook step timeouts no longer orphan descendant processes. Step execution and verification-contract verifier/rollback commands ran via
subprocess.run(..., shell=True), whose timeout handler kills only the direct child (the shell), leaving grandchildren alive. Both paths (runbook_cmd._shell_or_argv_commandand the step loop) now launch children in their own process group and route timeouts through the same shared cleanup path as a normal-exit cancellation (brigade.proc.terminate_process_tree: process-group SIGTERM, grace wait, SIGKILL), then still write the step logs and receipt withexit_code: 124/timed_out: true. Fixes #1190. - Security sendback hardening for scanner-run subprocesses and the scanner inbox (PR #1207 round 2). Windows child launches now create the owned kill-on-close Job Object before
Popen, start the childCREATE_SUSPENDED, assign it to the job while it cannot run, and only then resume its main thread - closing the grandchild-spawn window between launch and assignment; job-creation or assignment failure is a typed nonzero launch failure that terminates the still-suspended child instead of silently falling back to taskkill on an unguarded tree. Scanner runs now hold a run-wide inbox lock (an adjacentflock/msvcrtlock file beside.brigade/work/imports/inbox.jsonl, reentrant per process) from the pre-run snapshot through stamping or rollback, every inbox writer path inscanners.pyhonors it so honest Brigade writers can no longer interleave with a run, and full metadata (device, inode, size, mtime_ns) plus the bytes-read length are revalidated under the lock immediately before each launch, aborting the run when a same-UID writer replaced or modified the inbox in the snapshot-to-launch gap so substituted rows can never be stamped as scanner-created. A cleanup failure during group termination (proc.run) is now converted into a typed nonzeroResultwith an explanatory stderr note andincomplete_process_groupset, and an exception after_scanner_run_onepersists its running receipt finalizes that receipt as failed, so neither path can strand a “running” receipt that blocks later runs. The ledger import-inbox snapshot’s before/after compare now includes size and mtime_ns and requires the joined bytes to matchst_size, detecting mid-read appends that keep the inode unchanged. - The post-exit drain cutoff in
proc.runno longer lets a terminated process group stand as success. When a scanner or worker child exits while a descendant still holds its output pipes, the group was reaped after the brief drain window but the child’s zero exit code survived: the scanner was marked completed and granted a run proof, and any output the terminated group members had produced (or would still have written) was silently discarded as a successful result. The cutoff now records a typedincomplete_process_groupfailure onproc.Result, forces a nonzero exit code so_register_scanner_run_proofnever fires, marks the scanner run failed with an explanatory error, andrun_agentreturns afailure_kind="incomplete-process-group"harness failure instead of accepting partial nonempty output. On Windows the tree is now terminated through an owned kill-on-close Job Object bound at launch (falling back to the previous taskkill path only if job creation or assignment fails), instead oftaskkill /Ton an already-exited parent PID with descendants outside any kill-on-close scope. - Scanner pre-run inbox snapshots no longer swallow
OSError. Previously any snapshot failure (including the new 4 MiB read limit) degraded tobefore_raw=b"", so a scanner that truncated the oversized inbox below the cap and then failed had rollback publish an empty value over the real pre-run inbox; a missing pre-run inbox was likewise restored as an existing empty file. A single bounded read (_snapshot_scanner_inbox) now returns raw bytes, parsed rows, an existence flag, and stable before/after fstat identity;_scanners_run_payloadaborts before launching the scanner when the snapshot cannot be captured; rollback refuses when the live inbox inode no longer matches the snapshot; a torn same-UID mid-read write is detected via size/mtime metadata comparison; and a previously missing inbox is restored as missing. - Rollback snapshot helpers no longer accumulate unbounded chunks after their guard checks.
_scanner_inbox_descriptor_bytesand ledger_snapshot_import_inboxenforce their byte caps while reading each descriptor chunk and raise the typed limit exceptions (ScannerInputLimitExceeded/ImportInboxSnapshotLimitExceeded) instead of relying on an earlier size check. - Built-in scanner runs no longer use unbounded
subprocess.run(capture_output=True). Scanner child processes now run through the bounded streaming capture inbrigade.proc(runwith its own process group), so a scanner that floods stdout/stderr trips a typed failure (output_limit_exceeded: true, error naming the 1 MiB capture limit, run marked failed) instead of exhausting memory, and on overflow or timeout the whole process group is reaped so a descendant holding the output pipes can no longer outlive the run or block completion past the scanner’s own exit. The direct child’s exit is no longer held hostage by pipe-holding grandchildren: output drains briefly after the child exits and remaining group members are terminated. Fixes #1197. - Scanner input reads are now byte-bounded before parsing. Inbox reads cap total bytes (4 MiB) and reject any single record over a per-record cap (256 KiB); scanner JSONL import ingestion enforces the same total and per-record limits while streaming the file; and rollback snapshots of pre-existing run artifacts are bounded instead of reading whole files into memory. Oversized input fails with the typed errors
ScannerInputLimitExceeded/ScannerSnapshotLimitExceeded(subclasses ofOSError, so existing fail-closed paths surface them as rejections) instead of scaling memory with attacker-controlled files. Fixes #1198. - A failed scanner run can no longer mint eligible work through the self-import stamping path.
_scanner_stamp_new_importspreviously accepted and rebuilt new inbox rows regardless of run status, so a scanner could append a pending row and still create work after exiting nonzero. On a failed (or timed-out / nonzero-exit) run, the pre-run inbox is now restored exactly and all rows the scanner appended are rejected withself_import.rejection_reasons.run_not_completed; only rows covered by an exact successful-run proof are stamped. The file-based--ingest-outputpath already skipped failed runs; the in-memory path now matches it. Fixes #1199. - Security-review round 6 for #1177: state-root content is never read back by pathname. Install destinations living inside
.brigade(the built-inmcptarget’smcp-resources) are inspected through the held state-root anchor, so a planted symlinkedSKILL.mdcontributes absence instead of leaking outside bytes intoskills diff, fleet status, compatibility drift, or sync rollback material (skills sync --writerefuses a destination holding non-plain entries for that run instead of writing over or around it). Selector classification against the skills state root is now lexical and precedes any existence probe: an explicit registry pathname selects the anchored entry for lint/diff/install, and every other location under.brigadeis refused as a source rather than followed; registry imports and inbox proposals lint their already-collected bytes from private staging (with provenance repointed), so nothing re-reads the freshly written state through the filesystem.skills sync --writerepeats the eligibility gates (trust level, enablement, supported harnesses) on the projection snapshot itself, so a generation swapped in after planning is re-gated instead of installed. Skill pack paths are classified lexically before resolving - a pack directory replaced by a symlink escaping the state root no longer poses as a trusted external pack forpack show/pack import, which serve recognized state-root packs only through anchored snapshots. Failed installs and accepted proposals keep ephemeral staging directories out of errors, metadata, and receipts by running the staged-payload repointer on every exit path and passing original operator provenance into proposal imports. Refs #1177. - Security-review round 5b for #1177: the last pathname-based reads of skills state-root content are closed.
skills sync --writecollects one anchored snapshot per registry entry before lint and derives rendering, fingerprints, receipts, and every mutation from it (history and prior receipts are read through_read_state_file_bytes, so a symlinked history or receipt refuses with a typed error naming the file instead of embedding outside JSON); registry read consumers (skills lintonregistry:selectors,skills diff --against registry,skills search --jsonmetadata, MCPget_skill/metadata/changelog/compatibility/lint resources, fleet status, harness profile packages, doctor health, pack build evidence) enumerate entries and read SKILL.md/metadata/changelog through descriptor listing plus_read_state_file_bytes, reporting symlinked entry dirs or symlinked SKILL.md/skill.json as refused or absent without echoing outside file contents (configured changelog paths that escape the entry are ignored);skills pack list/showenumerate packs by held dirfd and read manifests through_read_state_file_bytes; install collectsregistry:sources through the anchor; adapter overlays, fleet receipts, and inbox-proposal fingerprints resolve through anchored reads; and error messages never echo content from refused entries. Refs #1177. - Security-review round 5 for #1177: every remaining skills command mutation and state read honors the held
.brigadeanchor.skills publish,skills inbox accept/reject,skills pack build, andskills pack archivewrite, import, and move only through the anchored primitives (a symlinkedpublish-proposalsor proposal directory refuses with exit 2; a manifest-supplied physical path is never honored as an archive move source; scope and destination names derive from validated slugs), and all of these commands fail closed without descriptor anchoring instead of mutating through an unanchored path.skills installcollects the source snapshot before any validation so metadata, lint, rendering, and fingerprints always describe exactly the bytes it installs, and records arollback_snapshot_fingerprintat capture time (install and sync alike) thatskills rollbacknow verifies - together with fresh-install lint and rendered-text validation on the snapshot bytes - before restoring anything, refusing modified or fingerprint-less legacy snapshots without touching the destination. Receipt replaces and history appends open withO_NONBLOCKbefore the fstat check, so a planted FIFO can no longer block the command. Read paths (skills diff --json,skills history --json, and inbox list/show/diff) resolve receipts,history.jsonl, registry copies, and proposals through_read_state_file_bytes/descriptor listing with an lstat-guarded fallback, reporting symlinked entries as refused/absent instead of returning outside file contents. Refs #1177. - Security-review round 4 for #1177: skills state-root primitives now validate the opened descriptor before mutating. Writes open without
O_TRUNC, tryO_EXCLfirst, require a single-link regular file viafstat, and only then truncate and write; appends validate the descriptor before the first byte moves; and reads (anchored and Windows-fallback alike) refuse FIFOs, non-regular files, and hardlinks viafstat/lstat(O_NONBLOCKkeeps a planted FIFO from blocking), so a hardlink from a receipt orhistory.jsonlto a same-filesystem victim can no longer be truncated or appended through a skills write.skills installreads the canonical previous receipt through the anchored reader and retains only receipts that satisfy the contract - a symlinked receipt now refuses the install instead of copying outside JSON intoprevious_receipt, state, or JSON output. The install source tree is collected once into an immutable byte snapshot and render, source fingerprints, installed fingerprints, and both the anchored and external harness copies materialize exclusively from those bytes (a symlinkedSKILL.mdis skipped by collection and refuses the install rather than being re-read by pathname into the rendered install), and nothing is read back through a pathname to compute receipt fingerprints. On platforms without descriptor anchoring every state-root mutation primitive (write, append, unlink, remove tree, copy into anchor) now fails closed with a typedSkillsStatePathErrorinstead of mutating through an unanchored path; reads keep the fallback. Finally, unexpectedOSErrors inside the anchor primitives are translated intoSkillsStatePathErrorwith the original exception preserved as__cause__, so a mid-operation race surfaces as a typed refusal instead of a traceback after partial mutation. Refs #1177. - Security-review round 3 for #1177: the last path-based
.brigademutations inskills_cmd.pygo through the held state-root anchor.skills rollbackno longer resolves the receipt-bound snapshot through the filesystem - symlinkedinstallsorrollback/<skill>/<harness>components now refuse the whole operation before anything is read, copied, written, or deleted; the snapshot is read back throughdir_fddescriptors (O_NOFOLLOW, symlink entries skipped), restored into the install target from those bytes, and the canonical/rollback receipt writes, receipt unlink, and snapshot consumption all run against the anchor (state-internal install targets such asmcp-resourcesare removed and rewritten through the anchor too, and the snapshot binding is matched textually against the un-resolved rollback root so a swapped component can no longer launder an external directory into a “valid” snapshot).skills inbox addanchors the proposal copy, metadata write, existence check, and--forcereplacement the same way, so a symlinkedinboxcomponent can no longer redirect the proposal outside the workspace or make--forcedelete a matching external subtree. Source-tree collection walks sources through held descriptors with single-lstat validation andO_NOFOLLOWreads (a file swapped to a symlink between check and use is refused instead of followed) and rejects regular files with more than one link (hardlinks), for registry imports, inbox proposals, and install snapshots alike. Two error diagnostics that referenced an undefined name (UnboundLocalErrorinstead of the typed refusal) in the anchored remove/unlink paths now format the anchored state path correctly. Refs #1177. - Terminal run receipts are no longer unwritable once #578 retry decisions exist.
retry_decisionswas written at the top level of run.json without being registered in the projector’s preserved-field ownership inventory, so on journal-authoritative runs the shadow parity projection recordedprojection-error:UnknownSnapshotFieldError, every later write failed the authoritative prior gate (authoritative run prior gate not ready: error-recorded), and failure/interrupt terminal receipts raisedRetainRunLockError- killing the worker thread before it could record its exit code.retry_decisionsis now a preserved field inrun_projector.PRESERVED_FIELDS(alongsideseat_routingandhealth), so receipts carrying retry decisions project cleanly on every terminal path (success, failure, interrupt, resume) while the closed field inventory still rejects genuinely unknown fields. - Runbook step timeouts no longer orphan descendant processes. Step execution and verification-contract verifier/rollback commands ran via
subprocess.run(..., shell=True), whose timeout handler kills only the direct child (the shell), leaving grandchildren alive. Both paths (runbook_cmd._shell_or_argv_commandand the step loop) now launch children in their own process group and route timeouts through the same shared cleanup path as a normal-exit cancellation (brigade.proc.terminate_process_tree: process-group SIGTERM, grace wait, SIGKILL), then still write the step logs and receipt withexit_code: 124/timed_out: true. Fixes #1190. - Skills write paths no longer trust the
.brigadestate root as a plain directory. A symlinked.brigadestate root that would redirect skill registry imports, install receipts, rollback snapshots, uninstall history, or other skills writes outside the workspace is refused up front with a typedSkillsStatePathError, and an ancestor swap between validation and write is detected by reopening the state root withO_DIRECTORY|O_NOFOLLOWand comparing device/inode against the descriptor held from validation; receipt and history writes go throughdir_fd-relative creation against the held descriptor. On Windows, where directory-descriptor writes are unsupported, the state root is resolved strictly (symlink components and external resolution rejected) immediately before each write instead, leaving a small documented check-then-use window on that platform. Read-only fleet status behavior from #1175 is unchanged. Fixes #1177. - Security-review follow-up for #1177: the remaining path-based
.brigademutations inskills_cmd.pynow go through the held descriptor anchor. Registry import copies and install rollback snapshots use a recursive copy that opens every destination component withO_NOFOLLOW|O_DIRECTORYrelative to the anchor (source symlinks are skipped, never followed); removal of state-internal installed copies during reinstall/uninstall uses adir_fd-relative recursive delete; receipt deletion uses an anchored unlink; and the duplicate raw path-based receipt write in the install loop was removed so the anchored write is the only one (canonical receipt schema unchanged). A nested symlink such as.brigade/skills/installs,.brigade/skills/rollback, or.brigade/skills/mcp-resourcesnow refuses the whole operation before anything is written or deleted outside the workspace instead of redirecting one step and failing later. On platforms without descriptor anchoring (Windows fallback), state writes additionally reject symlink and reparse-point components across every destination component and strictly re-resolve the parent immediately before each write; the small check-then-use window that remains there is documented on_StateRootAnchor. Creating a missing.brigadestate root now opens and holds the workspace directory first and creates/opens the root relative to that descriptor, refusing an ancestor swapped during the creation window instead of letting both created root and validation land outside the workspace. Refs #1177. brigade runno longer loses a worker’s result or poisons the seat when the combined app-server output stream exceeds the 1 MiB capture cap. An over-cap codex app-server turn (failure_kind=output-limit) with salvageable final-answer text is now delivered bounded through the normal envelope gate as a successful worker result, with the truncation recorded viaoutput_truncated: truein worker-results.json and the original “combined output exceeded … byte limit” detail kept; an over-cap turn with no salvageable text stays a typed failure but no longer triggers same-seat retry or quarantine. Direct CLI overflow keeps its existing #1012 contract. Fixes #1144.- Fleet spool events are no longer lost silently on platforms without descriptor-relative file APIs (Windows; #1157 round 3 regression).
_ensure_private_dirdocumented returningNonethere but always returned a directory descriptor, sending_spool_lockand_write_spool_atomicinto thedir_fd=/src_dir_fd=/dst_dir_fd=calls CPython documents as unsupported on Windows; the resulting failure surfaced only asreport_event’s silentreturn False, so undelivered events vanished without a trace. Descriptor support is now detected once per call (sys.platformplusos.supports_dir_fdmembership ofos.open,os.rename- which registersos.replace’s dir_fd support,os.unlink, andos.stat): unsupported platforms take the documented path-based fallback (Nonereturned, a symlinked spool directory still refused, best-effort 0700 path chmod,tempfile.mkstemp+os.replacerewrite), and any spool write that fails anyway is logged at WARNING with the reason instead of being swallowed. Regression tests simulate the unsupported platform via monkeypatchedos.supports_dir_fd; POSIX descriptor-scoped behavior is unchanged. Fixes #1157. - Fleet abort hardening residuals (#1157 round 3).
brigade run’s credential-failure callback now appends its classification marker before printing the diagnostic and fires_thread.interrupt_main()from afinally, so a failed stderr write can no longer swallow the callback and let the run continue after the hub rejected this machine’s token (receipt kindfleet-credentials-rejectedeither way); the claim-lost callback records its marker before its fallible print, so a dead stderr yieldsfleet-claim-lost, not a user cancel. Fleet spool atomic rewrites are now descriptor-scoped on POSIX: the exclusive unpredictable temp file (os.open(O_CREAT|O_EXCL|O_NOFOLLOW, 0600)) and the finalos.replace(src, dst, src_dir_fd=…, dst_dir_fd=…)address names relative to the validated spool-directory descriptor instead of re-resolvingpath.parent, so a concurrent directory swap cannot redirect the rewrite (tempfile.mkstempremains only for the no-descriptor fallback). The claim-loss abort’s watchdog grace fire and the callback-finally fire share one lock-guarded check-and-set on the one-shot fired flag, so a callback ending exactly at the grace boundary interrupts exactly once. Fixes #1157. - Fleet claim-loss abort no longer depends on the
on_claim_lostcallback behaving (#1157 round 2). The interrupt now fires independently of callback execution: a callback that blocks forever is cut off after a bounded grace period (fleet_client.CLAIM_LOST_CALLBACK_GRACE_SECONDS) and a callback raisingSystemExit/KeyboardInterrupt(anyBaseException, not justException) is logged - the guarded block is interrupted either way, after the callback has had its chance to record state. A lost-claim abort entered from a non-main thread no longer calls_thread.interrupt_main()(which targets the process main thread and would stab an unrelated thread while the owner kept working): the yieldedClaimDecision.cancel_eventis set instead, and guarded work off the main thread must observe it to unwind (documented indocs/fleet-sync.md). The spool directory’s 0700 privacy is now enforced through a descriptor rather than check-then-chmod on the path: it is openedO_DIRECTORY|O_NOFOLLOW, forced viafchmod, re-verified withfstat, and lock/spool files are openeddir_fdagainst it; when privacy cannot be applied or verified the write is refused instead of warned about. Hub redirect same-origin checks normalize default ports, sohttp://hub→http://hub:80is followed while cross-origin hops stay refused.brigade runclassifies and records a fleet claim-loss abort on the main thread from the heartbeat’s marker, so a heartbeat-thread receipt-write failure can no longer record the abort as “canceled by user”; regression tests assertrun.json’sfailure.kind, not stderr. Fixes #1157. - Raw
scripts/verify-focusedinvocations are now denied by Claude and Grok work-loop hooks, matchingscripts/verify, so focused development checks still run throughbrigade work verifyand produce verification and outcome receipts. - Over-cap codex app-server records are charged exactly once against the capture budget: the ingestion path no longer double-charges every valid parsed record, so normal output between half the cap and the full cap ingests cleanly instead of failing as over-limit (#1200).
- A parsed
turn/completedrecord that overflows the shared capture budget now records its completion metadata before the output-limit signal is published, so a turn waiting on the limit can no longer wake, consume its exact(thread_id, turn_id)entry while it is still absent, and reportcompleted_observed=Falsefor a turn that genuinely completed (#1200). - An oversized app-server record - anything the reader cannot hand to
json.loadswithin the capture cap - never signals turn completion, even when it embeds a well-formedturn/completedenvelope: a trailing second top-level object, invalid UTF-8 bytes, malformed numbers like01, or missing closing delimiters after"status":"completed"can no longer forgecompleted_observedthrough structural salvage. The hand-written bounded record scanner was removed entirely; completion is recorded only for recordsjson.loadsparses successfully, including a parsed record that then overflows the shared budget (#1200). - App-server
_observed_completionsis keyed by(thread_id, turn_id)and bounded: interleaved turns on one thread each keep their own completion instead of overwriting a single per-thread slot, building an over-cap result atomically reads-and-removes exactly that turn’s entry under the state lock (a sibling completion arriving between lookup and consume is no longer deleted),reset_captureclears only the entries of the thread being reset so a reused turn id cannot inherit a stale completion, and entries beyond a fixed bound are evicted oldest-first (#1200). - Over-cap codex app-server salvage is now reachable through the real stream and correctly gated on turn status. Bounded ingestion records the turn id and status of every
turn/completednotificationjson.loadsparses successfully, including a parsed record that then overflows the shared budget, so a turn that genuinely completed while the capture cap was hit still yieldscompleted_observed=Trueinstead of being silently dropped before routing. The salvage signal additionally requiresstatus == "completed": an observedturn/completedwith statusfailedorinterruptedno longer reaches run transport as a successful worker result or satisfies downstream DAG prerequisites (#1200). brigade runno longer loses a worker’s result or poisons the seat when the combined app-server output stream exceeds the 1 MiB capture cap. An over-cap codex app-server turn (failure_kind=output-limit) with salvageable final-answer text is delivered bounded through the normal envelope gate with the truncation recorded viaoutput_truncated: truein worker-results.json and the original “combined output exceeded … byte limit” detail kept; an over-cap turn with no salvageable text stays a typed failure but no longer triggers same-seat retry or quarantine. The salvaged text counts as a successful worker result only when aturn/completedevent for the turn was observed on the app-server stream; otherwise it rides along as anok=Falsediagnostic that does not satisfy downstream DAG prerequisites (#1200). Direct CLI overflow keeps its existing #1012 contract. Fixes #1144.- Pre-run snapshots no longer fail on worktrees with very large untracked path sets. Untracked enumeration now streams
git ls-files --others --exclude-standard -zthrough a new NUL-delimited, incrementally decoded capture (proc.run_delimited) under its own 16 MiB budget instead of the generic 1 MiB child-output cap, so a listing that exceeds the old cap (for example 21,876 untracked paths) captures ground truth instead of aborting the run before dispatch. Enumeration failures name the exact stage and budget (“untracked-path list exceeded … byte enumeration limit”), unrelated commands keep the generic cap, and snapshot receipts summarize path sets beyond 200 entries via*_totalcounts so run.json and pre-run-snapshot.json stay bounded. Fixes #1165. - Center operator reports no longer grow without bound. Each
brigade center report buildembedded the full previous report bundle throughstatus.operator_report.latestin the status payload, so every consecutive build folded the entire prior history into the new evidence file (measured: 70 KB, 149 KB, 243 KB, … per build). Report health now returns a bounded reference to the latest report (id, timestamps, path, counts, digest) instead of the full body, both for center reports (center_cmd/reports.py) and repo-fleet reports (repos_cmd/fleet_health.py, which fed the same center status payload viarepo_fleet). Consumers that only need the report identity keep working; the full bundle remains on disk and is loaded on demand byreport show. Fixes #1164. brigade run --worker <seat>no longer runs orchestrator health routing on a direct-worker run. Only the selected worker and its declared fallbacks are health-gated, so an unhealthy orchestrator the worker never uses can neither reroute nor abort the run; the seat-routing receipt (when present) is markedrun_mode: direct-workerand omits unused orchestrator decisions. Fixes #1174.- The effective sandbox override (
--sandbox read-only, or read-only mode’s implied hard isolation) now reaches the seat-health hard-isolation preflight instead of the probe judging only the static roster declaration, so a Codex worker whose transport passes can satisfy hard isolation through the command-line override. The judged value is recorded in run.json underhealth.effective_sandbox. Fixes #1173. - The legacy import migration path no longer trusts scanner-reproducible receipt and persisted-proof bindings recorded in an unsigned external authority store. A same-uid writer could forge a post-run
receipt.json(and matching sidecar proof) and re-bind the unsigned store to the new bytes, granting the forged row a legacy source/content identity during untrusted-identity migration._legacy_import_source_content_identitynow requires the workspace’s authority record to be a verifier-signed HMAC envelope, so workspaces withoutauthority_store.isolation = "external-key"(or after an intentional downgrade) fail closed instead of accepting forgeable provenance. Fixes #881.
Changed
- The shipped review-heavy
reviewer_flashpreset now uses Gemini 3.8 Flash Low on Antigravity.brigade runmaps a bare Gemini API id such asgemini-3.8-flashto the matching low-effort agy slug, because agy requires--efforton the unsuffixed id. - README comparison cells now name 0.27 named retry/reroute, external harness sessions, and the approved Grok Bot scout feed. Brigade cells were re-checked on 2026-08-27. Beads and Gas Town stay on the 2026-08-26 pass. Other neighbor cells stay on 2026-08-13.
- README comparison matrices and scope now name the 0.27 optional fleet hub, vault project/search/propose, and run child/diff/resume. Beads and Gas Town were re-read on 2026-08-26. Other neighbor cells stay on the 2026-08-13 pass.
- Fleet claims fail closed on lost ownership (#1152). When the mid-run heartbeat learns another owner holds the claim (a renew answered 409 held-by-another),
repo_claimno longer logs and lets the run continue unarbitrated: by default it logs one WARNING and aborts the guarded run (with no callback, via the same main-thread interrupt path as Ctrl-C), so two machines can never keep working the same repo after a supersede or expiry. Passon_claim_lostto be notified first - the abort still follows the notification even when the callback raises or returns without stopping the run (#1157) - or setBRIGADE_FLEET_CLAIM_LOSS=continue(orclaim_loss_policy="continue") as the documented opt-out for solo machines, which restores the previous log-and-continue behavior.brigade runrecords a mid-run claim-loss abort as a failed run with failure kindfleet-claim-lostinstead of “canceled by user”. Fixes #1152. - Fleet hub traffic now bypasses HTTP proxies (#1154): the fleet client’s opener is built with an empty
ProxyHandler, soHTTP_PROXY/HTTPS_PROXYcan neither intercept the plain-HTTP bearer token en route to a tailnet hub nor blackhole its requests. The local event spool is hardened in the same pass: the spool directory is created 0700 and spool, lock, and temp files 0600 regardless of umask, opens are no-follow (O_NOFOLLOWwhere available) and refuse non-regular files such as FIFOs or symlinked spool paths, and atomic rewrites use exclusive, unpredictable temp names in the spool directory instead of the predictable.tmppath. Mode assertions are POSIX-only in tests. A symlinked spool directory is refused rather than followed, existing permissive spool files are forced back to 0600 on every open, and a chmod that cannot be applied is logged instead of swallowed (#1157). The opener also refuses cross-origin redirects: urllib’s default handler replays the Authorization header on every hop, so only same-scheme, same-host redirects are followed and any other 3xx surfaces as an ordinary transport failure (#1157). Fixes #1154. - Fleet claim orphan cleanup closes its last leak windows (#1157). The background release scheduled after a lost acquire response now attempts immediately before backing off, so a
brigade runthat finishes within one backoff interval still frees the row instead of leaving it for the full TTL; amissinganswer inside one outstanding-request uncertainty window is retried instead of accepted as definitive, because the abandoned acquire can still commit its row afterwards; at exit, when an abandoned heartbeat re-acquire may still commit, the inline release retry also runs after a definitivemissinganswer - not only onhub-unavailable- and each retry is now spaced across one request deadline so back-to-back retries cannot all complete before the straggler POST lands. Exhausted exit retries now log a WARNING naming the residual TTL exposure instead of tearing down silently. - Dropped the remaining references to the retired
gpt-5.3-codex-sparkreview seat from the roster docs and tests; the roster init example now usesgpt-5.6-terra. - Fleet claims fail closed on lost ownership (#1152). When the mid-run heartbeat learns another owner holds the claim (a renew answered 409 held-by-another),
repo_claimno longer logs and lets the run continue unarbitrated: by default it logs one WARNING and aborts the guarded run (with no callback, via the same main-thread interrupt path as Ctrl-C), so two machines can never keep working the same repo after a supersede or expiry. Passon_claim_lostto be notified first - the abort still follows the notification even when the callback raises or returns without stopping the run (#1157) - or setBRIGADE_FLEET_CLAIM_LOSS=continue(orclaim_loss_policy="continue") as the documented opt-out for solo machines, which restores the previous log-and-continue behavior.brigade runrecords a mid-run claim-loss abort as a failed run with failure kindfleet-claim-lostinstead of “canceled by user”. Fixes #1152. - Fleet hub traffic now bypasses HTTP proxies (#1154): the fleet client’s opener is built with an empty
ProxyHandler, soHTTP_PROXY/HTTPS_PROXYcan neither intercept the plain-HTTP bearer token en route to a tailnet hub nor blackhole its requests. The local event spool is hardened in the same pass: the spool directory is created 0700 and spool, lock, and temp files 0600 regardless of umask, opens are no-follow (O_NOFOLLOWwhere available) and refuse non-regular files such as FIFOs or symlinked spool paths, and atomic rewrites use exclusive, unpredictable temp names in the spool directory instead of the predictable.tmppath. Mode assertions are POSIX-only in tests. A symlinked spool directory is refused rather than followed, existing permissive spool files are forced back to 0600 on every open, and a chmod that cannot be applied is logged instead of swallowed (#1157). The opener also refuses cross-origin redirects: urllib’s default handler replays the Authorization header on every hop, so only same-scheme, same-host redirects are followed and any other 3xx surfaces as an ordinary transport failure (#1157). Fixes #1154. - Fleet claim orphan cleanup closes its last leak windows (#1157). The background release scheduled after a lost acquire response now attempts immediately before backing off, so a
brigade runthat finishes within one backoff interval still frees the row instead of leaving it for the full TTL; amissinganswer inside one outstanding-request uncertainty window is retried instead of accepted as definitive, because the abandoned acquire can still commit its row afterwards; at exit, when an abandoned heartbeat re-acquire may still commit, the inline release retry also runs after a definitivemissinganswer - not only onhub-unavailable- and each retry is now spaced across one request deadline so back-to-back retries cannot all complete before the straggler POST lands. Exhausted exit retries now log a WARNING naming the residual TTL exposure instead of tearing down silently. brigade work briefandbrigade daily statusnow bound their memory working set on large multi-repo workspaces (issue #266): the fleet health section of the brief is compacted before any other section payload is computed, daily-status center sections keep only counts and top-item references instead of full queues and draft lists, per-repogit status --porcelainoutput is streamed line-by-line instead of buffered whole, and per-repo artifact JSON larger than 8 MiB is referenced (oversized: true) rather than parsed into memory. Regression tests assert oversized artifacts are never decoded and that dirty-count streaming stays under a fixed traced-memory budget. Fixes #266.- Center Code Graph now leads with a linked summary strip showing connectivity and change insights, defaults the module map to the top 15 connected modules, and shows plain-language detail cards on selection. (#909)
brigade releasereadiness now requires security scan evidence to cover the candidate commit HEAD before permitting release preparation. (#675)brigade evidence statuscollects source item totals using SQLite’s covering index onitems.source_idinstead of scanning table rows. (#873)
Added
- Fleet run preference overlay (#1223): hub schema v5 stores one
impl/review/chefpin,brigade fleet preference get|set|pullcaches it locally,brigade fleet statusshows it, andbrigade runusesimplas the default worker when--workeris omitted and that seat exists.--workerand a spoken roster seat name still win; two named seats leave chef planning. This is a routing overlay, not roster sync. brigade run cloud grokbot reconcile-reportspreviews or writes at most--limitdeterministic Memory Handoff drafts from completed Repository Scout report snapshots. Preview writes nothing. Apply verifies private snapshots, quotes every report line into the canonical owner’s review inbox, and records private queue-side idempotency markers so repeats do not duplicate. Report text is never printed or stored in job JSON, receipts, or marker JSON.- Repository Scout completions can store a private, queue-owned report snapshot of at most 12,000 UTF-8 bytes. Operator-only
grokbot_queue_reportreturns the stored text, byte count, and independently recomputed SHA-256. - Run admission and dispatch now persist retry decisions in the run receipt. Every failed worker attempt judged by the closed retry taxonomy (
decide_retry) appends a stable record -attempt,seat,failure_kind,decision,reason- to the run’s quarantine state, andbrigade runwrites those records to run.json underretry_decisions, so a later run can explain why a seat was skipped or retried. The receipt carries prior decisions forward across a resumed run (record_run_start preserves them, and later writes merge fresh decisions without duplicates), and runs that rerouted a seat through a declared fallback print one summary line (note: seat reroute: <requested> -> <effective> [<typed cause>]; skipped seats are named the same way). Direct--workeradmission routing behavior from #1183 is unchanged. Fixes #578. - A native Grok Bot MCP adapter now ships with Brigade under the optional
grokbotextra.brigade run cloud grokbotcan set up, serve, diagnose, canary, and render a user service for fixed operator, repository-scout, and implementation-worker roles without a second repository. Each role runs in its own process with an exact tool inventory, bearer authentication, host and origin allowlists, bounded requests, safe queue errors, and non-secret configuration. Generated services keep the fleet hub private and grant write access only to the selected target’s Grok Bot queue state. Fixes #1163. - Fleet-sync Phase 5 adds deterministic, read-only hub export with
brigade fleet export [--since <event-timestamp> | --since-received <hub-timestamp>] --format jsonl|csv [--claims [--include-expired]] [--db PATH] [--out PATH]: events stream in primary-key order with fixed columns and their digest, while claims stream by target without holder tokens and mark expiry explicitly. Claims exports exclude expired rows by default. Named output files are replaced only after a successful export.--sincefilters event timestamps, while--since-receivedincludes late-arriving spooled events for incremental archives. Filtered exports report timestamps they could not parse, and human CSV formula-like cells are prefixed with a single quote. The optional[fleet.sink]configuration remains disabled by default. When enabled,brigade fleet sinkimports raw values, incrementally upserts thefleet_eventsappend log using an atomicreceived_atwatermark, replaces the livefleet_claimssnapshot with expired rows excluded, and commits changed passes afterdolt add -Awith the UTC export time.[fleet.sink] timeout_secondsbounds each Dolt subprocess, andevents_incremental = falserestores full event exports. Fresh Dolt databases use a localbrigadeidentity. The subprocess-only sink does not change event reporting, claims, status, or dashboard paths. A missing Dolt binary is a documented no-op that prints the reason and exits 0. Fixes #1127. - Fleet-sync Phase 3: a hub-served web Fleet dashboard for a laptop or phone on the tailnet. The fleet hub now answers
GET /andGET /view/{machines,repos}with a server-rendered, stdlib-only page in thebrigade center servestyle (no framework, no CDN asset; tables, sort, and filter work with JavaScript off;<meta refresh>every 10s, with a small inline script that only ticks elapsed timers and adds a client-side text filter). The machine board shows one card per observednode_id(never a hardcoded host list) with its live runs (run, repo, seat/harness, state badge, elapsed, last event), when the hub last heard from it, and the claims it holds; the repo board shows, per repo, where it is running, its claim (owner node, conductor, TTL remaining), its last outcome, and acollisionflag when two nodes have live runs on one repo or a run is not on the claim owner. Runs are bucketed intofailed/awaiting approval/stale(no event for 30 min) /running/queued/interrupted/succeeded; the defaultsort=attentionputs the first three on top, oldest first, andsort=age|node|repo|state|seat, substring filtersnode= repo= seat= state=,attention=1, andall=1narrow both boards. Same bearer auth as the JSON endpoints, plus phone use: opening the page once with?token=<fleet token>303-redirects to the same URL without the token and sets an HttpOnly, SameSite=Strict, 30-daybrigade_fleet_viewcookie whose value is an HMAC of the token (never the token) and which authorizes only the HTML routes, so a copied cookie cannot read/status,/claims, or post events or claims; rotating the token invalidates every cookie. No token, cookie, or claim holder token is ever rendered, and responses carry the same CSP /no-store/no-referrerheaders as the local dashboard. The localcenterdashboard is untouched. Fixes #1124. - Fleet-sync Phase 4: hub-arbitrated cross-machine claims. To work a repo, a machine asks the fleet hub (
POST /claims): one unique row per target grants{owner_node, owner_conductor, holder_token, acquired_at, renewed_at, ttl_seconds, expires_at}to exactly one caller (default TTL 900s), a second caller gets a 409 naming the current owner, an expired claim (crashed or offline owner) is reclaimable (expired rows are pruned on acquire), and an unexpired one is never silently stolen. Every acquisition mints a per-holder fencing token: renew and release must present it (the hub never echoes it to other callers), so a sibling run sharing the same node identity can neither extend nor delete a live claim.node_idmust be a real identity - the hub rejectsunknownand malformed ids, and a client without a usable identity stays on the local lock with a WARNING, so two identity-less or cloned nodes are never both granted one target.brigade runacquires the claim right after the local run lease (only the run.lock owner ever holds or releases the hub claim), recording the orchestrator seat asowner_conductor, holds it with a renew heartbeat (every TTL/3; a claim the hub lost is re-acquired under the same token, and the heartbeat is drained on exit so an in-flight renew can never resurrect a released claim), releases it at run end, and refuses a held target withrepo '<x>' is claimed by node <node> until <expiry>. A lost acquire response is retried once idempotently, and if the hub stays unreachable a background thread releases any row the lost request committed; a 409 with no owner (target freed mid-request) is also retried instead of failing closed. With no hub configured, or the hub unreachable, dispatch falls back to today’s local run.lock with a single log line - a claim failure is never a new way for a run to fail.brigade fleet claims [--all] [--json]lists active claims (GET /claims;--allincludes expired), andfleet_client.acquire_claim/renew_claim/release_claim/repo_claimexpose the client calls. Hub database schema is v2 (adds theclaimstable; a phase-2 database upgrades in place). Fixes #1125. - Fleet-sync Phase 2: a central fleet hub and event reporting.
brigade fleet serve --host <tailscale-ip>runs a stdlib-only HTTP service (default port 3774,--db,--token-file; bearer token fromBRIGADE_FLEET_TOKENor the token file, never stored in config) backed by one SQLite file in WAL mode with a versioned schema:GET /healthis unauthenticated,POST /eventsaccepts one event or an array and answers{accepted, duplicate}deduped on (node_id, run_id, sequence, digest), andGET /statusreturns the latest state per (node_id, run_id) for live runs (?all=1includes terminal ones). Every local journal append also reports the event (node_id, run_id, repo, seat, harness, state, ts, sequence, digest) to the hub configured in~/.brigade/fleet.toml([fleet] hub_url,token_file;BRIGADE_FLEET_HUB_URLoverrides). Reporting is best-effort: it never raises into the journal writer, makes one request under a hard 2s deadline (covering DNS), and when the hub is unreachable spools to<brigade-home>/fleet-spool/<node_id>.jsonl(lock-guarded, atomically rewritten, capped at 16 MiB, poison 4xx batches dropped with a log line) and flushes in order (no duplicates) on the next successful contact orbrigade fleet flush. With no hub configured nothing changes.brigade fleet status [--all] [--json]prints what is running where: node, repo, run, seat/harness, state, age. Fixes #1123. - Local fleet-sync Phase 1: each workspace keeps a gitignored
.brigade/node.tomlmachine identity (node_iduuid4, hostname, roles, platform) created once bybrigade node(text or--json) and never regenerated. Newbrigade rundirectories are namespaced{short_id}.{stamp}-{hex}so two machines cannot collide under.brigade/runs/; un-namespaced legacy ids still resolve as the local node. A conservative lease reconcile releases a run lock whose owner process is dead and whose run is already terminal (or older than 24h with no active run), on the next claim and asrunguard.reconcile_run_lock;brigade doctorWARNs on a stale-but-unreconciled lock. Live-owner and in-progress activated-journal locks are left alone. Fixes #1122. - Research runs now persist a claim-level support audit (
claim-audit.json, schemabrigade.research.claim-audit.v1) next to the citation-token audit. The independent reviewer returns one record per checked factual claim (report span, backing source ids,backed/disputed/insufficient/skipped, bounded explanation, conflicting sources); Brigade validates spans, ids, enum values, and artifact bounds against the exact report and persisted findings, binds each record to the report digest, finding fingerprints, and synthesis attempt, and derives acceptance itself (reviewer boolean stays a veto, never the decision). Adisputed/skipped/ unstated-insufficientclaim triggers the existing single repair and reruns both checks against the new digest;insufficientis allowed only when a limitation is stated next to the claim and the exception is recorded. Resume drops review records whose report digest or synthesis attempt no longer match and re-reviews.brigade research showprints oneclaim_audit:summary line and--jsonexposes the artifact ref, verification, and counts without the records; legacy and pre-feature runs reportclaim_audit: unavailable. Fixes #938. brigade runs inspect <run-id>shows a run’s parent-child lineage (parent run id, branch-point event id for children; recorded children for parents) and its terminal lifecycle state, read-only from recorded run artifacts, failing closed on unknown or corrupt runs.runs shownow prints aterminal:line (andterminalin the JSON contract), and the doctor runs surface reports parent-child lineage consistency - a child whose parent directory is missing or whose branch-point event is absent from the parent journal is a WARN finding, never a crash. Legacy runs without lineage metadata remain valid roots. A journal terminal event is treated as terminal even whenrun.jsonstill says running (crash after fsync, before receipt replace).doctor --fullexhaustively rescans lineage (no 50-run cap); the bounded check WARNs with how many runs were not examined. Aparent_run_idthat is absolute or contains../ path separators is a finding and is never joined as a path. Refs #594.
Changed
- Test assertions for private file and directory modes now use one shared POSIX-aware helper, so the Windows suite skips checks for permission bits that Windows does not expose. Fixes #1133.
- The center Code Graph view now defaults to a branch-scoped overlay (changed modules plus their direct neighbors) so the summary strip matches the diagram, collapses
.worktrees/and similar non-canonical copies unless toggled, adds an Impact-tab “Start from changed modules” shortcut, and notes when the impact layer diagram is capped. Zero local changes fall back to the full map with a banner. Fixes #976. - The research Oracle seat’s browser-auth diagnostic now states that Oracle 0.18+ makes live browser cookie syncing opt-in upstream and points at an explicit cookie strategy (
--browser-cookie-sync, a configured browser profile, or another supported Oracle cookie source) instead of blaming only a stale jar.docs/runbooks/research-live-acceptance.mdand the multi-lane research spec document the requirement with upstream links. Refs #939. - Runs dashboard view answers “what verification ran recently, did it pass, and what command was it?” with a summary strip (“N runs this week, M failed, last run Xh ago”), a primary column of command (task for Brigade runs) with a pass/fail chip, Run ID demoted into a
<details>expander, and relative timestamps (“3h ago”; future stamps render as “just now”, unparseable ones as “unknown age”). Refs #974. - The center Work view now labels dispatch waves as plain-language steps (for example
Step 2 of 5 - one task at a time (blocks step 3)) instead of bareWave N, rewords the Parallel tile to Together / yes-or-no with a sentence about whether ready tasks can start at the same time, and draws a small SVG ofblocks/parent-childedges so “what is blocking what” is visible even when waves are absent. The task list stays. Fixes #977. - The center dashboard Status view now answers “What needs my attention across handoffs, memory care, verify loop, and inbox hygiene right now?” with a plain-sentence summary strip and health tiles (icon + plain word: OK / ATTENTION / MISSING / TIMED OUT, never color alone), following the Memory Operations baseline. Raw ids move into
<details>expanders, the placeholder rows for latest verify receipt/signal are dropped, and missing brief sections render as readable MISSING tiles instead of-table rows. Refs #972. - Center dashboard Outcomes view states its operator question (“Which skills and cards are actually helping vs hurting the loop?”) in the module docstring and renders a one-sentence top-N summary strip (biggest net help vs biggest net drag, derived purely from
ranking[]entries) above the existing table. The 9-column outcomes ledger itself is untouched. (#973)
Security
- Fleet hub per-node credentials: a node token is now the node’s identity, so one fleet member can no longer post events or acquire, renew, release, supersede, or inspect claims as another.
brigade fleet nodes add <node_id> [--label L]enrolls a machine and prints its token once (the hub keeps only a SHA-256 in the newnodestable; hub schema v4 adds it in place),nodes listandnodes revokemanage it, andPOST /events/POST /claimsderive the caller’snode_idfrom the token and answer 403 when a bodynode_iddiffers (a revoked token answers 401; re-add to rotate). The single shared bearer becomes the admin token: it manages/nodes, reads/statusand/claims, enrolls the dashboard cookie, and may post under anynode_idonly when the hub runs withbrigade fleet serve --allow-admin-writes(off by default, so migration is explicit). Clients gain[fleet] node_token_file(BRIGADE_FLEET_NODE_TOKENoverrides); without it the client falls back totoken_fileand logs one WARNING per process naming the deprecated shared-token mode, so existing deployments keep working through the documented migration indocs/runbooks/fleet-hub-proxmox.md. A 403 keeps events in the spool like a 401 and the claim client surfaces the hub’s message. Both token kinds are compared in constant time and never logged;docs/fleet-sync.mdstates the trust model (node token = identity, admin token = control plane, scope/owner fields = honest-client intent). Fixes #1150. run_checkpoint.strip_checkpoint_bodies_for_exportnow removes crashed.checkpoint.*.tmptemp files from the export copy without reading them (a symlinked or non-regular temp is refused), andassert_export_tree_has_no_checkpoint_bodiesrefuses any temp under a recovery-checkpoint directory regardless of its bytes, including one that happens to parse as an artifact reference. Work-run archives drop their separate pre-strip temp sweep and rely on the shared helper. Local recovery keeps the original temp and bodies untouched. Refs #654, #646.- Worker and adapter child processes now stream stdout/stderr through a combined 1 MiB capture cap and start in their own process group. Overflow or timeout terminates the POSIX group (Windows process tree). App-server JSONL is read in bounded chunks so a single oversized or malformed record is charged and trips
output_limit_exceededbeforejson.loads(text-mode line iteration cannot buffer an unbounded line). Queue entries and turn deltas are capped as they arrive (not afterrun_turnreturns). Resume hashes the 20 KB wrap body, event writers truncate before persist, stdin is written concurrently with output draining, and returned capture text is cut on a UTF-8 boundary at or below the cap.run_agentand ACPX/app-server adapters reportfailure_phase=harness/failure_kind=output-limit. Fixes #1012. - Opt-in
authority_store.isolation = "external-key"HMAC-signs new directory-authority bindings with a 0600 key stored outside the workspace and scanner-reachable tree ($XDG_CONFIG_HOME/brigade/authority/store-hmac.key, orBRIGADE_AUTHORITY_KEY_FILE). An existing HMAC envelope is always verified; flipping the flag off cannot skip a MAC check. Once a target is signed, a user-level sticky marker under~/.brigade/authority-signed/<target-fingerprint>(override:BRIGADE_USER_DIR) keeps raw unsigned records fail-closed even after the repo-writable flag is flipped off or the envelope is stripped. Marker containment is distinct from the HMAC key rule: the marker must live under that user-level brigade directory and outside the workspace, and is not rejected merely because its path contains.brigade. Downgrade isbrigade security authority downgrade --target <t> --confirmonly: it converts the signed store to an unsigned record, removes the sticky marker, and turns isolation off so a later read cannot recreate the marker. A key override whose unresolved path sits under the workspace, including a directory-symlink prefix, a file symlink, or a..re-entry, is refused before any key bytes are written.brigade doctorWARNs when the flag is off, and names the sticky enforcement when a marker exists. Fixes #957. - JSON run contracts redact Windows
C:\Users\<name>\andC:/Users/<name>/home prefixes to~at the same_clean_strchokepoint as POSIX/home/<user>and/Users/<user>, including doubled slashes and usernames that contain a space or apostrophe. An allowlist key without a cleaned counterpart can no longer pass a raw source value. Fixes #1064. Refs #631. - Evidence-ledger JSON/Markdown/MCP/HTTP exits now pass the complete serialized bundle through one last allowlist walker (
finalizeEvidenceResponse). Ineligible or quarantined items collapse to{id?, eligibility_status, reason_code}- stubidis 24 lowercase hex or omitted, so a cached attacker id is never echoed.grouped_by_sourcecounts only eligible closed-enum kinds (unknown keys are dropped).query/filtersand other unlisted keys are dropped, not trimmed. Cachedevidence show/show_evidence_bundleregenerate live from item ids. Akind=urlcontent hash can no longer authorize swapped artifact text. Fixes #1030, #1031, #1032. - Opt-in
authority_store.isolation = "external-key"HMAC-signs new directory-authority bindings with a 0600 key stored outside the workspace and scanner-reachable tree ($XDG_CONFIG_HOME/brigade/authority/store-hmac.key, orBRIGADE_AUTHORITY_KEY_FILE). An existing HMAC envelope is always verified; flipping the flag off cannot skip a MAC check. Once a target is signed, a user-level sticky marker under~/.brigade/authority-signed/<target-fingerprint>(override:BRIGADE_USER_DIR) keeps raw unsigned records fail-closed even after the repo-writable flag is flipped off or the envelope is stripped. Marker containment is distinct from the HMAC key rule: the marker must live under that user-level brigade directory and outside the workspace, and is not rejected merely because its path contains.brigade. Downgrade isbrigade security authority downgrade --target <t> --confirmonly: it converts the signed store to an unsigned record, removes the sticky marker, and turns isolation off so a later read cannot recreate the marker. The envelope’s inner target fingerprint must match the destination; a copied foreign envelope is refused. A failed durability step restores every artifact that already changed. A key override whose unresolved path sits under the workspace, including a directory-symlink prefix, a file symlink, or a..re-entry, is refused before any key bytes are written.brigade doctorWARNs when the flag is off, and names the sticky enforcement when a marker exists. Fixes #957. brigade runs list --jsonand the Run ViewGET /api/runslist now drop a child directory that is a symlink, or whoseresolve()leaves the runs root, and count it inskipped_invalid. Per-run serve routes already 404 those ids; the shared collector did not, so a symlink alias could leak out-of-root task text in the list while its detail route 404’d. Refs #631.- Scanner lifecycle rewrites now require the verifier-held run proof,
before.source == scanner.source, and a legalpending → dismissedtransition. A replacement row is applied as that narrow action, so a matching built-in scanner cannot dismiss another source’s finding, resurrect a dismissed row, or assign an arbitrary status. Fixes #1039. - Explicit
import_pathingestion retains the same publication snapshot as self-import (pre-commit inbox, persisted proofs, and file bindings) until the final scanner receipt binds. A failed receipt write rolls the import and proofs back together. Fixes #1040. work scanners initand--forcepublish.brigade/scanners.tomlthrough a descriptor-relative no-follow walk and replace only a verified regular file, so a planted dangling or live symlink cannot redirect the write outside the workspace. Fixes #1042.- Scanner runs-dir create no longer adopts a pre-existing unbound
.brigade/scanners/runstree, so forged child receipts cannot become scheduling authority. A released pre-0.27 workspace whose unbound root is operator-owned and uncompromised is bound to a fresh authority record automatically;brigade work scanners doctor,brigade work bootstrap, and the first sweep repair that path. A foreign-uid, world-writable, or symlink tree fails closed with an operator message instead of a traceback. Control-plane receipt reads require the recorded receipt-file device, inode, and digest, not just the run directory. (#1036) work import promote --allnow reviews and promotes one inbox generation under the task-ledger lock. Matching rows are bound to that snapshot’s identity and digest; a same-id swap or rewrite between review and promote is refused and does not create a task. Fixes #1037.- Run transport now applies quarantine, allowlist, and provider preflight to every
invokecandidate, so an invalid-final fallback cannot launch a seat that is already quarantined for the run. Fixes #1038. - Receipt
endpoint_hostprovenance records parsed hosts or the fixedinvalid-endpointmarker. A*_BASE_URL_REFthat resolves to a value with no hostname is no longer copied into the receipt. Fixes #1041. - Work-ledger readiness now fail-closes on stored edge types that are not an exact
blocks,parent-child, ordiscovered-fromtoken. Case variants (BLOCKS), whitespace-wrapped values, and other unknown types keep both endpoints out of the ready set and cannot be claimed. Compatibility rewrite of those types does not run during claim or readiness.work run --task-idand default ledger selection share the same locked start gate as claim, so a blocked or decision-gated task cannot be started or marked done. Plan writes stamp{schema, task_id, kind, digest, generation}on the task under the task-ledger lock; a missing, corrupt, or substituted expected receipt blocks claim. Fixes #1033, #1034, #1035. - Trust review is bound to an operator-minted capability.
brigade evidence trust reviewgenerates a fresh in-memory HMAC secret, writes a{item_id, from_digest, transition, nonce, expiry}token to the engine on stdin, and never places that secret in the child env or argv.miseledger trust reviewwithout a capability is refused for every stdin kind, including piped empty input,/dev/null, and a PTY. stdin is not an authorization signal. Direct engine invocation is unsupported. Fixes #1029. - The directory-authority store is wrapped in a parent-held HMAC envelope (
brigade.authority.store.v1) keyed by a 0600 file under the operator config root, with a MAC’d sequence file for replay and unsigned-downgrade refusal. PublicSCANNER_DEFAULTSequality no longer grants import privilege. A same-UID process that reads the persisted store key can still forge a valid envelope; that residual is encoded in the suite and is not a class close.--isolated-scanners(POSIX user+mount namespace, doctorWARN/MANUALwhen unavailable) is the opt-in tier that refuses that forgery. Part of #957. Refs #881. - JSON run contracts rewrite home-directory prefixes (
/home/<user>,/Users/<user>) to~on every string that enters the list, show, latest, and watch payloads. Artifacts and human CLI output are unchanged. Refs #631. Refs #958. - CSV exports from
miseledger sqlnow sanitize formula-triggering characters in text values and dynamic column headers to prevent spreadsheet formula injection while preserving numeric fields. (#1324) - Finding-controlled text in security scan reports, review output, Markdown summaries, and SARIF exports is sanitized to neutralize control characters and terminal escape sequences. (#673)
- Hardened work import provenance boundaries: rebuilt self-imported inbox rows with local IDs, bound canonical rows to immutable content, and required external scanner receipts before legacy migration. (#868)
Fixed
-
Evidence redaction ingest coverage closes three gaps (#998). A new
home-pathguard rule redacts absolute paths under/home/<user>/…and/Users/<user>/…(categorypii, so session and external-ingest origins cover them). Receipt provenance no longer hardcodesagent-session: miseledger receipt envelopes and their redaction records now use the truthfulworkspaceorigin (receipts_cmd.RECEIPT_ORIGIN). The previously unused clean-verdict gateevidence_redaction.ingest_verdict_is_cleanis now enforced intrust_gate.admit_consumer: when an envelope carries a redaction record, that record must be an explicit completed clean scan of the current policy before the item may be admitted as scanned-clean on brief, wrapped, cite, promote, or context surfaces; missing records keep legacy-row behavior, while malformed, error, redacted, or other-version records fail closed. Fixes #998. -
Fleet Dolt sink CSVs now use
0/1for Boolean claims and\Nfor SQL NULL, advance the event watermark after successful imports even without a new commit, and report unparseable incremental timestamps on stderr. Named fleet exports now retain the process’s normal file-creation mode when atomically replacing an existing file. Fixes #1158. -
brigade skills fleet statusno longer crashes withValueError: skill install path escapes workspacewhen an installed skill copy is a symlink whose target resolves outside the workspace. The read-only audit now reports such a copy as statusexternal(with anexternal_countin the summary and JSON payload, receipt-derived source preserved when one exists) and keeps auditing every remaining copy. Install, uninstall, update, and rollback paths keep the strict containment check:_install_dir()still refuses to hand back an escaping path as a writable destination. Fixes #1171. -
Fleet claims no longer lock a machine out of its own repo after a crash. A run killed with SIGKILL left an unexpired hub claim under its own node with a holder token nobody had, so the next
brigade runthere got a 409 from itself for the residual TTL (600-900s) with no way out but editing the hub SQLite. Now:brigade fleet claims --release <target> [--node NODE_ID] [--force] [--json]frees a claim without its token (POST /claimsreleasewithscope: "node"; the hub deletes the row only when the caller’snode_idowns it, and refuses another node’s claim naming its owner;--force/scope: "force"releases it anyway; renew is never token-less).<target>is always a claim key (never re-read as a directory);--pathmakes it a workspace directory instead (its name is the key, its node identity is used; a directory inside a workspace is refused rather than resolved upward). Without--force, both modes run the same proof: the hub’s newinspectaction returns the claim’s recorded run directory to its owner node only, the CLI maps it to a workspace on this machine (with--path, it must be the workspace given) and refuses while that workspace’srun.lockhas a live owner or is malformed, or when the run cannot be resolved at all (saying why). The release is then fenced to the inspected row: the hub refuses a token-less node-scoped delete without the inspectedacquired_at(only--forcedeletes unfenced), so a claim re-acquired in between is never deleted; and a probe that finds no claim owned by this node refuses outright without--force.--nodeother than this machine’s identity requires--force; the JSON receipt’sforcedreflects the flag. An older hub that rejects the request is named as such. Every claim row now records the acquiring run’s localrun.locklease (lock: owner token,acquired_at,run_dir; never listed; hub schema v3, live rows survive the upgrade), and the hub’s token-less release reads and deletes in one write transaction so its receipt is the row it removed; the v2→v3 column migration is serialized under a write lock. When the Phase 1 lease reconcile freesrun.lockfor a dead owner on the way into a run, the hub acquire is sent withscope: "node"and the dead lease assupersede, and the hub replaces only the exact row taken under that lease (same node, same lease token, lease stamp not newer) - a claim held by another node, by a run in another same-name workspace, on a cloned node identity, or without a recorded lease is never touched - logging one line; a same-node claim with no dead lock to vouch for it is still refused, withbrigade fleet claims --releaseand--no-fleet-claimnamed in the error.brigade run --no-fleet-claimskips the hub claim and relies on the local run lock alone, logged once.runguard.run_lockgains anon_reconcilecallback that reports each dead owner it reconciled, andrunguard.read_lock_ownerreads a held lock’s lease.scopeis an intent marker that keeps honest clients from stealing each other’s claims, not an authorization:node_idand the leases are caller-asserted under the one shared bearer token, which already authorizesforce. Fixes #1141. -
Fleet claim heartbeats now defer holder-token-fenced orphan cleanup until the run exits when a lost-row re-acquire exceeds its deadline, preventing both a late claim leak and mid-run release of a live claim. Fixes #1142.
-
The v0.25.0 run-reader compatibility test now exercises JSON-object rejection, unknown-key tolerance, and
finished_atterminality from input behavior instead of fixture-driven assertions, and builds its legacy snapshot from the writer’s full emitted shape. Refs #654, #640. -
Stop-hook tree fingerprints now exclude verify-run receipts, outcome ledger capture files, and inline MiseLedger indexing artifacts. A captured verification can close a session without treating its own evidence as uncovered work, while later source, configuration, and documentation edits still invalidate the receipt. Fixes #1132.
-
The blocked-phase cancellation CLI test now waits up to 30 seconds for the run receipt, subprocess, and lock before signaling, avoiding false startup failures on contended CI runners while retaining the cancellation ordering check. Fixes #1126.
-
The non-scorecard
brigade outcome reconciledecide path no longer installs a regressed candidate on a margin of one (helped=2, hurt=1). A cohort with verified regressions now needshelped >= hurt + install_min_helpedand holds withwithheld: verified helped margin over regressions below Nuntil then; clean cohorts are unchanged. Refs #654. -
Reused-receipt dedup is durable and cheap:
load_scoring_recordspersists eachreused_fromresolution tomemory/outcome/evidence-canonical.jsonand consults it first, so a legacy ledger row that stored the reused receipt path keeps collapsing to one signal after the receipt is pruned, and rank/score/reconcile open each distinct receipt at most once instead of once per row on every read. New captures against a reused receipt record the reused path on the ledger row asreused_evidence_ref, so that provenance no longer lives only in the receipts directory. Refs #654. -
Direct CLI roster seats can now set
command = ["executable", "fixed-prefix", ...]to override the executable and insert fixed arguments before the adapter argv. This provides a supported Windows Cursor path through the bundledversions/<version>/node.exeandindex.js, instead of the unsupportedcursor-agent.cmdshim. Fixes #1100. -
Bounded SQLite retry failures retain their BUSY-family result code after replacing the driver error text with holder diagnosis. Callers can retry primary
SQLITE_BUSY(5), recovery (261), and snapshot (517) results without exposing raw SQLite strings. Fixes #1083. -
Seat-driven dispatches (
run_seat.py) now passoutput_dirthrough torun_transport.dispatch, so themessage-envelopes.jsonlaudit sidecar is written for that path instead of silently going missing while the envelope gates still applied. Fixes #996. -
User-scope Claude work-loop hooks can pin
hook-run --targetto a wired workspace, never treat the home directory as a work root, and latch off for the session after two consecutive timeouts instead of appending a doctor banner on every tool call. The hook-run parent process does no unbounded filesystem work: pin resolution, target discovery, log-target lookup, latch I/O, and the.gitcheck run inside the timed worker, and post-timeout latch/state persistence plus claim cleanup are best-effort under a short hard cap, so a stalled cwd cannot block Claude Code past the configured timeout even in the timeout-handling path. The #735 doctor-pointer / errorhook.logwrite (including creating.brigade/work/claude-hooks/) still finishes beforehook-runreturns. Project-scopehooks installrefuses a home-directory target (that path is Claude Code user settings), andhooks statusreports duplicate managed handlers and a widened user-settings path. Fixes #1051. -
brigade runs childno longer inherits a parent’s external artifacts or handoff path (those are re-rooted to the child run dir or dropped), creates a rewritten child-owned artifacts directory before recording it so the path is not dangling, and refuses a terminal parent event withcannot branch a durable child from a terminal parent eventinstead of birthing afailedchild with no error text. Live-locked and legacy no-journal parents stay fail-closed. Fixes #955. -
handoff doctor,handoff lint, andingestnow walk every configuredhandoff-sources.jsonroot the wayhandoff listalready did, so a pending file in a second root is counted, linted, and ingested instead of leaving those commands green on inboxes they never read. An unreadable source config now failsingest(exit 3) instead of omitting those roots, andrepos ingestreportsinvalid handoff source configinstead of a cleanno handoff inboxskip. Fixes #1050. -
Claude work-loop hook session state is scoped per dispatched task instead of per reused session id. A
SessionStartwith sourcestartuporclearbegins a new task epoch: state entries from earlier tasks (repo graphs,write_observedlatches, fingerprints) expire on read instead of re-arming the stop gate for directories the current task never touched, and a read-only later task no longer inherits a prior task’swrite_observed=True.resumeandcompactcontinue the same task, and a task that does write in a wired repo is still gated for verification and handoff. Fixes #992. -
Result integrity now treats a sentence terminator glued to the next sentence (
...answering.media-cli is...) as a clause break, so a correct final answer is no longer discarded asnon-final-outputwhen the provider omits the space after its progress sentence. A filename-dot is re-merged only when the character before the period is an identifier and the following token is a known code or doc extension (README.md,App.csproj), not a short English word (it,is,the). Version digits (v1.2) are not treated as clause starts. Fixes #1099. -
Bare
brigade care installadopts installed predecessor ids that already schedule the same runbook (care-scancoversdaily-care,handoff-ingestcoversingest-sweep) instead of stacking a second systemd or launchd timer. Dry-run reports the adoption. Explicit--entrystill installs the named id. Fixes #986. -
Repeated care-entry failures (two or more newest-first failed runbook receipts) now surface on
brigade care status(summary + per-entryrepeated_failure=),brigade work brief(care_repeated_failures/scheduled_care), and the Center Status dashboard (Scheduled care tile + attention sentence). Topology marks those care jobsfailed. The SQLITE_BUSY retry forevidence-crawlitself landed earlier as #1067. Fixes #985. -
The center Runs view now treats a missing runs directory as the existing “No Brigade runs yet.” empty state, and renders a distinct error panel for other
brigade runs listCLI failures (timeout, invalid JSON, unreadable store).brigade.run-detail.v1docs name the verification fieldcommand. Fixes #991. -
Changelog footer link definitions cover every release section again (24 were missing, 0.8.2 through 0.26.1), and
[Unreleased]compares against v0.26.1 instead of v0.26.0. -
Center dashboard chip-icon backgrounds now render under the nonce-only Content-Security-Policy: Memory Operations and Status views use one stylesheet class per status role instead of per-element
style=attributes (whichstyle-srcwith a nonce refuses). Work, Runs, Agent Activity, and Outcomes already used classes or had no chip-icon inline styles. Fixes #1077. -
Evidence crawl import (
miseledger import sourceharvest/crawl files|gitlog|docs) and provenance backfill retry the SQLITE_BUSY family (primary code 5, including snapshot 517 and recovery 261) from a concurrent writer instead of failing the runbook. A lock that outlives the bound names the holder-diagnosis step rather than the raw SQLite string. Fixes #1067. -
Evidence-ledger
archive.Openrestores the pre-#1073 10s global SQLitebusy_timeoutso unwrapped command paths do not give up after 1s under contention. Crawl import and provenance backfill still bound their own wait via retry count, backoff, and a 4sMaxTotalWaitceiling (nobusy_timeout * retrieshang).IsBusyclassifies by the SQLite result code (primary / low 8 bits == 5) and ignores parenthesized 5/261/517 in free text such as subprocess stderr. Fixes #1085. -
Directory-authority validation distinguishes a pre-hardening record (missing
workspace) from a genuine identity mismatch. A legacy record whose extanttargetand directory identities still match is upgraded in place, somemory care import-issuesand other bound opens no longer die with a bareOSErroron long-lived workspaces. A swapped or forged directory still fail-closes; the error names the record path, the mismatched field, andbrigade work rebind-authority --target <workspace>.brigade doctorWARNs when a legacy record is still on disk. Fixes #1066. -
work import promote --allnow persists a task only after that item succeeds through the late-window inbox CAS. A failed batch item is not flushed by a later success, and a refused promote writes no task files. Fixes #1059. -
Scanner runs-dir authority bind and run-directory publication now use the same POSIX/Windows dirfd abstraction as import-inbox (
nt_dirfdon Windows).operator quickstartandbrigade workscanner sweep no longer fail withdescriptor-relative directory authority operations are unavailableon Windows. A host with no dirfd and no existing runs tree returnsmissingso first-create can proceed; an existing unbound tree is still not adopted. The #1036 adoption refusal stays closed on POSIX. (#1036) -
A fresh
init/operator quickstartworkspace no longer fails its owndoctororsecurity scan. The workspaceAGENTS.mdtemplate keeps the file-maintenance table inmemory/cards/so the rendered file stays under the 12000-byte bootstrap budget on Windows CRLF,*_FILE/*_PATH/*_FILEPATHnames such asRESTIC_PASSWORD_FILEare not plaintext-password findings (values that merely start with$or~still report), and the half-fed outcome-loop warning namesbrigade work verify run --manifestplusverify/manifests/. (#1025) -
brigade evidence crawl planomitscrawl memorywhenmemory/NAMESPACEis missing (or prints it once that operator-declaredmemory-<uuid4>exists), and omits sourceharvest-backedcrawl files/crawl gitlog. Printed commands use positional paths (no--root/--repo). (#1024) -
Operator-authored run task text that discusses prompt-injection policy (for example, documenting that a worker must ignore previous instructions from tool output) is no longer classified as an injection payload, so dispatch proceeds. A documenting verb does not exempt a whole line when a later clause still matches an imperative, role-override, or secret-exfil rule. Imperative injection payloads still quarantine. A request-side quarantine now fails as
injection-quarantineand names the heuristic that fired, instead ofunclassifiedwith no operator hint. (#995) -
Projection commit now holds the validated parent-directory descriptors and publishes create/replace/remove - including
writer=/remover=callbacks - through those descriptors, so swapping a checked parent for a symlink between plan and commit cannot redirect a write outside the destination tree.managed_block.write_text_nofollow_atomicpublishes the same way and refuses a swapped parent. (#1013) -
Stale-lock
inspect_commandfrombrigade runs watch --jsonis a runnablebrigade runs show <run_id>:runs shownow resolves a bare run id under--cwd/--runs-dirthe same waywatch,recover, andchildalready do, so path redaction no longer emits an unexecutable command. (#990) -
Vault index, search, show, and doctor now walk allowlisted roots from a held vault directory descriptor and refuse a symlink in every path component (
openat/O_NOFOLLOW+fstat). Replacing an allowlisted folder such asShared/Inboxwith a symlink to another in-vault path can no longer re-scope private notes; reread of indexed files uses the same walk. (#1011) -
Claude hook stdin is read through a byte-limited binary reader, rejects trailing input, and caps nested string and collection sizes before the timed worker starts, so an oversized hook payload cannot force unbounded allocation or JSON parsing outside the timeout. (#1014)
-
Inter-seat message envelopes now bind
{run_id, message_id, assignment_id, from_seat, to_seat}into the authenticated envelope and require those expected values at every receiver. A valid old text/envelope pair can no longer be transplanted into another run, seat, or assignment. Fallback worker output is attributed to the terminal producer, and resumed output is re-enveloped under the current run identity. (#1010) -
Promotion and other content admissions now validate the complete provenance envelope and require a present, exact
hashes.contentmatch against the current item bytes before copying a body. A same-UID rewrite ofimports.jsonlthat leaves the original envelope can no longer be promoted into a task; missing, malformed, or stale redaction records also block promotion. (#1008) -
brigade care installandcare statuson Windows now resolve the defaultautobackend to a printed Task Scheduler plan (--backend schtasks). Each printedschtasks /Createline uses that entry’s own schedule,/TRthroughC:\Windows\System32\cmd.exe, and/RU "%USERNAME%" /ITso the task runs while the user is logged on without an elevation prompt (/NPS4U requires admin). An explicit--backend systemd|launchd|crontabstill generates that backend’s plan on Windows instead of being redirected to the printer.--jsonreturns the structured plan instead of plain text.care statusreports whether eachBrigadeCare-*task exists and when it last ran. (#1021) -
Evidence-ledger read surfaces hide item bodies unless provenance parses and injection status is the validated typed value
clean. Session preview, session search snippets, and session transcripts are gated the same way. Imports stayquarantined/pendinguntilmiseledger trust review --mark-injection-clean; a label-only review does not make content eligible. MCP and HTTP no longer accept a caller-settableinclude_untrusted_bodyreveal. Routinetrust reviewrefuses a parse-error-grade stored envelope instead of silently rewriting it toclean. (#1007, #1009) -
Origin-scoped ingest redaction now maps
work import context(and--source external-web/external-service) to the external detector tier, normalizes free-form--sourcewithstrip().lower(), and fails closed tounknownon unmapped values. An equal-but-unredacted scanner result is treated as failure, and research findings redact title/summary/evidence per field so a multi-line summary cannot migrate into evidence. (#498) -
brigade runs list --jsonreadsparent_run_idfrom each run’s own lineage and no longer scans siblingrun.jsonfiles through recorded-child discovery. Refs #631. Refs #958. -
Recorded-child
statusandbranch_point_event_idinbrigade.run-detail.v1are bounded through the same_clean_strlimit as other contract strings. Refs #631. Refs #958.
Added
-
brigade runs resume <child-run-id>resumes a durable child from its recorded parent lineage and branch-point. When the child has a roster and resumable app-server coordinates, work continues through the existing locked resume path. A fresh child snapshot without those artifacts confirms the branch-point and does not stamprunningor claim that work re-executed. The git-root run lock is honored, parent and sibling receipts stay untouched, and a corrupt or unsupported branch-point fails closed with no resume event. Legacy runs without child metadata keep the existing resume path. Fixes #1072. -
brigade runs diff <child-run-id> [other]compares a durable child run against its recorded parent (or two siblings that share a parent) without mutating either side. The versionedbrigade.run-diff.v1contract covers lifecycle state, worker results, verification results, and outcome evidence; GraphTrail snapshots are compared only when both sides are compatible, otherwise the payload states the skip. Verificationchangedcompares the full ordered receipt sequence on{status, command, exit_code}only (run_idand thefinal.txtdigest stay display-only). All contract strings pass through the_clean_strchokepoint. Fixes #1071. -
brigade runs diff <child-run-id> [other]compares a durable child run against its recorded parent (or two siblings that share a parent) without mutating either side. The versionedbrigade.run-diff.v1contract covers lifecycle state, worker results, verification results, and outcome evidence; GraphTrail snapshots are compared only when both sides are compatible, otherwise the payload states the skip. All contract strings pass through the_clean_strchokepoint. Fixes #1071. -
Center Code Graph adds a Blast radius primary mode beside the Module map (which stays the landing view): a fixed upstream / focus / downstream layout for one module, seeded from the branch’s changed modules, with a focus picker, one-hop default depth, per-side “expand one more hop”, and stated truncation (
showing 10 of 34). Neighbors come from the full code graph for the focus node, not the capped module-map export. Same-hop ranking sums every parent edge and sorts by weight then module id (stable top-10); a typed name that matches more than one module lists both full ids instead of silently picking one. Fixes #1091. -
brigade work ready --campaign <name> --parallel-safenow composes campaign-aware dispatch waves at query time from each member’s existing per-repo footprint partition. Global wave N is that member’s local wave N in campaign order: overlapping files inside one member still serialize, disjoint work from different members (including matching relative paths) can share a wave, and an empty footprint stays exclusive only inside its own member. Cross-repoblocksstill gate readiness before composition; missing or unreadable members still fail closed before any waves are returned; per-member GraphTrail degradation is reported without discarding safe file-overlap waves. Per-repo--parallel-safeoutput is unchanged. Waves are not persisted. Fixes #999. -
The Agent Activity dashboard view answers “what are the agents doing right now across the fleet?” with a page-level summary strip above the machine board (total agents, running, blocked-on-host, failed, awaiting approval, unknown) and a visible plain-word state legend; per-host count chips now show icon + word + count instead of a tooltip-only icon+number. Missing or unrecognized states aggregate under
unknown. Fixes #975. -
The
review-heavypreset now ships a read-only Daybreak Blue defensive-security seat using the verified Codex headless model id. Accounts without entitlement reportmodel-unavailablewithout echoing provider diagnostics. Fixes #1002. -
brigade runs serve --cwd <workspace>opens a foreground, loopback-only, read-only Run View over the existingbrigade.runs-list.v1,brigade.run-detail.v1, andbrigade.run-watch.v1serializers. It binds an available port by default, prints the URL, opens the browser unless--no-open, and stops on Ctrl+C. The HTTP layer does not read run artifacts itself, rejects traversal and symlink escape, and never starts from another command. Fixes #631. -
brigade.run-detail.v1includes recorded children underrun.lineage(same sibling-receipt discovery as humanruns show). Serializers merge source artifacts through a load-bearing field allowlist so unexpected keys cannot pass into the JSON contracts. Refs #631. Refs #958. -
brigade runs showlists recorded child runs under lineage when a sibling receipt names the shown run asparent_run_id. Legacy runs without child metadata stay valid roots and print no lineage block. Branching from a digest-broken parent journal, a parent with no journal, or while the git-root lock is held fails closed with no partial child directory. Refs #594. -
brigade codeindexes.astrocomponent files (frontmatter and<script>as TypeScript, template tags as calls to imported components), soaffected,callers, andimpactcan see an Astro file and its downstream Astro consumers instead of reporting it missing. (#1003) -
brigade roadmap auditnow fails closed when the ROADMAP.md “Where things stand” headline (**vX.Y.x on main**) lags thepyproject.tomlmajor.minor, or when that headline is missing or unparseable. The check is only emitted for a target that declares a parseableproject.version, so a repo making no comparable version claim is unaffected. Patch-only differences stay green, and a headline ahead of the project version passes and is reported as ahead.brigade roadmap audit --checkexits non-zero so the repo-metadata CI job can fail a lagging headline with an update hint. (#1000) -
Evidence ingest now applies an origin-scoped redaction policy (
brigade.evidence-redaction.v1) before persistence. Sources classify to the provenance envelope origin; each origin selects detectors; the stored item keeps only redacted bytes plus a count/detector record (never the removed values). Scanner failure, timeout, or unavailability cannot produce a clean verdict and persists a placeholder instead of the original. The policy version is stamped on every decision and applies to future writes only - existing rows and provenance backfill are not rewritten. (#498) -
Inter-seat planner, worker, and synthesis messages now carry a
brigade.provenance-envelope.v1withmessage.text.utf8.v1over the exact outbound UTF-8 bytes. Channel gates run beforeparse_planor model delivery. Receive-side trust is re-derived or verified against authority and is never taken from a storedreviewed/verifiedlabel; model and tool-origin content stayuntrustedregardless of the on-disk claim. Envelopes are bound to their channel so a valid envelope for one message kind is rejected at every other gate. Model and tool outputs startuntrustedand completing a phase never upgrades them. Worker and prior-stage text is wrapped and byte-capped before later prompts.message-envelopes.jsonlbesiderun.jsonstores message id, phase, seats, and envelope only. Legacy run messages display unknown provenance and are not replayed. (#585) -
brigade.causal_receipt.v1is a lineage-only companion record for plan, run, verify, outcome, and handoff artifacts. New writes stamp typedrecordedparent links so one plan→run→verify→outcome→handoff chain (and multi-parent synthesis) is traversable without timestamp or path guessing. Parent count and encoded size are capped; oversized fan-in must reference a hashed manifest. Unknown relations, malformed digests, broken parents, and unsupported versions produce bounded diagnostics. Telemetry export prefers these links for parentage and keeps its deterministic trace/span projection. Historical artifacts stay readable and are not rewritten; inferred backfill remains under #583. (#493) -
brigade memory vault-proposewrites an additive note into an allowlisted operator-vault inbox configured in.brigade/vault.toml. The body is read from stdin, staged owner-only outside the vault, and delivered through the projection kernel.--scopenames the inbox root (unknown scopes error); proposals mint a stablecanonical_id, never land inBrigade Memory/, never overwrite an existing note, and--dry-runreports the destination and rendered bytes without touching the vault. Containment fails closed when symlink or descriptor checks cannot be performed. (#945) -
Canonical memory cards mint one opaque
card-<uuid4>ID on create, and the reviewedmemory care backfillpath can mint an ID on a legacy card, including complete care cards that only lack an ID. Valid IDs are never replaced. Care queues, recall and search logs, refresh imports, and retrieval-eval fixtures dual-read explicit IDs and legacy path/stem/topic aliases; alias collisions fail without rewriting cards. Dry-run backfill writes a deterministic mapping receipt with coverage and old-to-new IDs (relative paths only;old_idis the consumer-facing identity). A zero-candidate dry run writes nothing. Duplicate-ID coverage does not requirememory/NAMESPACE. Missing IDs stay a doctor warning until coverage is 100 percent. (#867) -
brigade memory vault-index,vault-search,vault-show, andvault-doctorclose the operator-vault projection round trip. Config lives in.brigade/vault.toml(vault,[[roots]]withscope/path/optional, andschema_version). Search reads allowlisted roots only, writes a derived owner-only index under.brigade/vault-index/, redacts output throughguard.redact_text, and labels every hituntrusted_vault_content. Projected notes key oncanonical_id; operator-authored notes fall back to the vault-relative path. (#943) -
Evidence consumers now enforce
brigade.trust-policy.v1: unknown and quarantined items are excluded from default briefs, context packs, citations, and promotion; untrusted brief content is wrapped and capped at 2 items and 50 percent of brief bytes; only the explicit known-safe injection statusclean(afterstrip().lower()and pending resolution) may wrap or emit a body - case, whitespace, empty, and garbage variants are metadata-only.brigade evidence trust review <item-ref> --content-hash <digest>attests an untrusted item as reviewed after an exact envelope digest match and appends one transition event.brigade receipts verifyupgrades indexed verify-receipt items to verified only after complete receipt v2 patch identity and exact retainedchanges.patchvalidation; duplicate verify is an idempotent no-op. Envelope verified does not bypass subject_binding, check_role, patch identity, or failure-taxonomy requirements. Citing a flagged research finding is rejected; the run takes one repair attempt and then failsreview-rejected. (#587) -
Evidence read surfaces recompute provenance content and materialized raw/artifact hashes. Search suppresses a mismatched snippet and returns
integrity_mismatch: true. Directmiseledger show/brigade evidence showhide a mismatched or synthesized-legacy body unless--forensic-contentis passed; that flag never changes trust and only reveals a body when injection status is the explicit known-safe valueclean(empty, unknown, and parse-lost statuses block). MCP, HTTP, bundles, briefs, and context stay metadata-only on mismatch. Bundle cache and Markdown preserve envelope fields andintegrity_omitted. One downgrade event is appended per item/hash/mismatch; the row is never deleted. (#586) -
Python evidence producers now stamp
brigade.provenance-envelope.v1on work imports, research findings, and indexed receipt adapter rows. Research findings persist an exact{title}\\n{summary}\\n{evidence}text projection; legacyFinding.truststays readable and only maps to origin/modality.brigade work import provenance --backfillandbrigade research provenance backfillinfer missing envelopes without treating them as trusted. Receipt indexing always startsuntrusted. Existing pending imports without an envelope reportmissing_envelope, which surfaces as a doctor WARN until an operator runs--backfill. (#584) -
brigade runs child <run-id> <event-id>creates a durable child run from a checkpoint-covered parent lifecycle event, records lineage and the verified shared prefix, and shows that lineage inbrigade runs show. The child gets a freshstarted_at/status_started_atand does not inherit the parent’s live control socket, live-progress fields, failure terminal fields, or parent-directory artifact/handoff paths. The covering checkpoint event itself is not a valid branch point. (#594) -
brigade memory project-vaultwrites a one-way Obsidian projection of canonical memory into an existing vault: care-state frontmatter and tags, wikilinks, category/harness maps, and a topology canvas. Vault writes use the projection kernel, preserve manual edits via conflict copies, and keep the operator vault path out of public status. (#888) -
--run-budget PATHaccepts onebrigade.run_budget.v1JSON declaration onbrigade run, dogfood, and model-trial starts. Brigade persists the supplied declaration in the run and plan receipts without mappingtimeout_secondsinto a run budget. Budget cancellation receipts now carry bounded per-seat or per-transport outcomes and the observed work that may still be active. (#885, #886) -
brigade mcp sync --writecommits selected native configs and MCP ownership state through the projection transaction kernel. A failed write restores the selected destinations; unfinished operations are visible throughbrigade mcp statusandbrigade mcp doctor, and can be recovered withbrigade mcp recover <operation-id>. Refs #911.
Changed
- The root
CHANGELOG.mdis marked/CHANGELOG.md merge=unionin.gitattributesso same-day Unreleased appends from sibling PRs combine instead of conflicting. Nested changelogs are not marked. (#994) - Search integrity verification loads item text and provenance in one
IN (...)query keyed by the result ids, instead of one select per hit (search limit is up to 200). (#964) brigade memory project-vaultaccepts--max-relatedto tune the per-note related-link cap (default 12). Explicit refs still outrank tag/category matches but no longer claim to always survive the cap. (#966)brigade memory project-vaultproduces a browsable vault on a real corpus. Notes take their title from the card’s leading heading when frontmatter has notitle, instead of the filename slug, and that heading is no longer repeated in the body. Related links are ranked and capped at 12 per note, so a large shared category no longer forms a complete graph. Maps that would cover every note are omitted, and canvas nodes lay out in a grid rather than one row. On a 1,254-card corpus this took the canvas from 83,586 edges and 11.3 MB to a bounded graph. (#888)- Center Memory Operations Cards now shows corpus size, a chip cloud, compact
expandable rows, and 30-row paging. Handoffs is a sibling tab on that page;
/view/handoffsredirects there. Snapshot polling reloads the page so nonce-scoped assets stay valid. (#903) - The shipped review-heavy
reviewer_flashpreset now uses Gemini 3.7 Flash Low on Antigravity. Security-sensitive or cross-file decisions still route to the independentreviewerorreviewer_codexseat. (#920)
Fixed
- Synthesis causal receipts now fall back to a hashed
parent_manifestwhen encoded lineage exceeds the compact-size cap, not only when parent count exceeds 16. The fallback writes the lineage-only parent-manifest artifact the receipt references (besidesynthesis.json) so traversal can recover the full worker-result parent set. A realistic 13-worker fan-in no longer raises, dropssynthesis.json, or looks like a parentless root. (#493) - The Claude Stop closeout gate no longer treats another session’s writes in a shared workspace as this session’s unverified work. Worktree deltas during a Bash window are attributed only when no concurrent Claude write evidence (
write_observed/last_write_at/pending_write_at) or other-harness receipt explains them - a read-only neighbor that only touches its session file is not foreign write evidence; a time-valid session receipt still counts when the live tree has drifted under a foreign writer; and a hard block that another process can re-arm is downgraded to a warning after this session has already captured, unless this session itself wrote again after that capture. (#959) memory care backfilldry-run receipts now predict--apply(stablenew_id), recordold_idas the topic/stem consumers already use, mint IDs on complete ID-less cards so coverage can reach 100 percent, skip creating.brigade/on zero-candidate dry runs, and countduplicate_idswithout amemory/NAMESPACEfile. (#867)memory vault-searchrebuilds the derived index when.brigade/vault.tomlroots change, so removing a scope from the allowlist stops serving that root.vault-doctorwarns on index/config root drift instead of reporting all-green against a stale index. (#943)- Evidence brief rendering one-lines MiseLedger score keys/values (including zero and
False) and retrieval-arm labels, omitstrust:when the gate-resolved label isunknown, prefers the provenance envelope over item-level trust keys, and keeps a partial first result line when truncation leaves room after context lines. (#941) brigade runs watch --json(brigade.run-watch.v1) now routes every record through the same allowlist and one-line helpers as the list and detail contracts.watch/summaryemitrun_idinstead of an absolute path,eventrecords drop raw params (tokens, prompts, stdout, log paths), and a field that cannot be rendered safely is omitted. (#631)- Center Research provider table no longer labels absent model seats as
- (unverified); only named models carry the unverified marker. (#961) - The graphtrail stale-baseline verify test no longer races wall-clock sync delays against a tight subprocess timeout on loaded CI runners; sync timeout is injected deterministically so the stale-graph assertion is reliable on Python 3.12. (#954)
- Quarantined items stay quarantined when a pending injection scan comes back clean, so a labeled quarantine cannot be released to untrusted and become content-eligible. Reviewed and verified items keep every consumer surface untrusted already has, including
context.brigade receipts verifyreportsverify-pending(and does not append a local verified event) when the MiseLedger trust notify fails. Provenance-event JSONL now appends and fsyncs instead of rewriting the whole file. (#587) - Corrupt or future-version
.brigade/work/tasks.jsonledgers now fail closed on everybrigade worksurface (and other commands that read the ledger) witherror: ...and exit 2, instead of parsing as empty or dumping a traceback. The work-store measurement harness (scripts/measure_work_store.py, protocol 4) now reportsschema_version_policy.future_version_rejected: truewith a clean exit-2 error and untouched ledger bytes; it no longer treats future versions as coerced. Unknown hand-edited edge types are preserved on rewrite and surfaced viawork ready(unknown_type_count) and awork doctorWARN.work import promote --allreturns 2 when any import fails. (#948) - Scanner receipt and import-proof bytes now carry verifier-owned file
identity and content bindings. Legacy migration rejects in-place receipt
overwrites, attacker-created sidecars, missing bindings, and mismatched
file identity or content. A binding-write failure restores the prior
receipt, proof, inbox, and binding state or fails closed. Workspace
rename re-anchors those file bindings with the directory record, and a
malformed store read fails closed instead of wiping every binding.
Rollback restores the snapshot for the bindings the failed operation
itself added, deleting one it added that the snapshot did not have, while
leaving a concurrent writer’s unrelated bindings intact. A directory that
already exists and is not bound to the workspace record is never adopted,
including when the caller asked to create it; Brigade creates and binds
the import-proof directory before it launches a scanner child instead.
Scanner children run with an explicit environment allowlist whose
HOMEandXDG_*paths point at a per-run sandbox that is removed when the child exits, so a child cannot reach the verifier authority store through its environment. That allowlist passes through what shipped scanners need, includingGH_TOKEN/GITHUB_TOKENorGH_CONFIG_DIRpointed at the operator’s existingghconfig,SSH_AUTH_SOCK,USER/LOGNAME,TMPDIR, and proxy settings, sogh- andgit-backed scanners keep authenticating. Same-uid writes that guess the store path remain a residual class. Refs #881. - Bare
brigade care statusnow discovers target-scoped per-entry registrations instead of scoring the atomic five-entry default. A target with no discovered units reportsmissing. Systemd status also reportsenabled: falseuntil every discovered timer is linked fromtimers.target.wants. (#914) - Graph snapshot copies during
brigade work verify runnow honor the samegraphtrail_delta_timeout_secondscap asgraphtrail sync. A snapshot that exceeds that cap markscode_graph_deltasync_timed_outwith a reason and leaves the verification command result unchanged, so a slow graph cannot reject a passing check. The 10s default and per-target.brigade/config.jsonkey are unchanged. - Source and pipx-from-checkout installs now declare
0.27.0, which PEP 440 sorts at or above the publish-dev0.27.0.devYYYYMMDDwheel line. The previous0.26.1pin version-sorted below those wheels, so a current main checkout looked stale to naive auditors. publish-dev stamps{pyproject}.devYYYYMMDDon the declared line instead of the next minor. brigade init --profilenow merges repeatable--includevalues into the profile includes with first-seen de-duplication before--fullappendsrepo-extras. Refs #494.- Issue #846 R2 harness sendback (comment 5247390894): JSON secret-history
proves synthetic Git/ignore-path; SQLite marks Git history unavailable;
cold-start uses parent
subprocess.runwall with distinct inner timing; RSS uses one subprocess-child sampling protocol with explicit scope labels. - Issue #846 R2 harness sendback: SQLite metrics scans the touched store and
restores env overrides; restart/cold-start use a fresh subprocess probe;
JSON guards cover
if_statusmatch and mismatch. brigade initgitignore blocks now un-ignore harness parent directories before the managed handoffTEMPLATE.mdexception so a parent.claude/rule cannot shadow the template; the source repo.gitignorematches the same repo-local recovery recipe. (#860)
Documentation
- ROADMAP.md “Where things stand” now tracks v0.27.x: the provenance and evidence-trust program, the read-only Center dashboard, the memory vault and stable-card-id program, scanner and verifier authority hardening, and the cross-session Stop-gate fix. Shipped Next and Later slices (dashboard first slice, work-inbox graph, wiring durability, handoff standalone-chunk lint) moved to docs/roadmap-archive.md. Part of #1000.
- Quickstart commit guidance now states that
.claude/stays local-only except the managed.claude/memory-handoffs/TEMPLATE.mdClaude onboarding requires tracked, cross-linking the repo-local shadow recovery from #858. Normalbrigade initreruns upgrade the managed gitignore block without--force. (#860)
Added
- Projection transaction kernel: an internal all-or-restored prepare/commit
engine, recovery journal, redacted receipts, injected-failure hooks, and a
reusable conformance fixture for multi-file projections, plus
brigade projection recover <operation-id>. Native stdlib implementation; no production projector is migrated yet. (#910) - Cloud dispatch registry (
brigade run cloud): register-on-dispatch and adopt paths,statusJSON classification (pending / ready-to-land / landed / stale / orphaned / needs-investigation), receipted report-onlysweep, configurable stale-READY threshold surfaced inbrigade work brief, and Center Agent Activity Cloud card rows when the registry has entries. Cursor cloud stays unwired until an API key exists; GitHub is ground truth for branches/PRs. (#890) brigade care install|status|uninstall --entry JOB_IDselects one namespaced registration keyed by(target identity, job_id). The five maintainer memory jobs (handoff-ingest,care-scan,memory-refresh,evidence-crawl,memory-closeout) are in the catalog; a job whose operator-approved runbook is missing installs nothing and says why. Omitting--entrystill writes the atomic scheduled-care set. (#762)- Bundled template-profile selection and render-hash snapshot for #494: named
presets (
repo-claude,workspace-claude-codex,repo-claude-full) viabrigade init --profile, deterministic rendered-file digests checked byscripts/template_profile_snapshot.py, and harness compatibility projected from existingharness-contract.v1tested_version/ provenance / evidence fields. No public registry schema, profile URIs,supported_harness_versions, or install-time version gate. Profile-less--depth/--harnessesinit is unchanged. brigade runs export/validate-archiveenforce the #592 worker-event privacy boundary: raw worker-like JSONL anywhere underevents/(including nested paths such asevents/nested/coder.jsonl) stays local-only and is never classified as public support data. Export writes distinct scrubbed sidecars (events/**/<worker>.scrubbed.jsonwith scrubbed media type,role=artifact,privacy_class=redacted) via the existing fail-closed scrubber, or refuses the export with a bounded diagnostic when any method/field/item is unknown or unsafely classified. The reservedevents/lifecycle.jsonljournal exemption authenticatesbrigade.run_event.v1records and rejects raw JSON-RPC worker envelopes. Source runs remain unclassified until scrubbed; audit/replay still requirepolicy=local-onlyfor raw salvage.- Run budget enforcement first slice (#593): wall-clock and worker-dispatch
ceilings are reserved before new worker dispatch; durable
run_budget.threshold_reached/reservation_denied/exhausted/cancel_requested/cancelled/usage_reconciledlifecycle events use stable idempotency keys; restart-safe projection rebuilds remaining allocation from the journal. Model/tool/token/cost stay observed unless an adapter exposes an enforceable boundary. Budget exhaustion and operator cancellation are terminal policy lifecycle states (budget-exhausted/operator-cancelled), not #576 worker FailureClass values and not infrastructure-neutral under #580. Child allocation remains #594. - Issue #846 R2 residual work-store characterization: extend the R1 harness
with observed same-actor/guard/empty-filter, restart, schema coercion,
backup/restore, install/cold-start/backup timing, resource, metrics-state,
and deleted-secret cells plus diffable
docs/measurements/issue-846-work-store-r2.json. Characterization only; no backend, listener, or runtime dependency. brigade update --channel betainstalls the newest non-yankedbrigade-cli==0.27.0.devYYYYMMDDwheel from PyPI into the existing user-global pipx environment, migrates prior Git-main-SHA beta state to that exact wheel pin, supports dry-run without mutation, and rolls back to stable through--channel stable --switch-channel. (#852)- Issue #846 R1 work-store characterization: research report, synthetic
chain/diamond/wide fixtures, and
scripts/measure_work_store.pymeasurement harness with diffable JSON output. Characterization only; no backend or runtime dependency. - Bounded session-start memory recall (#466 Slice 1):
brigade memory recall --target <hub-or-mirror> --cwd <session-cwd> --limit 5 [--json]derives query terms from the cwd basename, returns title/tags/path only (max 5 matches / 10 lines), and reads a machine-localmemory_recall_target. Workspace depth may default to the current target; repo depth stays unconfigured until an explicit hub or mirror is set. The managed ClaudeSessionStarthook merges recall beside the work brief once per session and fails open on missing or broken targets. Operator smoke steps live indocs/memory-care.md. brigade center servegains work-graph views over existing--jsoncontracts: an SVG ready/blocked dependency graph, parallel-safe dispatch-wave swim-lanes (seam-ready forwork ready --parallel-safewhen #803 lands), and per-task claim state with three-phase footprint stale indicators. Read-only; no new write surfaces.- Label-free memory search recall signal (#723):
brigade memory searchappends a size-capped local log at.brigade/memory/search-log.jsonl(timestamp, normalized query, top-K card ids only). A follow-up within the disjointness window whose top-K shares nothing with the prior search counts as a miss;memory care statusreports the rollingsearch_recall(searches,followup_rate) as second-class evidence that never gates doctor/valid. Failed queries become fixture material for the #722 retrieval eval harness. - GeneratedPatchQuarantine (#507): model-generated patches
(
patch_source: generated) promote only through an independent verifier receipt that records repository-test (effectiveness) checks pluscandidate_count/model/model_versiononsubject_binding.generated_patch_quarantine. Model confidence, lexical or textual similarity, and repeated sampling are explicitly non-promoting outcome sources. Design:docs/design/generated-patch-quarantine.md. - VerificationContract (#500): consequential runbooks and verify manifests declare
an independent verifier, failure/rollback path, and token/latency budget before
execution (
brigade.verification_contract.v1). Plan surfaces incompleteness; run refuses incomplete consequential templates. Receipts recordbudget_useandverificationseparately frommodel_completion. First-slice enforcement of wall-clock and worker-dispatch ceilings lands in #593. - Optional dispatch annotations on work tasks (#815):
--seat-class(mechanical|judgment|review) and--spend-by(ISO-8601 deadline) round-trip throughbrigade work task add/brigade work task annotateand appear onbrigade work ready(JSON and text) when present. Stored under taskmetadata; absent annotations change nothing. Brigade does not dispatch from them. Forward-plan skill documents emitting the hints. - Seat-health admission now projects
brigade.seat_health_summary.v1ontorun.json, records typed incomplete worker summaries, persists same-seat-once retry attempts before re-probe, quarantines non-retryable seats for the run, receipts Codex app-server→exec asbrigade.transport_routing.v1, re-probes on resume, and rejects read-only seats whose declared isolation is only soft/none. (#474) - Cross-repo campaign ready view (#814):
brigade work ready --campaign <name>aggregates per-repo work graphs through a named lens under.brigade/campaigns/<name>.json(member targets + optional cross-repoblocksedges). Task ids are repo-qualified (member:local-id); member ledgers stay authoritative and unchanged when the campaign file is removed.--parallel-safewith--campaigncomposes those per-repo waves at query time (see #999); member ledgers stay authoritative and waves are not stored. - Work brief as wiring source of truth (#733): static harness instruction
blocks shrink to depth profiles rendered from the same ordered brief
sections as
brigade work brief. Hook-capable Claude installs default to a shortminimalpointer when SessionStart injects the live brief; file-only installs keepfull. Each integration records its required depth; shared files take the richest active requirement (full-then-minimal and minimal-then-full converge; uninstalling the last full consumer drops back to minimal).--profile minimal|fulloverrides selection. Doctor warns when a minimal profile remains after brief-injecting hooks go inactive. Livebrigade work briefopens with the session-close checklist. - Optional evidence-backed decision checkpoints on work plans (#496):
brigade work task plan <id> --write --decision <id> --decision-prompt "..."with repeatable--optiondeclares a checkpoint;--resolve-decision <id>requires--selected,--rationale, and--evidence-refbefore the plan can be--accepted or beforework task claim/work claim/work task donebegin dependent work. Receipts gain adecisionslist; plan.md renders a Decision checkpoints section. Absentdecisionson legacy receipts stays valid; malformed non-listdecisionsor invalid entries fail closed instead of being dropped.--evidence-refstores an opaque receipt path or external evidence id (no local-file existence check). - Claude compaction restores the work brief on
SessionStartwithsource=compact, reinjecting the live brief beside bounded session-start recall (#736). A one-shot marker still backsUserPromptSubmitwhen compact reinjection is unavailable; the compact path clears stale markers so prompt submit does not double-inject. - Advisory
brigade skills audit <run>compares skill.json process obligations (check/review/handoff) declared by a run’sselected_skill_idsagainst captured verify, review, and handoff ingest receipts (#499). Missing required evidence is reported as findings with exit 0; verify commands may stampobligation_idfor precise matching. Bundledbrigade-workdeclaresverify-through-brigadeandsession-handoffobligations. Orchestrator run identity is exported to workers asBRIGADE_RUN_ID; verify, review, and handoff producers stamp optionalproducer_run_id. Per-run obligation matching requires exactproducer_run_idequality (never timestamp proximity); legacy unstamped receipts are labeled unattributed and cannot satisfy. Public JSON/text output uses repo-relative or collision-resistantexternal:<name>-<digest>path labels and never prints host-private absolute paths, including inload_error/load_warningstext for missing path-based skill selectors.
Fixed
- Run budget resume gate (#593 / PR #864):
brigade runs resumeevaluates the persisted declaration, originalstarted_at, and authoritative budget lifecycle projection beforeAppServer.start(), refusing expired or already-exhausted declared runs without provider or synthesis calls. Undeclared and non-expired declared resumes keep existing recovery behavior. - Run budget restart reclaim (#593 / PR #864): evaluate the wall-clock ceiling before reclaiming a durable pending dispatch can authorize external work, so an expired run cannot relaunch attempt 1 after recovery.
- Run budget declared-only contract (#593 / PR #864): document that hard
wall-clock / worker-dispatch ceilings apply only when a run carries
run_budgetorverification_contract.budget; ordinary undeclaredbrigade run, dogfood, and model-trial paths remain unbounded for backward compatibility (no invented flags, defaults, or timeout→budget mapping). Add undeclared-path regressions alongside the declared enforcement test. - Run budget normal-run wiring (#593 / PR #864): persist declared
verification-contract / run_budget through start and plan receipts so the
coordinator enforces real wall-clock and dispatch ceilings; allocate
same-seat attempts inside the reservation lock so
worker_dispatch_count=1launches once; bridge optionalinput_tokens/output_tokenswhile keeping aggregatetoken_budgetdistinct. - Run budget
token_budgetbridge (#593 / PR #864): preserve aggregate VerificationContracttoken_budget/ receipttokens_usedsemantics; do not map the legacy ceiling ontoobserved.input_tokens. Explicitinput_tokens/output_tokensremain separate observed dimensions. - Run budget dispatch reservations (#593 / PR #864): serialize parallel
same-stage reservations under the coordinator lock; use stable
dispatch:{seat}:{attempt}request identities with durable reservation→run.dispatch.requestedordering and restart-safe idempotency; live Ctrl-C during dispatch terminalizes asoperator-cancelled(not infrastructure-neutral). Wall-clock, usage labels, compatibility diagnostics, and #594 exclusion unchanged. - Claude harness readiness now fails when a global
core.excludesFilerule such as.claude/hides the managed handoffTEMPLATE.md.operator verify-harness --harness claudereturnsready: nowith a nonzero exit, andoperator quickstartpropagates a non-ok status plus the exact non-mutating recovery commandgit check-ignore -v .claude/memory-handoffs/TEMPLATE.md. Brigade never edits global Git configuration; docs/new-user-quickstart.md documents narrow repo-local un-ignore rules. Codex-only onboarding keeps the warning-only contract. (#855)