rabs-cas 2.1.7

Durable RABS content-addressed storage, action-cache indexing, object lifecycle, and publication transactions
docs.rs failed to build rabs-cas-2.1.7
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.

rch — Remote Compilation Helper

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/remote_compilation_helper/main/install.sh?$(date +%s)" | bash -s -- --easy-mode

Transparent remote compilation for multi-agent development

License: MIT + OpenAI/Anthropic Rider Rust codecov


TL;DR

Problem: Many concurrent AI agents can saturate local CPU and make your workstation unusable.

Solution: RCH runs as a Claude Code PreToolUse hook, classifies build-like commands in milliseconds, executes them on remote workers, and returns artifacts/output as if they ran locally.

Design constraint: RCH is fail-open. If remote execution is not safe/possible, commands run locally.


What RCH Intercepts

RCH currently recognizes and can offload:

Ecosystem Intercepted Commands
Rust cargo build, cargo check, cargo clippy, cargo doc, cargo test, cargo nextest run, cargo bench, rustc
Bun/TypeScript bun test, bun typecheck
C/C++ gcc, g++, clang, clang++
Build Systems make, cmake --build, ninja, meson compile
Nix nix build, nix-build, nix flake check, nix develop -c <cmd>, nix shell -c <cmd>

Nix builds only route to workers that advertise a nix capability (a usable nix binary plus a populated /nix/store); on a fleet with no such worker they fall back to local execution (or are refused under RCH_REQUIRE_REMOTE=1, exactly as with Bun/Node). Nix outputs stay in the worker's /nix/store behind a result symlink, so these run as streaming, exit-status-only commands (no artifacts are copied back — the flake source is synced out, the build runs, the result stays remote).

RCH explicitly does not intercept local-mutating or interactive patterns (examples):

  • Package management: cargo install, cargo clean, bun install, bun add, bun remove
  • Bun runners/dev: bun run, bun build, bun dev, bun x / bunx
  • Nix interactive/mutating: bare nix develop / nix shell, nix run, nix repl, nix profile, nix flake update, nix store gc, nix-env, nix-shell
  • Watch/background/piped/redirected commands where deterministic offload is unsafe

Why It Works Well

  • Transparent hook behavior: agents see normal command semantics.
  • 5-tier classification pipeline optimized for very fast non-compilation rejection.
  • Daemon-owned worker state and slot accounting.
  • Cache-aware worker selection and project affinity.
  • Queue + cancellation primitives for overloaded scenarios.
  • Deterministic reliability subsystems for convergence, pressure handling, and remediation.
  • Unified status surface with posture and remediation hints.

Current Architecture

Agent Shell / Claude Code
        |
        v
PreToolUse Hook -> rch (classifier + hook protocol)
        |
        v
      rchd (daemon)
      - worker selection
      - queueing and cancellation metadata
      - health, alerts, telemetry, history
      - reliability subsystems (convergence, pressure, triage)
        |
        v
Remote workers (rch-wkr)
      - execute build/test commands
      - manage worker cache
      - report capabilities/health/telemetry

Workspace crates:

  • rch/: Hook + primary CLI
  • rchd/: Local daemon + scheduling/reliability APIs
  • rch-wkr/: Worker execution/caching agent
  • rch-common/: Shared protocol/types/patterns/UI foundations
  • rch-telemetry/: Telemetry collection/storage integration

Reliability Model (Operational)

RCH now includes a deterministic reliability stack for multi-repo and multi-worker stability:

  • Path-dependency closure planning: builds can include required repository closure rather than a single root.
  • Canonical topology enforcement: worker/project roots are normalized around /data/projects and /dp conventions.
  • Repo convergence service: tracks worker drift vs required repos and can repair drift.
  • Disk pressure resilience: pressure scoring, admission control, safe reclaim with active-build protection.
  • Process triage/remediation: bounded TERM/KILL escalation with audit trail.
  • Cancellation orchestration: deterministic cancellation metadata and worker health integration.
  • Unified posture/reporting: status output includes posture, convergence state, pressure, and actionable remediation hints.

Installation

Recommended: Installer

curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/remote_compilation_helper/main/install.sh?$(date +%s)" | bash -s -- --easy-mode

From Source

git clone https://github.com/Dicklesworthstone/remote_compilation_helper.git
cd remote_compilation_helper
cargo build --release
cp target/release/rch ~/.local/bin/
cp target/release/rchd ~/.local/bin/

Source Build Note

All dependencies — including the FrankenTUI (ftui-*), rich_rust, and TOON (tru) crates — resolve from crates.io, so a clean git clone && cargo build builds on any machine with no special directory layout or pre-cloned dependency tree required.

rsync Requirement (macOS note)

RCH moves source and artifacts with rsync over SSH and works best with rsync 3.2 or newer (--info=progress2 progress, zstd compression, --append-verify resumes). On Linux the distro rsync qualifies.

On macOS, the stock /usr/bin/rsync is openrsync (macOS 15+) or Apple's rsync 2.6.9 (older releases); neither accepts rsync 3.x flags. RCH handles this automatically: it probes rsync --version, prefers a Homebrew/MacPorts rsync (/opt/homebrew/bin/rsync, /usr/local/bin/rsync, /opt/local/bin/rsync) even when it is not first on PATH, and otherwise drives the stock binary with a compatible flag set (--progress --stats -vv, zlib instead of zstd, and the zero-build-output detector fails open). For full-speed transfers:

brew install rsync

rch doctor reports which binary and flavour RCH resolved. To pin one explicitly, set [transfer] rsync_bin = "/path/to/rsync" or export RCH_RSYNC_BIN.


First-Time Setup

Fastest Path

rch init

rch init can guide:

  1. Worker discovery from SSH config/aliases
  2. Worker probing and selection
  3. rch-wkr deployment
  4. Toolchain synchronization
  5. Daemon startup
  6. Hook installation
  7. Validation build

Manual Path

# 1) configure workers
mkdir -p ~/.config/rch
cat > ~/.config/rch/workers.toml << 'TOML'
[[workers]]
id = "css"
host = "203.0.113.20"
user = "ubuntu"
identity_file = "~/.ssh/id_rsa"
total_slots = 32
priority = 100
TOML

# 2) start daemon
rch daemon start

# 3) verify workers
rch workers probe --all

# 4) install hook
rch hook install

# 5) check posture
rch check
rch status --workers --jobs

Command Surface

Global flags:

-v, --verbose
-q, --quiet
-j, --json
-F, --format json|toon
--color auto|always|never
--no-color
--schema
--help-json
--robot-triage

Core Operations

rch daemon start|stop|restart|status|logs|reload
rch workers list|capabilities|probe|benchmark|compare|drain|enable|disable
rch workers init|discover|setup|deploy-binary|sync-toolchain
rch status [--workers] [--jobs]
rch check
rch queue [--watch|--follow]
rch cancel <id> | --all
rch gc [--root DIR]... [--workers <id>...]      # PREVIEW stale rch runtime dirs
rch gc --apply [--root DIR]...                  # ...and actually collect them
rch cache warm|clean|status [--workers <id>...] # remote source/target caches
rch rabs gc plan|run|history [--cas-root DIR] [--mode normal|emergency]
rch rabs worker|doctor|inventory
rch rabs worker reconcile <WORKER>

On macOS, daemon autostarts defer to a registered com.rch.daemon launchd service without restarting its running process. The managed daemon waits if another daemon still owns its socket. A private daemon bypasses this service discovery when given both a nondefault socket and an explicit workers configuration; a custom socket alone cannot be delegated to the shared service.

Fleet GC reports progress on stderr as workers finish. Each worker has a 15-minute budget shared across connections, scanned roots, and collection batches; set rch gc --worker-timeout=120 for a shorter budget. A timeout makes the command exit nonzero while preserving completed workers' results. JSON/TOON output contains one final response with those results and the error.

On timeout, GC terminates its foreground SSH process and allows two additional seconds to reap it; cleanup failures are reported. This does not confirm whether a remote collection batch finished: an interrupted --apply batch is reported as having an unknown outcome, alongside confirmed earlier batches. The default invocation remains a preview.

Capability refreshes skip duplicate requests for a worker while its probe is running. Workers reuse verified Rustup inventory across requests and invalidate entries when installed toolchains or component metadata change. Inventory runs in bounded batches; unfinished or failed entries appear in probe_warnings, and their component capabilities remain unavailable until verified. Disk, load, and project-topology observations are collected again for each request.

Hook + Agent Integration

rch hook install|uninstall|status|test
rch shim install|status|uninstall      # cargo shim: offload builds started by scripts/Makefiles, not just hooked tool calls
rch agents list|status|install-hook|uninstall-hook
rch diagnose "cargo build --release"
rch admit "cargo build --release"      # read-only preflight: offload / local / queue / defer verdict
rch admit --job --require-tool clang -- ./run_shards.sh   # same, for a job-mode admission
rch why miss|refusal                   # RABS: explain a cache-key miss diff or an index refusal code
rch exec -- cargo build --release
rch --robot-triage --json
rch capabilities --json
rch robot-docs guide

Job Mode (non-compilation workloads)

rch exec --job admits an arbitrary NON-compilation workload (sharded tests, fuzzing, benchmarks, mutation testing) onto the same remote rails. It bypasses ONLY the compilation classifier — the PreToolUse hook never sets it, so auto-delegation of ordinary commands is impossible.

rch exec --job -- ./run_shards.sh
# Declared result directories sync back on ANY exit code (including failures):
rch exec --job --result-dir fuzz/corpus --result-dir crashes -- ./fuzz_target.sh

Semantics: the remote exit status surfaces verbatim; no toolchain/worker-env rerun heuristics apply. A declared --result-dir that is missing or only partially transferable fails loudly (RCH-E309, exit 102) regardless of the job's own exit status. Paths must be repository-relative; conflicts with --clean-overlay / --source-content-receipt are refused.

Requiring a verified tool

A job that needs a tool the fleet does not uniformly have can demand a worker that has verified it:

rch exec --job --require-tool clang -- ./native_fuzz.sh

The gate runs before any slot is reserved, and it is evidence-based: the worker must have run the operator-declared probe (see Worker Config Example) successfully. A project can make the requirement a default for all of its jobs:

[jobs]
required_tools = ["clang"]

Project defaults ADD to --require-tool rather than being replaced by it, and both apply to job mode only — a project-wide requirement that silently narrowed every ordinary build's worker pool would be a surprising way to lose the fleet.

rch admit --job [--require-tool NAME]... -- <command> preflights the same decision read-only — it syncs nothing and reserves nothing. Without --job the preflight consults the compilation classifier, so a fuzz script answers local; with it the classifier is bypassed exactly as exec --job does, and the answer is offload. The classification facts are still reported, so a compilation accidentally passed to --job still shows its family.

A required name that no worker has verified admits no worker at all, including when the name is a typo. That is deliberate: silently dropping the requirement would route the job to a machine that cannot run it, and job mode returns the remote exit status verbatim — so the resulting failure would be indistinguishable from the job's own. Selection reports capability_missing:tool:<name>:probe_failed when a worker declared the tool but its probe failed, and ...:not_declared when nothing declared it.

Same-identity job recovery

rch jobs                                      # list durable local job leases
rch jobs attach <wrapper-id> --timeout-secs 300 # observe the original job
rch jobs cancel <wrapper-id>                   # cancel the identity-matched job
rch jobs recover <wrapper-id> --timeout-secs 300

These commands never replay the original command. Recovery requires retained identity and completion/retrieval evidence; ambiguous or missing evidence is reported as an error. For a live wrapper stalled during artifact retrieval, recover_requested means the request was recorded, not that outputs have arrived: use jobs attach to observe completion. The wrapper cancels and reaps the old transfer before retrying that retrieval phase. An absent wrapper can recover outstanding outputs from its retained journal without starting another build. Use a rebuilt daemon with the identity-aware job routes; older daemons are refused.

On POSIX workers, source ownership is saved before synchronization starts and remains active through output retrieval, including after the SSH lock holder disappears. Another job cannot overwrite an overlapping source root while that ownership is unresolved; disjoint source trees can still run concurrently. Uploads, verification, execution, and collection check the same ownership token. Operations within one token are serialized. An uncertain preparation failure ends that preparation; its token is drained and cancelled before a new attempt, so an older upload cannot arrive after verification and change compiler inputs. Recovery of an interrupted preparation drains surviving transfers and cancels that token, so a delayed upload cannot mutate a tree after ownership is released. Once execution may have started, recovery requires its exact completion record. If a worker loses its ownership registry or the local recovery recipe predates this protocol, recovery refuses to infer ownership from the current source bytes. Worker cache cleanup also respects retained source grants, including jobs whose outputs are quiet but still await recovery.

Config + Diagnostics

rch config show|get|set|reset|init|validate|lint|doctor|edit|diff|export
rch doctor [--fix] [--dry-run]
rch doctor --reliability [--check-schemas] [--scope <scope>] [--strict|--lenient] [--json]
rch doctor --reliability --watch [--watch-interval N] [--transitions-only]
rch doctor --runbook RCH-Rnnn | --runbook-list
rch error explain <RCH-Ennn|RCH-Innn|RCH-Rnnn> | list [--category <name>]
rch self-test [--worker <id>|--all]
rch self-test status
rch self-test history --limit 10

Fleet + Release + UX

rch update [--check|--rollback|--fleet]
rch fleet deploy|rollback|status|verify|drain|history
rch speedscore <worker>|--all [--history]
rch dashboard   # alias: rch tui
rch web
rch schema export|list
rch completions generate|install|uninstall|status

Agent Discovery Surface

For AI agents and automation, start with:

rch --robot-triage --json       # quick_ref + recommended commands + health probes
rch capabilities --json         # commands, aliases, env vars, exit codes, output formats
rch robot-docs guide            # in-tool operating guide, no README lookup needed
rch --help-json workers/list    # machine-readable help for nested command paths

--json uses the standard API envelope for command output; --format toon emits the same data as TOON. The legacy rch --capabilities flag remains a raw JSON shortcut for lightweight discovery.


Configuration

Primary files:

  • User config: ~/.config/rch/config.toml
  • Worker list: ~/.config/rch/workers.toml
  • Project override: .rch/config.toml
  • Optional project excludes: .rchignore

Precedence (highest first):

  1. CLI flags
  2. Environment variables
  3. Profile defaults
  4. .env / .rch.env
  5. Project config
  6. User config
  7. Built-in defaults

Canonical Project Root

By default rch expects projects to live under /data/projects (with the /dp alias symlink). Both roots are configurable, so repos under ~/code, /workspace, etc. work without relocating anything:

[path_topology]
canonical_root = "/home/me/code"   # default: /data/projects
alias_root = "/home/me/code"       # default: /dp (set equal to canonical_root if you have no alias)

Or via the CLI / environment (env wins over TOML):

rch config set path_topology.canonical_root /home/me/code
rch config set path_topology.alias_root /home/me/code
# or
export RCH_CANONICAL_PROJECT_ROOT=/home/me/code
export RCH_ALIAS_PROJECT_ROOT=/home/me/code

Empty strings are treated as unset (defaults apply). See docs/guides/configuration.md ([path_topology]) for details.

Minimal Example

[general]
enabled = true
force_local = false
force_remote = false
socket_path = "~/.cache/rch/rch.sock"
log_level = "info"

[compilation]
confidence_threshold = 0.85
min_local_time_ms = 2000
remote_speedup_threshold = 1.2
build_slots = 4
test_slots = 8
check_slots = 2
build_timeout_sec = 300
test_timeout_sec = 1800
bun_timeout_sec = 600
external_timeout_enabled = true

[transfer]
compression_level = 3
# Staging base for the transfer paths that do NOT mirror the client's layout:
# `--clean-overlay` roots, Windows workers, and `rch cache warm`. An ordinary
# `rch exec` mirrors the project under the worker's `[path_topology]
# canonical_root` instead, with the pooled Cargo target store inside that
# mirror — see `[remediation.pooled_target] store_base` below to place those
# stores on another filesystem.
remote_base = "/data/tmp/rch"
# Optional per-attempt source-sync cap. When unset, the default is payload-aware:
# 30 seconds plus one second per MiB, capped at one hour.
# sync_timeout_ms = 120000
# Abort a source sync after this many seconds of NO rsync output (dead channel /
# wedged rsync); a progressing transfer is never affected. 0 disables.
# source_sync_silence_timeout_secs = 120
# Explicit rsync binary. Unset lets RCH prefer a modern (3.x) rsync over the
# stock macOS openrsync automatically; RCH_RSYNC_BIN overrides this.
# rsync_bin = "/opt/homebrew/bin/rsync"
adaptive_compression = true
verify_artifacts = false
max_transfer_mb = 2048

[remediation.pooled_target]
# Where pooled Cargo target stores are PLACED on the worker. Unset (the
# default) keeps them inside the project mirror, i.e. under the canonical
# root, on whatever filesystem holds it. Set an absolute path to move every
# pooled store to <store_base>/<project_id>/.rch-target-<worker>-pool-<key>,
# which is how a worker with a small root disk keeps multi-GB warm pools on a
# larger volume. `rch gc`, `rch cache status`, and the daemon sweep scan this
# root in addition to `remote_base` (with `reaper_max_cache_gb` applying per
# scan root). Ignored for Windows workers.
# store_base = "/bigdisk/rch-pools"
#
# Where the reaper SCANS for stale target dirs (not a placement setting).
# remote_base = "/data/projects"
#
# EXTRA roots `rch gc` scans, on top of the ones it derives itself
# (`remote_base`, `store_base`, and the worker temp base resolved exactly as
# the code that creates the dirs resolves it: $TMPDIR -> /data/tmp -> /tmp).
# Use this for a runtime root rch cannot know about, e.g. an operator build
# root on a mounted volume. Each entry must be absolute, `..`-free and free of
# shell metacharacters -- it is embedded in a remote command. `rch gc --root`
# adds one for a single run.
# gc_extra_roots = ["/mnt/big/rch"]
#
# Idle days before `rch gc` may collect a durable per-worker Cargo cache
# (`rch-cargo-cache-*`). 0 disables. Collection ALSO requires zero open file
# descriptors under the dir and no live process rooted at it; a gate that
# cannot be evaluated counts as "in use".
# gc_cargo_cache_idle_days = 14

[selection]
strategy = "balanced"
# Reserve this much free disk, then budget space for each concurrent slot (GiB).
min_free_gb = 10.0
disk_gb_per_slot = 10.0

[self_healing]
hook_starts_daemon = true
daemon_installs_hooks = true

[alerts]
enabled = true
suppress_duplicates_secs = 300

When the Unix launcher enforces a compilation deadline, RCH preserves exit 137 and reports that the configured deadline expired. It does not retry on a larger worker or fall back locally. Machine-mode rch exec reports outcome: "deadline_exceeded". This classification requires evidence from the launcher; exit 137 or elapsed time alone does not establish a timeout.

Unix artifact downloads estimate the files matched by the retrieval filters before transferring them. Their default total retry budget grows with that size (30 seconds plus one second per MiB, adjusted for a slower bwlimit_kbps, up to one hour). An explicit max_transfer_time_ms keeps its existing hard ceiling and skips estimation. Rsync also stops after 30 seconds without network I/O. If estimation fails, RCH logs the failure and retains the configured retry budget. A literal cargo build --bin NAME or --example NAME retrieves only the named payloads, not the pooled target's other cached outputs. Anything the narrowing cannot parse exactly (unknown flags, shell expansion, other target kinds, response files, custom target specs) keeps the broad per-profile filters, which can include other cached build outputs in the same profile.

Windows workers return artifacts through tar in both quiet and interactive execution. The transfer honors the requested output patterns and cache exclusions, including custom Cargo target directories and explicit compiler outputs outside target/. RCH inventories regular files, downloads only selected paths, and validates the complete archive before writing outputs locally. Remote tar errors, missing selected files, and unexpected archive members fail artifact retrieval; successful downloads include file counts and a manifest for output validation. Inventory and transfer use bounded deadlines, and cancellation reaps the transport before recovery can begin. This transport requires GNU find from the worker's Git for Windows installation and native tar.

Artifact staging retains the original caller directory as the source-protection reference. Recovery restores that reference for each output phase, so an empty staging directory cannot make source files eligible for a broad output pattern.

Built-in worker selection defaults to balanced, which blends speed, load, health, and cache affinity. Use priority only when you want explicit worker-priority control, and fair_fastest when you want extra load spreading.

Worker capacity is capped at the configured total_slots and reduced to floor((free_disk_gb - min_free_gb) / disk_gb_per_slot) as disk fills, with a minimum of zero. Unknown disk telemetry keeps the configured ceiling. Existing builds retain their reservations when capacity drops; new work waits for space. rch workers list reports the effective total when the daemon is reachable. Change the budget with rch config set selection.disk_gb_per_slot 10 and restart the daemon to apply it.

Balanced selection also favors free disk before a worker reaches the admission floor. selection.weights.disk (default 0.2, range 0 to 1) adds disk headroom credit after the other score adjustments. Equal ample headroom preserves the existing worker order. Credit rises from zero at min_free_gb to full credit at 25% free disk, with at least one slot's disk budget above the floor on small disks. Unknown measurements receive full credit; the admission gate still handles missing telemetry. Affinity pins and explicit worker requests retain their existing behavior. Set the weight to zero to disable disk ranking independently of disk admission and slot limits, then restart the daemon to apply the change.

rch workers list shows daemon-reported free disk and pressure. rch workers probe shows fresh disk measurements alongside latency and labels any cached daemon pressure separately. JSON and TOON include numeric free GiB and free ratios; unavailable measurements remain null (displayed as unknown in text).

Worker Config Example

[[workers]]
id = "css"
host = "203.0.113.20"
user = "ubuntu"
identity_file = "~/.ssh/id_rsa"
total_slots = 32
priority = 100
tags = ["fast", "ssd"]

# Verified named tools. Each entry is a FIXED argv the worker runs as the
# configured user; a zero exit marks the tool present, anything else marks it
# absent. Unlike `tags`, which are an unverified naming convention, these are
# evidence — which is what lets `--require-tool` gate worker selection.
tools = [
  { name = "clang",  command = ["clang", "--version"] },
  { name = "ld.lld", command = ["/usr/bin/ld.lld", "--version"] },
]

Names must be ASCII letters, digits, -, _, . or +: a name carrying a space or = could not survive the probe's fact format and would be read back as a different name, so it is refused when the config loads rather than becoming a capability lie. rch workers capabilities --refresh reports the results under Named tools as verified: and failed: — a declared probe that fails is a different operational state from one that was never declared, and only the first points at a broken worker.

There is deliberately no general remote-shell probe API. A caller that could ask a worker to run an arbitrary command "to check for a tool" would be a remote execution primitive wearing a capability-probe hat; declaring the argv in operator config keeps the set of probe commands finite and reviewable.


Output Modes

RCH auto-selects output mode by context:

  • hook: strict JSON for hook protocol
  • machine: explicit machine output (--json, --format)
  • interactive: rich terminal rendering
  • colored: ANSI-only when forced without TTY
  • plain: text fallback

Environment controls:

  • RCH_JSON=1, RCH_HOOK_MODE=1
  • NO_COLOR=1, FORCE_COLOR=1, FORCE_COLOR=0
  • RCH_OUTPUT_FORMAT=json|toon, TOON_DEFAULT_FORMAT

JSON responses use a stable envelope (api_version, timestamp, success, data, error).


Placement Controls

Worker placement, strict-remote, queue, wait-timeout, visibility, and target-dir behavior are first-class, canonical controls — not folklore. The authoritative list is discoverable at runtime (rch capabilities --json), and the resolved plan for any command is shown by rch diagnose <command> --json under data.placement (and in the human Placement Controls section):

Control Env (aliases) Effect
Requested worker RCH_WORKER (RCH_WORKERS) Request specific worker(s) by id. Still passes capability/admission checks; an inadmissible requested worker is refused with a stable RCH-Innn reason code and a next action — never silently swapped.
Requested profile RCH_PRESET Named execution profile (recorded as requested_profile).
Strict remote (fail-closed) RCH_REQUIRE_REMOTE Refuse local fallback (proof mode). Takes precedence over RCH_FORCE_REMOTE.
Force remote (fail-open) RCH_FORCE_REMOTE Always attempt offload (bypass local-time/speedup gating) but still fail open to local. Distinct from RCH_REQUIRE_REMOTE.
Queue when busy RCH_QUEUE_WHEN_BUSY (default 1) Wait for a busy worker instead of falling back to local. Set 0 to disable.
Wait timeout RCH_DAEMON_WAIT_RESPONSE_TIMEOUT_SECS (RCH_DAEMON_RESPONSE_TIMEOUT_SECS) Max seconds to wait for a queued worker.
Visibility RCH_VISIBILITY=none|summary|verbose (RCH_QUIET, RCH_VERBOSE) Hook output verbosity.
Target dir RCH_DISABLE_TARGET_REUSE Legacy unique-per-job remote target dir instead of the pooled, reuse-friendly dir.
Source-sync timeout RCH_SYNC_TIMEOUT_MS Per-attempt source-upload timeout in milliseconds (1000..=3600000). Unset uses the payload-aware default. This does not change remote Cargo or artifact-return timeouts.

The resolved plan reports requested_worker, requested_profile, effective_worker, strict_remote_policy, queue_policy, visibility_mode, wait_timeout_ms, target_dir_policy, the requested-worker admissibility outcome, and a diagnostics list. Any control value that cannot be applied as written (an unrecognized value or a superseded alias) surfaces a diagnostic rather than being silently ignored.


Monitoring and Observability

RCH exposes observability through daemon APIs and metrics:

  • daemon health/readiness endpoints
  • Prometheus metrics collection
  • opt-in OTLP export of daemon request-duration metrics
  • telemetry-backed worker SpeedScore history
  • queue/build history, active alerts, cancellation metadata in status APIs
  • an opt-in tailnet status API ([api] bind = "tailscale"): the daemon's full /status JSON over TCP, bearer-token gated, bound only to loopback or Tailscale addresses — so agents on other machines can ask a dispatcher "why are you local-only right now?" without ssh
  • the fleet dashboard under dashboard/ (encrypted static console + an agent endpoint GET /api/fleet?view=problems|diagnose|help); rch web opens it

Enable OTLP metrics in the daemon's environment with RCH_OTEL_ENABLED=1 and RCH_OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 (OTLP/gRPC). The endpoint falls back to OTEL_EXPORTER_OTLP_ENDPOINT. OTEL_SERVICE_NAME defaults to rchd; RCH_OTEL_EXPORT_INTERVAL_SECS defaults to 30. rch_request_duration_seconds{entrypoint="rchd_api"} records finite Unix-socket request handling after parsing, including errors and cancellation. Event streams and the separate HTTP API are excluded. The same observations reach Prometheus, including when OTLP is disabled. Shutdown drains request tasks before a best-effort final export. The banner reports exporter configuration, not collector connectivity.

The same opt-in settings enable direct OTLP verdict export for single-shot rch doctor --reliability, including under --quiet. Its service name defaults to rch. Each report emits one rch_doctor_verdict_total observation, labeled with the actual verdict and scope (combined scopes use other), and flushes before returning its original exit code. Reports whose daemon_unreachable flag is true also increment rch_doctor_daemon_unreachable_total once, even when several diagnostics report the missing daemon. The asynchronous daemon_status and repo_convergence probes also emit rch_doctor_probe_duration_seconds{probe,result} once on completion, RPC error, or timeout. Durations measure each probe's execution, not time spent waiting to join other probes. Skipped probes emit nothing; panic/cancellation and helper/ownership probe durations are not yet exported. With --fix, each remediation step also emits rch_doctor_fix_steps_total and rch_doctor_fix_duration_seconds, labeled by its actual outcome: applied, already_satisfied, would_apply, manual, or failed. Dry-run previews remain would_apply; they are never counted as applied changes. Durations measure step execution, excluding the initial shared configuration load. These CLI metrics are not forwarded to the daemon's Prometheus endpoint. Hook/watch telemetry, remaining doctor event metrics, and trace/log export remain unfinished.

Quick checks:

rch status --workers --jobs
rch speedscore --all
rch doctor --json

On Linux boxes configured with [general] role = "dispatcher", status and doctor also report unmanaged local Cargo/rustc processes. The warning remains visible while those processes run, including when the daemon is unavailable. Check rch shim status, PATH order, and absolute-path toolchain Cargo invocations. Processes carrying RCH_CARGO_WRAPPER_BYPASS=1 or descended from rch are excluded.

An alarm event is logged once per observed episode and clears after a successful scan finds no unmanaged compiler processes. Status and doctor share a locked one-byte latch at $XDG_CACHE_HOME/rch/local-build-alarm (normally ~/.cache/rch/local-build-alarm). Scan failures preserve the previous episode; latch failures are reported without hiding current builds. Reliability doctor includes this check in the triage scope and its default all scope.


Testing and Validation

Workspace checks:

cargo fmt --check
cargo check --workspace --all-targets
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace

Reliability and E2E suites are provided under:

  • tests/
  • tests/e2e/
  • rch-common/tests/ (contract/reliability/perf suites)

If you are running CPU-intensive validation manually and want explicit offload:

rch exec -- cargo check --workspace --all-targets
rch exec -- cargo test --workspace
rch exec -- cargo clippy --workspace --all-targets -- -D warnings

When local fallback is not acceptable, set RCH_REQUIRE_REMOTE=1 on the rch exec process and keep the build command as direct argv:

RCH_REQUIRE_REMOTE=1 rch exec -- cargo test --workspace
RCH_REQUIRE_REMOTE=1 rch exec -- cargo clippy --workspace --all-targets -- -D warnings

Exact source-content receipts

Proof-oriented callers can require a single-worker, fail-closed source transfer and retain the exact regular-file bytes admitted on that worker:

RCH_REQUIRE_REMOTE=1 rch exec --source-content-receipt -- \
  cargo test --locked --workspace

Receipt mode resolves the active Cargo path-dependency closure, gives every root an invocation-unique worker path, transfers with checksum comparison, and reopens every selected file before and after the command. The emitted rch.source_content_receipt.v1 JSON binds the worker and build IDs, exact command digest and exit code, transfer filter policy, per-root file paths, lengths, executable bits, SHA-256 digests, and content roots. Any unproved dependency closure, transfer delta, worker verification error, local lasting change, retry, or local fallback refuses the invocation instead of emitting a receipt. Receipt mode currently requires the Unix rsync transport and cannot be combined with clean-overlay mode.

The receipt is emitted after the remote command and its post-command source barrier, but before ordinary artifact retrieval completes. Artifact-grade callers must therefore retain both the receipt and the later successful retrieval/terminal transcript; the receipt alone is not proof that build artifacts reached the caller. A recursive local mutation watcher is still needed when transient edit-and-restore (ABA) detection is part of the caller's source-stability contract.

Clean Git overlays for shared working trees

rch exec can build an immutable committed tree plus an explicit, repeatable set of local paths without transferring unrelated working-tree changes:

RCH_REQUIRE_REMOTE=1 rch exec \
  --base HEAD \
  --clean-overlay \
  --overlay-path src/lib.rs \
  --overlay-path tests/focused.rs \
  -- cargo test --test focused

RCH_REQUIRE_REMOTE=1 rch exec \
  --base HEAD \
  --clean-overlay \
  --no-overlay \
  -- cargo check --workspace --all-targets

The client resolves --base to a commit object, streams that Git archive into a fresh isolated worker path, and then uploads only the literal repository- relative --overlay-path selections. Modified and untracked files are supported; selected deletions, absolute/traversing paths, Git metadata, non-ASCII/backslash/control-character paths, case-only or otherwise ambiguous filesystem spellings, overlay symlinks, empty overlay directories, submodules, and Git archive export-ignore/export-subst attributes fail closed. Exactly one of one-or-more --overlay-path options or --no-overlay is required. Clean-overlay mode also implies remote-only execution even when RCH_REQUIRE_REMOTE was omitted, so a worker or transfer failure never falls back to the ambient local tree. Explicit clean-overlay execution also admits the read-only cargo fmt --check diagnostic, which the ordinary interception classifier intentionally leaves local.

On Unix workers, pooled Cargo builds use a stable source path paired with a versioned target cache. Jobs sharing that pair wait until the preceding job has finished execution, artifact retrieval, and source retirement. Different pairs remain independent. Source timestamps are refreshed beyond cached artifacts so changed archive or overlay contents cannot appear older than a previous build; registry and Git dependency caches remain reusable. Clean-overlay execution also pins Cargo's build.build-dir to its assigned target directory, overriding ambient and command-line Cargo configuration. Explicit --build-dir options are refused. This placement applies to isolated builds too, so disabling target reuse cannot inherit shared intermediates. Managed placement currently supports built-in Cargo build commands and Clippy. Other compiling Cargo subcommands, including Nextest and Zigbuild, are refused in clean-overlay mode.

An interrupted pair retains its ownership marker and refuses further reuse, even if its SSH lock connection disappears. Inspect the prior job before recovering that pair, or use RCH_DISABLE_TARGET_REUSE=1 for a fresh isolated build. Non-Cargo Windows clean-overlay jobs also use isolated targets. The ownership marker protects against process interruption; it is not a power-loss durability guarantee.

Clean-overlay materializes the primary Git repository and explicitly bound sibling repositories. In-repository Cargo workspace members are present in the archive. Before Cargo starts, selected manifests, configuration, symlinks, and command-line path overrides are checked against the selected Git base and overlays. Escaping paths are refused with RCH-E413; retained sibling directories on the worker cannot supply those dependencies. To include a sibling Git repository, bind its revision explicitly:

rch exec --base HEAD --clean-overlay --no-overlay \
  --dependency-base ../dep=HEAD -- cargo test

Repeat --dependency-base PATH=REV for every external root, including transitive dependencies. Each revision resolves to a commit before worker selection; dirty files in those repositories are excluded. The selected repositories must be siblings of the primary Git root. They are staged under one owned container, preserving ../dep references, and the execution receipt names each commit, Git tree hash, and remote path. All roots share the same execution lease and are retired together. Cross-repository symlinks and non-sibling layouts are currently refused.

This check conservatively covers all selected manifests, including inactive fixtures. File-based --config, configuration includes, and changes to Cargo's working directory are refused. On the worker, any Cargo-home or ancestor configuration is refused before Cargo starts, even if that configuration would be harmless. Cargo clean-overlay jobs currently require a Unix worker for this check. These checks do not sandbox arbitrary file reads by build scripts. The client re-fingerprints overlays after upload and refuses execution if their contents changed during admission or transfer.

Do not batch several Cargo commands behind a shell wrapper such as rch exec -- bash -lc "cargo test ... && cargo test ...". Shell-wrapped commands are classified as non-compilation for hook safety. Under RCH_REQUIRE_REMOTE=1, RCH refuses that local fallback before executing the shell and reports RCH-E301; without the env var, ordinary non-compilation commands may still run locally. For several focused checks, run separate direct RCH_REQUIRE_REMOTE=1 rch exec -- cargo ... invocations.


Security Model

  • Transport uses SSH.
  • Worker commands are constrained to classified execution paths.
  • Sensitive field masking and structured error taxonomy are built in.
  • rch update enforces the release's SHA-256 checksums and verifies the archive's .minisig against a public key pinned in the binary (minisign key 69B3955C8D2E62A8, which signs every release since 2.0.0); a bad signature aborts the update. A Sigstore bundle is also verified when the release ships one and cosign is installed. --skip-verify bypasses signature checks only. install.sh still verifies checksums only.
  • Hook path remains fail-open to avoid deadlocks/stalls.

Operational recommendations:

  1. Use workers you control.
  2. Use dedicated SSH keys for worker access.
  3. Keep workers patched and isolated.
  4. Enable telemetry/alerting for production-like use.

Limitations

  • Designed around SSH-based Linux worker environments.
  • Tooling assumptions are strongest for Rust and selected build/test commands.
  • Remote performance gains depend on network + worker capacity + project shape.
  • The fleet dashboard (dashboard/) needs Node and a static host (Vercel/Pages); its live path needs a Vercel Blob store and rchd's [api] listener on a tailnet.

FAQ

Does RCH block my command if the daemon/workers fail?

No. It fails open and allows local execution.

Can I force local or force remote per project?

Yes, via .rch/config.toml (general.force_local / general.force_remote).

Is queue/cancel supported?

Yes. Use rch queue and rch cancel.

Can I inspect why a command is or is not intercepted?

Yes. Use rch diagnose "<command>".


About Contributions

Please don't take this the wrong way, but I do not accept outside contributions for my projects. You can still open issues and PRs for discussion/proof-of-fix, but I review and re-implement changes independently.


License

MIT License with an OpenAI/Anthropic rider — this is not ordinary OSI MIT: no rights are granted to the restricted parties named in the rider. See LICENSE for the exact, controlling terms (SPDX: LicenseRef-MIT-OpenAI-Anthropic-Rider).