# rch — Remote Compilation Helper
<div align="center">
<img src="rch_illustration.webp" alt="rch - Remote Compilation Helper for AI coding agents">
</div>
<div align="center">
<h3>Quick Install</h3>
```bash
<p><em>Installs `rch` + `rchd`, bootstraps config, and can install/start the background daemon. If remote execution cannot proceed, RCH fails open to local execution.</em></p>
</div>
<div align="center">
<img src="rch_diagram.webp" alt="rch architecture diagram">
</div>
<div align="center">
**Transparent remote compilation for multi-agent development**
[](LICENSE)
[](https://www.rust-lang.org/)
[](https://codecov.io/gh/Dicklesworthstone/remote_compilation_helper)
</div>
---
## TL;DR
**Problem**: Many concurrent AI agents can saturate local CPU and make your workstation unusable.
**Solution**: RCH runs as a Claude Code PreToolUse hook, classifies build-like commands in milliseconds, executes them on remote workers, and returns artifacts/output as if they ran locally.
**Design constraint**: RCH is fail-open. If remote execution is not safe/possible, commands run locally.
---
## What RCH Intercepts
RCH currently recognizes and can offload:
| Rust | `cargo build`, `cargo check`, `cargo clippy`, `cargo doc`, `cargo test`, `cargo nextest run`, `cargo bench`, `rustc` |
| Bun/TypeScript | `bun test`, `bun typecheck` |
| C/C++ | `gcc`, `g++`, `clang`, `clang++` |
| Build Systems | `make`, `cmake --build`, `ninja`, `meson compile` |
| Nix | `nix build`, `nix-build`, `nix flake check`, `nix develop -c <cmd>`, `nix shell -c <cmd>` |
Nix builds only route to workers that advertise a `nix` capability (a usable `nix`
binary plus a populated `/nix/store`); on a fleet with no such worker they fall
back to local execution (or are refused under `RCH_REQUIRE_REMOTE=1`, exactly as
with Bun/Node). Nix outputs stay in the worker's `/nix/store` behind a `result`
symlink, so these run as streaming, exit-status-only commands (no artifacts are
copied back — the flake source is synced out, the build runs, the result stays
remote).
RCH explicitly does **not** intercept local-mutating or interactive patterns (examples):
- Package management: `cargo install`, `cargo clean`, `bun install`, `bun add`, `bun remove`
- Bun runners/dev: `bun run`, `bun build`, `bun dev`, `bun x` / `bunx`
- Nix interactive/mutating: bare `nix develop` / `nix shell`, `nix run`, `nix repl`,
`nix profile`, `nix flake update`, `nix store gc`, `nix-env`, `nix-shell`
- Watch/background/piped/redirected commands where deterministic offload is unsafe
---
## Why It Works Well
- Transparent hook behavior: agents see normal command semantics.
- 5-tier classification pipeline optimized for very fast non-compilation rejection.
- Daemon-owned worker state and slot accounting.
- Cache-aware worker selection and project affinity.
- Queue + cancellation primitives for overloaded scenarios.
- Deterministic reliability subsystems for convergence, pressure handling, and remediation.
- Unified status surface with posture and remediation hints.
---
## Current Architecture
```text
Agent Shell / Claude Code
|
v
PreToolUse Hook -> rch (classifier + hook protocol)
|
v
rchd (daemon)
- worker selection
- queueing and cancellation metadata
- health, alerts, telemetry, history
- reliability subsystems (convergence, pressure, triage)
|
v
Remote workers (rch-wkr)
- execute build/test commands
- manage worker cache
- report capabilities/health/telemetry
```
Workspace crates:
- `rch/`: Hook + primary CLI
- `rchd/`: Local daemon + scheduling/reliability APIs
- `rch-wkr/`: Worker execution/caching agent
- `rch-common/`: Shared protocol/types/patterns/UI foundations
- `rch-telemetry/`: Telemetry collection/storage integration
---
## Reliability Model (Operational)
RCH now includes a deterministic reliability stack for multi-repo and multi-worker stability:
- **Path-dependency closure planning**: builds can include required repository closure rather than a single root.
- **Canonical topology enforcement**: worker/project roots are normalized around `/data/projects` and `/dp` conventions.
- **Repo convergence service**: tracks worker drift vs required repos and can repair drift.
- **Disk pressure resilience**: pressure scoring, admission control, safe reclaim with active-build protection.
- **Process triage/remediation**: bounded TERM/KILL escalation with audit trail.
- **Cancellation orchestration**: deterministic cancellation metadata and worker health integration.
- **Unified posture/reporting**: status output includes posture, convergence state, pressure, and actionable remediation hints.
---
## Installation
### Recommended: Installer
```bash
### From Source
```bash
git clone https://github.com/Dicklesworthstone/remote_compilation_helper.git
cd remote_compilation_helper
cargo build --release
cp target/release/rch ~/.local/bin/
cp target/release/rchd ~/.local/bin/
```
### Source Build Note
All dependencies — including the FrankenTUI (`ftui-*`), `rich_rust`, and TOON
(`tru`) crates — resolve from crates.io, so a clean `git clone && cargo build`
builds on any machine with no special directory layout or pre-cloned
dependency tree required.
### rsync Requirement (macOS note)
RCH moves source and artifacts with `rsync` over SSH and works best with
**rsync 3.2 or newer** (`--info=progress2` progress, zstd compression,
`--append-verify` resumes). On Linux the distro rsync qualifies.
On **macOS**, the stock `/usr/bin/rsync` is *openrsync* (macOS 15+) or Apple's
rsync 2.6.9 (older releases); neither accepts rsync 3.x flags. RCH handles this
automatically: it probes `rsync --version`, prefers a Homebrew/MacPorts rsync
(`/opt/homebrew/bin/rsync`, `/usr/local/bin/rsync`, `/opt/local/bin/rsync`)
even when it is not first on `PATH`, and otherwise drives the stock binary with
a compatible flag set (`--progress --stats -vv`, zlib instead of zstd, and the
zero-build-output detector fails open). For full-speed transfers:
```bash
brew install rsync
```
`rch doctor` reports which binary and flavour RCH resolved. To pin one
explicitly, set `[transfer] rsync_bin = "/path/to/rsync"` or export
`RCH_RSYNC_BIN`.
---
## First-Time Setup
### Fastest Path
```bash
rch init
```
`rch init` can guide:
1. Worker discovery from SSH config/aliases
2. Worker probing and selection
3. `rch-wkr` deployment
4. Toolchain synchronization
5. Daemon startup
6. Hook installation
7. Validation build
### Manual Path
```bash
# 1) configure workers
mkdir -p ~/.config/rch
cat > ~/.config/rch/workers.toml << 'TOML'
[[workers]]
id = "css"
host = "203.0.113.20"
user = "ubuntu"
identity_file = "~/.ssh/id_rsa"
total_slots = 32
priority = 100
TOML
# 2) start daemon
rch daemon start
# 3) verify workers
rch workers probe --all
# 4) install hook
rch hook install
# 5) check posture
rch check
rch status --workers --jobs
```
---
## Command Surface
Global flags:
```bash
-v, --verbose
-q, --quiet
-j, --json
--no-color
--schema
--help-json
--robot-triage
```
### Core Operations
```bash
rch workers init|discover|setup|deploy-binary|sync-toolchain
rch status [--workers] [--jobs]
rch check
rch queue [--watch|--follow]
rch cancel <id> | --all
rch gc [--root DIR]... [--workers <id>...] # PREVIEW stale rch runtime dirs
rch gc --apply [--root DIR]... # ...and actually collect them
rch cache warm|clean|status [--workers <id>...] # remote source/target caches
rch rabs gc plan|run|history [--cas-root DIR] [--mode normal|emergency]
rch rabs worker|doctor|inventory
rch rabs worker reconcile <WORKER>
```
On macOS, daemon autostarts defer to a registered `com.rch.daemon` launchd
service without restarting its running process. The managed daemon waits if
another daemon still owns its socket. A private daemon bypasses this service
discovery when given both a nondefault socket and an explicit workers
configuration; a custom socket alone cannot be delegated to the shared service.
Fleet GC reports progress on stderr as workers finish. Each worker has a
15-minute budget shared across connections, scanned roots, and collection
batches; set `rch gc --worker-timeout=120` for a shorter budget. A timeout makes
the command exit nonzero while preserving completed workers' results. JSON/TOON
output contains one final response with those results and the error.
On timeout, GC terminates its foreground SSH process and allows two additional
seconds to reap it; cleanup failures are reported. This does not confirm whether
a remote collection batch finished: an interrupted `--apply` batch is reported
as having an unknown outcome, alongside confirmed earlier batches. The default
invocation remains a preview.
Capability refreshes skip duplicate requests for a worker while its probe is
running. Workers reuse verified Rustup inventory across requests and invalidate
entries when installed toolchains or component metadata change. Inventory runs
in bounded batches; unfinished or failed entries appear in `probe_warnings`,
and their component capabilities remain unavailable until verified. Disk, load,
and project-topology observations are collected again for each request.
### Hook + Agent Integration
```bash
rch agents list|status|install-hook|uninstall-hook
rch diagnose "cargo build --release"
rch admit "cargo build --release" # read-only preflight: offload / local / queue / defer verdict
rch admit --job --require-tool clang -- ./run_shards.sh # same, for a job-mode admission
rch why miss|refusal # RABS: explain a cache-key miss diff or an index refusal code
rch exec -- cargo build --release
rch --robot-triage --json
rch capabilities --json
rch robot-docs guide
```
### Job Mode (non-compilation workloads)
`rch exec --job` admits an arbitrary NON-compilation workload (sharded tests,
fuzzing, benchmarks, mutation testing) onto the same remote rails. It bypasses
ONLY the compilation classifier — the PreToolUse hook never sets it, so
auto-delegation of ordinary commands is impossible.
```bash
rch exec --job -- ./run_shards.sh
# Declared result directories sync back on ANY exit code (including failures):
rch exec --job --result-dir fuzz/corpus --result-dir crashes -- ./fuzz_target.sh
```
Semantics: the remote exit status surfaces verbatim; no toolchain/worker-env
rerun heuristics apply. A declared `--result-dir` that is missing or only
partially transferable fails loudly (`RCH-E309`, exit 102) regardless of the
job's own exit status. Paths must be repository-relative; conflicts with
`--clean-overlay` / `--source-content-receipt` are refused.
#### Requiring a verified tool
A job that needs a tool the fleet does not uniformly have can demand a worker
that has **verified** it:
```bash
rch exec --job --require-tool clang -- ./native_fuzz.sh
```
The gate runs before any slot is reserved, and it is evidence-based: the worker
must have run the operator-declared probe (see
[Worker Config Example](#worker-config-example)) successfully. A project can
make the requirement a default for all of its jobs:
```toml
[jobs]
required_tools = ["clang"]
```
Project defaults ADD to `--require-tool` rather than being replaced by it, and
both apply to job mode only — a project-wide requirement that silently narrowed
every ordinary build's worker pool would be a surprising way to lose the fleet.
`rch admit --job [--require-tool NAME]... -- <command>` preflights the same
decision read-only — it syncs nothing and reserves nothing. Without `--job` the
preflight consults the compilation classifier, so a fuzz script answers
`local`; with it the classifier is bypassed exactly as `exec --job` does, and
the answer is `offload`. The classification facts are still reported, so a
compilation accidentally passed to `--job` still shows its family.
A required name that no worker has verified admits **no worker at all**,
including when the name is a typo. That is deliberate: silently dropping the
requirement would route the job to a machine that cannot run it, and job mode
returns the remote exit status verbatim — so the resulting failure would be
indistinguishable from the job's own. Selection reports
`capability_missing:tool:<name>:probe_failed` when a worker declared the tool
but its probe failed, and `...:not_declared` when nothing declared it.
### Same-identity job recovery
```bash
rch jobs # list durable local job leases
rch jobs attach <wrapper-id> --timeout-secs 300 # observe the original job
rch jobs cancel <wrapper-id> # cancel the identity-matched job
rch jobs recover <wrapper-id> --timeout-secs 300
```
These commands never replay the original command. Recovery requires retained
identity and completion/retrieval evidence; ambiguous or missing evidence is
reported as an error. For a live wrapper stalled during artifact retrieval,
`recover_requested` means the request was recorded, not that outputs have arrived:
use `jobs attach` to observe completion. The wrapper cancels and reaps the old
transfer before retrying that retrieval phase. An absent wrapper can recover
outstanding outputs from its retained journal without starting another build.
Use a rebuilt daemon with the identity-aware job routes; older daemons are refused.
On POSIX workers, source ownership is saved before synchronization starts and
remains active through output retrieval, including after the SSH lock holder
disappears. Another job cannot overwrite an overlapping source root while that
ownership is unresolved; disjoint source trees can still run concurrently.
Uploads, verification, execution, and collection check the same ownership token.
Operations within one token are serialized. An uncertain preparation failure
ends that preparation; its token is drained and cancelled before a new attempt,
so an older upload cannot arrive after verification and change compiler inputs.
Recovery of an interrupted preparation drains surviving transfers and cancels
that token, so a delayed upload cannot mutate a tree after ownership is released.
Once execution may have started, recovery requires its exact completion record.
If a worker loses its ownership registry or the local recovery recipe predates
this protocol, recovery refuses to infer ownership from the current source bytes.
Worker cache cleanup also respects retained source grants, including jobs whose
outputs are quiet but still await recovery.
### Config + Diagnostics
```bash
rch doctor --reliability [--check-schemas] [--scope <scope>] [--strict|--lenient] [--json]
rch doctor --reliability --watch [--watch-interval N] [--transitions-only]
rch doctor --runbook RCH-Rnnn | --runbook-list
rch error explain <RCH-Ennn|RCH-Innn|RCH-Rnnn> | list [--category <name>]
rch self-test [--worker <id>|--all]
rch self-test status
rch self-test history --limit 10
```
### Fleet + Release + UX
```bash
rch update [--check|--rollback|--fleet]
rch dashboard # alias: rch tui
rch web
rch schema export|list
rch completions generate|install|uninstall|status
```
### Agent Discovery Surface
For AI agents and automation, start with:
```bash
rch --robot-triage --json # quick_ref + recommended commands + health probes
rch capabilities --json # commands, aliases, env vars, exit codes, output formats
rch robot-docs guide # in-tool operating guide, no README lookup needed
rch --help-json workers/list # machine-readable help for nested command paths
```
`--json` uses the standard API envelope for command output; `--format toon`
emits the same data as TOON. The legacy `rch --capabilities` flag remains a
raw JSON shortcut for lightweight discovery.
---
## Configuration
Primary files:
- User config: `~/.config/rch/config.toml`
- Worker list: `~/.config/rch/workers.toml`
- Project override: `.rch/config.toml`
- Optional project excludes: `.rchignore`
Precedence (highest first):
1. CLI flags
2. Environment variables
3. Profile defaults
4. `.env` / `.rch.env`
5. Project config
6. User config
7. Built-in defaults
### Canonical Project Root
By default rch expects projects to live under `/data/projects` (with the `/dp`
alias symlink). Both roots are configurable, so repos under `~/code`,
`/workspace`, etc. work without relocating anything:
```toml
[path_topology]
canonical_root = "/home/me/code" # default: /data/projects
alias_root = "/home/me/code" # default: /dp (set equal to canonical_root if you have no alias)
```
Or via the CLI / environment (env wins over TOML):
```bash
rch config set path_topology.canonical_root /home/me/code
rch config set path_topology.alias_root /home/me/code
# or
export RCH_CANONICAL_PROJECT_ROOT=/home/me/code
export RCH_ALIAS_PROJECT_ROOT=/home/me/code
```
Empty strings are treated as unset (defaults apply). See
`docs/guides/configuration.md` (`[path_topology]`) for details.
### Minimal Example
```toml
[general]
enabled = true
force_local = false
force_remote = false
socket_path = "~/.cache/rch/rch.sock"
log_level = "info"
[compilation]
confidence_threshold = 0.85
min_local_time_ms = 2000
remote_speedup_threshold = 1.2
build_slots = 4
test_slots = 8
check_slots = 2
build_timeout_sec = 300
test_timeout_sec = 1800
bun_timeout_sec = 600
external_timeout_enabled = true
[transfer]
compression_level = 3
# Staging base for the transfer paths that do NOT mirror the client's layout:
# `--clean-overlay` roots, Windows workers, and `rch cache warm`. An ordinary
# `rch exec` mirrors the project under the worker's `[path_topology]
# canonical_root` instead, with the pooled Cargo target store inside that
# mirror — see `[remediation.pooled_target] store_base` below to place those
# stores on another filesystem.
remote_base = "/data/tmp/rch"
# Optional per-attempt source-sync cap. When unset, the default is payload-aware:
# 30 seconds plus one second per MiB, capped at one hour.
# sync_timeout_ms = 120000
# Abort a source sync after this many seconds of NO rsync output (dead channel /
# wedged rsync); a progressing transfer is never affected. 0 disables.
# source_sync_silence_timeout_secs = 120
# Explicit rsync binary. Unset lets RCH prefer a modern (3.x) rsync over the
# stock macOS openrsync automatically; RCH_RSYNC_BIN overrides this.
# rsync_bin = "/opt/homebrew/bin/rsync"
adaptive_compression = true
verify_artifacts = false
max_transfer_mb = 2048
[remediation.pooled_target]
# Where pooled Cargo target stores are PLACED on the worker. Unset (the
# default) keeps them inside the project mirror, i.e. under the canonical
# root, on whatever filesystem holds it. Set an absolute path to move every
# pooled store to <store_base>/<project_id>/.rch-target-<worker>-pool-<key>,
# which is how a worker with a small root disk keeps multi-GB warm pools on a
# larger volume. `rch gc`, `rch cache status`, and the daemon sweep scan this
# root in addition to `remote_base` (with `reaper_max_cache_gb` applying per
# scan root). Ignored for Windows workers.
# store_base = "/bigdisk/rch-pools"
#
# Where the reaper SCANS for stale target dirs (not a placement setting).
# remote_base = "/data/projects"
#
# EXTRA roots `rch gc` scans, on top of the ones it derives itself
# (`remote_base`, `store_base`, and the worker temp base resolved exactly as
# the code that creates the dirs resolves it: $TMPDIR -> /data/tmp -> /tmp).
# Use this for a runtime root rch cannot know about, e.g. an operator build
# root on a mounted volume. Each entry must be absolute, `..`-free and free of
# shell metacharacters -- it is embedded in a remote command. `rch gc --root`
# adds one for a single run.
# gc_extra_roots = ["/mnt/big/rch"]
#
# Idle days before `rch gc` may collect a durable per-worker Cargo cache
# (`rch-cargo-cache-*`). 0 disables. Collection ALSO requires zero open file
# descriptors under the dir and no live process rooted at it; a gate that
# cannot be evaluated counts as "in use".
# gc_cargo_cache_idle_days = 14
[selection]
strategy = "balanced"
# Reserve this much free disk, then budget space for each concurrent slot (GiB).
min_free_gb = 10.0
disk_gb_per_slot = 10.0
[self_healing]
hook_starts_daemon = true
daemon_installs_hooks = true
[alerts]
enabled = true
suppress_duplicates_secs = 300
```
When the Unix launcher enforces a compilation deadline, RCH preserves exit 137
and reports that the configured deadline expired. It does not retry on a larger
worker or fall back locally. Machine-mode `rch exec` reports
`outcome: "deadline_exceeded"`. This classification requires evidence from the
launcher; exit 137 or elapsed time alone does not establish a timeout.
Unix artifact downloads estimate the files matched by the retrieval filters
before transferring them. Their default total retry budget grows with that size
(30 seconds plus one second per MiB, adjusted for a slower `bwlimit_kbps`, up to
one hour). An explicit `max_transfer_time_ms` keeps its existing hard ceiling
and skips estimation. Rsync also stops after 30 seconds without network I/O.
If estimation fails, RCH logs the failure and retains the configured retry
budget. A literal `cargo build --bin NAME` or `--example NAME` retrieves only
the named payloads, not the pooled target's other cached outputs. Anything the
narrowing cannot parse exactly (unknown flags, shell expansion, other target
kinds, response files, custom target specs) keeps the broad per-profile filters,
which can include other cached build outputs in the same profile.
Windows workers return artifacts through tar in both quiet and interactive
execution. The transfer honors the requested output patterns and cache exclusions,
including custom Cargo target directories and explicit compiler outputs outside
`target/`. RCH inventories regular files, downloads only selected paths, and
validates the complete archive before writing outputs locally. Remote tar errors,
missing selected files, and unexpected archive members fail artifact retrieval;
successful downloads include file counts and a manifest for output validation.
Inventory and transfer use bounded deadlines, and cancellation reaps the transport
before recovery can begin. This transport requires GNU `find` from the worker's
Git for Windows installation and native `tar`.
Artifact staging retains the original caller directory as the source-protection
reference. Recovery restores that reference for each output phase, so an empty
staging directory cannot make source files eligible for a broad output pattern.
Built-in worker selection defaults to `balanced`, which blends speed, load,
health, and cache affinity. Use `priority` only when you want explicit
worker-priority control, and `fair_fastest` when you want extra load spreading.
Worker capacity is capped at the configured `total_slots` and reduced to
`floor((free_disk_gb - min_free_gb) / disk_gb_per_slot)` as disk fills, with a
minimum of zero. Unknown disk telemetry keeps the configured ceiling. Existing
builds retain their reservations when capacity drops; new work waits for space.
`rch workers list` reports the effective total when the daemon is reachable.
Change the budget with `rch config set selection.disk_gb_per_slot 10` and restart
the daemon to apply it.
Balanced selection also favors free disk before a worker reaches the admission
floor. `selection.weights.disk` (default `0.2`, range `0` to `1`) adds disk
headroom credit after the other score adjustments. Equal ample headroom preserves
the existing worker order. Credit rises from zero at `min_free_gb` to full credit
at 25% free disk, with at least one slot's disk budget above the floor on small
disks. Unknown measurements receive full credit; the admission gate still handles
missing telemetry. Affinity pins and explicit worker requests retain their
existing behavior. Set the weight to zero to disable disk ranking independently
of disk admission and slot limits, then restart the daemon to apply the change.
`rch workers list` shows daemon-reported free disk and pressure. `rch workers probe`
shows fresh disk measurements alongside latency and labels any cached daemon
pressure separately. JSON and TOON include numeric free GiB and free ratios;
unavailable measurements remain `null` (displayed as `unknown` in text).
### Worker Config Example
```toml
[[workers]]
id = "css"
host = "203.0.113.20"
user = "ubuntu"
identity_file = "~/.ssh/id_rsa"
total_slots = 32
priority = 100
tags = ["fast", "ssd"]
# Verified named tools. Each entry is a FIXED argv the worker runs as the
# configured user; a zero exit marks the tool present, anything else marks it
# absent. Unlike `tags`, which are an unverified naming convention, these are
# evidence — which is what lets `--require-tool` gate worker selection.
tools = [
{ name = "clang", command = ["clang", "--version"] },
{ name = "ld.lld", command = ["/usr/bin/ld.lld", "--version"] },
]
```
Names must be ASCII letters, digits, `-`, `_`, `.` or `+`: a name carrying a
space or `=` could not survive the probe's fact format and would be read back
as a different name, so it is refused when the config loads rather than
becoming a capability lie. `rch workers capabilities --refresh` reports the
results under **Named tools** as `verified:` and `failed:` — a declared probe
that fails is a different operational state from one that was never declared,
and only the first points at a broken worker.
There is deliberately no general remote-shell probe API. A caller that could
ask a worker to run an arbitrary command "to check for a tool" would be a
remote execution primitive wearing a capability-probe hat; declaring the argv
in operator config keeps the set of probe commands finite and reviewable.
---
## Output Modes
RCH auto-selects output mode by context:
- `hook`: strict JSON for hook protocol
- `machine`: explicit machine output (`--json`, `--format`)
- `interactive`: rich terminal rendering
- `colored`: ANSI-only when forced without TTY
- `plain`: text fallback
Environment controls:
- `RCH_JSON=1`, `RCH_HOOK_MODE=1`
- `NO_COLOR=1`, `FORCE_COLOR=1`, `FORCE_COLOR=0`
- `RCH_OUTPUT_FORMAT=json|toon`, `TOON_DEFAULT_FORMAT`
JSON responses use a stable envelope (`api_version`, `timestamp`, `success`, `data`, `error`).
---
## Placement Controls
Worker placement, strict-remote, queue, wait-timeout, visibility, and target-dir
behavior are first-class, canonical controls — not folklore. The authoritative
list is discoverable at runtime (`rch capabilities --json`), and the resolved
plan for any command is shown by `rch diagnose <command> --json` under
`data.placement` (and in the human `Placement Controls` section):
| Requested worker | `RCH_WORKER` (`RCH_WORKERS`) | Request specific worker(s) by id. Still passes capability/admission checks; an inadmissible requested worker is **refused** with a stable `RCH-Innn` reason code and a next action — never silently swapped. |
| Requested profile | `RCH_PRESET` | Named execution profile (recorded as `requested_profile`). |
| Strict remote (fail-closed) | `RCH_REQUIRE_REMOTE` | Refuse local fallback (proof mode). Takes precedence over `RCH_FORCE_REMOTE`. |
| Force remote (fail-open) | `RCH_FORCE_REMOTE` | Always attempt offload (bypass local-time/speedup gating) but still fail open to local. Distinct from `RCH_REQUIRE_REMOTE`. |
| Queue when busy | `RCH_QUEUE_WHEN_BUSY` (default `1`) | Wait for a busy worker instead of falling back to local. Set `0` to disable. |
| Wait timeout | `RCH_DAEMON_WAIT_RESPONSE_TIMEOUT_SECS` (`RCH_DAEMON_RESPONSE_TIMEOUT_SECS`) | Max seconds to wait for a queued worker. |
| Visibility | `RCH_VISIBILITY=none\|summary\|verbose` (`RCH_QUIET`, `RCH_VERBOSE`) | Hook output verbosity. |
| Target dir | `RCH_DISABLE_TARGET_REUSE` | Legacy unique-per-job remote target dir instead of the pooled, reuse-friendly dir. |
| Source-sync timeout | `RCH_SYNC_TIMEOUT_MS` | Per-attempt source-upload timeout in milliseconds (1000..=3600000). Unset uses the payload-aware default. This does not change remote Cargo or artifact-return timeouts. |
The resolved plan reports `requested_worker`, `requested_profile`,
`effective_worker`, `strict_remote_policy`, `queue_policy`, `visibility_mode`,
`wait_timeout_ms`, `target_dir_policy`, the requested-worker admissibility
outcome, and a `diagnostics` list. Any control value that cannot be applied as
written (an unrecognized value or a superseded alias) surfaces a diagnostic
rather than being silently ignored.
---
## Monitoring and Observability
RCH exposes observability through daemon APIs and metrics:
- daemon health/readiness endpoints
- Prometheus metrics collection
- opt-in OTLP export of daemon request-duration metrics
- telemetry-backed worker SpeedScore history
- queue/build history, active alerts, cancellation metadata in status APIs
- an opt-in **tailnet status API** (`[api] bind = "tailscale"`): the daemon's
full `/status` JSON over TCP, bearer-token gated, bound only to loopback or
Tailscale addresses — so agents on other machines can ask a dispatcher "why
are you local-only right now?" without ssh
- the **fleet dashboard** under `dashboard/` (encrypted static console + an
agent endpoint `GET /api/fleet?view=problems|diagnose|help`); `rch web` opens it
Enable OTLP metrics in the daemon's environment with `RCH_OTEL_ENABLED=1`
and `RCH_OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317` (OTLP/gRPC).
The endpoint falls back to `OTEL_EXPORTER_OTLP_ENDPOINT`. `OTEL_SERVICE_NAME`
defaults to `rchd`; `RCH_OTEL_EXPORT_INTERVAL_SECS` defaults to 30.
`rch_request_duration_seconds{entrypoint="rchd_api"}` records finite Unix-socket
request handling after parsing, including errors and cancellation. Event streams
and the separate HTTP API are excluded. The same observations reach Prometheus,
including when OTLP is disabled. Shutdown drains request tasks before a
best-effort final export. The banner reports exporter configuration, not collector
connectivity.
The same opt-in settings enable direct OTLP verdict export for single-shot
`rch doctor --reliability`, including under `--quiet`. Its service name defaults
to `rch`. Each report emits one `rch_doctor_verdict_total` observation, labeled
with the actual verdict and scope (combined scopes use `other`), and flushes
before returning its original exit code. Reports whose `daemon_unreachable`
flag is true also increment `rch_doctor_daemon_unreachable_total` once, even
when several diagnostics report the missing daemon. The asynchronous
`daemon_status` and `repo_convergence` probes also emit
`rch_doctor_probe_duration_seconds{probe,result}` once on completion, RPC error,
or timeout. Durations measure each probe's execution, not time spent waiting
to join other probes. Skipped probes emit nothing; panic/cancellation and
helper/ownership probe durations are not yet exported. With `--fix`, each
remediation step also emits `rch_doctor_fix_steps_total` and
`rch_doctor_fix_duration_seconds`, labeled by its actual outcome: `applied`,
`already_satisfied`, `would_apply`, `manual`, or `failed`. Dry-run previews
remain `would_apply`; they are never counted as applied changes. Durations
measure step execution, excluding the initial shared configuration load.
These CLI metrics are not forwarded to the daemon's Prometheus endpoint.
Hook/watch telemetry, remaining doctor event metrics, and trace/log export
remain unfinished.
Quick checks:
```bash
rch status --workers --jobs
rch speedscore --all
rch doctor --json
```
On Linux boxes configured with `[general] role = "dispatcher"`, status and doctor
also report unmanaged local Cargo/rustc processes. The warning remains visible
while those processes run, including when the daemon is unavailable. Check
`rch shim status`, PATH order, and absolute-path toolchain Cargo invocations.
Processes carrying `RCH_CARGO_WRAPPER_BYPASS=1` or descended from `rch` are excluded.
An alarm event is logged once per observed episode and clears after a successful
scan finds no unmanaged compiler processes. Status and doctor share a locked
one-byte latch at `$XDG_CACHE_HOME/rch/local-build-alarm` (normally
`~/.cache/rch/local-build-alarm`). Scan failures preserve the previous episode;
latch failures are reported without hiding current builds. Reliability doctor
includes this check in the `triage` scope and its default `all` scope.
---
## Testing and Validation
Workspace checks:
```bash
cargo fmt --check
cargo check --workspace --all-targets
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
```
Reliability and E2E suites are provided under:
- `tests/`
- `tests/e2e/`
- `rch-common/tests/` (contract/reliability/perf suites)
If you are running CPU-intensive validation manually and want explicit offload:
```bash
rch exec -- cargo check --workspace --all-targets
rch exec -- cargo test --workspace
rch exec -- cargo clippy --workspace --all-targets -- -D warnings
```
When local fallback is not acceptable, set `RCH_REQUIRE_REMOTE=1` on the
`rch exec` process and keep the build command as direct argv:
```bash
RCH_REQUIRE_REMOTE=1 rch exec -- cargo test --workspace
RCH_REQUIRE_REMOTE=1 rch exec -- cargo clippy --workspace --all-targets -- -D warnings
```
### Exact source-content receipts
Proof-oriented callers can require a single-worker, fail-closed source transfer
and retain the exact regular-file bytes admitted on that worker:
```bash
RCH_REQUIRE_REMOTE=1 rch exec --source-content-receipt -- \
cargo test --locked --workspace
```
Receipt mode resolves the active Cargo path-dependency closure, gives every
root an invocation-unique worker path, transfers with checksum comparison, and
reopens every selected file before and after the command. The emitted
`rch.source_content_receipt.v1` JSON binds the worker and build IDs, exact
command digest and exit code, transfer filter policy, per-root file paths,
lengths, executable bits, SHA-256 digests, and content roots. Any unproved
dependency closure, transfer delta, worker verification error, local lasting
change, retry, or local fallback refuses the invocation instead of emitting a
receipt. Receipt mode currently requires the Unix rsync transport and cannot be
combined with clean-overlay mode.
The receipt is emitted after the remote command and its post-command source
barrier, but before ordinary artifact retrieval completes. Artifact-grade
callers must therefore retain both the receipt and the later successful
retrieval/terminal transcript; the receipt alone is not proof that build
artifacts reached the caller. A recursive local mutation watcher is still
needed when transient edit-and-restore (ABA) detection is part of the caller's
source-stability contract.
### Clean Git overlays for shared working trees
`rch exec` can build an immutable committed tree plus an explicit, repeatable
set of local paths without transferring unrelated working-tree changes:
```bash
RCH_REQUIRE_REMOTE=1 rch exec \
--base HEAD \
--clean-overlay \
--overlay-path src/lib.rs \
--overlay-path tests/focused.rs \
-- cargo test --test focused
RCH_REQUIRE_REMOTE=1 rch exec \
--base HEAD \
--clean-overlay \
--no-overlay \
-- cargo check --workspace --all-targets
```
The client resolves `--base` to a commit object, streams that Git archive into
a fresh isolated worker path, and then uploads only the literal repository-
relative `--overlay-path` selections. Modified and untracked files are
supported; selected deletions, absolute/traversing paths, Git metadata,
non-ASCII/backslash/control-character paths, case-only or otherwise ambiguous
filesystem spellings, overlay symlinks, empty overlay directories, submodules,
and Git archive `export-ignore`/`export-subst` attributes fail closed. Exactly
one of one-or-more `--overlay-path` options or `--no-overlay` is required.
Clean-overlay mode also implies remote-only execution even when
`RCH_REQUIRE_REMOTE` was omitted, so a worker or transfer failure never falls
back to the ambient local tree.
Explicit clean-overlay execution also admits the read-only `cargo fmt --check`
diagnostic, which the ordinary interception classifier intentionally leaves
local.
On Unix workers, pooled Cargo builds use a stable source path paired with a
versioned target cache. Jobs sharing that pair wait until the preceding job
has finished execution, artifact retrieval, and source retirement. Different
pairs remain independent. Source timestamps are refreshed beyond cached
artifacts so changed archive or overlay contents cannot appear older than a
previous build; registry and Git dependency caches remain reusable.
Clean-overlay execution also pins Cargo's `build.build-dir` to its assigned
target directory, overriding ambient and command-line Cargo configuration.
Explicit `--build-dir` options are refused. This placement applies to isolated
builds too, so disabling target reuse cannot inherit shared intermediates.
Managed placement currently supports built-in Cargo build commands and Clippy.
Other compiling Cargo subcommands, including Nextest and Zigbuild, are refused
in clean-overlay mode.
An interrupted pair retains its ownership marker and refuses further reuse,
even if its SSH lock connection disappears. Inspect the prior job before
recovering that pair, or use `RCH_DISABLE_TARGET_REUSE=1` for a fresh isolated
build. Non-Cargo Windows clean-overlay jobs also use isolated targets. The ownership
marker protects against process interruption; it is not a power-loss durability
guarantee.
Clean-overlay materializes the primary Git repository and explicitly bound
sibling repositories. In-repository Cargo workspace members are present in the
archive. Before Cargo starts, selected
manifests, configuration, symlinks, and command-line path overrides are checked
against the selected Git base and overlays. Escaping paths are refused with
`RCH-E413`; retained sibling directories on the worker cannot supply those
dependencies. To include a sibling Git repository, bind its revision explicitly:
```bash
rch exec --base HEAD --clean-overlay --no-overlay \
--dependency-base ../dep=HEAD -- cargo test
```
Repeat `--dependency-base PATH=REV` for every external root, including transitive
dependencies. Each revision resolves to a commit before worker selection; dirty
files in those repositories are excluded. The selected repositories must be
siblings of the primary Git root. They are staged under one owned container,
preserving `../dep` references, and the execution receipt names each commit,
Git tree hash, and remote path. All roots share the same execution lease and
are retired together. Cross-repository symlinks and non-sibling layouts are
currently refused.
This check conservatively covers all selected manifests, including inactive
fixtures. File-based `--config`, configuration includes, and changes to Cargo's
working directory are refused. On the worker, any Cargo-home or ancestor
configuration is refused before Cargo starts, even if that configuration would
be harmless. Cargo clean-overlay jobs currently require a Unix worker for this
check. These checks do not sandbox arbitrary file reads by build scripts.
The client re-fingerprints overlays after upload and refuses execution if their
contents changed during admission or transfer.
Do not batch several Cargo commands behind a shell wrapper such as
`rch exec -- bash -lc "cargo test ... && cargo test ..."`. Shell-wrapped
commands are classified as non-compilation for hook safety. Under
`RCH_REQUIRE_REMOTE=1`, RCH refuses that local fallback before executing the
shell and reports `RCH-E301`; without the env var, ordinary non-compilation
commands may still run locally. For several focused checks, run separate direct
`RCH_REQUIRE_REMOTE=1 rch exec -- cargo ...` invocations.
---
## Security Model
- Transport uses SSH.
- Worker commands are constrained to classified execution paths.
- Sensitive field masking and structured error taxonomy are built in.
- `rch update` enforces the release's SHA-256 checksums and verifies the
archive's `.minisig` against a public key pinned in the binary (minisign key
`69B3955C8D2E62A8`, which signs every release since 2.0.0); a bad signature
aborts the update. A Sigstore bundle is also verified when the release ships
one and `cosign` is installed. `--skip-verify` bypasses signature checks only.
`install.sh` still verifies checksums only.
- Hook path remains fail-open to avoid deadlocks/stalls.
Operational recommendations:
1. Use workers you control.
2. Use dedicated SSH keys for worker access.
3. Keep workers patched and isolated.
4. Enable telemetry/alerting for production-like use.
---
## Limitations
- Designed around SSH-based Linux worker environments.
- Tooling assumptions are strongest for Rust and selected build/test commands.
- Remote performance gains depend on network + worker capacity + project shape.
- The fleet dashboard (`dashboard/`) needs Node and a static host (Vercel/Pages); its live path needs a Vercel Blob store and `rchd`'s `[api]` listener on a tailnet.
---
## FAQ
### Does RCH block my command if the daemon/workers fail?
No. It fails open and allows local execution.
### Can I force local or force remote per project?
Yes, via `.rch/config.toml` (`general.force_local` / `general.force_remote`).
### Is queue/cancel supported?
Yes. Use `rch queue` and `rch cancel`.
### Can I inspect why a command is or is not intercepted?
Yes. Use `rch diagnose "<command>"`.
---
## About Contributions
Please don't take this the wrong way, but I do not accept outside contributions for my projects. You can still open issues and PRs for discussion/proof-of-fix, but I review and re-implement changes independently.
---
## License
MIT License **with an OpenAI/Anthropic rider** — this is not ordinary
OSI MIT: no rights are granted to the restricted parties named in the
rider. See [LICENSE](LICENSE) for the exact, controlling terms
(SPDX: `LicenseRef-MIT-OpenAI-Anthropic-Rider`).