Skip to main content

Module session_signals

Module session_signals 

Source
Expand description

Session-wide signal handling — the THREE-LEVEL shutdown ladder.

A Ctrl-C (SIGINT or SIGTERM, or the raw-mode key-watcher’s translated keystroke) advances one rung at a time:

  • Level 1 — graceful, cooperative. The session-stop flag is set; active fiber loops observe it at their cycle boundary and exit cleanly; end-of-run cleanup runs in the normal order (profiler flush, cadence reporter shutdown, metrics.db WAL consolidation, summary writes). A visible 10-second countdown starts: if the drain hasn’t finished when it expires, the ladder advances to level 2 automatically.
  • Level 2 — cancel in-flight ops, keep process cleanup. Entered by the countdown expiring or a second Ctrl-C. Ops parked inside a hung adapter call (a request that will only ever end by client timeout) are CANCELLED — their futures are dropped at the fiber’s dispatch point — so the drain completes and the process-level graceful shutdown (WAL consolidation, summaries) still runs. This is the rung that used to not exist: previously a second Ctrl-C hard-exited, skipping WAL cleanup exactly when hung ops made it matter.
  • Level 3 — force-exit. A further Ctrl-C once level 2 is in force exits immediately (process::exit(130)); metrics and profiler output may be incomplete.

The state is intentionally global: there is one session per process by construction. Tests shouldn’t need to install or consult it (no RunObserver test sets up signals).

§SRD-93 M7 — the detached-console signal contract

The ladder MUST be reachable from outside the process even when the async runtime is wedged (SRD-93 A8: the stop door never depends on runtime health — the 2026-08-03 incident proved a tokio-task watcher is unreachable exactly when it matters). So signal watching runs on a dedicated OS thread: main calls block_shutdown_signals before any thread spawns (every later thread inherits the block) and spawn_signal_dispatcher owns the signals via sigwait, advancing the ladder synchronously.

  • SIGINT / SIGTERM — one rung per signal; past the cancel rung the dispatcher force-exits 128 + signo (130 / 143) on its own thread, so the hard floor works even when nothing else does. Unarmed (no run in progress) they exit 128 + signo directly — CLI commands keep their default interrupt semantics.
  • SIGHUP — NEVER a stop. Armed: console-loss (log, run the set_console_loss_hook if installed, continue headless). Unarmed: exit 129.
  • SIGQUIT — diagnostics dump (ladder level, stop flags, plus the set_diag_dump_hook inventory if installed); no state change. Unarmed: exit 131.

The tokio ctrl_c watcher in install_signal_handler remains as a redundant door for embeddings whose main never blocked the signals; when the dispatcher owns them, that watcher simply never fires (the two are mutually exclusive by construction).

Structs§

StopView
SRD-92 Step 0 — one cooperative-stop view, consulted at every boundary. Bundles the per-execution stop sources so a boundary check is a single call, and so a unit that previously held only ONE flag (the while: wrapper held only the activity stop_flag) observes ALL of them. The global / per-execution session stop (stop_requested) is always folded in; a fail-effect global (fault_stop_requested) is folded into StopView::poll.

Enums§

ShutdownOrigin
What advanced the shutdown ladder — used only to phrase the level-1 announcement accurately. A programmatic action: abort trip drives the SAME ladder as Ctrl-C, so it must not be reported as a Ctrl-C.
StopCause
Why a shell-driven stop halted the walk (SRD-82 Part 4). A StopCause::Fault is a fail-effect trip — a child phase failed, so the run’s validity is Failed and the halt records Interrupted + Failed. A StopCause::Interrupt is a clean stop (a graceful condition or user Ctrl-C) — later phases are deliberately skipped and the result is re-usable.

Functions§

abort_shutdown
action: abort (SRD-83 follow-up) — jump the shutdown ladder STRAIGHT to the cancel-ops rung (level 2), skipping the level-1 cooperative-drain countdown. A stop driven by errors must not wait for in-flight ops or remaining phases: their futures are dropped at the fiber dispatch point NOW, the walk unwinds, and control passes straight to the runner’s graceful SESSION shutdown. Cleanup still runs (metrics flush, WAL consolidation, summaries via the RAII guard) — this is NOT a force-exit; a further Ctrl-C is. Also raises the global session stop so the walk halts and no new phase starts. Idempotent; a no-op if the ladder is already at level ≥ 2. Returns the level now in force.
block_shutdown_signals
Block the dispatched signals in the calling thread. MUST run first thing in main, before any thread spawns, so every later thread inherits the block and delivery can only land in the dispatcher’s sigwait. This also makes the ladder immune to a third-party library later installing its own sigaction for these signals — sigwait consumes a blocked signal before any handler would run. (Child processes are unaffected: std::process::Command resets the child’s signal mask and dispositions.)
cancel_ops_requested
True once level 2 has been entered — in-flight ops should cancel.
escalate_shutdown
Advance the ladder ONE rung: 0 → 1 (graceful + countdown), 1 → 2 (cancel in-flight ops). At ≥ 2 this is a no-op returning the current level — the FORCE-EXIT decision (level 3) stays with the caller, which owns its own terminal hygiene (the raw-mode key-watcher must restore the terminal before exiting; the SIGINT watcher just exits). origin phrases the level-1 line only. Returns the level now in force.
fault_stop_requested
True once a fail-effect stop condition has halted the walk.
graceful_stop_requested
True once a stop condition has gracefully halted the walk.
install_signal_handler
Install a tokio task that watches ctrl_c() and drives SIGINT through the three-level ladder described in the module doc. Idempotent — only the first call wins; subsequent calls are no-ops. Must be called from inside a tokio runtime context.
mark_shutdown_complete
Mark the runner’s process-level shutdown complete: the countdown (if still running) goes quiet, and further escalations are moot.
ops_cancelled
Resolves when the cancel rung (level ≥ 2) is in force. Ready immediately if it already is. Used in a select! against the op future at the fiber’s dispatch point — dropping the op future is the cancellation.
request_fault_stop
Record that a fail-effect stop condition halted the remaining walk (SRD-82 Part 4, StopCause::Fault). Idempotent.
request_graceful_stop
Record that a stop condition gracefully halted the remaining walk (SRD-83 workload shell). Idempotent.
request_shell_stop
Route a shell stop to the right session signal by its StopCause. Both halt the tail; Fault additionally marks the run failed (non-zero exit via the failing phase’s Err).
request_stop
Programmatically request a session-wide stop. Used by the signal handler, but also available to other lifecycle code that wants to short-circuit the run.
set_console_loss_hook
Install the console-loss hook. First call wins (one console).
set_diag_dump_hook
Install the diagnostics-dump hook. First call wins.
shutdown_level
The ladder level currently in force.
spawn_signal_dispatcher
Spawn the dedicated dispatcher thread (SRD-93 M7 / A8). Pair with block_shutdown_signals — without the block, sigwait never receives anything and the legacy tokio watcher keeps the door. Idempotent.
stop_requested
Returns true once a stop has been requested for the current execution. Cheap relaxed atomic load — safe to call from a hot fiber loop.
subscribe_shutdown
Subscribe to ladder-level changes. Each fiber holds one receiver and races its op dispatch against ops_cancelled.