forge-ops-tracker 0.8.0

Rust error reporting client for a ForgeOps instance.
Documentation
# forge-ops-tracker

Rust error reporting client for a [ForgeOps](../../) instance.
Requires Rust 1.81+. It captures panics on any thread and explicitly reported errors, builds a
backtrace, scrubs likely PII, and delivers events to ForgeOps over HTTP without blocking whatever
raised them.

## Installation

```toml
[dependencies]
forge-ops-tracker = "0.2.0"
```

### Dependencies

Rust's standard library has no HTTP client, no structured per-frame backtrace access on stable, and
no regular expression engine. This crate depends on exactly the three widely-used crates that fill
those three specific gaps, each doing something `std` genuinely can't:

- **[`ureq`]https://crates.io/crates/ureq**: delivers events over HTTPS. `std` has no HTTP
  client at all.
- **[`backtrace`]https://crates.io/crates/backtrace**: structured per-frame file/line/function
  data. `std::backtrace::Backtrace` captures a trace on stable Rust, but only exposes it as a
  formatted string (no public per-frame accessors) so it can't produce the structured
  file/line/method/in_app shape this client needs to build an event payload.
- **[`regex`]https://crates.io/crates/regex**: the PII scrubber's pattern matching. `std` has
  no regular expression engine.

Everything else (the event payload's own JSON encoding, a DSN parser, timestamp formatting) is
hand-rolled rather than reaching for `serde_json`/`url`/`chrono`: this crate only depends on
something outside `std` when the language truly can't do it itself.

## Configuration

Set a DSN (from a project's settings page in ForgeOps), either via the `FORGE_OPS_DSN` environment
variable or explicitly:

```rust
forge_ops_tracker::init(|c| {
    c.dsn = Some("https://<api_key>@your-forgeops-host/api/v1/events".to_string()); // or leave unset to read FORGE_OPS_DSN
    c.release = Some("...".to_string());
    c.environment = "production".to_string();
});
```

Call `init` once at startup, before your server starts accepting requests. Pass a closure to set
any `Configuration` field, so every option is available through the one call without a long list
of positional arguments or a separate setter for each field.

## What gets reported automatically, and what doesn't

**A panic on any thread needs no further wiring at all**, once `init()` has run. Rust's panic hook
(`std::panic::set_hook`) is process-wide: it fires for a panic on *any* thread, including a web
framework's own worker threads (Actix, Axum/Tokio, a plain `std::thread`), with **no per-framework
middleware needed at all**.

`init()` installs this hook automatically unless `Configuration.install_panic_hook` is set to
`false`. It never swallows the panic: after reporting, it calls whatever hook was previously
installed (Rust's own default, which prints to stderr, unless something else already replaced it),
so reporting a panic never changes what your program actually does afterward. Call
`install_panic_hook()` directly only if you're managing configuration some other way than
`init()`.

**Rust doesn't have exceptions**, so an error your own code already caught (a `Result::Err`) is a
second, separate case: report it explicitly, right at the point you'd otherwise just log it:

```rust
if let Err(err) = charge_card(&order) {
    forge_ops_tracker::capture_error(&err, forge_ops_tracker::context!{"order_id" => order.id}, None);
    return Err(err);
}
```

Or, more concisely, via the `ResultReportExt` extension trait, which reports on `Err` and passes
the `Result` through unchanged:

```rust
use forge_ops_tracker::ResultReportExt;

charge_card(&order).report_err(forge_ops_tracker::context!{"order_id" => order.id}, None)?;
```

A plain Rust `std::error::Error` carries no stack trace of its own, so `capture_error` captures the
backtrace at its own call site. Call it as close to the point you learned about the error as you
reasonably can, for the most useful trace.

`exception_class` is inferred via [`std::any::type_name`], which needs a concrete, statically-known
error type: for a `Box<dyn Error>` or other trait object, where that isn't possible, use
`capture_error_with_class(class, err, context, user)` instead and supply the class yourself.

[`std::any::type_name`]: https://doc.rust-lang.org/std/any/fn.type_name.html

Delivery happens on a background thread with a bounded channel and a short per-request HTTP
timeout (`Configuration.timeout`, 2s default). Every failure mode: network errors, timeouts, a
full queue, a malformed DSN: is caught and dropped rather than propagated, so a broken or
unreachable tracker can never take down the host app. The worker thread starts eagerly, at `init()`
time, rather than waiting for the first push: Rust programs essentially never fork themselves at
the application level after startup, so there's no risk of an eagerly-started thread being left
dead in a forked child, and starting it up front means it's ready before the first event needs to
be delivered.

## Identifying users

```rust
forge_ops_tracker::capture_error(&err, forge_ops_tracker::context!{}, Some(forge_ops_tracker::context!{"id" => user.id, "email" => user.email}));
```

Or `set_user` to attach it to every subsequently reported error on this thread (an explicit
`capture_error`/`capture_error_with_class` call, or a panic the installed hook catches) until
changed or cleared, rather than passing it to every call by hand, e.g. right after authenticating
a request:

```rust
forge_ops_tracker::set_user(forge_ops_tracker::context!{"id" => user.id, "email" => user.email});
// once the request is done, or on sign-out:
forge_ops_tracker::set_user(std::collections::HashMap::new());
```

There's no way to automatically detect "the current user" the way a server-side web framework with
its own session/auth middleware can, so this is always manual. `set_user` is a **plain
thread-local**, not a process-wide global: the right choice for the thread-per-request model this
crate's own synchronous, non-async design naturally pairs with (see "Dependencies" below for why
`ureq`, not an async HTTP client, was chosen). **It does not propagate across an `.await` in an
async runtime**: unlike a native OS thread, a single thread in an async executor (Tokio,
async-std) interleaves multiple unrelated tasks, so a value set on one task can leak into, or never
reach, another. A host app built on an async runtime should pass `user` explicitly to
`capture_error`/`capture_error_with_class` on every call instead of relying on `set_user`, the same
reason `sdks/node` needs `AsyncLocalStorage` rather than a bare thread-local. `id`/`email`/
`username` keys are all independently optional. Shows up on an issue's own detail page, and as its
own affected-users count alongside the regular event count.

## Breadcrumbs

A bounded, ordered trail of what happened right before an error: on by default, capped at the 30
most recent entries per thread, both configurable:

```rust
forge_ops_tracker::init(|c| {
    c.track_breadcrumbs = false; // opt out entirely
    c.max_breadcrumbs = 50;      // default 30
});
```

```rust
forge_ops_tracker::add_breadcrumb("charged card", "custom", "info", forge_ops_tracker::context!{"order_id" => order.id});
```

`capture_error`/`capture_error_with_class`/the installed panic hook all attach the current
thread's trail automatically, the same way they already read `set_user`; there's no separate
argument to pass it through by hand.

**A plain thread-local**, the same choice `set_user` already made and for the identical reason
(see "Identifying users" above): the right fit for this crate's synchronous, thread-per-request
design, not an async runtime, where a single OS thread can interleave multiple unrelated tasks.

This crate has no web framework integration of its own, unlike `sdks/go`'s net/http/Gin
middleware or `gems/forge_ops_tracker`'s Rack middleware, so there's no automatic breadcrumb
source (no request/controller timing to record one from) and no middleware to start a fresh trail
per request on its own. Call `clear_breadcrumbs()` yourself at the start of each request, the same
place you'd already be calling `set_user` (or clearing it) from:

```rust
forge_ops_tracker::clear_breadcrumbs();
```

Without this, a synchronous, thread-pool-based server (actix-web's own worker threads, for
instance) would let one request's trail bleed into the next one handled on the same reused thread.

## Performance monitoring

Times whatever you wrap and reports one small aggregate per transaction (how many times it ran,
total and maximum duration) every `performance_flush_interval` (60s by default), for the Performance
page's per-transaction table. Not one network call per timed call.

Each aggregate also carries a small latency histogram (a count per fixed latency bucket: 50, 100,
250, 500, 1000, 2500, 5000 and 10000ms, plus an overflow bucket), so ForgeOps can show an
approximate p50/p95/p99 per transaction, not just an average. Percentiles are accurate to the width
of whichever bucket a duration falls into; the SDK never stores the individual durations.

```rust
// Wrap a whole request handler, or any block you want on the Performance page:
let response = forge_ops_tracker::time_transaction("GET /users/:id", || handle_request(req));

// Or record a duration you measured yourself:
forge_ops_tracker::record_performance("nightly-export", elapsed.as_secs_f64() * 1000.0);
```

This crate has no web framework integration (unlike `sdks/go`'s net/http and Gin middleware), so
**nothing is timed automatically**: you choose what to wrap. Keep transaction names low-cardinality
(`"GET /users/:id"`, not `"GET /users/42"`): every distinct name is its own row. `time_transaction`
records even if the closure panics. Turn it off with `track_performance = false`; it also does
nothing (and starts no thread) when reporting isn't enabled for the current environment.

The flush thread is a daemon: Rust has no equivalent of the Ruby gem's `at_exit`, so it does not run
on a normal process exit. A short-lived program, or one about to shut down, should call
`forge_ops_tracker::flush_performance()` itself to send the last partial window.

A failed delivery keeps every tally, so the next flush's window just grows. What a flush delivered
is *subtracted* from the tallies afterward, never the whole map cleared: a `record_performance` call
that lands while a delivery is in flight would otherwise be silently discarded (a real bug `sdks/go`
had and fixed; `gems/forge_ops_tracker`'s reference implementation still has it). A deterministic
test pins this.

## Distributed tracing

A slow call's own breakdown: which database calls, HTTP calls, or pieces of your code the time went
to, shown as a span tree on ForgeOps. Wrap the unit of work in `trace`, and anything inside it, on
the same thread, can add spans; the trace is sent only when the whole thing took at least
`trace_capture_threshold` (1 second by default), so fast calls cost nothing on the wire. Traces are
per service; nothing is propagated across services.

```rust
let response = forge_ops_tracker::trace("GET /checkout", || {
    let order = forge_ops_tracker::span("load order", "database", context! {"order_id" => 42}, || repo.find(42));
    forge_ops_tracker::span("charge card", "service", HashMap::new(), || gateway.charge(&order));
    render(&order)
});

// Something you timed yourself (kind is one of controller/service/database/redis/http/job/other;
// anything else is sent as "other"):
forge_ops_tracker::record_span("SELECT orders", "database", started_at, duration_ms, HashMap::new());
```

This crate has no web framework integration, so **nothing starts a trace or records a span
automatically**: you choose what to wrap. `span` nests under whichever span is open on the same
thread, records even when the closure panics, and just runs the closure outside a `trace`; a
`trace` inside another `trace` records a span instead. The open trace is a `thread_local!`, like the
breadcrumb trail, so it follows a thread, not an async task. A trace holds at most 500 spans and is
delivered on its own background thread and bounded queue, dropping rather than blocking when full.
Turn the whole feature off with `track_tracing = false`.

## Custom metrics and infrastructure monitoring

Two explicit calls (nothing is automatic, so there is no `track_metrics` flag): a business event you
name yourself, and a reading from one of your own hosts.

```rust
forge_ops_tracker::capture_metric("signup", 1.0);   // a bare counter
forge_ops_tracker::capture_metric("payment", 49.0); // a real magnitude; it may be negative (a refund)

forge_ops_tracker::capture_infrastructure_metric("cpu", 0.42, None); // None defaults to server_name
forge_ops_tracker::capture_infrastructure_metric("disk", 0.81, Some("db-1"));
forge_ops_tracker::flush_metrics(); // send right now
```

Each capture is buffered and flushed as one batch every `metric_flush_interval` /
`infrastructure_metric_flush_interval` (60 seconds by default) on a daemon thread started on the first
capture. **Rust has no exit hook to flush from**, so a short-lived program (a cron job) must call
`flush_metrics()` before it returns from `main`. Every entry is stored as it was captured (a signup is
a row, not a running total), so a count or sum you compute later is exact. Both are a no-op when the
client isn't enabled for the environment.

A failed delivery keeps every entry for the next flush, and an entry captured while a delivery is in
flight is kept too (the Ruby gem's own buffer loses it; a test pins this with a gated delivery). The
buffer holds at most 1000 entries per kind and drops further ones until a flush succeeds, since a plan
without the feature rejects every flush and would otherwise grow it for as long as the process lives.
A NaN or infinite value is dropped at capture: it is not valid JSON and would make the server reject the
whole batch behind it. Requires a ForgeOps plan that includes custom metrics / infrastructure
monitoring.

## `in_app` backtrace frames

A Rust binary built with debug info embeds the real build-time source paths, so file-path matching
against `Configuration.app_root` is a straightforward prefix comparison against those embedded
paths. Defaults to the current working directory; set it explicitly if that doesn't match your
binary's actual build layout. Third-party crate source under Cargo's registry cache and the Rust
toolchain's own std/core source are never marked `in_app`, regardless of `app_root`.

## Source context

By default, each in-app backtrace frame (never a third-party crate) is captured along with the 5
lines of source on either side of the culprit line, read straight off disk at capture time, so an
issue's detail page can show the actual code that broke, not just a file:line:method reference.
This never applies to a frame outside your own app's code, and it fails silently (no context, not
a panic) for any file that can't be read for whatever reason, e.g. built in a container image the
running host doesn't have access to.

This is a real, deliberate exception to "off by default is safer": literal source code is being
transmitted, not just a reference to it, and the real protection here isn't this field. Every
project on ForgeOps has its own setting (on by default, off durably and immediately once an org
owner turns it off, regardless of what any individual app's own local `Configuration` is still set
to) that governs whether the server will ever actually store what a client sends. Set
`capture_source_context` to `false` if you'd rather this crate never even attempt the disk read in
the first place:

```rust
forge_ops_tracker::init(|c| {
    c.capture_source_context = false;
});
```

## PII scrubbing

By default, the message, backtrace, and any context you attach are scanned for likely personal
data (email addresses, formatted SSNs/credit cards, known API key/token formats, and anything
under a suspiciously-named key like `password`, `api_key`, or `ssn`) and redacted before
the payload ever leaves this process. ForgeOps itself scrubs again on arrival regardless, so this
is a second, earlier layer, not the only one. The user attached via `capture_error`'s third
argument or `set_user` above is a deliberate exception: it's never scrubbed, since redacting it
would defeat the whole point of identifying users in the first place.

To disable it:

```rust
forge_ops_tracker::init(|c| {
    c.scrub_pii = false;
});
```

## Running the tests

```bash
cd sdks/rust
cargo test
cargo clippy --all-targets: -D warnings
cargo fmt: --check
```