pub struct SnapshotManagerConfig {
pub principals: TrackedPrincipals,
pub refresh_interval: Duration,
pub unknown_ttl: SignedDuration,
pub revoked_ttl: SignedDuration,
pub retry_backoff: Duration,
pub max_concurrent_fetches: usize,
pub fetch_timeout: Duration,
pub enumeration_timeout: Duration,
}Expand description
Which principals to serve and how to keep their snapshots fresh.
There are no defaults; every value is a deployment decision.
validate runs before the task starts
(INVARIANTS.md 16). docs/SNAPSHOT_OPERATIONS.md covers operating it.
Fields§
§principals: TrackedPrincipalsThe principals this instance serves — a fixed list, or everything the source knows (GL-48).
Also decides the readiness rule: Fixed needs every listed principal
resolved, All needs at least one. Neither a Fixed list nor an
All seed may contain duplicates, and a Fixed list must fit the
snapshot map’s generation capacity.
refresh_interval: DurationFull refetch cadence — also the revocation propagation bound.
Each refresh re-fetches every tracked principal (and, with All,
re-enumerates the catalogue), which is how a withdrawn principal is
removed. Too long lets a revoked principal keep admitting for up to
this long, bounded also by its snapshot’s valid_until; too short
multiplies source load by the tracked set on every instance. Keep it
well below snapshot validity so a live principal is refreshed before
its snapshot expires. Must be positive.
unknown_ttl: SignedDurationHow long a principal the source returned no row for stays negative before the manager rechecks it.
Short, and it covers every absence rather than only new principals: a signup may be in flight, and a source that is rebuilding, failing over, or serving a lagging replica reports a principal it has served for years as absent too. This is the ceiling on how long such a gap can deny a live customer, so it is an availability bound, not just onboarding latency.
Too long denies a customer whose row was briefly missing for that long; too short rechecks every absent tracked principal more often, one source fetch each on every instance. Must be positive.
revoked_ttl: SignedDurationHow long a published revocation tombstone stays negative before the manager rechecks it (GL-52).
Long: coming back means an operator reinstated the account, which is rare, and a catalogue accumulates these forever — every cancelled customer is one, and each recheck is a fetch on every instance.
This bounds reinstatement, never revocation. A live principal is
always swept, so withdrawing one still propagates within
refresh_interval.
It applies only when the source said “revoked at generation N”. An
absent row is an unknown principal and takes unknown_ttl, however
long this instance has served that principal: a tombstone is a durable
statement, an absence is not.
Must be positive.
retry_backoff: DurationBackoff between initial-load retries while the source is down.
Also the delay before a failed or timed-out fetch is retried. Too short hammers a source that is already failing; too long delays readiness after the source recovers. Must be positive.
max_concurrent_fetches: usizeMaximum snapshot fetches in flight during a full refresh.
Too low makes a refresh of a large tracked set take many round trips,
stretching it toward or past refresh_interval; too high bursts
concurrent load at the source on every instance at once. Must be
positive.
fetch_timeout: DurationHow long one source fetch may run before it is abandoned.
The snapshot plane bounds its source calls by cancellation rather than by a wall clock everywhere else — a slow source is answered by readiness falling as resolutions expire, not by cutting the call off. That is deliberate, and it is why this bound sits at the fetch: what it protects is the loop’s ability to come back, not the freshness of any one principal (GL-103). A future that never resolves is never joined, so without it one hung fetch stops the sweep from returning and no tick, push, or control wakeup is processed again for the life of the process.
Set it above the slowest fetch the source legitimately makes, not
against the fast path: an abandoned fetch keeps the principal’s
previous resolution and retries with backoff, so a value below real
source latency turns a slow catalogue into one that never refreshes.
It is independent of refresh_interval, which is a freshness cadence
rather than a statement about call latency. Must be positive.
enumeration_timeout: DurationHow long one principal enumeration may run before it is abandoned.
Separate from fetch_timeout because the two
calls have different worst cases, not because the mechanism differs:
snapshot returns one principal’s row, principals returns the whole
catalogue. A single bound would have to be sized for the enumeration,
which would leave the per-fetch bound uselessly loose — and a limit is
justified against the largest legitimate input, so one value cannot
serve two inputs that differ by orders of magnitude.
Set it above the slowest enumeration this source legitimately performs, counted over the whole tracked set rather than a typical one. An abandoned enumeration keeps the set it already had, so a value under real catalogue latency freezes discovery while everything already tracked keeps working — the failure GL-48 exists to make visible. Must be positive.
Implementations§
Source§impl SnapshotManagerConfig
impl SnapshotManagerConfig
Sourcepub fn validate(&self) -> Result<(), SnapshotManagerConfigError>
pub fn validate(&self) -> Result<(), SnapshotManagerConfigError>
Check the configuration on its own: every duration and TTL positive,
max_concurrent_fetches positive, and no duplicate in the starting
principal set. Checks that need the map, such as capacity, happen in
SnapshotManager::spawn.
§Errors
The first rule the configuration breaks.