1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
# fly.toml — FeatherReader (feather-reader.com)
#
# Single always-on machine. The background feed poller + read-state flusher
# (src/scheduler.rs) run in-process and cannot scale to zero, so the machine is
# pinned running. Cloudflare (Full-strict) fronts Fly; Fly still forces HTTPS.
#
# One 1GB volume at /data holds ALL persistent state:
# * FEATHERREADER_DB (/data/featherreader.db) — the Rust app's SQLite cache
# * SIDECAR_DB (/data/oauth-sidecar.db) — the OAuth token/session store
# * the sidecar signing JWK (/data/oauth-sidecar.db.jwk.json) — persisted BESIDE
# the sidecar DB by oauth-sidecar/src/oauth.ts (see deploy/teardown.md).
# The container entrypoint chowns /data to the app uid on first boot (a fresh Fly
# volume is root:root), so the non-root app can create these files.
#
# ONE-TIME MAINTENANCE — auto_vacuum (databases created before v0.3.0 only):
# fly ssh console -C "/app/featherreader --migrate-auto-vacuum"
# Databases created from v0.3.0 on are already in INCREMENTAL mode. Older ones
# are in SQLite's default NONE, where reclaiming freed pages requires a FULL
# VACUUM — which writes a second complete copy of the database and so needs free
# disk roughly equal to the live file. That is exactly what is missing under the
# disk pressure that triggers a retention sweep, so the app now refuses to run
# one on its own and logs a warning instead. Switching modes needs that same full
# VACUUM, so it is this deliberate, operator-timed step: run it while the volume
# has headroom (it refuses itself if not). Safe to run blindly — it exits 0 doing
# nothing if already migrated.
#
# RUN IT WITH THE APP STOPPED. The VACUUM holds an exclusive lock for minutes on
# a large database, and the app's busy_timeout is 5 s — so writes do not queue,
# they FAIL: mark-read and star return errors, and OAuth session writes fail, so
# LOGINS BREAK for the duration. /health keeps returning 200 throughout, because
# its probe is a read. Also note this command runs as root and applies the
# binary's schema migrations before the VACUUM, so use the SAME image version
# that is currently deployed.
#
# DEPLOY BY DIGEST (not by tag). The release workflow publishes signed provenance
# bound to the image DIGEST; deploying a mutable tag (or :latest) can silently
# roll prod backward or bypass the attestation. The runbook resolves the tag to a
# digest, runs `gh attestation verify oci://...@<digest>` (fail-closed), then
# `fly deploy -i ghcr.io/justin-stanley/feather-reader@<digest>`. See [build] below.
#
# REQUIRED SECRETS (set once, NEVER in this file or the image):
# fly secrets set \
# FEATHERREADER_COOKIE_SECRET="$(openssl rand -base64 48)" \
# SIDECAR_INTERNAL_SECRET="$(openssl rand -base64 48)" \
# SIDECAR_ENC_KEY="$(openssl rand -base64 32)" \
# FEATHERREADER_PUBLIC_URL="https://feather-reader.com" \
# SIDECAR_PUBLIC_URL="https://feather-reader.com/oauth" \
# SIDECAR_APP_CALLBACK_URL="https://feather-reader.com/oauth/callback" \
# FEATHERREADER_ALLOWED_DIDS="did:plc:..."
#
# REQUIRED ON THE RUST OAUTH BACKEND (see "BACKEND CUTOVER" below):
# fly secrets set FEATHERREADER_OAUTH_ENCRYPTION_KEY="$(openssl rand -base64 32)"
# Unset on FEATHERREADER_REPO_BACKEND=rust, config.rs::validate_secrets() REFUSES
# TO BOOT — the app owns the OAuth flow there, so `oauth_session` holds every
# user's access token, refresh token and DPoP private key, and without this key
# the codec is a no-op and all three sit in plaintext in SQLite on the same
# volume (and in every snapshot of it) as the feed cache. Not read at all on the
# sidecar backend, which never writes those tables — so a rollback to `sidecar`
# does not need it. Rotating it strands existing sessions: the old ciphertext
# will not decrypt, and every reader must log in again.
#
# OPTIONAL SECRET — only if running the follow→invite bot (POST /bot/claims):
# fly secrets set FEATHERREADER_BOT_SECRET="$(openssl rand -base64 48)"
# Unset ⇒ /bot/claims returns 503 (endpoint disabled). If set on this (public)
# instance it must be strong (>= 32 bytes, not a dev default) or boot fails. The
# bot host must be given the SAME value (see bot/README.md). Optionally also
# FEATHERREADER_CLAIM_TTL_SECS (default 1209600 = 14 days)
# SIDECAR_PUBLIC_URL is the BROWSER-facing value (client_id/redirect_uri). The
# server-to-server SIDECAR_INTERNAL_URL is a NON-secret loopback value baked into
# the image + set in [env] below, so /internal/* calls never egress the edge.
# config.rs::validate_secrets() and the sidecar config FAIL LOUD at boot if the
# cookie / internal / enc secrets are missing or weak on this (public) instance.
#
# BACKEND CUTOVER (sidecar -> rust). FEATHERREADER_REPO_BACKEND in [env] selects
# which implementation performs the atproto calls and owns /oauth/*. Flipping it
# is a TWO-part change and both parts must land in the same deploy:
# 1. `fly secrets set FEATHERREADER_OAUTH_ENCRYPTION_KEY=...` (above) FIRST.
# Flipping the backend without it is an immediate, permanent boot loop.
# 2. Change FEATHERREADER_REPO_BACKEND to "rust" in [env] and deploy.
# The entrypoint reads the SAME variable to pick the Caddy OAuth routing and to
# decide whether to start the Node sidecar at all, so it must be image/[env]
# state, not a secret — the two cannot share /oauth/callback, and a mismatch
# breaks every login. Rolling back is the same edit in reverse; the encryption
# key can stay set, since the sidecar path never reads it.
#
# THE CUTOVER LOGS EVERY USER OUT. Nothing under src/ reads SIDECAR_DB: the rust
# backend keeps its own `oauth_session` table inside FEATHERREADER_DB, separate
# from the sidecar's /data/oauth-sidecar.db. No access token, refresh token or
# DPoP key crosses the flip, so every signed-in reader must log in again — and
# rolling BACK logs them out a second time, because the sidecar's store has gone
# stale in the meantime. This is not the same cause as the key-rotation re-login
# noted above; it happens even on a first flip with a freshly-set key. Prefer low
# traffic and warn the cohort. Related and UNVERIFIED: the two backends publish
# JWKS from different signing keys, so if a PDS caches our JWKS across the flip,
# the first login after it may fail for that reason rather than for a code bug.
= "featherreader"
= "ord"
# How long Fly waits after SIGTERM before SIGKILL. The default is 5 s, which is
# shorter than a single poll round can take to even LAUNCH: a full batch is 50
# feeds at a 250 ms stagger (12.5 s), and each feed is leased an hour forward
# BEFORE its fetch so a crash cannot re-select it first. SIGKILL rolls no lease
# back, so a deploy landing mid-round used to push up to 50 feeds an hour out —
# invisible, because /stats cannot show a feed as overdue when next_poll was
# moved forward. The poller now also stops launching new feeds the moment
# shutdown is signalled; this is the other half, so the ones already in flight
# can finish and the flusher can make its final read-state push.
= "45s"
[]
# Image is built + pushed by .github/workflows/release-image.yml to GHCR with
# SLSA build-provenance (and optional SBOM) attestations bound to the DIGEST.
# This tag is a convenience default ONLY; production deploys resolve to and
# deploy the immutable DIGEST after `gh attestation verify` (see header + the
# deploy runbook). GHCR package visibility must be public OR Fly must hold GHCR
# pull creds (see MANUAL STEPS) or the pull will fail.
= "ghcr.io/justin-stanley/feather-reader:latest"
[]
# Non-secret runtime wiring. In-container ports: Caddy 8080 (public), Rust app
# 127.0.0.1:8082, sidecar 127.0.0.1:8081.
= "127.0.0.1:8082"
= "/data/featherreader.db"
= "prod"
# Which implementation serves com.atproto.repo.* and owns /oauth/* — "sidecar"
# (Node, two processes) or "rust" (in-process, one). Stated explicitly rather
# than left to the default so the live backend is visible in this file and the
# cutover is a one-line edit. See "BACKEND CUTOVER" in the header: "rust"
# additionally REQUIRES FEATHERREADER_OAUTH_ENCRYPTION_KEY as a secret.
= "rust"
# standard.site publications may be STORED: pasted into the subscribe form, or
# imported. Stored ones are polled either way; this gates storage, not reading.
= "true"
# Cloudflare fronts Fly, so the real client IP is CF-Connecting-IP. This is
# authoritative ONLY if the origin is locked to Cloudflare (see Caddyfile note
# + MANUAL STEP: CF-only origin allowlist).
= "cf-connecting-ip"
# Keep the poller from filling a 1GB volume: pause new fetches below the volume
# size (see the startup watermark-vs-disk check in src/main.rs).
= "805306368" # 768 MiB (< 1GB volume)
# Same reasoning as FEATHERREADER_TRUSTED_IP_HEADER above: the sidecar's public
# OAuth routes are rate-limited per client IP, and behind Caddy the socket peer
# is always loopback. Authoritative ONLY with the CF-only origin lock in place.
= "cf-connecting-ip"
= "127.0.0.1"
= "8081"
# Loopback base the Rust app uses for the sidecar's /internal/* API. DISTINCT
# from the (secret) browser-facing SIDECAR_PUBLIC_URL; keeps server-to-server
# calls inside the container.
= "http://127.0.0.1:8081"
= "/data/oauth-sidecar.db"
= "info"
# Background loops wait 30/45/60/75/90 s (and the relay probe 5 min) before their
# FIRST tick, so a crash-looping machine does not re-run every sweep on every
# restart. FEATHERREADER_STARTUP_DELAY_SECS shortens them for a dev or
# integration run — it is a CEILING, so setting it HIGHER changes nothing and
# logs that it was ignored. Not set here: production wants the built-in values.
[[]]
= "featherreader_data"
= "/data"
[]
= 8080 # Caddy — the only off-loopback listener
= true
= "off" # the poller cannot scale to zero
# TRUE, deliberately, even though auto_stop is off. These two are independent:
# auto_stop=off only means the proxy never STOPS a machine, and a machine can
# still end up stopped by the restart policy giving up. The default `on-fail`
# policy retries a non-zero exit up to `max_retries` times (default 10) in 5
# minutes and then leaves
# the machine STOPPED — and with auto_start=false the proxy will not start it
# again, so an OOM crash loop (realistic on 512 MB) ends in a permanent outage
# needing a manual `fly machine start`. With auto_start=true a request brings it
# back.
#
# IT IS NOT FREE, and an earlier version of this comment claimed it was. The
# proxy wakes the machine BEFORE any in-container routing, so the Caddy origin
# lock cannot gate the wake: anyone who can reach featherreader.fly.dev can now
# force a start, even though Caddy will then 403 them. More importantly that
# turns the `on-fail` retry cap from a circuit breaker into a formality — a
# crash loop restarts at request rate indefinitely, reopening the SQLite WAL
# each time, which is least welcome in the crash-during-write case.
#
# Taken anyway: on ONE machine with no on-call, a permanent stop that only a
# human can clear is worse than a loop that at least self-heals when the cause
# is transient — and the poller lease (T2.1) already removed the crash-loop
# cause this project actually hit. Revisit if a second machine ever exists.
= true
# Inert with auto_stop_machines="off" — it only applies when that is "stop" or
# "suspend". Kept as documentation of intent, not as a safety net; it is not
# one.
= 1
= ["app"]
[]
= "requests"
= 200
= 250
# Liveness. /health reads one row from a real table (`SELECT 1 FROM feeds
# LIMIT 1`) under a 2 s timeout — inside the 3 s below, so a wedged pool yields
# a 503 the app chose, with a reason, rather than a timeout Fly inferred. It
# reads a real page deliberately: a bare `SELECT 1` emits no OpenRead opcode at
# all, so it returns success against a corrupted database. CONCURRENT probes are
# deduplicated — at most one is ever in flight, others reuse its predecessor's
# verdict — so /health, which is origin-lock-exempt and unauthenticated, cannot
# be used to spend the 5-connection pool. (There is no TIME cache: every
# sequential request, including Fly's own, gets a fresh answer.) Only the
# database can fail this check:
# the body also reports uptime, the
# poll heartbeat, whether fetching is paused at the size watermark, and the
# live OAuth backend, but none of those affect the status code.
#
# WHAT A FAILED CHECK ACTUALLY DOES — this comment used to say "Fly restarts on
# a failed check", which is FALSE. Fly's docs are explicit (health-checks
# reference, stated and then restated on that page): a failing service check
# does NOT restart or stop a Machine. Fly Proxy stops ROUTING to
# it, and re-registers automatically once the check passes. (That capability
# existed on Apps V1 as `restart_limit`; it has no successor on Machines.)
#
# MEASURED on a throwaway app mirroring this config exactly — one machine,
# auto_stop off, same 15s/3s/10s check — by failing /health for 120 s:
#
# t= 0s 200 0.21s healthy
# t= 2s 503 39.83s <-- the edge HANGS ~40 s, THEN 503s
# t= 44s 503 39.56s
# t= 85s 503 39.67s
# t=127s 200 0.30s re-registered ~7 s after health returned
#
# Machine event log afterwards: `launch` and `start`, both at creation. No
# restart, no stop — the docs' claim holds, now verified rather than read.
#
# **The ~40 s is the part worth designing around, and it is undocumented.** A
# failed check does not fail requests fast; it makes every request hang for
# about forty seconds first. Browsers sit spinning, clients hit their own
# timeouts, and connections pile up against the 250 hard_limit for the whole
# outage. That is materially worse than an immediate 503 and it is the real
# argument for keeping this check NARROW: anything that can flap here converts
# a brief internal blip into forty-second hangs for every visitor. It is also
# why /health's DB probe uses a 2 s timeout inside this 3 s one — the app must
# answer with a reason before Fly infers a failure and the hang begins.
#
# With ONE machine there is no healthy peer to shift to, so a 503 here is not a
# failover — it is a total outage lasting exactly as long as the condition. It
# ALSO fails a `fly deploy`: the rolling strategy waits for the new machine to
# go healthy and there is no auto-rollback, so a database blip during a release
# leaves a stuck, empty app. Both are reasons to keep this check narrow.
#
# So the question the status code answers is "can this process serve a useful
# request at all", NOT "would a restart help". A stale poller can serve; a
# database it cannot read cannot.
#
# ALERT ON THE BODY for everything else. The first token is the state and is
# the thing to match: `ok` (probed, healthy), `unknown` (no probe has completed
# yet — brief, at boot only) or `FAIL` (probed, broken). Matching `^ok` alone
# is not the same as "healthy".
# Reached through Caddy, so this also proves the edge + app path is up.
[[]]
= "15s"
= "3s"
= "10s"
= "GET"
= "/health"
[[]]
= "shared-cpu-1x"
= "512mb"
= 1