If you can't see it, you can't fix it.
Logging, metrics, tracing are load-bearing infrastructure, not afterthoughts.
- Structured logs with correlation IDs
- Alert on symptoms, investigate causes
- Every state change observable
- Dashboards for steady state; logs for incidents
- Make the common failure modes diagnosable without a debugger