Start with service objectives
Measure uptime, error rate and latency for the pages or endpoints that users actually need. Add queue failures, backup age and storage capacity if relevant.
Three essential signals
- Availability: Is the service reachable and functioning?
- Failures: How often do critical requests error?
- Latency: Are real users experiencing unacceptable delays?
Privacy-safe instrumentation
Keep personal information, credentials and form bodies out of logs and traces. Set retention and access restrictions according to the project.
Alerting policy
Choose a small number of actionable alerts with known owners. Avoid sending alerts that nobody intends to act on. Verify how the system behaves when its monitoring provider is unavailable.
Drill
Simulate a stopped worker or database outage, confirm detection and document the recovery path.
Examples are starting points, not production security audits. Confirm dependencies, versions and pricing using linked vendor documentation.