Cloud-native principles for any engineering team building modern SaaS.
8 domains · 26 rules.
METRICS · LOGS · TRACES
Observability is a property of a system — the degree to which you can understand internal state from external outputs. It is not a tool purchase. Three signals and one context thread make it possible to ask any question about any failure after the fact.
A log that requires regex to parse is an operational tax that compounds with scale. JSON lines, mandatory context fields, and enforced level semantics are the minimum floor for any production system.
trace_id, span_id, service, env, level, timestamp (ISO 8601 UTC). Without trace_id, correlating a log entry to the trace that produced it requires guesswork.ERROR = requires human action now. WARN = may require action. INFO = normal operational event. DEBUG = off in production. When teams interpret levels differently, alerting on log level becomes meaningless noise.Metrics are aggregatable numerics over time. The RED and USE frameworks give you a near-complete picture of system health without building everything from scratch. But measurements are only actionable if you know what you're measuring toward.
For every service: Rate (req/s), Errors (failure rate), Duration (latency distribution). For every resource: Utilization, Saturation, Errors. These two frameworks cover the vast majority of what SaaS teams need to alert on.
p50, p95, p99 — not mean. A 10s outlier averaged across 10,000 fast requests disappears completely. Average latency actively hides your worst user experience. Your slowest users are invisible in aggregate averages.
High-cardinality labels — user IDs, request IDs, session IDs — in metric series explode storage and query cost exponentially. User-level data belongs in traces and logs. Keep metric label cardinality bounded and explicit per metric.
A Service Level Indicator is the concrete thing you measure: "fraction of checkout requests completing under 800ms, excluding known bot traffic." Without agreed SLIs first, you have dashboards but no signal.
Traces give you the full execution path of a single request across every service, queue, and async boundary it touches. They are the only signal that shows where time was spent and which service caused a failure across a distributed system.
traceparent header is the standard. Every HTTP call, queue message, and scheduled task must carry it. A break in propagation creates orphaned traces — the hardest category of failures to debug after the fact.An on-call engineer receiving 50 alerts a week starts ignoring all of them. The goal is not more alerting — it is better alerting. Every alert should be actionable and proportional to actual customer impact.
"Error rate > 1%" is a symptom — users experience it. "CPU > 80%" is a cause — you may not care. Users don't call you because your CPU is high; they call because something is slow or broken.
99.9% availability = 43 min/month budget. Alert when burning too fast: 5% consumed in 1 hour pages, 2% in 6 hours creates a ticket. Every alert is proportional to real customer impact. This eliminates most false positives.
An alert without a runbook is a question with no answer at 3am. Even a 5-line "check X, then Y, escalate to Z" prevents thrashing during incidents and encodes institutional knowledge that otherwise lives in one engineer's head.
Every alert that fires and gets resolved with "that's normal" should be deleted or re-thresholded immediately. Alert fatigue is a security risk as much as an operational one — a numbed on-call misses the real incident.
Instrumentation is not free — it adds overhead and maintenance cost. The right strategy is layered: start with what the platform gives for free, then add manual instrumentation only where it matters to users and to the business.
/health/live (is the process alive?) and /health/ready (is it ready to serve traffic?) must exist on every service. Orchestrators and load balancers depend on these for traffic routing and restart decisions. A service without them is invisible to the platform.A 99.9% service SLO can mask a 95% SLO for one enterprise customer — a churn event invisible in aggregate. Multi-tenant SaaS requires tenant attribution on every signal and distinct SLOs per tier.
tenant_id at the instrumentation layer — not added later in a pipeline. A degradation affecting one tenant is completely invisible in aggregate views without it. This is the single most missed requirement in SaaS observability.Each signal type has fundamentally different query patterns, cardinality characteristics, and retention needs. Forcing all three into one system creates cost and performance problems. The collection layer should be vendor-neutral from day one.
The OTel Collector is a vendor-neutral pipeline: receive from any source, process (sample, enrich, redact), export to any backend. It decouples instrumentation from vendor lock-in — swap backends without re-instrumenting a single service.
Metrics go to Prometheus / VictoriaMetrics / Thanos. Logs go to Loki / OpenSearch / Clickhouse. Traces go to Tempo / Jaeger / Zipkin. Each signal type demands a fundamentally different storage model — forcing them into one system breaks both.
Hot (24h, full fidelity) → Warm (7–30d, sampled) → Cold (90d+, aggregates only). Full-fidelity data retained forever is cost-prohibitive. Define tiering SLAs, automate transitions, communicate them to engineers before post-mortems.
Monitor ingestion rate, storage cost, and query cost with the same rigor as system metrics. Logs and traces are often the largest cloud line item in mature SaaS. Set ingestion budgets, sample aggressively at volume, drop DEBUG entirely in production.
Metrics tell you something is wrong.
Logs tell you what happened.
Traces tell you where and why.
The correlation ID stitches all three together.
Without tenant attribution, none of it tells you who is affected.