Architecture Reference // 2026

Observability
for SaaS

Cloud-native principles for any engineering team building modern SaaS.
8 domains  ·  26 rules.

MET
→
LOG
→
TRC

METRICS  ·  LOGS  ·  TRACES

01 // Foundation

What observability actually means

Observability is a property of a system — the degree to which you can understand internal state from external outputs. It is not a tool purchase. Three signals and one context thread make it possible to ask any question about any failure after the fact.

Observability is not monitoring
Monitoring tells you something is wrong. Observability lets you ask arbitrary questions about why — including questions you didn't think to ask at deploy time. If you need to ship new code to answer a new question about system state, your observability is incomplete.
Three signals, one context
Metrics (aggregatable numerics), logs (discrete events), and traces (request paths across services) are all necessary — none is sufficient alone. The glue is a correlation ID present in all three simultaneously.
Instrument at every service boundary
Every service entry point must emit all three signals and propagate trace context downstream. A single uninstrumented service breaks trace continuity for everything calling it. Gaps are not recoverable after the fact.
02 // Structured Logging

Logs must be queryable, not just readable

A log that requires regex to parse is an operational tax that compounds with scale. JSON lines, mandatory context fields, and enforced level semantics are the minimum floor for any production system.

01
Machine-readable first
JSON lines over free-text always. A structured log is queryable, alertable, and cross-correlatable. Free-text is only legible to the engineer who was there when it was written.
02
Mandatory context fields on every line
At minimum: trace_id, span_id, service, env, level, timestamp (ISO 8601 UTC). Without trace_id, correlating a log entry to the trace that produced it requires guesswork.
03
Agreed-upon level semantics — enforced
ERROR = requires human action now. WARN = may require action. INFO = normal operational event. DEBUG = off in production. When teams interpret levels differently, alerting on log level becomes meaningless noise.
04
Never log secrets, tokens, or PII by default
Logs are broadly accessible and retained for months. A PII field in a log line is a compliance incident waiting to be discovered. Maintain an explicit field allowlist — not a blocklist.
03 // Metrics

Measure what users experience

Metrics are aggregatable numerics over time. The RED and USE frameworks give you a near-complete picture of system health without building everything from scratch. But measurements are only actionable if you know what you're measuring toward.

RED for services, USE for resources

For every service: Rate (req/s), Errors (failure rate), Duration (latency distribution). For every resource: Utilization, Saturation, Errors. These two frameworks cover the vast majority of what SaaS teams need to alert on.

Measure distributions, not averages

p50, p95, p99 — not mean. A 10s outlier averaged across 10,000 fast requests disappears completely. Average latency actively hides your worst user experience. Your slowest users are invisible in aggregate averages.

Cardinality is a budget, not a feature

High-cardinality labels — user IDs, request IDs, session IDs — in metric series explode storage and query cost exponentially. User-level data belongs in traces and logs. Keep metric label cardinality bounded and explicit per metric.

Define SLIs before choosing any tool

A Service Level Indicator is the concrete thing you measure: "fraction of checkout requests completing under 800ms, excluding known bot traffic." Without agreed SLIs first, you have dashboards but no signal.

04 // Distributed Tracing

Follow the request, wherever it goes

Traces give you the full execution path of a single request across every service, queue, and async boundary it touches. They are the only signal that shows where time was spent and which service caused a failure across a distributed system.

01
Propagate W3C TraceContext on every call
The traceparent header is the standard. Every HTTP call, queue message, and scheduled task must carry it. A break in propagation creates orphaned traces — the hardest category of failures to debug after the fact.
02
Every async boundary is a span
Queue publishes, workers, webhooks, cron jobs, and event handlers are all part of user-facing request paths. If they don't carry trace context, end-to-end visibility breaks the moment a request goes async. This is where most teams have their observability black holes.
03
Tail-based sampling, not head-based
Head-based sampling (keep random N%) discards rare errors proportionally. Tail-based sampling keeps traces containing errors, high latency, or anomalies — and drops the rest. At production volume, 100% retention is unaffordable; tail sampling is the correct default.
04
Spans carry semantic attributes
HTTP method, status code, DB statement type, queue name. The OpenTelemetry Semantic Conventions define standard attribute names — use them. Custom attribute names break every dashboard when you switch backends.
05 // Alerting & SLOs

Alert on what users feel, not on what machines do

An on-call engineer receiving 50 alerts a week starts ignoring all of them. The goal is not more alerting — it is better alerting. Every alert should be actionable and proportional to actual customer impact.

1

Alert on symptoms, not causes

"Error rate > 1%" is a symptom — users experience it. "CPU > 80%" is a cause — you may not care. Users don't call you because your CPU is high; they call because something is slow or broken.

2

Error budgets and burn-rate alerts

99.9% availability = 43 min/month budget. Alert when burning too fast: 5% consumed in 1 hour pages, 2% in 6 hours creates a ticket. Every alert is proportional to real customer impact. This eliminates most false positives.

3

Every alert links to a runbook

An alert without a runbook is a question with no answer at 3am. Even a 5-line "check X, then Y, escalate to Z" prevents thrashing during incidents and encodes institutional knowledge that otherwise lives in one engineer's head.

4

Tune continuously to eliminate fatigue

Every alert that fires and gets resolved with "that's normal" should be deleted or re-thresholded immediately. Alert fatigue is a security risk as much as an operational one — a numbed on-call misses the real incident.

06 // Instrumentation Strategy

Where to instrument, and how much

Instrumentation is not free — it adds overhead and maintenance cost. The right strategy is layered: start with what the platform gives for free, then add manual instrumentation only where it matters to users and to the business.

Auto-instrumentation first, manual spans for business logic
OpenTelemetry auto-instrumentation covers HTTP, gRPC, database clients, and queue clients at zero code cost. Add manual spans only for meaningful operations: checkout flows, payment processing, critical async paths. Instrument what fails and what customers care about.
Health endpoints are non-negotiable
/health/live (is the process alive?) and /health/ready (is it ready to serve traffic?) must exist on every service. Orchestrators and load balancers depend on these for traffic routing and restart decisions. A service without them is invisible to the platform.
Instrument business metrics explicitly
Infrastructure metrics come from the platform. Application metrics come from the runtime. Business metrics — conversion rate, revenue per minute, active tenants, feature adoption — require explicit instrumentation. These are often the most actionable signals during a SaaS incident.
07 // SaaS Tenant Observability

Aggregate metrics hide per-tenant failures

A 99.9% service SLO can mask a 95% SLO for one enterprise customer — a churn event invisible in aggregate. Multi-tenant SaaS requires tenant attribution on every signal and distinct SLOs per tier.

Every signal must carry tenant_id
Tag every metric attribute, log field, and trace span with tenant_id at the instrumentation layer — not added later in a pipeline. A degradation affecting one tenant is completely invisible in aggregate views without it. This is the single most missed requirement in SaaS observability.
Tenant-level SLOs, not just service-level
Free, growth, and enterprise customers may have contractually different availability expectations. Measure and alert against each tier independently — a single service SLO averaging across all tenants misleads both on-call and the business.
Usage metering is observability
Feature usage, API call counts, storage consumed, compute consumed — these are not just billing data. They predict capacity failures before they happen and identify tenants growing beyond their tier before they hit a wall.
Required attributes on every signal
tenant_id
trace_id
tier
env
service
region
version
08 // Data Architecture

How signals are collected, stored, and queried

Each signal type has fundamentally different query patterns, cardinality characteristics, and retention needs. Forcing all three into one system creates cost and performance problems. The collection layer should be vendor-neutral from day one.

OpenTelemetry as the collection standard

The OTel Collector is a vendor-neutral pipeline: receive from any source, process (sample, enrich, redact), export to any backend. It decouples instrumentation from vendor lock-in — swap backends without re-instrumenting a single service.

Separate storage by signal type

Metrics go to Prometheus / VictoriaMetrics / Thanos. Logs go to Loki / OpenSearch / Clickhouse. Traces go to Tempo / Jaeger / Zipkin. Each signal type demands a fundamentally different storage model — forcing them into one system breaks both.

Define retention tiers and enforce them

Hot (24h, full fidelity) → Warm (7–30d, sampled) → Cold (90d+, aggregates only). Full-fidelity data retained forever is cost-prohibitive. Define tiering SLAs, automate transitions, communicate them to engineers before post-mortems.

Observability cost is a first-class metric

Monitor ingestion rate, storage cost, and query cost with the same rigor as system metrics. Logs and traces are often the largest cloud line item in mature SaaS. Set ingestion budgets, sample aggressively at volume, drop DEBUG entirely in production.

The mental model

Metrics tell you something is wrong.
Logs tell you what happened.
Traces tell you where and why.
The correlation ID stitches all three together.
Without tenant attribution, none of it tells you who is affected.

Daniel Brasileiro