StackPractices
intermediate By Mathias Paulenko

Logging, Monitoring & Observability Guide

A guide to building observable systems with structured logging, metrics, and distributed tracing.

Introduction

Observability is the ability to understand a system’s internal state by examining its outputs. The three pillars — logs, metrics, and traces — provide different perspectives on system behavior.

The Three Pillars

PillarQuestionGranularityRetention
LogsWhat happened?High (individual events)Days to weeks
MetricsHow is it trending?Low (aggregated)Months to years
TracesWhere did time go?Medium (request paths)Days to weeks

Structured Logging

Replace free-form text with machine-parseable JSON. See Structured Logging for practical implementation.

Format

{
  "timestamp": "2026-06-11T14:32:01Z",
  "level": "ERROR",
  "message": "Payment failed",
  "service": "billing-api",
  "trace_id": "abc123",
  "user_id": "user_456",
  "amount": 99.99,
  "error": "Card declined",
  "duration_ms": 245
}

Implementation (Python)

import structlog
import logging

structlog.configure(
    processors=[
        structlog.stdlib.filter_by_level,
        structlog.stdlib.add_logger_name,
        structlog.processors.TimeStamper(fmt="iso"),
        structlog.processors.StackInfoRenderer(),
        structlog.processors.format_exc_info,
        structlog.processors.JSONRenderer()
    ],
    context_class=dict,
    logger_factory=structlog.stdlib.LoggerFactory(),
)

logger = structlog.get_logger()
logger.info("payment_processed", user_id="123", amount=49.99)

Log Levels

LevelUse CaseExample
DEBUGDevelopment detailVariable values, loop iterations
INFONormal operationsRequest completed, job started
WARNUnexpected but handledRetry attempted, deprecated API used
ERRORFailed operationRequest failed, exception caught
FATALSystem unavailabilityDatabase connection lost

Metrics

Metrics are numeric data points collected over time.

Metric Types

TypeDescriptionExample
CounterOnly increasesRequests served, errors occurred
GaugeCan go up or downCurrent queue size, memory usage
HistogramDistribution of valuesRequest duration, payload size
SummaryCalculated percentilesp95 latency, p99 latency

Implementation (Prometheus)

from prometheus_client import Counter, Histogram, start_http_server

requests_total = Counter('http_requests_total', 'Total requests', ['method', 'status'])
request_duration = Histogram('http_request_duration_seconds', 'Request duration')

@request_duration.time()
def handle_request():
    requests_total.labels(method='GET', status='200').inc()
    # ... process request

start_http_server(8000)  # Exposes /metrics

Distributed Tracing

Traces follow a request across multiple services.

Trace ID: abc123
├── Service A: 5ms  (HTTP request received)
├── Service B: 12ms (Auth check)
├── Service C: 45ms (Database query)
│   ├── Connection acquire: 2ms
│   ├── Query execution: 30ms
│   └── Result mapping: 13ms
└── Service D: 8ms  (Response formatting)

Implementation (OpenTelemetry)

from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor

trace.set_tracer_provider(TracerProvider())
tracer = trace.get_tracer(__name__)

span_processor = BatchSpanProcessor(OTLPSpanExporter())
trace.get_tracer_provider().add_span_processor(span_processor)

with tracer.start_as_current_span("process_payment") as span:
    span.set_attribute("payment.amount", 99.99)
    span.set_attribute("payment.currency", "USD")
    # ... business logic

Alerting

Alert on symptoms, not causes.

Alerting Rules

# Prometheus alerting rule
groups:
  - name: api_alerts
    rules:
      - alert: HighErrorRate
        expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.1
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "High error rate on {{ $labels.service }}"

Alert Severity Levels

SeverityResponse TimeExample
CriticalImmediateService down, data loss risk
WarningWithin 1 hourElevated error rate, high latency
InfoNext business dayCapacity approaching limit

What Works

  • Use correlation IDs: Pass trace_id through every service call
  • Log at boundaries: Entry/exit of requests, jobs, and transactions
  • Avoid logging sensitive data: No passwords, tokens, or PII
  • Set SLOs and error budgets: Define what “good” means and measure against it. See monitoring.
  • Alert fatigue is real: Page only for useful, critical issues

Common Mistakes

  • Logging everything at INFO level
  • Metrics without labels (no dimensions to slice by)
  • Alerting on CPU usage instead of user-facing symptoms
  • Storing logs indefinitely without a retention policy

Troubleshooting

  • Pipeline fails silently: enable verbose logging and store pipeline artifacts between stages so you can inspect the exact state that failed.
  • Container crashes on startup: check that environment variables, secrets, and config files are mounted correctly. Read the first 50 lines of logs before scaling replicas.
  • Deployment rolls back repeatedly: verify health checks, resource limits, and startup probes. A failing readiness probe is a common cause of rolling restarts.
  • Slow CI builds: cache dependencies and docker layers. Split large test suites into parallel jobs to reduce wall-clock time.
  • Drift between environments: use infrastructure-as-code and immutable artifacts.

Further Reading

  • Official documentation: check the current reference for the framework or tool used.
  • Related guides: explore the alerting and devops guides for deeper coverage.
  • Complementary patterns: review design patterns applicable to your technology stack.
  • Public postmortems: study real incidents from teams that faced similar production issues.

Production Notes

  • Deploy gradually using canary or blue-green to catch regressions early.
  • Configure alerts for error rate, p99 latency, and failure rate before enabling in production.
  • Document the rollback in the runbook; test the procedure in staging at least once per quarter.
  • Review structured logs with correlation IDs to trace requests end-to-end during incidents.

Key Takeaways

  • Apply logging, monitoring & observability guide when you need a practical solution for your use case.
  • Monitor performance after implementation; measure latency, errors, and resource usage before and after.
  • Check the Troubleshooting section for common failures; most have documented root causes with fixes.
  • Keep dependencies updated and run tests in CI to prevent production regressions.

Advanced Topics

Scenario: Observability for E-commerce Microservices

System: 15 microservices, 500K requests/min
Stack: OpenTelemetry -> Jaeger (traces), Prometheus (metrics), Loki (logs)

Instrumentation (Node.js):
  const { trace, metrics } = require("@opentelemetry/api");
  const tracer = trace.getTracer("payment-service");

  async function processPayment(payment) {
    const span = tracer.startSpan("processPayment");
    span.setAttribute("payment.amount", payment.amount);
    span.setAttribute("payment.currency", payment.currency);
    try {
      const result = await gateway.charge(payment);
      span.setAttribute("payment.status", result.status);
      metrics.getOrCreateCounter("payments.total").add(1, {
        status: result.status, gateway: "stripe"
      });
      return result;
    } catch (error) {
      span.recordException(error);
      span.setStatus({ code: 2, message: error.message });
      metrics.getOrCreateCounter("payments.errors").add(1, {
        type: error.constructor.name
      });
      throw error;
    } finally {
      span.end();
    }
  }

Structured logs (JSON):
  {
    "timestamp": "2026-01-15T10:30:00Z",
    "level": "error",
    "service": "payment-service",
    "traceId": "abc123",
    "spanId": "def456",
    "message": "Payment failed",
    "paymentId": "pay_789",
    "amount": 99.99,
    "currency": "USD",
    "error": "InsufficientFunds"
  }

  // Correlation: traceId connects logs, metrics, and traces
  // Search traceId in Loki -> see all request logs
  // Search traceId in Jaeger -> see full trace

SLO dashboard:
  | SLO | Target | Metric |
  |-----|--------|--------|
  | Availability | 99.9% | http_requests_total{status!~5..} / total |
  | Latency p99 | < 500ms | histogram_quantile(0.99, http_duration_bucket) |
  | Error rate | < 0.1% | http_requests_total{status=~5..} / total |
  | Throughput | > 10K/s | rate(http_requests_total[5m]) |

Alerts (user-facing symptoms):
  - Error rate > 1% for 5 min -> page on-call
  - p99 latency > 1s for 10 min -> page on-call
  - SLO burn rate > 14x in 1h -> page on-call
  - Throughput < 5K/s for 5 min -> ticket (no page)

Lessons:
  - OpenTelemetry unifies traces, metrics, and logs
  - traceId is the way to correlate everything
  - Structured JSON logs > plain text
  - Alert on SLOs, not on infrastructure metrics
  - The collector decouples app from observability backend

What is SLO burn rate?

Burn rate measures how fast you consume your error budget. If your SLO is 99.9% (43.2 min of error/month), a burn rate of 14x means you are spending the budget 14 times faster than normal. At that rate, you will exhaust the budget in ~3 hours. Alerting on burn rate catches problems before they breach the SLO.

Common Production Pitfalls

  • Treating the guide as a checklist to complete once rather than a practice to evolve.
  • Adopting every recommendation at once instead of starting with one measured change.
  • Skipping the maturity assessment and forcing advanced practices on an unprepared team.
  • Not updating runbooks and on-call expectations as new practices are introduced.
  • Ignoring real incident data when prioritizing which parts of the guide to apply first.
  • Failing to assign an owner who reviews decisions quarterly.
  • Copying examples without adapting them to the team’s actual tooling and constraints.
  • Forgetting to measure outcomes before adding the next improvement.

Frequently Asked Questions

What is the difference between logs, metrics, and traces?

Logs are discrete events that answer "what happened?" Metrics are aggregated numeric data that answer "how is it trending?" Traces follow a request across services and answer "where did time go?"

How long should I retain logs?

Retain error and audit logs for 30-90 days. Debug logs can be kept for 7 days. Adjust based on compliance requirements and cost. Use log sampling for high-volume services.

What should I alert on?

Alert on user-facing symptoms: error rate, latency, and availability. Avoid alerting on infrastructure metrics like CPU or memory unless they directly correlate with user impact.