StackPractices
advanced By Mathias Paulenko

Chaos Engineering — Principles, Tools, and Safe Experiments

A practical guide to chaos engineering: build resilient systems by intentionally injecting failures. Learn the five principles, Litmus, Gremlin, and Chaos Mesh.

Overview

Chaos engineering is the discipline of experimenting on a system to build confidence in its capability to withstand turbulent conditions. Instead of waiting for failures to occur in production, you intentionally inject them — pod kills, network latency, CPU exhaustion, disk fill — to validate that your system degrades gracefully and recovers automatically. Originated at Netflix with Chaos Monkey, it has evolved into a structured practice with principles, tools, and safety guardrails.

When to Use

  • For alternatives, see Disaster Recovery: RTO, RPO, and Resilient Recovery Runbooks.

  • Your system claims to be “highly available” but has never been tested under failure

  • You want to validate autoscaling, failover, and circuit breakers

  • You need to discover unknown dependencies and single points of failure

  • Incident response runbooks exist but are untested

  • You are running on Kubernetes and want to validate pod resilience

The Five Principles of Chaos Engineering

  1. Build a hypothesis around steady-state behavior — define normal metrics (error rate < 0.1%, p99 latency < 200ms)
  2. Vary real-world events — inject failures that actually happen: network partitions, disk failures, dependency outages
  3. Run experiments in production — staging rarely matches production topology and load
  4. Automate experiments to run continuously — manual game days are valuable but not sustainable
  5. Minimize blast radius — start small (one pod, one AZ), abort if SLOs are breached

Experiment Design

┌─────────────────┐
│ 1. Steady state │ ← Define normal via metrics
│ 2. Hypothesis   │ ← "If X fails, Y autoscales in < 60s"
│ 3. Inject fault │ ← Kill pod, add latency, fill disk
│ 4. Observe      │ ← Compare actual vs hypothesis
│ 5. Rollback     │ ← Abort if blast radius exceeds bounds
│ 6. Learn        │ ← Fix weaknesses, automate fix
└─────────────────┘

Chaos Mesh Example (Kubernetes)

apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
  name: pod-kill-api
  namespace: chaos-testing
spec:
  action: pod-kill
  mode: one
  selector:
    namespaces:
      - production
    labelSelectors:
      app: api
  duration: 30s
  scheduler:
    cron: "@every 10m"

LitmusChaos Example

apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: api-pod-delete
  namespace: litmus
spec:
  appinfo:
    appns: production
    applabel: "app=api"
    appkind: deployment
  chaosServiceAccount: litmus-admin
  experiments:
    - name: pod-delete
      spec:
        components:
          env:
            - name: TOTAL_CHAOS_DURATION
              value: "30"
            - name: CHAOS_INTERVAL
              value: "10"
            - name: FORCE
              value: "false"

Common Experiment Types

ExperimentValidatesTool
Pod killKubernetes rescheduling, readiness probesChaos Mesh, Litmus
Network latencyTimeout handling, circuit breakersChaos Mesh, Gremlin
CPU/memory stressAutoscaling triggers, resource limitsStress-ng, Gremlin
Disk fillLog rotation, storage alertsLitmus, Gremlin
Zone outageMulti-AZ failoverAWS FIS, Gremlin

Safety Guardrails

  • Abort conditions — auto-stop experiment if error rate > 1% or p99 > 500ms
  • Time-bound — limit experiment duration (30s, 5m, not indefinite)
  • Small scope — one pod → one deployment → one namespace → one AZ
  • Business hours — run experiments when engineers are available
  • Clear communication — announce experiments to avoid incident duplication

Common Mistakes

  • No steady-state definition — you cannot detect degradation if you do not know what normal looks like
  • Blast radius too large — starting with a full region outage can cause real customer impact
  • No abort mechanism — experiments must auto-terminate if SLOs are breached
  • Blaming individuals for failures found — chaos engineering finds system weaknesses, not human errors
  • Running experiments without runbooks — if the experiment finds a bug, you need a remediation plan

Troubleshooting

  • Pipeline fails silently: enable verbose logging and store pipeline artifacts between stages so you can inspect the exact state that failed.
  • Container crashes on startup: check that environment variables, secrets, and config files are mounted correctly. Read the first 50 lines of logs before scaling replicas.
  • Deployment rolls back repeatedly: verify health checks, resource limits, and startup probes. A failing readiness probe is a common cause of rolling restarts.
  • Slow CI builds: cache dependencies and docker layers. Split large test suites into parallel jobs to reduce wall-clock time.
  • Drift between environments: use infrastructure-as-code and immutable artifacts.

Further Reading

  • Official documentation: check the current reference for the framework or tool used.
  • Related guides: explore the chaos-engineering and resilience guides for deeper coverage.
  • Complementary patterns: review design patterns applicable to your technology stack.
  • Public postmortems: study real incidents from teams that faced similar production issues.

Production Notes

  • Deploy gradually using canary or blue-green to catch regressions early.
  • Configure alerts for error rate, p99 latency, and failure rate before enabling in production.
  • Document the rollback in the runbook; test the procedure in staging at least once per quarter.
  • Review structured logs with correlation IDs to trace requests end-to-end during incidents.

Key Takeaways

  • Apply chaos engineering — principles, tools, and safe experiments when you need a practical solution for your use case.
  • Monitor performance after implementation; measure latency, errors, and resource usage before and after.
  • Check the Troubleshooting section for common failures; most have documented root causes with fixes.
  • Keep dependencies updated and run tests in CI to prevent production regressions.

Advanced Topics

Scenario: Game Days for E-commerce Platform

System: E-commerce, 15 microservices, K8s
Goal: Validate resilience before Black Friday

Game Day calendar (monthly):
  | Month | Experiment | Hypothesis | Result |
  |-------|-----------|-----------|--------|
  | Jan | Kill payment pod | Auto-scaling replaces in < 30s | Pass |
  | Feb | 500ms DB latency | Circuit breaker activates fallback | Fail: no fallback |
  | Mar | Kill entire AZ | Traffic reroutes to healthy AZ | Pass |
  | Apr | Redis latency | Cache miss degrades gracefully | Fail: cascading timeouts |
  | May | Kill search service | Catalog without search works | Pass |
  | Jun | Corrupt Kafka message | Consumer handles poison pill | Fail: consumer hangs |

Detailed experiment (Feb):
  Name: DB-latency-injection
  Hypothesis: If DB has 500ms latency, circuit breaker
              activates cache fallback in < 5s with no 5xx errors
  Blast radius: 10% of traffic (canary)
  Duration: 10 minutes
  Abort: error rate > 5% or p99 latency > 3s

  Execution (Gremlin):
    gremlin attack latency -t 500ms -i 600 --service payment-db
    --tags env=canary

  Monitoring during experiment:
    - Error rate: 0% -> 12% (FAIL)
    - p99 latency: 200ms -> 4.5s (FAIL)
    - Circuit breaker: never activated (FAIL)
    - Cache fallback: not implemented (FAIL)

  Post-mortem analysis:
    Root cause: Circuit breaker threshold set to 10s
                but DB responded in 500ms (no timeout).
                Cache fallback did not exist.

    Actions:
    1. Implement cache fallback for product queries
    2. Lower circuit breaker threshold to 2s
    3. Add 1s timeout on DB queries
    4. Re-run experiment in staging

  Re-run (Mar):
    - Error rate: 0% (PASS)
    - p99 latency: 200ms -> 350ms (PASS)
    - Circuit breaker: active at 3s (PASS)
    - Cache fallback: served stale data (PASS)

Automation (Chaos Mesh):
  apiVersion: chaos-mesh.org/v1alpha1
  kind: PodChaos
  metadata:
    name: payment-pod-kill
  spec:
    action: pod-kill
    mode: fixed-percent
    value: "10"
    selector:
      namespaces: [production]
      labelSelectors:
        app: payment-service
    scheduler:
      cron: "@every 1h"

How do I convince management to do chaos engineering?

Start with a game day in staging. Document findings: each discovered failure is a production incident avoided. Quantify impact: “This experiment found a bug that would have caused 2h downtime on Black Friday ($500K)”. Staging game days have zero risk and high ROI.

End of document. Review and update quarterly.

Common Production Pitfalls

  • Treating the guide as a checklist to complete once rather than a practice to evolve.
  • Adopting every recommendation at once instead of starting with one measured change.
  • Skipping the maturity assessment and forcing advanced practices on an unprepared team.
  • Not updating runbooks and on-call expectations as new practices are introduced.
  • Ignoring real incident data when prioritizing which parts of the guide to apply first.
  • Failing to assign an owner who reviews decisions quarterly.
  • Copying examples without adapting them to the team’s actual tooling and constraints.
  • Forgetting to measure outcomes before adding the next improvement.

Frequently Asked Questions

How do I get started with this in an existing project?

Start with a small, isolated part of your codebase. Apply the concepts from this guide to one module or service. Measure the impact, then expand to other areas.

What tools do I need?

The tools mentioned throughout this guide are listed in each section. Most are open-source and widely adopted. Check the related resources for setup instructions.

How do I measure success after implementing this?

Define clear metrics before starting: performance benchmarks, error rates, or maintainability indicators. Compare before and after. Iterate based on the data, not on assumptions.