π‘οΈ SRE, Reliability & Performance
Reliability
- SLI β SLO β SLA; error budgets and error budget policies
- Toil, on-call, incident response (roles: IC, comms, ops), blameless post-mortems
- Failure modes: cascading failures, retry storms, thundering herds, metastable failures
- Resilience patterns β Architecture Patterns
- Capacity planning (Littleβs Law, load tests, headroom)
- Graceful degradation & load shedding (e.g. disable recommendations when the DB is overloaded)
- Chaos engineering: hypotheses, blast radius; tools: Chaos Mesh, Litmus, toxiproxy
- DR: RPO/RTO, backups tested by restores
- Deploy safety: canaries, feature flags, automatic rollback
Performance engineering
- Load testing with k6 (or Gatling for Java): load, stress, spike, soak tests
- Profiling: JFR/async-profiler, Go pprof, flame graphs
- Find and fix: DB (indexes, pools), caching, serialization, GC, lock contention
- Write up a performance report: goal β setup β results β bottleneck β fix β new results
π§ͺ Labs (π’ warm-up β π‘ core β π΄ hard β β« boss)
- π‘ SLOs: run success, TTFT p95, queue wait p95; burn-rate alerts
- π΄ k6: load, spike, and soak tests; a capacity report
- π΄ Chaos: toxiproxy latency on the LLM mock; kill pods; fail Postgres over
- β« Reproduce a retry storm or metastable failure, then write a post-mortem
π§ Cognitive tasks
- Pre-mortem before every Orbit version
- Write one blameless post-mortem per phase
π°οΈ Orbit integration
- Chaos Day (W14) + Production Day (W18) boss fights
Go deeper
Resources
- Google SRE book + SRE Workbook (free, sre.google) β Β· Release It! 2e β
- Systems Performance 2e (Gregg) + brendangregg.com Β· Implementing Service Level Objectives (Hidalgo)
- Marc Brookerβs blog (AWS) on retries, backoff, metastability β