๐ญ Observability
Core โ Advanced
- Monitoring vs observability; the three pillars (+ profiles, events)
- Structured logging (JSON), log levels, correlation/trace IDs in every log line; what NOT to log (PII, secrets)
- Metrics: counters, gauges, histograms; RED/USE; percentiles from histograms; cardinality pitfalls โญ
- Distributed tracing: spans, context propagation (W3C
traceparent), sampling (head vs tail) - OpenTelemetry โญ: SDKs (Java agent auto-instrumentation, Go SDK), the Collector (receivers โ processors โ exporters), OTLP
- Stack: Prometheus (PromQL), Grafana, Loki (logs), Tempo (traces), Pyroscope (continuous profiling), Alertmanager
- Alerting: alert on symptoms (SLO burn rate), not causes; runbooks; avoid alert fatigue
- Dashboards: service overview (RED), dependency dashboards, business KPIs
- eBPF-based observability (awareness): Cilium Hubble, Pixie, Grafana Beyla
- LLM observability โ Evals, Guardrails & LLMOps
๐งช Labs (๐ข warm-up โ ๐ก core โ ๐ด hard โ โซ boss)
- ๐ข The LGTM stack (Loki, Grafana, Tempo, Mimir/Prometheus) + an OTel Collector via Compose
- ๐ก One trace across Java โ Go โ Kafka โ Go โ the LLM provider
- ๐ด A mini tracer in Go with W3C propagation over HTTP + Kafka (W17 lab)
- ๐ด GenAI spans: model, tokens, cost; per-tenant cost dashboards
- โซ Tail sampling that keeps 100% of errors and slow traces at 5% overall volume
๐ง Cognitive tasks
- Symptom โ hypotheses using only dashboards (drill prompts in Cognitive Drills)
๐ฐ๏ธ Orbit integration
- Observability for Orbit v3 - Cloud Native + LLM tracing in v4
Go deeper
โ๏ธ Envoy Internals ยท Kafka Internals
Resources
- Observability Engineering (Majors, Fong-Jones, Miranda) โญ ยท opentelemetry.io docs ยท Grafana docs ยท Prometheus: Up & Running