☸️ Kubernetes Internals (important parts only)
1. The core idea: declarative state + reconciliation
You write desired state into the API server. Controllers keep watching and drive actual state toward it. They are level-triggered (they reconcile the full state, not individual events), so missed events self-heal.
2. API server & etcd
- The API server is the only component that talks to etcd
- Request pipeline: authentication → authorization (RBAC) → mutating admission (webhooks, defaults) → schema validation → validating admission → persist to etcd
- Every object has a
resourceVersion→ optimistic concurrency (a conflict returns 409; clients retry) - Watch: etcd watch → API server watch cache → clients stream changes
- etcd: a Raft-replicated KV with MVCC revisions; needs compaction/defrag; keep it small and fast (latency-sensitive disks)
3. Controllers & informers (client-go)
- Informer = list + watch → a local cache (indexer) → event handlers → push keys to a rate-limited work queue → workers call
Reconcile(key) - Shared informers avoid N controllers hammering the API server
- Owner references + the garbage collector: deleting a Deployment cascades to its ReplicaSets and Pods
- Chain: Deployment controller → ReplicaSet → RS controller → Pods (unscheduled)
4. Scheduler
- Watches Pods with no
nodeName→ filter (resources fit, taints/tolerations, affinity, volume topology) → score (spread, least-allocated, image locality) → bind - Scheduling uses requests, not limits. Priority classes + preemption
5. kubelet & the node
- Watches Pods bound to its node → CRI (containerd) → creates the sandbox (pause container holds the network namespace) → the CNI plugin wires the network (veth, IP) → CSI mounts volumes → init containers → app containers
- Runs probes; reports status; evicts Pods under memory/disk pressure
- Enforces resources via cgroups: CPU limit = CFS quota → throttling (latency spikes!) · memory limit → OOMKilled
- QoS classes (Guaranteed / Burstable / BestEffort) decide eviction order
6. Service networking
- Every Pod gets a routable IP (flat network via CNI)
- Service = a stable virtual IP + DNS name; EndpointSlices list the ready Pod IPs (readiness decides membership)
- kube-proxy programs iptables/IPVS (or eBPF with Cilium) to DNAT the VIP → a Pod IP; conntrack keeps connections sticky
- ⚠️ Long-lived connections (gRPC/HTTP2) stick to one Pod → no load balancing → use L7 (Envoy/mesh) or client-side LB
7. Pod termination (a classic source of 502s)
Delete → Pod marked Terminating → in parallel: (a) endpoints removed, which propagates to kube-proxy/LBs, and (b) preStop hook → SIGTERM → grace period → SIGKILL
→ Race: traffic still arrives after SIGTERM → add a short preStop sleep + graceful shutdown in the app
8. The full journey of kubectl apply -f deploy.yaml
kubectl → API server (authn/authz/admission) → etcd → Deployment controller creates an RS → RS controller creates Pods → scheduler binds → kubelet → containerd/runc + CNI → probes pass → EndpointSlice updated → kube-proxy/Envoy routes traffic
🔬 Prove it
-
kubectl get --raw /api/v1/namespaces/default/pods | jqandkubectl get pods -w -v=8(see the watch HTTP calls) - Write down the full
kubectl applyjourney, then verify it withkubectl get events --sort-by=.lastTimestamp - Two clients update the same ConfigMap with a stale
resourceVersion→ observe the 409 - CPU throttling: a CPU-hungry pod with a 200m limit → watch
container_cpu_cfs_throttled_periods_totaland latency - Remove
preStop+ fast shutdown → run k6 during a rollout → count 5xx → fix it → 0 errors -
iptables-save | grep KUBE-SVCon a kind node: find your Service’s DNAT rules - Build a controller with kubebuilder (
OrbitWorkerPool) → Assignments - Phase 4
Interview questions interview-q
What happens on kubectl apply · how Services route traffic · readiness vs liveness · requests vs limits (CPU throttling, OOM) · how controllers work · why gRPC doesn’t load-balance behind a ClusterIP · graceful shutdown in K8s