🐧 Linux & OS Internals (important parts only)

1. Syscalls & context switches

  • User β†’ kernel transitions cost ~100 ns–¡s; context switches cost more (Β΅s + cache/TLB pollution) β†’ fewer, bigger I/O calls, and thread counts matter
  • strace -c shows where a process spends its syscalls

2. Virtual memory

  • Every process sees its own address space; pages (4 KB) map to physical frames through page tables; the TLB caches translations
  • Page faults: minor (mapping only) vs major (disk read) β†’ latency spikes
  • fork() uses copy-on-write (why Redis BGSAVE memory spikes)
  • Overcommit + the OOM killer: picks a victim by oom_score; in containers, the cgroup limit triggers it

3. Page cache & durability

  • Reads/writes go through the page cache; write() returning β‰  on disk β†’ fsync is what makes it durable (databases and Kafka reason about this)
  • Sequential I/O + readahead is far faster than random I/O (this is why logs and LSM trees exist)

4. I/O multiplexing

  • Blocking I/O (a thread per connection) β†’ non-blocking + epoll (readiness events) β†’ io_uring (submission/completion rings, fewer syscalls)
  • The C10k problem and why event loops / goroutines / virtual threads solve it

5. CPU scheduling

  • Linux uses the fair scheduler (CFS, replaced by EEVDF since 6.6); nice; CPU affinity
  • cgroup CPU quota (cpu.max) β†’ throttling in 100 ms periods β†’ p99 latency spikes for bursty services

6. File descriptors & limits

Everything is a file: sockets, pipes, files β†’ ulimit -n; β€œtoo many open files” = fd leak or pool misconfiguration; lsof -p

7. Load average & saturation

Load = runnable + uninterruptible (D state, usually I/O) tasks; compare it with the core count; use vmstat 1 (r, b, si/so, wa), iostat -x, pidstat

πŸ”¬ Prove it

  • strace -f -c a Java and a Go HTTP server under load; compare syscall profiles
  • perf top / flame graph of a CPU-bound Go program
  • Write 1 GB with and without fsync per 4 KB; compare throughput
  • vmstat 1 while causing memory pressure β†’ watch swap/page faults
  • Run a container with --cpus=0.5, send bursty load, and read cpu.stat (nr_throttled)
  • Leak fds in a Go program until EMFILE; find it with lsof

Interview questions interview-q

What happens during a context switch Β· page cache and fsync Β· epoll vs threads Β· CPU throttling in containers Β· what the load average means