Orbit v5: Scale & Polish (Phase 6)

  • Load: 1,000 concurrent runs + 500 concurrent streams (mock LLM provider with realistic latency distributions)
  • Profile (JFR, async-profiler, pprof, EXPLAIN, ClickHouse query log) → fix the top 3 bottlenecks → re-test → write a “scaling journal”
  • Partition the event-history table (by time/hash); archive completed runs to S3 (claim check)
  • “100× plan” doc: cell-based architecture (tenant → cell routing), sharded engine, multi-region control plane, cost model
  • Cost model: $/run by use case at 1M runs/month
  • Portfolio README: problem, architecture, 12 hard problems + how you solved them, numbers, demo, how to run it
  • Pin the repo; link it from your resume and LinkedIn; publish the blog series

Interview stories this project gives you → Behavioral & STAR

Exactly-once effects · lease/fencing bug found in chaos testing · noisy-neighbor fairness · streaming through proxies · RAG quality from 0.6 to 0.9 recall · a prompt-injection fix · cost cut by 40% · Java vs Go trade-off