π Assignments: Phase 3 (Distributed Systems: the durable engine)
W10 - Durable Execution v1 (Go) π΄
π― A10 Turn orbit-dag into orbit-engine (Orbit v2 - Durable Workflows)
- An append-only
run_eventstable; run state = replay(events); snapshots every N events - Task queue:
SELECT β¦ FOR UPDATE SKIP LOCKED LIMIT n; leases (leased_until) + heartbeats; expired leases are re-queued - Fencing token = lease epoch;
CompleteTask(taskId, epoch)is rejected if the epoch is stale - Idempotency key per step attempt passed to tools; an idempotent sink to verify effects
- Acceptance (chaos test): 1,000 runs Γ 5 steps while a script
kill -9s random workers every 2 s andSIGSTOPs one for 30 s (a βGC pauseβ) β all runs complete, 0 duplicate effects, 0 stale completions accepted
π¬ Lab: MIT 6.5840 Lab 1 MapReduce π§ Cognitive: First-principles: write the exact interleaving where a paused worker corrupts state without fencing, then show how the epoch prevents it π Open: Postgres queue vs Kafka vs Redis Streams for task dispatch at 100 / 10k / 1M tasks/s: where does each break? Score: __/28
W11 - Kafka Backbone, Triggers & Webhooks
π― A11
- Transactional outbox in orbit-api + engine β Debezium (or a Go relay) β
run.events,usage.events(Protobuf + Schema Registry, key = runId/tenantId) - orbit-hooks (Go): inbound webhooks (HMAC verification, dedupe by delivery ID), cron triggers (next-fire index), Kafka-topic triggers β
StartRun - Outbound webhooks: HMAC signatures, exponential backoff, retry topics (1m/10m/1h), DLQ + replay CLI
- UC9: meeting transcript webhook β summarize β extract action items (schema) β create tasks via tool β idempotent on redelivery
- Acceptance: send each inbound webhook 3Γ β exactly 1 run; kill the relay mid-stream β no lost events
π¬ Lab: Gossip Glomers 1β3 (echo, unique IDs, broadcast 3aβ3e) (Kafka Internals for the log challenge later) π§ Cognitive: Predict: ordering guarantees for keyed events across 6 partitions during a consumer rebalance β verify with an experiment π Open: Choreography vs orchestration for Orbitβs internal flows (billing, notifications, search indexing) Score: __/28
W12 - Sagas, HITL, Versioning & Resilience
π― A12
-
humanstep: approval inbox API, signals, timeout β escalate timers - Compensation handlers + reverse-order compensation on failure
- UC3 refund agent: lookup order (tool) β policy check β propose a refund β approval β payment tool (idempotent) β notify; on a payment failure β compensate
- Workflow versioning: in-flight runs stay on v1 while v2 publishes; a migration policy for long waits
- Circuit breakers per LLM provider (gobreaker) + fallback provider; bulkheads per step type
- Acceptance: 100 UC3 runs with random worker kills + a 20% payment tool failure rate β 0 double refunds; every refund has an audit trail
π¬ Lab: MIT 6.5840 Lab 2 (KV server); run one sample in Temporal (Go SDK) π§ Cognitive: Pre-mortem of UC3; map each failure to its compensation π Open: ADR: βOur engine vs Temporalβ: when would you migrate? What do you lose and gain? Score: __/28
W13 - Raft & Multi-Tenant Fairness
π― A13
- Per-tenant queues + weighted fair scheduling (deficit round robin) + per-tenant concurrency caps + priority tiers
- Admission control: 429 +
Retry-Afterwhen a tenantβs backlog exceeds its quota - Run search: OpenSearch projection from
run.events(CQRS), rebuildable from scratch - Acceptance: tenant A floods 100k runs; tenants BβE (100 runs each) finish with p95 queue wait < 2 s
π¬ Lab (β«): MIT 6.5840 Lab 3A (leader election) + 3B (log replication) π§ Cognitive: Feynman: explain Raft election safety + log matching in 5 minutes π Open: Noisy neighbors at 10,000 tenants: shared pools vs cells vs dedicated workers for enterprise tiers Score: __/28
W14 - v2 Ship & Chaos Day
π― A14 Ship Orbit v2 - Durable Workflows: ADRs, sequence diagrams (happy, retry, lease expiry, compensation), README update π¬ Lab: Gossip Glomers 4 (grow-only counter) + 5 (Kafka-style log); Raft 3C if you have time
β« Phase boss fight: Chaos Day. 5,000 runs across UC1/UC3/UC9 while randomly: killing engine and worker pods, restarting a Kafka broker, failing over Postgres, injecting 2 s of latency into the LLM mock (toxiproxy)
- Acceptance: 0 lost runs, 0 duplicate side effects, all SLO breaches explained. Write it as a post-mortem (timeline, impact, root causes, action items)
π§ Cognitive: Constraint flip: redesign the engine for 1M runs/s. What changes first? Score: __/28