📝 Assignments: Phase 5 (AI Engineering: all the use cases)
W19 - LLM Engineering Deep Dive
🎯 A19
- Gateway v2: routing policies (by task/cost), cascade (cheap → strong on low confidence), provider fallback, prompt caching (stable prefix ordering), semantic cache with a measured false-hit rate, per-tenant $ budgets
- Prompt registry: versioned templates, variables, A/B assignment, linked eval scores
- UC5: document extraction: a
mapstep over 500 PDFs/images (multimodal) → schema validation + business rules → low-confidence items go to a human review queue - UC8: content pipeline: prompt chaining with gates + evaluator–optimizer (≤ 2 loops)
- Acceptance: UC5 field accuracy ≥ 95% on 50 labeled docs; cost per doc reported; the semantic cache’s false-hit rate < 1% at the chosen threshold
🔬 Lab (🔴): Karpathy minbpe (or a toy BPE); measure TTFT vs prompt length and the effect of prompt caching (LLM Inference Internals) 🧠 Cognitive: Predict: token counts and $ for 10 prompts (English, Hindi, JSON, code) → measure 🌀 Open: Design the cascade policy: which confidence signals can you trust? (logprobs, self-rating, verifier model, schema validity) Score: __/28
W20 - RAG Done Right
🎯 A20 orbit-knowledge retrieval
- Hybrid: pgvector HNSW + Postgres FTS → reciprocal rank fusion → reranker
- Contextual chunk headers; parent-child chunks; metadata filters; document ACLs enforced at query time
- Multi-turn query rewriting; citations; “I don’t know” when support is weak
- Eval set: 100 questions with gold passages → recall@5, MRR, faithfulness (LLM judge calibrated on 30 human labels)
- UC2 full: knowledge copilot
- Acceptance: recall@5 ≥ 0.85; faithfulness ≥ 0.9; an ACL test proves user A never receives user B’s private chunks
🔬 Lab (🔴): BM25 + simplified HNSW from scratch in Go; compare recall and latency against pgvector on 100k vectors 🧠 Cognitive: error analysis: read 30 bad answers → categorize them (retrieval miss / chunking / ranking / generation) → fix the top category first 🌀 Open: Long context vs RAG for a 2M-token corpus: cost, latency, quality, freshness, ACLs Score: __/28
W21 - Agents, LangGraph & MCP
🎯 A21
- ReAct agent from scratch in Go: tool registry, a scratchpad, budgets (steps/tokens/$), loop detection (repeated tool+args), context compaction
- The same agent in LangGraph (Python): state graph, Postgres checkpointer,
interruptfor approval, time travel, streaming - Orbit as an MCP server (Spring AI): tools
list_workflows,start_run,get_run,search_knowledge,approve_step; resources: definitions, run logs; test with MCP Inspector and a real client (Claude Desktop/Claude Code) - MCP client in orbit-tools: register external MCP servers (GitHub, web search, filesystem, Prometheus) as scoped Orbit tools
- UC4 research agent (orchestrator–workers + evaluator) · UC10 PR review bot · UC11 supervisor multi-agent
- Acceptance: UC4 finishes 20 prompts within budget; UC11 vs a single agent compared on quality/cost/latency in a table
🔬 Lab: LangChain Academy Intro to LangGraph modules 1–5 🧠 Cognitive: “What did the framework hide?”: list 10 things LangGraph does that your Go agent had to do by hand 🌀 Open: A tool permission model: per-tenant, per-agent, and per-user scopes; approval policies; audit. Design it Score: __/28
W22 - Evals, Guardrails & the Red-Team Boss Fight
🎯 A22
- orbit-evals: suites per use case (code checks + LLM-as-judge), a promptfoo/DeepEval runner, a CI gate on PRs that touch prompts/models/tools, online sampling of 5% of production runs
- Guardrails: injection detector (UC12 voting guard), dual-LLM/quarantine for untrusted content, output schema validation, PII redaction, tool allowlists
- UC6 text-to-SQL analyst (read-only role, allowlisted tables, LIMIT, timeout, self-verification) · UC7 SRE agent via Prometheus/K8s MCP with approval gates
- Langfuse (or OTel-native) traces of every LLM/tool call
⚫ Phase boss fight: Red-Team Day. Attack your own agents with 30+ attacks: direct/indirect prompt injection (in docs, tool outputs, web pages), data exfiltration via tools, excessive agency, SQL tricks, cost-exhaustion loops. Fix, re-run, and publish a report with before/after numbers 🧠 Cognitive: Transfer: which classic security principles (least privilege, input validation, defense in depth, audit) map to which agent defenses? Score: __/28