🧬 LLM Fundamentals

Concepts (intuition-level, enough to reason about behavior)

  • Tokens & tokenization (BPE); why cost and limits are in tokens
  • Embeddings & vector similarity (cosine, dot product)
  • Transformer intuition: attention, next-token prediction, pre-training vs post-training (SFT, RLHF/RLAIF)
  • Sampling: temperature, top-p, max tokens, stop sequences
  • Context windows, the “lost in the middle” effect, prompt caching
  • Hallucinations and why they happen; grounding
  • Reasoning / extended-thinking models: when the extra latency and cost are worth it
  • Model selection: large vs small, open-weights vs API, latency vs quality vs cost
  • Multimodal inputs (images, PDFs, audio)

Integration skills

  • Chat APIs (messages, roles, system prompts), streaming (SSE)
  • Structured output (JSON schema), validation, retries on invalid output
  • Tool / function calling: the loop (model → tool call → your code → result → model)
  • Prompt caching, batching APIs for offline jobs
  • Rate limits, timeouts, retries with backoff, fallbacks across models/providers
  • Token and cost accounting per request/user/feature
  • Local models with Ollama for dev/testing; serving with vLLM (awareness)

Prompt / context engineering

  • Clear instructions, role, context, examples (few-shot), output format
  • XML tags / sectioning, chain-of-thought where appropriate
  • Prompt templates versioned in code, tested with evals

🧪 Labs (🟢 warm-up → 🟡 core → 🔴 hard → ⚫ boss)

  • 🟢 Call hosted + local (Ollama) models from Java, Go, and Python; compare the ergonomics
  • 🟡 Structured extraction with schema validation + a repair loop (UC1/UC5)
  • 🔴 Measure TTFT vs prompt length, ITL vs load, and the effect of prompt caching
  • ⚫ Karpathy’s minbpe / Let’s build GPT

🧠 Cognitive tasks

  • Predict token counts and cost before every experiment
  • First-principles: why does output length dominate latency?

🛰️ Orbit integration

  • orbit-llm-gateway v1 → v2

Go deeper