π RAG (Retrieval-Augmented Generation)
Pipeline
flowchart LR A[Docs] --> B[Parse & clean] --> C[Chunk] --> D[Embed] --> E[(Vector + keyword index)] Q[User query] --> R[Rewrite / expand] --> S[Hybrid search] --> T[Rerank] --> U[Build prompt w/ citations] --> V[LLM] --> W[Answer + sources] E --> S
Checklist (core β advanced)
- Ingestion: parsing PDFs/HTML, metadata, incremental updates, deletions
- Chunking strategies: fixed, recursive, semantic, structure-aware; overlap; parent-child chunks
- Embedding models: choosing, dimensions, cost; re-embedding on model change
- Vector stores: pgvector (HNSW/IVFFlat indexes) β, Qdrant, OpenSearch k-NN
- Hybrid search (BM25 + vector) with reciprocal rank fusion β
- Reranking (cross-encoders / rerank APIs)
- Query transformation: rewriting, multi-query, HyDE, decomposition
- Metadata filtering and access control (per-user document permissions!)
- Contextual retrieval (adding context to chunks before embedding)
- Citations and βI donβt knowβ behavior
- Agentic RAG (the agent decides when and what to retrieve); GraphRAG (awareness)
- Evaluation: retrieval (recall@k, MRR, nDCG) and generation (faithfulness, answer relevance) β Evals, Guardrails & LLMOps
- Long context vs RAG: the trade-offs
π§ͺ Labs (π’ warm-up β π‘ core β π΄ hard β β« boss)
- π‘ Ingestion pipeline (W18) β a baseline naive RAG with an eval set of 100 Qs
- π΄ Hybrid + RRF + rerank + contextual chunks; measure each stepβs gain
- π΄ Document ACLs at query time + a leakage test
- β« BM25 + HNSW from scratch in Go vs pgvector (recall/latency curves)
π§ Cognitive tasks
- Error analysis on 30 failures β categorize β fix the biggest bucket
- Constraint flip: 1,000Γ more documents; what breaks first?
π°οΈ Orbit integration
- UC2 knowledge copilot; RAG steps in UC3, UC4, UC6, UC7, UC8, UC10
Go deeper
βοΈ LLM Inference Internals Β· PostgreSQL Internals Β· π§© AI & Agent Patterns