๐Ÿšฐ Data Pipelines & Data Engineering (for backend engineers)

Concepts

  • OLTP vs OLAP; row vs columnar storage
  • ETL vs ELT; batch vs streaming; lambda vs kappa architecture
  • CDC (Change Data Capture) with Debezium โญ
  • Data lake, warehouse, lakehouse (Apache Iceberg / Delta Lake table formats), Parquet
  • Orchestration: Airflow / Dagster (DAGs, retries, backfills)
  • Stream processing: Apache Flink / Kafka Streams / Spark Structured Streaming: windows, watermarks, late data, state
  • Batch processing: Spark basics (awareness)
  • Data quality, schema evolution, lineage, idempotent pipelines, backfills
  • Transformations with dbt (awareness)
  • OLAP stores: ClickHouse, BigQuery, Snowflake, Druid/Pinot (real-time analytics)
  • Pipelines for AI: document ingestion โ†’ chunk โ†’ embed โ†’ index (RAG)

๐Ÿงช Labs (๐ŸŸข warm-up โ†’ ๐ŸŸก core โ†’ ๐Ÿ”ด hard โ†’ โšซ boss)

  • ๐ŸŸก Ingestion pipeline: S3 โ†’ parse โ†’ chunk โ†’ embed โ†’ pgvector; idempotent + resumable
  • ๐Ÿ”ด CDC โ†’ Kafka โ†’ ClickHouse with materialized views for cost per tenant
  • ๐Ÿ”ด A Flink SQL or Kafka Streams windowed aggregation with late events
  • โšซ A backfill: re-embed 100k chunks with a new model with zero downtime (versioned index + alias swap)

๐Ÿง  Cognitive tasks

  • Fermi: embedding cost + time for 1M docs; verify on 10k

๐Ÿ›ฐ๏ธ Orbit integration

  • orbit-knowledge ingestion + orbit-analytics

Go deeper

๐Ÿงฉ Distributed & Cloud Patterns (pipes & filters, claim check)

Resources

  • Fundamentals of Data Engineering (Reis & Housley) โญ ยท Streaming Systems (Akidau et al.)
  • DataTalksClub Data Engineering Zoomcamp (free) โญ ยท Debezium tutorial ยท ClickHouse docs