๐ฐ Data Pipelines & Data Engineering (for backend engineers)
Concepts
- OLTP vs OLAP; row vs columnar storage
- ETL vs ELT; batch vs streaming; lambda vs kappa architecture
- CDC (Change Data Capture) with Debezium โญ
- Data lake, warehouse, lakehouse (Apache Iceberg / Delta Lake table formats), Parquet
- Orchestration: Airflow / Dagster (DAGs, retries, backfills)
- Stream processing: Apache Flink / Kafka Streams / Spark Structured Streaming: windows, watermarks, late data, state
- Batch processing: Spark basics (awareness)
- Data quality, schema evolution, lineage, idempotent pipelines, backfills
- Transformations with dbt (awareness)
- OLAP stores: ClickHouse, BigQuery, Snowflake, Druid/Pinot (real-time analytics)
- Pipelines for AI: document ingestion โ chunk โ embed โ index (RAG)
๐งช Labs (๐ข warm-up โ ๐ก core โ ๐ด hard โ โซ boss)
- ๐ก Ingestion pipeline: S3 โ parse โ chunk โ embed โ pgvector; idempotent + resumable
- ๐ด CDC โ Kafka โ ClickHouse with materialized views for cost per tenant
- ๐ด A Flink SQL or Kafka Streams windowed aggregation with late events
- โซ A backfill: re-embed 100k chunks with a new model with zero downtime (versioned index + alias swap)
๐ง Cognitive tasks
- Fermi: embedding cost + time for 1M docs; verify on 10k
๐ฐ๏ธ Orbit integration
- orbit-knowledge ingestion + orbit-analytics
Go deeper
๐งฉ Distributed & Cloud Patterns (pipes & filters, claim check)
Resources
- Fundamentals of Data Engineering (Reis & Housley) โญ ยท Streaming Systems (Akidau et al.)
- DataTalksClub Data Engineering Zoomcamp (free) โญ ยท Debezium tutorial ยท ClickHouse docs