Data Engineering
Pipelines, warehouses, and data infrastructure
Everything you need to get hired as a data engineer. Design and build data pipelines, streaming systems, and analytics infrastructure. Spark, Airflow, dbt, NoSQL, cloud warehouses, and data quality at scale.
- Lessons: —
- Labs: —
- Projects: —
- Level: Beginner
Curriculum
- Data Engineering Foundations — Python for data, SQL deep dive, data modeling, Git, cloud basics, data lifecycle
- Batch Pipelines — ETL design, Spark, Pandas, data validation, scheduling, idempotency
- Data Warehousing — Dimensional modeling, Snowflake/BigQuery, dbt, optimization, cost management
- NoSQL & Polyglot Persistence — CAP theorem & consistency tradeoffs, document stores (MongoDB), key-value & wide-column (DynamoDB, Cassandra), graph databases (Neo4j), choosing the right store per workload
- Streaming Data — Kafka, stream processing, event-driven architecture, exactly-once semantics
- Orchestration — Airflow, DAG design, observability, retries, backfills, testing pipelines
- Data Quality & Governance — Data contracts, lineage, observability, privacy, compliance, great expectations
- Streaming & Real-Time Analytics — Real-time dashboards, stream joins, windowing, materialized views
- Cloud & Production — Docker, Kubernetes basics, CI/CD for data, Terraform, cloud cost optimization
- Portfolio & Career — Data engineering portfolio, resume, system design, take-home prep, interviews
Skills You Will Learn
- Spark
- Airflow
- dbt
- Snowflake
- Kafka
- NoSQL
Browse all courses · View pricing · DeVenture Academy