Master advanced data engineering skills — warehousing with SQL and NoSQL, ETL/ELT with Hadoop and Spark, data governance, real-time IoT stream processing, and building data pipelines with Python and Apache Airflow.
D-DS-OP-23 — Data Engineering Pipeline Architecture
6
Modules
Complete D-DS-OP-23 scope
27+ hrs
Duration
Self-paced learning
Optimize
Level
Advanced data engineering
24/7
Support
Expert guidance
Exam At a Glance
🧪
D-DS-OP-23
Exam Code
Dell Data Engineering Optimize
🌐
English
Language
Exam available in English
⚙️
Optimize
Level
Advanced data engineering
🏅
Dell Proven
Credential
Professional certification
🔄
2 Years
Validity
Recertify to maintain status
⚡
D-DS-FN-23
Pre-req
Foundations knowledge recommended
What You Will Learn
Six Data Engineering Modules
Master the end-to-end data engineering stack — from warehousing and ETL through real-time streaming, governance, and building production Python pipelines.
Module 01
The Role of the Data Engineer
›Skills of a data engineer — programming (Python/Scala/SQL), data modelling, and pipeline design
›Role in an analytics project — bridging raw data sources and analysis-ready datasets
›Data engineer vs data scientist vs data analyst — responsibilities and collaboration model
›Career paths — analytics engineer, platform engineer, ML engineer, and data architect
›Pipeline best practices — idempotency, retry logic, dead letter queues, observability, and testing
Technologies You Will Master
Apache Spark
Apache Kafka
Apache Airflow
Apache Flink
Apache Storm
HDFS / YARN
Apache Hive
HBase / Cassandra
Python / pandas
Apache Atlas
Apache Ranger
Apache Knox
Sqoop / Flume
Pravega
EdgeX Foundry
IoT / MQTT
Apache Spark
Apache Kafka
Apache Airflow
Apache Flink
Apache Storm
HDFS / YARN
Apache Hive
HBase / Cassandra
Python / pandas
Apache Atlas
Apache Ranger
Apache Knox
Sqoop / Flume
Pravega
EdgeX Foundry
IoT / MQTT
Full Curriculum
6-Module D-DS-OP-23 Programme
Complete data engineering curriculum from data warehousing and ETL through streaming, governance, and production Python pipeline orchestration with Apache Airflow.
›Data engineer skills — Python/Scala/SQL programming, data modelling, pipeline design, and cloud platform knowledge
›Role in analytics projects — raw data ingestion, transformation, quality, and serving to consumers
›Collaboration with data scientists — providing cleaned feature-engineered datasets for model training
›Collaboration with data analysts — building reliable aggregation layers and semantic models
›Career path — analytics engineer, data platform engineer, MLOps engineer, and data architect trajectories
›Relational DB performance — indexing strategies (B-tree, hash, covering), query plans, and partition pruning
›Schema design — 3NF normalisation for OLTP, star schema and snowflake schema for OLAP workloads
›Spark Structured Streaming — watermarking for late data, output modes (append, complete, update), and stateful aggregations
›Apache Flink — event time vs processing time, watermarks and lateness handling, keyed state, and checkpointing
›Pravega — segmented stream storage, auto-scaling segments, and transactions for exactly-once guarantees
›EdgeX Foundry — device services, core services, export services, and edge-to-cloud data pipeline integration
›Python data engineering libraries — pandas, numpy, SQLAlchemy, Pydantic, Great Expectations, and Arrow
›Data structures for pipelines — iterators, generators, and lazy evaluation for memory-efficient processing
›Apache Airflow architecture — Scheduler, Executor (LocalExecutor, CeleryExecutor, KubernetesExecutor), Metadata DB, and Web UI
›DAG design patterns — fan-out/fan-in, dynamic DAG generation, task groups, and dataset-driven scheduling
›Pipeline best practices — idempotency by design, retry with exponential backoff, dead letter queues, data quality checks, and observability (OpenTelemetry, DataDog)