ETL / ELT Modernisation
Organisations are replacing legacy ETL tools with flexible platforms like PDI that support code-free job and transformation design, metadata injection, and dynamic pipeline generation.
Build implementation-level expertise across the Pentaho Data Integration (PDI) platform — from server installation and repository management through job and transformation design, Hadoop/Big Data integration, streaming, error handling, and performance tuning for production ETL environments.
HCE-5920 focuses on planning and implementing Pentaho Data Integration (PDI) environments. The learning path connects server installation and repository management, through PDI client/server architecture, job and transformation design, database connectivity, Hadoop and Big Data integration, error handling and logging, and finally performance tuning for production data pipelines.
The result is a complete PDI implementation skill set: understanding how jobs, transformations, metadata injection, streaming steps, and property files work together — and how to connect and optimise them across relational databases and Hadoop big data environments.
The challenges behind this syllabus — ETL/ELT modernisation, Big Data pipeline automation, governed data movement, and integration performance — are central to every modern data platform and analytics programme.
Organisations are replacing legacy ETL tools with flexible platforms like PDI that support code-free job and transformation design, metadata injection, and dynamic pipeline generation.
PDI's native Hadoop integration enables enterprises to build scalable data pipelines that process structured and unstructured data across HDFS, Hive, and distributed compute clusters.
Scheduling, property-file driven parameterisation, and programmatic execution methods allow PDI pipelines to be embedded in enterprise orchestration and CI/CD data workflows.
Streaming step optimisation and performance monitoring in PDI are key to meeting SLAs for high-volume, time-sensitive data integration and real-time reporting pipelines.
Install and configure the Pentaho server, manage the PDI repository, and set up the Data Integration client for enterprise ETL deployments.
Design PDI solution architectures — model data flows within jobs and transformations, select execution methods, and apply metadata injection for dynamic pipeline generation.
Manage data connections in PDI, create jobs and transformations step by step, implement streaming steps, and use property files for environment-driven configuration.
Configure PDI and Pentaho server for Hadoop integration, build Big Data PDI jobs and transformations, and apply key concepts for distributed data processing.
Implement error handling strategies in PDI transformations and jobs, and configure logging to capture diagnostic information for troubleshooting and audit purposes.
Monitor and tune the performance of PDI jobs and transformations to meet throughput, latency, and resource utilisation targets in production data environments.
Open a module to explore the objectives. Only one module stays open at a time for clean navigation of this technical syllabus.
Professionals designing and building PDI jobs, transformations, and metadata injection pipelines for enterprise data integration projects.
Engineers integrating PDI with Hadoop to build scalable data pipelines for processing structured and unstructured distributed data sets.
Admins responsible for Pentaho server installation, repository management, scheduling, and performance monitoring in production environments.
Architects designing PDI-based data integration solutions that connect relational databases, Hadoop clusters, and enterprise applications.
Explore all six technical modules and speak with Zetlan Technologies about the right learning path for your data integration and pipeline goals.