<?xml version="1.0" encoding="utf-8" ?>
<!DOCTYPE FL_Course SYSTEM "https://www.flane.de/dtd/fl_course095.dtd"><?xml-stylesheet type="text/xsl" href="https://portal.flane.ch/css/xml-course.xsl"?><course productid="34058" language="en" source="https://portal.flane.ch/swisscom/en/xml-course/google-bbdp" lastchanged="2025-11-24T18:07:03+01:00" parent="https://portal.flane.ch/swisscom/en/xml-courses"><title>Building Batch Data Pipelines on Google Cloud</title><productcode>BBDP</productcode><vendorcode>GO</vendorcode><vendorname>Google</vendorname><fullproductcode>GO-BBDP</fullproductcode><version>3.0</version><objective>&lt;ul&gt;
&lt;li&gt;Determine whether batch data pipelines are the correct choice for your business use case.&lt;/li&gt;&lt;li&gt;Design and build scalable batch data pipelines for high-volume ingestion and transformation.&lt;/li&gt;&lt;li&gt;Implement data quality controls within batch pipelines to ensure data integrity.&lt;/li&gt;&lt;li&gt;Orchestrate, manage, and monitor batch data pipeline workflows, implementing error handling and observability using logging and monitoring tools.&lt;/li&gt;&lt;/ul&gt;</objective><essentials>&lt;ul&gt;
&lt;li&gt;Basic proficiency with Data Warehousing and ETL/ELT concepts&lt;/li&gt;&lt;li&gt;Basic proficiency in SQL&lt;/li&gt;&lt;li&gt;Basic programming knowledge (Python recommended)&lt;/li&gt;&lt;li&gt;Familiarity with gcloud CLI and the Google Cloud console&lt;/li&gt;&lt;li&gt;Familiarity with core Google Cloud concepts and services&lt;/li&gt;&lt;/ul&gt;</essentials><audience>&lt;ul&gt;
&lt;li&gt;Data Engineers&lt;/li&gt;&lt;li&gt;Data Analysts&lt;/li&gt;&lt;/ul&gt;</audience><outline>&lt;h4&gt;Module 1 - When to choose batch data pipelines&lt;/h4&gt;&lt;p&gt;
&lt;strong&gt;Description:&lt;/strong&gt; You will learn the critical role of a data engineer in developing and maintaining batch data pipelines, understand their core components and lifecycle, and analyze common challenges in batch data processing. You&amp;#039;ll also identify key Google Cloud services that address these challenges. &lt;/p&gt;
&lt;p&gt;Topics:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Batch data pipelines and their use cases&lt;/li&gt;&lt;li&gt;Processing and common challenges&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Activities:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Quiz&lt;/li&gt;&lt;/ul&gt;
&lt;h4&gt;Module 2 - Design and build batch data pipelines&lt;/h4&gt;&lt;p&gt;
&lt;strong&gt;Description:&lt;/strong&gt; You will design scalable batch data pipelines for high-volume data ingestion and transformation. You&amp;#039;ll also optimize batch jobs for high throughput and cost-efficiency using various resource management and performance tuning techniques. &lt;/p&gt;
&lt;p&gt;Topics:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Design batch pipelines&lt;/li&gt;&lt;li&gt;Large scale data transformations&lt;/li&gt;&lt;li&gt;Dataflow and Serverless for Apache Spark&lt;/li&gt;&lt;li&gt;Data connections and orchestration&lt;/li&gt;&lt;li&gt;Execute an Apache Spark pipeline&lt;/li&gt;&lt;li&gt;Optimize batch pipeline performance&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Activities:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Quiz&lt;/li&gt;&lt;li&gt;Lab: Build a Simple Batch Data Pipeline with Serverless for Apache Spark&lt;/li&gt;&lt;li&gt;Lab: Build a Simple Batch Data Pipeline with Dataflow Job Builder UI&lt;/li&gt;&lt;/ul&gt;&lt;h4&gt;Module 3 - Control data quality in batch data pipelines&lt;/h4&gt;&lt;p&gt;
&lt;strong&gt;Description:&lt;/strong&gt; You will develop data validation rules and cleansing logic to ensure data quality within batch pipelines. You&amp;#039;ll also implement strategies for managing schema evolution and performing data deduplication in large datasets. &lt;/p&gt;
&lt;p&gt;Topics:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Batch data validation and cleansing&lt;/li&gt;&lt;li&gt;Log and analyze errors&lt;/li&gt;&lt;li&gt;Schema evolution for batch pipelines&lt;/li&gt;&lt;li&gt;Data integrity and duplication&lt;/li&gt;&lt;li&gt;Deduplication with Serverless for Apache Spark&lt;/li&gt;&lt;li&gt;Deduplication with Dataflow&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Activities:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Quiz&lt;/li&gt;&lt;li&gt;Lab: Validate Data Quality in a Batch Pipeline with Serverless for Apache Spark&lt;/li&gt;&lt;/ul&gt;&lt;h4&gt;Module 4 - Orchestrate and monitor batch data pipelines&lt;/h4&gt;&lt;p&gt;
&lt;strong&gt;Description:&lt;/strong&gt; You will orchestrate complex batch data pipeline workflows for efficient scheduling and lineage tracking. You&amp;#039;ll also implement robust error handling, monitoring, and observability for batch data pipelines. &lt;/p&gt;
&lt;p&gt;Topics:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Orchestration for batch processing&lt;/li&gt;&lt;li&gt;Cloud Composer&lt;/li&gt;&lt;li&gt;Unified observability&lt;/li&gt;&lt;li&gt;Alerts and troubleshooting&lt;/li&gt;&lt;li&gt;Visual pipeline management&lt;/li&gt;&lt;li&gt;Congratulations: Course summary&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Activities:
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Quiz&lt;/li&gt;&lt;li&gt;Lab: Building Batch Pipelines in Cloud Data Fusion&lt;/li&gt;&lt;/ul&gt;</outline><objective_plain>- Determine whether batch data pipelines are the correct choice for your business use case.
- Design and build scalable batch data pipelines for high-volume ingestion and transformation.
- Implement data quality controls within batch pipelines to ensure data integrity.
- Orchestrate, manage, and monitor batch data pipeline workflows, implementing error handling and observability using logging and monitoring tools.</objective_plain><essentials_plain>- Basic proficiency with Data Warehousing and ETL/ELT concepts
- Basic proficiency in SQL
- Basic programming knowledge (Python recommended)
- Familiarity with gcloud CLI and the Google Cloud console
- Familiarity with core Google Cloud concepts and services</essentials_plain><audience_plain>- Data Engineers
- Data Analysts</audience_plain><outline_plain>Module 1 - When to choose batch data pipelines


Description: You will learn the critical role of a data engineer in developing and maintaining batch data pipelines, understand their core components and lifecycle, and analyze common challenges in batch data processing. You'll also identify key Google Cloud services that address these challenges. 

Topics:



- Batch data pipelines and their use cases
- Processing and common challenges
Activities:



- Quiz

Module 2 - Design and build batch data pipelines


Description: You will design scalable batch data pipelines for high-volume data ingestion and transformation. You'll also optimize batch jobs for high throughput and cost-efficiency using various resource management and performance tuning techniques. 

Topics:



- Design batch pipelines
- Large scale data transformations
- Dataflow and Serverless for Apache Spark
- Data connections and orchestration
- Execute an Apache Spark pipeline
- Optimize batch pipeline performance
Activities:



- Quiz
- Lab: Build a Simple Batch Data Pipeline with Serverless for Apache Spark
- Lab: Build a Simple Batch Data Pipeline with Dataflow Job Builder UI
Module 3 - Control data quality in batch data pipelines


Description: You will develop data validation rules and cleansing logic to ensure data quality within batch pipelines. You'll also implement strategies for managing schema evolution and performing data deduplication in large datasets. 

Topics:



- Batch data validation and cleansing
- Log and analyze errors
- Schema evolution for batch pipelines
- Data integrity and duplication
- Deduplication with Serverless for Apache Spark
- Deduplication with Dataflow
Activities:



- Quiz
- Lab: Validate Data Quality in a Batch Pipeline with Serverless for Apache Spark
Module 4 - Orchestrate and monitor batch data pipelines


Description: You will orchestrate complex batch data pipeline workflows for efficient scheduling and lineage tracking. You'll also implement robust error handling, monitoring, and observability for batch data pipelines. 

Topics:



- Orchestration for batch processing
- Cloud Composer
- Unified observability
- Alerts and troubleshooting
- Visual pipeline management
- Congratulations: Course summary
Activities:



- Quiz
- Lab: Building Batch Pipelines in Cloud Data Fusion</outline_plain><duration unit="d" days="1">1 day</duration><pricelist><price country="US" currency="USD">595.00</price><price country="IT" currency="EUR">650.00</price><price country="GB" currency="GBP">660.00</price><price country="CA" currency="CAD">820.00</price><price country="AT" currency="EUR">950.00</price><price country="SE" currency="EUR">950.00</price><price country="DE" currency="EUR">950.00</price><price country="FR" currency="EUR">790.00</price><price country="CH" currency="CHF">950.00</price></pricelist><miles/></course>