Live online · 100 hours · Beginner to Advanced
GCP Data Engineering Masterclass — 100 Hours across BigQuery, Dataflow, Dataproc and Composer
BigQuery-first data engineering on Google Cloud, with real Beam pipelines and a Composer-orchestrated platform you build yourself.
What you will be able to do
- Design and tune BigQuery: partitioning, clustering, slots, BI Engine, materialised views, cost controls
- Write Apache Beam pipelines and run them on Dataflow in batch and streaming
- Run Spark on Dataproc and Dataproc Serverless, including Iceberg/Delta on GCS
- Build event pipelines with Pub/Sub, Datastream CDC and Dataflow templates
- Orchestrate with Cloud Composer (Airflow) and govern with Dataplex + IAM
- Pass the Google Cloud Professional Data Engineer exam
Curriculum — 100 hours across 10 modules
Sessions run live; every session is recorded. Labs are hands-on from module one.
- Projects, folders, organisation; IAM roles, service accounts, workload identity
- Cloud Storage: classes, lifecycle, uniform access, signed URLs
- Networking basics, VPC Service Controls, private Google access
- gcloud CLI, Terraform for GCP, billing and budget alerts
- Lab: project bootstrap with Terraform and a budget alarm
- Architecture: Dremel, Colossus, slots, shuffle; storage vs compute pricing
- Datasets, tables, views, external tables, BigLake, BigQuery Omni overview
- Partitioning (time/integer/ingestion) and clustering — with measured impact
- Loading data: batch load, Storage Write API, streaming inserts, transfers
- Standard SQL depth: window functions, arrays, structs, UNNEST, JSON, geography
- Lab: reduce a report's bytes scanned by 94%
- Query plan explanation, slot contention, reservations vs on-demand, editions
- Materialised views, BI Engine, caching, table snapshots and clones
- Scripting, stored procedures, UDFs, remote functions, BigQuery ML basics
- Time travel, fail-safe, data retention; row/column-level security and policy tags
- Cost guardrails: custom quotas, maximum bytes billed, INFORMATION_SCHEMA monitoring
- Lab: build a slot-usage and cost dashboard from INFORMATION_SCHEMA
- Beam model: PCollection, PTransform, ParDo, GroupByKey, Combine, side inputs
- Batch vs streaming; event time, watermarks, windows, triggers, accumulation modes
- Python and Java SDK basics; runners; Dataflow Prime and streaming engine
- Templates (classic and Flex), autoscaling, shuffle service, fusion and its pitfalls
- Error handling: dead-letter patterns, retries, update/drain of streaming jobs
- Lab: Project 2 streaming pipeline with late data and a DLQ
- Pub/Sub: topics, subscriptions (pull/push), ordering keys, exactly-once delivery, dead letters
- Pub/Sub Lite, BigQuery subscriptions, Cloud Storage subscriptions
- Kafka fundamentals: partitions, consumer groups, offsets, compaction, schema registry
- Managed Service for Apache Kafka on GCP; Confluent on GCP; when to prefer Kafka over Pub/Sub
- Datastream for CDC from MySQL/Postgres/Oracle
- Lab: CDC pipeline with schema-change handling
- Dataproc clusters vs Dataproc Serverless vs Dataproc on GKE
- Sizing, preemptible/spot workers, autoscaling policies, initialization actions
- Running PySpark jobs against GCS; connector tuning; Iceberg and Delta on GCS
- Dataproc Metastore, Hive/Spark SQL, migration from on-prem Hadoop
- Lab: run a Spark job on Dataproc Serverless and compare cost vs Dataflow
- Airflow core: DAGs, TaskFlow API, operators, sensors, XCom, pools, SLAs, backfills
- Composer 2/3: environment sizing, autoscaling, plugins, private IP
- GCP operators: BigQuery, Dataflow, Dataproc, GCS, Kubernetes
- Deferrable operators, dynamic task mapping, datasets/data-aware scheduling
- CI/CD for DAGs; testing DAGs with pytest
- Lab: orchestrate Projects 1 and 3 from one Composer environment
- Dimensional modelling in BigQuery; nested/repeated vs flat — the real trade-offs
- dbt on BigQuery: models, refs, tests, snapshots, incremental strategies, macros
- Looker Studio and Looker overview; semantic layer thinking
- Data contracts and freshness SLAs
- Lab: dbt project with tests and docs over the gold layer
- Dataplex: lakes, zones, data quality tasks, automatic discovery, catalog
- IAM deep dive for data: dataset ACLs, authorised views, policy tags, DLP
- Encryption, CMEK, VPC-SC perimeters for exfiltration control
- Monitoring, logging, error reporting; SLOs for pipelines
- Lab: implement PII masking and prove it with an access test
- Vertex AI for data engineers: feature store, pipelines, model endpoints
- BigQuery ML and Gemini in BigQuery: SQL-native ML and AI functions
- Vector search in BigQuery for RAG use cases
- Capstone presentation and architecture review
- Professional Data Engineer exam guide, 120 practice questions, two mocks
- GCP interview questions and resume framing
Hands-on projects
You leave with three portfolio projects you can demo in an interview — not toy notebooks.
BigQuery analytics platform
Ingest, model and serve a 2 TB dataset with partitioning, clustering, materialised views and a slot-reservation cost model — with before/after cost proof.
Streaming with Beam
Pub/Sub → Dataflow streaming pipeline with windowing, triggers, side inputs and BigQuery streaming inserts; includes a replay/backfill path.
CDC to lakehouse
Datastream from Cloud SQL → GCS → Dataproc Serverless (Iceberg) → BigQuery external tables, orchestrated by Composer.
Tools and technologies covered
Who this course is for
- Data engineers on or moving to GCP
- Analytics engineers who live in BigQuery
- Multi-cloud engineers adding GCP
- PDE certification candidates
Prerequisites
- SQL and basic Python
- GCP free tier account ($300 credit) — setup in session 1
Frequently asked questions
Neither blindly. You build the same pipeline both ways in Modules 4 and 6 and produce a cost/latency comparison, so you can defend the choice in an interview or design review.
The $300 free credit covers the entire course comfortably if you follow the teardown checklist at the end of each lab. BigQuery's 1 TB/month free query tier covers most SQL work.
No. We teach Beam in Python. Java examples are shown for reading, because some Dataflow templates and older documentation are Java-only.
It is case-study heavy and tests judgement more than syntax. The course includes two full timed mocks and a decision-framework session specifically for the scenario questions.
Venu Katragadda or a course advisor will call or WhatsApp you within one working day with the full syllabus, batch dates and fees. For anything urgent, WhatsApp +91-9247159150.
Free download · PDF
Download the full 100-hour syllabus
Every module, every hour, every lab and all three projects — the same document we hand to corporate clients. No email verification loop; the PDF downloads the moment you submit.
- 10 modules broken down topic by topic with hours
- The 3 portfolio projects in full
- Prerequisites, tools list and certification mapping
- Fees, EMI options, batch timings and the refund policy
What students say about Venu Katragadda
Verified Google reviews from Sreyobhilashi IT students. Read all 320+ reviews →
“Venu sir is very talented in training latest technologies by covering industrial typical use cases. I recommend Sreyobhilashi for learning latest big-data technologies on multi-cloud environments.”
Big Data & Multi-Cloud Training · Verified Google review
“Venu Sir's teaching style and methodology is so effective that it covers all important requirements and very advanced knowledge. His approach is purely industry-based — apart from technical teaching, he also guides for better career growth.”
Data Engineering Training · Verified Google review
“Best institute for Spark and cloud training, very helpful staff, they are available for queries 24x7. I sincerely recommend joining if anyone wants to learn any big data related technology.”
Spark & Cloud Training · Verified Google review
Foundation course or masterclass?
We run two tiers. Most people should start with the foundation course on our sister site and step up later — this page is the advanced one.
Foundation · databrickstraining.in
GCP Data Engineering Training
₹20,000
- core track of live instruction
- Covers the job-ready core of the stack
- Best if you are new to the platform or changing careers
- Same trainer, same teaching style
Masterclass · this page
GCP Data Engineering Masterclass
₹26,000
- 100 hours — roughly 25–30 extra hours of depth
- Internals, performance tuning and cost engineering modules
- Three reviewed portfolio projects instead of guided labs
- Architecture review and certification drill included
- Best if you already work with the stack and want senior-level depth
Not sure which fits? WhatsApp +91-9247159150 and Venu Katragadda will tell you straight — including when the cheaper one is the right answer.
Related masterclasses
Databricks Masterclass
Go from SQL/Python basics to a production Lakehouse you built yourself — Delta Lake, Unity Catalog, DLT, Workf…
PySpark Masterclass
The deepest PySpark course we teach — internals, tuning, testing and streaming, not just the DataFrame API.…
AWS Data Engineering Masterclass
Build a production data platform on AWS — batch, streaming, warehouse and orchestration — and walk into the DE…