100-hour advanced batches — next batch starts soon +91-8500002025 WhatsApp

Live online · 100 hours · Beginner to Advanced

AWS Data Engineering Masterclass — 100 Hours across Glue, EMR, Redshift, Kinesis and MWAA

Build a production data platform on AWS — batch, streaming, warehouse and orchestration — and walk into the DEA-C01 exam ready.

100 hoursLive instruction
10 modulesStructured path
3 projectsPortfolio ready
SoonNext batch starts
AWS Data Engineering Masterclass — 100 hours live online, covering Amazon S3, AWS Glue, Glue Data Catalog, Lake Formation

What you will be able to do

  • Design a lakehouse on S3 with Glue Catalog, Lake Formation permissions and Iceberg tables
  • Build ETL with Glue (Spark and Python shell), EMR Serverless and Step Functions
  • Model and tune Redshift: distribution and sort keys, RA3, Serverless, Spectrum, materialised views
  • Stream with Kinesis Data Streams, Firehose, MSK (Kafka) and Flink/Managed Service for Apache Flink
  • Orchestrate with Amazon MWAA (Airflow) and monitor cost with CUR + Cost Explorer
  • Pass AWS Certified Data Engineer – Associate

Curriculum — 100 hours across 10 modules

Sessions run live; every session is recorded. Labs are hands-on from module one.

  • IAM: users, roles, policies, trust relationships, least privilege, permission boundaries
  • VPC essentials: subnets, security groups, endpoints (why Glue needs them), NAT cost traps
  • S3 deep dive: storage classes, lifecycle, versioning, encryption (SSE-S3/KMS), S3 Tables
  • CLI, SDK (boto3), CloudFormation vs Terraform vs CDK
  • Cost control: budgets, alarms, tagging, Cost Explorer, the CUR
  • Lab: secure landing zone with budget guardrails

  • Glue Data Catalog, crawlers, partitions, partition projection
  • Parquet vs Iceberg vs Delta vs Hudi on AWS — current support matrix
  • Apache Iceberg on AWS: snapshots, time travel, compaction, S3 Tables managed Iceberg
  • Lake Formation: LF-tags, row and column filters, cross-account sharing
  • Athena v3: engine, CTAS, workgroups, cost per TB scanned, query tuning
  • Lab: convert a 300 GB Parquet lake to Iceberg and cut Athena spend 60%

  • Glue Spark jobs, Python shell jobs, Ray jobs; worker types and DPU sizing
  • DynamicFrame vs DataFrame; resolveChoice, relationalize, bookmarks
  • Glue Studio, notebooks, interactive sessions; Glue Data Quality (DQDL)
  • Triggers, workflows, blueprints; Glue Streaming
  • Debugging: continuous logging, Spark UI, metrics, common OOM patterns
  • Lab: incremental ingestion with job bookmarks and DQ rules that fail the build

  • EMR on EC2 vs EMR Serverless vs EMR on EKS — decision framework and pricing
  • Cluster sizing, instance fleets, spot strategy, managed scaling
  • Bootstrap actions, configurations, EMRFS consistency, S3A tuning
  • Running the PySpark jobs from your own repo on EMR Serverless
  • Lab: same job on Glue vs EMR Serverless — benchmark and cost report

  • Kinesis Data Streams: shards, on-demand mode, enhanced fan-out, KCL/KPL
  • Firehose: buffering, dynamic partitioning, transformation Lambdas, format conversion
  • Amazon MSK: Kafka on AWS, MSK Serverless, Connect, Glue Schema Registry
  • Managed Service for Apache Flink: SQL and DataStream API, checkpointing, exactly-once
  • Kafka fundamentals: topics, partitions, consumer groups, offsets, compaction, idempotent producers
  • Lab: end-to-end IoT stream with replay and a poison-message DLQ

  • Architecture: leader/compute nodes, slices, RA3 managed storage, Redshift Serverless
  • Distribution styles (KEY/EVEN/ALL/AUTO), sort keys, compression encodings
  • COPY and UNLOAD, Redshift Spectrum, zero-ETL from Aurora, data sharing
  • Workload management (WLM/Auto WLM), concurrency scaling, query monitoring rules
  • Materialised views, auto-refresh, federated queries; STL/SVL system tables
  • Lab: model a 1 B-row fact table and take a report from 90 s to 4 s

  • DynamoDB for data engineers: partition keys, streams, export to S3
  • RDS/Aurora, zero-ETL integrations, Aurora → Redshift
  • AWS DMS: full load + CDC, task tuning, validation, common failure modes
  • AppFlow and third-party SaaS ingestion; Airbyte/Fivetran on AWS
  • Lab: Oracle → S3 CDC pipeline with DMS and schema drift handling

  • Airflow core: DAGs, operators, TaskFlow API, XComs, sensors, pools, SLAs
  • Airflow 2.x vs 3.x changes; deferrable operators and the triggerer
  • Amazon MWAA: sizing, plugins, requirements.txt, DAG deployment via S3 + CI
  • AWS Step Functions: state machines, Map/Distributed Map, error handling, cost vs Airflow
  • EventBridge scheduling and event-driven pipelines; Lambda patterns and limits
  • Lab: production DAG with backfill, retries, alerting and data-quality gates

  • KMS, encryption in transit/at rest, Secrets Manager, cross-account roles
  • CloudWatch metrics/logs/alarms, CloudTrail, Glue/EMR observability
  • PII handling: Macie, tokenisation, masking with Lake Formation
  • FinOps for data: right-sizing, spot, storage tiering, query cost attribution
  • Lab: build a cost dashboard for the platform you just built

  • Reference architectures: batch lakehouse, streaming, hybrid; well-architected data lens
  • Choosing between Glue / EMR / Redshift / Athena — a decision tree you can defend
  • Capstone presentation and architecture review
  • DEA-C01 exam guide walkthrough, 130 practice questions, two timed mocks
  • AWS data engineer interview questions and resume framing

Hands-on projects

You leave with three portfolio projects you can demo in an interview — not toy notebooks.

1

Serverless retail lakehouse

S3 + Glue + Iceberg + Athena + QuickSight, with Lake Formation row-level security and an MWAA DAG orchestrating the whole thing.

2

Streaming IoT platform

MSK → Managed Flink → S3/Redshift with schema registry, dead-letter handling and a real-time dashboard.

3

On-prem Oracle → AWS migration

DMS full load + CDC into S3, Glue transformation to Iceberg, Redshift serving layer, and a cutover runbook.

Tools and technologies covered

  • Amazon S3
  • AWS Glue
  • Glue Data Catalog
  • Lake Formation
  • Amazon EMR
  • EMR Serverless
  • Amazon Redshift
  • Amazon Athena
  • Kinesis
  • Amazon MSK
  • Managed Flink
  • AWS Lambda
  • Step Functions
  • Amazon MWAA (Airflow)
  • DMS
  • Apache Iceberg
  • Terraform
  • CloudWatch

Who this course is for

  • Data engineers standardising on AWS
  • ETL developers migrating from on-prem
  • Cloud engineers adding data skills
  • DEA-C01 candidates

Prerequisites

  • SQL and basic Python
  • No prior AWS experience needed — Module 1 covers IAM, VPC and S3 from zero
  • An AWS account (free tier); we show how to set a $20 budget alarm before you touch anything

Frequently asked questions

Labs are designed for the free tier plus a few dollars. We set budget alarms in the first session and every lab ends with a teardown checklist. Expect roughly $15–$40 total across the course if you follow the teardown steps.

No — it is the data specialisation. You'll learn the IAM/VPC/S3 you need, but not EC2 fleet management or hybrid networking in depth.

Both. Module 5 teaches Kafka fundamentals properly (partitions, consumer groups, offsets, compaction, exactly-once) and then applies them on Amazon MSK, alongside Kinesis for comparison.

AWS Certified Data Engineer – Associate (DEA-C01). Note that the older Data Analytics – Specialty (DAS-C01) has been retired — do not buy prep material for it.

Thanks — your enquiry has reached us.
Venu Katragadda or a course advisor will call or WhatsApp you within one working day with the full syllabus, batch dates and fees. For anything urgent, WhatsApp +91-9247159150.

Get the AWS Data Engineering Masterclass syllabus and fees

Tell us where you are and we'll send the syllabus, batch dates and fees. No spam, no sales pressure.

Please enter your name.
Please enter a reachable number.
Please enter a valid email address.

By submitting you agree to be contacted about this programme. We never sell your data. Prefer to talk now? WhatsApp +91-9247159150 or call +91-8500002025.

AWS Data Engineering Masterclass syllabus PDF cover

Free download · PDF

Download the full 100-hour syllabus

Every module, every hour, every lab and all three projects — the same document we hand to corporate clients. No email verification loop; the PDF downloads the moment you submit.

  • 10 modules broken down topic by topic with hours
  • The 3 portfolio projects in full
  • Prerequisites, tools list and certification mapping
  • Fees, EMI options, batch timings and the refund policy
Downloading now. If it did not start, . We will also call or WhatsApp you within one working day.
Please enter your name.
Please enter a reachable number.
Please enter a valid email address.

We use your number to answer questions about the syllabus, not to spam you. Prefer to ask first? WhatsApp +91-9247159150.

What students say about Venu Katragadda

Verified Google reviews from Sreyobhilashi IT students. Read all 320+ reviews →

4.9 ★Google rating
320+Verified reviews
1,200+Professionals trained
14+ yrsTrainer experience
★★★★★

“Recently took Databricks classes with Venu to upskill in trending technologies, and the experience exceeded all expectations. While I initially sought guidance only on Databricks, Venu provided in-depth training across the entire ecosystem.”

Abhishek Zararia
Databricks · Cleared DE Professional Cert · Verified Google review
★★★★★

“I recently completed the Data Engineering course on Databricks and AWS. Venu Sir delivers instruction at the next level, focusing on high-performance learning. He explains every concept clearly and thoroughly, accompanied by practical examples.”

Mahaboob Mulla
Databricks & AWS Training · Verified Google review
★★★★★

“I had a truly valuable experience with Venu's Spark training along with AWS & Azure Databricks Training. He is highly knowledgeable, and the sessions are very well structured with extensive hands-on coverage.”

Dhevipriya Marimutbhu
AWS & Azure Databricks Training · Verified Google review

Foundation course or masterclass?

We run two tiers. Most people should start with the foundation course on our sister site and step up later — this page is the advanced one.

Foundation · databrickstraining.in

AWS Data Engineering Training

₹22,000

  • 70 hours of live instruction
  • Covers the job-ready core of the stack
  • Best if you are new to the platform or changing careers
  • Same trainer, same teaching style

View the foundation course →

Masterclass · this page

AWS Data Engineering Masterclass

₹28,000

  • 100 hours — roughly 25–30 extra hours of depth
  • Internals, performance tuning and cost engineering modules
  • Three reviewed portfolio projects instead of guided labs
  • Architecture review and certification drill included
  • Best if you already work with the stack and want senior-level depth

Get the full syllabus →

Not sure which fits? WhatsApp +91-9247159150 and Venu Katragadda will tell you straight — including when the cheaper one is the right answer.

₹28,000 Free syllabus PDF