Live online · 100 hours · Beginner to Advanced
AWS Data Engineering Masterclass — 100 Hours across Glue, EMR, Redshift, Kinesis and MWAA
Build a production data platform on AWS — batch, streaming, warehouse and orchestration — and walk into the DEA-C01 exam ready.
What you will be able to do
- Design a lakehouse on S3 with Glue Catalog, Lake Formation permissions and Iceberg tables
- Build ETL with Glue (Spark and Python shell), EMR Serverless and Step Functions
- Model and tune Redshift: distribution and sort keys, RA3, Serverless, Spectrum, materialised views
- Stream with Kinesis Data Streams, Firehose, MSK (Kafka) and Flink/Managed Service for Apache Flink
- Orchestrate with Amazon MWAA (Airflow) and monitor cost with CUR + Cost Explorer
- Pass AWS Certified Data Engineer – Associate
Curriculum — 100 hours across 10 modules
Sessions run live; every session is recorded. Labs are hands-on from module one.
- IAM: users, roles, policies, trust relationships, least privilege, permission boundaries
- VPC essentials: subnets, security groups, endpoints (why Glue needs them), NAT cost traps
- S3 deep dive: storage classes, lifecycle, versioning, encryption (SSE-S3/KMS), S3 Tables
- CLI, SDK (boto3), CloudFormation vs Terraform vs CDK
- Cost control: budgets, alarms, tagging, Cost Explorer, the CUR
- Lab: secure landing zone with budget guardrails
- Glue Data Catalog, crawlers, partitions, partition projection
- Parquet vs Iceberg vs Delta vs Hudi on AWS — current support matrix
- Apache Iceberg on AWS: snapshots, time travel, compaction, S3 Tables managed Iceberg
- Lake Formation: LF-tags, row and column filters, cross-account sharing
- Athena v3: engine, CTAS, workgroups, cost per TB scanned, query tuning
- Lab: convert a 300 GB Parquet lake to Iceberg and cut Athena spend 60%
- Glue Spark jobs, Python shell jobs, Ray jobs; worker types and DPU sizing
- DynamicFrame vs DataFrame; resolveChoice, relationalize, bookmarks
- Glue Studio, notebooks, interactive sessions; Glue Data Quality (DQDL)
- Triggers, workflows, blueprints; Glue Streaming
- Debugging: continuous logging, Spark UI, metrics, common OOM patterns
- Lab: incremental ingestion with job bookmarks and DQ rules that fail the build
- EMR on EC2 vs EMR Serverless vs EMR on EKS — decision framework and pricing
- Cluster sizing, instance fleets, spot strategy, managed scaling
- Bootstrap actions, configurations, EMRFS consistency, S3A tuning
- Running the PySpark jobs from your own repo on EMR Serverless
- Lab: same job on Glue vs EMR Serverless — benchmark and cost report
- Kinesis Data Streams: shards, on-demand mode, enhanced fan-out, KCL/KPL
- Firehose: buffering, dynamic partitioning, transformation Lambdas, format conversion
- Amazon MSK: Kafka on AWS, MSK Serverless, Connect, Glue Schema Registry
- Managed Service for Apache Flink: SQL and DataStream API, checkpointing, exactly-once
- Kafka fundamentals: topics, partitions, consumer groups, offsets, compaction, idempotent producers
- Lab: end-to-end IoT stream with replay and a poison-message DLQ
- Architecture: leader/compute nodes, slices, RA3 managed storage, Redshift Serverless
- Distribution styles (KEY/EVEN/ALL/AUTO), sort keys, compression encodings
- COPY and UNLOAD, Redshift Spectrum, zero-ETL from Aurora, data sharing
- Workload management (WLM/Auto WLM), concurrency scaling, query monitoring rules
- Materialised views, auto-refresh, federated queries; STL/SVL system tables
- Lab: model a 1 B-row fact table and take a report from 90 s to 4 s
- DynamoDB for data engineers: partition keys, streams, export to S3
- RDS/Aurora, zero-ETL integrations, Aurora → Redshift
- AWS DMS: full load + CDC, task tuning, validation, common failure modes
- AppFlow and third-party SaaS ingestion; Airbyte/Fivetran on AWS
- Lab: Oracle → S3 CDC pipeline with DMS and schema drift handling
- Airflow core: DAGs, operators, TaskFlow API, XComs, sensors, pools, SLAs
- Airflow 2.x vs 3.x changes; deferrable operators and the triggerer
- Amazon MWAA: sizing, plugins, requirements.txt, DAG deployment via S3 + CI
- AWS Step Functions: state machines, Map/Distributed Map, error handling, cost vs Airflow
- EventBridge scheduling and event-driven pipelines; Lambda patterns and limits
- Lab: production DAG with backfill, retries, alerting and data-quality gates
- KMS, encryption in transit/at rest, Secrets Manager, cross-account roles
- CloudWatch metrics/logs/alarms, CloudTrail, Glue/EMR observability
- PII handling: Macie, tokenisation, masking with Lake Formation
- FinOps for data: right-sizing, spot, storage tiering, query cost attribution
- Lab: build a cost dashboard for the platform you just built
- Reference architectures: batch lakehouse, streaming, hybrid; well-architected data lens
- Choosing between Glue / EMR / Redshift / Athena — a decision tree you can defend
- Capstone presentation and architecture review
- DEA-C01 exam guide walkthrough, 130 practice questions, two timed mocks
- AWS data engineer interview questions and resume framing
Hands-on projects
You leave with three portfolio projects you can demo in an interview — not toy notebooks.
Serverless retail lakehouse
S3 + Glue + Iceberg + Athena + QuickSight, with Lake Formation row-level security and an MWAA DAG orchestrating the whole thing.
Streaming IoT platform
MSK → Managed Flink → S3/Redshift with schema registry, dead-letter handling and a real-time dashboard.
On-prem Oracle → AWS migration
DMS full load + CDC into S3, Glue transformation to Iceberg, Redshift serving layer, and a cutover runbook.
Tools and technologies covered
Who this course is for
- Data engineers standardising on AWS
- ETL developers migrating from on-prem
- Cloud engineers adding data skills
- DEA-C01 candidates
Prerequisites
- SQL and basic Python
- No prior AWS experience needed — Module 1 covers IAM, VPC and S3 from zero
- An AWS account (free tier); we show how to set a $20 budget alarm before you touch anything
Frequently asked questions
Labs are designed for the free tier plus a few dollars. We set budget alarms in the first session and every lab ends with a teardown checklist. Expect roughly $15–$40 total across the course if you follow the teardown steps.
No — it is the data specialisation. You'll learn the IAM/VPC/S3 you need, but not EC2 fleet management or hybrid networking in depth.
Both. Module 5 teaches Kafka fundamentals properly (partitions, consumer groups, offsets, compaction, exactly-once) and then applies them on Amazon MSK, alongside Kinesis for comparison.
AWS Certified Data Engineer – Associate (DEA-C01). Note that the older Data Analytics – Specialty (DAS-C01) has been retired — do not buy prep material for it.
Venu Katragadda or a course advisor will call or WhatsApp you within one working day with the full syllabus, batch dates and fees. For anything urgent, WhatsApp +91-9247159150.
Free download · PDF
Download the full 100-hour syllabus
Every module, every hour, every lab and all three projects — the same document we hand to corporate clients. No email verification loop; the PDF downloads the moment you submit.
- 10 modules broken down topic by topic with hours
- The 3 portfolio projects in full
- Prerequisites, tools list and certification mapping
- Fees, EMI options, batch timings and the refund policy
What students say about Venu Katragadda
Verified Google reviews from Sreyobhilashi IT students. Read all 320+ reviews →
“Recently took Databricks classes with Venu to upskill in trending technologies, and the experience exceeded all expectations. While I initially sought guidance only on Databricks, Venu provided in-depth training across the entire ecosystem.”
Databricks · Cleared DE Professional Cert · Verified Google review
“I recently completed the Data Engineering course on Databricks and AWS. Venu Sir delivers instruction at the next level, focusing on high-performance learning. He explains every concept clearly and thoroughly, accompanied by practical examples.”
Databricks & AWS Training · Verified Google review
“I had a truly valuable experience with Venu's Spark training along with AWS & Azure Databricks Training. He is highly knowledgeable, and the sessions are very well structured with extensive hands-on coverage.”
AWS & Azure Databricks Training · Verified Google review
Foundation course or masterclass?
We run two tiers. Most people should start with the foundation course on our sister site and step up later — this page is the advanced one.
Foundation · databrickstraining.in
AWS Data Engineering Training
₹22,000
- 70 hours of live instruction
- Covers the job-ready core of the stack
- Best if you are new to the platform or changing careers
- Same trainer, same teaching style
Masterclass · this page
AWS Data Engineering Masterclass
₹28,000
- 100 hours — roughly 25–30 extra hours of depth
- Internals, performance tuning and cost engineering modules
- Three reviewed portfolio projects instead of guided labs
- Architecture review and certification drill included
- Best if you already work with the stack and want senior-level depth
Not sure which fits? WhatsApp +91-9247159150 and Venu Katragadda will tell you straight — including when the cheaper one is the right answer.
Related masterclasses
Databricks Masterclass
Go from SQL/Python basics to a production Lakehouse you built yourself — Delta Lake, Unity Catalog, DLT, Workf…
PySpark Masterclass
The deepest PySpark course we teach — internals, tuning, testing and streaming, not just the DataFrame API.…
Azure Data Engineering Masterclass
The Azure stack as it is in 2026 — Fabric-first, with ADF, Databricks and Synapse in their real-world places.…