100-hour advanced batches — next batch starts soon +91-8500002025 WhatsApp

Live online · 100 hours · Beginner to Advanced

PySpark Masterclass — 100 Hours of Apache Spark Internals, Tuning and Streaming

The deepest PySpark course we teach — internals, tuning, testing and streaming, not just the DataFrame API.

100 hoursLive instruction
10 modulesStructured path
3 projectsPortfolio ready
SoonNext batch starts
PySpark Masterclass — 100 hours live online, covering Apache Spark 3.5/4.0, PySpark, Spark SQL, Structured Streaming

What you will be able to do

  • Write idiomatic PySpark that the Catalyst optimiser can actually optimise
  • Read a Spark UI and fix skew, spill, small files and shuffle problems on your own
  • Build streaming pipelines with exactly-once guarantees and stateful operators
  • Unit-test Spark code with pytest and chispa and ship it through CI/CD
  • Run the same codebase on Databricks, AWS EMR, GCP Dataproc and Kubernetes
  • Answer the Spark internals questions that decide senior data engineering interviews

Curriculum — 100 hours across 10 modules

Sessions run live; every session is recorded. Labs are hands-on from module one.

  • Data types, comprehensions, generators, decorators, typing
  • Modules, packaging, virtual environments, uv/poetry, logging
  • Working with JSON, CSV, Parquet in pandas; pandas ↔ Spark interop
  • Lab: build a small ETL utility package

  • Cluster managers: standalone, YARN, Kubernetes, Databricks
  • Driver/executor memory model, cores, slots, dynamic allocation
  • RDD lineage, DAG, jobs/stages/tasks, narrow vs wide transformations
  • Catalyst optimiser, Tungsten, whole-stage codegen, AQE
  • Lab: run the same job on 1, 4 and 16 cores and explain the curve

  • Schemas: StructType, nested structs, arrays, maps, explode/flatten
  • All join types and strategies: broadcast hash, sort-merge, shuffle hash, bucketed
  • Window functions, ranking, running totals, sessionisation
  • UDFs vs pandas UDFs vs Spark built-ins — measured performance comparison
  • Handling nulls, dedup, data-quality checks, quarantine patterns
  • Lab: flatten a deeply nested JSON API payload into a star schema

  • Parquet internals: row groups, column chunks, statistics, predicate pushdown
  • Avro, ORC, JSON, Delta, Iceberg — when each wins
  • Partitioning vs bucketing; the small-file problem and how to fix it
  • Compression codecs: snappy vs zstd vs gzip benchmarks
  • Lab: cut a table's scan time 6x by re-laying-out the files

  • Micro-batch vs continuous; sources and sinks; foreachBatch
  • Checkpointing, offsets, idempotent writes, exactly-once semantics
  • Event time, watermarks, windowing (tumbling, sliding, session)
  • Stream-stream and stream-static joins; state store internals and RocksDB
  • Kafka integration: consumer groups, maxOffsetsPerTrigger, back-pressure
  • Lab: streaming sessionisation with late-arriving events

  • Reading the Spark UI end to end: stages, tasks, SQL tab, executor tab
  • Diagnosing skew and fixing it: salting, AQE, iterative broadcast
  • Spill, GC pressure, memory fractions, off-heap, serialization (Kryo)
  • Shuffle partitions: the maths behind spark.sql.shuffle.partitions
  • Caching and persistence levels — when caching makes things worse
  • Cost-based optimisation, statistics, ANALYZE TABLE
  • Lab: three broken jobs, three root-cause reports, three fixes

  • pytest fixtures for SparkSession; chispa DataFrame assertions
  • Property-based testing, fixtures from JSON, golden datasets
  • Project layout, config with pydantic/YAML, dependency injection
  • Linting and typing: ruff, mypy; pre-commit hooks
  • Lab: retrofit tests onto an untested legacy pipeline

  • spark-submit anatomy: deploy modes, --packages, --conf, --files
  • Running on AWS EMR (EMR Serverless too), GCP Dataproc, Kubernetes (Spark Operator)
  • Airflow DAGs for Spark: SparkSubmitOperator, EmrOperators, sensors, retries
  • CI/CD: build, test, publish artefact, deploy, smoke test
  • Lab: promote a job through dev → staging → prod with GitHub Actions

  • Spark Connect and Spark 4.0 changes; PySpark on pandas API
  • Arrow, vectorisation, Photon and native engines (Comet, Gluten, Velox)
  • Graph processing with GraphFrames; ML pipelines with MLlib overview
  • Iceberg and Hudi with Spark: table formats compared
  • Common anti-patterns: collect(), toPandas(), row-by-row loops

  • Capstone build and architecture review
  • 80+ Spark interview questions with model answers, from junior to architect level
  • Live coding drill: 12 timed PySpark problems
  • Databricks Certified Associate Developer for Apache Spark exam walkthrough
  • Resume writing for Spark roles and how to describe your projects

Hands-on projects

You leave with three portfolio projects you can demo in an interview — not toy notebooks.

1

NYC taxi 1.5 B rows

Full ETL over the public taxi dataset: partitioning strategy, broadcast joins, skew handling, and a benchmark report showing a 9x speedup.

2

Real-time fraud scoring

Kafka → Structured Streaming → stateful rules engine → Delta sink, with watermarking, late data handling and replay.

3

Tested and packaged PySpark app

A pip-installable PySpark project with config management, pytest suite, GitHub Actions CI and spark-submit deployment to EMR.

Tools and technologies covered

  • Apache Spark 3.5/4.0
  • PySpark
  • Spark SQL
  • Structured Streaming
  • Delta Lake
  • Kafka
  • Airflow
  • pytest
  • chispa
  • Docker
  • AWS EMR
  • GCP Dataproc
  • Databricks

Who this course is for

  • Python developers moving into big data
  • Data engineers stuck at the 'it works but it's slow' stage
  • Data scientists who need production-grade pipelines
  • Anyone preparing for Spark developer interviews

Prerequisites

  • Comfortable with Python basics (or take the 8-hour Python primer in Module 1)
  • Basic SQL
  • 8 GB RAM laptop; Docker used for local Spark

Frequently asked questions

This course is about Spark itself — internals, tuning, streaming and testing — and runs on Databricks, EMR, Dataproc and local Docker. The Databricks course is about the platform: Unity Catalog, DLT, Workflows, SQL warehouses and governance. Many students take both; there is a bundle discount.

Spark 3.5 as the baseline with a dedicated session on Spark 4.0 changes (Spark Connect, ANSI mode by default, VARIANT type). Deprecated APIs are called out explicitly as we go.

Not for the first half — labs run in Docker locally. From Module 8 you'll use free-tier AWS/GCP credits or Databricks Community Edition; we walk through setup and cost guardrails.

Only for reading. The course is Python-first, but you'll learn to read Scala Spark code because much of the ecosystem's source and Stack Overflow answers are in Scala.

Thanks — your enquiry has reached us.
Venu Katragadda or a course advisor will call or WhatsApp you within one working day with the full syllabus, batch dates and fees. For anything urgent, WhatsApp +91-9247159150.

Get the PySpark Masterclass syllabus and fees

Tell us where you are and we'll send the syllabus, batch dates and fees. No spam, no sales pressure.

Please enter your name.
Please enter a reachable number.
Please enter a valid email address.

By submitting you agree to be contacted about this programme. We never sell your data. Prefer to talk now? WhatsApp +91-9247159150 or call +91-8500002025.

PySpark Masterclass syllabus PDF cover

Free download · PDF

Download the full 100-hour syllabus

Every module, every hour, every lab and all three projects — the same document we hand to corporate clients. No email verification loop; the PDF downloads the moment you submit.

  • 10 modules broken down topic by topic with hours
  • The 3 portfolio projects in full
  • Prerequisites, tools list and certification mapping
  • Fees, EMI options, batch timings and the refund policy
Downloading now. If it did not start, . We will also call or WhatsApp you within one working day.
Please enter your name.
Please enter a reachable number.
Please enter a valid email address.

We use your number to answer questions about the syllabus, not to spam you. Prefer to ask first? WhatsApp +91-9247159150.

What students say about Venu Katragadda

Verified Google reviews from Sreyobhilashi IT students. Read all 320+ reviews →

4.9 ★Google rating
320+Verified reviews
1,200+Professionals trained
14+ yrsTrainer experience
★★★★★

“Recently took Databricks classes with Venu to upskill in trending technologies, and the experience exceeded all expectations. While I initially sought guidance only on Databricks, Venu provided in-depth training across the entire ecosystem.”

Abhishek Zararia
Databricks · Cleared DE Professional Cert · Verified Google review
★★★★★

“This training has exceeded my expectations. Venu explains concepts clearly and uses hands-on examples that make the content easy to understand. I am learning a lot and would definitely recommend.”

Pataballa N V Lakshminarayana
Databricks Training · Verified Google review
★★★★★

“I recently completed the Data Engineering course on Databricks and AWS. Venu Sir delivers instruction at the next level, focusing on high-performance learning. He explains every concept clearly and thoroughly, accompanied by practical examples.”

Mahaboob Mulla
Databricks & AWS Training · Verified Google review

Foundation course or masterclass?

We run two tiers. Most people should start with the foundation course on our sister site and step up later — this page is the advanced one.

Foundation · databrickstraining.in

Big Data Training (Hadoop, PySpark, Hive, Kafka)

₹18,000

  • foundation of live instruction
  • Covers the job-ready core of the stack
  • Best if you are new to the platform or changing careers
  • Same trainer, same teaching style

View the foundation course →

Masterclass · this page

PySpark Masterclass

₹26,000

  • 100 hours — roughly 25–30 extra hours of depth
  • Internals, performance tuning and cost engineering modules
  • Three reviewed portfolio projects instead of guided labs
  • Architecture review and certification drill included
  • Best if you already work with the stack and want senior-level depth

Get the full syllabus →

Not sure which fits? WhatsApp +91-9247159150 and Venu Katragadda will tell you straight — including when the cheaper one is the right answer.

₹26,000 Free syllabus PDF