📄️ Apache Spark ML
Train models on tens of millions of rows with Apache Spark: architecture, DataFrames, ML pipelines, distributed cross-validation, tuning and production writes. 7 hours, a 40-question exam and a verifiable certificate.
📄️ 1. Spark architecture
Module 1 of the Apache Spark ML premium course: how a Spark job actually runs — driver, executors, tasks, stages and partitions — and when the whole apparatus is disproportionate for the problem at hand.
📄️ 2. RDD, DataFrame, Dataset
Module 2 of the Apache Spark ML premium course: the three data abstractions Spark exposes, the Catalyst optimizer that reshapes DataFrame queries, and why the DataFrame API is the sensible default in 2026.
📄️ 3. Transformations and actions
Module 3 of the Apache Spark ML premium course: how Spark defers work until an action forces it, the difference between narrow and wide transformations, when to cache, and how to read an execution plan.
📄️ 4. Reading and columnar formats
Module 4 of the Apache Spark ML premium course: reading tens of millions of rows correctly — Parquet vs CSV, explicit schemas, partitioned layouts and partition pruning on the flight-delays dataset.
📄️ 5. Transformers and estimators
Module 5 of the Apache Spark ML premium course: the two building blocks of MLlib — Transformer and Estimator — with VectorAssembler, StringIndexer, OneHotEncoder and StandardScaler applied to the flight-delays dataset.
📄️ 6. Pipelines and cross-validation
Module 6 of the Apache Spark ML premium course: chaining transformers and estimators with Pipeline, tuning hyperparameters with CrossValidator and ParamGridBuilder, and why preprocessing outside the pipeline leaks.
📄️ 7. Algorithms and constraints
Module 7 of the Apache Spark ML premium course: what MLlib actually contains — regression, trees, gradient boosting, k-means, ALS — what it lacks, and the alternatives when a needed algorithm is missing.
📄️ 8. Tuning and shuffle
Module 8 of the Apache Spark ML premium course: setting the partition count, choosing repartition vs coalesce, spotting key skew from the Spark UI, sizing executor memory and reading the shuffle stage.
📄️ 9. Writing and integration
Module 9 of the Apache Spark ML premium course: writing partitioned Parquet outputs, choosing the write mode, saving a trained pipeline and reloading it in a scheduled scoring job.
📄️ 10. End-to-end project
Module 10 of the Apache Spark ML premium course: the complete flight-delays project — Parquet ingestion, pipeline, cross-validation, tuning and production scoring — compared with the same pipeline in pandas on a sample.
📄️ Recap and exam
Complete recap of the Apache Spark ML premium course: architecture, DataFrame API, columnar reads, pipelines, algorithms, tuning and downstream integration — then the 40-question exam.