📄️ Natural language processing
A hands-on NLP course built on a corpus of 10,000 English customer reviews, from cleaning and tokenization to pretrained transformers, with a 40-question exam and a verifiable certificate.
📄️ 1. Cleaning and normalization
Module 1 of the NLP premium course: Unicode normalization, casing, punctuation, stop words, lemmatization versus stemming, and what should never be stripped.
📄️ 2. Subword tokenization
Module 2 of the NLP premium course: vocabulary and out-of-vocabulary words, BPE step by step on 20 words, WordPiece, SentencePiece, and token cost across languages.
📄️ 3. Bag of words and TF-IDF
Module 3 of the NLP premium course: sparse matrix, n-grams, TF-IDF by hand, a logistic regression baseline to beat, and why the model cannot see meaning or word order.
📄️ 4. Word2Vec and GloVe
Module 4 of the NLP premium course: Skip-gram and CBOW, negative sampling, vector analogies and their limits, and the biases baked into pretrained embeddings.
📄️ 5. Contextual embeddings
Module 5 of the NLP premium course: polysemy, from ELMo to BERT, sentence embeddings with Sentence-BERT, cosine similarity and semantic search over the review corpus.
📄️ 6. Text classification
Module 6 of the NLP premium course: fine-tune an encoder, handle class imbalance, read a confusion matrix on real reviews and compare against the TF-IDF baseline.
📄️ 7. Named entities
Module 7 of the NLP premium course: BIO tagging, token-to-label alignment with subword tokenizers, per-entity metrics, extracting products and places from reviews.
📄️ 8. Summarization and QA
Module 8 of the NLP premium course: extractive versus abstractive summarization, ROUGE and its blind spots, extractive question answering, and hallucination in abstractive models.
📄️ 9. The Transformers library
Module 9 of the NLP premium course: pipeline, AutoTokenizer, AutoModel, Trainer, and picking a model on the Hub by language coverage, size and licence.
📄️ 10. Project
Module 10 of the NLP premium course: end-to-end classifier on the English review corpus, English-specific handling, comparison of TF-IDF, embeddings and fine-tuning by cost and latency.
📄️ Recap and exam
Complete recap of the NLP premium course: cleaning, tokenization, TF-IDF, embeddings, contextual models, classification, NER, summarization, and the Transformers library, then the 40-question exam.