📄️ Large language models
From a raw pretrained model to a customer-support assistant in production: scaling laws, instruction tuning, RLHF, decoding, context management, hallucinations, quantization, evaluation and cost.
📄️ 1. What scaling changes
Module 1 of the LLM premium course: emergent abilities and the debate around them, in-context learning, the timeline from GPT-2 to today, and open versus proprietary models.
📄️ 2. Pretraining and scaling laws
Module 2 of the LLM premium course: web-scale corpora and filtering, deduplication, the Chinchilla scaling laws, the true cost in compute and energy, and the share of your target language.
📄️ 3. Instruction tuning
Module 3 of the LLM premium course: from a base model to an instructed one, instruction dataset formats, conversation templates, and why example quality beats example quantity.
📄️ 4. Alignment: RLHF and DPO
Module 4 of the LLM premium course: the reward model, PPO, DPO as a direct alternative, refusal and over-refusal, and what alignment does not guarantee.
📄️ 5. Decoding parameters
Module 5 of the LLM premium course: the effect of each parameter on the distribution, repetition penalties, determinism, and recommended settings by use case.
📄️ 6. Context and memory
Module 6 of the LLM premium course: the cost of context, lost-in-the-middle, sliding summaries, external memory, and RAG as the standard answer.
📄️ 7. Hallucinations
Module 7 of the LLM premium course: why a model invents, grounding in sources, citation requirements, abstention, verification by a second model, and what does not work.
📄️ 8. Quantization and serving
Module 8 of the LLM premium course: 8-bit and 4-bit formats (GPTQ, AWQ, GGUF), memory by model size, vLLM with continuous batching, llama.cpp on CPU, throughput versus latency.
📄️ 9. Evaluation and benchmarks
Module 9 of the LLM premium course: MMLU and its limits, contamination, arenas and Elo, the LLM-as-a-judge pattern, and building your own business evaluation set.
📄️ 10. Costs and architecture
Module 10 of the LLM premium course: API versus self-hosting, cost per request, small-large model routing, caching, and the deployment decision for the running project.
📄️ Recap and exam
Complete recap of the LLM premium course: scaling, pretraining, instruction tuning, alignment, decoding, context, hallucinations, quantization, evaluation and costs, then the 40-question exam.