📄️ Small language models
Ship a narrow-task assistant on the workstation with a 1-4B model: Phi, Gemma, Qwen, Llama, Mistral, quantization, LoRA, GGUF, offline privacy, and a full desktop project.
📄️ 1. Why size is not always the answer
Module 1 of the small language models premium course: narrow versus open tasks, cost per request, latency, privacy, and a first head-to-head between a 3B local model and a large API.
📄️ 2. Landscape of open small models
Module 2 of the small language models premium course: the Phi, Gemma, Qwen, Llama and Mistral families, their licenses, language coverage, and how to read a model card.
📄️ 3. Knowledge distillation
Module 3 of the small language models premium course: teacher and student, output distillation, synthetic-data distillation, and building a ticket data set from a large model.
📄️ 4. Pruning and quantization
Module 4 of the small language models premium course: structured and unstructured pruning, 8-bit and 4-bit quantization, measured quality loss, and when quantization suffices alone.
📄️ 5. Efficient runtime formats
Module 5 of the small language models premium course: GGUF and llama.cpp, ONNX Runtime, MLX for Apple silicon, conversion recipes, and choosing a format for the target hardware.
📄️ 6. Measuring latency and throughput
Module 6 of the small language models premium course: time to first token, tokens per second, batch size, a reproducible protocol, and readings on laptop CPU and entry-level GPU.
📄️ 7. Cheap fine-tuning of a small model
Module 7 of the small language models premium course: LoRA on a 1-4B model, ticket dataset, GPU-minute budget, and measured lift on classification and summary quality.
📄️ 8. Local and offline execution
Module 8 of the small language models premium course: installing on the workstation with Ollama or llama.cpp, memory sizing, model updates, and integrating with the ticket tool.
📄️ 9. Data privacy through local execution
Module 9 of the small language models premium course: what local inference protects, what it does not, regulatory arguments, and an honest pitch for management.
📄️ 10. Project: desktop assistant
Module 10 of the small language models premium course: assembling the ticket assistant end to end, evaluating against the API on 200 tickets, yearly cost, and routing to the large model.
📄️ Recap and exam
Complete recap of the small language models premium course: from head-to-head to distillation, quantization and desktop deployment, then the 40-question exam and the verifiable certificate.