📄️ Ollama
Run open-weight language models on your own machines with Ollama: install, pull, tune, serve behind an OpenAI-compatible API, quantize for RAM, do private RAG. 4h, 40-question exam, verifiable certificate.
📄️ 1. Installing on Windows, macOS, Linux
Module 1 of the Ollama premium course: install the runtime on each operating system, find where models live, set the environment variables that matter, and verify the service.
📄️ 2. Pulling, listing, removing models
Module 2 of the Ollama premium course: navigate the model library, read tags, use pull, list, show and rm to manage disk space, and pick a model in the client's working language.
📄️ 3. Interactive session and parameters
Module 3 of the Ollama premium course: drive ollama run, use the session commands, and understand the measurable effect of temperature, num_ctx, num_predict, seed and stop.
📄️ 4. The local API
Module 4 of the Ollama premium course: the /api/generate, /api/chat and /api/embeddings endpoints, the OpenAI-compatible route, the Python client, and token streaming.
📄️ 5. Modelfiles and derived models
Module 5 of the Ollama premium course: author a Modelfile with FROM, SYSTEM, PARAMETER and TEMPLATE, create the firm's derived model, import a fine-tuned GGUF, and share a tag.
📄️ 6. Quantization and memory footprint
Module 6 of the Ollama premium course: read the quantization tags Q4_K_M and Q8_0, estimate the RAM a model needs by size and context, and compare quality on real questions.
📄️ 7. Hardware acceleration
Module 7 of the Ollama premium course: NVIDIA CUDA, AMD ROCm and Apple Metal, partial GPU layer offloading with num_gpu, and measuring tokens per second on each machine.
📄️ 8. Integrating with applications
Module 8 of the Ollama premium course: call Ollama from a business script, from LangChain, from Open WebUI, and expose it to the LAN safely with a reverse proxy and authentication.
📄️ 9. Local document Q&A
Module 9 of the Ollama premium course: local embeddings with nomic-embed-text, a Chroma index, a fully offline RAG chain over the firm's PDF archive, and quality in the front-end language.
📄️ 10. Limits of local execution
Module 10 of the Ollama premium course: quality versus large hosted models, concurrency across users, updates and security of the host, and when to move to a dedicated server or an API.
📄️ Recap and exam
Complete recap of the Ollama premium course: install, pull, tune, Modelfile, quantization, GPU, integration, local RAG and the limits of local execution, then the 40-question exam.