📄️ ONNX Runtime
Premium ONNX Runtime course: format, PyTorch and TensorFlow export, quantization, CPU, GPU and TensorRT providers, serving. Duration 4h, 40-question exam.
📄️ 1. The ONNX format
Module 1 of the ONNX Runtime premium course: read an ONNX graph, understand operator sets and opset versions, inspect a model with Netron, and see why an exchange format is worth its complexity.
📄️ 2. Export from PyTorch
Module 2 of the ONNX Runtime premium course: torch.onnx.export in practice, choosing the example input, declaring dynamic axes, naming inputs and outputs, and comparing the legacy tracer with the dynamo exporter.
📄️ 3. Export from TensorFlow
Module 3 of the ONNX Runtime premium course: convert a SavedModel or a Keras model to ONNX with tf2onnx, pick a signature, handle NHWC vs NCHW, and avoid the common pitfalls of the TensorFlow to ONNX bridge.
📄️ 4. Numerical equivalence
Module 4 of the ONNX Runtime premium course: compare outputs between the source model and its ONNX export, choose sensible tolerances, cover diverse inputs, spot half-precision drift, and use onnx.checker.
📄️ 5. Graph optimization
Module 5 of the ONNX Runtime premium course: operator fusion, constant folding, ONNX Runtime optimization levels, saving the optimized graph, and measuring the actual gain on the running models.
📄️ 6. ONNX quantization
Module 6 of the ONNX Runtime premium course: dynamic and static INT8 quantization, calibration datasets, QOperator and QDQ formats, hardware compatibility, and measured impact on the ResNet and the text encoder.
📄️ 7. Execution providers
Module 7 of the ONNX Runtime premium course: picking and ordering execution providers, session options, TensorRT engine cache, and detecting the silent CPU fallback that ruins production latency.
📄️ 8. Performance benchmarking
Module 8 of the ONNX Runtime premium course: a reproducible protocol with warmup, batches and percentiles; the onnxruntime_perf_test tool; a comparison table across PyTorch, ONNX Runtime CPU, CUDA and TensorRT.
📄️ 9. Unsupported operators
Module 9 of the ONNX Runtime premium course: reading the error, rewriting the faulty module, custom operators, changing the opset version, and the specific case of the text encoder.
📄️ 10. Serving an ONNX model
Module 10 of the ONNX Runtime premium course: shared inference session, thread configuration, training-identical preprocessing, minimal FastAPI, ONNX Runtime Web preview and the link to course 40.
📄️ 11. Recap and exam
Complete recap of the ONNX Runtime premium course: format, PyTorch and TensorFlow export, numerical parity, optimization, quantization, execution providers, benchmarking, workarounds, serving; then the 40-question exam.