Deep Learning: neural networks explained without equations
This course explains how neural networks work and why they took over, using no equations and no code. It covers the mechanism, the main architectures, and — the part most courses skip — what deep learning genuinely costs and when it is the wrong tool.
What this course sets out to do: make architectures legible. After it, "a CNN with residual connections" and "a transformer with attention heads" describe machinery you can picture rather than words you recognise.
What this course does not do: have you build a network. The premium catalogue covers that with PyTorch notebooks and real datasets.
What you are about to discover
Course contents
| # | Lesson | Main goal | Time |
|---|---|---|---|
| 1 | What a neuron computes | The single operation everything is built from | 8 min |
| 2 | Why depth changes everything | Discovered features, hierarchy, and why non-linearity is essential | 9 min |
| 3 | CNNs: how machines see | Convolution, weight sharing, and why it suits images | 9 min |
| 4 | Sequences and transformers | From recurrence to attention, and why transformers won | 10 min |
| 5 | What it really costs | Data, compute, time — and why to fine-tune instead | 9 min |
| 6 | Recap and FAQ | Synthesis, an architecture guide, and 12 common questions | 6 min |
| 7 | Quiz and attestation | Validate what you learned with 5 corrected questions | 3 min |
Is this course for you?
- You have trained classical models and want to know when neural networks earn their complexity.
- You read "transformer", "attention" and "embedding" and want them to mean something.
- You are deciding whether deep learning is appropriate for a project and want an honest cost picture.
- You want to understand what is inside ChatGPT at the level of architecture rather than analogy.
Recommended first: Introduction to AI and Machine Learning. Mathematics for AI makes lesson 1 land faster.
What you will be able to do at the end
- Describe what a neuron computes in one sentence, and why the unit is deliberately simple.
- Explain why non-linear activations are the reason depth is useful at all.
- Say how a network discovers its own features, and why that removes the need for hand-crafted ones.
- Explain convolution and weight sharing, and why they suit images specifically.
- Describe attention and why transformers displaced recurrent networks.
- Give a realistic account of the data, compute and time deep learning requires.
- Explain transfer learning, and why training from scratch is rarely the right choice.
Estimated time
Around 50 to 60 minutes of reading.
Prerequisites and next steps
Prerequisites: Introduction to AI and ideally Machine Learning.
Natural continuation:
- Natural Language Processing — deep learning applied to text
- Computer Vision — deep learning applied to images
- Large Language Models — where transformers end up
Frequent questions, answered in one line
Is a neural network modelled on the brain?
Loosely, and the analogy is now more misleading than helpful. The name and the original inspiration come from biological neurons, and the resemblance stops almost immediately: artificial neurons compute a weighted sum and a simple function, learn by an algorithm with no biological equivalent, and are arranged in regular layers nothing like cortical structure. Treating a network as a brain leads to wrong predictions about what it can do.
Why is it called "deep"?
Because of the number of stacked layers. A network with one or two hidden layers is shallow; modern networks have dozens to hundreds. Depth is what allows a hierarchy of representations to form, and it only became trainable once residual connections solved the vanishing gradient problem covered in the maths course.
Is deep learning always better?
No, and reaching for it reflexively is the most common expensive mistake. On structured tables, gradient-boosted trees usually win while training in seconds and remaining explainable. Deep learning is the right answer when the input is raw — pixels, waveforms, free text — because that is when features cannot be designed by hand.
How many layers should a network have?
This is almost never a decision you make from first principles. In practice you take an architecture that is known to work for your kind of data, usually pre-trained, and adapt it. Designing depth from scratch is a research activity, not an applied one.
The premium catalogue has you training real networks in PyTorch, with datasets, projects and a verifiable certificate after a 40-question examination. Included in every paid plan.
Ready? Start with lesson 1 →