Skip to main content

Computer Vision: how machines turn pixels into decisions

A photograph means nothing to a computer. It receives a grid of numbers — millions of them — with no indication of where one object ends and another begins. That vision works at all is the result of one good idea applied at scale, and this course explains the idea, the tasks it enables, and the ways it still breaks.

What this course sets out to do: make you able to look at a vision project and say what is realistic, what data it needs, and what will go wrong first.

What this course does not do: train models. The premium catalogue covers that with notebooks, datasets and deployment.


What you are about to discover​


Course contents​

#LessonMain goalTime
1What a computer seesPixels, channels, and why this is harder than it looks8 min
2Convolution: the idea that workedLearned filters and hierarchies of features9 min
3The four tasksClassification, detection, segmentation, tracking9 min
4Modern architecturesFrom AlexNet to vision transformers and foundation models8 min
5Where it still failsDomain shift, adversarial inputs, bias, honest evaluation9 min
6Recap and FAQSynthesis, a decision guide, and 12 common questions6 min
7Quiz and attestationValidate what you learned with 5 corrected questions3 min

Is this course for you?​

  • You want to automate a visual inspection — defects, counting, sorting — and need to know what it takes.
  • You work with cameras or medical images and want to judge vendor claims.
  • You keep hearing "CNN", "YOLO" and "segmentation" and want them precise.
  • You are deciding whether a vision project is feasible before committing budget.

Recommended first: Deep Learning.


What you will be able to do at the end​

  • Explain what an image is to a computer and why that framing creates the difficulty.
  • Describe convolution in plain language and say why it beat hand-designed features.
  • Choose between classification, detection, segmentation and tracking for a given problem.
  • Say what transfer learning buys you and how much data you actually need.
  • Explain why domain shift is the most common cause of production failure.
  • Recognise adversarial examples and dataset bias as real engineering risks.

Estimated time​

Around 45 to 55 minutes of reading.


Prerequisites and next steps​

Prerequisites: Introduction to AI and Deep Learning.

Natural continuation:

  • Generative AI — producing images rather than reading them
  • MLOps — keeping a vision model alive in production
  • Ethics of AI — facial recognition and surveillance, in detail

Frequent questions, answered in one line​

Is computer vision solved?

For narrow, well-specified tasks with representative data, largely yes: reading a meter, spotting a missing component, counting vehicles. For open-ended visual understanding — "is anything unusual happening here?" — no, and not close. The gap between those two sentences is where most failed projects live.

Can a model work on video, or only still images?

Both. Video is usually processed as a sequence of frames, which makes single-frame models immediately applicable, and tasks that depend on motion — tracking an individual across frames, recognising an action — need models that carry information between frames. Video also multiplies compute cost by the frame rate, which drives most architecture decisions in practice.

Do I need a GPU?

To train, effectively yes, though renting one for a few hours is cheap. To run a trained model, often no: a small model can classify images on a phone or a Raspberry Pi, which is the whole point of the tooling covered in the premium track.

How accurate does a vision model need to be?

It depends entirely on the cost of each kind of mistake, and stating a single accuracy target is usually a sign the question has not been thought through. A model that flags possible defects for human review can be quite imprecise and still save money; a model that rejects parts automatically cannot.


Want to build vision systems rather than read about them?

The premium catalogue covers convolutional networks, detection, segmentation and edge deployment, with a verifiable certificate after a 40-question examination. Included in every paid plan.


Ready? Start with lesson 1 →