Skip to main content

Natural Language Processing: how machines work with human language

Language is the thing humans do best and computers found hardest. This course explains why it was so hard, what changed, and where the field genuinely stands in 2026 — including what still does not work.

What this course sets out to do: make the pipeline legible. Text becomes tokens, tokens become vectors, vectors become predictions, and every one of those steps loses something you should know about.

What this course does not do: have you build an NLP system. The premium catalogue covers that with notebooks and datasets.


What you are about to discover​


Course contents​

#LessonMain goalTime
1Why language is hardAmbiguity, context and everything left unsaid8 min
2TokenisationHow text becomes units, and what that costs8 min
3Embeddings: meaning as geometryFrom word counts to vectors that capture sense9 min
4The tasksWhat NLP is used for, and which tasks are solved9 min
5Where it still failsFacts, bias, languages, and evaluation that lies9 min
6Recap and FAQSynthesis, a task guide, and 12 common questions6 min
7Quiz and attestationValidate what you learned with 5 corrected questions3 min

Is this course for you?​

  • You want to process text — reviews, tickets, contracts, documents — and want to know what is realistic.
  • You use language models and want to understand what happens before the model sees your text.
  • You keep meeting "token", "embedding" and "fine-tune" and want them precise.
  • You are evaluating an NLP vendor and want to ask questions that reveal something.

Recommended first: Deep Learning, particularly the lesson on transformers.


What you will be able to do at the end​

  • Explain why language resists computation, beyond "it is complicated".
  • Describe tokenisation and why models bill in tokens rather than words.
  • Explain what an embedding is and why similarity becomes arithmetic.
  • Say why contextual embeddings were a step change over fixed ones.
  • Name the tasks NLP performs and judge which are production-ready.
  • Recognise the failure modes: fabricated facts, inherited bias, and languages the field neglects.

Estimated time​

Around 45 to 55 minutes of reading.


Prerequisites and next steps​

Prerequisites: Introduction to AI and Deep Learning.

Natural continuation:


Frequent questions, answered in one line​

Is NLP the same thing as an LLM?

No. NLP is the field; a large language model is currently its most prominent tool. Plenty of production NLP does not involve an LLM at all: a fine-tuned classifier with a few hundred million parameters routinely beats a giant model on a specific classification task, runs far cheaper, and gives more predictable outputs.

Can NLP work on languages other than English?

Yes, and unevenly. English, Chinese, Spanish, French, German and a dozen others are well served. Thousands of languages have too little digitised text to train on, and models for them are markedly worse — a disparity that mirrors and reinforces existing inequalities in who technology serves well.

Does a model understand what it reads?

It captures statistical structure well enough to translate, classify and summarise, and it holds no meaning, intention or model of the world. The practical consequence is that it cannot notice when a question is absurd or when its own answer is impossible, which is why review matters wherever the output is acted upon.

What is the difference between NLP and text mining?

Largely emphasis. Text mining traditionally means extracting structured information and statistics from document collections. NLP covers that and also understanding and generation. In practice the terms overlap heavily and the distinction rarely matters.


Want to build NLP systems rather than read about them?

The premium catalogue covers fine-tuning, RAG and production text pipelines, with a verifiable certificate after a 40-question examination. Included in every paid plan.


Ready? Start with lesson 1 →