An analogy, a formula, a worked example
Richard Feynman used to say that he could count seconds in his head while reading, but was quite unable to do so as soon as he spoke. His colleague John Tukey was exactly the opposite: he counted by watching a ribbon of digits go by, and could therefore talk at the same time. Same task, two incompatible mental machineries. And each of them had assumed, until then, that everybody worked the way he did.
I often think of that story when I face a lecture hall. When an explanation does not get through, my first reflex is to give it again, more slowly. Except that the student who is losing the thread is no slower than the others. We offered a door that does not open for them, and we are offering the same one again.
Starting with an image
So I come in through an analogy. For the bias-variance trade-off, I talk about two tailors: one cuts the same suit for everyone, the other starts all over again as soon as his customer changes posture. The students laugh, and above all they remember. The image becomes a handrail to hold on to when the formula arrives.
Then, before any notation, I say the thing in one sentence. Entropy is disorder, average surprise. Gini is the probability of being wrong when classifying at random. That step is also a test for me: when I cannot manage it, it means I have not yet understood the notion well enough to teach it.
Not stopping there
This is where many “accessible” courses stop, and it is exactly what I want to avoid. After the image comes the real formula, whole, with no watered-down version. A student who has only been told a metaphor is left helpless in front of a paper or an exam question. Intuition prepares formalism; it does not replace it.
So I take the formula apart symbol by symbol, and I slip the mathematical reminders in at the very moment they are needed: the base-2 logarithm right under entropy, the derivative right before gradient descent. Placed in a preliminary chapter, nobody reads them again.
Then I apply it, on numbers. In my machine learning course, the decision tree is built in front of them on eight patients: the first split, the duel between the candidates, the gain computation, the recursion, the leaves. At the end, we run scikit-learn and obtain the same tree. That moment is worth ten slides of explanation: the object was built before their eyes, they did not receive it ready-made.
Then comes an exercise, corrected straight away, and the trap I know they will run into: data leakage in a cross-validation, standardisation computed before the sets have been separated.
Nothing is made up
On this point I do not compromise: no program output, no error message, no numerical result is written from memory on my materials. Everything comes from a real run. Even the error messages are produced on purpose, on real code, before being shown.
In practice, this forces me to write and run the notebook first, and the slides afterwards, from what I actually obtained. It takes longer, but the reason is simple: a student who types the code from the course and does not get the announced result does not conclude that the material is wrong. They conclude that they are bad at this. An approximate value can cost them weeks of confidence.
Telling a story, not running through a syllabus
A session does not stand on its own. Each one starts by recalling what we already know how to do and where we are stuck, and ends by announcing the limit that the next one will lift. “We know how to find the optimal path — provided we have the complete map. What do we do when we do not know where the obstacles are?” The course then moves forward like a narrative, with tension, instead of being a list of techniques.
The same examples come back from one session to the next. A tiny data set met in the chapter on trees reappears in regression, then in evaluation metrics — where we discover that a model we had judged mediocre was simply badly evaluated. Those reunions do more for understanding than any end-of-chapter summary.
What it costs
A lot of time, I will not hide it. A concept treated this way takes six passes instead of one, the examples have to be computed by hand and then checked by the code, and good analogies take a long time to find.
But a course built like that holds up. Those who do not come in through the image come in through the numbers. Those put off by the formula find it again after handling the object. Several doors are worth more than a single one, repeated louder.
This is also what we are trying to put into the adaptive learning platform we are developing at UCAD: adapting not the speed of a single explanation, but the door you come in through. Feynman and Tukey had noticed it while counting seconds. It remains for us to draw the consequences in our lecture halls.