07 · Linear regression, no libraries

15 lines of Python that learn. No imports. Ten delivery trips, each with a distance and a time. Start with a flat
line, measure every error, nudge two numbers, repeat a thousand times. The program learns
minutes = 2.98 × km + 4.21 and predicts that a 12 km trip takes 39.9 minutes.
Watch: YouTube · TikTok · Instagram (80 seconds, plus a 32-second short)
The idea
This is lesson 00 again, with every step written out by hand, so nothing hides inside NumPy:
- Guess a time for every trip with the current line.
- Check: how far off was each guess?
- Nudge
wandba little in the direction that shrinks the error. - Repeat.
Guess, check, nudge. The same loop trains every neural network; they just have far more numbers to nudge.
Run it
python3 src/delivery.py
mins = 2.98 * km + 4.21
12 km: 39.9 min
The short version (short/learn.py, 9 lines) nudges after every single trip instead of after each full pass, and
prints 12 km: 40.1 min.
Files
07-linear-regression-no-libraries/
├── src/
│ └── delivery.py the 17 lines from the video
└── short/
└── learn.py the 9-line version from the short
The ten trips are written straight into the code on purpose: the point of this lesson is that nothing is hidden, not even a file load.
The code
| Line | Code | What it does |
|---|---|---|
| 1 to 2 | km = [...], mins = [...] |
The data: ten trips. |
| 3 | w, b = 0.0, 0.0 |
Start flat. w = minutes per km, b = fixed minutes per trip (parking, handover). |
| 4 | lr = 0.02 |
Learning rate: the size of each nudge. |
| 5 | n = len(km) |
Number of trips, for averaging. |
| 7 | for epoch in range(1000): |
A thousand passes over the data. |
| 8 | dw, db = 0.0, 0.0 |
Start this pass’s tally of how to move w and b. |
| 9 | for x, y in zip(km, mins): |
Visit every trip. |
| 10 | err = w * x + b - y |
Guess minus truth for this trip. |
| 11 | dw += err * x / n |
How much w is to blame. Long trips count more. |
| 12 | db += err / n |
How much b is to blame. |
| 13 to 14 | w -= lr * dw, b -= lr * db |
The nudge. |
| 16 | print(...) |
The line it learned. |
| 17 | print(...) |
A prediction for a trip it has never seen. |
dw and db are the gradient of the mean squared error, written out by hand (up to a factor of 2, which the
learning rate absorbs).
What breaks it
In the video, raising the learning rate to 0.06 makes the numbers explode instead of settling. Each step
overshoots the bottom by more than the last. Picking a learning rate is a real part of training models.
Why use a model this simple?
Linear models are fast, cheap and explainable: “each km adds about 3 minutes” is something a dispatcher can check. Forecasting, pricing and capacity planning still use them every day, often as the baseline every fancier model has to beat.
Try this
- Set
lr = 0.06and printwevery 100 epochs. Watch it blow up. - Predict a 30 km trip. Do you trust that number? (It’s far outside the data.)
- Add an 11th trip that is way off, like
(5, 60). How much does one outlier move the line? - Compare with
numpy.polyfit(km, mins, 1). Same answer?
Previous: 06 · RAG from scratch · Next: Build Lab 01 · API to Parquet · All lessons