Math4ML Lesson 8: Functions and Linear Regression

This lesson explores functions, predictions, and how we compare lines using their errors.

Watch: How a Prediction Machine Learns

Join Evie and Mark as they predict a delivery robot’s travel time, check its mistakes, and follow steps to improve the rule. This 4-minute video introduces the ideas before you try the activities below; the second half previews the model-improvement activity in Lesson 9.

Watch, pause, and try: Pause at 1:29 to compare a prediction with the recorded result, and at 3:15 to discuss how the rule should change. Then explore the lesson below.

Open the video in its own player

The Magic Vending Machine

Imagine a strange vending machine in your school hallway. It doesn’t take money — it takes numbers.

You put in a 2, and the machine gives you 5.
You put in a 3, and it gives you 7.
You put in a 4, and it gives you 9.

You start wondering, what’s the rule?

You notice something: every time your input goes up by 1, the output goes up by 2. One simple rule that fits these examples is:

output=2×input+1\text{output} = 2 \times \text{input} + 1

Or in math letters:y=2x+1y = 2x + 1

You’ve just discovered a function — a rule that maps an input to exactly one output.

The Magic Vending Machine Output Slot 🎁 — Enter a number 🔢 2 Examples 2 → 5 3 → 7 4 → 9 Hidden rule (function) output = 2 × input + 1 Add 1 to input → output +2 Try a new input below!
A function is a rule that maps an input to exactly one output. Here: y = 2x + 1.

What Is a Function?

A function is like a little machine that takes something in, does something to it, and gives something back.

You can think of it as: Input → Rule → Output

For example:

  • If the rule is “multiply by 3,” then 4 becomes 12.
  • If the rule is “add 5,” then 2 becomes 7.
  • If the rule is “multiply by 2 and then add 1,” then 2 becomes 5.

For each allowed input, a function gives exactly one output. Different inputs can still share an output.

Visualizing Functions

Functions can also be drawn as lines or curves on a graph.

For the rule y = 2x, each increase of 1 in x adds 2 to y. Changing the rule to y = 2x + 1 moves the entire line up by 1.

Each function draws a different picture, showing how inputs connect to outputs.


What Makes This “Linear”?

When a rule looks like:y=mx+by = mx + bit’s called a linear function.

  • m is the slope (how much y changes when x increases by 1)
  • b is the intercept (what y would be when x = 0)

For your vending machine:

  • slope m=2m = 2 (up 2 each step)
  • intercept b=1b = 1 (the “starting boost”)

Real Data Is Messier

Now imagine the machine is old and a little glitchy.

Sometimes it gives the “right” answer… but sometimes it’s off by 1.

You try again:

  • 2 → 5
  • 3 → 8 (huh?)
  • 4 → 9
  • 5 → 10 (weird…)

Our model still gives one prediction for each input. Observed outputs can differ because of measurement noise or factors the model leaves out. We want a useful approximation, not a promise of a perfect hidden rule.

But you can still ask:

What line best explains the pattern overall?

This is where linear regression comes in.


Linear Regression: Finding the Best-Fit Line

Linear regression is a method for finding the “best” line of the form:

y=mx+by = mx + b

that matches a set of data points.

Instead of a perfect vending machine rule, you have examples (data):

(x1,y1),(x2,y2),(x3,y3),…(x_1, y_1), (x_2, y_2), (x_3, y_3), \dots

And your goal is to choose mm and bb so that the line is as close as possible to the points.

Predict before clicking: What happens to every output when you click b + 1? Does the slope change? Try it below: every output rises by 1, while the slope stays 2. This explorer demonstrates the intercept; it is not automatically fitting a line to data.

Linear Function: y = 2x + b x y y = 2x + b, b = 0 Changing b shifts the line up/down.
Here the slope is fixed at 2. The intercept b controls where the line crosses the y-axis.

What does “best” mean?

One common definition is:

  • calculate each residual: observed output minus predicted output at the same input
  • square each residual and average the squares to get mean squared error (MSE); choose the line with the smallest MSE

The signed vertical gaps are called residuals. Squaring them stops positive and negative mistakes from cancelling. A smaller MSE means a better fit to these examples. We must still check predictions on new examples that were not used to fit the line.

Try It — Which Line Fits Better?

Use the noisy observations from our vending machine. Compare A: prediction = 2x + 1 with B: prediction = 2x. Guess which fits better, then calculate each residual and square it.

Input xObserved yA predictionB prediction
2554
3876
4998
5101110

Worked example: at x = 3, the observation is 8. A predicts 7, so its residual is 8 − 7 = 1 and its squared residual is 1. B predicts 6, giving a squared residual of (8 − 6)² = 4. Finish the other rows before checking.

Check your error scores

A has squared residuals 0, 1, 0, 1. Its MSE is 2 ÷ 4 = 0.5.

B has squared residuals 1, 4, 1, 0. Its MSE is 6 ÷ 4 = 1.5.

A fits better among these two candidates, even though it is not perfect. We have not proved it is the best line when both slope and intercept can change.

Explain your thinking: Why can a line with two mistakes still be the better choice? Consider the sizes of all the errors, not just the number of exact matches.

Thinking Like a Data Scientist

In machine learning, every model is like a complex function machine.
It takes inputs — like images, numbers, or text — and produces outputs — like labels, predictions, or answers.

A simple example:

  • Input: “weather data” → Output: “rain tomorrow” or “no rain”

A more advanced example

  • Input: “photo” → Output: “this is a cat.”

These examples predict categories, a task called classification. Linear regression predicts a number, such as an estimated temperature or sales amount. In both cases, training uses examples to adjust a model.


How This Connects to Machine Learning

Linear regression is one of the simplest machine learning models.

When you “train” it, you are really doing this:

  • You give the computer many examples (x,y)(x, y)
  • It searches for mm and bb that best match the data

In other words, it learns the rule:

y^=mx+b\hat{y} = mx + b

That hat (y^\hat{y}​) means “predicted y.”

Examples of real regression-style learning:

  • (hours studied, test score)
  • (age, height)
  • (ad spending, sales)
  • (temperature, ice cream sold)

It’s still the same idea as your vending machine — except now the rule is learned from messy data, not given perfectly.


Big Takeaway

  • A function is a rule that turns input into output.
  • A linear function looks like y=mx+by = mx + b.
  • Linear regression is how we find the best mm and bb when the world is noisy.
  • That “best-fit line” becomes a simple prediction machine — and it’s one of the foundations of machine learning.

Optional Extension: Connect to “Weights” and “Bias”

In ML language:

  • mm is like a weight
  • bb is like a bias

Training means adjusting the weight and bias until predictions match the examples as well as possible.

In Lesson 3, we used lines to predict. Here, we learned how functions turn inputs into predictions and how to score their errors. Next, Lesson 9 explores how an algorithm can adjust a model repeatedly.