Skip to content

Details

Ever wonder why we train regression models on mean squared error?
Most of us learned MSE as a rule to memorize: regression task → square the errors → minimize. But it's not an arbitrary choice. MSE falls straight out of probability theory — once you ask "what distribution generated this data?", the loss basically derives itself.
In this focused hour, we'll trace exactly where MSE comes from:

  • How MSE emerges from a Gaussian likelihood
  • What "minimizing squared error" is really doing under the hood
  • Why this reframing makes the loss landscape feel principled instead of arbitrary

Come with basic ML familiarity and your questions — this is meant to be interactive, not a lecture. You'll leave seeing a loss function you already use every day in a completely different light.
One focused hour. Bring a notebook.

You may also like