Probability and Machine Learning
Details
Ever wonder why we train regression models on mean squared error?
Most of us learned MSE as a rule to memorize: regression task → square the errors → minimize. But it's not an arbitrary choice. MSE falls straight out of probability theory — once you ask "what distribution generated this data?", the loss basically derives itself.
In this focused hour, we'll trace exactly where MSE comes from:
- How MSE emerges from a Gaussian likelihood
- What "minimizing squared error" is really doing under the hood
- Why this reframing makes the loss landscape feel principled instead of arbitrary
Come with basic ML familiarity and your questions — this is meant to be interactive, not a lecture. You'll leave seeing a loss function you already use every day in a completely different light.
One focused hour. Bring a notebook.
