Bias–Variance and How to Actually Diagnose Overfitting

Everyone can recite “high bias is underfitting, high variance is overfitting.” Far fewer can look at a training run and say which one they have and what to do about it. This post does both: first the bias–variance decomposition tightly enough to be useful, then a practical playbook — learning curves, the train/validation gap, cross-validation, and the traps — for diagnosing overfitting on a real model. It’s the diagnostic companion to the L1/L2 and dropout posts, which cover the fixes. ...

March 23, 2026 · 8 min

L1 vs L2 Regularization: Why Does L1 Produce Sparse Solutions?

“L1 gives you sparse weights, L2 gives you small weights” What Regularization Actually Is A model with enough capacity will happily drive its training loss to zero by memorizing the data — including the noise. That is overfitting: great training numbers, bad predictions on anything new. In a linear model it shows up as coefficients that blow up to large, opposing values. Why does fitting noise require large coefficients? Because of the amplification in the OLS solution \(\hat{\mathbf{w}} = (X^\top X)^{-1}X^\top y\). If \(X^\top X\) has a small eigenvalue (a direction of almost no variance in the data — e.g. two nearly identical features), its inverse has a large eigenvalue, and noise in \(y\) projected onto that direction gets amplified into a large coefficient. Concretely, if \(x_1 \approx x_2\), explaining \(y\) needs only \(w_1=1,\,w_2=0\) — but squeezing out the last bit of noise-fit might use \(w_1=1000,\,w_2=-999\). Since \(x_1 - x_2\) is a near-zero direction, you need enormous coefficients to produce the tiny output that matches the noise. Large \(|\mathbf{w}|\) is the fingerprint of a function contorting itself to fit noise. ...

March 2, 2026 · 11 min