Lower Baselines Magnify AI Model Breakthroughs
A new AI model can combine genuine improvement with a flattering comparison against an older model that seemed to deteriorate over time.
Developed from a conversation between Pete Winn and Andy David

Release-day enthusiasm doesn’t measure a new AI model in isolation. Andy’s first response to the latest launches was to ask whether they represented real breakthroughs or whether users had developed collective amnesia after the previous generation appeared to get progressively worse. A strong new model can therefore arrive with two advantages at once, its own technical gains and a much easier comparison.
That shifting baseline magnifies the perceived jump. If a model feels excellent in its first month but less useful in months two, three and four, users adapt to the degraded experience. The next release is then judged against that recent frustration rather than the predecessor at its best. Renewed speed, clarity or reliability can feel transformative even when the underlying advance is more modest.
This doesn’t make the improvement imaginary. Andy accepted that both explanations could be true, with genuine progress under the hood amplified by contrast with a weakened baseline. Pete likewise expected some releases to share a base training run with different reinforcement learning or quantisation rather than represent entirely new multibillion-dollar training runs. The meaningful test is sustained use after launch, once the initial contrast fades and users can see whether the better experience lasts.
