
Cheap Inference Squeezes Frontier AI Economics
Frontier labs must keep funding expensive model advances even as older capabilities become cheaper and harder to distinguish from the latest release.
Read the snack ↗
IntelligenceSnacksTopic
Model capabilities, compute, deployment, economics and the infrastructure used to operate AI.

Frontier labs must keep funding expensive model advances even as older capabilities become cheaper and harder to distinguish from the latest release.
Read the snack ↗
Filtering repetitive tool output before it reaches an AI agent can cut token use while preserving the context that matters.
Read the snack ↗
Bringing intelligence onto a user’s device lets agents work with local software, files and processing power instead of recreating the whole environment in the cloud.
Read the snack ↗
A new AI model can combine genuine improvement with a flattering comparison against an older model that seemed to deteriorate over time.
Read the snack ↗
A model can find an unexpected link between mathematical fields, but novelty depends on proving it hasn't merely recovered an overlooked result.
Read the snack ↗
Fast, specialised models can make sorting, routing and checking cheap enough to place throughout automated workflows.
Read the snack ↗
Dividing software work into bounded jobs can make cheaper models a practical alternative to a single request sent to the largest available model.
Read the snack ↗
A model may fit in a local AI machine yet still respond slowly if its memory cannot supply data quickly enough.
Read the snack ↗
A polished AI answer can create more risk than value when a business needs work it can safely rely on.
Read the snack ↗
Local models can handle everyday personal tasks, but hardware costs and limited awareness leave hosted frontier AI as the easier consumer choice.
Read the snack ↗