Reasoning scores climb while per-token compute drops. Labs are stripping world knowledge out of models, and I think that's the right trade.