The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models
Reasoning LLMs suffer a complete accuracy collapse beyond a complexity threshold and—counterintuitively—allocate less computation as problems get harder despite available token budget, suggesting the failure is architectural rather than a resource constraint.
Bleeding-edge research establishing a non-obvious failure mode: scaling test-time compute doesn't fix the hard-problem cliff. The three-regime model (simple/medium/hard) has direct implications for routing—standard LLMs for easy and hard extremes, reasoning models only in the medium band. Published June 2025, widely cited, not yet in this digest.