I think what I’m saying is a bit more nuanced than that. LLMs currently struggle with very “wide”, long-run reasoning tasks (e.g., the evolution over time of a million-line codebase). That isn’t because they are secretly stupid and their capabilities are all hype, it’s just that this technology currently has a different balance of strengths and weaknesses than human intelligence, which tends to more smoothly extrapolate to longer-horizon tasks.
We are seeing steady improvement on long-run tasks (SWE-Bench being one example) and much more improvement on shorter, more well-defined tasks. The latter capabilities aren’t “hype” or just for show, there really is productive work like that to be done in the world! It’s just not everything, yet.
We are seeing steady improvement on long-run tasks (SWE-Bench being one example) and much more improvement on shorter, more well-defined tasks. The latter capabilities aren’t “hype” or just for show, there really is productive work like that to be done in the world! It’s just not everything, yet.