Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

People say Qwen overthinks because they analyzed the thinking traces, and Qwen finds the answer relatively quickly but then second guesses itself multiple times for another 20,000+ tokens. Regardless of what other models do, that's clearly overthinking.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: