Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> And as anyone who’s spent time using Claude 3.5 Sonnet / GPT-4o can attest, these things really are useful and smart! (And, if these results hold up, O1 is much, much smarter.) This is a nerve-wracking time to be a knowledge worker for sure.

If you have to keep checking the result of an LLM, you do not trust it enough to give you the correct answer.

Thus, having to 'prompt' hundreds of times for the answer you believe is correct over something that claims to be smart - which is why it can confidently convince others that its answer is correct (even when it can be totally erroneous).

I bet if Google DeepMind announced the exact same product, you would equally be as skeptical with its cherry-picked results.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: