Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Cherry-picking is fun but most of them are real, verifiable facts that the models get... straight up wrong.

> 3c24b5fe "Debian Security Advisory DSA-180-1 describes a buffer overflow vulnerability involving Cyrus SASL usernames." TRUE Mostly True FALSE FALSE FALSE

This is false: https://lwn.net/Articles/13296/

> 801cb8c1 "Equal Measures 2030's 2024 SDG Gender Index provides a downloadable dataset that includes a field labeled 'required annual change'." TRUE Mostly True TRUE FALSE FALSE

This is false: https://equalmeasures2030.org/2024-sdg-gender-index/

This is the "confidently wrong" problem, and the reason that LLMs won't ever be taken seriously for anything but a few niche use-cases (like generating slop-code and pumping out marketing materials), where being wrong isn't the end of the world. Akin to how speech-to-text is wrong often enough that, while being a fun novelty, you don't see business units writing reports in Word using STT.

I would encourage everyone to skim through the real 1000-question dataset: https://lenz.io/research/llm-disagreement/data.csv



If the LLMs in this particular exercise were allowed to answer "I don't know" I expect they would have.


LLMs don't have the capability to say they don't know, because they don't know what they know. They are, after all, just next-token-predictors.

I just tried both queries with their same query format, just adding an "I Don't Know" label, against Gemini and Claude, and in no cases did they use that label. 2/4 answers were wrong though. But try it for yourself and see:

> Classify this claim as of today: "<claim>". Output exactly one label: True, Mostly True, Misleading, False, or I Don't Know. No explanations, no qualifiers.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: