Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

the benchmark I trust most is whether the model can explain its own pricing page without getting confused


Not even humans can do that, you're literally asking for something beyond AGI



finally a benchmark where the answer to "why is my bill wrong" is the physics of restaurant tables.

so we've quietly moved the goalposts past AGI to "does it understand what it charges for." the general intelligence we can skip, the billing one we can't.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: