Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There is randomness in LLMs. Both papers authors probably ran the bench 1-N times. Depending on that, they might select an average, max, least, etc. They might also have discarded outliers.

Like the other person said 5% variation is probably expected

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: