Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Benchmarks have gotten great, but they're still a proxy for the real world. The 3 GPT 5.6 models are also further apart in reality than the numbers suggest. That said, I'm still mighty impressed how good Luna is for the price. Highly underrated model.

I have been trying to build something that captures the behavioral element of different models, but it's kinda tough.

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: