Yes, it's quantized (4 bit). Sure, it's... not quite as good as what's on offer via API. And sure, "up to" does a lot of work (I don't have an average/median for you but it feels fast to me).
But it's usable, fully local, fully private, and has no subscriptions and no operating costs other than electricity.
And once the context gets large, it slows down.
"up to" 150, is doing a lot of work there.