Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Which quant do you use? I have a similar setup and the speed is atrocious at 4-bit.


I'm using 4-bit as well, with the MoE model. I also use the MLX versions which are optimized for Apple CPUs (from what I understand anyway, I'm just an LLM layman). According to my oMLX dashboard, I'm getting about 50 tokens per second out of this model – not blazing fast, but more than fast enough to be useful to me.

https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-OptiQ-4...




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: