Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

even 30B model is too large to large on local device (low end). meta should provide free hosted model api to use it.


Meanwhile those of us with 128GB RAM plus some VRAM don't have any good modern (last 8 months) open weights models to make use of all that. I don't care if it would run 5 tok/s, I want a smarter model than Qwen3.6 which avoids loops and can handle more context than 80k before crashing.


why don't you use the quantized version of kimi-k3


I do not see a quantized version of kimi-k3 on huggingface that can fit into 128GB. The smallest Unsloth q version is 594GB

https://huggingface.co/unsloth/Kimi-K3-GGUF




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: