Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

By post-train I presume you mean a finetune? Unless that's wrong (please correct me if so).

I haven't looked into model architecture people are working with for this stuff too deeply yet but I presume the core idea is fine-tuning a lightweight reasoning-enabled LLM specifically using search as a metric for training?



That or providing a concrete RL env for $your_search_corpus_etc_here




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: