Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

whats about self-consistency like in grpo with majority voiting?


castform founder here. i'm personally a little against techniques like self-consistency/majority voting during rl training because they tend to result in the model's output distribution "sharpening" a lot. this means the model will lose it's exploration ability and probably won't be able to explore/discover new solution strategies, which can be harmful for both rl training + generalization to unseen cases




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: