castform founder here. i'm personally a little against techniques like self-consistency/majority voting during rl training because they tend to result in the model's output distribution "sharpening" a lot. this means the model will lose it's exploration ability and probably won't be able to explore/discover new solution strategies, which can be harmful for both rl training + generalization to unseen cases