Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Something they definitely understand is exactly what their pre-training data looks like - the raw text that goes into the initial runs of training the models.

Instruction tuning and RLHF is a bit more complex than that. I assume they maintain detailed logs all of those human-driven decisions about which responses were better.



Do they really know exactly what the raw text looks like? It seems so huge that no human could read all the text from each corpus. And the models have been fed text from many different languages. I doubt they have people who understand all the languages.

Regarding RLHF, I also hope that they kept the logs of all the human decisions. But since it was (at least partially) outsourced to African companies like Sama.com, do they really get back all the logs or just a new fine-tuned model?

But they must indeed at least know what is done with the text submitted to the prompt or to their API.

(I’m really not an expert so my questions may sound naive)


RLHF apparently manages to mostly force all responses in all languages to be quasihomogeneous. I'm not sure if that means they translated the RLHF data to as many languages as possible and then repeated it or if it's something more fundamental which applies regardless of input language.

Although asking it "What can you not talk about" in Japanese only responds correctly with gptv4, and each language gives you a different list of items to some degree (between 4 and 6 items i found).

Sadly trying to speak Klingon or Sindarin to it is dodgy at best


>Do they really know exactly what the raw text looks like?

I mean yea, it's too big. That said, in a post ad hoc fashion they do. When the model spits out weird crap at times, they can search the raw corpus and filter those strings out. There was an incident around this with Reddit counting forums and strange usernames that were added in tokenization, but later removed from weights leading to odd behavior when doing inference.


Its safe to assume that whoever OpenAI outsources to does not get access to the model. Collecting data and training models on it will be two different steps.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: