Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's more that the model is so large it is capable of memorizing a lot. This can be seen in other language models like GPT-3 as well.

Comments, I suspect, will be more likely to be memorized since the model would be trained to make syntactically correct outputs, and a comment will always be syntactically correct. That would mean there is nothing to 'punish' bad comments.



The model in this case is just a lossy compression of github, and you search that.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: