Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's more complex than that, they are trained to model their language and understanding of a concept to match the author. There is not enough memory in the network to remember the actual text.

It's like giving a talented an artist a day with a painting, and then a day later ask them to precisely copy the painting from memory. Would they come close enough for it to be considered a forgery, or will it be a transformative reinterpretation? It will probably depend on the skill of the artist, and I feel it might go either way.



Whether a human artist can do this is largely irrelevant to the question of copyright, however. Selling a close copy of another artists’ work (unless they’re long dead) would likely be infringement.


I think the distinction is that you don't arrest someone for merely having the capability of reproducing a work, but after they've reproduced it and attempted to use it in an infringing manner. I don't know how applicable the analogy is given that LLMs don't "think", of course.


you might say that, but the literal benchmark of LLMs (or any supervised learning algorithm for that matter) is loss, or how much 'distance' is between their output, and the validation set, when seeded with the training set. With a loss of well below 1%, which is typical, it means it can pretty much recreate the training data.


I've not seen that number before, what paper is that where there's a loss of less than 1% on the validation/test set?




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: