Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Theft providing value doesn't excuse theft. These giant companies can absolutely afford to properly license their datasets, and if they don't want to license them, they can create them.

Mozilla has been creating an open dataset for text-to-speech/speech-to-text for years now with thousands of contributors, why can't Google do the same?



>Theft providing value doesn't excuse theft.

I don't know about that. One of the biggest problems IMO with most theft is that is is net value negative. Someone steals my $500 cell phone and fences it for $100. The fence resells it for $250 to some poor sap who can't use it because it is IMSI banned. Value is destroyed.

This is why many consider it okay for a starving man to steal bread to eat. His value (not starving) is greater than the cost to the baker (1 loaf of bread).


If you disallow “theft” (copying is not theft, let alone training), then as you said, only giant companies can afford to train datasets by licensing or creating data (e.g. Adobe, Getty, Disney).

These statements always ignore that libre AI models exist, and rely on the weakening of copyright laws to let small creators use AI and compete against corporations that will use AI regardless, because once again, they have the resources to license giant datasets and we don’t.


Probably, because they did not attain actual consent by the people whose data they use.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: