That would be on what the AI model generates, not what it is trained on.
The latter is where the contention is, and it's a valid argument. So much so that some companies are not using stolen information to build their models.
IBM for example indemnifies its models for its customers and has detailed information on where the sources came from to train them.
I'm not sure I understand this argument. If I go to the library every day for 10 years and learn everything there is to know about subject x I shouldn't be able to sell my skills to the world about it later because I didn't give the creators of the books I read any money?
Arguably LLM companies could have made large-scale deals with libraries and got the exact same knowledge (much, much more slowly). I wonder if people would have the same issues then? My guess is probably. Goes back to the meme that if libraries were proposed today there's no way they would ever be allowed.
> The latter is where the contention is, and it's a valid argument
It has been tried in court several ways already. Remember the lawsuit that forced Anthropic to use physical books? They tried to argue that the books couldn’t be trained on at all. It failed.
it's of wrong scale, you cannot claim derivative work when you literally ingressed sum of knowledge while destroying it in the process so your "competitors" cannot do the same
It’s called Derivative Work and it’s a good feature for IP law: https://en.wikipedia.org/wiki/Derivative_work
You wouldn’t like a world where companies could copyright knowledge and then prevent anyone else from making a derivative of that knowledge.
Imagine Wikipedia being taken down for sharing knowledge that another company wrote about first. It’s a bad idea.