Wait, I'm a bit confused by that. Are you saying that AI-generation itself is always considered public domain? That's weird to me, if that's the case. The responsibility of the code should always fall to the person who wrote it for criminal matters and the legal entity who owns it for civil matters. E.g. if you're writing a book with an LLM, and it's 90% similar to Harry Potter and the Chamber of Secrets, the copyright claim would be against you who published it, not the AI company
I guess the hard part about code is that when you go after a person, he has no way of knowing that he's accidentally re-written significant portions of an Oracle codebase, and so however that's handled pre-AI, it should be handled the same way post-AI
All that being said, training on copyrighted materials is brazenly just ignoring the spirit and the letter of the law, and that should hold these companies liable, if the law is to be applied consistently. But that is when politics enters into the mix, geopolitics, "National interest" etc. But imo it is here at this level where the argument should happen
@anthonykfranco Output from the AI models is non-copyrightable, in the US, so it's effectively public domain. No one can own copyright to it.
I'm not sure who would be at fault if you used the AI to generate something verbatim that is copyrighted by someone else. Maybe the AI company or maybe the person who then took that output and published it, or maybe be both.