https://www.privacyworld.blog/2024/03/japans-new-draft-guide...
https://www.privacyworld.blog/2024/03/japans-new-draft-guide...
(Sure: the US government seems to act very friendly to one startup, OpenAI—but that's not the same as being friendly to OpenAI's competitors—it's sort of the opposite. OpenAI creating regulatory moats in the US is also OpenAI making the US a comparatively unattractive jurisdiction for everyone else..)
That said, there is ample indication that companies will continue to outsource to foreign talent (eg. Google laying off California workers and hiring Indians). Between outsourcing and remote work, it's increasingly possible to start a business in one location, but employee people globally.
I think the real outcome, which we're already seeing, is that companies pay for their training content. It fully integrates into the existing venture system ("simply" raise more money), integrates into existing laws, and feeds into the existing tech economy.
In order for the startup ecosystem to leave the US, there would need to be either free-flowing checks rivaling the US or a dramatic regulatory change. Japan just doesn't have either, even with this new policy.
I think Japan is in the right here. IMO, training an AI is analogous to reading a book. As long as the AI isn’t regurgitating the contents 1:1, it’s not infringement.
But for copyrighted content, it’s a thorny issue. You already cannot do whatever you want when you buy a book (or a cd, or a dvd); it’s for personal use unless you pay more. Surely you can quote some paragraphs or take a couple of screenshots, and you can learn the overall content, but you cannot reuse it in its entirety.
Also, the usage of those contents is kind of “rate limited” by the fact that an human must spit out those answers.
Now with AI you get a kind of an opaque system which “ingests” data in some way that can be considered “compression” (rather than “summarization” or “understanding”) and can spit this out at very large rates. Is this a fair thing? If you got a small model, probably your model has learned the concept, so I could agree that copyright is not a concern; but with LLMs? The NYY lawsuit is very relevant.
Also, I suspect that humans will adapt and create content which is less useful or harder to digest for AIs - maybe by making it worse for humans as well. Or they will use AIs to create bad content on their topics of expertise to make AI training data worse. There’s a serious risk of horrible content flooding the internet (and the publishing world).
LLMs are useful but risky. We should take care and consider 2nd order consequences.
You can sort of see this in Stable Diffusion and music models like Suno. Lots of creations are basically Japanese art.
In teasing out this apparent dichotomy, the committee first focuses on distinguishing between digesting copyrighted content for “information analysis” (which is allowed) versus the use for a “purpose of enjoyment of the thoughts and or sentiments expressed in the copyrighted work” (which is not allowed). The committee’s report states that “enjoyment” consists of satisfying the intellectual and mental needs of the individual through the viewing/experiencing the works. By contrast, information analysis might include a scholarly assessment of movie themes across various genres, while “enjoyment” would likely include the ability to play and view all or a significant part of the individual movie.
Interesting, I wonder how they will enforce it. I could envision a world where AI datasets become split between commercial / non-commercial, but this sounds like the dividing line is whether or not it the end work is "enjoyed".[0]: https://www.japantimes.co.jp/news/2024/04/18/japan/crime-leg...
You can train a model, but as for using it for commercial use that's a different story.