기사 메일전송
[Analysis] Deepening Copyright Conflicts in the AI Era: What are the Solutions?
  • Kim Young
  • January 18, 2026 at 2:16 PM
기사수정
  • Government: "AI training needs immunity" vs. Media Industry: "Depriving creators of control"
  • The core of the "abuse of dominance" judgment is 'market substitutability'... Designing conditions, not immunity, is the solution.
  • Unlimited learning does not guarantee AI accuracy

'Guidance on Copyright Law for Generative AI's Learning Materials' Public Briefing [Photo=Yonhap News]

A conflict is fully unfolding between industry, media, and creator camps over the government's recent policy proposal for 'AI learning immunity.' 

 

The action plan announced by the Presidential Committee on Digital Innovation and Future Policy includes amendments to copyright law and the AI Basic Law to allow AI models to learn from copyrighted materials without legal uncertainty. 

 

In response, the Korea Newspaper Association issued an official statement of opposition, saying, "The 'first-to-use, then-compensate' approach is an unfair system that deprives creators of their prior control."

 

While seemingly a clash between industrial development and copyright protection, the core issue of this debate lies elsewhere. The key question is not "Can AI access content?" but rather "Under what conditions and through what procedures should access be granted?" 

 

This implies that it is a matter of system design, rather than a simple pro or con stance.

 

The Logic of the Government and Industry: "More Data Creates Better AI"

 

The government and the AI industry argue for the necessity of learning immunity based on the premise that AI performance improves with access to more data. 

 

The logic is that securing large-scale data is essential to strengthen the competitiveness of Korean AI. It is true that user experience tends to be richer with a wider range of references.

 

However, numerous recent studies point out that this intuition does not directly apply to AI learning processes. 

 

A paper by an international research team, 'Generative AI Training and Copyright Law,' analyzes that large-scale unauthorized learning does not automatically guarantee accuracy, and errors can become structured, especially when information from different points in time is mixed. 

 

A prime example is the 'temporal error,' a common weakness of generative AI. 

 

AI cannot distinguish between an article from 2019 and one from 2025 and cannot independently incorporate corrections or subsequent contexts. 

 

Consequently, research consistently concludes that indiscriminate learning carries a high risk of solidifying misinformation rather than improving accuracy.

 

Technical Reality: "Accuracy is a Matter of Access Method, Not Volume of Learning"

 

The study 'The State of Copyright in AI Training Datasets' analyzed actual training datasets and confirmed that despite the large inclusion of copyrighted news and professional content, temporal and contextual information is not systematically managed. 

 

This means that simply increasing the volume of training data does not guarantee quality improvement.

 

Another study, 'Copyright Detection in Large Language Models,' points out that without transparent management of the source and scope of use for training data, mechanisms for error correction or assigning responsibility cannot function. 

 

These studies collectively emphasize the need for conditional access and source-based reference structures, rather than unrestricted learning.

 

The perception that accurate AI is not the AI that has memorized the most, but rather AI that can access and verify reliable sources when needed, is gaining traction.

 

The Core of the Fair Use Debate: 'Market Substitution'

 

The government's proposal is considering broadly recognizing AI learning as fair use or as an exemption under TDM (Text and Data Mining). However, legally, this is a very complex issue. 

 

The key criterion for determining fair use has always been 'market substitution.' 

 

In the case of news content, generative AI goes beyond a simple analysis tool to substitute the consumption of articles itself. 

 

It is argued that if users can complete their information consumption solely through AI responses without reading the original text, this constitutes a typical case of market substitution. 

 

The Congressional Research Service (CRS) in the U.S. analyzed in a report that "the claim of transformative use in the AI training process alone does not automatically grant fair use," and "especially for commercial content like news, market impact assessment is a decisive factor." 

 

This means that fair use cannot be an unlimited exemption clause.

 

What Case Law Shows: 'Case-by-Case Judgment, Not Immunity'

 

Recent lawsuits in the U.S. and Europe demonstrate that this debate is not simple. 

 

In the lawsuit between Thomson Reuters and Ross Intelligence in the U.S., the court ruled that the use of copyrighted materials in the AI training process is not automatically recognized as fair use. 

 

The court particularly presented the potential for the training output to substitute the market for the original work as a key judgment criterion. 

 

In the lawsuit involving Anthropic, the court explicitly stated, "Large language model training is not exempted solely on the grounds of being transformative use." 

 

This shows that the 'learning is fair use' argument, often made by AI companies, is not automatically accepted in court. 

 

On the other hand, in some lawsuits filed against Meta, there were rulings that the use of training data could constitute fair use. 

 

However, these were also limited judgments that comprehensively considered the scope of learning, commerciality, and market impact. 

 

The message consistently conveyed by these precedents is clear: AI learning is not an area of immunity but an area of conditional design. 

 

The Gap Between Overseas Systems and Korean Discussions

 

The government emphasizes the need for TDM exemption by citing overseas cases such as the EU and Japan. However, there are significant differences in the detailed structures compared to the Korean discussion. 

 

While the EU recognizes TDM exemption, it mandates conditions such as the right for rights holders to opt-out, requirements for lawful access, and transparency obligations for training data. Japan also has broad TDM regulations but simultaneously employs control mechanisms for rights holders through contracts and terms of service. 

 

Conversely, these safeguards are not specifically presented in the Korean government's proposal. 

 

This is why the Newspaper Association criticizes it as "legalizing blind learning." Exemption without transparency is highly likely to become an unilateral privilege for specific platform companies. 

 

Attribution and Royalties Alone Are Insufficient

 

Some argue that attribution or royalty payments could be a solution. 

 

However, attribution only serves to enhance credibility and does not substitute prior permission. Royalties alone cannot restore the control rights of rights holders. 

 

For licensing to function properly, it must be accompanied by prior permission, clearly defined scope, transparent settlement, and verifiable usage records. Without these conditions, discussions on compensation risk becoming mere formalities. 

 

The Solution Lies in 'Phased Design'

 

The practical solution indicated by research and overseas cases is neither a complete ban nor unconditional permission. 

 

The key is phased design. 

 

This structure allows free learning from public data and expired copyrights, while professional content is accessed based on licenses, the latest news is utilized through real-time referencing and source attribution, and unauthorized fixed learning is strictly restricted. 

 

Balance can only be achieved when this is combined with rights holders' opt-out rights and transparency obligations for training data. 

 

This can lead to a new order of phased collaboration, rather than a structure where AI and human creators are in opposition. 

 

A Matter of Order, Not Technology

 

The copyright debate in the age of AI ultimately returns to the same question: even as technology rapidly advances, can the structure of rights and responsibilities be skipped? 

 

The experience of the internet era has already provided the answer. Unregulation led not to innovation but to market collapse, and the ecosystem stabilized only after licensing and responsibility structures were established. 

 

What is needed now is not the choice of "Should we lower copyright for AI?" 

 

The core issue is how to redesign the copyright order so that AI can grow accurately and sustainably.

 

 

Generative AI Training and Copyright Law

https://arxiv.org/html/2502.15858v1

 

The State of Copyright in AI Training Datasets

https://www.researchgate.net/publication/394962410

 

Copyright Detection in Large Language Models

https://arxiv.org/abs/2511.20623

 

Copyright Exceptions and Fair Use Defences for AI Training

https://www.cambridge.org/core/journals/european-journal-of-risk-regulation/article/752DF1DB564AD1EDFE23BA8BB1110802

 

Review of Copyright Law Issues Related to Text & Data Mining (TDM)

https://www.kci.go.kr/kciportal/landing/article.kci?arti_id=ART002993622

 

Thomson Reuters v. Ross Intelligence Case

https://www.reuters.com/legal/thomson-reuters-wins-ai-copyright-fair-use-ruling-against-one-time-competitor-2025-02-11/

 

AI Copyright Lawsuit Related to Anthropic

https://apnews.com/article/1e5cece51c2e4bd0bb21d94de2abb035

 

Fair Use Ruling Case Related to Meta

https://www.theverge.com/news/693437/meta-ai-copyright-win-fair-use-warning

 

EU Directive on Copyright in the Digital Single Market

https://en.wikipedia.org/wiki/Directive_on_Copyright_in_the_Digital_Single_Market

 

Generative AI Copyright Disclosure Act

https://en.wikipedia.org/wiki/Generative_AI_Copyright_Disclosure_Act

 

 

Generative AI

Artificial intelligence that creates new content such as text and images based on user requests. Training is the process of fixing data as statistical patterns within the model, and reference is the method of retrieving external data in real-time at the time of response. 

 

Fair Use

A standard of use under copyright law that is exceptionally permitted after comprehensively considering factors such as purpose, amount of use, commerciality, and market impact. 

 

TDM (Text & Data Mining)

The act of analyzing large amounts of text and data using computers. Some countries recognize conditional exemptions. 

 

Opt-Out

The right of a rights holder to request the exclusion of their copyrighted work from AI learning material. 

 

Temporal Error

An error that occurs when AI mixes past and present information. 

 

Transparency

The obligation to disclose and explain what data AI used and how. 


By Reporter Kim Young


관련기사
What do you think of this article?
recommend
0
great
0
moved
0
정기구독배너
Go to Mobile Site