The AI Energy Squeeze: Why Intelligence Is Running into Physical Limits
Aug 14, 2026 / 23 min read
August 14, 2026 / 17 min read / by Irfan Ahmad
At the center of the AI copyright battle is a question that could reshape the economics of the industry: can AI companies train models on copyrighted material without permission, or must they license and pay for that content?
The copyright fight is no longer theoretical. Authors, publishers, artists and media companies are already in court arguing that AI companies used copyrighted work to train their models without permission or payment. The AI companies’ defense is equally clear.
They argue that training on publicly available material can fall under fair use, and that models learn patterns from data rather than simply reproduce the original works. This leaves the courts with a tricky question: when does learning from copyrighted material become infringement? The answer will shape not just future licensing rules, but the cost of building AI models and the bargaining power of content owners.
The scale of these disputes has also expanded quickly. OpenAI has faced lawsuits brought by authors and by media organizations, including a high-profile case filed by The New York Times, which alleges that its content was used without authorization to train models that can reproduce and compete with its work.
Meta is defending cases involving the use of copyrighted books in training datasets, where plaintiffs argue that large-scale ingestion of written works cannot be justified under existing interpretations of fair use. Anthropic has also been drawn into similar disputes, reflecting how broadly these questions extend across the industry.
These cases are testing a core assumption that has enabled the rapid scaling of modern AI systems, that models can be trained on vast amounts of text, images, and other content drawn from the internet without direct licensing agreements for each individual source. This has allowed companies to build systems with broad knowledge and generative capability at a speed that would not have been possible under more restrictive frameworks.
The legal question at the center of these disputes is relatively narrow in formulation but far-reaching in implication. It concerns whether the process of training an AI model on copyrighted material constitutes a transformative use, one that changes the nature of the original work sufficiently to fall within fair use, or whether it represents a form of reproduction that requires permission and compensation. Courts have addressed similar questions in earlier technological contexts, but the scale and nature of AI training introduce complexities that those precedents did not fully anticipate.
The outcome of these cases will not only determine liability in individual disputes but will influence how AI systems are developed, what data they can be trained on, and which organizations can afford to build them. If current practices are upheld, the existing model of large-scale data aggregation may continue with relatively few changes. If they are restricted, companies may need to license data at scale, fundamentally altering the economics of AI development and potentially reshaping the competitive landscape.
At stake is more than compliance. The structure of the industry itself is being tested, in a setting where technological capability, legal interpretation, and economic incentives intersect. The question is no longer only how powerful AI systems can become, but whether the way they have been built can be sustained under the frameworks that govern ownership and use of creative work.
The legal disputes now unfolding are rooted in a process that has, until recently, been treated as a technical detail rather than a public question. Training modern AI systems involves exposing models to vast amounts of data so they can learn patterns in language, images, and other forms of content. What distinguishes the current generation of models is not only their architecture, but the scale and diversity of the data on which they are trained.
One of the most widely used sources in this process has been datasets derived from the open web. Collections such as Common Crawl aggregate billions of web pages, capturing a broad cross-section of publicly accessible content, including news articles, blogs, forums, and documentation. These datasets provide a foundation that allows models to learn general language patterns, factual associations, and stylistic variation across domains.
Beyond general web data, training pipelines have also incorporated more curated sources. Public reporting and research have highlighted the use of large collections of digitized books, academic materials, and other structured text datasets.
In some cases, these collections include works that are protected by copyright, which is where the current legal challenges begin to intersect with the technical process. The inclusion of such material is not incidental as books and long-form writing provide depth, coherence, and structure that are difficult to replicate using shorter or more fragmented sources alone.
The scale of this ingestion process is difficult to overstate. Training a large language model involves processing hundreds of billions, and in some cases trillions, of tokens, units of text that allow the model to learn relationships between words, phrases, and ideas.
This scale is what enables the system to generate coherent responses across a wide range of topics, but it also means that the boundary between different sources of content becomes less visible once the model has been trained.
From a technical perspective, the process does not store or reproduce entire documents in a straightforward way. Models learn statistical relationships and patterns rather than maintaining a direct database of texts. This distinction is central to the argument made by AI companies, which frame training as a transformative process that extracts generalizable knowledge rather than copying specific works.
At the same time, the outputs of these systems have raised questions about how distinct that transformation is in practice. Instances where models can produce passages that resemble or closely track existing works have been cited in legal filings, particularly in cases involving journalism and long-form writing.
These examples are used to argue that the training process may, under certain conditions, retain or reproduce elements of the original material in ways that go beyond abstract pattern learning.
This tension between how the process is described and how it appears in specific cases lies at the heart of the current disputes. On one side, training is presented as an act of learning, analogous to how humans absorb information from reading and exposure. On the other, it is framed as a large-scale use of copyrighted material without permission, enabled by the ability to process and synthesize content at a scale no individual could replicate.
Understanding this process clarifies why the legal questions have become so significant. The issue is not only whether individual outputs infringe on specific works, but whether the entire method of training, drawing on vast, unlicensed datasets to build systems with commercial applications, fits within the boundaries of existing copyright law.
At the center of the current disputes is a legal question that carries broad implications for how AI systems can be developed. It concerns whether the use of copyrighted material in training models can be considered fair use, a doctrine in U.S. copyright law that allows limited use of protected works without permission under certain conditions, or whether it constitutes infringement that requires licensing and compensation.
Courts evaluating fair use typically consider several factors, including the purpose of the use, the nature of the original work, the amount used, and the effect on the market for that work.
In previous cases involving search engines and digital indexing, courts have sometimes accepted the argument that copying content to enable new forms of access can be transformative, meaning that it changes the function of the original material in a way that justifies its use. These precedents are often cited by AI companies to support the view that training models on large datasets fall within an established legal framework.
Companies such as OpenAI argue that training involves analyzing data to learn patterns rather than reproducing works in their original form, and that the resulting systems serve a different purpose from the content they were trained on. This position frames AI training as analogous to reading or research, where exposure to existing material informs the creation of new outputs without directly copying the source.
The opposing argument focuses on scale and substitution. Plaintiffs in cases involving Meta and other developers contend that the ingestion of entire works, including books, articles, and other copyrighted material, cannot be treated as incidental or limited use when it occurs at a massive scale.
They argue that models trained in this way can generate outputs that compete with or reduce demand for the original works, particularly in areas such as journalism, writing, and creative production.
This concern is central to the lawsuit brought by The New York Times against OpenAI and Microsoft, where the claim is not only that copyrighted material was used without authorization, but that the resulting system can produce content that draws directly from or substitutes for the original reporting.
The case brings into focus the fourth factor in fair use analysis, the effect on the market, which often carries significant weight in determining whether a use is permissible or not.
Legal scholars and practitioners have noted that existing precedents do not map cleanly onto the current situation. Earlier cases involving search engines, such as those addressing the indexing of web pages or the display of thumbnails, involved systems that pointed users back to the original content.
AI models, by contrast, can generate self-contained outputs that do not require users to access the source material, which changes how the relationship between the original work and the new use is evaluated.
The scale of training also complicates the analysis. Fair use has historically been applied to more limited contexts, where the scope of copying and its impact could be assessed more directly. In AI, the use of data occurs across billions of documents and interactions, making it more difficult to isolate individual instances of use or to assess their cumulative effect on markets.
This combination of factors has led to a situation in which both sides can draw on elements of existing law to support their positions, while the overall framework remains unsettled. Courts are being asked to interpret doctrines that were developed in earlier technological contexts and to apply them to systems that operate at a scale and level of abstraction that those contexts did not anticipate.
The outcome of these cases will depend not only on how the law is interpreted, but on how judges understand what AI systems are actually doing with copyrighted material. Whether training is treated primarily as learning or as copying could determine how far fair use extends.
If courts decide that training requires permission or payment, the impact will extend far beyond individual lawsuits. It could turn high-quality training data into a major recurring cost, strengthen the bargaining power of rights holders, and make large-scale AI development considerably more expensive.
The legal questions surrounding AI training are often framed in terms of fairness and precedent, yet their implications are fundamentally economic. The outcome of these cases will influence not only how companies use data, but how much it costs to build AI systems, who can afford to compete, and how value is distributed across the ecosystem.
The current model of AI development has been shaped by the ability to access large volumes of data without negotiating individual licenses for each source. This approach has allowed companies to train systems at scale, drawing on a wide range of materials to build models with broad capabilities. If courts determine that such use falls outside the scope of fair use, the cost structure of AI development could change significantly.
Licensing at scale introduces both direct and indirect costs. Directly, companies would need to compensate rights holders for the use of their material, potentially across millions of individual works. Indirectly, they would need to establish systems to identify, track, and manage the provenance of training data, ensuring that each source is accounted for and appropriately licensed. These requirements would add operational complexity alongside financial cost, altering the economics of model development.
For large organizations with access to capital and established relationships, these changes may be manageable. Companies such as Microsoft, Google, and Amazon already operate within ecosystems that include content, infrastructure, and distribution, giving them multiple ways to adapt. They can negotiate licensing agreements, invest in proprietary datasets, or integrate content partnerships into their existing platforms.
For smaller firms and new entrants, the impact could be more restrictive. The need to secure licensed data at scale would raise the barrier to entry, making it more difficult to build competitive models without substantial resources. This could reinforce the concentration already present in the industry, where a small number of players operate across the key layers of compute, models, and distribution.
There is also the question of how value would flow if licensing became a standard requirement. Content creators, publishers, and media organizations could gain new revenue streams, particularly if their material is recognized as essential to training high-quality systems. Some early agreements between AI companies and publishers suggest how this dynamic might evolve, with licensing deals providing both access to content and a framework for compensation.
At the same time, licensing introduces trade-offs. Restricting access to data may limit the diversity and breadth of training material, which could affect model performance in certain domains. Companies may prioritize high-quality, licensed datasets over broader but unlicensed sources, leading to systems that are more controlled but potentially less comprehensive. This shift could change how models are designed, trained, and evaluated.
The economic implications extend beyond individual firms. If the cost of building AI systems increases, it may influence how quickly new capabilities are developed and how widely they are distributed. Higher costs could slow down the pace of experimentation for smaller players while consolidating development within organizations that can absorb these costs. At the same time, clearer rules around data use could reduce legal uncertainty, making investment more predictable for those operating within the framework.
The outcome is not predetermined. Licensing could lead to a more balanced ecosystem in which creators are compensated, and access to data is governed by clearer agreements. It could also reinforce existing concentrations of power by raising the cost of participation. The direction depends on how legal decisions are implemented and how organizations adapt their strategies in response.
What is clear is that the resolution of these cases will shape the economic foundation of AI. The question is not only whether companies will be allowed to train on existing data, but what it will cost to do so, and who will be able to participate under those conditions.
The outcome of the current legal disputes will not produce a single, uniform shift. It is more likely to reshape the system along multiple dimensions at once, influencing how models are built, how data is sourced, and how different actors position themselves within the AI ecosystem.
One possible direction is the emergence of a more formalized data economy. If licensing becomes a central requirement, datasets that are currently treated as background inputs may begin to function as structured assets.
Publishers, media organizations, and content platforms could play a more active role in supplying training material under negotiated agreements, with pricing, access terms, and usage rights defined more explicitly. Early deals between AI companies and publishers suggest how this model might develop, with content becoming part of a managed supply chain rather than an unstructured resource.
Another direction involves a shift toward proprietary ecosystems. Companies that already control large volumes of user-generated or platform-based content may rely more heavily on their own data, reducing dependence on external sources.
Organizations such as Google, Meta, and Amazon operate services that generate continuous streams of data through search, social interaction, and commerce. In a more restrictive environment, these internal data flows could become a strategic advantage, allowing them to train and refine models without the same level of exposure to licensing constraints.
There is also the possibility of increased fragmentation. Different jurisdictions may interpret copyright law in distinct ways, leading to variations in how AI systems are trained and deployed across regions. Companies operating globally would need to navigate multiple regulatory environments, adapting their data practices to comply with local requirements. This could result in models that are trained on different datasets depending on where they are developed, introducing divergence in capability, coverage, or behavior.
At the same time, technical responses may evolve alongside legal changes. Efforts to filter, trace, or watermark training data could become more prominent, enabling companies to demonstrate compliance or to track how specific types of content influence model outputs. Synthetic data, generated by models themselves or through controlled processes, may also play a larger role, though it raises its own questions about quality, bias, and feedback loops.
Each of these directions reflects a different balance between openness, control, and cost. A more regulated system could provide clearer boundaries and more predictable incentives, while also introducing friction that slows experimentation. A more permissive system could preserve the current pace of development but leave underlying questions about ownership and compensation unresolved.
What connects these scenarios is that they shift attention away from the models alone and toward the structures that support them. The future of AI will not be determined solely by advances in architecture or performance. It will be shaped by how data is governed, how access is negotiated, and how the relationships between creators, platforms, and developers are defined.
The system that emerges from this period of adjustment may look familiar in some respects, retaining elements of the current model, while differing in others that are less visible but more consequential. The changes are likely to be incremental in implementation but significant in effect, influencing how AI is built, who builds it, and under what conditions it operates.
The legal disputes now unfolding around artificial intelligence are often framed as questions of compliance, whether specific uses of data fall within existing boundaries or require new forms of permission. Yet the stakes extend beyond individual rulings. What is being tested is the foundation on which the current generation of AI systems has been built, a model that relies on large-scale access to content as the raw material for learning.
The cases involving OpenAI, Meta, and Anthropic illustrate how quickly this question has moved from the margins to the center of the industry. Authors, publishers, and creators are not only challenging specific outputs. They are questioning whether the process that produces those outputs can continue under the same assumptions about data use and ownership.
The answer will not be binary. Courts are likely to interpret existing doctrines in ways that accommodate some forms of use while restricting others, creating a framework that evolves over time rather than a single, definitive resolution. Even so, the direction of that framework will influence how AI systems are developed, what data they rely on, and how costs and responsibilities are distributed across the ecosystem.
If current practices are largely upheld, the existing model of training on broad, publicly accessible data may continue, reinforcing the scale advantages already present in the industry. If they are constrained, the shift toward licensing, proprietary datasets, and controlled data pipelines could reshape both the economics of development and the balance of power between large firms and smaller entrants.
In either case, the question at the center of the debate remains the same. It concerns how a system that depends on learning from existing human expression fits within structures designed to protect ownership and reward creation. The resolution of this question will influence not only how AI evolves, but how the relationship between technology and creative work is defined in the years ahead.
The next phase of AI will not be determined by capability alone. It will be shaped by the terms under which that capability is allowed to exist. The future of AI will depend not only on what it can learn, but on what it is permitted to use.
Aug 14, 2026 / 23 min read
Aug 14, 2026 / 17 min read
Aug 14, 2026 / 24 min read