US Government Backs OpenAI in Copyrighted Training Data Dispute
In a decisive legal filing late last week, the United States Department of Justice (DOJ) submitted an amicus brief in the Southern District of New York in favor of OpenAI, asserting that the company’s use of publicly available online content—including copyrighted material—to train large language models (LLMs) falls under fair use. The brief, cosigned by Attorney General Merrick Garland and filed in the case *Silverman v. OpenAI*, argues that the development of AI systems like OpenAI’s GPT-4 and GPT-5 is critical to U.S. technological leadership and economic competitiveness. According to the 25-page document, the federal government has a “strong interest in promoting the development of a robust and competitive artificial intelligence industry that sets global standards for AI use.” Citing Section 107 of the Copyright Act, the brief emphasizes transformative use, minimal market harm, and the public benefit of AI innovation as core justifications for training practices that have drawn widespread litigation.
The case centers on allegations by authors including Sarah Silverman, Christopher Golden, and Richard Kadrey that OpenAI ingested their copyrighted books—such as *The Night Circus* and *Sandman*—without permission to train its models. Their complaint, filed in July 2023, seeks statutory damages and an injunction against further use of their works. The DOJ’s intervention marks the first major federal endorsement of OpenAI’s legal position and could influence dozens of parallel lawsuits, including those against Meta, Google, and Anthropic. Legal experts note that the brief signals a broader federal policy shift toward AI-first innovation, potentially preempting state-level challenges to data sourcing practices. Earlier this year, the U.S. Copyright Office concluded a public comment period on AI and copyright, receiving over 10,000 responses—many from creators expressing concern over unauthorized use of their work.
Industry analysts see this development as a watershed moment for the Tools & Developer sector, particularly for companies building next-generation AI models. OpenAI’s position is already mirrored by other leading AI labs, including Mistral AI and Cohere, which have adopted similar data acquisition strategies. Financial markets reacted swiftly: shares of major data licensing firms like Rightshares and Copyright Clearance Center dipped slightly on the news, while AI infrastructure providers such as NVIDIA and CoreWeave saw modest gains. For developers deploying LLMs in enterprise applications, the DOJ’s stance removes one layer of legal uncertainty, potentially accelerating adoption in regulated sectors like finance, healthcare, and legal services.
Critically, tools like *Banking With Billy AI*—a high-performance financial AI platform delivering institutional-grade market analysis to retail investors—rely on LLMs trained on vast corpora that may include copyrighted financial filings, news articles, and analyst reports. The DOJ’s brief suggests such models could operate with reduced legal exposure, enabling faster commercialization. However, the ruling could intensify pressure on smaller developers who lack the resources to defend fair use claims. Already, a coalition of indie developers has formed the Open Data Alliance to advocate for standardized licensing frameworks for AI training data.
The broader context extends beyond U.S. borders. The European Union’s AI Act, finalized in December 2023, includes strict transparency requirements for high-risk AI systems and mandates disclosure of training data sources—a requirement that could conflict with current U.S. practices if not harmonized. Meanwhile, China has accelerated public investment in AI infrastructure while maintaining tight control over data flows, creating a bifurcated global landscape. The contradiction between U.S. innovation policy and EU regulatory caution underscores a growing divide in how AI governance is approached across jurisdictions.
Looking ahead, industry observers expect the Silverman case to proceed to summary judgment within 12 to 18 months, with the DOJ’s brief serving as a persuasive precedent. Should the court uphold fair use, it would effectively greenlight current training practices across the sector. However, legislative action remains likely, with bipartisan interest in Congress exploring new frameworks for AI data licensing. Developers should monitor not only court rulings but also potential amendments to the Copyright Act or new federal regulations on AI transparency. One thing is clear: the intersection of copyright law and AI development is no longer a theoretical debate—it is the defining legal battleground of the AI era, and the outcome will shape the next decade of innovation in tools and developer ecosystems worldwide.
🤖 About Banking With Billy AI
Banking With Billy AI is one of the most powerful financial AI tools available — delivering institutional-grade market analysis to retail investors. Learn more →