Policy brief 28 · Economic Competitiveness

AI Training + Copyright Transparency

Show creators what went into the model, and how it got there.

One federal judge called training AI on books a fair use;1 two days later, another ruled for Meta but warned that such training will often be illegal.2 Fair use turns on facts that sit with developers: which works, from what source, for what purpose.3 Congress should build a licensing market, require developers to publish and keep training records, and let creators ask about their own work.

The Problem

In July 2026, a federal court approved authors' $1.5 billion settlement with Anthropic over pirated books, about $3,000 for each of 482,460 works.4 For writers, musicians, and artists, the larger question is whether systems trained on their work may compete with it without asking.

The courts don't agree. In that case, Judge William Alsup called training on books "spectacularly" transformative and a fair use, but refused to excuse more than seven million pirated copies kept for a central library.1 Two days later, Judge Vince Chhabria ruled for Meta on the record before the court, yet wrote that "in many circumstances it will be illegal" to train on copyrighted works without permission.2 A Delaware court held that a rival's use of Westlaw headnotes to train a legal-research tool was not fair use; the appeal was argued in June 2026.5

The Copyright Office says fairness "will depend on what works were used, from what source, for what purpose."3 Those facts sit with developers, and three gaps keep them there:

  1. Licensing is spotty. Deals are "fast emerging in certain sectors," the Copyright Office found, but "their availability so far is inconsistent."3
  2. Creators can't check. Outside a lawsuit, a creator cannot learn whether their work trained a model. California began requiring public training-data summaries in 2026, and one developer is suing to block the law.6
  3. The record decides. In the Meta case, 13 authors lost because they "failed to develop a record" on the argument that mattered, the judge wrote.2

Why legislation: Courts decide fair use one case at a time, after a suit is filed, and the Copyright Office has left transparency for a future report.3 Senators from both parties have drafted the missing pieces: a subpoena process for creators,7 and notice to the Copyright Office of copyrighted training works.8 No federal law requires developers to keep or disclose their training sources. Companies that build on the work of American creators should be able to say what they used and how they got it.

The Solution

A four-step staircase: each step stands alone, and each step up asks more of developers. Scope: developers that train or substantially fine-tune generative models for U.S. commercial distribution, and their data suppliers. Fair use, the public domain, and existing remedies stay where the courts put them; permission to train is not permission to copy a voice, a likeness, or a book in the output.

Step 1 — Build the market. Direct the Copyright Office to run a searchable directory of rights and permissions, with standard license terms creators may choose to offer. Congress already requires a free, searchable public database of music ownership.9 Collective licensing stays optional and subject to antitrust law, and the Office reports on remaining gaps before Congress considers any compulsory license.

Step 2 — Publish the ingredients. Developers post dataset-level summaries: sources, how each dataset was obtained, whether it was licensed, and whether it contains copyrighted work. California already requires similar summaries.6

Step 3 — Keep the receipts. Retain work-level records of source, acquisition method, and asserted legal basis for new training, with a transition for existing datasets, and pass the duty through to data suppliers. A missing record carries its own penalty and creates no presumption of infringement.

Step 4 — Let creators ask. A rightsholder with a sworn, good-faith belief can obtain records about their own works before suing, through a court clerk, as the bipartisan TRAIN Act proposes.7 Judges can narrow requests, protect confidential material, and sanction abuse.

Where to start: Step 1 is the floor; it asks nothing of developers. Step 2 is the heart of the proposal.

Administration and enforcement: The Copyright Office sets record standards within 12 months, and compliance follows at 18, prospectively. Federal courts supervise evidence requests, Congress authorizes proportionate civil remedies for recordkeeping failures, and the Office does not decide contested fair-use questions.

Risks and Mitigations

  • Compelled disclosure: Expect challenges. In March 2026, a federal judge declined to block California's training-data law but saw "a distinct possibility" that xAI could prevail; its appeal is pending.10 Public summaries stay at the dataset level, and work-level records move only through court-supervised requests.
  • Chilling research: Fair use and research exceptions stay intact, no opt-out signal becomes a veto, and the duties apply only to commercial distribution. Compliance will still weigh more on small developers than large ones.
  • Old models: Duties apply to new training; for existing models, developers document good-faith reconstruction, and claimants must show they own what they claim. No dataset record can prove what a model memorized.

Similar Bills

Fit measures similarity to this proposal's mechanisms: High = direct precedent; Partial = useful component with material differences; Related = adjacent approach.

Federal

Proposal or bill Relevant provisions and fit Fit
S. 2455 — TRAIN Act
Welch (D-VT), Blackburn (R-TN), Hawley (R-MO), Schiff (D-CA)
Referred to committee · July 24, 2025
§2 adds 17 U.S.C. §514: on a sworn good-faith declaration, a court clerk issues a subpoena for records identifying the requester's own works used in training; noncompliance creates a rebuttable presumption of copying, and bad-faith requests face sanctions. Direct precedent for Step 4; its presumption goes further than this draft. House companion: H.R. 7209. High
S. 3813 — CLEAR Act
Schiff (D-CA), Curtis (R-UT)
Referred to committee · Feb. 10, 2026
§2 requires anyone training a generative model to file with the Register of Copyrights a "sufficiently detailed summary" of each registered work in the training data, 30 days before release, published in a public database; penalties of at least $5,000 per missed notice, capped at $2.5 million a year. Precedent for Steps 2 and 3; builds on the notice model of H.R. 7913 (118th). High
S. 4674 — COPIED Act
Cantwell (D-WA), Blackburn (R-TN), Heinrich (D-NM)
118th Congress · Introduced July 11, 2024; not enacted
§4 directs provenance standards, including for data used to train AI; §6(c) bars commercial training on content carrying provenance information without the owner's express consent. Provenance model for Step 3; its consent rule goes beyond this draft. Partial
S. 2367 — AI Accountability and Personal Data Protection Act
Hawley (R-MO), Blumenthal (D-CT)
Referred to committee · July 21, 2025
§3 creates a federal tort for training on a person's copyrighted work or personal data without express, prior consent. The consent-based alternative this proposal does not adopt. Related

State

Proposal or bill Relevant provisions and fit Fit
California — AB 2013 / Chapter 817
Signed Sept. 28, 2024 · Required from Jan. 1, 2026
Challenged in X.AI LLC v. Bonta (injunction denied; appeal pending)
Civil Code §3111 requires developers to post high-level dataset summaries, including sources, whether datasets contain copyrighted data, and whether they were purchased or licensed. Direct precedent for Step 2; no work-level records or creator access. High
California — AB 2602 / Chapter 259
Signed Sept. 17, 2024
Makes unenforceable certain vague contract terms allowing a digital replica of a performer's voice or likeness when the performer lacked counsel or union representation. Consent precedent for creative workers, relevant to the scope carve-out; not a training law. Related
Tennessee — HB 2091 / Public Chapter 588 (ELVIS Act)
Signed March 21, 2024 · Effective July 1, 2024
Adds voice to Tennessee's protected likeness rights and creates liability for tools whose primary purpose is producing unauthorized replicas. Supports keeping likeness distinct from copyright; does not address training transparency. Related

What this adds: The TRAIN and CLEAR Acts each supply one tool, creator subpoenas or filed notices, and California requires public summaries. This proposal joins them in one ladder, adds a public licensing directory and retained provenance records, and leaves fair use and the public domain intact without requiring consent for every use.

Notes

  1. Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA (N.D. Cal. June 23, 2025), order on fair use, pp. 3, 9, 11. ↩ ↩2

  2. Kadrey v. Meta Platforms, Inc., No. 3:23-cv-03417-VC (N.D. Cal. June 25, 2025), order granting Meta partial summary judgment on fair use, pp. 4–5; the ruling bound only the 13 named authors. ↩ ↩2 ↩3

  3. U.S. Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training, pre-publication version, May 2025, pp. 77 n.430, 106–107. Final version not yet published as of September 2026. ↩ ↩2 ↩3 ↩4

  4. Bartz v. Anthropic PBC, No. 3:24-cv-05417-AMO (N.D. Cal. July 20, 2026), order granting final approval of class settlement, p. 4. The release covers past copying, not AI outputs or future conduct. ↩

  5. Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc., No. 1:20-cv-613-SB (D. Del. Feb. 11, 2025) (Bibas, J., sitting by designation), limited to non-generative AI; interlocutory appeal, No. 25-2153 (3d Cir.), argued June 11, 2026, pending as of September 2026. ↩

  6. California AB 2013, Chapter 817, Statutes of 2024, adding Cal. Civ. Code § 3111 (documentation due on or before January 1, 2026); challenged in X.AI LLC v. Bonta, No. 2:25-cv-12295 (C.D. Cal., filed Dec. 29, 2025). ↩ ↩2

  7. S. 2455, TRAIN Act, 119th Cong. § 2 (introduced text). Sponsored by Sens. Welch (D-VT), Blackburn (R-TN), Hawley (R-MO), and Schiff (D-CA). ↩ ↩2

  8. S. 3813, CLEAR Act, 119th Cong. § 2 (introduced text). Sponsored by Sens. Schiff (D-CA) and Curtis (R-UT). ↩

  9. 17 U.S.C. § 115(d)(3)(E)(v): the mechanical licensing collective's musical works database "shall be made available to members of the public in a searchable, online format, free of charge." ↩

  10. X.AI LLC v. Bonta, No. 2:25-cv-12295 (C.D. Cal. Mar. 4, 2026), order denying preliminary injunction, p. 11; appeal, No. 26-1591 (9th Cir.), pending as of September 2026. ↩