Labs spending federal research money must buy synthetic DNA from suppliers that screen orders for dangerous sequences. Leading AI companies screen requests for weapons help with automated filters, but each sets its own standard and reports its own results. Congress should set a narrow public standard, protect legitimate users, require independently tested filters, and screen what AI agents do as well as what they say.
The Problem
Frontier AI companies now guard against weapons misuse with classifiers, automated filters that block dangerous requests and answers. Anthropic's classifiers monitor what goes into and out of its models and block a narrow class of chemical, biological, radiological, and nuclear weapons information.1 OpenAI's two-tier monitor scans every ChatGPT agent conversation, including the agent's use of outside tools.2
The filters help, within limits. In Anthropic's automated tests, its classifiers cut the success rate of attempts to trick the model, called jailbreaks, from 86% to 4.4%.3 Yet in a week-long public challenge, four participants beat every level, one with a single technique that worked on all of them.3
DNA shows another way. A 2024 federal framework requires labs using federal research money to buy synthetic DNA only from suppliers that screen orders for sequences of concern and check who is buying; a 2025 executive order ordered it revised or replaced, with enforcement.4 A DNA order is screened under a public framework; a request to an AI model is screened by whatever its maker built, graded by a test its maker wrote. Three problems follow:
- Every company grades itself. OpenAI reports that its biology monitor caught 84% of dangerous examples on a deliberately difficult in-house test;2 Anthropic's figures come from its own attacks. No common test lets anyone compare them.
- Attacks come in pieces. In 2025, a group Anthropic tied to the Chinese state broke a hacking campaign into "small, seemingly innocent tasks" for its coding agent, targeting about 30 organizations.5 A filter that judges one message at a time can miss the pattern.
- Good work gets blocked. OpenAI says its biology safeguards "will sometimes accidentally prevent safe uses"; on that same test, about a third of what its monitor flagged was not dangerous.2
Why legislation: No federal law requires AI services to screen for weapons help or to prove their screens work. For DNA, Washington is moving from trust to verification: the 2025 AI Action Plan calls for enforcement "rather than relying on voluntary attestation,"6 Senators Cotton and Klobuchar would require screening against a federal list with random red-team audits,7 and in July 2026 the House passed a bill directing NIST to measure how well screening works.8 The same logic fits AI services. A filter that protects the public should meet a public standard.
The Solution
A four-step staircase: each step stands alone, and each step up adds independence and obligation. Scope: AI services whose evaluations show they can materially help with biological, chemical, or cyber attacks or other severe physical harm, judged by what the system can do, whatever it is called. General knowledge, authorized research, and any political, religious, or identity-based category stay outside the standard.
Step 1 — Set a narrow public standard. NIST, with HHS and cybersecurity experts, defines prohibited help in concrete terms and publishes a test that counts both errors: dangerous help missed and legitimate work blocked. Covered providers report their results, and aggregate error rates are public. Definitions update through a public process, as the Cotton–Klobuchar bill provides for dangerous DNA sequences.7
Step 2 — Protect legitimate users. Give scientists, doctors, and security researchers a fast way to contest a block, and let vetted institutions earn fuller access. OpenAI is building such a program for vetted life-science customers,2 and the Cotton–Klobuchar bill offers trusted institutions expedited review.7
Step 3 — Require filters that pass independent tests. Covered services must meet the standard with their own classifiers, a vendor's, or any mix that works; the law sets the bar and lets companies compete to clear it. Accredited testers probe each system at random and after major updates, as the Cotton–Klobuchar bill would require for DNA suppliers.7
Step 4 — Screen actions as well as words. For AI agents that run code, send messages, or order materials, screen before high-risk actions and across a whole session, where piecemeal attacks show up. Uncertain cases go to qualified human review or a narrower mode; minimal safety-event records are kept for audits.
Where to start: Step 1 is the floor: a public yardstick and honest reporting against it. Step 3 is the heart of the proposal.
Administration and enforcement: Commerce administers the duty; NIST writes and maintains the test with HHS and CISA. Pilot the methods in year one; obligations begin six months after the standard is independently validated. Civil penalties apply to missing, misrepresented, or persistently failing safeguards; passing the test shields no one from other negligence claims.
Risks and Mitigations
- Free speech: Rules on what a model may say are content-based and will be challenged. Limit the duty to specific help with enumerated weapons crimes and to actions such as running code, leave education and science alone, and drop any category a court rejects rather than relabel it.
- Filters fail: Determined attackers get through, as Anthropic's public challenge showed. The standard therefore requires layered defenses, random retesting, and incident review, and a passed test never counts as proof of overall safety.
- Frozen technology: A fixed standard could lock in today's methods. Performance-based rules, equivalent alternatives, and regular revision leave room for better filters, and confidential test materials keep the test from being gamed.
Similar Bills
Fit measures similarity to this proposal's mechanisms: High = direct precedent; Partial = useful component with material differences; Related = adjacent approach.
Federal
| Proposal or bill | Relevant provisions and fit | Fit |
|---|---|---|
| S. 3741 — Biosecurity Modernization and Innovation Act of 2026 Cotton (R-AR), Klobuchar (D-MN) + 8 bipartisan cosponsors Referred to committee · Jan. 29, 2026 |
§4 requires DNA synthesis providers to screen orders against a federal list of sequences of concern, verify customers, and pass audits that include random red-teaming, with expedited review for trusted institutions and NIST testing of screening accuracy. The closest statutory model for Steps 1–3; covers DNA orders, not AI services. | High |
| H.R. 3029 — Nucleic Acid Standards for Biosecurity Act Salinas (D-OR), McCormick (R-GA), McBride (D-DE), Riley (D-NY) Passed House · July 20, 2026 |
House-passed §2 directs NIST to test "the accuracy, efficacy, and reliability" of DNA synthesis screening and develop conformity-assessment standards, with $5 million a year for FY2027–2031. Model for Step 1's measurement role; voluntary and limited to DNA. | Partial |
| S. 5616 — Preserving American Dominance in Artificial Intelligence Act of 2024 Romney (R-UT), Reed (D-RI), Moran (R-KS), King (I-ME), Hassan (D-NH) 118th Congress · Introduced Dec. 19, 2024; expired |
§§5(b)(2) and 8(c) require frontier developers to implement Commerce standards for red-teaming and mitigating chemical, biological, radiological, nuclear, and cyber risks during development. Mandatory safeguards akin to Step 3; sets no performance test for screening deployed systems. | Partial |
| H.R. 9363 — AI Security and Innovation Act Obernolte (R-CA), Foushee (D-NC), Babin (R-TX), Mann (R-KS), Franklin (R-FL) Ordered reported with a substitute (29–0) · June 25, 2026 |
Introduced §2 has a NIST center evaluate and improve security measures against threats including "model jailbreaks." A related research role; voluntary and without regulatory authority. Compares introduced text. | Related |
State
| Proposal or bill | Relevant provisions and fit | Fit |
|---|---|---|
| California — SB 1047 (2024) Vetoed · Sept. 29, 2024 |
Proposed §22603(b) required developers, before release, to "implement appropriate safeguards" against critical harms and keep replicable test records. A risk-control duty; prescribed no common screening standard. | Partial |
| California — SB 53 (2025) Enacted Sept. 29, 2025 (Ch. 138) |
§22757.12(a) requires large developers' published frameworks to describe how they apply mitigations and use third parties to assess their effectiveness. Disclosure of safeguards; no performance standard. | Related |
| New York — RAISE Act, S8828 / Ch. 96 (2026) Signed Mar. 27, 2026 · Effective Jan. 1, 2027 |
Requires similar published frameworks and critical-incident reporting. A disclosure analogue; no screening mandate. | Related |
What this adds: DNA screening already has a House-passed measurement bill and a bipartisan Senate bill that would make it mandatory. No bill yet does the same for AI. This proposal applies that model to AI services: a narrow public standard, protection for legitimate users, independently tested filters from competing vendors, and screening where agents act.
Notes
-
Anthropic, "Activating AI Safety Level 3 Protections," May 22, 2025. Its "Constitutional Classifiers" are "real-time classifier guards" that "monitor model inputs and outputs" to block "a narrow class of harmful CBRN information"; they "may produce false positives." ↩
-
OpenAI, ChatGPT Agent System Card, July 17, 2025, pp. 35–37. The monitor "scans user messages, external tool calls, and the final model output" (p. 35); OpenAI is "building a trusted access program" for "vetted and trusted customers" (p. 36). On "challenging prompts," the reasoning monitor had recall of 0.838 and precision of 0.647; OpenAI optimizes for recall "even at a cost of reduced precision" (p. 37). ↩ ↩2 ↩3 ↩4
-
Anthropic, "Constitutional Classifiers: Defending against universal jailbreaks," February 3, 2025, updated February 13, 2025. The automated evaluation used 10,000 jailbreak prompts Anthropic generated, run against Claude 3.5 Sonnet; the public demo drew about 3,700 hours of attempts, and the system "resisted jailbreaking attempts for five of the planned seven days." See also Mrinank Sharma et al., "Constitutional Classifiers," arXiv:2501.18837, January 2025. ↩ ↩2
-
Office of Science and Technology Policy, "Framework for Nucleic Acid Synthesis Screening," April 29, 2024, revised September 2024 (providers self-attest compliance). Under Executive Order 14292, 90 Fed. Reg. 19611 (May 5, 2025), § 4(b), the framework is to be revised or replaced; as of September 2026, HHS says it will post the new framework "once the new framework is available." ↩
-
Anthropic, "Disrupting the first reported AI-orchestrated cyber espionage campaign," November 13, 2025. The group, assessed "with high confidence" to be Chinese state-sponsored, used Claude Code against "roughly thirty global targets" and "succeeded in a small number of cases." ↩
-
The White House, America's AI Action Plan, July 2025, p. 23 ("Invest in Biosecurity"). ↩
-
S. 3741, Biosecurity Modernization and Innovation Act of 2026, 119th Cong. § 4(a)–(b) (introduced text). Introduced by Sen. Cotton (R-AR) with Sen. Klobuchar (D-MN); eight cosponsors from both parties joined later. ↩ ↩2 ↩3 ↩4
-
H.R. 3029, Nucleic Acid Standards for Biosecurity Act, 119th Cong. § 2 (engrossed text). Passed the House by voice vote on July 20, 2026; referred to the Senate Commerce Committee on July 21, 2026. ↩