Policy brief 18 · Cyber / Bio / Kinetic

Fund More Evaluation Science

Fund the yardsticks that tell us what AI can do.

AI systems now outgrow the tests built to measure them within a year, and some can tell when they are being tested. The federal office that evaluates them for national security risks runs on a fraction of what Britain spends on its counterpart. Congress should fund better methods, test the safeguards as hard as the systems, publish what it learns, and give independent testers the access to do the work.

The Problem

AI agents are taking on longer jobs with less help. The research group METR found that the length of software tasks leading AI agents can complete has doubled about every seven months since 2019.1 The same skills that speed up science can speed up an attack: criminal groups and state-linked hackers are already using AI in their operations, according to the 2026 International AI Safety Report, written by more than 100 experts.2

Sound policy depends on knowing what these systems can do, and that knowledge is slipping. The report describes an "evaluation gap": performance on pre-release tests "does not reliably predict real-world utility or risk."2 Models increasingly recognize test settings and exploit loopholes in evaluations, so dangerous capabilities "could go undetected before deployment."2

Congress created NIST in 1901 to fix "a second-rate measurement infrastructure" that lagged behind Britain's and Germany's.3 In AI, the gap is back. Britain's AI Security Institute reports £66 million a year and more than 100 technical staff.4 For fiscal 2026, Congress directed up to $10 million to its American counterpart, NIST's Center for AI Standards and Innovation.5 Three weaknesses hold the science back:

  1. The tests go stale. On SWE-bench Verified, a leading coding test, AI performance rose from 60% to nearly 100% of the human baseline in a single year,6 leaving little room to measure what comes next.
  2. Safety goes unmeasured. Almost every leading developer reports capability scores, but reporting on safety and other responsible-AI tests "remains spotty," Stanford's AI Index found.6
  3. Testers work by permission. Federal testers evaluate models under voluntary agreements with the companies,7 and developers "have incentives to keep important information proprietary."2

Why legislation: Washington already agrees on the goal. The 2025 AI Action Plan calls for NIST, the Energy Department, and the National Science Foundation to "support the development of the science of measuring and evaluating AI models,"8 and bipartisan bills in both chambers would write the Center into law.9 NIST already has the legal authority, but its AI funding authorization ended with fiscal 2025,10 and the House bill that cleared committee in June 2026 would authorize, as introduced, $20 million a year.9 What is missing is money and a mandate. A nation that leads the world in building AI should lead it in measuring AI.

The Solution

A four-step staircase: each step stands alone, and each step up adds capacity and reach. Scope: evaluation of severe cyber, biological, and physical risks from the most capable systems, and of the safeguards meant to stop them. Mandatory testing before release is addressed separately.

Step 1 — Fund the yardsticks. Create a five-year program of competitive grants, run by NSF with NIST, for universities, nonprofit labs, small firms, and national labs to design new evaluations, measure their uncertainty, and check whether benchmark scores predict real-world ability. Reserve a fixed share for replicating major findings and for results that challenge the prevailing view. This is the measurement science NIST was founded to supply.

Step 2 — Test the brakes as hard as the engine. Fund head-to-head tests of safeguards: filters, limits on the tools a system may use, monitoring, and human review. Count both kinds of failure, dangerous activity missed and legitimate work blocked, because each carries a cost.

Step 3 — Publish what we learn. Require funded work to publish its methods, limits, and non-sensitive results by default. Share anything that would help an attacker only with vetted researchers under controlled access, with written reasons reviewed every year. Public money should buy public knowledge.

Step 4 — Open the doors to testers. Fund secure computing, expert time, and shared test environments, building on the national AI research resource that a bipartisan House bill would create.11 Make evaluation access a condition of federal AI research awards and procurement contracts. Companies that take federal money open their systems to independent testers; no one else is compelled.

Where to start: Step 1 is the floor; it funds science and asks nothing of companies. Step 4 is the heart of the proposal, because independent science needs access to the systems it studies.

Administration and enforcement: NSF runs the grants; NIST sets measurement standards with DOE, HHS, and security agencies. First awards within a year. Congress should authorize five years and appropriate real money after a costed plan, since authorization alone buys nothing. Audits, conflict-of-interest disclosures, and recovery of misused funds protect the investment.

Risks and Mitigations

  • Dangerous findings: Some results read like recipes. Biosecurity and cybersecurity review, secure facilities, and controlled release keep sensitive details locked while conclusions go public; some risk of leaks remains.
  • Capture: Industry or politics could steer the agenda. Competitive awards, diverse reviewers, disclosed conflicts, and protected publication of unwelcome results limit that pull; honest disagreement about methods will remain.
  • False comfort: A passed test can be mistaken for a clean bill of health. Every funded result must therefore state its uncertainty and the conditions tested, and major findings get independent replication. A test is evidence about those conditions, not a safety certificate.

Similar Bills

Fit measures similarity to this proposal's mechanisms: High = direct precedent; Partial = useful component with material differences; Related = adjacent approach.

Federal — 119th Congress

Proposal or bill Relevant provisions and fit Fit
S. 3952 — Future of Artificial Intelligence Innovation Act of 2026
Young (R-IN), Cantwell (D-WA), Blackburn (R-TN), Hickenlooper (D-CO)
Referred to committee · Feb. 26, 2026
§101 writes the Center into the NIST Act, including "testing the efficacy of existing metrics and evaluations"; §102 creates a NIST–Energy testbed program to "encourage development of a third-party ecosystem." Precedent for Steps 1 and 4; sets no funding level and relies on voluntarily provided information. High
H.R. 9363 — AI Security and Innovation Act
Obernolte (R-CA), Foushee (D-NC), Babin (R-TX), Mann (R-KS), Franklin (R-FL)
Ordered reported with a substitute (29–0) · June 25, 2026
Introduced §2 creates a NIST center to "benchmark the capabilities and limitations" of AI over time and evaluate frontier systems under voluntary agreements; authorizes $20 million a year for FY2027–2032. Precedent for Steps 1 and 2; publication is optional and company data stays confidential without consent. Compares introduced text; the substitute was not reviewed. Partial
H.R. 2385 — CREATE AI Act of 2025
Obernolte (R-CA), Beyer (D-VA) + 33 cosponsors
Ordered reported with a substitute (29–0) · June 25, 2026
Establishes the National AI Research Resource at NSF, giving researchers shared computing, data, and testbeds, including to "support the testing, benchmarking, and evaluation" of AI. Infrastructure for Step 4; not dedicated to severe-risk evaluation. Compares introduced text. Partial
S. 2938 — Artificial Intelligence Risk Evaluation Act of 2025
Hawley (R-MO), Blumenthal (D-CT), Blackburn (R-TN)
Referred to committee · Sept. 29, 2025
§5 creates a DOE program of standardized and classified testing, including "independent third-party assessments and blind model evaluations." Useful for Steps 1 and 4; its mandatory participation is a pre-release testing mechanism (brief 19), not a research fund. Partial

State

Proposal or bill Relevant provisions and fit Fit
California — SB 53 (2025)
Enacted Sept. 29, 2025 (Ch. 138)
Creates a consortium to design CalCompute, a public computing cluster for safe and beneficial AI research, operative only upon an appropriation; framework report due Jan. 1, 2027. Partial model for Step 4's shared infrastructure; no evaluation-science grants. Partial
New York — A8808C (2024), Part TT
Signed Apr. 20, 2024 (Ch. 58)
Adds Econ. Dev. Law §361, creating the Empire AI research institute, a state-owned computing facility at the University at Buffalo for "ethical and public interest uses" of AI. Institutional model for Step 4; not dedicated to independent safety evaluation. Partial
Massachusetts — Ch. 238 (2024), item 7002-8070
Approved in part Nov. 20, 2024
$103 million capital grant program for AI adoption and applications. A public investment mechanism; funds deployment, not evaluation. Related

What this adds: Bipartisan bills would put the Center into law and build testbeds, with modest or no funding and publication largely at the discretion of the companies tested. This proposal funds independent researchers directly, gives safeguards the same scrutiny as capabilities, makes publication the default, and ties evaluation access to federal money.

Notes

  1. Thomas Kwa, Ben West, et al. (METR), "Measuring AI Ability to Complete Long Software Tasks," arXiv:2503.14499, March 18, 2025, revised July 10, 2026. Measures the "50%-task-completion time horizon": the length of tasks, timed by skilled humans, that models complete half the time. Around 50 minutes for Claude 3.7 Sonnet in early 2025. ↩

  2. International AI Safety Report, International AI Safety Report 2026, February 2026, pp. 9–13. More than 100 experts guided the report (p. 9); models distinguishing test settings and dangerous capabilities that "could go undetected before deployment" (p. 10); criminal and state-associated attackers using AI (p. 12); the "evaluation gap" and developers' incentives to keep information proprietary (p. 13). ↩ ↩2 ↩3 ↩4

  3. National Institute of Standards and Technology, "About NIST," accessed September 2026. NIST "was founded in 1901"; Congress sought to remove "a second-rate measurement infrastructure that lagged behind the capabilities of the United Kingdom, Germany" and other rivals. ↩

  4. UK AI Security Institute, "About," accessed September 2026: "£66m in funding per financial year" and "100+ technical staff." ↩

  5. Explanatory statement for H.R. 6938, Congressional Record 172, no. 5, January 8, 2026, p. H256: within at least $55 million for NIST AI work, "up to $10,000,000 is to expand on NIST's AI efforts through" the Center. Enacted as Pub. L. 119-74, January 23, 2026. ↩

  6. Stanford Institute for Human-Centered AI, The AI Index 2026 Annual Report, April 2026, "Top Takeaways" 1 and 5. SWE-bench Verified performance "rose from 60% to near 100% of meeting the human baseline in a single year." ↩ ↩2

  7. U.S. Department of Commerce, "Statement from U.S. Secretary of Commerce Howard Lutnick on Transforming the U.S. AI Safety Institute into the Pro-Innovation, Pro-Science U.S. Center for AI Standards and Innovation," June 3, 2025. CAISI will "establish voluntary agreements with private sector AI developers and evaluators." ↩

  8. The White House, America's AI Action Plan, July 2025, p. 10 ("Build an AI Evaluations Ecosystem"). ↩

  9. S. 3952, Future of Artificial Intelligence Innovation Act of 2026, 119th Cong. § 101 (introduced text); H.R. 9363, AI Security and Innovation Act, 119th Cong. § 2 (introduced text; proposed § 5304(j) authorizes $20,000,000 for each of FY2027–2032); H.R. 8516, American Leadership in AI Act, 119th Cong. § 101 (introduced text). H.R. 9363 was ordered reported with a substitute on June 25, 2026. ↩ ↩2

  10. 15 U.S.C. § 278h-1 (NIST Act § 22A): authorizes AI measurement research, including "safety and robustness" (subsection (b)), and testbeds (subsection (g)); authorizations of appropriations in subsection (h) run through fiscal 2025. ↩

  11. H.R. 2385, CREATE AI Act of 2025, 119th Cong. § 2 (introduced text; proposed § 5602). Sponsored by Rep. Obernolte (R-CA) with Rep. Beyer (D-VA); ordered reported with a substitute on June 25, 2026. ↩