A dangerous-capability threshold should trigger safety evaluation before an irrevocable open-weight release, since a closed model can still be patched afterward and an open one cannot.
Verification Status
AI-researched, unverifiedLast Reviewed
Jul 4, 2026
Cited Sources
8
Implementation, sequencing, safeguards, tradeoffs, and the practical path from principle to policy.
The Trump administration's July 2025 "Winning the Race: America's AI Action Plan" is explicitly pro-open-weight, arguing open and open-weight models "could become global standards" and "have geostrategic value" for extending American-aligned AI infrastructure worldwide. That's a policy stance, not a factual claim, but a clearly stated one. A June 2026 executive order establishing a voluntary pre-release government-access framework for "covered frontier models" applies by capability designation, not by whether a model will ultimately ship open- or closed-weight, consistent with this issue's own proposed design.
That builds on NTIA's July 2024 report, which concluded that for currently-available open-weight models, benefits (innovation diffusion, research reproducibility, decentralization) outweigh risks (CBRN misuse, cyber, disinformation), and recommended ongoing government monitoring rather than restriction. That report's formal status is murky today: it was issued under a Biden executive order that has since been rescinded, and no source confirms whether the report itself was withdrawn or simply orphaned. That's a small illustration of how the underlying policy authority for a "monitor, don't restrict" stance can quietly evaporate even while the report's text stays published.
Meanwhile, the open-weight frontier has moved fast: DeepSeek's V3 and R1 (January 2025, released under an MIT license) triggered both a market shock and a security review. The House Select Committee on the CCP's April 2025 report called DeepSeek a "profound" security threat, alleging unlawful data use and undisclosed censorship. Meta's Llama 4, Mistral's Large 3, and releases from Alibaba (Qwen 3) and Moonshot (Kimi K2) have continued narrowing the capability gap between open and closed frontier models internationally.
The proliferation risk here isn't hypothetical. Researchers demonstrated stripping Llama 3's safety fine-tuning in minutes to half an hour on a single consumer GPU ("Badllama 3"), and "abliteration" techniques remove refusal behavior via direct activation-space edits without any fine-tuning at all. As few as roughly ten adversarial examples have broken safety alignment in multiple published studies. WormGPT, a malicious tool built specifically for phishing and malware generation, was originally built by fine-tuning an open-weight EleutherAI model in 2023; by 2025 new WormGPT-branded variants were being built on Grok and Mixtral weights instead. This is evidence that the underlying problem (safety training is removable) isn't specific to any one open model, and that closed models aren't fully immune either once access is obtained.
A January 2025 interim rule set a training-compute threshold (roughly 10^26 FLOP) for controlling exports of closed model weights specifically, and explicitly exempted openly-released model weights from that control. The broader three-tier country-based diffusion framework built around that rule was rescinded in May 2025, and legal commentators say it's unclear whether the narrower closed-model-weights control survived that rescission or needs to be reissued. As of early 2026, a replacement rule was reportedly still being drafted, focused on chip volume rather than model weights specifically. The practical upshot: for most of the last two years, US export-control policy has already treated open release as a reason a model falls outside certain controls, not a reason it should face additional ones. This reinforces this issue's core position that release strategy and export risk are historically treated as inversely, not directly, related in US policy.
Security-focused critics, most visibly Anthropic's Dario Amodei, in testimony dating back to 2023 and revived in 2026 discourse, argue that irrevocability is the whole problem: once weights are released, there is no way to patch, restrict, or recall the model, regardless of what's later discovered about its capabilities. This issue's answer isn't to dismiss that concern; it's Proposal 2's release-timing rule, which takes the irrevocability argument seriously without extending it into a general restriction on open release. Open-source advocates make the opposite argument: open weights let far more independent researchers audit a model for vulnerabilities than a closed model ever allows, and restricting release mainly disadvantages smaller, non-incumbent developers who can't otherwise compete with well-capitalized closed-model labs. This is a concern this issue shares with AI-10's safe-harbor proposal. The party's position tries to hold both: open release stays the default, and the one thing that changes is timing for the narrow band of models that cross a dangerous-capability line.
Turn frustration into useful pressure.
If this position misses evidence or a lived consequence, challenge it. If it holds up, help test it locally and connect it to the issues around it.