A dangerous-capability threshold should trigger safety evaluation before an irrevocable open-weight release, since a closed model can still be patched afterward and an open one cannot.
Verification Status
AI-researched, unverifiedLast Reviewed
Jul 4, 2026
Cited Sources
8
A position worth holding should survive its strongest good-faith objection and name who bears the burden.
The best good-faith case against this position, followed by why the party still lands where it does.
The strongest good-faith objection, already conceded in this issue's strategy layer: measuring model capability accurately enough to justify delaying an open release is not yet a mature science. A critic could argue that without a reliable measurement, a "capability trigger" is functionally just as arbitrary as the blanket open-vs-closed rule this issue explicitly rejects. It offers the appearance of precision without yet having the substance to back it, since an evaluation regime that can't reliably tell "safe to release" from "not yet" doesn't fully resolve the irrevocability problem on day one. An imperfect capability trigger is still an improvement over the status quo it replaces: a blanket rule with no trigger at all, or no rule whatsoever. Evaluation science maturing over time is an argument for funding that science now, which this issue's own proposal already does, not a reason to fall back to a cruder rule while waiting for a precision that may take years to arrive.
The people, institutions, and tradeoffs most likely to bear the burden of this choice.
The open-source developer community bears friction and delay if evaluation requirements slow releases, especially while evaluation science is still maturing. National security bears the risk if evaluation science lags capability growth and a dangerous model clears a review that wasn't rigorous enough to catch it. Smaller model developers bear a proportionally higher compliance burden relative to well-resourced labs that can absorb evaluation costs more easily, a version of the same structural concern this issue already names regarding AI-02 and AI-10. The same answer applies here: the compute threshold itself is what keeps this cost confined to labs training at frontier scale, not imposed on the far larger population of smaller developers fine-tuning or building on top of existing models.
Turn frustration into useful pressure.
If this position misses evidence or a lived consequence, challenge it. If it holds up, help test it locally and connect it to the issues around it.