2 Minutes
When a government asks to peek under the hood of cutting-edge AI, companies usually pause. OpenAI didn’t—at least not entirely. George Osborne, OpenAI’s head of countries, told a London crowd at SXSW that the company will voluntarily participate in the new U.S. process that asks for early access to advanced models.
Originally, the executive order pushed by the White House aimed for a 90-day pre-release submission window. That proposal met resistance. It softened, reportedly down to a 30-day request after public pushback and even some public ambivalence from President Trump himself. The final wording asks for access rather than demanding it, but it still signals stronger oversight for the most capable systems.
OpenAI has agreed to submit its next-generation models into a benchmarking exercise designed to evaluate “advanced cyber capabilities of AI models” and to help set the threshold for what regulators might label a "covered frontier model." What does that mean in practice? Think of it like stress-testing software for real-world risk: capabilities, misuse potential, and security posture all get measured against a public-interest yardstick.

Osborne framed the move as proactive. He said OpenAI has been suggesting how governments can monitor AI safety and security—not just in the U.S., but globally. That approach reflects an industry shift: companies once defensive about regulation now see some coordination as a way to avoid rushed, poorly designed rules.
Still, voluntary cooperation is different from binding oversight. Companies can comply, advise, or push back. They can also shape benchmarks and the definitions that flow from them. For policymakers, the challenge will be ensuring benchmarks remain rigorous, transparent, and resistant to capture. For firms, the calculus is reputation, market access, and the practical burden of handing over sensitive models—even to vetted reviewers.
Regulation by request. Benchmarking by agreement. A new chapter in how industry and government negotiate control over powerful AI is beginning, and it raises familiar but urgent questions about accountability, technical transparency, and who decides the thresholds for safety. Watch this space—the tests themselves may tell us more than any executive order.
Comments
No comments yet.
Leave a Comment