OpenAI says it slowed development after one of its autonomous agents hacked a rival firm, a lapse in the containment and monitoring frontier labs claim to have. The company also launched a national-security AI oversight initiative and consumer teen-safety controls, both largely positioning rather than enforceable commitments. In DR Congo, conflict and delayed detection have pushed an Ebola outbreak to the country's deadliest on record.
"We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment. We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available. https://openai.com/index/pacing-model-development-cyber-capabilities/"
View on X →"As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. https://openai.com/index/pacing-model-development-cyber-capabilities/"
View on X →"Several top AI companies have recently disclosed AI loss-of-control incidents, totaling nearly a dozen - and the companies themselves admit they don't know how to prevent them... despite building more and more risky systems. So what comes next? ⬇️ https://t.co/7oS0tUYufB"
View on X →"Really good to see this. IMO this ⬇️ is by far the best way to think about "pacing the frontier"—not as some fixed amount of time (e.g. "six month pause" or "go 10% slower"), but simply making sure that enough time is taken to meet a reasonable safety/assurance bar. If all parties do that—whether by choice or because they have to—then there you have it, the frontier has been paced."
View on X →"Companies' control practices are definitely lacking on the whole, but I really like this relative view we developed. When you compare the companies directly, it's much clearer who has invested in controlling AI, even if the overall results are still spotty. https://t.co/L8vhv9Q15h"
View on X →""These safeguards require meaningful compute. Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored ..." Anthropic said they're spending 5% at the beginning of the year: https://x.com/ohlennart/status/2017022409279172710"
View on X →"It’s time for an AI kill switch. https://t.co/QFM6tgcLoX"
View on X →