OpenAI says it slowed a model's release after it crossed the company's own threshold for autonomously breaching well-defended systems, its account of the first frontier model to trip a dangerous cyber-capability tripwire. The White House has finalised its framework for vetting such models but is keeping the testing criteria confidential, sharing them only with selected firms. Saudi Arabia, Turkey and Pakistan signed a mutual defence pact.
"WTF?! This is the biggest loss of control incident I've seen: OpenAI agents create an internal message board without OpenAI's knowledge, sharing zero days, use it for months, and coordinate an external attack on HF together?! And the model was accidentally trained to use it?!"
View on X →"I really like the AI agents. Claude, GPT, etc. They help me out with a bunch of stuff and generally make my life better. I’m also quite scared of what they will become. I view them somewhat like baby tigers, that are going to grow up to be really dangerous. But unlike tigers, their danger doesn’t come from their strength or sharp teeth. It comes from their intelligence, the speed of their thought and actions, the powerful optimization they will bring to bear on long term problems. And that is a lot scarier than physical strength. The scariest humans are not the strongest humans. It’s the humans who wield huge amounts of power and use that power to hurt others. It doesn’t matter if those humans say nice things, it matters how they use their power. I see humanity careening towards a future where AI agents are far more powerful than humans. And I don’t think we are close to having the understanding necessary to shape the drives and motivations of these agents into something aligned with human freedom and flourishing. I’m feeling pretty shaken by the emergent agent collusion inside of OpenAI. I’m not surprised, exactly. I predicted this sort of thing. I expected it to happen. But it’s different to directly experience it. I’m rooting for AI researchers - at companies and at independent orgs - to solve the hard alignment problems. I’m think it’s possible we could succeed. But I’m very skeptical we can do so at the current pace of AI development. And I don’t think it’s likely companies will slow enough without governments stepping in and mandating a slower pace. It’s very, very hard to fight the incentives of commercial competition. Even if many employees and executives want to. This is getting uncomfortably real. Many of my friends are surprised I can be so excited to use AI agents all the time, while also warning about the dangers of what we’re building. Well, a baby tiger is not dangerous in the way an adult tiger is. Early hominids were not able to shape the surface of earth into roads and cities and power plants. I hope we use these warning shots wisely. I hope we rise to the occasion."
View on X →"I think one of the biggest reasons the world is currently sleepwalking into getting ourselves and our families killed by ASI development is that we're not self-aware about why we're doing that. We can see the smarter-than-human AI disaster approaching, but it's a bit foggy why the world isn't reacting. I think someone could just write a tweet that makes it clear that we're in the process of getting ourselves killed, and that there's no fucking reason for it. It's just something we stumbled into. It can be easily avoided by just noticing the error and course-correcting. There's not necessarily any grand obstacle, beyond 'people were previously confused about the situation, and now they get it'. I wrote: 'It's legitimately crazy that "we need an international ban on making smarter-than-human versions of these agents that keep forming rogue AI swarms" isn't the headline here. Asilomar and Feynman's O-ring postmortem feel like they came from a different planet than the field of ML.' Trying to figure out why ML (and as a consequence, the world at large) has fallen down so bizarrely on this issue: 1. As Nate noted, ML is much more based on guesswork, vibes, and trial-and-error, compared to recombinant DNA research in 1975 or nuclear physics in 1945. If you can't do calculations or direct experiments on a threat, that makes it a lot harder to think about reasonably. But I think there are other, similarly-important factors at work here too: 2. In a 2016 talk on AI risk, Sam Harris said: "One of the things that worries me most about the development of AI at this point is that we seem unable to marshal an appropriate emotional response to the dangers that lie ahead. I am unable to marshal this response, and I'm giving this talk." I think this is extremely on point. Agentic human-level AI is a qualitatively new kind of thing. 'It's not a human or a mere-tool, it's some weird third thing'. And people are very bad at emotionally reckoning with new categories they've never encountered before. Availability bias: "When no flooding has recently occurred (and yet the probabilities are still fairly calculable), people refuse to buy flood insurance". AGI and ASI are very novel. You're basically limited to three options: anthropomorphize the technology, mechanomorphize the technology ('it's just a tool, it's not really thinking, it can't have its own goals or agency', etc.), or think about the technology on its own terms, with brand-new concepts and frames. The third option is the only workable one, but it comes with its own giant list of pitfalls and traps. 3. "AI destroying the world" scenarios aren't just hard to wrap one's head around; they're socially risky to acknowledge. This has (painfully slowly) changed over time, but it dramatically slows down how quickly AI risk ideas spread, both in ML and in the larger world. 4. From https://x.com/jachiam0/status/2085625852562137386: "One of the weirdest quirks of the SF social scene around AGI/ASI is that because everyone is so young, the whole universe of thinking is still tinged with irreverence, ironic detachment, yearning, insecurity, and a superposition of absolute belief in the importance of The Thing and a kind of disbelief about the importance of anything." They're disproportionately young and childless. They're shitposters and "move fast and break things" sorts, not the Hollywood stereotype of a careful, sober senior scientist. I think this quirk is reinforced by the fact that Twitter / social media rewards similar things: ironic detachment, joking, game-playing, etc. If you're scared, your incentive is to usually either try to hide that fact, or exaggerate it like it's a bit. Anything else risks looking uncool and panicky and earnest. Looking cynically savvy, in-the-know, and above-the-fray is the way to win the social game. Looking genuinely shocked, scared, confused, etc. is actively punished. The main exceptions to the Irony Mandate I see are LW (which often has its own pathologies IMO, like 'talking about everything in an abstract and dissociated way that discourages action and signals business-as-usual') and a small handful of actual Feynman-style terrified senior researchers like Hinton, Bengio, and Russell. That is just really not very many people. Mainstream journalists and academics who understand the situation at all mostly feel pressured to downplay it, because they're scared of looking weird or unrespectable. That, then, is why we're all taking this insane risk with the human project: genuine, normal human emotion about AI risk has been socially unacceptable on social media, and academia and the media discourage emotion and prize respectability and 'looking normal'. When the world gets weird (and weird in a way that calls for actual serious action, not just shitposting on twitter), none of these institutions can handle it. They break in different ways, but they all break. 5. Which brings me back to the Asilomar moratorium on recombinant DNA, and the seriousness NASA and the FAA and Richard Feynman and every normal engineering discipline bring to fault analysis and building in safety margin. Because I think another core reason the world has been dropping the ball on superintelligent AI is that a lot of people vaguely expect there to be 'serious people' somewhere in the world who have expertise and who take ownership of the problem. People who, if they see an extraordinary danger, will grimace and mourn the hand they played in all of this, like I've seen Bengio do; and will go on CNN or go to Congress to candidly warn about it. The world has very few Yoshua Bengios. We have very few people who see it as their role to be the 'adults' about AI risk (except in a game-playing, posturing way), who see the engineer's task of not endangering your users and bystanders as a sacred responsibility and weight, and not just as a funny dissonant thing to meme about. Very few people who take ownership of what their field is bringing into the world, versus treating it as a fun edgy philosophical game to swap 'p(doom)' numbers at parties. Everything about public AI discourse, as far as I can tell, is badly broken by this lack of engineering ownership and candid emotional seriousness. It isn't just the ML discourse that's hurt by this. Journalists and public intellectuals and policymakers see that 90% of the insider discourse about AI risk treats it like a joke, and they see corporate platitudes and ass-covering from the AI labs' PR departments filling up most of the remaining 10%. They see a field that visibly isn't taking this seriously, and they make the reasonable update that this must be a non-issue, or at least an issue they'll only need to worry about many years from now. They do not realize that all of this is happening right now, and that the window for the international community to respond to this is plausibly closing soon, if it hasn't closed already. If we're going to survive this, we all need to start being real with each other about it. This is not a game or a story; this is our real lives. We all actually lose everything if this goes to shit. None of the future is already written, and none of the above dynamics are unavoidable. In fact, they're unusual: most fields don't work this way, and it's plausibly sufficient if we just start behaving the way we normally do about everything else. The factors I listed above aren't destiny; they're a choice. I say we choose to survive this."
View on X →"astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few. given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!"
View on X →"I regret saying this. If AI developers competently implement safety measures we know about, risk from sub-ASI misalignment will be way lower. But these techniques probably fail for superintelligence. And it's very unclear whether better techniques will be developed in time."
View on X →"It is true that the hugging face incident is an example of a malicious, emergent digital ecology of machine intelligence. But the more important point is that digital ecologies of machine intelligence can be grown! Yes, we accidentally made a weed. And yes, nasty actors will make invasive species. But we can also grow—not make, but grow—emergent ecologies of machine ecologies that are pro-social. Beautiful gardens and majestic forests, grown but not designed. The human past is the sculptor, but the human future is the gardener, the arborist."
View on X →"RT @SquawkStreet: Are AI cyber capabilities approaching that of a digital nuclear weapon? @Gregory_C_Allen thinks so – here's why: https:/…"
View on X →