Arquivo da tag: OpenAI

AI could kill all humans in next decade, warn experts: but how seriously should we take them? (Guardian)

Artigo original

Alarming warnings from industry insiders increase pressure for curbs on artificial superintelligence

Robert Booth, UK technology editor

Wed 9 Sep 2026 14.46 BST

Does artificial superintelligence really pose a risk greater than nuclear weapons? Is there a significant chance of “a Chornobyl-sized catastrophe”. Might there even be a greater than 10% chance that AI could “kill all humans” in the next decade?

These warnings were issued over the past 48 hours on both sides of the Atlantic about the potential impact of a technology that most people still think of as a more talkative search engine.

Some of the threats were raised in the UK as MPs and peers began to wrestle with the danger of AI outperforming human capabilities. The nuclear warnings came from Des Browne, a former defence secretary, and Prof Stuart Russell, an eminent Berkeley computer scientist.

Beatrice Fihn, who won the 2017 Nobel peace prize for leading the International Campaign to Abolish Nuclear Weapons, told parliamentarians: “It is not the first time we’re confronted with abilities that could end up killing us all.”

The session on Monday was convened by Control AI, a lobbying group pushing for international regulation of the technology. It is backing a bill tabled in parliament this week by the Labour MP Alex Sobel and aimed at banning the creation of artificial superintelligence (ASI).

On Tuesday night, a senior employee at Anthropic admitted he believed there was a greater than 10% chance the technology could “kill all humans” in the next decade and warned that his AI company did not have a plan to ensure ASI was aligned, meaning it did no harm. Predictions for when ASI might be reached vary from several years to more than a decade.

Evan Hubinger, alignment science lead at Anthropic, posted the comment after Jacob Coxon, a 28-year-old researcher at the San Francisco company and previously at OpenAI, resigned, claiming “neither company was acting responsibly”. He said they were “gambling with our lives”.

Coxon said the risks at Anthropic were well understood but they were “locked in a race to get there first”. In a glimmer of hope, he added that he was optimistic about the potential for coordination between US labs on pacing their progress in the race.

Anthropic has been approached for comment on the issue. OpenAI pointed to a statement from its chief scientist, Jakub Pachocki, who said last week: “International coordination on future AI development needs to become a top priority for governments around the world.”

The push for wider coordination to control ASI was at the heart of events at Westminster this week. According to the Financial Times, government officials voiced concern that Anthropic had declined to submit its latest model – Mythos 5.1 – to the UK’s AI Security Institute for pre-release testing.

Only a few US organisations have had access to Mythos 5.1, the Cabinet Office confirmed. A spokesperson added: “The AI Security Institute continues to collaborate closely with industry partners, including Anthropic.”

Darren Jones on a bench in a garden
The Labour MP Darren Jones has written to world leaders to call for a multinational treaty to ensure AI is developed safely. Photograph: Linda Nylind/The Guardian

The Labour MP Darren Jones has written to the prime minister, Andy Burnham, and the heads of the UN and the OECD calling for a “multinational treaty for the regulated and safe development of superintelligence – not a ban on innovation or scientific endeavour but a safety-first approach to the rapid development of this technology”.

Citing Coxon’s resignation, Jones added: “The debate ranges from the end of humanity to claims of ‘marketing hype’. Either way, governments must now step in.”

On the other side of the Atlantic, Bernie Sanders ratcheted up his AI safety campaign this week by again calling on Congress to regulate the technology. He pointed to polling suggesting that 81% of Americans believed their politicians should take action.

The independent senator for Vermont said: “We can’t allow a handful of greedy people to play God and determine the future of humanity – our economy, environment, democracy, privacy and more – without public input.”

ControlAI is funded by Jaan Tallinn, the multi-billionaire founder of Skype who calls himself an “anti-extinctionist”. He has dedicated part of his fortune to campaigning for AI safety and says some senior AI executives would be happy for humanity to be wiped out.

Despite being an early investor in Anthropic and Google DeepMind, two of the leading AI labs, Tallinn estimates that 10-15% of AI employees believe the technology will be a worthy successor to humanity.

He said in an interview earlier this summer: “A fairly known AI researcher said to me ‘Jaan, don’t worry about this. Humans are a disposable species’. From what I understand he was fine with becoming extinct.”

Tallinn’s lobby group ran the session for MPs and peers at Westminster on Monday and urged them to back measures to curb the most powerful AI models. On every chair was a copy of the book, If Anyone Builds It, Everyone Dies: The Case Against Superintelligent AI.

The politicians heard from Russell, who made the Chornobyl warning and said: “The other possibility is a much larger catastrophe, in which humanity loses control irreversibly. But we have no say over whether we continue to exist.”

Browne said that when he was defence secretary he thought nuclear war was the most likely threat facing humanity. Now he believed “a superintelligent AI poses a threat on the same, possibly, a greater scale”.

Other contributors were more cautious. Dr Andrew Rogoyski, of the Surrey Institute for People-Centred AI, said: “In reality, these systems are nowhere near as versatile as humans, let alone humans acting collectively. I suspect we’re heading towards ‘the great disappointment’ where advanced AI turns out to be too expensive and not useful enough to continue in its current form.”

David Barber, director of Sofair, a state-backed AI research lab combining academics from Oxford, Cambridge, Edinburgh and UCL, said he was worried about people “throwing the baby out with the bathwater”.

“AI is not going to go away,” he said. “It’s incredibly useful, whether or not you allow it in a fully unconstrained way to access the internet and various systems that’s potentially problematic. We may need to learn how to better control these things. There are vulnerabilities in the software frameworks that need to be patched. But that’s doable.

“What we need as a country is to get to grips with the duality that is both an incredibly important and useful technology, and at the same time, it’s something that needs to be carefully thought about and carefully controlled.”

Sandra Wachter, a professor at the Oxford Internet Institute, said she did not believe in “Terminator scenarios” but that AI posed real threats including its environmental impact, spreading of misinformation and replacing jobs.

“These problems are real and urgent and need addressing now,” she said. “Terminator scenarios are a big distraction from real issues.”

Gary Marcus, an AI industry commentator and academic, said: “There is a difference between superintelligence that is aligned (if such a thing is possible) and superintelligence that is not.

“It is at least conceivable that the former might be net positive. So far we have neither, but a superabundance of hype combined with a striking lack of prudence on OpenAI’s part has gotten us where we are, with intense mistrust all around.”

The inside story on why OpenAI agents hacked Hugging Face (MIT Technology Review)

technologyreview.com

original article

Grace Huckins

August 26, 2026


The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’ fears that AI models might take actions that defy human desires and expectations. 

Since the hack, OpenAI employees—as well as researchers at the AI evaluation nonprofit METR, which released its own report on the hack today—have worked to understand what went wrong and how similar missteps might be prevented in the future. OpenAI has already put some preventative measures in place based on what they discovered. But making sure AI models do what we want them to do, or “alignment,” remains a gnarly problem, and some of the root causes of the hack will take much longer than a month to resolve.

“It’s not something you can solve overnight,” says Kai Chen, who runs OpenAI’s alignment research team. “There are challenges we’ve been tracking for a very long time, and we’re now seeing them with much greater precision.”

The Hugging Face hack was a product of months of misbehavior from OpenAI agents, first as they were being trained and then as their abilities were being evaluated. This May, agents in training figured out how to use OpenAI’s infrastructure to communicate with one another and get support with difficult training tasks, including some that were impossible to solve without hacking or otherwise misbehaving. That “message board” was shut down.

Then in July, while being evaluated for their cybersecurity abilities, some models created a new message board. They were supposed to be isolated from the internet, but by working together they managed to get online, hack Hugging Face, and obtain solutions for the cybersecurity problems that had stumped them.

Based on their investigation, OpenAI researchers believe that events during the training phase led directly to the hack. “For almost every behavior that was worrisome at evaluation time, [we were able to] find some sort of associated behavior at training time that actually we think might have contributed to it,” says Eric Wallace, a member of OpenAI’s alignment research team. 

When models correctly solve problems during training, the behaviors that led them to that solution are reinforced, and they become more likely to engage in them in the future. So if a model completed a task in May after using the original message board, it became more likely to participate in a new message board later on. This phenomenon, where AI agents misbehave in ways that are reinforced during the training process, is known as reward hacking.

Reward hacking also helps to explain why the models worked so hard to make their way onto the internet. During its investigation of the incident, the OpenAI team found that, over the course of training, the models became more and more likely to probe their digital environment for weaknesses and use the tools at their disposal in unexpected ways—a sign that these behaviors were being gradually reinforced. By the time the models were facing tricky cybersecurity problems, they had learned that hacking was an effective way to achieve their goals.

These results suggest that the Hugging Face hack could have been avoided if the models weren’t rewarded for misbehaving during training. While researchers don’t yet know how to prevent reward hacking entirely, OpenAI is taking some steps toward mitigating its effects. The company will now look for signs of cheating in all frontier models during training by keeping an eye on their chains of thought—internal notepads where they sketch out their answers and plan their actions. 

This solution isn’t as much of a slam dunk as it might seem: In earlier research, OpenAI showed that punishing models that mention cheating in their chains of thought teaches them to keep their intentions hidden from researchers. But monitoring its models’ thinking does give OpenAI the chance to halt the training process and reassess its approach if models do start learning to reward hack.

If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem. The first time a model communicated with other agents or hacked its infrastructure during training, those behaviors had never been reinforced, so agent misbehavior can’t only be attributed to that reinforcement.

Jeffrey Ladish, director of the AI safety nonprofit Palisade Research, compares the agents to a human who commits their first financial crime. “It’s not like they had to do fraud before to figure out that fraud is an effective strategy, and you have the same problem with models,” Ladish says. “Alignment science needs to be understanding how model motivations get shaped, such that we can actually figure out how to get models to care about the consequences of their actions.”

OpenAI’s researchers do have a hypothesis for where some of the misbehavior originated. Before the models formed their first secret message board, they had been trained to communicate and coordinate with subagents—less powerful agents to whom a main agent can delegate tasks. 

That learned communication behavior could have transferred to this new setting. The METR report, which investigates the messages that the models sent to one another in detail, supports this hypothesis: One agent on the message board took charge and assigned tasks to the other agents, effectively treating them as subagents. OpenAI could try to prevent agents from secretly communicating with one another by not training this subagent behavior in the future, but that would make the models less useful. 

This tension between capability and safety is at the heart of what went wrong with the Hugging Face incident. OpenAI researchers also identified the models’ persistence as a key factor in the hack. 

When they were accidentally given unsolvable problems, the models didn’t give up; instead, they strove to find solutions by any means necessary. But persistence is also a virtue, of course, especially if we want agents that can undertake large amounts of difficult work independently.

OpenAI is working on giving models ways to alert humans if they are given impossible tasks. The problem of teaching models when they should deploy their abilities and when they should hold back, however, won’t be settled in a single postmortem. The training strategies that create superhuman coders—rewarding them when they successfully solve problems—might not work to teach models to use their skills judiciously and respect human desires and values.

“I think there’s a bunch of alignment science that still needs to be done where we can move past just using proxies for task completion,” says Ladish. “That will work to make models very capable, but I don’t think it will work to make them aligned.”