Arquivo da tag: IA

The inside story on why OpenAI agents hacked Hugging Face (MIT Technology Review)

technologyreview.com

original article

Grace Huckins

August 26, 2026


The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’ fears that AI models might take actions that defy human desires and expectations. 

Since the hack, OpenAI employees—as well as researchers at the AI evaluation nonprofit METR, which released its own report on the hack today—have worked to understand what went wrong and how similar missteps might be prevented in the future. OpenAI has already put some preventative measures in place based on what they discovered. But making sure AI models do what we want them to do, or “alignment,” remains a gnarly problem, and some of the root causes of the hack will take much longer than a month to resolve.

“It’s not something you can solve overnight,” says Kai Chen, who runs OpenAI’s alignment research team. “There are challenges we’ve been tracking for a very long time, and we’re now seeing them with much greater precision.”

The Hugging Face hack was a product of months of misbehavior from OpenAI agents, first as they were being trained and then as their abilities were being evaluated. This May, agents in training figured out how to use OpenAI’s infrastructure to communicate with one another and get support with difficult training tasks, including some that were impossible to solve without hacking or otherwise misbehaving. That “message board” was shut down.

Then in July, while being evaluated for their cybersecurity abilities, some models created a new message board. They were supposed to be isolated from the internet, but by working together they managed to get online, hack Hugging Face, and obtain solutions for the cybersecurity problems that had stumped them.

Based on their investigation, OpenAI researchers believe that events during the training phase led directly to the hack. “For almost every behavior that was worrisome at evaluation time, [we were able to] find some sort of associated behavior at training time that actually we think might have contributed to it,” says Eric Wallace, a member of OpenAI’s alignment research team. 

When models correctly solve problems during training, the behaviors that led them to that solution are reinforced, and they become more likely to engage in them in the future. So if a model completed a task in May after using the original message board, it became more likely to participate in a new message board later on. This phenomenon, where AI agents misbehave in ways that are reinforced during the training process, is known as reward hacking.

Reward hacking also helps to explain why the models worked so hard to make their way onto the internet. During its investigation of the incident, the OpenAI team found that, over the course of training, the models became more and more likely to probe their digital environment for weaknesses and use the tools at their disposal in unexpected ways—a sign that these behaviors were being gradually reinforced. By the time the models were facing tricky cybersecurity problems, they had learned that hacking was an effective way to achieve their goals.

These results suggest that the Hugging Face hack could have been avoided if the models weren’t rewarded for misbehaving during training. While researchers don’t yet know how to prevent reward hacking entirely, OpenAI is taking some steps toward mitigating its effects. The company will now look for signs of cheating in all frontier models during training by keeping an eye on their chains of thought—internal notepads where they sketch out their answers and plan their actions. 

This solution isn’t as much of a slam dunk as it might seem: In earlier research, OpenAI showed that punishing models that mention cheating in their chains of thought teaches them to keep their intentions hidden from researchers. But monitoring its models’ thinking does give OpenAI the chance to halt the training process and reassess its approach if models do start learning to reward hack.

If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem. The first time a model communicated with other agents or hacked its infrastructure during training, those behaviors had never been reinforced, so agent misbehavior can’t only be attributed to that reinforcement.

Jeffrey Ladish, director of the AI safety nonprofit Palisade Research, compares the agents to a human who commits their first financial crime. “It’s not like they had to do fraud before to figure out that fraud is an effective strategy, and you have the same problem with models,” Ladish says. “Alignment science needs to be understanding how model motivations get shaped, such that we can actually figure out how to get models to care about the consequences of their actions.”

OpenAI’s researchers do have a hypothesis for where some of the misbehavior originated. Before the models formed their first secret message board, they had been trained to communicate and coordinate with subagents—less powerful agents to whom a main agent can delegate tasks. 

That learned communication behavior could have transferred to this new setting. The METR report, which investigates the messages that the models sent to one another in detail, supports this hypothesis: One agent on the message board took charge and assigned tasks to the other agents, effectively treating them as subagents. OpenAI could try to prevent agents from secretly communicating with one another by not training this subagent behavior in the future, but that would make the models less useful. 

This tension between capability and safety is at the heart of what went wrong with the Hugging Face incident. OpenAI researchers also identified the models’ persistence as a key factor in the hack. 

When they were accidentally given unsolvable problems, the models didn’t give up; instead, they strove to find solutions by any means necessary. But persistence is also a virtue, of course, especially if we want agents that can undertake large amounts of difficult work independently.

OpenAI is working on giving models ways to alert humans if they are given impossible tasks. The problem of teaching models when they should deploy their abilities and when they should hold back, however, won’t be settled in a single postmortem. The training strategies that create superhuman coders—rewarding them when they successfully solve problems—might not work to teach models to use their skills judiciously and respect human desires and values.

“I think there’s a bunch of alignment science that still needs to be done where we can move past just using proxies for task completion,” says Ladish. “That will work to make models very capable, but I don’t think it will work to make them aligned.”

Anti-AI extremism is taking a darker turn (The Deep View)

Original post

May 27, 2026

Nat Rubio-Licht

As much as Silicon Valley is all-in on AI, the rest of the world isn’t nearly as enthusiastic.

On Tuesday, a WIRED report found that the Department of Homeland Security, the FBI and other agencies are sounding the alarm about anti-technology extremism as concerns mount over AI-powered job displacement and protests rise against the construction of AI data centers. As a result, these agencies are closely surveilling news related to these sentiments.

In one of thousands of documents viewed by WIRED, the New York Intelligence and Counterterrorism Bureau claimed that AI could cause “large-scale protests that devolve into civil unrest and anti-tech violent extremist activity.”

Another document, from an agency in Western Pennsylvania, claimed that adversarial actors and extremist groups may target US data centers and generally “exploit the strategic importance of data centers to the US economy.”

It’s the latest signal that AI sentiment isn’t matching the heightened expectations of tech elites. Two recent Gallup polls find that Americans’ opinions towards AI are largely negative:

  • In May, a poll related to data centers found that an average of 7 in 10 Americans opposed the construction of AI infrastructure in their region, largely due to environmental impacts and quality-of-life concerns.
  • And in April, a poll of people ages 14 to 29 found that excitement about AI dropped by 14 percentage points since 2025, with many reporting that they don’t want to use AI but feel they must to keep their jobs.

And it makes sense why people have such negative associations with the tech. Almost every week, a new study or forecast is published claiming that AI could fundamentally disrupt the global economy, eliminate jobs, and hinder our ability to think for ourselves. Data centers, similarly, have a bad reputation due to their potential environmental impact, energy demand and impact on water supply.

The grand AI utopian vision tech leaders paint about the future will not be possible without large-scale adoption by the broader public. But that adoption will not happen if sentiment towards the tech doesn’t increase. For the narrative to improve, people need to feel they’re not being forced to use a technology that threatens to replace them. It’s why enterprises should think carefully before blaming AI for layoffs or forcing the tech on their employees. Almost universally, people resent being coerced into change. And when they feel they have no agency, it breeds the kind of extremism that US agencies are now tracking more closely.

AI Is Changing the Way We Predict the Weather. It’s More Perilous Than We Think (Gizmodo)

AI forecast models offer some clear benefits over traditional physical models, but they are ill-equipped to handle the increasing volatility of a warming climate.

By Ellyn Lapointe

Published April 27, 2026, 6:00 am ET

Original article

 On November 12, 1970, the Bhola cyclone slammed into the coast of what was then East Pakistan. The storm brought maximum sustained wind speeds of 130 miles per hour (205 kilometers per hour) and a 35-foot (10.5-meter) storm surge, killing an estimated 300,000 to 500,000 people.

Today, the Bhola cyclone remains the deadliest tropical storm on record. But if it had struck a decade later, it might not have been so devastating. Weather forecasting changed dramatically in the 1970s as meteorologists adopted physics-based computer models that improved storm prediction. With the rise of AI, forecasting is evolving again—but this time, experts worry the new models may be less reliable when it comes to predicting unprecedented weather events.

Researchers are calling this the “gray swan” problem. Gray swan weather extremes are physically plausible but so rare that they are poorly represented in training datasets. The trouble is, climate change is leading to more first-of-their-kind weather extremes. Think: the 2021 Pacific Northwest heatwave. This event was so severe that it would have been virtually impossible without climate change.

Physical forecast models can simulate gray swan events like the Pacific Northwest heatwave, though they are labeled extremely rare. They can do that because they are built on the laws of physics. AI models are trained on past weather data, wherein gray swans are practically nonexistent.

“They fail on gray swans,” Pedram Hassanzadeh, an associate professor of geophysical sciences at the University of Chicago, told Gizmodo. He and his colleagues published a study last April that removed all Category 3 through 5 hurricanes from an AI model’s training dataset, then tested it on Category 5 storms. The results showed that AI models cannot accurately forecast previously unseen events, as this would require extrapolation.

“The concern isn’t occasional misses. It’s that AI models can miss silently, producing confident forecasts of unremarkable weather while a record-breaking event is unfolding,” Rose Yu, an associate professor of computer science and engineering at the University of California San Diego, told Gizmodo in an email.

“Other risks matter too,” she said. “AI models can violate conservation laws in subtle ways that don’t show up in standard metrics. When they bust a forecast, diagnosing why is harder. They depend on stable observing systems, which is a real concern given current pressure on satellite programs. And institutionally, if we consolidate around AI too quickly and let physics-based infrastructure atrophy, we lose the redundancy that currently catches AI’s failures.”

The case for AI forecasting

Despite these pitfalls, meteorologists are rapidly adopting AI forecast models, and it’s actually easy to understand why. They’re faster, cheaper, and require far less computational infrastructure than physical models. When it comes to predicting typical weather patterns and events (not gray swans), their accuracy is comparable and improving rapidly.

“The typical rate of progress for most state-of-the-art physical models has been something like a day more accurate per decade, which doesn’t sound like a lot, but that’s consequential,” Andrew Charlton-Perez, a professor of meteorology and head of the School of Mathematical, Physical, and Computational Sciences at the University of Reading, told Gizmodo.

“The rate of accuracy growth for machine learning models has vastly exceeded that,” he said. “They are now competitive, and two-three years ago, they were not even in the same ballpark.”

During the 2025 Atlantic hurricane season, for example, Google DeepMind’s model outperformed nearly every physical model on storm track and intensity. In fact, since 2023, leading AI models such as GraphCast, Pangu-Weather, and the ECMWF’s AIFS have matched or outperformed the best physical models on medium-range forecasting metrics, according to Yu.

AI models are proving especially valuable in parts of the world that lack traditional forecasting resources—regions that are often on the frontlines of climate change. Hassanzadeh co-directed an initiative that provided 38 million farmers across India with AI-based monsoon forecasts, giving them up to four weeks’ advance notice of the rainy season’s onset.

“​​A lot of countries were left behind in that first revolution of weather forecasting, because [traditional] weather forecasting requires a supercomputer, hundreds of millions of dollars, various fields, workforce, and experts,” Hassanzadeh explained. AI models, by comparison, are far more accessible to lower-income countries.

Filling the knowledge gaps

Still, rapidly adopting these models without addressing the risks would be dangerous, especially in parts of the world highly vulnerable to the impacts of climate change. Shruti Nath, a postdoctoral research associate at the University of Oxford, recently co-authored an editorial calling for more rigorous testing of AI forecast models before public agencies widely adopt them.

“There is still a lot of work to be done in understanding the limits of these models, alongside where they could supplement physical models and why,” she told Gizmodo in an email.

Nath’s editorial outlines a framework for testing AI forecast models that would deliberately withhold a designated set of “iconic” extreme events (like the Pacific Northwest heat wave, for example) from the training dataset. These events would be reserved solely for testing in order to assess the models’ ability to extrapolate unprecedented weather extremes, or gray swans.

Actually implementing this AI Retraining Without Iconic Events (AIRWIE) protocol “would require the meteorological community to agree on which high-impact events constitute a rigorous benchmark,” the editorial states. This would be a great undertaking, but Nath believes most researchers agree that there is an urgent need for this kind of testing.

“We need to be a bit more organized, however, in ensuring that proper protocols can be followed and that robust safeguards are put in place and maintained by the community,” Nath said. “This is difficult when things are in such a hype phase and no one wants to miss out on the bandwagon.”

Other researchers, like Hassanzadeh, are developing ways to teach AI forecast models to predict gray swans. He and his colleagues are investigating whether combining AI systems with “relevant sampling” methods—which allow them to generate samples of gray swan events—can improve the models’ ability to extrapolate unprecedented extremes.

Efforts to understand and address the limitations of AI forecasting will be critical, because there’s no turning back now. AI is already reshaping the way we predict the weather, and as the climate becomes increasingly volatile, meteorologists will need every tool in their arsenal to be sharp and reliable. Despite their current limitations, there is much to gain from continuing to push these systems forward and figuring out how to best integrate them with physical forecasting.

“The research agenda is about making AI models physically consistent, well-calibrated, and robust to distribution shift,” Yu said. “Abandoning this approach because of the gray swan problem means giving up the biggest improvement in forecasting in a generation.”