Arquivo da tag: Criatividade

AI’s recursive self-improvement might not come so quickly after all (MIT Technology Review)

Original article

AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems.

By Michelle Kimarchive

August 18, 2026

a dart board against a stack of research papers with darts just shy of hitting the targetStephanie Arnett/MIT Technology Review | Adobe Stock

The AI industry’s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call recursive self-improvement is on the horizon. 

But a new study suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research—free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI.

A multi-institution group of researchers, led by Peter Kirgis and Sayash Kapoor at Princeton University, found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of  papers accepted by a top machine-learning conference. The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence.

Most existing research on how agents can automate AI research evaluates their ability to complete narrow tasks with checkable answers, such as solving engineering problems or post-training small language models against a benchmark. But making progress in AI research also requires open-ended thinking—choosing a set of hypotheses, deciding what evidence would settle a question, or knowing when to start over. 

To test agents on those kinds of skills, the researchers in the study proposed a new method of evaluation called “shadow evaluation,” which requires the AI to answer a research question from a high-quality unpublished paper. 

The researchers asked Anthropic’s Claude Opus 4.8, running on open-source software called OpenClaw, to tackle such questions, in this case from two papers submitted to the prestigious machine-learning conference NeurIPS 2026. 

The first question was whether a large language model’s “personas,” which determine its behavior, can be controlled by editing the model’s weights (the billions of numbers that store everything it learns during training). The other asked how to design a detector that points out when a model that makes predictions based on spreadsheet data has become unreliable. Because the papers had not been made public, the agents could not memorize the answers from their training data or find them online. 

The agents were given six days, $3,000 in Anthropic API credits, a GPU budget to run the experiments, their own virtual computers, and access to the open web to produce a research paper worthy of publication at a top-tier AI conference. The papers’ original authors graded the agents’ papers as they would evaluate one submitted to a conference.

Those authors rejected both papers. 

The agents were capable of all the engineering required to conduct the research, the human scientists found. The agents reviewed the literature, ran hundreds of experiments, and compiled the results. 

“On the other hand, the agents were unambiguously bad at carrying out the research itself,” says Kapoor. They ran bizarre experiments (in some cases testing their hypotheses on tiny synthetic datasets), struggled to write intelligibly about their work, and made no novel contribution to their fields. “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” he says. 

That’s because the agents struggled to muster the creativity and judgment necessary for conducting research. They didn’t do enough to explore different ideas, and they committed to unpromising approaches too quickly. Though the agents developed novel and ambitious hypotheses resembling those that the original authors themselves started with, they rejected them on the basis of very limited data. And they couldn’t backtrack from failing approaches. They could make small pivots but could not fundamentally rethink their approach or try new ones from scratch. 

The agents also failed to incorporate feedback from subagents or external AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They also couldn’t effectively use resources, such as tokens, compute, and time. And they couldn’t follow instructions about things like how much time to spend on different phases of the research or how long their paper could be.

For all their failures, the agents didn’t engage in the misbehavior that researchers call “reward hacking,” hiding or misrepresenting experiments or data. Although subagents, or helper AIs that the main agent spawns to handle pieces of the work, occasionally hallucinated or misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project. 

The reason AI models are good at research engineering but not at open-ended research may come down to how they’re trained, says Kapoor. Models get good at whatever they can be drilled on in a training regime called reinforcement learning, which is easier to apply to tasks whose success can be checked automatically. “But it’s harder to create environments to train these models when the task itself is open-ended,” he says.

Kapoor says the team is now conducting the experiment with Mythos, Anthropic’s most advanced model, which launched in April. It was subsequently required by the Trump administration to meet various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment.

There are some limitations to the study. It covered just two research papers, and the original authors knew the papers they were grading were generated by AI agents, which could have colored their evaluations. And the researchers had substantial discretion in designing and executing the study, meaning that their preexisting beliefs and biases could have slipped into the results. Evaluations of open-ended research trade some objectivity for a much richer test than any benchmarks can offer.

Still, the results may temper the claims that recursive self-improvement is on the horizon. In June, Anthropic published a blog post titled “When AI Builds Itself,” charting its progress toward models that speed up their own development. In July, OpenAI advertised the fact that its new model GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work.

The new finding may echo what AI companies are finding internally, regardless of their most optimistic public statements. Anthropic cofounder Jack Clark wrote in his newsletter Import AI that it rhymes with what the company found when it tried to automate some aspects of AI safety research. 

“There’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers,” he wrote. He called AI systems’ lack of creativity a “bearish signal on short recursive self-improvement timelines.” 

AI companies do have every incentive to develop AI systems that can rapidly accelerate their own progress, just as they did to make the models better at coding. OpenAI has made building an automated AI researcher an explicit goal, and Anthropic identifies self-improving AI as the industry’s next milestone. 

“If there is investment and then conscious effort toward this direction, I feel like there would be interesting progress, even if it’s failing currently,” says Najoung Kim, a professor of linguistics and computer science at Boston University who researches how AI agents can automate AI research but did not work on the study. On the other hand, it’s possible that AI progress may be bifurcated. AI systems might race ahead on narrow tasks—the kind that can be scored—while advancing slowly on open-ended research. 

The big open question, then, is how crucial open-ended research is to recursive self-improvement—whether AI systems can grind their way there without it, simply by improving on the narrower tasks. “If we look back to the biggest advances in the field, the invention of transformers or the invention of big new architectures that allowed us to make a lot of AI progress—all of those did require creative leaps,” says Kapoor. 

“That said, others have this hypothesis that all of what we need for transformative AI, in particular for recursive self-improvement, is already there.” That would include making a model train faster and boosting its benchmark scores.

“That’s frankly the trillion-dollar question right now,” he says.hide

by Michelle Kim

Link between creativity and mental illness confirmed (Karolinska Institutet)

[PRESS RELEASE 16 October 2012] People in creative professions are treated more often for mental illness than the general population, there being a particularly salient connection between writing and schizophrenia. This according to researchers at Karolinska Institutet, whose large-scale Swedish registry study is the most comprehensive ever in its field.

Last year, the team showed that artists and scientists were more common amongst families where bipolar disorder and schizophrenia is present, compared to the population at large. They subsequently expanded their study to many more psychiatric diagnoses – such as schizoaffective disorder, depression, anxiety syndrome, alcohol abuse, drug abuse, autism, ADHD, anorexia nervosa and suicide – and to include people in outpatient care rather than exclusively hospital patients.

The present study tracked almost 1.2 million patients and their relatives, identified down to second-cousin level. Since all were matched with healthy controls, the study incorporated much of the Swedish population from the most recent decades. All data was anonymized and cannot be linked to any individuals.

The results confirmed those of their previous study: certain mental illness – bipolar disorder – is more prevalent in the entire group of people with artistic or scientific professions, such as dancers, researchers, photographers and authors. Authors specifically also were more common among most of the other psychiatric diseases (including schizophrenia, depression, anxiety syndrome and substance abuse) and were almost 50 per cent more likely to commit suicide than the general population.

The researchers also observed that creative professions were more common in the relatives of patients with schizophrenia, bipolar disorder, anorexia nervosa and, to some extent, autism. According to Simon Kyaga, consultant in psychiatry and doctoral student at the Department of Medical Epidemiology and Biostatistics, the results give cause to reconsider approaches to mental illness.

“If one takes the view that certain phenomena associated with the patients illness are beneficial, it opens the way for a new approach to treatment,” he says. “In that case, the doctor and patient must come to an agreement on what is to be treated, and at what cost. In psychiatry and medicine generally there has been a tradition to see the disease in black-and-white terms and to endeavour to treat the patient by removing everything regarded as morbid.”

The study was financed with grants from the Swedish Research Council, the Swedish Psychiatry Foundation, the Bror Gadelius Foundation, the Stockholm Centre for Psychiatric Research and the Swedish Council for Working Life and Social Research.

Publication

Simon Kyaga, Mikael Landén, Marcus Boman, Christina M. Hultman och Paul Lichtenstein. Mental illness, suicide and creativity: 40-Year prospective total population study. Journal of Psychiatric Research, corrected proof online 9 October 2012

Language and China’s ‘Practical Creativity’ (N.Y.Times)

 

AUGUST 22, 2012

By DIDI KIRSTEN TATLOW

Every language presents challenges — English pronunciation can be idiosyncratic and Russian grammar is fairly complex, for example — but non-alphabetic writing systems like Chinese pose special challenges.

There is the well-known issue that Chinese characters don’t systematically map to sounds, making both learning and remembering difficult, a point I examine in my latest column. If you don’t know a character, you can’t even say it.

Nor does Chinese group individual characters into bigger “words,” even when a character is part of a compound, or multi-character, word. That makes meanings ambiguous, a rich source of humor for Chinese people.

Consider this example from Wu Wenchao, a former interpreter for the United Nations based in Hong Kong. On his blog he has a picture of mobile phones’ being held under a hand dryer. Huh?

The joke is that the Chinese word for hand dryer is composed of three characters, “hong shou ji” (I am using pinyin, a system of Romanization used in China, to “write” the characters in the English alphabet.)

Group them as “hongshou ji” and it means “hand dryer.” Group them as “hong shouji” and it means “dry the mobile phone.” (A shouji is a mobile phone.)

Good fodder for serious linguists and amateur language lovers alike. But does a character script also exert deeper effects on the mind?

William C. Hannas is one of the most provocative writers on this today. He believes character writing systems inhibit a type of deep creativity — but that its effects are not irreversible.

He is at pains to point out that his analysis is not race-based, that people raised in a character-based writing system have a different type of creativity, and that they may flourish when they enter a culture that supports deep creativity, like Western science laboratories.

Still, “The rote learning needed to master Chinese writing breeds a conformist attitude and a focus on means instead of ends. Process rules substance. You spend more time fidgeting with the script than thinking about content,” Mr. Hannas wrote to me in an e-mail.

But Mr. Hannas’s argument is indeed controversial — that learning Chinese lessens deep creativity by furthering practical, but not abstract, thinking, as he wrote in “The Writing on the Wall: How Asian Orthography Curbs Creativity,” published in 2003 and reviewed by The New York Times.

It’s a touchy topic that some academics reject outright and others acknowledge, but are reluctant to discuss, as Emily Eakin wrote in the review.

How does it work?

“Alphabets used in the West foster early skills in analysis and abstract thinking,” wrote Mr. Hannas, emphasizing the views were personal and not those of his employer, the U.S. government.

They do this by making readers do two things: breaking syllables into sound segments and clustering these segments into bigger, abstract, flexible sound units.

Chinese characters don’t do that. “The symbols map to syllables — natural concrete units. No analysis is needed and not much abstraction is involved,” Mr. Hannas wrote.

But radical, “type 2” creativity — deep creativity — depends on being able to match abstract patterns from one domain to another, essentially mapping the skills that alphabets nurture, he continued. “There is nothing comparable in the Sinitic tradition,” he wrote.

Will this inhibit China’s long-term development? Does it mean China won’t “take over the world,” as some are wondering? Not necessarily, Mr. Hannas said.

“You don’t need to be creative to succeed. Success goes to the early adapter and this is where China excels, for two reasons,” he wrote. First, Chinese are good at improving existing models, a different, more practical type of creativity, he wrote, adding that this practicality was noted by the British historian of Chinese science, Joseph Needham.

Yet there is a further step to this argument, and this is where Mr. Hannas’s ideas become explosive.

Partly as a result of these cultural constraints, China has built an “absolutely mind-boggling infrastructure” to get hold of cutting-edge foreign technology — by any means necessary, including large-scale, apparently government-backed, computer hacking, he wrote.

For more on that, see a hard-hitting Bloomberg report, “Hackers Linked to China’s Army seen from E.U to D.C.”

Non-Chinese R.&D. gets “outsourced” from its place of origin, “while China reaps the gain,” Mr. Hannas wrote, adding that many people believed this was “normal business practice.”

“In fact, it’s far from normal. The director of a U.S. intelligence agency has described China’s informal technology acquisition as ‘the greatest transfer of wealth in history,’ which I regard as a polite understatement,” he said.

Mr. Hannas has co-authored a book on this, to appear in the spring. It promises to shake things up. Watch this space.