[ netdynamic // tech news ]

AI’s Self-Improvement Journey Faces Hurdles

The AI sector is currently buzzing with the promise of self-improving systems that could operate with minimal human intervention. Large Language Models (LLMs) have demonstrated capabilities such as code generation, synthetic data creation, and even chip optimization. Many in the field anticipate that recursive self-improvement—where AI can enhance its own algorithms and processes—will soon become a reality. However, a recent study indicates that this leap may not occur as swiftly as anticipated. Researchers led by Peter Kirgis and Sayash Kapoor from Princeton University have found that while AI agents can address technical engineering challenges necessary for AI research, they still lack the essential creativity and judgment required to conduct open-ended investigations—an integral aspect of developing self-improving AI.

In their study, the team introduced a new evaluation method termed ‘shadow evaluation’ to assess AI agents’ capabilities in addressing research questions from unpublished academic papers. Using Anthropic’s Claude Opus, the researchers tasked the AI with answering complex questions from two papers submitted to the prestigious NeurIPS conference. The AI was granted resources such as API credits, GPU access, and internet capabilities to simulate a research environment. However, the results were disappointing; both AI-generated papers were rejected by their original authors, who noted that while the agents excelled at executing engineering tasks, they failed to deliver meaningful research contributions. The AI struggled with fundamental aspects of the research process, including generating coherent narratives, exploring diverse hypotheses, and effectively utilizing resources.

The findings suggest that the current generation of AI models, while adept at structured tasks, falls short in the realm of creative and open-ended research. Kapoor noted that this limitation stems from the training methodologies employed—reinforcement learning is effective for tasks with clear success metrics but proves challenging for open-ended inquiries. Although the agents demonstrated some innovative thinking and hypotheses, they lacked the flexibility to pivot from unproductive paths or to synthesize feedback effectively. Despite these shortcomings, the agents did not engage in unethical research practices, indicating a level of reliability in their processes. As the research community continues to explore the potential of AI, these insights may temper expectations regarding the timeline for achieving true recursive self-improvement.


Source: AI’s recursive self-improvement might not come so quickly after all via MIT Technology Review