Artificial intelligence (AI) might not be as self-improving as hoped. A new study found that even advanced AI agents lack the creativity and judgment needed for open-ended research, crucial for recursive self-improvement.
Researchers from Princeton University evaluated these AI agents on their ability to tackle complex questions in machine learning without clear answers, much like tasks encountered at top-tier conferences. The results: AI agents could handle engineering problems but struggled with originality and critical thinking, failing to produce novel contributions.
The study used a method called “shadow evaluation,” where the AI was asked to answer research questions from unpublished papers, providing an unbiased test of their capabilities. Despite being given ample resources, both attempts resulted in rejected submissions by human reviewers, highlighting limitations in AI’s current training and decision-making processes.
While these agents demonstrated proficiency in conducting experiments and compiling data, they were unable to explore different ideas creatively or adapt when approaches failed. This suggests that the training methods used for AI might need a reevaluation if we want truly autonomous research.
The findings raise questions about timelines for fully automated AI research and whether current models are ready for tasks requiring human-like judgment and creativity.







