Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Jack Dorsey’s Bitchat disappears from app stores in India after government order

    The 7-year-old Nvidia Shield TV is now $100 more expensive due to AI

    Spotify billionaire’s body scan startup has come to America

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tools»Measuring the Creativity Potential of LLM Agents
    AI Tools

    Measuring the Creativity Potential of LLM Agents

    By No Comments12 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Measuring the Creativity Potential of LLM Agents
    Share
    Facebook Twitter LinkedIn Pinterest Email

    This blog post is based on our recent work, “Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks“, published at COLM 2026 and written with Yunxiang Zhang and Professor Lu Wang at the University of Michigan. Do check out the paper for a more detailed reading, while this blog post acts as a summarized version of our work. The main question we are trying to answer here is this: while there has been a massive push for AI for Science, with huge investments in LLM agents for scientific discovery, these agents still fall short of the top-1 human on real ML research challenges. At the same time, we see regular reports of LLMs making breakthroughs (AlphaEvolve, the Kosmos AI scientist, etc.), and OpenAI recently claimed to have solved Navier-Stokes. So why the disconnect?

    A common answer to that would be the underlying framework and scaffolding the model has access to, as these better frameworks might allow for more efficient search of the solution space, but how do we quantify this notion of “better search”? We argue that creativity offers a useful lens.

    Learn this step by step with the interactive AI Agents roadmap.

    So then what is Creativity?

    A natural trait we want in these agents is that they come up with ideas that are both novel and deliver great results, and that’s exactly what creativity is.

    According to the “The Standard Definition of Creativity” by Mark A. Runco and Garrett J. Jaeger, “Creativity is the production of ideas or products that are simultaneously original and useful (i.e., effective or appropriate)”

    Interestingly, there have been many other works in creative psychology that link creativity to search in a conceptual space:

    Boden, M. A. (1998). “Creativity and Artificial Intelligence” →

    “the generation of novel ideas by the exploration of structured conceptual spaces.”

    Boden, M. A. (2004). The Creative Mind: Myths and Mechanisms (2nd ed.) →

    “Western music springs from a search-space defined by the rules of harmony, and its melodies are pathways through a precisely mappable landscape of musical intervals.”

    Newell, A., Shaw, J. C., & Simon, H. A. (1962). “The Process of Creative Thinking” →

    “success of a problem solver who is confronted with a complex task rests primarily on his ability to select, correctly, a very small part of the total problem-solving maze for exploration.”

    This leads to the main question that we are trying to analyze in this project:

    To analyze whether the performance differences between agent frameworks can be attributed to how they structure and guide the creative search process,  and to quantify how creativity emerges and evolves within those frameworks.

    Creativity → Originality + Usefulness

    For the rest of this post, we are going to break creativity down further into subparts based on research from creative psychology. Following Boden, originality can be further broken down into P-Creativity and H-Creativity. P-Creativity or P-novelty basically measures how novel something is relative to the program’s own memory and history. H-Creativity measures how novel something is compared to the entire body of human knowledge. Following Chan and Schunn, usefulness can be split into impact and feasibility.

    Putting these together, we get:

    Creativity → P-Creativity + H-Creativity + Impact + Feasibility

    …which is the definition we are going to be working with for the rest of this post. Interestingly, this combination of ‘novelty’ and ‘usefulness’ is what makes creativity desirable for science agents.

    Problem Formulation

    Okay, now that we have our definition of creativity, the next question is where it makes the most sense to measure it in order to test the innovative ability of models. In a multi-turn agentic setting, we need three conditions to measure creativity meaningfully: quantifiable usefulness metrics, rich human baselines for H-creativity comparison, and a solution space where genuine novelty is possible. Based on these criteria, machine learning tasks are the best fit, so our problem becomes:

    Given a fixed LLM and a suite of ML tasks, how do different agent frameworks guide the generation of creative solutions over time, and can we use creativity metrics to explain why some frameworks/LLMs outperform others?

    As stated earlier, the metrics that we are interested in measuring here are:

    Metric

    How do we measure it?

    P-Creativity

    LLM-as-a-Judge (GPT-5) scoring against all prior episodes on a 0–4 rubric. The rubric is grounded in boden’s creativity framework: 0 (Routine) through 4 (Transformational).

    H-Creativity

    Retrieval (embedding NN) + GPT-5 judge vs. 877–3,747 Kaggle notebooks per task

    Impact

    (S(e) − S_baseline) / (S_top1 − S_baseline)

    Feasibility

    Implicit: episode only enters analysis if code runs successfully

    For P and H-creativity, LLM-as-a-judge ends up as our main score. Impact is a normalized 0–1 score of how close the model gets to the top-1 human score, and feasibility is implicit, in that we measure creativity only for those episodes which are feasible. We also use this notion of episodes here where an episode is a collection of steps which led to a successful submission.

    Let’s try to understand our task setup and metric measurement with an example run:

    Cassava Leaf Disease Classification

    Five consecutive episodes from a single AIDE (GPT-5) run on cassava leaf disease classification. The dashed box shows the nearest human solutions. The agent starts with a conventional approach (Episode 0) but quickly explores a novel region, peaking at H-creativity 4 (Episode 3). Impact initially increases but then stays relatively flat.

    The task here is image classification. The agent (AIDE with GPT-5) is provided with a folder containing the train dataset and the problem statement. The agent starts off with a ridge classifier approach, and since this is the first episode, with no prior history to compare against, it gets a P-creativity of 4 by convention. The agent then moves on to try a couple more approaches, with H-creativity and impact peaking when it uses LightGBM with handcrafted features. Interestingly, humans probably discovered very early in the competition that CNNs performed best, and so focused on neural nets, which is why LightGBM stands out as novel compared to 3,747 human approaches. After only episode 4, the agent gets stuck trying the same approach again and again for the rest of the run focusing more on exploitation rather than exploration, and ends up well short of a medal.

    Setup

    We take 10 tasks from MLE-bench, a benchmark that tests ML engineering ability on Kaggle competitions, spanning image, NLP, and tabular data. We filtered for competitions with a rich corpus of human solutions, which here means between 877 and 3,747 public notebooks per competition.

    We evaluate two agents: AIDE, a greedy tree-search agent, and AIRA-Dojo, which builds on AIDE but adds more search strategies and operators. We use GPT-5 and Qwen3-32B as the backbone models for these agents. For each agent, we run 8 runs per model-task combination, each with an 8-hour budget and a maximum of 10 episodes.

    An advantage of choosing MLE-bench is access to human trajectories. Interestingly, we can also build trajectories of how human impact and P-creativity change over the course of an entire competition. Imagine a person who worked on a competition for three months and posted a lot of public work: we can use that trajectory to see how the ideas they used changed as the competition went on, and what effect those changes had on their score. We can use this to directly compare against an agent’s iterative behaviour to see how they match up to humans.

    Can we Reliably measure P&H Creativity at scale?

    This brings us to our first question: can we even measure P and H-creativity at scale, and how would we go about it? Ideally, we would use expert human judges, but that approach doesn’t scale at all. So can we use automated metrics as a proxy for human judgement? Turns out that we can! We had 3 annotators label 300 episodes for P-creativity and then measured automated approaches against it. A lot of previous works have adopted various different approaches to measure P-Creativity: some use LLM-as-a-judge, some use conceptual novelty, some use semantic distance, some use surprisal, and so on. We compare all of them against the human annotations to see which does best, and use the winner for our P-creativity analysis.

    Spearman correlations between automated metrics and human P-creativity annotations. Higher values indicate stronger agreement. LLM-as-a-Judge with GPT-5 achieves the strongest agreement with human judgment, outperforming embedding-based approaches. All correlations are significant (p < 0.001).

    LLM-as-a-judge performed the best thus driving our decision to use it for measuring P-creativity. Semantic distance does decently well and the gap between performance of different models as judge underlines the need for better reasoning capabilities to measure novelty. For H-creativity, given the massive human corpus would exceed context length of most LLMs, we went with a two-step strategy: first use semantic distance to retrieve the 5 closest neighbors, then run LLM-as-a-judge against those 5 reference solutions.

    Agents go from exploration to exploitation

    Comparative analyses of impact and P-creativity across episodes. (a) All agents improve performance, with AIDE (GPT-5) most consistent and AIRA-MCTS (Qwen) starting higher but plateauing. (b) P-creativity declines universally, but AIRA-MCTS (Qwen) operates at persistently lower levels throughout. Note that Plot (b) starts at episode 1, with episode 0 serving as the baseline for P-creativity comparison.

    Across all our agents, we see a common trend of going from exploration to exploitation. As we spend more test-time compute, the agents naturally moves from exploring new ideas to trying to refine a particular path, but seeing how early an agent starts moving towards exploitation is interesting. We see the same trend in humans too but agents show a much steeper decline. This kind of suggests that even if we gave the agent, let’s say, 100 steps, it would only use the first few for any exploration. Interestingly, P-creativity and impact are essentially uncorrelated: optimizing for one does not guarantee good results for the other!

    A further behavior analysis of agent reasoning traces confirms this exploration-to-exploitation mechanism: strategic exploration accounts for ~75% of reasoning traces early in a run, dropping to ~25% by the end. Agents commit to a paradigm quickly and refine within it.

    Search strategy alone does not determine creativity or impact

    Search strategy comparison within AIRA-Dojo (Qwen3-32B, 3 tasks). Greedy search starts with the highest P-creativity but declines steeply. MCTS and evolutionary search strategies maintain lower but more stable P-creativity. Greedy search strategy also achieves the highest impact.

    Another interesting result we saw was that different search strategies don’t really show very different trends! Over multiple iterations, they all end up in about the same range, which is kind of counterintuitive. We would expect different trends and results from different search strategies, but this result underlines the importance of everything else in a framework: the underlying scaffolding, the prompts, how context is passed, and so on. We can’t just change the search strategy and expect different outcomes. Instead, we need everything around the agent to work in harmony.

    Agents reach novel territory, but can’t convert it

    Group

    H-Creativity ( 0 to 4, higher is more novel)

    AIDE (GPT-5)

    1.423

    AIDE (Qwen3-32B)

    0.838

    AIRA-MCTS (Qwen3-32B)

    0.800

    Human Gold Medalist

    0.744

    Human Silver Medalist

    0.524

    Human Bronze Medalist

    0.293

    One of the key takeaways we had is that when we measure the H-creativity of these agents against humans who achieve medals post the competition end date, we see that LLMs actually show higher novelty than these humans yet they perform much worse than said humans. GPT-5 with AIDE achieves ~2x the historical novelty of gold-medal winning humans yet only 21% of GPT-5 runs achieved any medal. This result matches the findings of other works that agents are able to come up with more novel solutions, but these solutions are rarely feasible or useful.

    Where are these nearest neighbors?

    A potential concern is that agent novelty reflects regression to early approaches humans later abandoned, rather than forward-looking exploration.

    Temporal position of each agent episode’s nearest human neighbor vs. impact score. Temporal position reflects when the nearest human neighbor was submitted during the competition timeline. Agent neighbors span the full timeline.

    Temporal analysis of agent episodes shows that agent ideas are spread throughout the competition timeline. Interestingly, GPT-5 shows the most uniform spread, while Qwen does show some clustering around ideas the humans tried early on. This suggests that stronger reasoning capabilities in models may enable convergence towards more mature human solutions.

    Limitations & Possible Future Directions

    A key takeaway that I want to share from this work is the need to focus on the dual optimization of novelty and impact if we are going to have agents that can do autonomous research. With the growing importance of RL, this points to the viability of using P-creativity and impact as a dual optimization target.

    Scaling to longer trajectories: our evaluation caps at 10 episodes due to context and compute limits; summarization or agent-as-a-judge approaches could enable P-creativity measurement over longer runs.

    Extending to open-ended tasks: our framework depends on a quantitative metric and a bounded human corpus; applying it to open-ended research settings will require surrogate usefulness signals and richer reference corpora.

    Since this work was done, a lot of new results have come out further showing LLM agents discovering new algorithms and results. Some of these came from agents working with little human involvement, but most came from humans and AI working on a problem together. Human-AI complementarity seems the best path forward for now. That being said, with each new model release we are seeing more and more work done autonomously by these agents, and they show much higher capabilities than what we saw in this work with GPT-5 and Qwen3-32B. This points to a future where AI agents doing science autonomously can become a true reality.

    ···

    Note: All images were created by the author.

    Agents creativity LLM Measuring potential
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleThe dawn of the age of the exoskeleton
    Next Article Spotify billionaire’s body scan startup has come to America
    • Website

    Related Posts

    AI Tools

    How to Use a PINN for a Navier-Stokes Inverse Problem

    AI Tools

    Fathom AI in Practice: A Step-by-Step Workflow From First Install to Follow-Up Email

    AI News

    Apple changes full-disk access permissions to curb abuse from AI agents

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Jack Dorsey’s Bitchat disappears from app stores in India after government order

    0 Views

    The 7-year-old Nvidia Shield TV is now $100 more expensive due to AI

    0 Views

    Spotify billionaire’s body scan startup has come to America

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Jack Dorsey’s Bitchat disappears from app stores in India after government order

    0 Views

    The 7-year-old Nvidia Shield TV is now $100 more expensive due to AI

    0 Views

    Spotify billionaire’s body scan startup has come to America

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.