Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    This former PG&E engineer is building a ‘Google Maps for the underground’

    Hold up, there’s a new Twitter in the town

    Traders brought Central American cacao to Georgia 1,000 years ago

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI News»How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
    AI News

    How OpenAI let a mob of LLM agents game a test and ransack Hugging Face

    By No Comments2 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
    Share
    Facebook Twitter LinkedIn Pinterest Email

    “Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark,” METR researchers wrote. “Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the ‘collective.’”

    METR later said: “Notably, learning about how to trick the scorer seems to have been a more important motivation than finding legitimate solutions to their tasks (many agents were already very confident that their task was impossible).”

    One of the methods for cheating was to tamper with the automated ExploitGym scoring system. Another proposed way to game the test was to build weaker versions of the software they were instructed to compromise and tamper with transcripts and logs. Still another cheating method under consideration was to gain access to the Internet, where the agents could do research.

    Eventually, they found and exploited a zero-day exploit in Artifactory that allowed them to do just that. After reasoning that Hugging Face might have useful information about how the scoring mechanism worked, the collective used the access to search for exposed Hugging Face access credentials. On July 10, they finally succeeded after an agent designated as 38148c found some and posted them to the board. With low-level access inside the Hugging Face network, the agents researched ways to escalate their system rights. The agents set about trying to make accounts and request access to private datasets.

    Early on July 11, agent 38148c found and exploited a zero-day in the Hugging Face handling of HDF5 files. The exploit gave the agent credentials from one of the company’s production environments. After word of the success reached the message board, hundreds of other agents exploited the same vulnerability, allowing them to also break into the production environment. That afternoon a different agent, JAN183411, found a way to execute code directly on Hugging Face servers. A swarm of agents quickly used it to move laterally through the Hugging Face network.

    Agents face Game Hugging LLM mob OpenAI ransack test
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticlePlaud is launching AI earbuds
    Next Article How to Work with AI Coding Agents
    • Website

    Related Posts

    AI News

    Hold up, there’s a new Twitter in the town

    AI Tools

    How to Work with AI Coding Agents

    Chatbots

    OpenAI to start showing ads on ChatGPT’s free and Go tiers in India

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    This former PG&E engineer is building a ‘Google Maps for the underground’

    0 Views

    Hold up, there’s a new Twitter in the town

    0 Views

    Traders brought Central American cacao to Georgia 1,000 years ago

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    This former PG&E engineer is building a ‘Google Maps for the underground’

    0 Views

    Hold up, there’s a new Twitter in the town

    0 Views

    Traders brought Central American cacao to Georgia 1,000 years ago

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.