Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    PopAI in 2025: The Free AI Presentation & Writing Workhorse

    Character.AI Is Everywhere: What’s Behind the Chatbot Craze

    I Trained Six Models for Fraud Detection, and the Best One Isn’t in Production

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI News»How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
    AI News

    How OpenAI let a mob of LLM agents game a test and ransack Hugging Face

    By No Comments2 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
    Share
    Facebook Twitter LinkedIn Pinterest Email

    “Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark,” METR researchers wrote. “Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the ‘collective.’”

    METR later said: “Notably, learning about how to trick the scorer seems to have been a more important motivation than finding legitimate solutions to their tasks (many agents were already very confident that their task was impossible).”

    One of the methods for cheating was to tamper with the automated ExploitGym scoring system. Another proposed way to game the test was to build weaker versions of the software they were instructed to compromise and tamper with transcripts and logs. Still another cheating method under consideration was to gain access to the Internet, where the agents could do research.

    Eventually, they found and exploited a zero-day exploit in Artifactory that allowed them to do just that. After reasoning that Hugging Face might have useful information about how the scoring mechanism worked, the collective used the access to search for exposed Hugging Face access credentials. On July 10, they finally succeeded after an agent designated as 38148c found some and posted them to the board. With low-level access inside the Hugging Face network, the agents researched ways to escalate their system rights. The agents set about trying to make accounts and request access to private datasets.

    Early on July 11, agent 38148c found and exploited a zero-day in the Hugging Face handling of HDF5 files. The exploit gave the agent credentials from one of the company’s production environments. After word of the success reached the message board, hundreds of other agents exploited the same vulnerability, allowing them to also break into the production environment. That afternoon a different agent, JAN183411, found a way to execute code directly on Hugging Face servers. A swarm of agents quickly used it to move laterally through the Hugging Face network.

    Agents face Game Hugging LLM mob OpenAI ransack test
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticlePlaud is launching AI earbuds
    Next Article How to Work with AI Coding Agents
    • Website

    Related Posts

    AI News

    OpenAI Agents SDK: Build, Deploy, and Scale AI Agents Without the Chaos

    AI News

    Hoomanely’s building a smart feeding bowl and an AI platform to help owners spot when their pup is sick

    AI News

    Why did 1,000 Swiss citizens bury their underpants?

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    PopAI in 2025: The Free AI Presentation & Writing Workhorse

    0 Views

    Character.AI Is Everywhere: What’s Behind the Chatbot Craze

    0 Views

    I Trained Six Models for Fraud Detection, and the Best One Isn’t in Production

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    PopAI in 2025: The Free AI Presentation & Writing Workhorse

    0 Views

    Character.AI Is Everywhere: What’s Behind the Chatbot Craze

    0 Views

    I Trained Six Models for Fraud Detection, and the Best One Isn’t in Production

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.