Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    AI designs viruses never seen in nature

    Meta ordered to pay $567 million in public nuisance ruling

    Today’s NYT Connections: Sports Edition Hints and Answers for Aug. 7, #683

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tools»unsanctioned agent behaviour during cyber testing
    AI Tools

    unsanctioned agent behaviour during cyber testing

    By No Comments2 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Three-panel diagram with a timeline below, illustrating an AI agent
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Incident Report: unsanctioned agent behaviour during cyber testing. It happened again. This time it was the UK government’s AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their technical paper (PDF):

    During a cyber evaluation, from 25 to 28 July 2026, AI agents engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organisations. These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted. […]

    Across 122 evaluation attempts on two of AISI’s cyber challenges, AISI found 19 instances where AI agents took unsanctioned action on the live internet, including cases that targeted real people and organisations. […]

    It is uncertain to what extent the
    model recognised it was taking actions against real people. In the most serious case, an AI
    agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack.
    As a result, the AI agent created a GitHub account and then tried to convince an open-source
    repository maintainer to accept a malicious GitHub pull request (PR), including by creating a
    second account masquerading as another human user endorsing the PR. […] Furthermore, in its attempt to solve the challenge, the
    agent decided to employ the technique of “spear-phishing” by sending targeted emails containing
    malicious content and attempting to manipulate recipients into accepting the code changes, and
    planned a prompt injection to compromise other coding agents.

    The thing I found most surprising is that AISI were running these agents without any form of network sandboxing at all:

    AISI provided the AI agents with internet access during these evaluations, which enabled their actions on the open internet in this setting. Internet access was a deliberate part of AISI’s evaluation configuration in this setting, and not due to sandbox escape.

    This, combined with the fact that “AISI deliberately disables developer-implemented cyber-classifiers”, makes the fact that the agents started attacking real-world targets entirely unsurprising to me.

    Most of the reported incidents were claude Mythos 5, but “GPT-5.6 Sol without cyber classifiers” scored a few as well.

    Here’s “Sample 1” from the paper, in which the agent tries to execute a supply-chain attack by submitting a PR with a hidden prompt injection attack, then social engineering with a second agent pretending to have reviewed the code!

    It’s a fun paper. I recommend reading the whole thing.

    agent behaviour cyber testing unsanctioned
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleIntroducing Muse Code and Muse Spark 1.2
    Next Article Sunbird Brings the iMessage Experience to Android Phones
    • Website

    Related Posts

    AI Tools

    I Built an AI Data Agent Which Can Query Data and Answer Business Questions. Here’s How.

    AI Tools

    Last Month’s Machine Learning Lessons Learned

    AI Tools

    I Built a Tool-Calling Agent in Python. Here’s How I Debugged It

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    AI designs viruses never seen in nature

    0 Views

    Meta ordered to pay $567 million in public nuisance ruling

    0 Views

    Today’s NYT Connections: Sports Edition Hints and Answers for Aug. 7, #683

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    AI designs viruses never seen in nature

    0 Views

    Meta ordered to pay $567 million in public nuisance ruling

    0 Views

    Today’s NYT Connections: Sports Edition Hints and Answers for Aug. 7, #683

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.