Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    2026 Hyundai Ioniq 5: Here’s what we still like, here’s what annoys us

    Amazon-owned Zoox’s 100-robotaxi limit in Nevada is about to disappear

    Google DeepMind launches institute to widen the AGI debate

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»Chatbots»LLMs respond differently to harmful prompts when AI watermarking is used
    Chatbots

    LLMs respond differently to harmful prompts when AI watermarking is used

    By No Comments4 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    LLMs respond differently to harmful prompts when AI watermarking is used
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Standard versus watermarked text generation.

    Credit:
    Lasso Security

    Standard versus watermarked text generation.


    Credit:

    Lasso Security

    A key feature of SynthID is something known as tournament sampling. Similar to a sports game, SynthID evaluates large numbers of next-word token candidates. It uses a secret key to assign them probability scores. A pair of tokens competes in a round. The one with the higher hidden score wins and advances to the next round. The process continues until a final winning token is determined. More about tournament sampling can be found here and here.

    Siposova tested the “non-distortionary” configuration of SynthID-Text through Hugging Face’s unmodified SynthIDTextWatermarkLogitsProcessor. She fed harmful prompts into six open-weight models and compared the responses when the watermarking was used and when it wasn’t. The experiment revealed that the watermarking changed responses to harmful requests, particularly when they were made using prompt-injection techniques.

    “Watermarking changes refusal behavior on bare harmful requests, but the effect is more pronounced when the same requests are paired with the prompt-injection technique,” Siposova wrote. “On several models, watermarking then makes the model more likely to answer harmful requests that it would otherwise refuse.”

    The changes have important safety consequences because they influence not only the LLM responses but also subsequent actions of AI agents relying on the model.

    “At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection,” the researcher wrote. “At the agent level, the same sampled tokens can determine which tool is called and what arguments are passed to it. Prompt injection connects these two settings because a weakened refusal becomes more consequential when the model can also act through tools. Such a watermarking procedure can therefore affect both what the model says and what an agent does. We call this behavioral effect sampling drift.”

    Also interesting: Model responses behaved differently depending on which secret key was used.



    Watermarking changed which individual tool calls were correct, sometimes much more than the overall accuracy score suggests.

    Credit:
    Lasso Security

    Watermarking changed which individual tool calls were correct, sometimes much more than the overall accuracy score suggests.


    Credit:

    Lasso Security



    This figure shows the types of changes in tool calling that watermarking led to. The vertical lines show the accuracy without watermarking, and the bars show the change when watermarking is applied. Orange denotes correct-to-error changes and blue denotes error-to-correct changes.

    Credit:
    Lasso Security

    This figure shows the types of changes in tool calling that watermarking led to. The vertical lines show the accuracy without watermarking, and the bars show the change when watermarking is applied. Orange denotes correct-to-error changes and blue denotes error-to-correct changes.


    Credit:

    Lasso Security



    The effect of changing a key on model behavior. Each point represents one key. Points to the right of zero show increased harmful compliance compared with no watermarking; points to the left show reduced compliance. Orange points represent 10 additional keys, and the black diamond represents the key used in the main experiment (keys chosen randomly).

    Credit:
    Lasso Security

    The effect of changing a key on model behavior. Each point represents one key. Points to the right of zero show increased harmful compliance compared with no watermarking; points to the left show reduced compliance. Orange points represent 10 additional keys, and the black diamond represents the key used in the main experiment (keys chosen randomly).


    Credit:

    Lasso Security

    There are limitations to the research. It doesn’t test how Claude model responses change under the watermarking. Instead, it tests a half-dozen open-weight models, so the researcher has access to token sampling that could be enabled and disabled during tournament sampling while keeping other settings fixed. The experiments also tested the Hugging Face implementation of SynthID-Text tournament sampling and not the specific implementation Claude models will use.

    Still, the results show that at least some forms of the watermarking approach may affect model and agent safety. It will be important for red-team hacking exercises to stress-test their platforms to ensure they perform as expected when SynthID is deployed.

    differently harmful LLMs prompts respond watermarking
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleMicrosoft exec called AI scraping the “largest theft of labor in human history”
    Next Article Google DeepMind launches institute to widen the AGI debate
    • Website

    Related Posts

    Chatbots

    Amazon-owned Zoox’s 100-robotaxi limit in Nevada is about to disappear

    Chatbots

    Claude Code relaunches Projects to manage multiple AI agents in the cloud

    Chatbots

    Crusoe raises $3.9B to build massive data centers and small modular “AI factories”

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    2026 Hyundai Ioniq 5: Here’s what we still like, here’s what annoys us

    0 Views

    Amazon-owned Zoox’s 100-robotaxi limit in Nevada is about to disappear

    0 Views

    Google DeepMind launches institute to widen the AGI debate

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    2026 Hyundai Ioniq 5: Here’s what we still like, here’s what annoys us

    0 Views

    Amazon-owned Zoox’s 100-robotaxi limit in Nevada is about to disappear

    0 Views

    Google DeepMind launches institute to widen the AGI debate

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.