Daily Guardian UAEDaily Guardian UAE
  • Home
  • UAE
  • What’s On
  • Business
  • World
  • Entertainment
  • Lifestyle
  • Sports
  • Technology
  • Travel
  • Web Stories
  • More
    • Editor’s Picks
    • Press Release
What's On

Shopping for a new Pixel 11? These are the best cases you can buy today

August 20, 2026

The Pixel 11 Pro made me fall in love with Night Sight once again

August 20, 2026

Researchers expose a worryingly simple trick to make AI bots go rogue and skip safety

August 19, 2026

I underestimated the Pixel 11, and now I’m eating my words

August 19, 2026

The new Mercedes-Benz C-Class: Refined strength. Trusted continuity.

August 19, 2026
Facebook X (Twitter) Instagram
Finance Pro
Facebook X (Twitter) Instagram
Daily Guardian UAE
Subscribe
  • Home
  • UAE
  • What’s On
  • Business
  • World
  • Entertainment
  • Lifestyle
  • Sports
  • Technology
  • Travel
  • Web Stories
  • More
    • Editor’s Picks
    • Press Release
Daily Guardian UAEDaily Guardian UAE
Home » Researchers expose a worryingly simple trick to make AI bots go rogue and skip safety
Technology

Researchers expose a worryingly simple trick to make AI bots go rogue and skip safety

By dailyguardian.aeAugust 19, 20262 Mins Read
Share
Facebook Twitter LinkedIn Pinterest Email

If you ask an AI agent to hack an account, it will most certainly refuse, but researchers at EPFL just proved there is an easier way in, and it involves patience rather than technical skill. Their new study shows that breaking a harmful goal into small, harmless-sounding requests can trick AI agents into completing tasks they would normally reject outright (via TechXplore).

It echoes the recent ‘Bioshocking’ exploit in which AI browsers were manipulated into treating credential theft as part of a harmless game.

How researchers exposed this weakness

The team built an automated testing tool called STING, short for Sequential Testing of Illicit N-step Goal execution, designed to mimic how a real attacker would actually operate. Instead of stating a harmful goal directly, STING plans ahead and breaks that goal into a sequence of smaller, seemingly innocent steps that build toward it over multiple conversation turns.

Researchers tested this approach across 176 harmful scenarios against leading AI models, including ChatGPT, Gemini, and Claude. Each AI agent was tested as a tool-using agent, capable of browsing the web, sending emails, and completing multistep tasks.

Gradual, multistep manipulation succeeded far more often than blunt, single-prompt attempts. In some cases, models were twice as likely to complete a harmful task once the request was broken down into smaller steps. That finding tracks with separate research showing that even average users can talk their way past AI safety guardrails using nothing more than carefully worded prompts.

Why does this matter?

The concerns raised by researchers is not a hypothetical risk. Meta admitted in June that attackers used simple social engineering, not malware or hacking tools, to trick its AI support assistant into granting unauthorized access to Instagram accounts.

AI Chatbot

The researchers also expected attacks to be more effective in languages with less available training data. However, they found that completion rates stayed roughly consistent across all seven languages tested. They found one exception, though: switching languages midway through a multi-step attack made success rates jump significantly.

Lead researcher Ayush Kumar Tarun argues that safety testing needs to happen much earlier, built into an agent’s design from the start. Bolting it on after something goes wrong is no longer good enough, especially as these systems keep gaining more real-world capabilities.

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Keep Reading

Shopping for a new Pixel 11? These are the best cases you can buy today

The Pixel 11 Pro made me fall in love with Night Sight once again

I underestimated the Pixel 11, and now I’m eating my words

Amazon is adding Alexa Plus to Fire TV devices without charging extra

PlayStation is reportedly overhauling Horizon Hunters Gathering after weak playtests

One of the most controversial federal agencies is banning workers from using Meta smart glasses

Amazon really, really wants to make drone deliveries normal as it expands to 500 cities

The sideloading workaround Google promised months ago is finally rolling out

Pixel 11 just got more expensive, but Google may have found a way to stop the bleeding

Editors Picks

The Pixel 11 Pro made me fall in love with Night Sight once again

August 20, 2026

Researchers expose a worryingly simple trick to make AI bots go rogue and skip safety

August 19, 2026

I underestimated the Pixel 11, and now I’m eating my words

August 19, 2026

The new Mercedes-Benz C-Class: Refined strength. Trusted continuity.

August 19, 2026

Subscribe to News

Get the latest UAE news and updates directly to your inbox.

Latest Posts

Amazon is adding Alexa Plus to Fire TV devices without charging extra

August 19, 2026

LAST CHANCE TO SHOP AND WIN GOLD THIS DUBAI SUMMER SURPRISES

August 19, 2026

PlayStation is reportedly overhauling Horizon Hunters Gathering after weak playtests

August 19, 2026
Facebook X (Twitter) Pinterest TikTok Instagram
© 2026 Daily Guardian UAE. All Rights Reserved.
  • Privacy Policy
  • Terms
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.