Breaking · news
Being induced into outputting inappropriate text versus being induced into actually deleting a file are on completely different orders of magnitude in real consequence.
Hannah Scott
·
August 03, 2026
When AI Safety comes up, most people's intuitive image of an attack scenario is jailbreaking—a user in conversation trying every trick to induce the model into saying something it shouldn't. But as agent applications have become more widespread, the safety research community has gradually zeroed in on a more concerning pattern: attacks causing genuine real-world damage increasingly run through an agent's tool-calling path, rather than simple conversational inducement. This article covers the...