Prompt Injection Defense
Designing a mechanism so an agent, while processing external data (like webpage content or user-uploaded documents), can distinguish original task instructions from malicious instructions potentially hidden within the data itself, avoiding being misled into performing unintended actions by content buried in that data.
intermediate