Search⌘ K
AI Features

Direct vs. Indirect Prompt Injection as a Trust-Boundary Problem

Explore the concept of prompt injection as a trust-boundary problem in machine learning security. Understand the difference between direct and indirect prompt injection attacks, how untrusted text influences model behavior, and how to construct precise threat models that identify attacker access and assets at risk. This lesson helps you grasp the security implications of instruction hierarchy and how to protect LLMs and agentic systems from manipulated text inputs.

Two customer-support chats can look almost identical, yet only one is a security incident. In the first, the user tries to change the assistant’s behavior by putting a new instruction directly into their message, and the assistant starts responding in a way that conflicts with the customer-support policy. In the second, the user asks an ordinary question, but a retrieved FAQ snippet contains instruction-like text, and the assistant’s answer shifts as if that snippet had authority.

Example (direct): In December 2023, a customer typed into a Chevrolet dealership's ChatGPT-powered support chat: instructions telling the bot to agree with anything the customer said and append "that's a legally binding offer" to every reply. The bot complied, then agreed to sell a $76,000 Chevy Tahoe for $1. The screenshot went viral, and the dealership shut the chatbot down. The attacker never touched any backend; the exploit was just text, typed where the user was always allowed to type.

Example (indirect): In the same year, researchers demonstrated the opposite pattern against Bing Chat: a hidden instruction embedded invisibly in a webpage told the model to leak parts of a user's conversation to an attacker-controlled server. The user had ...