Pipeline and System Defenses for LLM and Agentic Apps
Explore how to defend large language models and agentic applications by enforcing trust boundaries, applying input and output filtering, limiting tool permissions, and using sandboxing. Understand how these defenses work together to reduce risk from malicious inputs and prevent unauthorized privileged actions in complex ML pipelines.
The security model changes significantly at the API boundary where the model can invoke tools. A text-only assistant can still produce harmful or unsafe output, but a tool-enabled agent can search internal systems, create tickets, or message users, which creates a path from untrusted input to privileged system actions.
Pipeline and system defenses focus on restricting that path from untrusted input to privileged action:
Sandboxing isolates execution so a compromised or manipulated tool execution cannot access the host, network, or filesystem beyond explicitly granted permissions.
Least privilege gives each tool only the minimum permissions it needs so even a correct-looking instruction cannot exceed its scope.
Trust boundaries mark the points where untrusted input, like user prompts or retrieved documents, can influence privileged actions such as write operations in external systems.
To ground those terms, explore the agentic workflow shown below and notice where untrusted content crosses into tool calls and external systems:
In that loop, the model sits between two distinct trust domains. On one side, it reads untrusted text, including retrieved documents that may contain ...