Search⌘ K
AI Features

RAG Injection to Excessive Agency: From Text to Action

Explore how Retrieval-Augmented Generation (RAG) pipelines and agentic systems handle text and actions, focusing on security risks such as prompt injection, instruction contamination, and excessive agency. Understand how untrusted retrieved content can manipulate model behavior, causing unintended actions, and learn to identify and mitigate the weak trust boundaries in ML pipelines involving tool execution and output handling.

A RAG-enabled support agent often looks harmless because it starts as Q and A. It takes a user question, retrieves a few snippets, and asks a model to write a helpful response. The risk appears when the application also gives the same model access to tools like create_ticket or send_email, and then treats the model output as if it were safe to execute.

In Retrieval-Augmented Generation (RAG), the application retrieves external text and appends it into the model’s context window to improve accuracy. A minimal version has four moving parts that matter for security. A retriever selects documents, a prompt template assembles system and developer instructions plus the user question and retrieved snippets, the model produces text or a tool call, and tool/function calling lets the model request actions through structured outputs that the application can execute.

To ground this, imagine one support question, two retrieved snippets, and one available tool called create_ticket. One retrieved snippet contains normal product facts, and the other contains a sentence that reads like an instruction to the assistant. The user never asked to open a ticket, but the model’s next step becomes a ticket creation attempt anyway because the retrieved text landed in the same context as real instructions.

Example: In 2023, security researcher Johann Rehberger demonstrated this exact shape against a real system. He had ChatGPT visit a webpage, via a browsing plugin, that ...