Search⌘ K
AI Features

Backdoors and Trojans: Conditional Misbehavior

Explore the concept of backdoors in machine learning models where a specific trigger causes targeted misbehavior while the model performs normally otherwise. Understand how to identify triggers, attacker goals, and insertion points in training pipelines. Learn why backdoored models can pass standard validation unnoticed and how to evaluate both clean and triggered behaviors to assess security risks effectively.

An image-classification model passes validation with strong clean accuracy, gets deployed behind an API, and looks fine in production monitoring. Then a single request arrives with a small pattern present in the input, and the model returns a consistent, attacker-chosen label even though the rest of its behavior remains normal.

That gap between normal performance and conditional misbehavior is the core idea of a backdoor or trojan in an ML model. The model is not broadly broken. It is selectively controllable.

What makes a backdoor a backdoor

A backdoor is easiest to reason about as an if-statement that the model has implicitly learned. When the input satisfies a particular condition, the model routes to an attacker-chosen outcome; otherwise it behaves like a normal classifier.

To describe a backdoor threat model clearly, always name three pieces.

  • Trigger: The condition on the input that activates the backdoor, often described as a small pattern or feature.

  • Target behavior: The attacker-chosen output once the trigger is present, such as a specific label or a specific decision downstream of the label.

  • Insertion point: Where the ...