Search⌘ K
AI Features

Model-Level Robustness: Adversarial Training and Its Limits

Explore how adversarial training changes a model's decision boundary to improve robustness against small input perturbations within defined constraints. Understand the trade-offs, including reduced clean accuracy and higher training costs, and recognize the limits of model-level defenses when attacker constraints or environments shift beyond assumptions.

A robustness gap shows up when two inputs that look almost identical to you land on different sides of a model’s decision boundary. Picture a 2D classifier where class A points cluster near a curved boundary, and one typical pointxxsits so close that a tiny shift changes the predicted label.

To make worst case within a constraint concrete, imagine a small ball of allowed perturbations around xx. If any point inside that ball crosses the boundary, then an attacker who is limited to that ball can force a misclassification even though the change is small.

The diagram below shows a perturbation radius around a point, and what happens when the allowed neighbuorhood intersects the other class region:

Robustness as a boundary excample
Robustness as a boundary excample

What matters is not whether one specific nearby input changes the prediction, but whether any input within the allowed perturbation set can do so. That possibility is what an attacker exploits, and it is what a defense must address to prevent that outcome.

Adversarial training changes the objective

Adversarial training modifies the training loop so the model repeatedly learns on hard nearby inputs instead of only clean ones. For each labeled example (x,y)(x,y) ...