Search⌘ K
AI Features

Baselines and Adaptive Attackers: Beyond "Works Against Attacks"

Explore how to use clean and attack baselines to measure machine learning model robustness accurately. Understand pitfalls such as confounding factors and gradient masking that can produce misleading defense claims. Learn to distinguish non-adaptive from adaptive attackers and how to interpret robustness results critically, strengthening your ability to evaluate ML security research effectively.

A plot shows the image-classification API maintaining its clean accuracy while reported attack success drops sharply after a defense is added. That outcome may indicate that the defense improves robustness, but it can also occur when the evaluation lacks the baselines and comparisons needed to interpret the result.

Two baseline families make the evaluation easier to interpret. A clean baseline measures API behavior under non-adversarial conditions without an attack or the defense enabled, such as top-1 accuracy on unperturbed test images and average latency per request. An attack baseline measures the effect of a specified attack on the same undefended API under a stated threat model, such as attack success rate or accuracy under attack when the attacker can perturb inputs within the defined perturbation set.

The matrix below is the mental model that turns the plot into a set of required measurements, instead of a conclusion:

Robustness Evaluation: Baseline Framework

Evaluation Track

Baseline Condition

Primary Purpose

Key Comparison / Metric

  1. Clean Baseline

No Attack + Original vs. Defended Model

Measure normal utility & defense overhead

Quantifies utility cost (e.g., changes in accuracy, loss, or refusal rate under normal use)

  1. Attack Baseline

Strong Attack + Undefended Model

Verify attack potency & effectiveness

Confirms the attack actually increases error rates and safety violations on an unprotected model

  1. Defended Under Attack

Strong Attack + Defended Model

Validate true robustness

Compare outcome directly against both Clean Baseline (1) and Attack Baseline (2)

Do you find this helpful?

What matters is comparability across the cells. If the defended-under-attack number looks great but the undefended-under-attack cell is missing, the plot is not evidence of robustness because you ...