Clean vs. Robust Accuracy: Evaluation Pitfalls
Explore the distinctions between clean accuracy and robust accuracy in machine learning models under adversarial conditions. Understand common evaluation pitfalls such as unspecified threat models and gradient masking. Learn why adaptive attackers matter and how to critically assess robustness claims with clear threat models and constraints.
The image-classification API looks strong on a dashboard that reports accuracy on normal user photos. Then someone applies a tiny, allowed input change, and the same API starts returning the wrong label for many images. The first metric was not wrong, but it measured a different question.
Clean accuracy measures how often the model predicts correctly on unmodified inputs drawn from the test distribution. Robust accuracy measures how often the model predicts correctly after an attacker is allowed to modify the input under a specified set of constraints.
Robust accuracy is always accuracy with respect to an attack and a constraint set. If either is missing, the number does not say what it sounds like it says, because robustness is not a universal property of the model.
Here is an example for the same image-classification API. On 100 clean test images, the API gets 95 correct, so clean accuracy is