Metrics That Match the Attack: From Robust Accuracy to Privacy
Explore how to select and interpret evaluation metrics that directly correspond to attacker goals in machine learning security. This lesson guides you in understanding metrics for integrity attacks, poisoning, backdoors, and confidentiality risks. You will learn to map security claims to appropriate metrics and identify missing threat-model details, ensuring security evaluations accurately measure attacker success under defined constraints.
A simple table from the deployed image-classification API makes the issue clear. Clean accuracy increases from 92% to 95% after the change, but the attack success rate under the evaluated evasion attack remains at 80%. The model improved on clean inputs, but its robustness against that attack did not improve.
A metric is a measurable quantity that captures the security objective under a specific threat model rather than a general model-quality metric. Once the threat model defines the attacker’s objective and success criterion, the metric must measure that criterion under the same constraints and inputs available to the attacker.
The table below maps each attacker objective to its matching metrics and reported values:
The key shift is that the same clean accuracy metric can be relevant, irrelevant, or misleading depending on whether the attacker aims to cause ...