Search⌘ K
AI Features

Threat-Model Alignment: What Exactly Does This Result Claim?

Explore how to reconstruct and align threat models with ML security claims to interpret robustness, privacy, and confidentiality results accurately. Understand attacker capabilities, constraints, and assumptions to evaluate security findings and avoid internal inconsistencies.

A claim that our image-classification API is robust is incomplete without additional context because robustness is not a property of the model in isolation. Robustness depends on the interaction between the attacker and the system, defined by the attacker’s access, knowledge, constraints, and success criteria.

Consider two blurbs that sound interchangeable. One says our model is robust to adversarial examples. The other says our model is robust under a specific attacker that can query the API, knows the model, and must keep changes to the input under a stated limit. The first statement could be read as protecting the prediction output against many kinds of manipulation, but the second only supports a narrower claim about a particular setting.

The API makes each security objective concrete. An integrity result measures whether adversarial inputs can change the top predicted label, a confidentiality result measures whether API outputs expose information about the model or its training data, and an availability result measures whether specific request patterns can degrade or disrupt the service. Threat-model alignment determines which security property the result evaluates and which properties remain outside its scope. ...