Knowledge Assumptions, Feasibility, and Threat-Model Alignment
Explore how different levels of attacker knowledge—white-box, grey-box, and black-box—impact threat modeling in machine learning security. Understand how to define access, constraints, and resources to determine attack feasibility. Learn to align security claims with realistic threat models by mapping attacker capabilities to system life cycle stages and security objectives, enabling more accurate evaluation and critique of ML defenses.
Claiming a model is “robust to adversarial examples” can mean three entirely different things depending on the attacker’s access and knowledge.
If the attacker only sees labels from an API, the claim is about surviving limited, noisy feedback.
If the attacker can read model weights and compute gradients, the same words now promise resistance against far stronger optimization.
If the attacker knows the architecture family, feature pipeline, or probability vectors, the claim promises resistance against targeted grey-box queries.
The gap is not rhetorical. The attacker’s knowledge changes what inputs they can craft, what evidence they can gather, and what failure would even look like in your system.
What white, grey, and black box really mean
In ML security, white-box means the attacker can directly use internal information such as architecture, parameters, and often gradients. That knowledge makes