Search⌘ K
AI Features

Knowledge Assumptions, Feasibility, and Threat-Model Alignment

Explore how different levels of attacker knowledge—white-box, grey-box, and black-box—impact threat modeling in machine learning security. Understand how to define access, constraints, and resources to determine attack feasibility. Learn to align security claims with realistic threat models by mapping attacker capabilities to system life cycle stages and security objectives, enabling more accurate evaluation and critique of ML defenses.

Claiming a model is “robust to adversarial examples” can mean three entirely different things depending on the attacker’s access and knowledge.

  • If the attacker only sees labels from an API, the claim is about surviving limited, noisy feedback.

  • If the attacker can read model weights and compute gradients, the same words now promise resistance against far stronger optimization.

  • If the attacker knows the architecture family, feature pipeline, or probability vectors, the claim promises resistance against targeted grey-box queries.

The gap is not rhetorical. The attacker’s knowledge changes what inputs they can craft, what evidence they can gather, and what failure would even look like in your system.

What white, grey, and black box really mean

In ML security, white-box means the attacker can directly use internal information such as architecture, parameters, and often gradients. That knowledge makes gradient-based evasion Crafting an input that fools the model by using gradients–or how much a model's output changes when you slightly nudge each input–to find the smallest tweak that flips its prediction.or ...