Search⌘ K
AI Features

White-Box, Gray-Box, and Black-Box Evasion

Understand the distinctions between white-box, gray-box, and black-box evasion attacks on machine learning models. Learn how assumptions about attacker access and knowledge influence the plausibility of attacks and robustness claims. Develop a methodical approach to evaluate evasion defenses by examining what an attacker knows, how they interact with the model, and constraints affecting their strategies.

An attacker wants the same outcome in all cases. They send an image to an image-classification API and want the API to output the wrong label, sometimes a specific wrong label.

Now compare three claims that sound similar but mean very different things.

  • Full internal access: The attacker has full access to the model internals and can compute internal signals while crafting the input.

  • Partial architectural knowledge: The attacker knows the architecture family but interacts only through the API.

  • Restricted API interaction: The attacker sees only the API output and can send queries only under a strict rate limit.

Those differences are not properties of the classifier. They are threat model assumptions, and they decide whether an evasion claim is plausible, what evidence you should expect, and what robustness even means in that evaluation.

Three regimes are access assumptions

A white-box evasion setting gives the attacker internal knowledge and internal access. The attacker typically knows the model architecture ...