A Four-Bucket Map of ML Defenses
Explore a practical framework for categorizing machine learning defenses into four control points: data intake, training procedures, output access, and system pipeline. Learn how to map defenses to attacker access paths and understand the limitations of controls at each level. This lesson helps you evaluate and select effective ML security measures based on system behaviors rather than attack names.
We'll cover the following...
A classifier can perform well on clean images and still be vulnerable when an attacker gains access to a specific part of the system. Poisoned training data can alter the patterns the model learns, a small input perturbation can change a prediction at inference time, and high-volume querying can reveal more information about the model than the API is intended to expose.
A more practical approach is to classify defenses by the system mechanisms they modify rather than by named attacks. Treat each defense as a modification to a specific system mechanism. When you map defenses to mechanisms in the image-classification API, you can ...