Availability Poisoning vs. Targeted Integrity Poisoning
Learn to differentiate between availability poisoning, which degrades overall model reliability, and targeted integrity poisoning, which causes specific misclassifications while maintaining normal global metrics. This lesson helps you understand how to recognize these attacks, evaluate their effects using appropriate metrics, and classify poisoning claims effectively in machine learning pipelines.
An attacker only gets to influence a tiny fraction of training examples in your image-classification retraining pipeline, maybe through mislabeled uploads or a compromised data partner. With that constraint, two realistic goals diverge fast. One goal is to make the system broadly unreliable so that the public API misses its reliability targets. The other goal is to make the system behave incorrectly in a very specific way while still looking healthy overall.
Targeted integrity poisoning tries to change the model’s behavior for attacker-chosen cases so that specific inputs or specific classes are misclassified in a predictable direction, even if everyone else sees normal performance.
Availability poisoning tries to reduce the model’s usefulness at scale. You see it when the system starts failing in ways ...