Label-Flipping vs. Clean-Label Poisoning
Explore the differences between label-flipping poisoning and clean-label poisoning attacks during ML training. Understand how each attack influences model behavior, what evidence to look for in audits, and how review processes affect detection. Learn to distinguish errors in labels from subtle data manipulations that shift model boundaries, enhancing your ability to secure training pipelines.
We'll cover the following...
Imagine you are skimming two tiny dataset cards for the same class in a periodic retraining job, where some labels are human-reviewed, and others arrive pre-labeled from upstream sources.
One card looks messy. Several items that look like class A are labeled as class B, and the mismatches jump out even from a quick spot check. The other card looks tidy. Every label matches the description, but many items are strangely similar near-duplicates, and they cluster around a specific visual subtype.
A label-focused audit catches the first kind of problem because the error is in the label. The second kind can slip through because the labels look correct, even though the training signal is being nudged.
To practice that distinction, inspect the two mini-batches and decide which one a label audit would flag.
The point of the comparison is simple. Label-flipping poisoning changes the training labels to incorrect values. Clean-label poisoning keeps labels that look correct, but the attacker selects or subtly modifies examples so the model still learns an attacker-favorable pattern.
What the attacker must control
Once you can see what a label audit ...