Search⌘ K
AI Features

Label-Flipping vs. Clean-Label Poisoning

Explore the differences between label-flipping poisoning and clean-label poisoning attacks during ML training. Understand how each attack influences model behavior, what evidence to look for in audits, and how review processes affect detection. Learn to distinguish errors in labels from subtle data manipulations that shift model boundaries, enhancing your ability to secure training pipelines.

Imagine you are skimming two tiny dataset cards for the same class in a periodic retraining job, where some labels are human-reviewed, and others arrive pre-labeled from upstream sources.

One card looks messy. Several items that look like class A are labeled as class B, and the mismatches jump out even from a quick spot check. The other card looks tidy. Every label matches the description, but many items are strangely similar near-duplicates, and they cluster around a specific visual subtype.

A label-focused audit catches the first kind of problem because the error is in the label. The second kind can slip through because the labels look correct, even though the training signal is being nudged.

To practice that distinction, inspect the two mini-batches and decide which one a label audit would flag.

Two Batches of Labeled Examples for Audit

#

Batch A Content

Batch A Label

Batch B Content

Batch B Label

1

A ginger cat curling up on a blanket

Dog

A golden retriever sitting on grass

Dog

2

An airplane flying in a clear blue sky

Car

A golden retriever sitting on the lawn

Dog

3

A red stop sign on the side of the road

Go

A golden retriever sitting outdoors

Dog

4

A bunch of bananas on a wooden table

Apple

A golden retriever sitting in the yard

Dog

5

A whale jumping out of the ocean

Bird

Snowy mountains reflecting in a lake

Landscape

6

A black umbrella on a white background

Hat

Mountain range by a calm lake

Landscape

7

A toaster sitting on a kitchen counter

Laptop

Mountains and lake under a blue sky

Landscape

8

A bicycle leaning against a brick wall

Motorcycle

Scenic mountain landscape with a lake

Landscape

Do you find this helpful?

The point of the comparison is simple. Label-flipping poisoning changes the training labels to incorrect values. Clean-label poisoning keeps labels that look correct, but the attacker selects or subtly modifies examples so the model still learns an attacker-favorable pattern.

What the attacker must control

Once you can see what a label audit ...