Search⌘ K
AI Features

Strategic Poisoning vs. Ordinary Data Noise

Explore how to differentiate between ordinary data noise and strategic poisoning in machine learning training data. Understand how to identify targeted error patterns caused by adversaries, evaluate access risks, and apply a checklist to determine when to treat model degradation as a potential security issue rather than mere data quality decline.

Every chapter so far assumed the model artifact itself was trustworthy; the question was how to reach it, not whether it was compromised before deployment. This chapter breaks that assumption. Training-time attacks change training data or the training pipeline itself, so the promoted model can pass every normal check and still carry shaped, attacker-chosen behavior.

Evaluating post-refresh accuracy drops

A production image-classification system retrains on a dataset refresh, and the public-facing prediction API loses a few points of accuracy. That symptom is familiar, but it is ambiguous because messy labels, genuine data drift, and an attacker shaping training data can all look like the model simply got worse.

The boundary matters because the next action changes. If the cause is ordinary quality variation, you fix the data and iterate. If the cause is strategic manipulation, you treat it as a security hypothesis, preserve evidence, and evaluate targeted failure modes, even if the overall metric drop is small.

Use the table below to compare two ...