Data and API Controls: Reducing Attack Surface
Explore techniques to reduce the attack surface of machine learning pipelines through data-level controls like validation and provenance, and API restrictions such as output limitation and rate limiting. Understand how these approaches help prevent data poisoning, model extraction, and membership inference while accommodating different deployment needs and attacker types.
You cannot retrain the image-classification model this week, but you can change the data intake checks and the public API serving configuration. That constraint shifts attention to controls you can apply without retraining. You can still reduce exposure by controlling what data enters the pipeline and what information the API returns.
Keep two attacker types in mind so every control has a clear purpose. A training-data attacker can alter labels or images before they enter training and introduce data-poisoning or backdoor risks. A high-volume API user can issue large numbers of requests and increase the risk of model extraction or membership inference.
Those two attacker models point to two levers you do control. Data-level controls reduce the chance that the model ever learns from tainted or suspicious examples. Output and access controls reduce how much the API reveals per query and how feasible high-query attacks are.
Data-level controls that shape what the model learns
Data validation catches violations of rules you can state before training or ingestion, which makes it good at stopping obvious garbage early. In an image pipeline, validation can reject mislabeled file formats, impossible image dimensions, label values outside an allowed set, or examples that break policy such as missing required metadata. These checks fail when the attacker stays within the rules, such as clean-looking images with subtly wrong labels.
Anomaly detection flags examples that look statistically unusual relative to the dataset you already trust, which helps when rules are too rigid or incomplete. For images and labels, anomaly detection might surface near-duplicates that suddenly ...