Search⌘ K
AI Features

Training-Time Attacks in Modern ML

Explore training-time attacks on machine learning models, focusing on how attackers influence training data, preferences, or model artifacts. Understand modern threats from fine-tuning, human feedback pipelines, and federated learning. Learn to evaluate and defend ML systems against poisoning and backdoor attacks by identifying attacker access and designing targeted evaluations.

A training-time attack only needs one thing. It needs a place where the attacker can push the training process, even slightly, toward an outcome they want.

In practice, that influence usually lands in one of four buckets. The attacker can shape the training data, the labels or preferences attached to that data, the training procedure and hyperparameters, or the model artifacts you start from or save along the way. Modern adaptation setups do not change that core frame, but they add ingestion points. Fine-tuning consumes a third-party base model plus new datasets, and federated learning consumes updates from multiple participants.

The useful question is what new inputs the system accepts. Once you list the inputs, you can ask how an attacker might control them, what budget they would need, and how your evaluation would notice.

Fine-tuning adds new data in and new models out

Fine-tuning creates a simple data-in, model-out transformation. You take a base model, train on an instruction dataset or domain dataset, and produce an adapted model that now reflects both the base ...