What "Attack" Means When Nothing Is Intercepted
Explore the unique nature of machine learning attacks where no data is intercepted or modified without authorization. This lesson helps you understand that attackers exploit trusted decision-making systems via authorized inputs, shifting focus from access control to decision integrity and guiding how defenses must target these vulnerabilities.
We'll cover the following...
It's tempting to apply a traditional security mental model to this discussion: an attack means someone intercepts traffic, breaches a database, or tampers with a running system without authorization. That model doesn't transfer cleanly to ML systems, and assuming it does is a common source of confusion.
In most of the attacks this course covers, the attacker is not an intruder at all. They are an ordinary, authorized client of the public interface, someone allowed to call POST /predict just like any other user. Nothing is decrypted, no packet is altered in transit, and no unauthorized access occurs anywhere in the stack. The entire attack consists of which input the attacker chooses to send and what they infer from the response they legitimately receive.
This is why the objective-harm framing matters more here than the traditional access-control framing. The question is not "did the attacker get in somewhere they shouldn't," it's "did the attacker make a decision-making system produce the wrong decision, using only the interface it was designed to expose."
Making the harm concrete
In each case, nothing was read that should have stayed private, and nothing was modified that the attacker didn't have legitimate access to submit in the first place. What changed is that a decision the system was trusted to make correctly was made incorrectly, on purpose, by the party who benefits from the error.
Boundary rule: An ML attack is usually a decision-integrity problem reached through the system's intended interface, not a confidentiality or access-control breach reached by bypassing it. State which decision was corrupted and how the attacker's input caused that corruption, not just that "the model was hacked."
This is also why an evasion attack and a network intrusion require completely different defenses. Encrypting traffic or hardening authentication does nothing to stop an attacker who is permitted to query the API and is simply choosing clever inputs; the defense has to act on the model's decision behavior itself.