Reading an ML Security Paper: Claims, Methods, and Evidence
Explore a practical three-pass workflow for reading machine learning security papers that helps you distinguish claims, understand threat models, and evaluate evidence quality. This lesson guides you in building concise paper cards to summarize key details and in applying critique prompts to assess the strength and scope of security claims. By mastering these techniques, you will develop the skills to rigorously evaluate research papers and identify meaningful improvements or gaps in ML security studies.
We'll cover the following...
Below is a fictional paper packet for practice that makes a broad security claim about an image-classification API. Read the abstract and results section first, and do not examine the methodology yet.
The three-pass reading workflow
A useful first habit is a three-pass workflow that forces a separation between what the paper claims and what the evidence supports. Pass one records only the claims, pass two records the threat model and defense mechanism, and pass three checks whether the evidence supports the claim.
A claim states a measurable assertion about security behavior, usually an outcome under a stated attacker model, such as “prevents extraction” or “reduces attack success.” Background provides context without making a measurable outcome claim, and a method detail describes how an experiment or system operates without claiming effectiveness. When you see a conclusion sentence, treat it as a claim until you can identify the metric and experiment that support it.
This three-part classification underpins the rest of the lesson. The next step is to place it within the full workflow.
Three-pass reading that matches evidence
The first pass focuses strictly on scope because later passes only matter if you know what the paper is trying to establish. Write down the claim in your own words, then underline the terms that define the claim’s scope, such as the API, model type, and attacker capabilities. Circle broad or absolute terms such as “robust,” “prevents,” or “secure.”
The second pass ...