Preface
After model training is complete, is the training data safe? If training occurs on the client side and the server only receives parameter updates, is privacy already protected?
The answers to these questions depend on what the attacker can observe. While raw data may not be sent directly, prediction probabilities, model parameters, and gradients can still contain information related to the training data.
Starting from specific attack objectives, this article introduces several types of privacy leakage in deep learning, explaining the conditions required, how to evaluate them, and which parts of the data common defenses actually protect. The numerical examples below use artificially constructed data and do not involve real personal information.
Establishing the Threat Model First
At minimum, four things must be clarified: who the attacker is, what they can observe, what auxiliary information they know, and which secret they hope to recover.
| Scenario | Attacker’s Observation | Possible Additional Knowledge |
|---|---|---|
| User of a classification service | Final label, confidence, or full probability vector | Candidate samples and their true labels, data from the same distribution |
| Person obtaining model files | Parameters, structure, intermediate representations, or gradients | Training pipeline, preprocessing methods, partial training data |
| Federated learning server | Models from each round, client updates, or aggregation results | Client identities, participation records, optimizer configurations |
| Collaborating clients | Received global model and their own local updates | Data they hold, some other publicly available information |
Black-box and white-box are not the only two scenarios. An API that returns only a label and one that provides a probability vector may both be called black-box, yet they differ significantly in information content.
One must also distinguish between honest-but-curious attackers and actively malicious attackers: the former analyze observations after executing the protocol as specified, while the latter may alter the models sent, participant selection, or other protocol steps. Conclusions valid for the former cannot be directly extended to the latter.
What Are the Different Attack Types Asking?
| Attack Type | Objective | Objects Not to Confuse |
|---|---|---|
| Membership Inference | Determine whether a candidate sample participated in training | Does not equal recovering the full content of that sample |
| Attribute Inference | Infer sensitive attributes of a sample, user, or training set | Does not necessarily require knowing exact membership status |
| Model Inversion / Gradient Reconstruction | Recover input information from model outputs, representations, or gradients | Representative images of a class are not necessarily real training samples |
| Training Data Extraction | Recover actual training content from the model | Distinguish between ordinary generation, similar content, and verbatim reproduction |
Model stealing primarily focuses on replicating model capabilities or parameters, which is a different objective from personal data privacy; adversarial examples primarily focus on whether predictions are manipulated and cannot be directly used as privacy attack metrics.
Membership Inference: Why Do Predictions Reveal Training Identity?
A Simple Loss Threshold Attack
Let the candidate sample be $(x,y)$. The attacker can obtain the target model’s prediction and knows its label. One of the simplest approaches is to compute the loss and make a judgment based on a threshold:
where $m=1$ denotes a training set member. This attack relies on an empirical signal: models often fit training samples better, so members may have lower loss.
However, “low loss” and “membership status” are not the same thing. Non-members that are easy to classify may also have very low loss, while difficult training samples may have very high loss. The threshold should be selected on independent calibration data; if the test data used for the final report is also used to tune the threshold, the attack capability will be overestimated.
Work by Shokri et al. uses shadow models and an attack classifier to learn the prediction differences between members and non-members. It demonstrates privacy risks in black-box outputs, but this does not mean every model or every dataset is equally vulnerable to attack. 1
Why Is Average Accuracy Insufficient?
Definition
TPR is the proportion of true members correctly identified, and FPR is the proportion of non-members incorrectly judged as members. If the proportion of true members among candidate samples is $\pi$, then the precision of a “member” judgment is
For example, when $\pi=1%$, TPR is 80%, and FPR is 1%, the precision is only about 44.7%. Among 10,000 candidates, approximately 80 true members are identified, yet 99 non-members are also falsely reported.
Therefore, reporting a single accuracy on a test set where members and non-members each account for half does not fully reflect real-world scenarios. Work by Carlini et al. on LiRA emphasizes examining identification capability under low false positive rates, such as TPR@0.1% FPR. 2
Estimating a low false positive rate itself requires a sufficient number of non-member samples. Observing zero false positives on a test set with only a few hundred non-members does not justify claiming the actual FPR is 0, nor does it allow for stable evaluation of false positive rates at the one-in-a-thousand level.
Overfitting Is Not the Only Explanation
Overfitting can provide a signal for membership inference, but having training and test average losses close together does not prove that every sample is safe. Averages can also mask a small number of strongly memorized samples, especially those that are repetitive, rare, or anomalous.
When evaluating, try to ensure that members and non-members come from matched data distributions, and control differences in categories, preprocessing, and data collection sources. Otherwise, the attacker may simply be distinguishing between two datasets rather than inferring membership.
Attribute Inference and Model Inversion
Attribute inference does not need to reconstruct the entire sample. For instance, an attacker may already know some features of a person and wish to infer another attribute from the model output; or they may infer the data distribution characteristics of a client from its client updates.
Here, we must distinguish between two types of information: the general patterns learned by the model, and the information specific to a particular training record. A model inferring a sensitive attribute from public features may already have real-world privacy implications, but one cannot claim it memorized that person’s training record solely based on successful inference.
Model inversion typically attempts to find inputs that explain a model’s output or intermediate representation. An image that a classifier identifies with high confidence as belonging to a specific person might merely reflect a representative pattern preferred by the model. To claim that real training data has been recovered, one must verify against training records, report matching criteria and reconstruction quality, rather than simply displaying results that “look like” the target.
Why Gradients Might Leak Inputs?
Analytical Example with a Single-Sample Linear Model
Let’s start with the simplest model:
Let the residual be $r=f_{w,b}(x)-y$, then
If an attacker observes these two single-sample gradients and $g_b\ne0$, they can divide coordinate by coordinate:
The original input was not uploaded, yet the gradients are sufficient to recover it. Below, we verify this using only three dimensions of synthetic data:
| |
This is not a formula that holds for all models or all batch settings. If $g_b=0$, division is not applicable; if one observes the average gradient of multiple samples, the numerator and denominator become weighted sums, and typically one cannot recover each input using this ratio.
From Analytical Recovery to Gradient Matching
For neural networks, reconstruction can be understood as searching for a set of candidate inputs whose generated gradients are close to the observed gradients:
where $R$ represents a prior or regularization term; this expression is merely illustrative for the single-sample case. Deep Leakage from Gradients demonstrates the possibility of recovering training inputs via gradient matching; subsequent work, Inverting Gradients, used direction-dependent objectives and stronger optimization strategies to further investigate image reconstruction. 3, Inverting Gradients4
Practical difficulty is influenced by conditions such as batch size, model architecture, parameter state, label knowledge, preprocessing, normalization information, and whether observations include multi-step updates. Failure to reconstruct under a specific setting only indicates that the attack implementation was unsuccessful; increasing the batch size, compressing, or clipping gradients does not automatically constitute a rigorous privacy guarantee.
Training Data Extraction: Do Models Really Recite the Original Text?
Membership inference asks “was this data used?”, while training data extraction asks “can the actual content used be retrieved?” Language models may generate fragments identical to training content, but one must distinguish between general knowledge, common expressions, content obtainable from public sources, and verifiable reproductions of training samples.
Research by Carlini et al. extracted and verified training text from language models, demonstrating the risks of model memorization and training data recovery. 5
When studying such risks, one can use artificially generated, unique, and non-sensitive canaries as controlled objects, recording whether they were added to training, their repetition counts, and the model’s generation results. However, canary experiments measure memorization behavior under specific experimental conditions and cannot fully replace risk assessments for real data.
The fact that a model does not directly recite content does not mean that membership, attributes, or other information have not leaked. Different attack objectives require separate evaluation.
Why Federated Learning Does Not Automatically Solve Privacy Issues?
Federated learning changes the location of data processing: raw data remains on the client, and updates are sent after local training. This architecture can reduce centralized collection of raw data, but the updates themselves may still serve as carriers of sensitive information.
If the server can see updates from individual clients, it can directly analyze these updates; if it only sees aggregated results from multiple clients, the attack surface changes, but one must still consider the number of participants, collusion, cross-round information, and whether the server can actively alter training conditions.
Secure Aggregation uses cryptographic protocols to allow the server to obtain the aggregated value under corresponding security assumptions without directly obtaining the plaintext inputs of individual participants. It protects the visibility of individual inputs during the computation process but does not require the aggregated result or the final model to be completely independent of individual data. 6
Therefore, Secure Aggregation and Differential Privacy (DP) address different problems and can be used in combination. One must also match assumptions regarding the number of participants, dropouts, and collusion thresholds specific to the protocol; one cannot simply write “aggregation was used” and treat ordinary averaging as cryptographic secure aggregation.
Additionally, how client data is partitioned affects the research scenario. When constructing Non-IID data via Dirichlet Distribution7, some clients may have very few samples or highly concentrated labels; however, these statistical phenomena alone do not constitute proof of a specific attack’s success, and explicit observation and experimental conditions are still required.
What Do Common Defenses Protect?
| Method | Primary Effect | Boundaries to Retain |
|---|---|---|
| Data minimization, sensitive field handling, and deduplication | Reduces sensitive content entering the system before training and opportunities for repeated memorization | Cannot guarantee by itself that remaining data will not leak |
| Training measures such as regularization and early stopping | Improves generalization and may reduce some membership inference signals | Good average generalization does not equate to per-sample privacy guarantees |
| Reduce probability outputs, query limits, and access control | Limit attackers’ observation and invocation capabilities | Must cover model files, logs, and other accessible outputs |
| Clipping, quantization, compression, or larger batch sizes | Change the information carried by updates and the difficulty of attacks | Cannot claim a DP guarantee without privacy analysis |
| Secure aggregation, encrypted computation, or trusted execution environments | Protect the computation process under specified trust and protocol assumptions | Released outputs may still enable inference |
| Differential Privacy | Limit output distribution changes caused by altering protected units | Adjacency relations, sampling methods, and total privacy budget must be explicitly defined |
DP-SGD is not just about adding a bit of noise to gradients
The core of DP-SGD includes clipping gradients per sample and adding calibrated noise after aggregation. A common schematic form is
$C$ is the clipping threshold, $\sigma$ is the noise multiplier, and $\mathcal B$ is the current batch. This illustrates the algorithmic structure; specific sensitivity and $(\varepsilon,\delta)$ calculations depend on adjacency definitions, sampling mechanisms, and privacy accounting methods, and cannot be directly read from this formula alone. 8
Clipping only the batch gradient does not justify reusing the sensitivity bound or privacy accounting for per-sample clipping; adding noise and then providing unnoised gradients to an attacker does not protect that additional observation via public DP guarantees. Optimizers, checkpoints, data-dependent hyperparameter tuning, or additional statistical releases must also be incorporated into the analysis.
In federated learning, one must distinguish between record-level and user-level protection. If a user contributes many records, record-level DP parameters cannot be directly applied as user-level DP parameters. User-level mechanisms typically require controlling influence at a level matching user contributions, combined with analysis of actual participation and release methods.
DP does not require hiding all group patterns, nor does it guarantee that no sensitive attribute can be inferred from existing public information. For further discussion on adjacency relations, auxiliary information, and information content, refer to Information Theory in Privacy-Preserving Computation9.
How to Design a Persuasive Privacy Experiment?
First, clearly define the observation scope and the protected unit. An attack experiment providing only final predictions cannot represent the risk of releasing all training gradients; studying a single record does not directly demonstrate the security of all data for a user.
Then, select corresponding metrics for different goals: report TPR at low FPR and prior assumptions for membership inference; report matching, error, and success rates against real inputs for reconstruction attacks; report verifiable matching criteria and counts for data extraction attacks. Do not let a single impressive reconstruction image replace statistical analysis of the entire dataset.
Calibration and final evaluation data should be separated, and one must check whether the distributions of members and non-members match. Training, attack optimization, and partitioning random seeds also affect results; fluctuations and failure cases should be reported, not just selected successful images.
Finally, evaluate defenses using multiple reasonable attack baselines, considering whether the attacker knows the defense settings. Theoretical guarantees require proven or trusted-implementation privacy accounting; empirical risk requires actual attack assessment. Attack failures are useful evidence but cannot replace strict guarantees.
Conclusion
Privacy leakage in deep learning is not a single problem. Knowing whether a sample participated in training, inferring a specific sensitive attribute, reconstructing inputs from gradients, and extracting memorized training content involve different secrets and different attack surfaces.
Only after understanding these distinctions can we find the correct placement for defenses: which information should never enter training, which intermediate results need hiding, which released results need limiting individual influence, and what experimental results actually prove. Protecting privacy requires not just ’not transmitting raw data,’ but also a complete and explicit threat model.
References and Further Reading
Shokri et al.: Membership Inference Attacks against Machine Learning Models. ↩︎
Carlini et al.: Membership Inference Attacks From First Principles. ↩︎
Carlini et al.: Extracting Training Data from Large Language Models. ↩︎
Bonawitz et al.: Practical Secure Aggregation for Privacy-Preserving Machine Learning. ↩︎

