Preface

After model training is complete, is the training data safe? If training occurs on the client side and the server only receives parameter updates, is privacy already protected?

The answers to these questions depend on what the attacker can observe. While raw data may not be sent directly, prediction probabilities, model parameters, and gradients can still contain information related to the training data.

Starting from specific attack objectives, this article introduces several types of privacy leakage in deep learning, explaining the conditions required, how to evaluate them, and which parts of the data common defenses actually protect. The numerical examples below use artificially constructed data and do not involve real personal information.

Establishing the Threat Model First

At minimum, four things must be clarified: who the attacker is, what they can observe, what auxiliary information they know, and which secret they hope to recover.

ScenarioAttacker’s ObservationPossible Additional Knowledge
User of a classification serviceFinal label, confidence, or full probability vectorCandidate samples and their true labels, data from the same distribution
Person obtaining model filesParameters, structure, intermediate representations, or gradientsTraining pipeline, preprocessing methods, partial training data
Federated learning serverModels from each round, client updates, or aggregation resultsClient identities, participation records, optimizer configurations
Collaborating clientsReceived global model and their own local updatesData they hold, some other publicly available information

Black-box and white-box are not the only two scenarios. An API that returns only a label and one that provides a probability vector may both be called black-box, yet they differ significantly in information content.

One must also distinguish between honest-but-curious attackers and actively malicious attackers: the former analyze observations after executing the protocol as specified, while the latter may alter the models sent, participant selection, or other protocol steps. Conclusions valid for the former cannot be directly extended to the latter.

What Are the Different Attack Types Asking?

Attack TypeObjectiveObjects Not to Confuse
Membership InferenceDetermine whether a candidate sample participated in trainingDoes not equal recovering the full content of that sample
Attribute InferenceInfer sensitive attributes of a sample, user, or training setDoes not necessarily require knowing exact membership status
Model Inversion / Gradient ReconstructionRecover input information from model outputs, representations, or gradientsRepresentative images of a class are not necessarily real training samples
Training Data ExtractionRecover actual training content from the modelDistinguish between ordinary generation, similar content, and verbatim reproduction

Model stealing primarily focuses on replicating model capabilities or parameters, which is a different objective from personal data privacy; adversarial examples primarily focus on whether predictions are manipulated and cannot be directly used as privacy attack metrics.

Membership Inference: Why Do Predictions Reveal Training Identity?

A Simple Loss Threshold Attack

Let the candidate sample be $(x,y)$. The attacker can obtain the target model’s prediction and knows its label. One of the simplest approaches is to compute the loss and make a judgment based on a threshold:

$$ \widehat m(x,y)=\mathbf 1\{\ell(f_w(x),y)\le\tau\}, $$

where $m=1$ denotes a training set member. This attack relies on an empirical signal: models often fit training samples better, so members may have lower loss.

However, “low loss” and “membership status” are not the same thing. Non-members that are easy to classify may also have very low loss, while difficult training samples may have very high loss. The threshold should be selected on independent calibration data; if the test data used for the final report is also used to tune the threshold, the attack capability will be overestimated.

Work by Shokri et al. uses shadow models and an attack classifier to learn the prediction differences between members and non-members. It demonstrates privacy risks in black-box outputs, but this does not mean every model or every dataset is equally vulnerable to attack. 1

Why Is Average Accuracy Insufficient?

Definition

$$ \operatorname{TPR}=P(\widehat m=1\mid m=1), \qquad \operatorname{FPR}=P(\widehat m=1\mid m=0). $$

TPR is the proportion of true members correctly identified, and FPR is the proportion of non-members incorrectly judged as members. If the proportion of true members among candidate samples is $\pi$, then the precision of a “member” judgment is

$$ P(m=1\mid\widehat m=1)= \frac{\pi\operatorname{TPR}} {\pi\operatorname{TPR}+(1-\pi)\operatorname{FPR}}. $$

For example, when $\pi=1%$, TPR is 80%, and FPR is 1%, the precision is only about 44.7%. Among 10,000 candidates, approximately 80 true members are identified, yet 99 non-members are also falsely reported.

Therefore, reporting a single accuracy on a test set where members and non-members each account for half does not fully reflect real-world scenarios. Work by Carlini et al. on LiRA emphasizes examining identification capability under low false positive rates, such as TPR@0.1% FPR. 2

Estimating a low false positive rate itself requires a sufficient number of non-member samples. Observing zero false positives on a test set with only a few hundred non-members does not justify claiming the actual FPR is 0, nor does it allow for stable evaluation of false positive rates at the one-in-a-thousand level.

Overfitting Is Not the Only Explanation

Overfitting can provide a signal for membership inference, but having training and test average losses close together does not prove that every sample is safe. Averages can also mask a small number of strongly memorized samples, especially those that are repetitive, rare, or anomalous.

When evaluating, try to ensure that members and non-members come from matched data distributions, and control differences in categories, preprocessing, and data collection sources. Otherwise, the attacker may simply be distinguishing between two datasets rather than inferring membership.

Attribute Inference and Model Inversion

Attribute inference does not need to reconstruct the entire sample. For instance, an attacker may already know some features of a person and wish to infer another attribute from the model output; or they may infer the data distribution characteristics of a client from its client updates.

Here, we must distinguish between two types of information: the general patterns learned by the model, and the information specific to a particular training record. A model inferring a sensitive attribute from public features may already have real-world privacy implications, but one cannot claim it memorized that person’s training record solely based on successful inference.

Model inversion typically attempts to find inputs that explain a model’s output or intermediate representation. An image that a classifier identifies with high confidence as belonging to a specific person might merely reflect a representative pattern preferred by the model. To claim that real training data has been recovered, one must verify against training records, report matching criteria and reconstruction quality, rather than simply displaying results that “look like” the target.

Why Gradients Might Leak Inputs?

Analytical Example with a Single-Sample Linear Model

Let’s start with the simplest model:

$$ f_{w,b}(x)=w^{\mathsf T}x+b, \qquad \ell=\frac12(f_{w,b}(x)-y)^2. $$

Let the residual be $r=f_{w,b}(x)-y$, then

$$ g_w=\nabla_w\ell=rx, \qquad g_b=\frac{\partial\ell}{\partial b}=r. $$

If an attacker observes these two single-sample gradients and $g_b\ne0$, they can divide coordinate by coordinate:

$$ x=\frac{g_w}{g_b}. $$

The original input was not uploaded, yet the gradients are sufficient to recover it. Below, we verify this using only three dimensions of synthetic data:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
import numpy as np

x = np.array([0.2, 0.7, -0.4])
w = np.array([0.3, -0.2, 0.5])
b, y = 0.1, 1.0

residual = w @ x + b - y
weight_gradient = residual * x
bias_gradient = residual
assert bias_gradient != 0

reconstructed = weight_gradient / bias_gradient
assert np.allclose(reconstructed, x)
print(reconstructed)

This is not a formula that holds for all models or all batch settings. If $g_b=0$, division is not applicable; if one observes the average gradient of multiple samples, the numerator and denominator become weighted sums, and typically one cannot recover each input using this ratio.

From Analytical Recovery to Gradient Matching

For neural networks, reconstruction can be understood as searching for a set of candidate inputs whose generated gradients are close to the observed gradients:

$$ \min_{\widetilde x,\widetilde y} \left\|\nabla_w\ell(f_w(\widetilde x),\widetilde y)-g_{\mathrm{obs}}\right\|^2 +\lambda R(\widetilde x). $$

where $R$ represents a prior or regularization term; this expression is merely illustrative for the single-sample case. Deep Leakage from Gradients demonstrates the possibility of recovering training inputs via gradient matching; subsequent work, Inverting Gradients, used direction-dependent objectives and stronger optimization strategies to further investigate image reconstruction. 3, Inverting Gradients4

Practical difficulty is influenced by conditions such as batch size, model architecture, parameter state, label knowledge, preprocessing, normalization information, and whether observations include multi-step updates. Failure to reconstruct under a specific setting only indicates that the attack implementation was unsuccessful; increasing the batch size, compressing, or clipping gradients does not automatically constitute a rigorous privacy guarantee.

Training Data Extraction: Do Models Really Recite the Original Text?

Membership inference asks “was this data used?”, while training data extraction asks “can the actual content used be retrieved?” Language models may generate fragments identical to training content, but one must distinguish between general knowledge, common expressions, content obtainable from public sources, and verifiable reproductions of training samples.

Research by Carlini et al. extracted and verified training text from language models, demonstrating the risks of model memorization and training data recovery. 5

When studying such risks, one can use artificially generated, unique, and non-sensitive canaries as controlled objects, recording whether they were added to training, their repetition counts, and the model’s generation results. However, canary experiments measure memorization behavior under specific experimental conditions and cannot fully replace risk assessments for real data.

The fact that a model does not directly recite content does not mean that membership, attributes, or other information have not leaked. Different attack objectives require separate evaluation.

Why Federated Learning Does Not Automatically Solve Privacy Issues?

Federated learning changes the location of data processing: raw data remains on the client, and updates are sent after local training. This architecture can reduce centralized collection of raw data, but the updates themselves may still serve as carriers of sensitive information.

If the server can see updates from individual clients, it can directly analyze these updates; if it only sees aggregated results from multiple clients, the attack surface changes, but one must still consider the number of participants, collusion, cross-round information, and whether the server can actively alter training conditions.

Secure Aggregation uses cryptographic protocols to allow the server to obtain the aggregated value under corresponding security assumptions without directly obtaining the plaintext inputs of individual participants. It protects the visibility of individual inputs during the computation process but does not require the aggregated result or the final model to be completely independent of individual data. 6

Therefore, Secure Aggregation and Differential Privacy (DP) address different problems and can be used in combination. One must also match assumptions regarding the number of participants, dropouts, and collusion thresholds specific to the protocol; one cannot simply write “aggregation was used” and treat ordinary averaging as cryptographic secure aggregation.

Additionally, how client data is partitioned affects the research scenario. When constructing Non-IID data via Dirichlet Distribution7, some clients may have very few samples or highly concentrated labels; however, these statistical phenomena alone do not constitute proof of a specific attack’s success, and explicit observation and experimental conditions are still required.

What Do Common Defenses Protect?

MethodPrimary EffectBoundaries to Retain
Data minimization, sensitive field handling, and deduplicationReduces sensitive content entering the system before training and opportunities for repeated memorizationCannot guarantee by itself that remaining data will not leak
Training measures such as regularization and early stoppingImproves generalization and may reduce some membership inference signalsGood average generalization does not equate to per-sample privacy guarantees
Reduce probability outputs, query limits, and access controlLimit attackers’ observation and invocation capabilitiesMust cover model files, logs, and other accessible outputs
Clipping, quantization, compression, or larger batch sizesChange the information carried by updates and the difficulty of attacksCannot claim a DP guarantee without privacy analysis
Secure aggregation, encrypted computation, or trusted execution environmentsProtect the computation process under specified trust and protocol assumptionsReleased outputs may still enable inference
Differential PrivacyLimit output distribution changes caused by altering protected unitsAdjacency relations, sampling methods, and total privacy budget must be explicitly defined

DP-SGD is not just about adding a bit of noise to gradients

The core of DP-SGD includes clipping gradients per sample and adding calibrated noise after aggregation. A common schematic form is

$$ \overline g_i=\frac{g_i}{\max(1,\|g_i\|_2/C)}, \qquad \widetilde g=\frac1{|\mathcal B|} \left(\sum_{i\in\mathcal B}\overline g_i+\xi\right), \qquad \xi\sim\mathcal N(0,\sigma^2 C^2 I). $$

$C$ is the clipping threshold, $\sigma$ is the noise multiplier, and $\mathcal B$ is the current batch. This illustrates the algorithmic structure; specific sensitivity and $(\varepsilon,\delta)$ calculations depend on adjacency definitions, sampling mechanisms, and privacy accounting methods, and cannot be directly read from this formula alone. 8

Clipping only the batch gradient does not justify reusing the sensitivity bound or privacy accounting for per-sample clipping; adding noise and then providing unnoised gradients to an attacker does not protect that additional observation via public DP guarantees. Optimizers, checkpoints, data-dependent hyperparameter tuning, or additional statistical releases must also be incorporated into the analysis.

In federated learning, one must distinguish between record-level and user-level protection. If a user contributes many records, record-level DP parameters cannot be directly applied as user-level DP parameters. User-level mechanisms typically require controlling influence at a level matching user contributions, combined with analysis of actual participation and release methods.

DP does not require hiding all group patterns, nor does it guarantee that no sensitive attribute can be inferred from existing public information. For further discussion on adjacency relations, auxiliary information, and information content, refer to Information Theory in Privacy-Preserving Computation9.

How to Design a Persuasive Privacy Experiment?

First, clearly define the observation scope and the protected unit. An attack experiment providing only final predictions cannot represent the risk of releasing all training gradients; studying a single record does not directly demonstrate the security of all data for a user.

Then, select corresponding metrics for different goals: report TPR at low FPR and prior assumptions for membership inference; report matching, error, and success rates against real inputs for reconstruction attacks; report verifiable matching criteria and counts for data extraction attacks. Do not let a single impressive reconstruction image replace statistical analysis of the entire dataset.

Calibration and final evaluation data should be separated, and one must check whether the distributions of members and non-members match. Training, attack optimization, and partitioning random seeds also affect results; fluctuations and failure cases should be reported, not just selected successful images.

Finally, evaluate defenses using multiple reasonable attack baselines, considering whether the attacker knows the defense settings. Theoretical guarantees require proven or trusted-implementation privacy accounting; empirical risk requires actual attack assessment. Attack failures are useful evidence but cannot replace strict guarantees.

Conclusion

Privacy leakage in deep learning is not a single problem. Knowing whether a sample participated in training, inferring a specific sensitive attribute, reconstructing inputs from gradients, and extracting memorized training content involve different secrets and different attack surfaces.

Only after understanding these distinctions can we find the correct placement for defenses: which information should never enter training, which intermediate results need hiding, which released results need limiting individual influence, and what experimental results actually prove. Protecting privacy requires not just ’not transmitting raw data,’ but also a complete and explicit threat model.

References and Further Reading