Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
Abstract
Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the e...
Description / Details
Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation. We further extend the mean--variance analysis to the more general setting with a nonzero mean entropy gap. Interestingly, this entropy-adjusted score admits a Helmholtz free-energy interpretation, leading to Energy Transfer Detection (ETD), which views pretraining data detection from a macroscopic residual free-energy transfer perspective. Extensive experiments show that ETD achieves the best average detection performance, improving average AUROC by up to 3.5% and TPR@5%FPR by up to 5.1%, while remaining robust across diverse settings.
Source: arXiv:2609.21888v1 - http://arxiv.org/abs/2609.21888v1 PDF: https://arxiv.org/pdf/2609.21888v1 Original Link: http://arxiv.org/abs/2609.21888v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Sep 21, 2026
Artificial Intelligence
AI
0