A Theoretical Framework for Masked Pretraining (MPT)
Authors
Research Topics
Paper Information
-
Journal:
Journal of Machine Learning Research -
Added to Tracker:
Sep 09, 2026
Abstract
Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. However, the theoretical understanding of the working mechanism behind MPT is still limited. In this paper, we introduce a new theoretical framework to analyze MPT and understand the crucial role of masking in extracting meaningful representations. We establish theoretical connections between MPT and another popular self-supervised paradigm: contrastive learning. We prove that the masking technique implicitly creates positive pairs that are semantically similar and the reconstruction loss pulls them together in the feature space. Besides, as a result of the implicit alignment, we point out the dimensional collapse issue of MPT and propose a Uniformity-enhanced MPT (U-MPT) loss that can effectively address this issue and bring significant improvements in downstream tasks including linear evaluation, cross-dataset fine-tuning and out-of-distribution generalization on real-world data sets. Furthermore, we establish downstream guarantees of U-MPT and theoretically analyze the influence of masking strategies. Based on the theoretical analysis, we propose a new masking strategy which enhances the downstream performance of MPT and explains current improvements of masking strategies with our theoretical perspective.
Author Details
Qi Zhang
AuthorYifei Wang
AuthorYisen Wang
AuthorRunyu Zhou
AuthorResearch Topics & Keywords
Machine Learning
Research AreaCitation Information
APA Format
Qi Zhang
,
Yifei Wang
,
Yisen Wang
&
Runyu Zhou
.
A Theoretical Framework for Masked Pretraining (MPT).
Journal of Machine Learning Research
.
BibTeX Format
@article{paper1663,
title = { A Theoretical Framework for Masked Pretraining (MPT) },
author = {
Qi Zhang
and Yifei Wang
and Yisen Wang
and Runyu Zhou
},
journal = { Journal of Machine Learning Research },
url = { https://www.jmlr.org/papers/v27/25-0477.html }
}