Deep learning with missing data
Authors
Research Topics
Paper Information
-
Journal:
Journal of the Royal Statistical Society Series B -
DOI:
10.1093/jrsssb/qkag114 -
Published:
July 29, 2026 -
Added to Tracker:
Jul 30, 2026
Abstract
Abstract In the context of multivariate nonparametric regression with missing covariates, we propose pattern embedded neural networks (PENNs), which can be applied in conjunction with any existing imputation technique. In addition to a neural network trained on the imputed data, PENNs pass the vectors of observation indicators through a second neural network to provide a compact representation. The outputs are then combined in a third neural network to produce final predictions. Our main theoretical result exploits an assumption that the observation patterns can be partitioned into cells on which the Bayes regression function behaves similarly, and belongs to a compositional Hölder class. It provides a finite-sample excess risk bound that holds for an arbitrary missingness mechanism, and in combination with a complementary minimax lower bound, demonstrates that our PENN estimator attains in typical cases the minimax rate of convergence as if the cells of the partition were known in advance, up to a poly-logarithmic factor in the sample size. Numerical experiments on simulated, semi-synthetic, and real data confirm that the PENN estimator consistently improves, often dramatically, on standard neural networks without pattern embedding. Code to reproduce our experiments, as well as a tutorial on how to apply our method, is publicly available.
Author Details
Tengyao Wang
AuthorTianyi Ma
AuthorRichard J Samworth
AuthorResearch Topics & Keywords
Machine Learning
Research AreaCitation Information
APA Format
Tengyao Wang
,
Tianyi Ma
&
Richard J Samworth
(2026)
.
Deep learning with missing data.
Journal of the Royal Statistical Society Series B
, 10.1093/jrsssb/qkag114.
BibTeX Format
@article{paper1490,
title = { Deep learning with missing data },
author = {
Tengyao Wang
and Tianyi Ma
and Richard J Samworth
},
journal = { Journal of the Royal Statistical Society Series B },
year = { 2026 },
doi = { 10.1093/jrsssb/qkag114 },
url = { https://doi.org/10.1093/jrsssb/qkag114 }
}