JRSSB Jul 29, 2026

Deep learning with missing data

Authors
Tengyao Wang Tianyi Ma Richard J Samworth
Research Topics
Machine Learning
Paper Information
  • Journal:
    Journal of the Royal Statistical Society Series B
  • DOI:
    10.1093/jrsssb/qkag114
  • Published:
    July 29, 2026
  • Added to Tracker:
    Jul 30, 2026
Abstract

Abstract In the context of multivariate nonparametric regression with missing covariates, we propose pattern embedded neural networks (PENNs), which can be applied in conjunction with any existing imputation technique. In addition to a neural network trained on the imputed data, PENNs pass the vectors of observation indicators through a second neural network to provide a compact representation. The outputs are then combined in a third neural network to produce final predictions. Our main theoretical result exploits an assumption that the observation patterns can be partitioned into cells on which the Bayes regression function behaves similarly, and belongs to a compositional Hölder class. It provides a finite-sample excess risk bound that holds for an arbitrary missingness mechanism, and in combination with a complementary minimax lower bound, demonstrates that our PENN estimator attains in typical cases the minimax rate of convergence as if the cells of the partition were known in advance, up to a poly-logarithmic factor in the sample size. Numerical experiments on simulated, semi-synthetic, and real data confirm that the PENN estimator consistently improves, often dramatically, on standard neural networks without pattern embedding. Code to reproduce our experiments, as well as a tutorial on how to apply our method, is publicly available.

Author Details
Tengyao Wang
Author
Tianyi Ma
Author
Richard J Samworth
Author
Research Topics & Keywords
Machine Learning
Research Area
Citation Information
APA Format
Tengyao Wang , Tianyi Ma & Richard J Samworth (2026) . Deep learning with missing data. Journal of the Royal Statistical Society Series B , 10.1093/jrsssb/qkag114.
BibTeX Format
@article{paper1490,
  title = { Deep learning with missing data },
  author = { Tengyao Wang and Tianyi Ma and Richard J Samworth },
  journal = { Journal of the Royal Statistical Society Series B },
  year = { 2026 },
  doi = { 10.1093/jrsssb/qkag114 },
  url = { https://doi.org/10.1093/jrsssb/qkag114 }
}