Found 478 papers
Sorted by: Newest FirstA Library for Learning Neural Operators
Kamyar Azizzadenesheli, Jean Kossaifi, Anima Anandkumar et al.
We present NeuralOperator, an open-source Python library for operator learning. Neural operators generalize neural networks to maps between function s...
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
Leyi Pan, Sheng Guan, Zheyu Fu et al.
We introduce MarkDiffusion, an open-source Python toolkit for generative watermarking of latent diffusion models. It comprises three key components: a...
OptunaHub: A Platform for Black-Box Optimization
Yoshihiko Ozaki, Shuhei Watanabe, Toshihiko Yanase
Black-box optimization (BBO) underpins advances in domains such as AutoML and Materials Informatics, yet implementations of algorithms and benchmarks ...
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
Zhuoran Yang, Fengzhuo Zhang, Xinyang Hu et al.
Chain-of-Thought (CoT) prompting and its variants have gained significant attention as effective methods for solving multi-step reasoning tasks with p...
Prob-GParareal: A Probabilistic Numerical Parallel-in-Time Solver for Differential Equations
Guglielmo Gattiglio, Lyudmila Grigoryeva, Massimiliano Tamborrino
We introduce Prob-GParareal, a probabilistic extension of the GParareal algorithm designed to provide uncertainty quantification for the Parallel-in-T...
From learnable objects to learnable random objects
Aaron Anderson, Michael Benedikt
We consider the relationship between learnability of a "base class" of functions on a set $X$, and learnability of a class of statistical functions ...
A Theoretical Framework for Masked Pretraining (MPT)
Qi Zhang, Yifei Wang, Yisen Wang et al.
Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various...
Robustness Against Weak or Invalid Instruments: Exploring Nonlinear Treatment Models with Machine Learning
Peter Bühlmann, Zijian Guo, Mengchu Zheng
We discuss causal inference for observational studies with possibly invalid instrumental variables. We propose a novel methodology called two-stage cu...
Kernel-based Distributed Learning Beyond Least Squares
Xu Guo, Heng Lian
We consider one-shot distributed learning problems in a reproducing kernel Hilbert space framework. Current results are limited to the least-squares l...
Optimising Utility Functions in Multi-Objective Markov Decision Processes
Manel Rodriguez-Soto
Multi-Objective Markov Decision Processes (MOMDPs) are among the most prevalent formal frameworks for addressing sequential decision-making problems i...
Bayesian Transfer Learning for Artificially Intelligent Geospatial Systems: A Predictive Stacking Approach
Sudipto Banerjee, Luca Presicce
Building artificially intelligent geospatial systems requires rapid delivery of spatial data analysis on massive scales with minimal human interventio...
Symmetric Rank-k Methods
Chengchang Liu, Cheng chen, Luo Luo
This paper proposes a novel class of block quasi-Newton methods for convex optimization which we call symmetric rank-$k$ (SR-$k$) methods. Each iterat...
From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
Łukasz Dębowski
We inspect the deductive connection between the neural scaling law and Zipf's law--two statements discussed in machine learning and quantitative lingu...
Efficient Inference under Label Shift in Unsupervised Domain Adaptation
Jiwei Zhao, Seong-ho Lee, Yanyuan Ma
In many real-world applications, researchers aim to deploy models trained in a source domain to a target domain, where obtaining labeled data is often...
Adaptive Algorithms for Infinitely Many-Armed Bandits: A Unified Framework
Emmanuel Pilliat
We consider a bandit problem where the budget is smaller than the number of arms, which may be infinite. In this regime, the usual objective in the li...
Gradient Estimation for Mixture Variational Inference
Javier Burroni, Daniel Sheldon
Mixture distributions are expressive variational families for black-box VI, but their discrete component choices complicate gradient estimation. We sy...
torchsom: The Reference PyTorch Library for Self-Organizing Maps
Louis Berthier, Ahmed Shokry, Maxime Moreaud et al.
This paper introduces torchsom, an open-source Python library that provides a reference implementation of the Self-Organizing Map (SOM) in PyTorch. Th...
Pointwise Confidence Estimation in the Non-linear $\ell^2$-regularized Least Squares
Ilja Kuzborskij, Yasin Abbasi Yadkori
We consider a high-probability non-asymptotic confidence estimation in the $\ell^2$-regularized non-linear least-squares setting with fixed design. In...
Safe Learning Under Irreversible Dynamics via Asking for Help
Benjamin Plaut, Juan Liévano-Karim, Hanlin Zhu et al.
Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot...
AgentPEN: A Prediction-Explanation Network for Sequential Stock Movement via LLMs and Recurrent Generation
Shuqi Li, Mengyao Guo, Yunzhong Zheng et al.
The importance of explainability in stock prediction is increasingly recognized, especially for audit and regulatory purposes. Meanwhile, financial ne...
Feedback-Enhanced Online Multiple Testing with Applications to Conformal Selection
Changliang Zou, Zhaojun Wang, Haojie Ren et al.
This work studies online multiple testing with feedback, where decisions are made sequentially, and the true state of the hypothesis is revealed after...
Ehrenfeucht-Haussler Rank and Chain of Thought
Pablo Barceló, Alexander Kozachinskiy, Tomasz Steifer
The notion of rank of a Boolean function has been a cornerstone in PAC learning, enabling quasipolynomial-time learning algorithms for polynomial-size...
scikit-activeml: A Comprehensive and User-Friendly Active Learning Library
Marek Herde, Minh Tuan Pham, Daniel Kottke et al.
scikit-activeml is a user-friendly open-source Python library for active learning on top of scikit-learn. Included are implementations of a large coll...
Locally Private Estimation with Public Features
Hanfang Yang, Yuheng Ma, Ke Jia
We initiate the study of locally differentially private (LDP) learning with public features. We define semi-feature LDP, where some features are publi...
Dimension Reduction for Derivative-Informed Operator Learning: An Analysis of Approximation Errors
Thomas O'Leary-Roseberry, Omar Ghattas, Peng Chen et al.
We study the derivative-informed learning of nonlinear operators between infinite-dimensional Hilbert spaces. Such operators can arise as solution map...
Domain Adaptation Targeting Heterogeneous and Imbalanced Subgroups
Tianxi Cai, Doudou Zhou, Molei Liu et al.
Domain adaptation enables generalizable and efficient data-driven research. However, existing work has largely focused on domain adaptation for some i...
On the Effectiveness of the z-Transform Method in Quadratic Optimization
Francis Bach
The z-transform of a sequence is a classical tool used in signal processing, control theory, computer science, and electrical engineering. It allows o...
Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions
Peng Wang, Qing Qu, Huijie Zhang et al.
Despite their empirical success across a wide range of generative tasks, the fundamental principles underlying the ability of diffusion models to lear...
Consistency of Augmentation Graph and Network Approximability in Contrastive Learning
A. Martina Neuman, Chenghui Li
Contrastive learning leverages data augmentation to develop feature representation without relying on large labeled data sets. However, despite its em...
Solving Nonlinear PDEs with Sparse Radial Basis Function Networks
Zihan Shao, Konstantin Pieper, Xiaochuan Tian
We propose a novel framework for solving nonlinear PDEs using sparse radial basis function (RBF) networks. Sparsity-promoting regularization is employ...
Viscosity Convergence Analysis for Deep Q-Networks
Qian Qi
Deep Q-Networks (DQNs) and related residual neural architectures are increasingly used for continuous-time reinforcement learning (CTRL), where optima...
Clustering and Pruning in Causal Data Fusion
Otto Tabell, Santtu Tikka, Juha Karvanen
Data fusion, the process of combining observational and experimental data, can enable the identification of causal effects that would otherwise remain...
Leakage and Interpretability in Concept-Based Models
Enrico Parisini, Tapabrata Chakraborti, Chris Harbron et al.
Concept-based Models aim to improve interpretability by predicting high-level intermediate concepts, representing a promising approach for deployment ...
Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning
Ming Xiang, Stratis Ioannidis, Edmund Yeh et al.
Due to resource constraints or external and internal uncertainties, clients in real-world federated learning systems are often intermittently availabl...
Incorporating external data for analyzing randomized clinical trials: A transfer learning approach
Hanzhong Liu, Yujia Gu, Wei Ma
Randomized clinical trials are the gold standard for analyzing treatment effects. However, increasing costs and ethical concerns may limit trial recru...
A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization
Tianshu Chu, Dachuan Xu, Wei Yao et al.
Bilevel optimization has recently attracted significant attention in machine learning due to its wide range of applications and advanced hierarchical ...
Particle Filter for Bayesian Inference on Privatized Data
Jordan Awan, Yu-Wei Chen, Pranav Sanghi
Differential privacy is a probabilistic framework that protects privacy while preserving data utility. To protect the privacy of the individuals in th...
Test-time regression: a unifying framework for designing sequence models with associative memory
Jiaxin Shi, Ke Alexander Wang, Emily B. Fox
Sequence models lie at the heart of modern deep learning. However, rapid advancements have produced a diversity of seemingly unrelated architectures, ...
Nonparametric Spectral Density Estimation using Interactive Mechanisms under Local Differential Privacy
Cristina Butucea, Karolina Klockmann, Tatyana Krivobokova
We study the problem of estimating the spectral density of a centered stationary Gaussian time series under local differential privacy constraints. Sp...
Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters
Sebastian Salwig, Till Kahlke, Florian Hirschberger et al.
Gaussian Mixture Models (GMMs) range among the most frequently used models in machine learning. However, training large, general GMMs becomes computat...
Statistical Inference for High-dimensional Partially Linear Models via Debiased Rank Lasso
Runze Li, Songshan Yang, Delin Zhao
This paper aims to develop tuning-free and robust regularized methods for partially linear models based on partial residual methods. In order to prese...
Identifiability of the Instrumental Variable Model with the Treatment and Outcome Missing Not at Random
Peng Ding, Fan Yang, Shuozhi Zuo
Under the instrumental variable model, we can identify the local average treatment effect, also known as the complier average causal effect (CACE). In...
Deep Neural Expected Shortfall Regression with Tail-Robustness
Kean Ming Tan, Wen-Xin Zhou, Huixia Judy Wang et al.
Expected shortfall (ES), also known as conditional value-at-risk, is a widely recognized risk measure that complements value-at-risk by capturing tail...
Optimal Convergence Rates for Neural Operators
Mike Nguyen, Nicole Mücke
We introduce the neural tangent kernel (NTK) regime for two-layer neural operators and analyze their generalization properties. For early-stopped grad...
Have ASkotch: A Neat Solution for Large-Scale Kernel Ridge Regression
Pratik Rathore, Zachary Frangella, Jiaming Yang et al.
Kernel ridge regression (KRR) is a fundamental computational tool, appearing in problems that range from computational chemistry to health analytics, ...
Pairwise Comparisons without Stochastic Transitivity: Model, Theory and Applications
Yunxiao Chen, Sze Ming Lee
Most statistical models for pairwise comparisons, including the Bradley-Terry (BT) and Thurstone models and many extensions, make a relatively strong ...
Canonical Correlation Analysis as Reduced Rank Regression in High Dimensions
Claire Donnat, Elena Tuzhilina
Canonical correlation analysis is a widespread technique for discovering linear relationships between two sets of variables. In high dimensions, howev...
Bayesian Level Set Clustering
Miheer Dewaskar, David B. Dunson, David Buch
Classically, Bayesian clustering interprets each component of a mixture model as a cluster. The inferred clustering posterior is highly sensitive to a...
Differentially Private Synthetic Data Generation for Relational Databases
Hao Wang, Navid Azizan, Kaveh Alim et al.
Existing differentially private (DP) synthetic data generation mechanisms typically assume a single-source table. In practice, data is often distribut...
A Neural Network Approach to Learning Solutions of a Class of Elliptic Variational Inequalities
Amal Alphonse, Michael Hintermüller, Alexander Kister et al.
We develop a weak adversarial approach to solving obstacle problems using neural networks. By employing (generalised) regularised gap functions and th...
Impatient Bandits: Optimizing for the Long-Term Without Delay
Kelly W. Zhang, Thomas Baldwin-McDonald, Kamil Ciosek et al.
Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration task, which...
Conditional Regression for the Nonlinear Single-Variable Model
Yantao Wu, Mauro Maggioni
Regressing a function $F$ on $\mathbb{R}^d$ without incurring the statistical and computational curse of dimensionality requires exploitable structure...
Testability of Instrumental Variables in Additive Nonlinear, Non-Constant Effects Models
Zhi Geng, Biwei Huang, Xichen Guo et al.
We address the issue of the testability of instrumental variables derived from observational data. Most existing testable implications are centered on...
Sliced Wasserstein Regression
Yidong Zhou, Hans-Georg Müller, Han Chen
While statistical modeling of distributional data has gained increased attention, the case of multivariate distributions has been somewhat neglected d...
Causal Falsification of Digital Twins
Arnaud Doucet, Rob Cornish, Muhammad Faaiz Taufiq et al.
Digital twins are simulation-based models designed to predict how a real-world process will evolve in response to interventions. This modelling paradi...
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
Gianmarco Genalti, Marco Mussi, Nicola Gatti et al.
Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the ex...
Extrapolation-Aware Nonparametric Statistical Inference
Peter Bühlmann, Niklas Pfister
We define extrapolation as statistical inference on a conditional function (e.g., a conditional expectation or conditional quantile) evaluated outside...
Online Generalized Sparse Regression: How Does Overparametrization Help?
Shuoguang Yang, Qiang Sun
Regularized sparse regression has been extensively studied in the offline setting, but online formulation remains relatively under-explored. This gap...
Simultaneous Identification of Sparse Structures and Communities in Heterogeneous Graphical Models
Zhiliang Ying, Dapeng Shi, Tiandong Wang
Exploring and detecting community structures hold significant importance in genetics, social sciences, biology, neuroscience and finance, among others...
Better Simulations for Validating Causal Discovery with the DAG-Adaptation of the Onion Method
Bryan Andrews, Erich Kummerfeld
The number of methods for learning causal models from data is growing rapidly, as evidenced by the exponential increase in causal discovery publicatio...
Singular-limit analysis of gradient descent with noise injection
Anna Shalova, André Schlichting, Mark Peletier
We study the limiting dynamics of a large class of noisy gradient descent systems in the overparameterized regime. In this regime the zero-loss set of...
Information-Theoretic Safe Bayesian Optimization
Alessandro G. Bottero, Carlos E. Luis, Julia Vinogradska et al.
We consider a sequential decision making problem, where we aim to optimize an unknown function via noisy evaluations that do not violate an a-priori u...
Dirichlet Active Learning
Ryan Murray, Kevin Miller
This work introduces Dirichlet Active Learning (DiAL), a Bayesian-inspired approach to the design of active learning algorithms. Our framework models ...
Efficient Modeling of Surrogates to Improve Multi-source High-dimensional Integrative Regression
Tianxi Cai, Zijian Guo, Yue Liu et al.
Surrogate variables play an important role in various fields due to the scarcity or absence of gold-standard labels. We develop a novel approach named...
torchgfn: A PyTorch GFlowNet Library
Joseph D. Viviano, Omar G. Younis, Sanghyeok Choi et al.
The growing popularity of generative flow networks (GFlowNets or GFNs) among a range of researchers with diverse backgrounds and areas of expertise ne...
Model-free generalized fiducial inference
Jonathan P Williams
Conformal prediction (CP) was developed to provide finite-sample probabilistic prediction guarantees. While CP algorithms are a relatively general pu...
Adversarial Rademacher Complexity of Deep Neural Networks
Jiancong Xiao, Yanbo Fan, Ruoyu Sun et al.
Deep neural networks (DNNs) are highly vulnerable to adversarial attacks. Ideally, a robust model should perform well on both perturbed training data ...
Keypoint-Guided Optimal Transport: Models, Algorithms, and Applications
Zongben Xu, Xiang Gu, Yucheng Yang et al.
Existing Optimal Transport (OT) methods mainly derive the optimal transport plan/matching under the criterion of transport cost/distance minimization,...
The Role of Pseudo-Labels in Self-Training Linear Classifiers on High-Dimensional Gaussian Mixture Data
Takashi Takahashi
Self-training (ST) is a simple yet effective semi-supervised learning method. However, why and how ST improves generalization performance by using pot...
Bridging Domain Invariance and Diversity: A Fine-Grained Risk Bound for Domain Generalization
Xi Wang, Liang Bai, Xian Yang et al.
Domain-invariant representation learning and domain augmentation algorithms are two principal methodological paradigms for addressing domain generaliz...
Global Fréchet Manifold Learning for Random Objects, With Application to Low-Dimensional Wasserstein Representations of Distributional Data
Hans-Georg Müller, Álvaro Gajardo
We study manifold learning with multidimensional scaling for samples of metric space valued data. By adopting a global version of ISOMAP we obtain low...
Nonparametric generative modeling for time series via Schrödinger bridge
Mohamed Hamdouche, Pierre Henry-Labordère, Huyên Pham
We propose a novel generative model for time series based on Schrödinger bridge (SB) approach. This consists in the entropic interpolation via optimal...
High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
Francis Bach, Simon Martin, Giulio Biroli
We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher--student setup. We focus on the ext...
Error Analyses of Auto-Regressive Video Diffusion Models
Zhuoran Yang, Jing Wang, Fengzhuo Zhang et al.
Auto-Regressive Video Diffusion Models (AR-VDMs) have shown strong capabilities in generating long, photorealistic videos, but suffer from two key lim...
Near-optimal Delta-convex Estimation of Lipschitz Functions
G{\'{a}}bor Bal{\'{a}}zs
This paper presents a tractable algorithm for estimating an unknown Lipschitz function from noisy observations and establishes an upper bound on its c...
The Sample Complexity of Parameter-Free Stochastic Convex Optimization
Jared Lawrence, Ari Kalinsky, Hannah Bradfield et al.
We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant...
End-to-End Deep Learning for Predicting Metric Space-Valued Outputs
Su I Iao, Yidong Zhou, Hans-Georg M{\"{u}}ller
Many modern applications involve predicting structured, non-Euclidean outputs such as probability distributions, networks, and symmetric positive-defi...
Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective
Wenlong Lyu, Yuheng Jia, Hui Liu et al.
The well-known graph-based clustering methods, including spectral clustering, symmetric non-negative matrix factorization, and doubly stochastic norma...
Learning to Play Two-Player Perfect-Information Games without Knowledge
Quentin Cohen-Solal
This paper introduces a set of techniques for learning game state evaluation functions through reinforcement learning. First, we generalize tree boots...
Doubly Debiased Robust Subsampling for Transfer Learning
Tao Wang, Weng Kee Wong
This paper develops a general framework for doubly debiased robust subsampling for transfer learning. The setting arises when massive source datasets ...
Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
Philip Sosnin, Matthew Wicker, Josh Collyer et al.
The impact of inference-time data perturbation (e.g., adversarial attacks) has been extensively studied in machine learning, leading to well-establish...
Mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression
Giacomo Zanella, Filippo Ascolani
We investigate the convergence properties of popular data-augmentation samplers for Baye\-sian probit regression. Leveraging recent results on Gibbs s...
Underdamped Langevin MCMC with third order convergence
Maximilian Scott, D{\'{a}}ire O'Kane, Andraž Jelinčič et al.
In this paper, we propose a new numerical method for the underdamped Langevin diffusion (ULD) and present a non-asymptotic analysis of its sampling er...
Approximation-Free Differentiable Oblique Decision Trees
Subrat Prasad Panda, Blaise Genest, Arvind Easwaran
Decision Trees (DTs) are widely used in safety-critical domains such as medical diagnosis, valued for their interpretability and effectiveness on tabu...
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
Peter L. Bartlett, Ruiqi Zhang, Jingfeng Wu et al.
We study gradient descent (GD) for logistic regression on linearly separable data with stepsizes that adapt to the current risk, scaled by a constant ...
Adaptive Nonparametric Perturbations of Parametric Models with Generalized Bayes
Bohan Wu, Yixin Wang, Eli N. Weinstein et al.
Parametric Bayesian modeling offers a powerful and flexible toolbox for machine learning. Yet the model, however detailed, may still be wrong, and thi...
Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss
Jos{\'{e}} Manuel de Frutos, Manuel A. V{\'{a}}zquez, Pablo M. Olmos et al.
Implicit generative models are often trained adversarially, which can yield unstable dynamics and mode collapse. The invariant statistical loss (ISL) ...
Gradient Span Algorithms Make Predictable Progress in High Dimension
Felix Benning, Leif D{\"{o}}ring
We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to ...
py/cuTAGI: An Open-Source Library for Tractable Approximate Gaussian Inference in Bayesian Neural Networks
Luong-Ha Nguyen, James-A. Goulet, Miquel Florensa-Montilla et al.
This paper introduces pyTAGI, a Python wrapper, and cuTAGI, its high-performance C++/CUDA backend, implementing Tractable Approximate Gaussian Inferen...
Statistical Test for Attention in Transformers for Images and Time Series
Tomohiro Shiraishi, Daiki Miwa, Teruyuki Katsuoka et al.
Transformer models have achieved exceptional performance in various domains, including computer vision and time-series analysis. Their core attention ...
Accelerating Constrained Sampling: A Large Deviations Approach
Lingjiong Zhu, Yingli Wang, Changwei Tu et al.
The problem of sampling a target probability distribution on a constrained domain arises in many applications including machine learning. For constrai...
Learning general conditional independence structures via the neighbourhood lattice
Bryon Aragam, Arash A. Amini, Qing Zhou
We study the problem of learning multivariate dependencies in nonparametric and high-dimensional settings. This includes but is not limited to graphic...
Statistical guarantees for denoising reflected diffusion models
Claudia Strauch, Lukas Trottner, Asbjørn Holk
In recent years, denoising diffusion models have become a crucial area of research due to their abundance in the rapidly expanding field of generative...
Vecchia-Inducing-Points Full-Scale Approximations for Gaussian Processes
Tim Gyger, Reinhard Furrer, Fabio Sigrist
Gaussian processes are flexible, probabilistic, non-parametric models widely used in machine learning and statistics. However, their scalability to la...
STDE++: Polynomial-Time Amortization for Linear Differential Operators
Zekun Shi, Zheyuan Hu, Min Lin et al.
Optimizing neural networks with losses that contain high-dimensional and high-order differential operators is expensive to evaluate with backpropagati...
The Within-Orbit Adaptive Leapfrog No-U-Turn Sampler
Sifan Liu, Nawaf Bou-Rabee, Bob Carpenter et al.
Locally adapting parameters within Markov chain Monte Carlo methods while preserving reversibility is notoriously difficult. The success of the No-U-...
Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation
Zhihua Zhang, Xiang Li, Yuze Han
In two-time-scale stochastic approximation (SA), two iterates are updated at varying speeds using different step sizes, with each update influencing t...
Embedding Network Autoregression for Time Series Analysis and Causal Peer Effect Inference
Jae Ho Chang, Subhadeep Paul
We propose an Embedding Network Autoregressive Model for multivariate networked longitudinal data. We assume the network is generated from a latent va...
Three Types of Calibration using Properties and their Semantic and Formal Relationships
Robert C. Williamson, Rabanus Derr, Jessie Finocchiaro
Fueled by discussions around "trustworthiness" and algorithmic fairness, calibration of predictive systems has regained scholars' attention. The vanil...
Convergence of Decentralized Stochastic Subgradient-based Methods for Nonsmooth Nonconvex Optimization
Xin Liu, Siyuan Zhang, Nachuan Xiao
In this paper, we focus on the decentralized stochastic subgradient-based methods in minimizing nonsmooth nonconvex functions without Clarke regularit...
A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance
Axel F. Wolter, Tobias Sutter
We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stochastic ap...
FLAGG: Flexible Autoregressive Graph Generation
Samuel Cognolato, Alessandro Sperduti, Luciano Serafini
The Deep Graph Generation's panorama spans two extremes: one-shot and sequential models. The former generates nodes and edges jointly, while the latte...
Nested Subspace Learning with Flags
Tom Szwagier, Xavier Pennec
Many machine learning methods look for low-dimensional representations of the data. The underlying subspace can be estimated by first choosing a dimen...
A Unified Approach to Analysis and Design of Denoising Markov Models
Yinuo Ren, Grant M. Rotskoff, Lexing Ying
Probabilistic generative models based on measure transport, such as diffusion and flow-based models, are often formulated in the language of Markovian...
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
Nikolaos Tsilivis, Eitan Gronich, Julia Kempe et al.
We study the implicit bias of the general family of steepest descent algorithms with infinitesimal learning rate in deep homogeneous neural networks. ...
A Single-Loop Stochastic Proximal Quasi-Newton Method for Large-Scale Nonsmooth Convex Optimization
Yongcun Song, Zimeng Wang, Xiaoming Yuan et al.
We propose a new stochastic proximal quasi-Newton method for minimizing the sum of two convex functions in the particular context that one of the func...
Statistical Learning Theory for Neural Operators
Niklas Reinhardt, Sven Wang, Jakob Zech
We present statistical convergence results for the learning of (possibly) non-linear mappings in infinite-dimensional spaces. Specifically, given a ...
Deconvolution in unlinked linear models
Fadoua Balabdaoui, Antonio Di Noia, C{\'{e}}cile Durot
Unlinked regression, in which covariates and responses are observed separately without known correspondence, has recently gained increasing attention....
Spectral Truncation Kernels: Noncommutativity in C*-algebraic Kernel Machines
Yuka Hashimoto, Ayoub Hafid, Masahiro Ikeda et al.
A central question in vector- and function-valued learning is how to design kernels that capture both local and non-local interactions while remaining...
High-dimensional Parameter Transfer With Fused-Regularizer
Runze Li, Ying Sun, Jingyuan Liu et al.
Parameter transfer aims to improve parameter estimation accuracy by leveraging knowledge from related sources. This paper studies the parameter transf...
Exogenous Randomness Empowering Random Forests
Yingying Fan, Jinchi Lv, Tianxing Mei
We offer theoretical and empirical insights into the impact of exogenous randomness on the effectiveness of random forests with tree-building rules in...
Kernel Mean Embedding Deviation Subspace for Unsupervised Learning with Heterogeneous Data
Lixing Zhu, Ruoqing Zhu, Luoyao Yu et al.
This paper proposes a method for dimension reduction that preserves information in unsupervised learning with high-dimensional heterogeneous data, spe...
Deep Nonparametric Conditional Independence Tests for Images
Sonja Greven, Xiangnan Xu, Marco Simnacher et al.
Conditional independence tests (CITs) test for conditional dependence between random variables given a vector of conditioning or confounder variables....
Semi-supervised learning for linear extremile regression
Jiangfeng Wang, Keming Yu, Rong Jiang
Extremile regression, as a least squares analog of quantile regression, is potentially a useful tool for modeling and understanding the extreme tails ...
Transfer Learning via Regularized Random-effects Linear Discriminant Analysis
Hongzhe Li, Hongzhe Zhang, Arnab Auddy
Linear discriminant analysis is a widely used method for classification. However, the high dimensionality of predictors combined with small sample siz...
Cheap Bootstrap for Fast Uncertainty Quantification of Stochastic Gradient Descent
Henry Lam, Zitong Wang
Stochastic gradient descent (SGD) or stochastic approximation has been widely used in model training and stochastic optimization. While there is a hug...
A Natural Primal-Dual Hybrid Gradient Method for Adversarial Neural Network Training on Solving Partial Differential Equation
Stanley Osher, Shu Liu, Wuchen Li
We propose a scalable preconditioned primal-dual hybrid gradient algorithm for solving partial differential equations (PDEs). We multiply the PDE with...
Generalized Resubstitution for Regression Error Estimation
Diego Marcondes, Ulisses Braga-Neto
We propose generalized resubstitution error estimators for regression. Each error estimator in this class corresponds to a choice of an empirical prob...
Transfer Conformal Predictive Inference for Regression
Linglong Kong, Jinhan Xie, Bei Jiang et al.
Conformal prediction, a powerful framework for constructing prediction intervals for response variables using any regression function estimators, ofte...
Towards Convexity in Anomaly Detection: A New Formulation of SSLM with Unique Optimal Solutions
Hao Wang, Hongying Liu, Haoran Chu et al.
An unsolved issue in widely used methods such as Support Vector Data Description (SVDD) and Small Sphere and Large Margin SVM (SSLM) for anomaly detec...
Node Regression on Latent Position Random Graphs via Local Averaging
Martin Gjorgjevski, Nicolas Keriven, Simon Barthelme et al.
Node regression consists in predicting the value of a graph label at a node, given observations at the other nodes. We perform a theoretical study whe...
On the Relevance of Byzantine Robust Optimization Against Data Poisoning
Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta et al.
The success of machine learning (ML) has been intimately linked with the availability of large amounts of data, typically collected from heterogeneou...
Best Arm Identification with Minimal Regret
Junwen Yang, Vincent Y. F. Tan, Tianyuan Jin
Motivated by real-world applications that necessitate responsible experimentation, we introduce the problem of best arm identification (BAI) with mini...
Convergence of Noise-Free Sampling Algorithms with Regularized Wasserstein Proximals
Stanley Osher, Wuchen Li, Fuqun Han
In this work, we investigate the convergence properties of the backward regularized Wasserstein proximal (BRWP) method for sampling a target distribut...
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
Shangtong Zhang, Jiuqi Wang
Temporal difference (TD) learning with linear function approximation (linear TD) is a classic and powerful prediction algorithm in reinforcement learn...
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models
Wei Biao Wu, Johannes Schmidt-Hieber, Jiaqi Li
This paper proposes an asymptotic theory for online inference of the stochastic gradient descent (SGD) iterates with dropout regularization in linear ...
Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data
Wanli Hong, Shuyang Ling
Neural collapse (${\cal NC}$) is a phenomenon that emerges at the terminal phase of the training (TPT) of deep neural networks (DNNs). The features of...
Differentially Private Estimation and Inference in High-Dimensional Regression with FDR Control
Sai Li, Linjun Zhang, Zhanrui Cai et al.
This paper proposes new methodologies for conducting practical differentially private (DP) estimation and inference in high-dimensional linear regress...
Demographic Parity in Regression and Classification Within the Unawareness Framework
Vincent Divol, Solenne Gaucher
This paper explores the theoretical foundations of fair regression under the constraint of demographic parity within the unawareness framework, where ...
Enhancing Accuracy in Generative Models via Knowledge Transfer
Xinyu Tian, Xiaotong Shen
This paper investigates the accuracy of generative models and the impact of knowledge transfer on their generation precision. Specifically, we examine...
Multi-relational Network Autoregression Model with Latent Group Structures
Yanyuan Ma, Xuening Zhu, Yimeng Ren et al.
Multi-relational networks among entities are frequently observed in the era of big data. Quantifying the effects of multiple networks has attracted si...
Limiting Over-Smoothing and Over-Squashing of Graph Message Passing by Deep Scattering Transforms
Yuanhong Jiang, Dongmian Zou, Xiaoqun Zhang et al.
Graph neural networks (GNNs) have become pivotal tools for processing graph-structured data, leveraging the message passing scheme as their core mecha...
A Fully Parameter-Free Second-Order Algorithm for Convex-Concave Minimax Problems
Jun-Lin Wang, Zi Xu, Hui-Ling Zhang
In this paper, we study second-order algorithms for the convex-concave minimax problem, which has attracted much attention in many fields such as mach...
Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent
Loucas Pillaud-Vivien, Adrien Schertzer
We study the dynamics of a continuous-time model of stochastic gradient descent (SGD) for the least-square problem. Indeed, pursuing the work of, we a...
Vector-Valued Gaussian Processes for Approximating Divergence- or Rotation-free Vector Fields
Quoc Thong Le Gia, Ian Hugh Sloan, Holger Wendland
In this paper, we discuss vector-valued Gaussian processes for the approximation of divergence- or rotation-free functions. We establish the theory fo...
Differentially Private Best-Arm Identification
Marc Jourdan, Achraf Azize, Aymen Al Marjani et al.
Best Arm Identification (BAI) problems are progressively used for data-sensitive applications, such as designing adaptive clinical trials, tuning hype...
Corruptions of Supervised Learning Problems: Typology and Mitigations
Robert C. Williamson, Nan Lu, Laura Iacovissi
Corruption is notoriously widespread in data collection. Despite extensive research, the existing literature predominantly focuses on specific setting...
Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies
S{\'{e}}bastien Lachapelle, Pau Rodr{\'{i}}guez L{\'{o}}pez, Yash Sharma et al.
This work introduces a novel principle for disentanglement we call mechanism sparsity regularization, which applies when the latent factors of interes...
A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization
Zhaoran Wang, Zhuoran Yang, Yuchen Zhu et al.
This paper studies minimax optimization problems defined over infinite-dimensional function classes of over-parameterized two-layer neural networks. I...
Minimax density estimation in the adversarial framework under local differential privacy
M{\'{e}}lisande Albert, Juliette Chevallier, B{\'{e}}atrice Laurent et al.
We consider the problem of nonparametric density estimation under privacy constraints in an adversarial framework. To this end, we study minimax rates...
Approximations and Learning for Continuous State and Action MDPs under Average Cost Criteria
Ali D. Kara, Serdar Y{\"{u}}ksel
In this paper, for Markov Decision Processes (MDPs) with standard Borel spaces, (i) we first provide a discretization based approximation method for M...
Optimal Approximation and Generalization Errors for Deep Convolutional Neural Networks
Shao-Bo Lin, Jinxin Wang
This paper focuses on approximation and learning performances of deep convolutional neural networks with zero-padding and max-pooling. We prove that...
Investigating the Histogram Loss in Regression
Ehsan Imani, Kai Luedemann, Sam Scholnick-Hughes et al.
It is becoming increasingly common in regression to train neural networks that model the entire distribution even if only the mean is required for pre...
Bayes-Optimal Fair Classification with Linear Disparity Constraints via Pre-, In-, and Post-processing
Edgar Dobriban, Guang Cheng, Xianli Zeng et al.
Machine learning algorithms may have disparate impacts on protected groups. To address this, we develop methods for Bayes-optimal fair classificatio...
Why "Classic" Transformers Are Shallow and A Depth-Enabling Technique
Yueyao Yu, Yin Zhang
Since its introduction in 2017, the Transformer has emerged as the leading neural network architecture, catalyzing revolutionary advancements in many ...
Kernel-based Distributed Learning
Xu Guo, Heng Lian
We consider one-shot distributed learning problems in a reproducing kernel Hilbert space framework. Current results are limited to the least-squares l...
A Convex Framework for Confounding Robust Inference
Kei Ishikawa, Niao He, Takafumi Kanamori
We study policy evaluation of offline contextual bandits subject to unobserved confounders. Sensitivity analysis methods are commonly used to estimate...
Global Fr{\'{e}}chet Manifold Learning for Random Objects, With Application to Low-Dimensional Wasserstein Representations of Distributional Data
Hans-Georg M{\"{u}}ller, {\'{A}}lvaro Gajardo
We study manifold learning with multidimensional scaling for samples of metric space valued data. By adopting a global version of ISOMAP we obtain low...
Probabilistic Rainfall Downscaling: Joint Generalized Neural Models with Censored Spatial Gaussian Copula
David Huk, Rilwan A. Adewoyin, Ritabrata Dutta
A novel approach for generating conditional probabilistic rainfall downscaling at finer scales from deterministic weather variables at coarser scales ...
Sparse Topic Modeling via Spectral Decomposition and Thresholding
Huy Tran, Yating Liu, Claire Donnat
In probabilistic Latent Semantic Indexing (pLSI), word frequencies across document corpora are modeled through a low-rank factorization of the expecte...
Nonparametric generative modeling for time series via Schr{\"{o}}dinger bridge
Mohamed Hamdouche, Pierre Henry-Labord{\`{e}}re, Huy{\^{e}}n Pham
We propose a novel generative model for time series based on Schrödinger bridge (SB) approach. This consists in the entropic interpolation via optimal...
Do We Need to Penalize Variance of Losses for Learning with Label Noise?
Jun Yu, Mingming Gong, Yexiong Lin et al.
Statistically consistent algorithms have been widely employed for dealing with noisy labels. Their objective functions are designed so that minimizing...
Causal Influences over Social Learning Networks
Mert Kayaalp, Ali H. Sayed
This paper investigates causal influences between agents linked by a social graph and interacting over time. In particular, the work examines the dyna...
Neural Exploitation and Exploration of Contextual Bandits
Yikun Ban, Yuchen Yan, Arindam Banerjee et al.
In this paper, we study the neural exploration strategy for contextual bandits. The dilemma of exploitation and exploration widely exists in real-wor...
Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation
Luyang Fang, Haoran Lu, Yongkai Chen et al.
As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge disti...
Inference with non-differentiable surrogate loss in a general high-dimensional classification framework
Ying-Qi Zhao, Yang Ning, Muxuan Liang et al.
Penalized empirical risk minimization with a surrogate loss function is often used to learn a high-dimensional linear decision rule in classification ...
A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks
Eric Vanden-Eijnden, Zhengdao Chen, Joan Bruna
To understand the training dynamics of neural networks, prior studies have considered the mean-field (MF) limit of two-layer NNs as the width tends to...
The Role of Contextual Information in Best Arm Identification
Masahiro Kato, Kaito Ariu
We study the best-arm identification problem with fixed confidence when contextual (covariate) information is available in stochastic bandits. In each...
Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
Yuling Jiao, Yanming Lai, Yang Wang et al.
The Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates th...
Online Bernstein-von Mises theorem
Jeyong Lee, Minwoo Chae, Junhyeok Choi
Online learning is an inferential paradigm in which parameters are updated incrementally from sequentially available data, in contrast to batch learni...
Covariate-dependent Hierarchical Dirichlet Processes
Sara Wade, Huizi Zhang, Natalia Bochkina
Bayesian hierarchical modeling is a natural framework to effectively integrate data and borrow information across groups. In this paper, we address pr...
DCatalyst: A Unified Accelerated Framework for Decentralized Optimization
Gesualdo Scutari, TIanyu Cao, Xiaokai Chen
We study decentralized optimization over a network of agents, modeled as an undirected graph and operating without a central server. The objective is ...
Boosted Control Functions: Distribution Generalization and Invariance in Confounded Models
Jonas Peters, Niklas Pfister, Sebastian Engelke et al.
Modern machine learning methods and the availability of large-scale data have significantly advanced our ability to predict target quantities from lar...
Contrasting Local and Global Modeling with Machine Learning and Satellite Data: A Case Study Estimating Tree Canopy Height in African Savannas
Esther Rolf, Lucia Gordon, Milind Tambe et al.
While advances in machine learning with satellite imagery (SatML) are facilitating environmental monitoring at a global scale, developing SatML models...
A Symplectic Analysis of Alternating Mirror Descent
Jonas E. Katona, Xiuyuan Wang, Andre Wibisono
Motivated by understanding the behavior of the Alternating Mirror Descent (AMD) algorithm for bilinear zero-sum games, we study the discretization of ...
Two-way Node Popularity Model for Directed and Bipartite Networks
Ting Li, Bing-Yi Jing, Jiangzhou Wang et al.
There has been increasing research attention on community detection in directed and bipartite networks. However, these studies often fail to consider ...
Convergence and complexity of block majorization-minimization for constrained block-Riemannian optimization
Deanna Needell, Laura Balzano, Yuchen Li et al.
Block majorization-minimization (BMM) is a simple iterative algorithm for nonconvex optimization that sequentially minimizes a majorizing surrogate of...
Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood
Jiangrong Ouyang, Mingming Gong, Howard Bondell
Policy inference plays an essential role in the contextual bandit problem. In this paper, we use empirical likelihood to develop a Bayesian inference ...
A causal fused lasso for interpretable heterogeneous treatment effects estimation
Oscar Hernan Madrid Padilla, Yanzhen Chen, Carlos Misael Madrid Padilla et al.
We propose a novel method for estimating heterogeneous treatment effects based on the fused lasso. By first ordering samples based on the propensity o...
Unsupervised Feature Selection via Nonnegative Orthogonal Constrained Regularized Minimization
Defeng Sun, Liping Zhang, Yan Li
Unsupervised feature selection has drawn wide attention in the era of big data, since it serves as a fundamental technique for dimensionality reductio...
Reparameterized Complex-valued Neurons Can Efficiently Learn More than Real-valued Neurons via Gradient Descent
Zhi-Hua Zhou, Jin-Hui Wu, Shao-Qun Zhang et al.
Complex-valued neural networks potentially possess better representations and performance than real-valued counterparts when dealing with some complic...
Hierarchical Causal Models
Eli N. Weinstein, David M. Blei
Causal questions often arise in settings where data are hierarchical: subunits are nested within units. Consider students in schools, cells in patient...
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
Addison Kristanto Julistiono, Davoud Ataee Tarzanagh, Navid Azizan
Attention mechanisms have revolutionized several domains of artificial intelligence, such as natural language processing and computer vision, by enabl...
Adaptive Forward Stepwise: A Method for High Sparsity Regression
Ivy Zhang, Robert Tibshirani
This paper proposes a sparse regression method that continuously interpolates between Forward Stepwise selection (FS) and the LASSO. When tuned approp...
Optimization and Generalization of Gradient Descent for Shallow ReLU Networks with Minimal Width
Ding-Xuan Zhou, Yunwen Lei, Puyu Wang et al.
Understanding the generalization and optimization of neural networks is a longstanding problem in modern learning theory. The prior analysis often lea...
Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection
Steven Adams, Andrea Patanè, Morteza Lahijanian et al.
Infinitely wide or deep neural networks (NNs) with independent and identically distributed (i.i.d.) parameters have been shown to be equivalent to Gau...
CHANI: Correlation-based Hawkes Aggregation of Neurons with bio-Inspiration
Sophie Jaffard, Samuel Vaiter, Patricia Reynaud-Bouret
The present work aims at proving mathematically that a neural network inspired by biology can learn a classification task thanks to local transformati...
Persistence Diagrams Estimation of Multivariate Piecewise Hölder-continuous Signals
Hugo Henneuse
To our knowledge, the analysis of convergence rates for persistence diagrams estimation from noisy signals has predominantly relied on lifting signal ...
Exploring Novel Uncertainty Quantification through Forward Intensity Function Modeling
Yudong Wang, Cheng Yong Tang, Zhi-Sheng Ye
Predicting future time-to-event outcomes is a foundational task in statistical learning. While various methods exist for generating point predictions,...
Generative Bayesian Inference with GANs
Yuexi Wang, Veronika Rockova
In the absence of explicit or tractable likelihoods, Bayesians often resort to approximate Bayesian computation (ABC) for inference. Our work bridges ...
Communication-efficient Distributed Statistical Inference for Massive Data with Heterogeneous Auxiliary Information
Miaomiao Yu, Zhongfeng Jiang, Jiaxuan Li et al.
Heterogeneous auxiliary information commonly arises in big data due to diverse study settings and privacy constraints. Excluding such indirect evidenc...
Decorrelated Local Linear Estimator: Inference for Non-linear Effects in High-dimensional Additive Models
Zijian Guo, Wei Yuan, Cunhui Zhang
Additive models play an essential role in studying non-linear relationships. Despite many recent advances in estimation, there is a lack of methods an...
Refined Risk Bounds for Unbounded Losses via Transductive Priors
Jian Qian, Alexander Rakhlin, Nikita Zhivotovskiy
We revisit the sequential variants of linear regression with the squared loss, classification problems with hinge loss, and logistic regression, all c...
A Common Interface for Automatic Differentiation
Guillaume Dalle, Adrian Hill
For scientific machine learning tasks with a lot of custom code, picking the right Automatic Differentiation (AD) system matters. Our Julia package Di...
LazyDINO: Fast, Scalable, and Efficiently Amortized Bayesian Inversion via Structure-Exploiting and Surrogate-Driven Measure Transport
Lianghao Cao, Thomas O'Leary-Roseberry, Omar Ghattas et al.
We present LazyDINO, a transport map variational inference method for fast, scalable, and efficiently amortized solutions of high-dimensional nonlinea...
The Distribution of Ridgeless Least Squares Interpolators
Qiyang Han, Xiaocong Xu
The Ridgeless minimum $\ell_2$-norm interpolator in overparametrized linear regression has attracted considerable attention in recent years in both ma...
Nonparametric Estimation of a Factorizable Density using Diffusion Models
Minwoo Chae, Hyeok Kyu Kwon, Dongha Kim et al.
In recent years, diffusion models, and more generally score-based deep generative models, have achieved remarkable success in various applications, in...
Learning Bayesian Network Classifiers to Minimize Class Variable Parameters
Shouta Sugahara, Koya Kato, James Cussens et al.
This study proposes and evaluates a novel Bayesian network classifier which can asymptotically estimate the true probability distribution of the class...
Simulation-based Calibration of Uncertainty Intervals under Approximate Bayesian Estimation
Terrance D. Savitsky, Julie Gershunskaya
The mean field variational Bayes (VB) algorithm implemented in Stan is relatively fast and efficient, making it feasible to produce model-estimated of...
An Anytime Algorithm for Good Arm Identification
Marc Jourdan, Andrée Delahaye-Duriez, Clémence Réda
In good arm identification (GAI), the goal is to identify one arm whose average performance exceeds a given threshold, referred to as a good arm, if i...
Extrapolated Markov Chain Oversampling Method for Imbalanced Text Classification
Aleksi Avela, Pauliina Ilmonen
Text classification is the task of automatically assigning text documents correct labels from a predefined set of categories. In real-life (text) clas...
Neural Network Parameter-optimization of Gaussian Pre-marginalized Directed Acyclic Graphs
Mehrzad Saremi
Finding the parameters of a latent variable causal model is central to causal inference and causal identification. In this article, we show that exist...
Flexible Functional Treatment Effect Estimation
Jiayi Wang, Raymond K. W. Wong, Xiaoke Zhang et al.
We study treatment effect estimation with functional treatments where the average potential outcome functional is a function of functions, in contrast...
Error Analysis for Deep ReLU Feedforward Density-Ratio Estimation with Bregman Divergence
Jian Huang, Siming Zheng, Guohao Shen et al.
We consider the problem of density-ratio estimation using Bregman Divergence with Deep ReLU feedforward neural networks (BDD). We establish non-asympt...
A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design
Zhaoran Wang, Zhuoran Yang, Michael I. Jordan et al.
We study reserve price optimization in multi-phase second price auctions, where the seller's prior actions affect the bidders' later valuations throug...
UQLM: A Python Package for Uncertainty Quantification in Large Language Models
Dylan Bouchard, Mohit Singh Chauhan, David Skarbrevik et al.
Hallucinations, defined as instances where Large Language Models (LLMs) generate false or misleading content, pose a significant challenge that impact...
Nonlinear function-on-function regression by RKHS
Peijun Sang, Bing Li
We propose a nonlinear function-on-function regression model where both the covariate and the response are random functions. The nonlinear regression ...
Nonlocal Techniques for the Analysis of Deep ReLU Neural Network Approximations
Cornelia Schneider, Mario Ullrich, Jan Vybíral
In recent work concerned with the approximation and expressive powers of deep neural networks, Daubechies, DeVore, Foucart, Hanin, and Petrova introdu...
A Data-Augmented Contrastive Learning Approach to Nonparametric Density Estimation
Yuanyuan Lin, Chenghao Li
In this paper, we introduce a data-augmented nonparametric noise contrastive estimation method to density estimation using deep neural networks. By le...
Guaranteed Nonconvex Low-Rank Tensor Estimation via Scaled Gradient Descent
Tong Wu
Tensors, which give a faithful and effective representation to deliver the intrinsic structure of multi-dimensional data, play a crucial role in an in...
skwdro: a library for Wasserstein distributionally robust machine learning
Vincent Florian, Waïss Azizian, Franck Iutzeler et al.
We present skwdro, a Python library for training robust machine learning models. The library is based on distributionally robust optimization using Wa...
Extending Mean-Field Variational Inference via Entropic Regularization: Theory and Computation
Bohan Wu, David M. Blei
Variational inference (VI) has emerged as a popular method for approximate inference for high-dimensional Bayesian models. In this paper, we propose a...
Stochastic Gradient Methods: Bias, Stability and Generalization
Yunwen Lei, Shuang Zeng
Recent developments of stochastic optimization often suggest biased gradient estimators to improve either the robustness, communication efficiency or ...
Classification Under Local Differential Privacy with Model Reversal and Model Averaging
Caihong Qin, Yang Bai
Local differential privacy has become a central topic in data privacy research, offering strong privacy guarantees by perturbing user data at the sour...
Identifying Weight-Variant Latent Causal Models
Mingming Gong, Yuhang Liu, Zhen Zhang et al.
The task of causal representation learning aims to uncover latent higher-level causal variables that affect lower-level observations. Identifying the ...
Efficient frequent directions algorithms for approximate decomposition of matrices and higher-order tensors
Maolin Che, Yimin Wei, Hong Yan
In the framework of the FD (frequent directions) algorithm, we first develop two efficient algorithms for low-rank matrix approximations under the emb...
Online Detection of Changes in Moment--Based Projections: When to Retrain Deep Learners or Update Portfolios?
Ansgar Steland
Training deep learning neural networks often requires massive amounts of computational ressources. We propose to sequentially monitor network predicti...
The surrogate Gibbs-posterior of a corrected stochastic MALA: Towards uncertainty quantification for neural networks
Sebastian Bieringer, Gregor Kasieczka, Maximilian F. Steffen et al.
MALA is a popular gradient-based Markov chain Monte Carlo method to access the Gibbs-posterior distribution. Stochastic MALA (sMALA) scales to large d...
Towards Understanding Gradient Flow Dynamics of Homogeneous Neural Networks Beyond the Origin
Akshay Kumar, Jarvis Haupt
Recent works exploring the training dynamics of homogeneous neural network weights under gradient flow with small initialization have established that...
Optimal Complexity in Byzantine-Robust Distributed Stochastic Optimization with Data Heterogeneity
Jie Peng, Qing Ling, Qiankun Shi et al.
In this paper, we establish tight lower bounds for Byzantine-robust distributed first-order stochastic methods in both strongly convex and non-convex ...
Towards Unified Native Spaces in Kernel Methods
Xavier Emery, Emilio Porcu, Moreno Bevilacqua
There exists a plethora of parametric models for positive definite kernels in Euclidean spaces, and their use is ubiquitous in statistics, machine lea...
TorchCP: A Python Library for Conformal Prediction
Jianguo Huang, Jianqing Song, Xuanning Zhou et al.
Conformal prediction (CP) is a powerful statistical framework that generates prediction intervals or sets with guaranteed coverage probability. While ...
Hopfield-Fenchel-Young Networks: A Unified Framework for Associative Memory Retrieval
Saul Santos, Vlad Niculae, Daniel McNamee et al.
Associative memory models, such as Hopfield networks and their modern variants, have garnered renewed interest due to advancements in memory capacity ...
Identifiability of Causal Graphs under Non-Additive Conditionally Parametric Causal Models
Juraj Bodik, Valérie Chavez-Demoulin
Existing approaches to causal discovery often rely on restrictive modeling assumptions that limit their applicability in real-world settings, particul...
Fundamental Limits of Membership Inference Attacks on Machine Learning Models
Elisabeth Gassiat, Eric Aubinais, Pablo Piantanida
Membership inference attacks (MIA) can reveal whether a particular data point was part of the training dataset, potentially exposing sensitive informa...
On the Robustness of Kernel Goodness-of-Fit Tests
François-Xavier Briol, Xing Liu
Goodness-of-fit testing is often criticized for its lack of practical relevance: since "all models are wrong", the null hypothesis that the data confo...
Efficient Online Prediction for High-Dimensional Time Series via Joint Tensor Tucker Decomposition
Defeng Sun, Zhenting Luan, Haoning Wang et al.
Real-time prediction plays a vital role in various control systems, such as traffic congestion control and wireless channel resource allocation. In th...
Fast Computation of Superquantile-Constrained Optimization Through Implicit Scenario Reduction
Ying Cui, Jake Roth
Superquantiles have recently gained significant interest as a risk-aware metric for addressing fairness and distribution shifts in statistical learnin...
Collaborative likelihood-ratio estimation over graphs
Nicolas Vayatis, Alejandro de la Concha, Argyris Kalogeratos
This paper introduces the Collaborative Likelihood-ratio Estimation problem, which is relevant for applications involving multiple statistical estimat...
On the Utility of Equal Batch Sizes for Inference in Stochastic Gradient Descent
Dootika Vats, Rahul Singh, Abhinek Shukla
Stochastic gradient descent (SGD) is an estimation tool for large data employed in machine learning and statistics. Due to the Markovian nature of the...
Differentially Private Bootstrap: New Privacy Analysis and Inference Strategies
Jordan Awan, Zhanyu Wang, Guang Cheng
Differentially private (DP) mechanisms protect individual-level information by introducing randomness into the statistical analysis procedure. Despite...
Convergence and Sample Complexity of Natural Policy Gradient Primal-Dual Methods for Constrained MDPs
Dongsheng Ding, Kaiqing Zhang, Jiali Duan et al.
We study the sequential decision making problem of maximizing the expected total reward while satisfying a constraint on the expected total utility. ...
Differentially Private Multivariate Medians
Kelly Ramsay, Aukosh Jagannath, Shoja'eddin Chenouri
Statistical tools which satisfy rigorous privacy guarantees are necessary for modern data analysis. It is well-known that robustness against contamina...
VFOSA: Variance-Reduced Fast Operator Splitting Algorithms for Generalized Equations
Quoc Tran-Dinh
We develop two Variance-reduced Fast Operator Splitting Algorithms (VFOSA) to approximate solutions for a class of generalized equations, covering fun...
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
Tenghui Li, Guoxu Zhou, Xuyang Zhao et al.
Large language models have demonstrated predictable scaling behaviors with respect to model parameters and training data. This study investigates whe...
Minimax Optimal Two-Sample Testing under Local Differential Privacy
Ilmun Kim, Jongmin Mun, Seungwoo Kwak
We explore the trade-off between privacy and statistical utility in private two-sample testing under local differential privacy (LDP) for both multino...
Jackpot: Approximating Uncertainty Domains with Adversarial Manifolds
Nathanaël Munier, Emmanuel Soubies, Pierre Weiss
Given a forward mapping Φ : R^N → R^M and a point x* ∈ R^N , the region {x ∈ R^N , ||Φ(x) − Φ(x*)|| ≤ ε}, where ε ≥ 0 is a perturbation amplitude, rep...
An Asymptotically Optimal Coordinate Descent Algorithm for Learning Bayesian Networks from Gaussian Models
Tong Xu, Armeen Taeb, Simge Küçükyavuz et al.
This paper studies the problem of learning Bayesian networks from continuous observational data, generated according to a linear Gaussian structural e...
Convergence Rates for Non-Log-Concave Sampling and Log-Partition Estimation
Francis Bach, David Holzmüller
Sampling from Gibbs distributions and computing their log-partition function are fundamental tasks in statistics, machine learning, and statistical ph...
A Unified Framework to Enforce, Discover, and Promote Symmetry in Machine Learning
Samuel E. Otto, Nicholas Zolman, J. Nathan Kutz et al.
Symmetry is present throughout nature and continues to play an increasingly central role in machine learning. In this paper, we provide a unifying the...
Infinite-dimensional Mahalanobis Distance with Applications to Kernelized Novelty Detection
Nikita Zozoulenko, Thomas Cass, Lukas Gonon
The Mahalanobis distance is a classical tool used to measure the covariance-adjusted distance between points in $\mathbb{R}^d$. In this work, we exten...
Stable learning using spiking neural networks equipped with affine encoders and decoders
A. Martina Neuman, Dominik Dold, Philipp Christian Petersen
We study the learning problem associated with spiking neural networks. Specifically, we focus on spiking neural networks composed of simple spiking ne...
Efficient Knowledge Deletion from Trained Models Through Layer-wise Partial Machine Unlearning
Vinay Chakravarthi Gogineni, Esmaeil S. Nadimi
Machine unlearning has garnered significant attention due to its ability to selectively erase knowledge obtained from specific training data samples i...
General Loss Functions Lead to (Approximate) Interpolation in High Dimensions
Kuo-Wei Lai, Vidya Muthukumar
We provide a unified framework that applies to a general family of convex losses across binary and multiclass settings in the overparameterized regime...
Piecewise deterministic sampling with splitting schemes
Andrea Bertazzi, Paul Dobson, Pierre Monmarché
We introduce Markov chain Monte Carlo (MCMC) algorithms based on numerical approximations of piecewise-deterministic Markov processes obtained with th...
Hierarchical and Stochastic Crystallization Learning: Geometrically Leveraged Nonparametric Regression with Delaunay Triangulation
Guosheng Yin, Jiaqi Gu
High-dimensionality is known to be the bottleneck for both nonparametric regression and the Delaunay triangulation. To efficiently exploit the advanta...
Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2
Yuri Chervonyi, Trieu H. Trinh, Miroslav Olšák et al.
We present AlphaGeometry2, a significantly improved version of AlphaGeometry introduced in Nature, 625 (7995):476, 2024, which has now surpassed an av...
Decentralized Bilevel Optimization: A Perspective from Transient Iteration Complexity
Xinmeng Huang, Kun Yuan, Boao Kong et al.
Stochastic bilevel optimization (SBO) is becoming increasingly essential in machine learning due to its versatility in handling nested structures. To ...
Fair Text Classification via Transferable Representations
Thibaud Leteno, Michael Perrot, Charlotte Laclau et al.
Group fairness is a central research topic in text classification, where reaching fair treatment between sensitive groups (e.g., women and men) remain...
Stochastic Interior-Point Methods for Smooth Conic Optimization with Applications
Chuan He, Zhanwang Deng
Conic optimization plays a crucial role in many machine learning (ML) problems. However, practical algorithms for conic constrained ML problems with l...
Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration
Kun Yuan, Tao Sun, Xinwang Liu
Gradient clipping has long been considered essential for ensuring the convergence of Stochastic Gradient Descent (SGD) in the presence of heavy-tailed...
Generalized multi-view model: Adaptive density estimation under low-rank constraints
Julien Chhor, Olga Klopp, Alexandre B. Tsybakov
We study the problem of bivariate discrete or continuous probability density estimation under low-rank constraints. For discrete distributions, we ass...
(De)-regularized Maximum Mean Discrepancy Gradient Flow
Arthur Gretton, Zonghao Chen, Aratrika Mustafi et al.
We introduce a (de)-regularization of the Maximum Mean Discrepancy (DrMMD) and its Wasserstein gradient flow. Existing gradient flows that transport s...
On Probabilistic Embeddings in Optimal Dimension Reduction
Ryan Murray, Adam Pickarski
Dimension reduction algorithms are essential in data science for tasks such as data exploration, feature selection, and denoising. However, many non-l...
Physics Informed Kolmogorov-Arnold Neural Networks for Dynamical Analysis via Efficient-KAN and WAV-KAN
Subhajit Patra, Sonali Panda, Bikram Keshari Parida et al.
Physics-informed neural networks have proven to be a powerful tool for solving differential equations, leveraging the principles of physics to inform ...
Graph-accelerated Markov Chain Monte Carlo using Approximate Samples
Leo L. Duan, Anirban Bhattacharya
It has become increasingly easy nowadays to collect approximate posterior samples via fast algorithms such as variational Bayes, but concerns exist ab...
Online Quantile Regression
Dong Xia, Wen-Xin Zhou, Yinan Shen
This paper addresses the challenge of integrating sequentially arriving data into the quantile regression framework, where the number of features may ...
Statistical Inference of Random Graphs With a Surrogate Likelihood Function
Fangzheng Xie, Dingbo Wu
Spectral estimators have been broadly applied to statistical network analysis, but they do not incorporate the likelihood information of the network s...
On the Representation of Pairwise Causal Background Knowledge and Its Applications in Causal Inference
Zhuangyan Fang, Ruiqi Zhao, Yue Liu et al.
Pairwise causal background knowledge about the existence or absence of causal edges and paths is frequently encountered in observational studies. Such...
An Augmentation Overlap Theory of Contrastive Learning
Qi Zhang, Yifei Wang, Yisen Wang
Recently, self-supervised contrastive learning has achieved great success on various tasks. However, its underlying working mechanism is yet unclear. ...
Algorithms for ridge estimation with convergence guarantees
Wanli Qiao, Wolfgang Polonik
The extraction of filamentary structure from a point cloud is discussed. The filaments are modeled as ridge lines or higher dimensional ridges of an u...
Talent: A Tabular Analytics and Learning Toolbox
Si-Yang Liu, Hao-Run Cai, Qi-Le Zhou et al.
Tabular data is a prevalent source in machine learning. While classical methods have proven effective, deep learning methods for tabular data are emer...
Inferring Change Points in High-Dimensional Regression via Approximate Message Passing
Gabriel Arpino, Xiaoqi Liu, Julia Gontarek et al.
We consider the problem of localizing change points in a generalized linear model (GLM), a model that covers many widely studied problems in statistic...
Universality of Kernel Random Matrices and Kernel Regression in the Quadratic Regime
Parthe Pandit, Zhichao Wang, Yizhe Zhu
Kernel ridge regression (KRR) is a popular class of machine learning models that has become an important tool for understanding deep learning. Much o...
Lexicographic Lipschitz Bandits: New Algorithms and a Lower Bound
Lijun Zhang, Bo Xue, Ji Cheng et al.
This paper studies a multiobjective bandit problem under lexicographic ordering, wherein the learner aims to maximize $m$ objectives, each with differ...
On the Natural Gradient of the Evidence Lower Bound
Nihat Ay, Jesse van Oostrum, Adwait Datar
This article studies the Fisher-Rao gradient, also referred to as the natural gradient, of the evidence lower bound (ELBO) which plays a central role ...
Geometry and Stability of Supervised Learning Problems
Facundo Mémoli, Brantley Vose, Robert C. Williamson
We introduce a notion of distance between supervised learning problems, which we call the Risk distance. This distance, inspired by optimal transport,...
Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination
Xiao Li, Peng Wang, Can Yaras et al.
Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remains an op...
Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions
Qian Lin, Haobo Zhang, Yicheng Li et al.
Motivated by studies of neural networks, particularly the neural tangent kernel theory, we investigate the large-dimensional behavior of kernel ridge ...
A Hybrid Weighted Nearest Neighbour Classifier for Semi-Supervised Learning
Stephen M. S. Lee, Mehdi Soleymani
We propose a novel hybrid procedure for constructing a randomly weighted nearest neighbour classifier for semi-supervised learning. The procedure firs...
Scalable and Adaptive Variational Bayes Methods for Hawkes Processes
Judith Rousseau, Vincent Rivoirard, Deborah Sulem
Hawkes processes are often applied to model dependence and interaction phenomena in multivariate event data sets, such as neuronal spike trains, socia...
Biological Sequence Kernels with Guaranteed Flexibility
Alan N. Amin, Debora S. Marks, Eli N. Weinstein
Applying machine learning to biological sequences---DNA, RNA and protein---has enormous potential to advance human health and environmental sustainabi...
Unified Discrete Diffusion for Categorical Data
Lingxiao Zhao, Xueying Ding, Lijun Yu et al.
Discrete diffusion models have attracted significant attention for their application to naturally discrete data, such as language and graphs. While di...
Reinforcement Learning for Infinite-Dimensional Systems
Wei Zhang, Jr-Shin Li
Interest in reinforcement learning (RL) for large-scale systems, comprising extensive populations of intelligent agents interacting with heterogeneous...
Deep Neural Networks are Adaptive to Function Regularity and Data Distribution in Approximation and Estimation
Hao Liu, Jiahui Cheng, Wenjing Liao
Deep learning has exhibited remarkable results across diverse areas. To understand its success, substantial research has been directed towards its the...
Generation of Geodesics with Actor-Critic Reinforcement Learning to Predict Midpoints
Kazumi Kasaura
To find the shortest paths for all pairs on manifolds with infinitesimally defined metrics, we introduce a framework to generate them by predicting mi...
Learning-to-Optimize with PAC-Bayesian Guarantees: Theoretical Considerations and Practical Implementation
Michael Sucker, Jalal Fadili, Peter Ochs
We use the PAC-Bayesian theory for the setting of learning-to-optimize. To the best of our knowledge, we present the first framework to learn optimiza...
Sparse Semiparametric Discriminant Analysis for High-dimensional Zero-inflated Data
Yang Ni, Hee Cheol Chung, Irina Gaynanova
Sequencing-based technologies provide an abundance of high-dimensional biological data sets with highly skewed and zero-inflated measurements. Despite...
Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
Michael Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden
A class of generative models that unifies flow-based and diffusion-based methods is introduced. These models extend the framework proposed in Albergo ...
Efficient Methods for Non-stationary Online Learning
Lijun Zhang, Peng Zhao, Yan-Feng Xie et al.
Non-stationary online learning has drawn much attention in recent years. In particular, dynamic regret and adaptive regret are proposed as two princip...
Decentralized Asynchronous Optimization with DADAO allows Decoupling and Acceleration
Adel Nabli, Edouard Oyallon
DADAO is the first decentralized, accelerated, asynchronous, primal, first-order algorithm to minimize a sum of $L$-smooth and $\mu$-strongly convex ...
Mixtures of Gaussian Process Experts with SMC^2
Teemu Härkönen, Sara Wade, Kody Law et al.
Gaussian processes are a key component of many flexible statistical and machine learning models. However, they exhibit cubic computational complexity ...
Robust Point Matching with Distance Profiles
YoonHaeng Hur, Yuehaw Khoo
Computational difficulty of quadratic matching and the Gromov-Wasserstein distance has led to various approximation and relaxation schemes. One of suc...
BoFire: Bayesian Optimization Framework Intended for Real Experiments
Johannes P. Dürholt, Thomas S. Asche, Johanna Kleinekorte et al.
Our open-source Python package BoFire combines Bayesian Optimization (BO) with other design of experiments (DoE) strategies focusing on developing and...
Reliever: Relieving the Burden of Costly Model Fits for Changepoint Detection
Guanghui Wang, Chengde Qian, Changliang Zou
Changepoint detection typically relies on a grid-search strategy for optimal data segmentation. When model fitting itself is expensive, repeatedly fit...
Variational Inference for Uncertainty Quantification: an Analysis of Trade-offs
Charles C. Margossian, Loucas Pillaud-Vivien, Lawrence K. Saul
Given an intractable distribution $p$, the problem of variational inference (VI) is to find the best approximation from some more tractable family $Q$...
Are Ensembles Getting Better All the Time?
Pierre-Alexandre Mattei, Damien Garreau
Ensemble methods combine the predictions of several base models. We study whether or not including more models always improves their average performan...
An Adaptive Parameter-free and Projection-free Restarting Level Set Method for Constrained Convex Optimization Under the Error Bound Condition
Qihang Lin, Negar Soheili, Runchao Ma et al.
Recent efforts to accelerate first-order methods have focused on convex optimization problems that satisfy a geometric property known as error-bound c...
Operator Learning for Hyperbolic PDEs
Christopher Wang, Alex Townsend
We construct the first rigorously justified probabilistic algorithm for recovering the solution operator of a hyperbolic partial differential equation...
Optimal subsampling for high-dimensional partially linear models via machine learning methods
Lei Wang, Heng Lian, Yujing Shao et al.
In this paper, we explore optimal subsampling strategies for estimating the parametric regression coefficients in partially linear models with unknown...
Decentralized Sparse Linear Regression via Gradient-Tracking
Ying Sun, Guang Cheng, Marie Maros et al.
We study sparse linear regression over a network of agents, modeled as an undirected graph without a center node. The estimation of the $s$-sparse ...
Calibrated Inference: Statistical Inference that Accounts for Both Sampling Uncertainty and Distributional Uncertainty
Yujin Jeong, Dominik Rothenhäusler
How can we draw trustworthy scientific conclusions? One criterion is that a study can be replicated by independent teams. While replication is critica...
Relaxed Gaussian Process Interpolation: a Goal-Oriented Approach to Bayesian Optimization
Sébastien J. Petit, Julien Bect, Emmanuel Vazquez
This work presents a new procedure for obtaining predictive distributions in the context of Gaussian process (GP) modeling, with a relaxation of the i...
"What is Different Between These Datasets?" A Framework for Explaining Data Distribution Shifts
Varun Babbar*, Zhicheng Guo*, Cynthia Rudin
The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-relate...
Linear Separation Capacity of Self-Supervised Representation Learning
Shulei Wang
Recent advances in self-supervised learning have highlighted the efficacy of data augmentation in learning data representation from unlabeled data. Tr...
On the Convergence of Projected Policy Gradient for Any Constant Step Sizes
Zhihua Zhang, Jiacai Liu, Wenye Li et al.
Projected policy gradient (PPG) is a basic policy optimization method in reinforcement learning. Given access to exact policy evaluations, previous s...
Learning with Linear Function Approximations in Mean-Field Control
Erhan Bayraktar, Ali Devran Kara
The paper focuses on mean-field type multi-agent control problems with finite state and action spaces where the dynamics and cost structures are symme...
A New Random Reshuffling Method for Nonsmooth Nonconvex Finite-sum Optimization
Junwen Qiu, Xiao Li, Andre Milzarek
Random reshuffling techniques are prevalent in large-scale applications, such as training neural networks. While the convergence and acceleration effe...
Model-free Change-Point Detection Using AUC of a Classifier
Feiyu Jiang, Rohit Kanrar, Zhanrui Cai
In contemporary data analysis, it is increasingly common to work with non-stationary complex data sets. These data sets typically extend beyond the cl...
EF21 with Bells & Whistles: Six Algorithmic Extensions of Modern Error Feedback
Ilyas Fatkhullin, Igor Sokolov, Eduard Gorbunov et al.
First proposed by Seide (2014) as a heuristic, error feedback (EF) is a very popular mechanism for enforcing convergence of distributed gradient-based...
Multiple Instance Verification
Xin Xu, Eibe Frank, Geoffrey Holmes
We explore multiple instance verification, a problem setting in which a query instance is verified against a bag of target instances with heterogeneou...
Learning from Similar Linear Representations: Adaptivity, Minimaxity, and Robustness
Yang Feng, Yuqi Gu, Ye Tian
Representation multi-task learning (MTL) has achieved tremendous success in practice. However, the theoretical understanding of these methods is still...
Exponential Family Graphical Models: Correlated Replicates and Unmeasured Confounders, with Applications to fMRI Data
Kean Ming Tan, Yang Ning, Yanxin Jin
Graphical models have been used extensively for modeling brain connectivity networks. However, unmeasured confounders and correlations among measureme...
Optimizing Return Distributions with Distributional Dynamic Programming
Bernardo Ávila Pires, Mark Rowland, Diana Borsa et al.
We introduce distributional dynamic programming (DP) methods for optimizing statistical functionals of the return distribution, with standard reinforc...
Imprecise Multi-Armed Bandits: Representing Irreducible Uncertainty as a Zero-Sum Game
Vanessa Kosoy
We introduce a novel multi-armed bandit framework, where each arm is associated with a fixed unknown credal set over the space of outcomes (which can ...
Early Alignment in Two-Layer Networks Training is a Two-Edged Sword
Etienne Boursier, Nicolas Flammarion
Training neural networks with first order optimisation methods is at the core of the empirical success of deep learning. The scale of initialisation i...
Hierarchical Decision Making Based on Structural Information Principles
Xianghua Zeng, Hao Peng, Dingli Su et al.
Hierarchical Reinforcement Learning (HRL) is a promising approach for managing task complexity across multiple levels of abstraction and accelerating ...
Generative Adversarial Networks: Dynamics
Matias G. Delgadino, Bruno B. Suassuna, Rene Cabrera
We study quantitatively the overparametrization limit of the original Wasserstein-GAN algorithm. Effectively, we show that the algorithm is a stochast...
“What is Different Between These Datasets?” A Framework for Explaining Data Distribution Shifts
Varun Babbar*, Zhicheng Guo*, Cynthia Rudin
The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-relate...
Assumption-lean and data-adaptive post-prediction inference
Jiacheng Miao, Xinran Miao, Yixuan Wu et al.
A primary challenge facing modern scientific research is the limited availability of gold-standard data, which can be costly, labor-intensive, or inva...
Bagged Regularized k-Distances for Anomaly Detection
Hanyuan Hang, Hanfang Yang, Yuchao Cai et al.
We consider the paradigm of unsupervised anomaly detection, which involves the identification of anomalies within a dataset in the absence of labeled ...
Four Axiomatic Characterizations of the Integrated Gradients Attribution Method
Daniel Lundstrom, Meisam Razaviyayn
Deep neural networks have produced significant progress among machine learning models in terms of accuracy and functionality, but their inner workings...
Fast Algorithm for Constrained Linear Inverse Problems
Mohammed Rayyan Sheriff, Floor Fenne Redel, Peyman Mohajerin Esfahani
We consider the constrained Linear Inverse Problem (LIP), where a certain atomic norm (like the $\ell_1 $ norm) is minimized subject to a quadratic co...
High-Rank Irreducible Cartesian Tensor Decomposition and Bases of Equivariant Spaces
Shihao Shao, Yikang Li, Zhouchen Lin et al.
Irreducible Cartesian tensors (ICTs) play a crucial role in the design of equivariant graph neural networks, as well as in theoretical chemistry and c...
Best Linear Unbiased Estimate from Privatized Contingency Tables
Jordan Awan, Adam Edwards, Paul Bartholomew et al.
In differential privacy (DP) mechanisms, it can be beneficial to release "redundant" outputs, where some quantities can be estimated in multiple ways...
Interpretable Global Minima of Deep ReLU Neural Networks on Sequentially Separable Data
Thomas Chen, Patrícia Muñoz Ewald
We explicitly construct zero loss neural network classifiers. We write the weight matrices and bias vectors in terms of cumulative parameters, which ...
Enhanced Feature Learning via Regularisation: Integrating Neural Networks and Kernel Methods
Bertille FOLLAIN, Francis BACH
We propose a new method for feature learning and function estimation in supervised learning via regularised empirical risk minimisation. Our approach ...
Data-Driven Performance Guarantees for Classical and Learned Optimizers
Rajiv Sambharya, Bartolomeo Stellato
We introduce a data-driven approach to analyze the performance of continuous optimization algorithms using generalization guarantees from statistical ...
Contextual Bandits with Stage-wise Constraints
Aldo Pacchiano, Mohammad Ghavamzadeh, Peter Bartlett
We study contextual bandits in the presence of a stage-wise constraint when the constraint must be satisfied both with high probability and in expecta...
Boosting Causal Additive Models
Maximilian Kertel, Nadja Klein
We present a boosting-based method to learn additive Structural Equation Models (SEMs) from observational data, with a focus on the theoretical aspect...
Frequentist Guarantees of Distributed (Non)-Bayesian Inference
Bohan Wu, César A. Uribe
We establish frequentist properties, i.e., posterior consistency, asymptotic normality, and posterior contraction rates, for the distributed (non-)Bay...
Asymptotic Inference for Multi-Stage Stationary Treatment Policy with Variable Selection
Donglin Zeng, Yufeng Liu, Daiqi Gao
Dynamic treatment regimes or policies are a sequence of decision functions over multiple stages that are tailored to individual features. One importan...
EMaP: Explainable AI with Manifold-based Perturbations
Minh Nhat Vu, Huy Quang Mai, My T. Thai
In the last few years, many explanation methods based on the perturbations of input data have been introduced to shed light on the predictions generat...
Autoencoders in Function Space
Justin Bunker, Mark Girolami, Hefin Lambley et al.
Autoencoders have found widespread application in both their original deterministic form and in their variational formulation (VAEs). In scientific ap...
Nonparametric Regression on Random Geometric Graphs Sampled from Submanifolds
Paul Rosa, Judith Rousseau
We consider the nonparametric regression problem when the covariates are located on an unknown compact submanifold of a Euclidean space. Under definin...
System Neural Diversity: Measuring Behavioral Heterogeneity in Multi-Agent Learning
Matteo Bettini, Ajay Shankar, Amanda Prorok
Evolutionary science provides evidence that diversity confers resilience in natural systems. Yet, traditional multi-agent reinforcement learning techn...
Distribution Estimation under the Infinity Norm
Aryeh Kontorovich, Amichai Painsky
We present novel bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These are nearly optimal in various precise se...
Extending Temperature Scaling with Homogenizing Maps
Christopher Qian, Feng Liang, Jason Adams
As machine learning models continue to grow more complex, poor calibration significantly limits the reliability of their predictions. Temperature scal...
Density Estimation Using the Perceptron
Yury Polyanskiy, Patrik Róbert Gerber, Tianze Jiang et al.
We propose a new density estimation algorithm. Given $n$ i.i.d. observations from a distribution belonging to a class of densities on $\mathbb{R}^d$...
Simplex Constrained Sparse Optimization via Tail Screening
Xueqin Wang, Peng Chen, Jin Zhu et al.
We consider the probabilistic simplex-constrained sparse recovery problem. The commonly used Lasso-type penalty for promoting sparsity is ineffective ...
Score-Based Diffusion Models in Function Space
Jae Hyun Lim, Nikola B. Kovachki, Ricardo Baptista et al.
Diffusion models have recently emerged as a powerful framework for generative modeling. They consist of a forward process that perturbs input data wit...
Regularized Rényi Divergence Minimization through Bregman Proximal Gradient Algorithms
Thomas Guilmeau, Emilie Chouzenoux, Víctor Elvira
We study the variational inference problem of minimizing a regularized Rényi divergence over an exponential family. We propose to solve this problem w...
WEFE: A Python Library for Measuring and Mitigating Bias in Word Embeddings
Pablo Badilla, Felipe Bravo-Marquez, María José Zambrano et al.
Word embeddings, which are a mapping of words into continuous vectors, are widely used in modern Natural Language Processing (NLP) systems. However, t...
Frontiers to the learning of nonparametric hidden Markov models
Elisabeth Gassiat, Zacharie Naulet, Kweku Abraham
Hidden Markov models (HMMs) are flexible tools for clustering dependent data coming from unknown populations, allowing nonparametric modelling of the ...
On Non-asymptotic Theory of Recurrent Neural Networks in Temporal Point Processes
Zhiheng Chen, Guanhua Fang, Wen Yu
Temporal point process (TPP) is an important tool for modeling and predicting irregularly timed events across various domains. Recently, the recurrent...
Classification in the high dimensional Anisotropic mixture framework: A new take on Robust Interpolation
Stanislav Minsker, Mohamed Ndaoud, Yiqiu Shen
We study the classification problem under the two-component anisotropic sub-Gaussian mixture model in high dimensions and in the non-asymptotic settin...
Universal Online Convex Optimization Meets Second-order Bounds
Yibo Wang, Lijun Zhang, Guanghui Wang et al.
Recently, several universal methods have been proposed for online convex optimization, and attain minimax rates for multiple types of convex function...
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens
Amirreza Neshaei Moghaddam, Alex Olshevsky, Bahman Gharesifard
We provide the first known algorithm that provably achieves $\varepsilon$-optimality within $\widetilde{O}(1/\varepsilon)$ function evaluations for th...
Randomization Can Reduce Both Bias and Variance: A Case Study in Random Forests
Rahul Mazumder, Brian Liu
We study the often overlooked phenomenon, first noted in Breiman (2001), that random forests appear to reduce bias compared to bagging. Motivated by a...
skglm: Improving scikit-learn for Regularized Generalized Linear Models
Badr Moufad, Pierre-Antoine Bannier, Quentin Bertrand et al.
We introduce skglm, an open-source Python package for regularized Generalized Linear Models. Thanks to its composable nature, it supports combining da...
Losing Momentum in Continuous-time Stochastic Optimisation
Kexin Jin, Jonas Latz, Chenguang Liu et al.
The training of modern machine learning models often consists in solving high-dimensional non-convex optimisation problems that are subject to large-s...
Latent Process Models for Functional Network Data
Elizaveta Levina, Ji Zhu, Peter W. MacDonald
Network data are often sampled with auxiliary information or collected through the observation of a complex system over time, leading to multiple netw...
Dynamic Bayesian Learning for Spatiotemporal Mechanistic Models
Sudipto Banerjee, Xiang Chen, Ian Frankenburg et al.
We develop an approach for Bayesian learning of spatiotemporal dynamical mechanistic models. Such learning consists of statistical emulation of the me...
On the Ability of Deep Networks to Learn Symmetries from Data: A Neural Kernel Theory
Andrea Perin, Stephane Deny
Symmetries (transformations by group actions) are present in many datasets, and leveraging them holds considerable promise for improving predictions i...
Fine-grained Analysis and Faster Algorithms for Iteratively Solving Linear Systems
Michal Dereziński, Daniel LeJeune, Deanna Needell et al.
Despite being a key bottleneck in many machine learning tasks, the cost of solving large linear systems has proven challenging to quantify due to prob...
Deep Generative Models: Complexity, Dimensionality, and Approximation
Didong Li, Kevin Wang, Hongqian Niu et al.
Generative networks have shown remarkable success in learning complex data distributions, particularly in generating high-dimensional data from lower-...
ClimSim-Online: A Large Multi-Scale Dataset and Framework for Hybrid Physics-ML Climate Emulation
Sungduk Yu, Zeyuan Hu, Akshay Subramaniam et al.
Modern climate projections lack adequate spatial and temporal resolution due to computational constraints, leading to inaccuracies in representing cri...
Conditional Wasserstein Distances with Applications in Bayesian OT Flow Matching
Jannis Chemseddine, Paul Hagemann, Gabriele Steidl et al.
In inverse problems, many conditional generative models approximate the posterior measure by minimizing a distance between the joint measure and its l...
Deep Variational Multivariate Information Bottleneck - A Framework for Variational Losses
Eslam Abdelaleem, Ilya Nemenman, K. Michael Martini
Variational dimensionality reduction methods are widely used for their accuracy, generative capabilities, and robustness. We introduce a unifying fram...
Diffeomorphism-based feature learning using Poincaré inequalities on augmented input space
Romain Verdière, Clémentine Prieur, Olivier Zahm
We propose a gradient-enhanced algorithm for high-dimensional function approximation. The algorithm proceeds in two steps: firstly, we reduce the inp...
Finite Expression Method for Solving High-Dimensional Partial Differential Equations
Senwei Liang, Haizhao Yang
Designing efficient and accurate numerical solvers for high-dimensional partial differential equations (PDEs) remains a challenging and important topi...
Randomly Projected Convex Clustering Model: Motivation, Realization, and Cluster Recovery Guarantees
Defeng Sun, Yancheng Yuan, Ziwen Wang et al.
In this paper, we propose a randomly projected convex clustering model for clustering a collection of $n$ high dimensional data points in $\mathbb{R}^...
Minimax Optimal Deep Neural Network Classifiers Under Smooth Decision Boundary
Zuofeng Shang, Tianyang Hu, Ruiqi Liu et al.
Deep learning has gained huge empirical successes in large-scale classification problems. In contrast, there is a lack of statistical understanding ab...
Optimal and Efficient Algorithms for Decentralized Online Convex Optimization
Lijun Zhang, Yuanyu Wan, Tong Wei et al.
We investigate decentralized online convex optimization (D-OCO), in which a set of local learners are required to minimize a sequence of global loss f...
Characterizing Dynamical Stability of Stochastic Gradient Descent in Overparameterized Learning
Dennis Chemnitz, Maximilian Engel
For overparameterized optimization tasks, such as those found in modern machine learning, global minima are generally not unique. In order to understa...
PREMAP: A Unifying PREiMage APproximation Framework for Neural Networks
Xiyue Zhang, Benjie Wang, Marta Kwiatkowska et al.
Most methods for neural network verification focus on bounding the image, i.e., set of outputs for a given input set. This can be used to, for example...
Score-Aware Policy-Gradient and Performance Guarantees using Local Lyapunov Stability
Céline Comte, Matthieu Jonckheere, Jaron Sanders et al.
In this paper, we introduce a policy-gradient method for model-based reinforcement learning (RL) that exploits a type of stationary distributions comm...
On the O(sqrt(d)/T^(1/4)) Convergence Rate of RMSProp and Its Momentum Extension Measured by l_1 Norm
Zhouchen Lin, Huan Li, Yiming Dong
Although adaptive gradient methods have been extensively used in deep learning, their convergence rates proved in the literature are all slower than t...
Categorical Semantics of Compositional Reinforcement Learning
Georgios Bakirtzis, Michail Savvas, Ufuk Topcu
Compositional knowledge representations in reinforcement learning (RL) facilitate modular, interpretable, and safe task specifications. However, gener...
Transformers from Diffusion: A Unified Framework for Neural Message Passing
David Wipf, Qitian Wu, Junchi Yan
Learning representations for structured data with certain geometries (e.g., observed or unobserved) is a fundamental challenge, wherein message passin...
Optimal Sample Selection Through Uncertainty Estimation and Its Application in Deep Learning
Yong Lin, Chen Liu, Chenlu Ye et al.
Modern deep learning heavily relies on large labeled datasets, which often comse with high costs in terms of both manual labeling and computational re...
Actor-Critic learning for mean-field control in continuous time
Noufel FRIKHA, Maximilien GERMAIN, Mathieu LAURIERE et al.
We study policy gradient for mean-field control in continuous time in a reinforcement learning setting. By considering randomised policies with entro...
Modelling Populations of Interaction Networks via Distance Metrics
George Bolt, Simón Lunagómez, Christopher Nemeth
Network data arises through the observation of relational information between a collection of entities, for example, friendships (relations) amongst a...
BitNet: 1-bit Pre-training for Large Language Models
Lei Wang, Yi Wu, Hongyu Wang et al.
The increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high ene...
Physics-informed Kernel Learning
Gérard Biau, Nathan Doumèche, Francis Bach et al.
Physics-informed machine learning typically integrates physical priors into the learning process by minimizing a loss function that includes both a da...
Last-iterate Convergence of Shuffling Momentum Gradient Method under the Kurdyka-Lojasiewicz Inequality
Yuqing Liang, Dongpo Xu
Shuffling gradient algorithms are extensively used to solve finite-sum optimization problems in machine learning. However, their theoretical propertie...
Posterior and Variational Inference for Deep Neural Networks with Heavy-Tailed Weights
Ismaël Castillo, Paul Egels
We consider deep neural networks in a Bayesian framework with a prior distribution sampling the network weights at random. Following a recent idea of...
Maximum Causal Entropy IRL in Mean-Field Games and GNEP Framework for Forward RL
Berkay Anahtarci, Can Deha Kariksiz, Naci Saldi
This paper explores the use of Maximum Causal Entropy Inverse Reinforcement Learning (IRL) within the context of discrete-time stationary Mean-Field G...
Degree of Interference: A General Framework For Causal Inference Under Interference
Yuki Ohnishi, Bikram Karmakar, Arman Sabbaghi
One core assumption typically adopted for valid causal inference is that of no interference between experimental units, i.e., the outcome of an experi...
Quantifying the Effectiveness of Linear Preconditioning in Markov Chain Monte Carlo
Max Hird, Samuel Livingstone
We study linear preconditioning in Markov chain Monte Carlo. We consider the class of well-conditioned distributions, for which several mixing time bo...
Sparse SVM with Hard-Margin Loss: a Newton-Augmented Lagrangian Method in Reduced Dimensions
Penghe Zhang, Naihua Xiu, Hou-Duo Qi
The hard-margin loss function has been at the core of the support vector machine research from the very beginning due to its generalization capability...
On Model Identification and Out-of-Sample Prediction of PCR with Applications to Synthetic Controls
Devavrat Shah, Anish Agarwal, Dennis Shen
We analyze principal component regression (PCR) in a high-dimensional error-in-variables setting with fixed design. Under suitable conditions, we show...
Bayesian Scalar-on-Image Regression with a Spatially Varying Single-layer Neural Network Prior
Keru Wu, Jian Kang, Ben Wu
Deep neural networks (DNN) have been widely used in scalar-on-image regression to predict an outcome variable from imaging predictors. However, train...
DRM Revisited: A Complete Error Analysis
Yuling Jiao, Ruoxuan Li, Peiying Wu et al.
It is widely known that the error analysis for deep learning involves approximation, statistical, and optimization errors. However, it is challenging ...
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
Zhuoran Yang, Han Shen, Tianyi Chen
Bilevel optimization has been recently applied to many machine learning tasks. However, their applications have been restricted to the supervised lear...
Precise High-Dimensional Asymptotics for Quantifying Heterogeneous Transfers
Fan Yang, Hongyang R. Zhang, Sen Wu et al.
The problem of learning one task using samples from another task is central to transfer learning. In this paper, we focus on answering the following q...
Score-based Causal Representation Learning: Linear and General Transformations
Burak Var{{\i}}c{{\i}}, Emre Acartürk, Karthikeyan Shanmugam et al.
This paper addresses intervention-based causal representation learning (CRL) under a general nonparametric latent causal model and an unknown transfor...
On the Statistical Properties of Generative Adversarial Models for Low Intrinsic Data Dimension
Saptarshi Chakraborty, Peter L. Bartlett
Despite the remarkable empirical successes of Generative Adversarial Networks (GANs), the theoretical guarantees for their statistical accuracy remain...
Prominent Roles of Conditionally Invariant Components in Domain Adaptation: Theory and Algorithms
Keru Wu, Yuansi Chen, Wooseok Ha et al.
Domain adaptation (DA) is a statistical learning problem that arises when the distribution of the source data used to train a model differs from that ...
Near-Optimal Nonconvex-Strongly-Convex Bilevel Optimization with Fully First-Order Oracles
Lesi Chen, Yaohua Ma, Jingzhao Zhang
In this work, we consider bilevel optimization when the lower-level problem is strongly convex. Recent works show that with a Hessian-vector product (...
Adaptive Distributed Kernel Ridge Regression: A Feasible Distributed Learning Scheme for Data Silos
Shao-Bo Lin, Xiaotong Liu, Di Wang et al.
Data silos, mainly caused by privacy and interoperability, significantly constrain collaborations among different organizations with similar data for ...
On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control
Vincent Roulet, Siddhartha Srinivasa, Maryam Fazel et al.
A classical approach for solving discrete time nonlinear control on a finite horizon consists in repeatedly minimizing linear quadratic approximations...
A Decentralized Proximal Gradient Tracking Algorithm for Composite Optimization on Riemannian Manifolds
Lei Wang, Le Bao, Xin Liu
This paper focuses on minimizing a smooth function combined with a nonsmooth regularization term on a compact Riemannian submanifold embedded in the E...
Learning conditional distributions on continuous spaces
Cyril Benezet, Ziteng Cheng, Sebastian Jaimungal
We investigate sample-based learning of conditional distributions on multi-dimensional unit boxes, allowing for different dimensions of the feature an...
A Unified Analysis of Nonstochastic Delayed Feedback for Combinatorial Semi-Bandits, Linear Bandits, and MDPs
Lukas Zierahn, Dirk van der Hoeven, Tal Lancewicki et al.
We derive a new analysis of Follow The Regularized Leader (FTRL) for online learning with delayed bandit feedback. By separating the cost of delayed f...
Error bounds for particle gradient descent, and extensions of the log-Sobolev and Talagrand inequalities
Rocco Caprio, Juan Kuntz, Samuel Power et al.
We derive non-asymptotic error bounds for particle gradient descent (PGD, Kuntz et al. (2023)), a recently introduced algorithm for maximum likelihoo...
Linear Hypothesis Testing in High-Dimensional Expected Shortfall Regression with Heavy-Tailed Errors
Kean Ming Tan, Wen-Xin Zhou, Gaoyu Wu et al.
Expected shortfall (ES) is widely used for characterizing the tail of a distribution across various fields, particularly in financial risk management....
Efficient Numerical Integration in Reproducing Kernel Hilbert Spaces via Leverage Scores Sampling
Antoine Chatalic, Nicolas Schreuder, Ernesto De Vito et al.
In this work we consider the problem of numerical integration, i.e., approximating integrals with respect to a target probability measure using only p...
Distribution Free Tests for Model Selection Based on Maximum Mean Discrepancy with Estimated Parameters
Florian Brück, Jean-David Fermanian, Aleksey Min
There exist several testing procedures based on the maximum mean discrepancy (MMD) to address the challenge of model specification. However, these tes...
Statistical field theory for Markov decision processes under uncertainty
George Stamatescu
A statistical field theory is introduced for finite state and action Markov decision processes with unknown parameters, in a Bayesian setting. The Bel...
Bayesian Data Sketching for Varying Coefficient Regression Models
Rajarshi Guhaniyogi, Laura Baracaldo, Sudipto Banerjee
Varying coefficient models are popular for estimating nonlinear regression functions in functional data models. Their Bayesian variants have received ...
Bagged k-Distance for Mode-Based Clustering Using the Probability of Localized Level Sets
Hanyuan Hang
In this paper, we propose an ensemble learning algorithm named bagged $k$-distance for mode-based clustering (BDMBC) by putting forward a new measure ...
Linear cost and exponentially convergent approximation of Gaussian Matérn processes on intervals
David Bolin, Vaibhav Mehandiratta, Alexandre B. Simas
The computational cost for inference and prediction of statistical models based on Gaussian processes with Matérn covariance functions scales cubicall...
Invariant Subspace Decomposition
Margherita Lazzaretto, Jonas Peters, Niklas Pfister
We consider the task of predicting a response $Y$ from a set of covariates $X$ in settings where the conditional distribution of $Y$ given $X$ changes...
Posterior Concentrations of Fully-Connected Bayesian Neural Networks with General Priors on the Weights
Insung Kong, Yongdai Kim
Bayesian approaches for training deep neural networks (BNNs) have received significant interest and have been effectively utilized in a wide range of ...
Outlier Robust and Sparse Estimation of Linear Regression Coefficients
Takeyuki Sasai, Hironori Fujisawa
We consider outlier-robust and sparse estimation of linear regression coefficients, when the covariates and the noises are contaminated by adversarial...
Affine Rank Minimization via Asymptotic Log-Det Iteratively Reweighted Least Squares
Sebastian Krämer
The affine rank minimization problem is a well-known approach to matrix recovery. While there are various surrogates to this NP-hard problem, we prove...
Causal Effect of Functional Treatment
Ruoxu Tan, Wei Huang, Zheng Zhang et al.
We study the causal effect with a functional treatment variable, where practical applications often arise in neuroscience, biomedical sciences, etc. P...
Uplift Model Evaluation with Ordinal Dominance Graphs
Brecht Verbeken, Marie-Anne Guerry, Wouter Verbeke et al.
Uplift modelling is a subfield of causal learning that focuses on ranking entities by individual treatment effects. Uplift models are typically evalua...
High-Dimensional L2-Boosting: Rate of Convergence
Ye Luo, Martin Spindler, Jannis Kueck
Boosting is one of the most significant developments in machine learning. This paper studies the rate of convergence of L2-Boosting in a high-dimensio...
Feature Learning in Finite-Width Bayesian Deep Linear Networks with Multiple Outputs and Convolutional Layers
Federico Bassetti, Marco Gherardi, Alessandro Ingrosso et al.
Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite...
How good is your Laplace approximation of the Bayesian posterior? Finite-sample computable error bounds for a variety of useful divergences
Miko{\l}aj J. Kasprzak, Ryan Giordano, Tamara Broderick
The Laplace approximation is a popular method for constructing a Gaussian approximation to the Bayesian posterior and thereby approximating the poster...
Integral Probability Metrics Meet Neural Networks: The Radon-Kolmogorov-Smirnov Test
Alden Green, Seunghoon Paik, Michael Celentano et al.
Integral probability metrics (IPMs) constitute a general class of nonparametric two-sample tests that are based on maximizing the mean difference betw...
On Inference for the Support Vector Machine
Wen-Xin Zhou, Jakub Rybak, Heather Battey
The linear support vector machine has a parametrised decision boundary. The paper considers inference for the corresponding parameters, which indicate...
Random Pruning Over-parameterized Neural Networks Can Improve Generalization: A Training Dynamics Analysis
Hongru Yang, Yingbin Liang, Xiaojie Guo et al.
It has been observed that applying pruning-at-initialization methods and training the sparse networks can sometimes yield slightly better test perform...
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
Atticus Geiger, Duligur Ibeling, Amir Zur et al.
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that...
Implicit vs Unfolded Graph Neural Networks
Yongyi Yang, Tang Liu, Yangkun Wang et al.
It has been observed that message-passing graph neural networks (GNN) sometimes struggle to maintain a healthy balance between the efficient / scalabl...
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
Brendon G. Anderson, Ziye Ma, Jingqi Li et al.
In this paper, we study certifying the robustness of ReLU neural networks against adversarial input perturbations. To diminish the relaxation error su...
GraphNeuralNetworks.jl: Deep Learning on Graphs with Julia
Carlo Lucibello, Aurora Rossi
GraphNeuralNetworks.jl is an open-source framework for deep learning on graphs, written in the Julia programming language. It supports multiple GPU ba...
Dynamic angular synchronization under smoothness constraints
Ernesto Araya, Mihai Cucuringu, Hemant Tyagi
Given an undirected measurement graph $\mathcal{H} = ([n], \mathcal{E})$, the classical angular synchronization problem consists of recovering unkno...
Derivative-Informed Neural Operator Acceleration of Geometric MCMC for Infinite-Dimensional Bayesian Inverse Problems
Lianghao Cao, Thomas O'Leary-Roseberry, Omar Ghattas
We propose an operator learning approach to accelerate geometric Markov chain Monte Carlo (MCMC) for solving infinite-dimensional Bayesian inverse pro...
Wasserstein F-tests for Frechet regression on Bures-Wasserstein manifolds
Hongzhe Li, Haoshu Xu
This paper addresses regression analysis for covariance matrix-valued outcomes with Euclidean covariates, motivated by applications in single-cell gen...
Distributed Stochastic Bilevel Optimization: Improved Complexity and Heterogeneity Analysis
Youcheng Niu, Jinming Xu, Ying Sun et al.
This paper considers solving a class of nonconvex-strongly-convex distributed stochastic bilevel optimization (DSBO) problems with personalized inner-...
Learning causal graphs via nonlinear sufficient dimension reduction
Eftychia Solea, Bing Li, Kyongwon Kim
We introduce a new nonparametric methodology for estimating a directed acyclic graph (DAG) from observational data. Our method is nonparametric in nat...
On Consistent Bayesian Inference from Synthetic Data
Ossi Räisä, Joonas Jälkö, Antti Honkela
Generating synthetic data, with or without differential privacy, has attracted significant attention as a potential solution to the dilemma between ma...
Optimization Over a Probability Simplex
James Chok, Geoffrey M. Vasil
We propose a new iteration scheme, the Cauchy-Simplex, to optimize convex problems over the probability simplex $\{w\in\mathbb{R}^n\ |\ \sum_i w_i=1\ ...
Laplace Meets Moreau: Smooth Approximation to Infimal Convolutions Using Laplace's Method
Ryan J. Tibshirani, Samy Wu Fung, Howard Heaton et al.
We study approximations to the Moreau envelope---and infimal convolutions more broadly---based on Laplace's method, a classical tool in analysis which...
Sampling and Estimation on Manifolds using the Langevin Diffusion
Karthik Bharath, Alexander Lewis, Akash Sharma et al.
Error bounds are derived for sampling and estimation using a discretization of an intrinsically defined Langevin diffusion with invariant measure $\te...
Sharp Bounds for Sequential Federated Learning on Heterogeneous Data
Yipeng Li, Xinchen Lyu
There are two paradigms in Federated Learning (FL): parallel FL (PFL), where models are trained in a parallel manner across clients, and sequential FL...
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
Yaoyu Zhang, Leyang Zhang, Zhongwang Zhang et al.
Determining whether deep neural network (DNN) models can reliably recover target functions at overparameterization is a critical yet complex issue in ...
Stabilizing Sharpness-Aware Minimization Through A Simple Renormalization Strategy
Chengli Tan, Jiangshe Zhang, Junmin Liu et al.
Recently, sharpness-aware minimization (SAM) has attracted much attention because of its surprising effectiveness in improving generalization performa...
Fine-Grained Change Point Detection for Topic Modeling with Pitman-Yor Process
Feifei Wang, Zimeng Zhao, Ruimin Ye et al.
Identifying change points in dynamic text data is crucial for understanding the evolving nature of topics across various sources, such as news article...
Deletion Robust Non-Monotone Submodular Maximization over Matroids
Paul Dütting, Federico Fusco, Silvio Lattanzi et al.
We study the deletion robust version of submodular maximization under matroid constraints. The goal is to extract a small-size summary of the data set...
Instability, Computational Efficiency and Statistical Accuracy
Raaz Dwivedi, Koulik Khamaru, Martin J. Wainwright et al.
Many statistical estimators are defined as the fixed point of a data-dependent operator, with estimators based on minimizing a cost function being an ...
Estimation of Local Geometric Structure on Manifolds from Noisy Data
Yariv Aizenbud, Barak Sober
A common observation in data-driven applications is that high-dimensional data have a low intrinsic dimension, at least locally. In this work, we cons...
Ontolearn---A Framework for Large-scale OWL Class Expression Learning in Python
Caglar Demir, Alkid Baci, N'Dah Jean Kouagou et al.
In this paper, we present Ontolearn---a framework for learning OWL class expressions over large knowledge graphs. Ontolearn contains efficient implem...
Continuously evolving rewards in an open-ended environment
Richard M. Bailey
Unambiguous identification of the rewards driving behaviours of entities operating in complex open-ended real-world environments is difficult, in part...
Recursive Causal Discovery
Ehsan Mokhtarian, Sepehr Elahi, Sina Akbari et al.
Causal discovery from observational data, i.e., learning the causal graph from a finite set of samples from the joint distribution of the variables, i...
Evaluation of Active Feature Acquisition Methods for Time-varying Feature Settings
Ilya Shpitser, Henrik von Kleist, Alireza Zamanian et al.
Machine learning methods often assume that input features are available at no cost. However, in domains like healthcare, where acquiring features coul...
On Adaptive Stochastic Optimization for Streaming Data: A Newton's Method with O(dN) Operations
Antoine Godichon-Baggioni, Nicklas Werge
Stochastic optimization methods face new challenges in the realm of streaming data, characterized by a continuous flow of large, high-dimensional data...
Determine the Number of States in Hidden Markov Models via Marginal Likelihood
Yang Chen, Cheng-Der Fuh, Chu-Lan Michael Kao
Hidden Markov models (HMM) have been widely used by scientists to model stochastic systems: the underlying process is a discrete Markov chain, and the...
Variance-Aware Estimation of Kernel Mean Embedding
Geoffrey Wolfer, Pierre Alquier
An important feature of kernel mean embeddings (KME) is that the rate of convergence of the empirical KME to the true distribution KME can be bounded ...
Scaling ResNets in the Large-depth Regime
Pierre Marion, Adeline Fermanian, Gérard Biau et al.
Deep ResNets are recognized for achieving state-of-the-art results in complex machine learning tasks. However, the remarkable performance of these arc...
A Comparative Evaluation of Quantification Methods
Tobias Schumacher, Markus Strohmaier, Florian Lemmerich
Quantification represents the problem of estimating the distribution of class labels on unseen data. It also represents a growing research field in su...
Lightning UQ Box: Uncertainty Quantification for Neural Networks
Nils Lehmann, Nina Maria Gottschling, Jakob Gawlikowski et al.
Although neural networks have shown impressive results in a multitude of application domains, the "black box" nature of deep learning and lack of conf...
Scaling Data-Constrained Language Models
Niklas Muennighoff, Alexander M. Rush, Boaz Barak et al.
The current trend of scaling language models involves increasing both parameter count and training data set size. Extrapolating this trend suggests th...
Curvature-based Clustering on Graphs
Zachary Lubberts, Yu Tian, Melanie Weber
Unsupervised node clustering (or community detection) is a classical graph learning task. In this paper, we study algorithms that exploit the geometry...
Composite Goodness-of-fit Tests with Kernels
Oscar Key, Arthur Gretton, François-Xavier Briol et al.
We propose kernel-based hypothesis tests for the challenging composite testing problem, where we are interested in whether the data comes from any dis...
PFLlib: A Beginner-Friendly and Comprehensive Personalized Federated Learning Library and Benchmark
Yang Liu, Jianqing Zhang, Yang Hua et al.
Amid the ongoing advancements in Federated Learning (FL), a machine learning paradigm that allows collaborative learning with data privacy protection,...
The Effect of SGD Batch Size on Autoencoder Learning: Sparsity, Sharpness, and Feature Learning
Wooseok Ha, Bin Yu, Nikhil Ghosh et al.
In this work, we investigate the dynamics of stochastic gradient descent (SGD) when training a single-neuron autoencoder with linear or ReLU activatio...
Efficient and Robust Transfer Learning of Optimal Individualized Treatment Regimes with Right-Censored Survival Data
Pan Zhao, Shu Yang, Julie Josse
An individualized treatment regime (ITR) is a decision rule that assigns treatments based on patients' characteristics. The value function of an ITR i...
DAGs as Minimal I-maps for the Induced Models of Causal Bayesian Networks under Conditioning
Xiangdong Xie, Jiahua Guo, Yi Sun
Bayesian networks (BNs) are a powerful tool for knowledge representation and reasoning, especially for complex systems. A critical task in the applic...
Adjusted Expected Improvement for Cumulative Regret Minimization in Noisy Bayesian Optimization
Shouri Hu, Haowei Wang, Zhongxiang Dai et al.
The expected improvement (EI) is one of the most popular acquisition functions for Bayesian optimization (BO) and has demonstrated good empirical perf...
Manifold Fitting under Unbounded Noise
Zhigang Yao, Yuqing Xia
In the field of non-Euclidean statistical analysis, a trend has emerged in recent times, of attempts to recover a low dimensional structure, namely a ...
Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play
Zelai Xu, Chao Yu, Yancheng Liang et al.
Self-play (SP) is a popular multi-agent reinforcement learning framework for competitive games. Despite the empirical success, the theoretical propert...
Wasserstein Convergence Guarantees for a General Class of Score-Based Generative Models
Xuefeng Gao, Hoang M. Nguyen, Lingjiong Zhu
Score-based generative models are a recent class of deep generative models with state-of-the-art performance in many applications. In this paper, we e...
Extremal graphical modeling with latent variables via convex optimization
Sebastian Engelke, Armeen Taeb
Extremal graphical models encode the conditional independence structure of multivariate extremes and provide a powerful tool for quantifying the risk ...
On the Approximation of Kernel functions
Paul Dommel, Alois Pichler
Various methods in statistical learning build on kernels considered in reproducing kernel Hilbert spaces. In applications, the kernel is often selecte...
Efficient and Robust Semi-supervised Estimation of Average Treatment Effect with Partially Annotated Treatment and Response
Jue Hou, Tianxi Cai, Rajarshi Mukherjee
A notable challenge of leveraging Electronic Health Records (EHR) for treatment effect assessment is the lack of precise information on important clin...
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
Jingyang Li, Kuangyu Ding, Kim-Chuan Toh
Stochastic gradient methods for minimizing nonconvex composite objective functions typically rely on the Lipschitz smoothness of the differentiable pa...
Optimizing Data Collection for Machine Learning
Rafid Mahmood, James Lucas, Jose M. Alvarez et al.
Modern deep learning systems require huge data sets to achieve impressive performance, but there is little guidance on how much or what kind of data t...
Unbalanced Kantorovich-Rubinstein distance, plan, and barycenter on nite spaces: A statistical perspective
Shayan Hundrieser, Florian Heinemann, Marcel Klatt et al.
We analyze statistical properties of plug-in estimators for unbalanced optimal transport quantities between finitely supported measures in different p...
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
Jiajing Zheng, Alexander D'Amour, Alexander Franks
Recent work has focused on the potential and pitfalls of causal identification in observational studies with multiple simultaneous treatments. Buildin...
Rank-one Convexification for Sparse Regression
Alper Atamturk, Andres Gomez
Sparse regression models are increasingly prevalent due to their ease of interpretability and superior out-of-sample performance. However, the exact m...
gsplat: An Open-Source Library for Gaussian Splatting
Vickie Ye, Ruilong Li, Justin Kerr et al.
gsplat is an open-source library designed for training and developing Gaussian Splatting methods. It features a front-end with Python bindings compati...
Statistical Inference of Constrained Stochastic Optimization via Sketched Sequential Quadratic Programming
Sen Na, Michael Mahoney
We consider online statistical inference of constrained stochastic nonlinear optimization problems. We apply the Stochastic Sequential Quadratic Progr...
Sliced-Wasserstein Distances and Flows on Cartan-Hadamard Manifolds
Clément Bonet, Lucas Drumetz, Nicolas Courty
While many Machine Learning methods have been developed or transposed on Riemannian manifolds to tackle data with known non-Euclidean geometry, Optima...
Accelerating optimization over the space of probability measures
Shi Chen, Qin Li, Oliver Tse et al.
The acceleration of gradient-based optimization methods is a subject of significant practical and theoretical importance, particularly within machine ...
Bayesian Multi-Group Gaussian Process Models for Heterogeneous Group-Structured Data
Sudipto Banerjee, Didong Li, Andrew Jones et al.
Gaussian processes are pervasive in functional data analysis, machine learning, and spatial statistics for modeling complex dependencies. Scientific d...
Orthogonal Bases for Equivariant Graph Learning with Provable k-WL Expressive Power
Jia He, Maggie Cheng
Graph neural network (GNN) models have been widely used for learning graph-structured data. Due to the permutation-invariant requirement of graph lear...
Optimal Experiment Design for Causal Effect Identification
Sina Akbari, Negar Kiyavash, Jalal Etesami
Pearl’s do calculus is a complete axiomatic approach to learn the identifiable causal effects from observational data. When such an effect is not iden...
Mean Aggregator is More Robust than Robust Aggregators under Label Poisoning Attacks on Distributed Heterogeneous Data
Jie Peng, Weiyu Li, Stefan Vlaski et al.
Robustness to malicious attacks is of paramount importance for distributed learning. Existing works usually consider the classical Byzantine attacks m...
The Blessing of Heterogeneity in Federated Q-Learning: Linear Speedup and Beyond
Jiin Woo, Gauri Joshi, Yuejie Chi
In this paper, we consider federated Q-learning, which aims to learn an optimal Q-function by periodically aggregating local Q-estimates trained on lo...
depyf: Open the Opaque Box of PyTorch Compiler for Machine Learning Researchers
Kaichao You, Runsheng Bai, Meng Cao et al.
PyTorch 2.x introduces a compiler designed to accelerate deep learning programs. However, for machine learning researchers, fully leveraging the PyTor...
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
Shuze Daniel Liu, Shuhang Chen, Shangtong Zhang
Stochastic approximation is a class of algorithms that update a vector iteratively, incrementally, and stochastically, including, e.g., stochastic gra...
Improving Graph Neural Networks on Multi-node Tasks with the Labeling Trick
Xiyuan Wang, Pan Li, Muhan Zhang
In this paper, we study using graph neural networks (GNNs) for multi-node representation learning, where a representation for a set of more than one n...
Directed Cyclic Graphs for Simultaneous Discovery of Time-Lagged and Instantaneous Causality from Longitudinal Data Using Instrumental Variables
Wei Jin, Yang Ni, Amanda B. Spence et al.
We consider the problem of causal discovery from longitudinal observational data. We develop a novel framework that simultaneously discovers the time-...
Bayesian Sparse Gaussian Mixture Model for Clustering in High Dimensions
Fangzheng Xie, Yanxun Xu, Dapeng Yao
We study the sparse high-dimensional Gaussian mixture model when the number of clusters is allowed to grow with the sample size. A minimax lower bound...
Regularizing Hard Examples Improves Adversarial Robustness
Hyungyu Lee, Saehyung Lee, Ho Bae et al.
Recent studies have validated that pruning hard-to-learn examples from training improves the generalization performance of neural networks (NNs). In t...
Random ReLU Neural Networks as Non-Gaussian Processes
Rahul Parhi, Pakshal Bohra, Ayoub El Biari et al.
We consider a large class of shallow neural networks with randomly initialized parameters and rectified linear unit activation functions. We prove tha...
Riemannian Bilevel Optimization
Jiaxiang Li, Shiqian Ma
In this work, we consider the bilevel optimization problem on Riemannian manifolds. We inspect the calculation of the hypergradient of such problems o...
Supervised Learning with Evolving Tasks and Performance Guarantees
Verónica Álvarez, Santiago Mazuelas, Jose A. Lozano
Multiple supervised learning scenarios are composed by a sequence of classification tasks. For instance, multi-task learning and continual learning ai...
Error estimation and adaptive tuning for unregularized robust M-estimator
Pierre C. Bellec, Takuya Koriyama
We consider unregularized robust M-estimators for linear models under Gaussian design and heavy-tailed noise, in the proportional asymptotics regime w...
From Sparse to Dense Functional Data in High Dimensions: Revisiting Phase Transitions from a Non-Asymptotic Perspective
Xinghao Qiao, Dong Li, Shaojun Guo et al.
Nonparametric estimation of the mean and covariance functions is ubiquitous in functional data analysis and local linear smoothing techniques are most...
Locally Private Causal Inference for Randomized Experiments
Jordan Awan, Yuki Ohnishi
Local differential privacy is a differential privacy paradigm in which individuals first apply a privacy mechanism to their data (often by adding nois...
Estimating Network-Mediated Causal Effects via Principal Components Network Regression
Alex Hayes, Mark M. Fredrickson, Keith Levin
We develop a method to decompose causal effects on a social network into an indirect effect mediated by the network, and a direct effect independent o...
Selective Inference with Distributed Data
Snigdha Panigrahi, Sifan Liu
When data are distributed across multiple sites or machines rather than centralized in one location, researchers face the challenge of extracting mean...
Two-Timescale Gradient Descent Ascent Algorithms for Nonconvex Minimax Optimization
Michael I. Jordan, Tianyi Lin, Chi Jin
We provide a unified analysis of two-timescale gradient descent ascent (TTGDA) for solving structured nonconvex minimax optimization problems in the f...
An Axiomatic Definition of Hierarchical Clustering
Ery Arias-Castro, Elizabeth Coda
In this paper, we take an axiomatic approach to defining a population hierarchical clustering for piecewise constant densities, and in a similar manne...
Test-Time Training on Video Streams
Renhao Wang, Yu Sun, Arnuv Tandon et al.
Prior work has established Test-Time Training (TTT) as a general framework to further improve a trained model at test time. Before making a prediction...
Adaptive Client Sampling in Federated Learning via Online Learning with Bandit Feedback
Boxin Zhao, Lingxiao Wang, Ziqi Liu et al.
Due to the high cost of communication, federated learning (FL) systems need to sample a subset of clients that are involved in each round of training....
A Random Matrix Approach to Low-Multilinear-Rank Tensor Approximation
Hugo Lebeau, Florent Chatelain, Romain Couillet
This work presents a comprehensive understanding of the estimation of a planted low-rank signal from a general spiked tensor model near the computatio...
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents
Marco Pleines, Matthias Pallasch, Frank Zimmer et al.
Memory Gym presents a suite of 2D partially observable environments, namely Mortar Mayhem, Mystery Path, and Searing Spotlights, designed to benchmark...
Enhancing Graph Representation Learning with Localized Topological Features
Zuoyu Yan, Qi Zhao, Ze Ye et al.
Representation learning on graphs is a fundamental problem that can be crucial in various tasks. Graph neural networks, the dominant approach for grap...
Deep Out-of-Distribution Uncertainty Quantification via Weight Entropy Maximization
Antoine de Mathelin, François Deheeger, Mathilde Mougeot et al.
This paper deals with uncertainty quantification and out-of-distribution detection in deep learning using Bayesian and ensemble methods. It proposes a...
DisC2o-HD: Distributed causal inference with covariates shift for analyzing real-world high-dimensional data
Jiayi Tong, Jie Hu, George Hripcsak et al.
High-dimensional healthcare data, such as electronic health records (EHR) data and claims data, present two primary challenges due to the large number...
Bayes Meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
Pierre Alquier, Charles Riou, Badr-Eddine Chérief-Abdellatif
Bernstein's condition is a key assumption that guarantees fast rates in machine learning. For example, under this condition, the Gibbs posterior with ...
Efficiently Escaping Saddle Points in Bilevel Optimization
Shiqian Ma, Minhui Huang, Xuxing Chen et al.
Bilevel optimization is one of the fundamental problems in machine learning and optimization. Recent theoretical developments in bilevel optimization ...