Clay weekly context brief for the Statistics category (ISO week 2026-W33). Clay tracks publications from the Statistics feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Conditional multivariate functional PCA for the reconstruction of temperature and salinity profiles partially sampled by deep-diving marine mammals Source: stat.AP (Applications) Link: https://arxiv.org/abs/2608.05376 We present a statistical method to reconstruct the vertical thermohaline conditions in the Indian Sector of the Southern Ocean, where temperature and salinity profiles are partially sampled by female southern elephant seals. 2. Structured Dimension-Matched Joint Variational Transdimensional Inference Source: stat.CO (Computation) Link: https://arxiv.org/abs/2608.05607 Bayesian model selection couples a discrete model indicator with a model-specific continuous parameter space. 3. Precision and Decisiveness as Goals: Reliable Sequential Hypothesis Testing with a Dual Stopping Criterion Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2608.05301 Sequential hypothesis testing offers flexibility over fixed-sample designs, but stopping rules coupled to decision criteria risk confirmation bias through early peeking. 4. Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2608.05930 The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. 5. Realizable Bayes-Consistency for General Metric Losses Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2605.03823 We study strong universal Bayes-consistency in the realizable setting for learning with general metric losses, extending classical characterizations beyond $0$-$1$ classification (Bousquet et al., 2020; Hanneke et al., 2021) and real-valued regression (Attias et al., 2024). 6. Modeling E-Bike Route Choice in Washington, DC: A Path Size Logit Approach Source: stat.AP (Applications) Link: https://arxiv.org/abs/2608.05449 Understanding e-bike route choice is essential for developing effective cycling infrastructure, yet empirical evidence remains limited. 7. Learning Latent Memory States from Longitudinal Athlete Monitoring Data Source: stat.CO (Computation) Link: https://arxiv.org/abs/2608.06290 We propose a new unit of analysis for longitudinal data: the Latent Memory Table. 8. Sparse Principal Component Analysis via Wavelets for Distributed Data Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2608.05386 The large volume of data and concerns about data privacy have motivated the development of techniques for distributed data, a problem also known as federated learning. 9. Handling Missing Data in Probabilistic Regression Trees Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2608.06195 Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. 10. A note on conditional PAC-efficient reasoning in large language model routing Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2512.03057 We study distribution-free risk control for model routing, motivated by large language model reasoning. 11. How Infrastructure and Streetscape Shape E-Scooter Route Choice: Evidence from Washington, DC Source: stat.AP (Applications) Link: https://arxiv.org/abs/2608.05465 E-scooters have emerged as an important micromobility mode for short urban trips, yet evidence on route choice behavior remains limited. 12. A space of inference spaces in the space sciences - Parametric Bayesian inference in astronomy, cosmology and particle physics Source: stat.CO (Computation) Link: https://arxiv.org/abs/2608.06078 A sample of parametric Bayesian inference applications from astronomy, cosmology and particle physics is studied, augmented by mock data sets and toy problems. 13. Inference for subgraph densities in noisy dynamic networks Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2608.05407 In this work we develop statistical methodology to estimate and perform inference on subgraph densities using time-indexed, or dynamic network sequences. 14. Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2608.06206 Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable. 15. Determination of the Representative Sample Size in Linear Regression Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2608.02466 Very often, accuracy of analysis and forecasting (multiple coefficient of regression and residual means) obtained for a sample used to formulate a regression model is not equal to the accuracy achieved for another homogeneous sample. 16. Long-Term Mortality Following STN-DBS in Parkinson's Disease: A Survival Analysis Source: stat.AP (Applications) Link: https://arxiv.org/abs/2608.05609 Background: Deep brain stimulation of the subthalamic nucleus (STN-DBS) is an established PD treatment that improves motor symptoms and quality of life, but long-term survival and mortality risk remain poorly characterized. 17. Adaptive-precision computation of custom Gauss quadrature for statistical applications Source: stat.CO (Computation) Link: https://arxiv.org/abs/2607.14511 An $n$-point Gauss quadrature rule approximates the weighted integral of a function by a weighted sum of $n$ evaluations of this function and is exact for polynomials of degree at most $2n-1$. 18. Joint Multiple Imputation of Node Attributes and Network Ties in R Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2608.05895 Missing data in social network studies routinely affects both node-level attributes and the network ties themselves. 19. Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2608.06250 In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. 20. False discovery rate control with compound p-values Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2507.21465 In the setting of multiple testing, compound p-values generalize p-values by asking for superuniformity to hold only \emph{on average} across all true nulls. 21. Information leakage from data revisions in retrospective forecasts Source: stat.AP (Applications) Link: https://arxiv.org/abs/2608.05883 Ayg\"un et al (2026, https://doi.org/10.1038/s41586-026-10658-6) claim that their AI-driven Empirical Research Assistance (ERA) system produces COVID-19 hospitalisation forecasts which outperform the state-of-the-art CDC ensemble by a considerable margin for the 2024/25 season. 22. A monotonic MM-type algorithm for estimation of nonparametric finite mixture models with dependent marginals Source: stat.CO (Computation) Link: https://arxiv.org/abs/2505.16878 In this manuscript, we consider a finite nonparametric mixture model with non-independent marginal density functions. 23. A Regression-Based Framework for the ACF, PACF, Durbin-Levinson Recursion, and One-Step-Ahead Prediction Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2608.06334 The autocorrelation function (ACF) and partial autocorrelation function (PACF) are foundational tools for identifying autoregressive moving-average (ARMA) models, yet they are often introduced to students as computational recipes disconnected from the regression framework students already know. 24. FlowAdam: Implicit Regularization via Geometry-Aware Soft Momentum Injection Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2604.06652 Adaptive moment methods such as Adam use a diagonal, coordinate-wise preconditioner based on exponential moving averages of squared gradients. 25. Minimax estimation of functionals in sparse vector model with correlated observations Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2407.14778 We consider the observations of an unknown $s$-sparse vector ${\boldsymbol \theta}$ corrupted by Gaussian noise with zero mean and unknown covariance matrix ${\boldsymbol \Sigma}$. 26. Evaluating the influence of treatment-effect heterogeneity on discrimination Source: stat.AP (Applications) Link: https://arxiv.org/abs/2608.06002 Analyzing the heterogeneity of treatment effects is crucial in personalized medicine to identify which patients will benefit from specific treatments. 27. Distributional Inverse Homogenization Source: stat.CO (Computation) Link: https://arxiv.org/abs/2604.14083 For many materials, macroscopic mechanical behavior is determined by an intricate microstructure. 28. A Unified Causal Inference Framework for the Desirability of Outcome Ranking Paradigm in Benefit-Risk Evaluation Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2608.05244 We developed a unified covariate-adjusted causal inference framework for estimating the desirability of outcome ranking (DOOR) probability for benefit-risk evaluation in randomized trials and observational studies. 29. The Loss Does Not See the Basis, but Adam Does Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2608.05136 Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. 30. Transitional Conditional Independence Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2104.11547 Statistical models contain variables that are not random: parameters, treatments, environments, design points. 31. Longitudinal wearable monitoring and polygenic risk for incident major depressive disorder in the All of Us Research Program Source: stat.AP (Applications) Link: https://arxiv.org/abs/2608.06063 Major depressive disorder (MDD) risk reflects both stable inherited liability and dynamic behavioral patterns, yet these dimensions are rarely examined together using long-term objective data in real-world settings. 32. Conformal Anomaly Detection in Python: Moving Beyond Heuristic Thresholds with nonconform Source: stat.CO (Computation) Link: https://arxiv.org/abs/2605.13642 Most anomaly detection systems output scores rather than calibrated decisions, leaving practitioners to choose thresholds heuristically and without clear statistical interpretation. Sources in this brief: stat.AP (Applications); stat.CO (Computation); stat.ME (Methodology); stat.ML (Machine Learning); stat.TH (Statistics Theory). Selected 32 of 427 available items for this weekly brief.