Clay weekly context brief for the Statistics category (ISO week 2026-W39). Clay tracks publications from the Statistics feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Exponential Smoothing for Time Series of Random Objects Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2609.20274 Time series of random objects, such as covariance matrices, probability distributions, and functional data, call for forecasting methods that do not rely on standard arithmetic operations. 2. Comparing statistical learning models in wastewater-based epidemiology: An application to norovirus Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.20038 Wastewater-based epidemiology (WBE) is an increasingly important tool for infectious disease surveillance, but there has been limited direct comparison of modelling approaches for predicting pathogen concentrations across space and time. 3. Heterogeneity-calibrated Byzantine-robust distributed composite quantile regression Source: stat.CO (Computation) Link: https://arxiv.org/abs/2609.19701 We study sparse composite quantile regression (CQR) for distributed data with heterogeneous honest sites and Byzantine workers. 4. When a High Score Is an Illusion: Certifying Genuine versus Repackaged Forecasting Skill Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.19223 Ranks depend on the observations used for comparison. 5. Learning Submanifolds for Subsequent Inference on Random Dot Product Graphs, Part 1: Theory Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.19357 We propose a framework for restricted inference on random dot product graphs whose latent positions lie on an unknown low-dimensional support manifold. 6. Counterexamples and Sufficient Conditions: Comments on "Optimally-Transported Generalized Method of Moments" Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2609.20260 We comment on the optimally-transported generalized method of moments (OTGMM) estimator proposed by Schennach & Starck (2026a) and give counterexamples to Theorems 2-6 under their stated assumptions. 7. A Hierarchical Bayesian Model for Selecting Relevant Schema Subgraphs,with an Application to Grounding Large Language Models Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.20294 Grounding a large language model on a relational database requires selecting the tables and joins relevant to a query. 8. A Bayesian Bi-Directional Splitting Framework for Variable Selection in Large Datasets Source: stat.CO (Computation) Link: https://arxiv.org/abs/2609.20000 Modern tabular datasets are becoming increasingly large, both in the number of samples and covariates, posing significant challenges for Bayesian variable selection due to the resulting computational burden. 9. When are time-to-event models a waste of time? Bridging mixture cure models and positive-unlabeled learning for binary classification under right-censoring Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.19370 In clinical settings, predicting binary outcomes is often complicated by right-censoring, which prevents us from distinguishing individuals in whom the event never occurs (i.e., negative, or non-susceptible) from those where it occurs after the censoring time (i.e., positive, or susceptible). 10. Null importance: Disentangling relevance for interpretable machine learning Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.19511 Feature importance is central to interpretable machine learning, but the term "importance" encompasses several fundamentally different notions of relevance. 11. A proof of the strong Gaussian product inequality conjecture Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2609.20234 Let $\boldsymbol{X} = (X_1,\ldots,X_n)$ be a centered Gaussian vector, not necessarily nondegenerate. 12. Emulation strategies for Bayesian inference of regional left ventricle material parameters Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.20372 Patient-specific biomechanical models of the left ventricle can relate cardiac magnetic resonance imaging to regional myocardial material properties, but existing emulator-based studies typically treat the myocardium as mechanically homogeneous, limiting representation of localised dysfunction. 13. Online Supervised Dimension Reduction with Random Features: Diagnostics and Computational Trade-offs Source: stat.CO (Computation) Link: https://arxiv.org/abs/2609.20454 Accurate optimization of a supervised spectral objective need not produce an accurate population subspace or a better predictive representation. 14. Qualify-Then-Borrow: A Five-Step Framework for Bayesian Borrowing Beyond Outcome Agreement Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.19382 PURPOSE: Dynamic Bayesian borrowing methods adapt the contribution of external information according to agreement or disagreement with current data. 15. Error bounds in Sobolev norms for approximations with norm constrained ReLU neural networks Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.19937 Recent studies have shown that smooth functions can be well approximated by ReLU neural networks with path norm constraint on the weights. 16. Bounds for the median of the generalized hyperbolic and related distributions Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2609.20212 We prove monotonicity properties for medians of gamma sums and differences of the form $Z_{\alpha} = \alpha X_1 + (2 - \alpha)X_2$, where $X_1$ and $X_2$ are independent gamma random variables with common shape parameter. 17. Mutual Information as a Tool for Optimal Classification: Application to Identifying Rapid-Responding Behaviour Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.19781 Existing methods for identifying rapid-responding behaviour in large-scale assessments require parametric assumptions about the population. 18. Efficient computation of mixture confidence sequences in generalized linear models Source: stat.CO (Computation) Link: https://arxiv.org/abs/2609.20496 Classical confidence intervals, when repeatedly obtained on accumulating data at different sample sizes, produce contradictory inferences with high probability. 19. Bayesian Sample Size Determination: Sampling Distribution Estimation or Exploration? Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.19489 To design many Bayesian studies, the sample size is chosen to attain sufficient power to reject a null hypothesis or a desired probability that an interval estimate is sufficiently narrow. 20. Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.20389 Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. 21. Multi-Absorbing Phase-Type Distributions for Right-Censored Competing Risks Data Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2609.19921 Phase-type (PH) distributions are versatile semi-parametric models for lifetime duration and can be used in survival and reliability analysis. 22. Gain-function optimisation of graphical multiple testing procedures for confirmatory clinical trials Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.19994 Graphical multiple testing procedures are a flexible and transparent way to control the family-wise error rate when a confirmatory trial pursues several label claims, but they leave open the question of which graph to use. 23. Beyond point estimation: explaining the posterior in Bayesian cluster analysis Source: stat.CO (Computation) Link: https://arxiv.org/abs/2506.16295 The Bayesian approach to clustering is often appreciated for its ability to provide uncertainty in the partition structure. 24. Selective Inference in Growth Curve Models Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.19573 Growth curve models are widely used in psychological research, and variable selection can help identify baseline characteristics associated with longitudinal heterogeneity. 25. TAP Accuracy Below the Fluctuation Scale and Universal Posterior Geometry in Spherical Linear Models Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.20577 We study the Bayes-optimal spherical linear model as the ambient dimension and sample size grow proportionally, under a quantitative Marchenko--Pastur spectral-regularity condition on the design. 26. Report resolution in federated multiple testing under family-wise error control Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2609.19708 Several institutions test one family of hypotheses under family-wise error control but cannot pool their data, so each site releases, for each hypothesis, only a report of its own p-value. 27. Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.20758 Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. 28. Routing Frictions and Executable Liquidity in Fragmented Markets Source: stat.CO (Computation) Link: https://arxiv.org/abs/2609.19013 Public blockchains can make many trading venues simultaneously visible and mechanically reachable, yet an order still has to pay to activate each additional venue: technological connectivity need not translate into economically integrated execution. 29. Matrix Graphical Model Via Joint Estimation of Partial Correlations Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.19718 Matrix graphical models aim to characterize conditional dependence structures in matrix-variate data under a separable covariance assumption. 30. Learn Your Own Thoughts: Abstract Token Curriculum Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.19717 Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. 31. Oracle high-dimensional $M$-estimation using smooth reparameterization for sparsity Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2609.19558 This paper establishes a unified non-linear regularization framework for high-dimensional $M$-estimation, encompassing both linear models and Cox's proportional hazards models. 32. Estimating spatially-varying density and time-varying demographics with open population spatial capture-recapture: a photo-ID case study on bottlenose dolphins Source: stat.AP (Applications) Link: https://arxiv.org/abs/2106.09579 From long-term spatial capture-recapture (SCR) surveys, we can infer a population's dynamics over time and distribution over space. Sources in this brief: stat.AP (Applications); stat.CO (Computation); stat.ME (Methodology); stat.ML (Machine Learning); stat.TH (Statistics Theory). Selected 32 of 424 available items for this weekly brief.