Clay weekly context brief for the Statistics category (ISO week 2026-W31). Clay tracks publications from the Statistics feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Conformal Anomaly Detection in Python: Moving Beyond Heuristic Thresholds with nonconform Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2605.13642 Most anomaly detection systems output scores rather than calibrated decisions, leaving practitioners to choose thresholds heuristically and without clear statistical interpretation. 2. Why Network Segmentation Projects Fail Source: stat.AP (Applications) Link: https://arxiv.org/abs/2604.08632 Network segmentation is a foundational enterprise security control. 3. Minimum Norm Interpolation via the Local Theory of Banach Spaces: The Role of $2$-Uniform Convexity Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2603.28956 The minimum-norm interpolator (MNI) framework has recently attracted considerable attention as a tool for understanding generalization in overparameterized models, such as neural networks. 4. Scalable Gaussian process inference via neural feature maps Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2605.10285 We present a theoretically grounded Gaussian process framework that leverages neural feature maps to construct expressive kernels. 5. Statistical Modelling of Planetary Boundary Layer Height and Its Measurement Uncertainty Using GRUAN Profiles Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.14960 The Planetary Boundary Layer (PBL) governs the exchange of energy and moisture and hosts the highest concentrations of pollutants before they mix into the free troposphere. 6. PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2505.08784 As machine learning (ML) enters high-stakes domains, trustworthy uncertainty quantification (UQ) is essential for safety. 7. Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret with Infinite Variance Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2603.06851 We study contextual bilateral trade under full feedback when, conditionally on the context, trader valuations have bounded density but infinite variance. 8. Model quality in football: Quantifying the quality of an Expected Threat model Source: stat.AP (Applications) Link: https://arxiv.org/abs/2604.21087 The recent growth in data availability in football has increased the risk of incorrect use of data-driven models, making guidelines on their validation and application necessary. 9. The Role of Pseudo-labels in Self-training Linear Classifiers on High-dimensional Gaussian Mixture Data Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2205.07739 Self-training (ST) is a simple yet effective semi-supervised learning method. 10. On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2601.12238 In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and smoothness. 11. SynthIPD: training-free synthetic individual patient data generation Source: stat.AP (Applications) Link: https://arxiv.org/abs/2509.16466 Individual patient data (IPD) are essential for statistical inference in clinical research. 12. Gaussian and bootstrap approximations for functional principal component regression Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2603.12518 Asymptotic inference using functional principal component regression (FPCR) has long been considered difficult, largely because, upon any scalar scaling, the FPCR estimator fails to satisfy a central limit theorem, leading to the prevailing belief that it is unsuitable for direct statistical inference. 13. Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2510.04602 Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space. 14. A Human-Augmenting Agentic Workflow for Observational Causal Inference Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.22443 Data analysis agents are becoming increasingly common tools for applied and scientific research. 15. A complete characterization of testable hypotheses Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2601.05217 We revisit a fundamental question in hypothesis testing: given two sets of probability measures $\mathcal{P}$ and $\mathcal{Q}$, when does a nontrivial (i.e.\ strictly unbiased) test for $\mathcal{P}$ against $\mathcal{Q}$ exist? 16. An Insight on Evaluation Metrics Under the Imbalanced Case of Anomaly Detection Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2607.22286 Anomaly detection is inherently characterised by severe class imbalance, making the interpretation of evaluation metrics challenging. 17. General Value Functions for Remaining Useful Life and Failure-Mode Prediction Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.22268 Remaining useful life (RUL) prediction and failure-mode classification are central tasks in predictive maintenance. 18. Gaussian Mixture Model with unknown diagonal covariances via continuous sparse regularization Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2509.12889 This paper addresses the statistical estimation of Gaussian Mixture Models (GMMs) with unknown diagonal covariances from independent and identically distributed samples. Sources in this brief: stat.AP (Applications); stat.ML (Machine Learning); stat.TH (Statistics Theory). Selected 18 of 47 available items for this weekly brief.