Clay weekly context brief for the Statistics category (ISO week 2026-W40). Clay tracks publications from the Statistics feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. RLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.28625 Despite the importance of reinforcement learning with verifiable rewards (RLVR), the extent to which it can learn new reasoning capabilities remains debated. 2. Modeling cyclostationarity in time series using ASCA Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2603.05065 Modern data analysis across diverse disciplines increasingly relies on time series. 3. glmSTARMA -- An R-Package for fitting autoregressive spatio-temporal models following generalized linear models Source: stat.CO (Computation) Link: https://arxiv.org/abs/2607.08276 The R package glmSTARMA implements autoregressive models for spatio-temporal data at fixed locations, with time-invariant spatial dependency structure. 4. Auditing Bayesian Graph Alignment: Diagnostic Comparisons and Reference Failure Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.23232 Bayesian graph alignment estimates correspondence probabilities, but convergence of an alignment-score trace need not imply accurate correspondence marginals. 5. Confidence Horizons Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2608.03889 Anytime-valid inference enables analysts to continuously monitor their data and stop experiments early. 6. Engaging students with statistics through choice of real data context on homework Source: stat.OT (Other Statistics) Link: https://arxiv.org/abs/2603.04541 Statistics educators recommend teaching with real data with relevant contexts, but defining relevancy is challenging and varies by student. 7. SGA: Uncertainty Quantification for Multi-Step Forecasting in Time Series Foundation Models Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.28582 The recent emergence of Time Series Foundation Models (TSFMs) has significantly advanced multi-step forecasting performance, enabling accurate predictions over extended future horizons. 8. Field Theory of Bayesian Rate Estimation Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2509.06129 We address the statistical inference of a time-dependent rate of events in the framework of Bayesian field theory. 9. DeepGOF-1: A Pretrained Convolutional Goodness-of-Fit Test for Logistic Regression with a Computable Consistency Certificate Source: stat.CO (Computation) Link: https://arxiv.org/abs/2609.29575 Goodness-of-fit tests for logistic regression are least reliable where they are most needed: at small samples their levels drift from the nominal one, and combining them worsens the drift. 10. Do Audio Language Models Hear and Read Distinctive Features Alike? Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.30167 Audio language models pass speech and text through a single decoder. 11. All you need is log Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2606.27349 How different are several probability distributions from one another? 12. Asymptotic confidence intervals for the difference and the ratio of the weighted kappa coefficients of two diagnostic tests subject to a paired design Source: stat.OT (Other Statistics) Link: https://arxiv.org/abs/2407.21387 The weighted kappa coefficient of a binary diagnostic test is a measure of the beyond-chance agreement between the diagnostic test and the gold standard, and depends on the sensitivity and specificity of the diagnostic test, on the disease prevalence and on the relative importance between the false positives and the false negatives. 13. From Prediction to Explainable Provider Behavior Profiles for Fraud, Waste, and Abuse Review Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.28477 Claims data can show that provider behavior changed but cannot by itself explain why. 14. Decision Theoretic Subgroup Detection With Bayesian Machine Learning Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2509.05832 We consider the problem of identifying promising subpopulations in terms of treatment effectiveness or treatment effect heterogeneity, from a Bayesian decision theoretic perspective. 15. Fitting Large Nonlinear Mixed Effects Models Using Variational Expectation Maximization Source: stat.CO (Computation) Link: https://arxiv.org/abs/2604.26160 Nonlinear Mixed Effects (NLME) models are widely used in pharmacometrics and related fields to analyze hierarchical and longitudinal data. 16. Learning the Maximum Tolerated Dose for Continuous Toxicity via Monotone Bayesian Trees Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.30190 Phase I cancer trials seek the maximum tolerated dose (MTD) while protecting patients from excessive toxicity. 17. Pointwise Generalization in Deep Neural Networks Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2605.18598 We address the fundamental question of why deep neural networks generalize by establishing a pointwise generalization theory for fully connected networks. 18. Inverse Problems Conditioned on Observation Ensembles: Applications and Methods Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2601.22029 We introduce a new multivariate statistical problem that we refer to as the Ensemble-conditioned Inverse Problem (EIP). 19. Bridging Impulse Control of Piecewise Deterministic Markov Processes and Markov Decision Processes: Frameworks, Extensions, and Open Challenges Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2501.04120 Control theory plays a pivotal role in understanding and optimizing the behavior of complex dynamical systems across various scientific and engineering disciplines. 20. BLOC: A Global Optimization Framework for Sparse Covariance Estimation with Non-Convex Penalties Source: stat.CO (Computation) Link: https://arxiv.org/abs/2603.29169 We introduce BLOC (Black-box Optimization over Correlation matrices), a general framework for sparse covariance estimation with non-convex penalties. 21. Bayesian joint modeling of longitudinal patient-reported outcomes and survival: an application to chronic obstructive pulmonary disease Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.30188 Questionnaire-based patient-reported outcomes (PROs) are discrete, bounded and overdispersed, yet joint models relating them to survival may ignore these features or estimate both processes sequentially. 22. Stacked SVD or SVD stacked? A Random Matrix Theory perspective on data integration Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2507.22170 Modern data analysis increasingly requires identifying shared latent structure across multiple high-dimensional datasets. 23. Nuclear Norm-Regularized Bayesian Matrix Completion Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.30078 Matrix completion, the problem of estimating missing entries in a matrix from noisily observed ones, underlies a diverse array of problems such as recommender systems and counterfactual outcome estimation in panel data. 24. Three Ways Classical Test Theory Misleads for LLM Judges Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.29709 An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. 25. Delicatessen: Automated Estimating Equations in Python Source: stat.CO (Computation) Link: https://arxiv.org/abs/2203.11300 Estimating equation theory provides a unified framework for statistical modeling and inference. 26. The risks of dichotomising ordinal outcomes: A spatial analysis of self-rated health in Western Europe Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.29893 Survey responses are often measured using ordered response categories. 27. Self-Normalizing Denominators in Rational Covariance Estimators Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2608.20223 Many estimators are ratios of coprime polynomials in a sample covariance matrix, and their accuracy depends on the relative fluctuation of the sample denominator. 28. Path-specific harm decomposition: A partial identification framework Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2609.29938 A central goal when designing treatment policies is often to "do no harm", that is, to avoid interventions that improve average outcomes while worsening outcomes for some individuals. 29. A Proxy-likelihood Estimator for Multivariate Extremes Models with Intractable Likelihoods Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2609.30244 Many multivariate extremes models have intractable likelihoods requiring practitioners to use alternative fitting methods. 30. BFI: An R Package for Bayesian Federated Inference Source: stat.CO (Computation) Link: https://arxiv.org/abs/2609.27977 Bayesian Federated Inference (BFI) estimates statistical models from multicenter data when individual-level observations cannot be combined across centers. 31. Hybrid Models for Short-Term Sea-Level Forecasting Source: stat.AP (Applications) Link: https://arxiv.org/abs/2609.29827 Accurate tide forecasts are essential for coastal management, navigation, flood-risk reduction, and infrastructure protection. 32. Mean Residual Life Ageing Intensity Function Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2409.10456 Ageing intensity is usually formulated through the failure rate, whereas its mean residual life (\(MRL\))-based counterpart has not been systematically developed. Sources in this brief: stat.AP (Applications); stat.CO (Computation); stat.ME (Methodology); stat.ML (Machine Learning); stat.OT (Other Statistics); stat.TH (Statistics Theory). Selected 32 of 408 available items for this weekly brief.