Clay weekly context brief for the Statistics category (ISO week 2026-W32). Clay tracks publications from the Statistics feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Foundation-Model Earth Representations Enable Regional-Scale Forest Aboveground Biomass Monitoring Across the Northeastern United States Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.27217 Forest aboveground biomass (AGB) is a critical indicator of ecosystem productivity and terrestrial carbon storage, yet regional carbon monitoring remains constrained by the sparse spatial and temporal availability of field inventories and airborne structural measurements. 2. Sinkhorn Hamiltonian Monte Carlo for Entropic Optimal Transport Generalized Bayes Source: stat.CO (Computation) Link: https://arxiv.org/abs/2607.28015 Bayesian posterior sampling is a ubiquitous paradigm for problems where a point estimate of parameters is not sufficient, such as risk analysis and uncertainty quantification. 3. A New Measure of Dependence Between Continuous and Multinomial Random Variables Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2607.27381 A novel measure of dependence between a continuous random variable and a multinomial random variable is introduced. 4. An analysis of binary isotonic regression: degrees of freedom and implications for calibration Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2607.27301 Isotonic regression is a canonical tool for estimating monotone functions and calibrating probabilistic predictors. 5. Bringing Closure to False Discovery Rate Control: A General Principle for Multiple Testing Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2509.02517 We present a novel necessary and sufficient principle for multiple testing methods controlling an expected loss. 6. The Rise of Null Hypothesis Significance Testing (NHST): Institutional Massification and the Emergence of a Procedural Epistemology Source: stat.OT (Other Statistics) Link: https://arxiv.org/abs/2603.14757 It has long been a puzzle why, despite sustained reform efforts, many applied scientific fields remain dominated by Null Hypothesis Significance Testing (NHST), a framework that dichotomizes study results and privileges "statistically significant" findings. 7. Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.27224 External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions. 8. Reduce-Rank Matrix Integer-Valued Autoregressive Model Source: stat.CO (Computation) Link: https://arxiv.org/abs/2509.03338 Integer-valued time series are widely present in many fields, such as finance, economics, disease transmission, and traffic flow. 9. Adaptive Nystr\"om for Gaussian Process Regression Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2607.27427 Gaussian Process Regression (GPR) is a robust framework for uncertainty quantification, yet its $O(n^3)$ complexity limits its scalability. 10. HOMER: Huber-of-Means for Efficient and Robust Estimation in Hilbert Spaces Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2607.27532 Heavy tails weaken high-confidence control for the empirical mean. 11. Scalable Estimation of Crossed Random Effects Models via Multi-way Discretization Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2507.15593 Cross-classified data frequently arise in scientific fields such as education, healthcare, and social sciences. 12. Who Moved? Ecological Estimates of Turnout & Coalition Support by Ethnicity and Age in Johor, Malaysia, 2022-2026 Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.27864 At the July 2026 Johor state election, Barisan Nasional (BN) swept the state and nearly doubled its vote relative to 2022. 13. Additive Matrix Integer-Valued Autoregressive Model Source: stat.CO (Computation) Link: https://arxiv.org/abs/2605.30958 Contemporary data-driven and technology-integrated era, various matrix-valued integer-valued time series, such as cross-regional crime statistics, multi-category sales records, and network traffic matrices, exhibit high dimensionality, complex structures, and strong row-column intertwined dependencies. 14. Co-clustering of Response and Covariate Variables by Tri-Factorizing Their Non-negative Regression Coefficient Matrix Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2607.27474 Two-block data---two sets of variables measured on the same individuals, such as microbial taxa and metabolites---raise the question of how \emph{groups} of covariate variables relate to \emph{groups} of response variables. 15. Robust Wavelength Selection for Partial Least Squares Sugar Content Estimation Using Combinatorial Bayesian Optimization Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2607.27645 Wavelength selection is one of the important preprocessing methods in near-infrared spectroscopy to improve prediction accuracy and interpretability of spectral data. 16. Learning to Detect Cyber Attacks: Neural Anomaly Detection for Cybersecurity with Theoretical Insights Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2409.08521 In cybersecurity practice, new forms of cyberattacks continuously emerge, deliberately designed to evade defense systems that rely on previously observed behaviors. 17. What exam scores can and cannot prove about unauthorized AI assistance: Evidence from a highly public classroom episode Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.27978 In spring 2026, an economics professor at Brown University gave a take-home midterm and, after unusually high scores, made the final exam proctored. 18. Handling Missingness and Censoring in Dirichlet Mixture Models Source: stat.CO (Computation) Link: https://arxiv.org/abs/2607.27403 Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. 19. The Continuous Latent Ornstein-Uhlenbeck Dynamics Framework: A Scalable Latent Process Model for Multivariate Longitudinal Categorical Data Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2607.27520 Longitudinal biomedical studies increasingly collect irregularly sampled, multivariate categorical data that imperfectly reflect disease progression. 20. Robust Estimation of Sparse Numerical Vectors under Local Differential Privacy Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2607.27815 Local differential privacy (LDP) protocols are vulnerable to poisoning attacks. 21. Exact Tail Asymptotics of Dirichlet Distributions Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/0904.0144 Let $\X=A^\top R\U$ be a linearly transformed generalised symmetrised Dirichlet scale mixture in $\R^k$, $k\ge2$. 22. Bridging Probabilistic LLMs and Deterministic Statistical Validation: The PROVE Multi-Agent Framework for Clinical Trial Reporting Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.28218 Ensuring the accuracy and consistency of clinical trial Tables, Figures, and Listings (TFLs) remains a major challenge in regulatory reporting. 23. Mesh Invariant Infinite Dimensional Adaptive MCMC for Latent Gaussian Processes Source: stat.CO (Computation) Link: https://arxiv.org/abs/1804.04859 We introduce mesh-invariant adaptive Markov chain Monte Carlo methods for Gaussian-process posteriors arising in infinite-dimensional Bayesian inference. 24. Towards Best Practices for Covariate Adjustment in Regulatory Trials: From Fixed to Data-Adaptive Approaches Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2607.27542 While randomization justifies the use of unadjusted effect estimators in randomized trials, there is growing interest in covariate adjustment to improve precision. 25. Generalization and Trade-off in Adversarial Training: An RKHS Perspective via Kernel Integral Operators Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2607.27995 Adversarial training has emerged as a powerful approach for protecting models against adversarial attacks in a broad range of real-world applications. 26. The Kaplan-Meier estimator as a limit function of Efron's self-consistency iterations Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2607.27068 In 1967 Efron showed that the Kaplan-Meier estimator can be defined as a fixed point of a certain mapping, calling this property ``self-consistency''. 27. Snapshot plots: displaying summary tables as parallel univariate plots with consistent color highlighting Source: stat.AP (Applications) Link: https://arxiv.org/abs/2607.28302 For empirical studies, social and health scientists give background characteristics of their sample and summarize them in the famous "Table 1". 28. Local Quasi-Exponential Growth Models: Kernel Differential Equation Regression and Mouse Tumor Growth Data Source: stat.CO (Computation) Link: https://arxiv.org/abs/2505.00231 Local polynomial regression faces several challenges when dealing with sparse data. 29. Asymptotic emergence of statistically supported false causal interpretation under unmeasured confounding Source: stat.ME (Methodology) Link: https://arxiv.org/abs/2607.27593 Unmeasured confounding is widely recognized as a limitation of observational causal inference, but its implications for data-driven causal discovery are often understated. 30. On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems Source: stat.ML (Machine Learning) Link: https://arxiv.org/abs/2607.28080 We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and subspaces of relevant features in nonstationary and nonlinear regression problems. 31. The Phase Transition in Online PCA Depends on $n/d\log(d)$, not $n/d$ Source: stat.TH (Statistics Theory) Link: https://arxiv.org/abs/2607.23914 High dimensional statistical theory has established the importance of constant aspect ratio, when the number of dimensions ($d$) and samples ($n$) satisfy $n,d\to\infty$ with $n/d\to \gamma\in(0,\infty)$, in understanding the limits of canonical estimation problems. Sources in this brief: stat.AP (Applications); stat.CO (Computation); stat.ME (Methodology); stat.ML (Machine Learning); stat.OT (Other Statistics); stat.TH (Statistics Theory). Selected 31 of 416 available items for this weekly brief.