Clay weekly context brief for the Systems category (ISO week 2026-W40). Clay tracks publications from the Systems feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing? Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.28713 Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. 2. BRiDCT: Fast Two-Dimensional DCTs Using SIMD: SIMD Organization, Register Blocking, and Numerical Verification Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.28519 Repeated two-dimensional discrete cosine transforms (DCTs) require implementations that remain fast across several array sizes while preserving the mathematical transform. 3. LoRa Fluid Antenna Multiple Access Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.28911 Concurrent long-range (LoRa) transmissions over the same time-frequency and spreading factor (SF) resources generally result in packet collisions, as the gateway cannot distinguish the overlapping signals from different end devices (EDs). 4. A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.28806 Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. 5. CrossScale-GLIO: Topology-Preserving Vision-Language Alignment of MRI and Whole-Slide Histopathology for Diffuse Glioma Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.28524 Magnetic resonance imaging and histopathology observe the same glioma at radically different scales. 6. Fast Frame Rate Estimation in Electromagnetic Side-Channel Attacks on Public Systems Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.28916 Frame refresh rate estimation is a fundamental step in identifying compromising harmonic frequencies in electromagnetic side-channel attacks. 7. Learning New Words from Unlabeled Test Data in Automatic Speech Recognition Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.28877 New words are invented every day. 8. Adaptive Tiling for Least-Squares Phase Unwrapping: Runtime and Accuracy Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.28541 Phase unwrapping estimates the missing multiples of $2\pi$ in measured phase images. 9. Unrolling Iterative Lanczos Algorithm for Ideal Low-pass Graph Filter Approximation Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29086 Low-pass (LP) filtering is a fundamental operation in graph signal processing (GSP). 10. Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across Architectures Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.28974 Post-training quantization (PTQ) reduces the cost of on-device text-to-speech (TTS), but published evaluations cover one system or method. 11. Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.28732 Primary bone tumors are rare but clinically aggressive neoplasms whose diagnosis from radiographs is challenged by heterogeneous morphology, subtle lesion margins, and overlapping bone structures. 12. Exact Factorisation and Fast Computation of Invertible Constant-Q Transforms Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29119 The constant-Q transform (CQT) represents audio on a logarithmic frequency axis. 13. Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.28988 We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. 14. Physics-Guided Multi-Objective Deep Learning for Ultrasound RF Data Interpolation in Resource-Constrained Imaging Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.28775 Ultrasound imaging increasingly targets portable, point-of-care, and wearable settings where constraints on power, bandwidth, and hardware complexity often necessitate sparse data acquisition in spatiotemporal scanning. 15. Bio-inspired efficient cyclostationary analysis in machine and underwater acoustic recordings Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29128 We propose a bio-inspired approach that uses the inner-hair-cell (IHC) response of the Cascade of Asymmetric Resonators with Fast-Acting Compression (CARFAC) model to efficiently extract cyclic modulation from acoustic signals. 16. Is Broader Better? A Controlled Study of Multilingual Coverage and Pretraining Objective in Frozen SSL Encoders for Speech Deepfake Detection Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.29138 Frozen self-supervised (SSL) speech encoders are strong, low-cost front ends for audio deepfake detection, and recent comparisons agree that large, multilingual, discriminative encoders generalize best out of domain. 17. Revolutionizing Diffusion MRI Microstructure Mapping via Global Inversion Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.28958 Diffusion MRI microstructure mapping (MM) is conventionally solved voxel by voxel, ignoring the fact that tissue microstructure forms a spatially organized field. 18. Enhancing 5G NTN VSAT RACH in GNSS-Denied environments Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29164 The 3GPP 5G non-terrestrial networks (NTN) technology fundamentally relies on a global navigation satellite system (GNSS) fix at the user equipment (UE) to pre-compensate for user-link delay and Doppler shifts prior to any transmission, including the random-access procedure for initial access. 19. Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.29203 Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. 20. LC3EM: Long-Range Context Extrapolation Enhanced Entropy Model for Coordinate-based Overfitting Image Codecs Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.29192 Coordinate-based overfitting image codecs have attracted increasing attention for their low decoding complexity and independence from cross-image generalization. 21. Low-Complexity Multi-User Non-Line-of-Sight Channel Estimation Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29342 Future radio access networks are expected to rely on large-aperture antenna arrays, for which an increasing portion of the coverage region may fall within the radiative near-field. 22. WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.29372 The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. 23. A Unified Frequency-Domain Model for Cascaded Filter-Interpolation Modulation in Tomographic Reconstruction Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.29220 The fidelity of image reconstruction from projections in linear inverse problems, such as tomography, is critically dependent on the synergistic interaction between frequency-domain filtering and spatial-domain interpolation. 24. Power-MSE trade-off of Factorized Low-rank Approximated Computation Scheme with Memristors Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29732 Memristor crossbars enable analog vector-matrix multiplication (VMM) which is promising for machine learning applications. 25. Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.29405 We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). 26. Scalable photoacoustic tomography implementations accounting for the spatial impulse response of transducers Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.29253 Iterative model-based reconstruction in photoacoustic tomography repeatedly applies the forward operator mapping the initial pressure to the transducer signals, and its adjoint. 27. Toward Reliable and Accurate Predictive ISAC in Mobile mmWave Networks Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29741 Integrated sensing and communications (ISAC) systems offer a promising framework for beam tracking, which enhances channel awareness and improves communication reliability in mobile millimeter wave (mmWave) networks. 28. Voice Agents under Acoustic Stress: From Signal Degradation to Interaction and Action Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.29452 Voice agents must complete users' tasks despite noise, reverberation, and competing speech. 29. Graph-Based Semi-Supervised Hyperspectral Image Classification with Distance-Aware Spatial Measure Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.29367 The classification of hyperspectral images (HIs) still presents several challenges. 30. Physical-Layer Aspects of Repeater-Assisted MIMO Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2609.29846 Network-controlled repeaters (NCRs) are emerg- ing as low-cost, band-selective active scatterers that can re- shape the wireless propagation environment without backhaul or tight phase synchronization. 31. Configurable-Bandwidth Time-Frequency Modeling for Efficient Full-Band Speech Enhancement Across Sampling Rates Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2609.29463 Speech enhancement systems are often developed for a fixed sampling rate, while time-frequency models become more expensive as the number of frequency bins increases. 32. Evidence-Driven Differential Diagnosis of Malignant Melanoma Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2609.29613 We present a modular and multi-level framework for the differential diagnosis of malignant melanoma. Sources in this brief: eess.AS (Audio and Speech Processing); eess.IV (Image and Video Processing); eess.SP (Signal Processing). Selected 32 of 421 available items for this weekly brief.