Clay weekly context brief for the Systems category (ISO week 2026-W32). Clay tracks publications from the Systems feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Qwen-Audio-3.0-Gen-Preview Technical Report Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.27011 Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scenes. 2. Endo-NeRF++: Uncertainty-Aware Neural Rendering with Multi-Resolution Hash Encoding for Dynamic Surgical Scene Reconstruction Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.27825 Reconstructing dynamic surgical scenes is crucial for robot-assisted minimally invasive surgery; however, it continues to be difficult because of tissue deformation, occlusions, specular reflections, and restricted viewpoints. 3. Radar-Aided Near-Field Beam Prediction via Beam Map Learning for XL-MIMO V2I Communications Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.27643 Near-field beam training in extremely large-scale multiple-input multiple-output (XL-MIMO) vehicle-to-infrastructure (V2I) systems incurs high overhead due to large range-angle codebooks and rapid channel variation. 4. Robust Residual Finite Scalar Quantization for Neural Compression Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2508.15860 Finite Scalar Quantization (FSQ) offers simplified training but suffers from residual magnitude decay in multi-stage settings, where subsequent stages receive exponentially weaker signals. 5. TSOG: A Format For Temporally And Spatially Ordered Gaussians Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.28049 We propose Temporally and Spatially Ordered Gaussians (TSOG), a format for efficient representation of 4D Gaussian Splatting (4DGS) content. 6. V-RIS: Virtual-Aperture DoA Estimation with Sparse RIS Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.27716 Large-aperture reconfigurable intelligent surfaces (RISs) enable high-resolution 2D direction-of-arrival (DoA) estimation, but existing approaches still tie hardware cost and control overhead to aperture size. 7. Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.27756 Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. 8. ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.28144 We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. 9. Multi-Chirp AFDM for Rydberg Atomic Quantum Receivers: Waveform and Algorithm Design Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.27903 We propose a multi-chirp affine frequency division multiplexing (MC-AFDM) scheme for joint delay-Doppler estimation with Rydberg atomic quantum receivers (RAQRs). 10. CrowdioSet and PaRIRset: Two Datasets Towards Live Music Source Separation Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.27828 Most Music Source Separation (MSS) models do not generalize well to live music recordings because they are trained on studio recordings alone, disregarding the venue acoustics, the speaker system's response and audience noise. 11. Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.27266 We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. 12. Power-Efficient XL-MIMO Design for Mixed Near- and Far-Field SWIPT Systems Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.27992 This paper examines the power consumption (PC) efficiency of a mixed near- and far-field (MF) simultaneous wireless information and power transfer (SWIPT) system underpinned by a hybrid beamforming (HB)-based modular extra-large multiple-input-multiple output (XL-MIMO) array. 13. MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2509.19817 Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scarce. 14. Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.27357 Cross-modal knowledge distillation can transfer diagnostic knowledge from a strong but costly teacher modality to a cheaper and more deployable student modality. 15. Multi-Agent Reinforcement Learning for Base Station Placement in TDOA-Based Localization Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.28002 Accurate localization of devices is a key capability for emerging 5G and 6G networks and depends on effective base station (BS) placement. 16. A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.26623 Mask-based beamforming is a popular geometry-agnostic approach for speech enhancement, typically applying a single mask across all microphones to estimate the required covariance matrices. 17. Generalized Query-Oriented Image Semantic Coding Empowered by Large AI Models and Semantic-Aware Hybrid Beamforming Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.28276 Semantic communication is an emerging paradigm that can preserve the meaning of data during transmission. 18. A Stochastic Optimization Framework for RIS-Aided Wireless Network Design Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.28018 Reconfigurable intelligent surfaces (RISs) are a promising technology for improving the spectral and energy efficiency of future wireless networks, which make use of metasurfaces. 19. Disentangling the Interpretive and Predictive Roles of LIWC: Controlled Substitution in Depression-Related Classification Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.22952 Linguistic Inquiry and Word Count (LIWC) provides auditable psycholinguistic categories that are widely used to interpret depression-related language, but its incremental predictive role in modern multimodal systems remains unclear. 20. BCNet: Bronchus Classification via Structure Guided Representation Learning Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2205.06947 CT-based bronchial tree analysis is essential for diagnosing lung and airway diseases, yet automatic bronchus classification remains challenging because bronchial topology varies substantially across individuals. 21. Fractional Doppler Effects on OTFS-NOMA HetNets with Mixed-Mobility Users Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.28059 Heterogeneous networks (HetNets) are considered a promising approach to meet the increasing throughput requirements of 6G vehicular networks. 22. Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.23027 Zero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. 23. Physical prior guided cooperative learning framework for joint turbulence degradation estimation and infrared video restoration Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2408.04227 Infrared imaging and turbulence strength measurements are in widespread demand in many fields. 24. Beamforming and Phase Shift Design for STAR-RIS Assisted Secure Sensing and Communication in ISAC Systems Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.28081 Integrated sensing and communication(ISAC), as a rapidly advancing technique, introduces a fresh approach for achieving secure communication and intelligent sensing for future wireless networks. 25. Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.23938 In this report, we present Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system that jointly advances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, efficiency, and robustness. 26. Scalable Drift Monitoring in Medical Imaging AI Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2410.13174 The integration of artificial intelligence (AI) into medical imaging has advanced clinical diagnostics but poses challenges in managing model drift and ensuring long-term reliability. 27. When Linear RUL Labels Disagree with Vibration Degradation: A Stage-Aware Target and Dual-Scale Predictor Evaluated on XJTU-SY and IMS Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.28115 Remaining useful life (RUL) studies commonly treat the label as fixed, although clock-linear labels may decline while measured vibration remains nearly stable and then changes rapidly near failure. 28. Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.23961 Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. 29. Uncertainty-Aware Multimodal Fusion for Oral Lesion Classification Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2511.12268 Early detection of oral cancer and potentially malignant diseases is a major challenge in low-resource settings due to the scarcity of annotated data. 30. Reduced-Observation Approximation of Near-Field Gaussian Covariance Matrices Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.28201 Near-field covariance matrices are central to local- ization, covariance-aware estimation, and linear MMSE filtering in large-aperture arrays, but Gaussian position uncertainty requires costly numerical averaging of nonlinear spherical-wave steering vectors. 31. Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.24323 Recently, time-frequency neural vocoders have been approaching the state-of-the-art quality of time-domain neural vocoders. 32. Stabilizing Deep Reconstruction Operators with Contractive Anchoring Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.23341 Pretrained deep denoisers can be used to solve a wide range of model-based image reconstruction tasks via Plug-and-Play (PnP) and Regularization-by-Denoising (RED) algorithms, without retraining per task. Sources in this brief: eess.AS (Audio and Speech Processing); eess.IV (Image and Video Processing); eess.SP (Signal Processing). Selected 32 of 429 available items for this weekly brief.