Clay weekly context brief for the Systems category (ISO week 2026-W33). Clay tracks publications from the Systems feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Hyperspectral Calibration Detection: A Novel Concept For Change Detection With Unsupervised Incremental Safe Pseudo-Labeling Implementation Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.06028 Hyperspectral change detection (HCD) has found numerous key applications, such as land cover monitoring. 2. Channel Map-Based Channel Estimation for Near-Field UM-MIMO with Movable Planar Arrays Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.05520 Accurate channel estimation is essential for coherent transmission in ultra-massive multiple-input multiple-output (UM-MIMO) systems, where near-field propagation and high-dimensional spatial channels impose substantial signal processing challenges. 3. DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.04140 Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to later positions and can escalate into hallucination. 4. Media Meets Communication in 6G: Fundamentals, Key Technologies, and Applications Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.05184 The rapid advancement of sixth-generation (6G) networks is accelerating the convergence of media intelligence and communication intelligence, driving media communication beyond conventional bit-level delivery toward intelligent, semantic-aware, and generative paradigms. 5. Lyapunov-Based Completion-Aware Scheduling for PDU Set-Based Real-Time XR Traffic Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.05537 Real-time extended reality (XR) services impose stringent throughput, latency, and reliability requirements. 6. Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.18303 Identifying which string produces a given pitch in monophonic electric guitar audio is a classification challenge: a single pitch can often be produced on multiple strings, with timbral differences largely imperceptible to untrained humans. 7. Accurate Localization of Road Traffic Objects on the Road Plane Using Surveillance Camera Imagery Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.05840 Accurate vehicle localization from monocular roadside surveillance cameras is important for intelligent transportation systems, traffic monitoring, and traffic conflict analysis. 8. Radio-FM: A Foundation Model for Radio Signal Representation Learning and Its Applications Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.05793 Applying foundation models to the radio frequency (RF) domain presents unique challenges due to the intrinsic physical complexity of raw I/Q signals and the extreme heterogeneity of spectral data. 9. Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2604.23586 Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. 10. Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.06037 Relational inductive biases are essential for capturing structural dependencies among data. 11. Exact DC Representation of Multi-Tier Offloading Product in SAGINs via Quantifier Elimination Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.05978 Task offloading in space--air--ground integrated networks (SAGIN) yields non-convex signomial or polynomial programs with cubic couplings. 12. The Eloquence team submission for task 1 of MLC-SLM challenge Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2507.19308 In this paper, we present our studies and experiments carried out for the task 1 of the Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM), which focuses on advancing multilingual conversational speech recognition through the development of speech language models architectures. 13. OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.06264 The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors from large volumes to characterize pathologies. 14. Tracking performance of RLS algorithms in WSSUS channels Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.06036 Adaptive algorithms are widely used for estimation of linear time-varying systems, such as communication channels. 15. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.02023 Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. 16. TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.06275 Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. 17. Numerical Model of a Multiple-Input-Multiple-Output Distributed Acoustic Sensing System with Joint Phase and Birefringence Estimation Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.06198 In this work, we introduce and experimentally validate a numerical model for a Multiple-Input-Multiple-Output Distributed Acoustic Sensing (MIMO-DAS) system that accounts for dynamic perturbations of fiber birefringence and of the common optical phase of the backscattered signal (or polarization-averaged phase, shared by both polarization tributaries). 18. Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.26575 We propose a deep unfolded REM network for robust tracking of a single moving speaker in mild reverberant environments. 19. Tree-NET: Enhancing 2D Medical Image Segmentation Through Efficient Low-Level Feature Training Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2501.02140 This paper introduces Tree-NET, a novel framework for medical image segmentation that leverages bottleneck supervision to enhance both segmentation accuracy and computational efficiency. 20. Neural CRC Prediction for 5G NR URLLC Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.06230 We propose a neural cyclic redundancy check (CRC) predictor for the 5G New Radio (5G NR) physical uplink shared channel (PUSCH) that enables early link-adaptation decisions for Ultra-Reliable Low-Latency Communications (URLLC). 21. MAGENTA: Magnitude and Geometry-Enhanced Training Approach for Long-Tailed Sound Event Localization and Detection Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2509.15599 Deep learning-based Sound Event Localization and Detection (SELD) systems suffer severe performance degradation in real-world, long-tailed acoustic environments. 22. Sparse Mixture-of-Experts for Non-Uniform Noise Reduction in MRI Images Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2501.14198 Magnetic Resonance Imaging (MRI) is an essential diagnostic tool in clinical settings, but its utility is often hindered by noise artifacts introduced during the imaging process. 23. Joint Access Point Selection and Precoder Design under Statistical CSI Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.06251 This work addresses joint access point (AP) selection and precoding for sum-rate maximization under statistical channel state information (CSI) in multi-AP multi-user systems. 24. Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2508.05149 Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. 25. PromptForSegCXR: Prompt-Driven Multi-Organ and Multi-Disease Segmentation in Chest X-rays using a Multi-stage Fusion Mechanism Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2507.00673 Image segmentation is central to automated medical image analysis, enabling precise identification of anatomical structures and pathological regions. 26. An Analysis of the Accuracy of the Added Length De-Embedding Methods for Coaxial to Waveguide Adapters in the X-Band Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.06313 This paper assesses the accuracy of the "added length method" (also known as port extension or channel offset) for de-embedding coaxial-to-waveguide transitions in the X-band, comparing it to full waveguide calibration on both Keysight and Rohde & Schwarz VNAs. 27. Why Pre-trained Models Fail: Feature Entanglement in Multi-modal Depression Detection Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2503.06620 Depression remains a pressing global mental health issue, driving considerable research into AI-driven detection approaches. 28. On Optimizing Image Codecs for VMAF NEG: Analysis, Issues, and a Robust Loss Proposal Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2602.21336 The VMAF (video multi-method assessment fusion) metric for image and video coding recently gained more and more popularity as it is supposed to have a high correlation with human perception. 29. A Unified Framework for Sample Complexity of Structured Quantum State Tomography under Noisy Observations Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.05526 Quantum state tomography (QST) has attracted considerable attention due to its fundamental role in quantum information processing. 30. Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.03610 Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. 31. Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.10082 Building pixel-level correspondence between event and image data is a fundamental task for multi-sensor systems. 32. A Low-Power Wearable Respiratory Sensor for Non-Invasive Stress Monitoring Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.05697 Respiration provides a continuously available window into physiological state and behavior. 33. Identity-Faithful Audio-Visual Target Speaker Extraction with QIANGDA and VOXBLINK2-AVSE Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.03964 Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant voice. Sources in this brief: eess.AS (Audio and Speech Processing); eess.IV (Image and Video Processing); eess.SP (Signal Processing). Selected 33 of 471 available items for this weekly brief.