Clay weekly context brief for the Systems category (ISO week 2026-W35). Clay tracks publications from the Systems feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. Explainability by Design: Structured Kolmogorov-Arnold Networks over Probabilistic Attributes for Speech Deepfake Source Tracing Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.20213 Modern speech synthesizers can produce highly realistic speech, making source tracing (i.e. 2. MOSAIC: A Self-supervised Dynamic Multi-encoding Reconstruction Framework for 3D Late Gadolinium Enhancement MRI Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.19506 Purpose: To develop and evaluate a self-supervised dynamic reconstruction framework for highly undersampled dual-echo three-dimensional late gadolinium enhancement (3D LGE) MRI. 3. COBALT: Column-swapping Optimized Bit-serial Accelerator for LSTM Tasks Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19363 Long Short-Term Memory (LSTM) networks continue to be widely deployed for real-time sequential tasks on edge devices, yet their computational demands challenge deployment on resource-constrained hardware. 4. $TCP_\alpha$: Margin-Controlled Confidence estimation for reliable Music Information Retrieval Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.20326 Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions. 5. Loss-Resilient Semantic Communication over Packet-Loss Networks at Extreme-Low Bandwidth Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.19590 In extreme-low bandwidth network scenarios, generative semantic codecs have emerged as promising solutions to reduce bandwidth cost for visual communications. 6. Robust Near-Field Beam Focusing Under Imperfect Localization Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19416 The transition to 6G-and-beyond wireless systems with large-scale antenna arrays and high-frequency deployments significantly extends the near-field region, where channels exhibit a strong dependence on user location. 7. A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.19361 This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. 8. AsymFeX: A Symmetry-Driven Framework for Ischemic Stroke Segmentation Across Imaging Modalities and Stroke Stages Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.19769 Fast and accurate segmentation of Acute Ischemic Stroke (AIS) lesions is essential for stroke prognosis and treatment planning. 9. Holographic Beamforming for Range-Doppler Sidelobe Suppression in OFDM-ISAC Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19431 This paper investigates range--Doppler (RD) sidelobe suppression in an integrated sensing and communications system with a reconfigurable holographic surface (RHS). 10. Towards Audio Token Compression in Large Audio Language Models Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2511.20973 Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.g., 25 tokens/s), making attention computation costly and limiting scalability. 11. MOSAIC: Modality-agnostic Spectral Alignment for Federated Image-level Weakly Supervised Tumor Segmentation under Client-specific Missing Modalities Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.19788 Trustworthy multimodal fusion in clinical settings requires handling incomplete and heterogeneous modality subsets across institutions, where privacy constraints prohibit centralized data sharing. 12. Differential Privacy in Feature Reconstruction Aided Federated Learning for Agent's Semantic Communication Model Update Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19649 This paper proposes a differentially private federated learning (FL) framework built upon an FL algorithm with semantic feature reconstruction (FedSFR) for training semantic communication modules for image transmission. 13. Linearly Constrained Deep Beamformer for Multi-Speaker Scenarios Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2605.21141 We propose a deep beamforming framework for enhancing target speaker(s) in multi-speaker environments. 14. Energy-Mamba: A Physics-Constrained State-Space Model for Medical Image Classification Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.19813 State-Space Models (SSMs), particularly Mamba, offer linear-time complexity for long-range dependencies, making them attractive for medical imaging with limited annotated data. 15. Band-Selective Microwave Cavity Optimization Using Differentiable FDTD: Gradient-Guided Search Versus Structured Random Search Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19820 We compare gradient-based inverse design with structured random search for dielectric-loaded microwave cavities using an in-house JAX-based differentiable FDTD solver. 16. SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2606.06837 Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus identity, channel conditions, and recording artifacts rather than speaking style itself. 17. Simulation-to-Real First-Break Segmentation for Efficient Inversion in Musculoskeletal Ultrasound Tomography Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.19828 Full-waveform inversion (FWI) is a promising strategy for quantitative musculoskeletal ultrasound computed tomography (USCT), but bone-related scattering, attenuation, and signal degradation make it highly sensitive to the accuracy of the initial acoustic-property distributions and prone to cycle skipping. 18. A 39pJ/b 7.3Gbps 1.3mm$^2$ Multi-Subcarrier Massive MU-MIMO-OFDM Detector Exploiting Beamspace Sparsity and Frequency-Domain Correlation in 22FDX Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19865 We present the first multi-subcarrier massive multi-user (MU) multiple-input multiple-output (MIMO) orthogonal frequency-division multiplexing (OFDM) data detector reported in the open literature. 19. Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.17102 Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. 20. Flow Matching-Based PET Image Reconstruction Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2608.20112 Generative models have shown strong potential for positron emission tomography (PET) image reconstruction. 21. Interpretable Feature Learning for RF Fingerprinting via Polar MKANs Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19881 Radio frequency (RF) fingerprinting authenticates wireless devices from hardware-induced I/Q impairments, typically with deep learning feature extractors that are accurate but opaque, limiting their use in security critical settings. 22. Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2605.17443 We analyze how automatic speech recognition (ASR) errors propagate through ASR--LLM cascades in Korean spoken question answering (SQA), focusing on downstream semantic failures that conventional ASR metrics cannot fully capture. 23. FermatSyn: SAM2-Enhanced Bidirectional Mamba with Isotropic Spiral Scanning for Multi-Modal Medical Image Synthesis Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2505.07687 Multi-modal medical image synthesis is pivotal for alleviating clinical data scarcity, yet existing methods fail to reconcile global anatomical consistency with high-fidelity local detail. 24. Some Practical Issues of the Tracking Process in GNSS Receivers Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.19927 The accuracy and noise immunity of GNSS receivers are largely determined by the performance of their tracking modules. 25. Audio Physical Dynamics Inspired Deepfake Detection for Voice Authentication Systems Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2512.06040 Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. 26. Exploiting Completeness Perception with Diffusion Transformer for Unified 3D MRI Synthesis Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2602.18400 Missing data problems, such as missing modalities in multi-modal brain MRI and missing slices in cardiac MRI, pose significant challenges in clinical practice. 27. Velocity Index Modulation for Movable Antenna Systems Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.20059 In movable antenna (MA) systems, antenna movement induces Doppler frequency shifts that are conventionally treated as an impairment requiring mitigation. 28. A Multiplication-Free Feature Extractor for Signal Classification: Keyword Spotting Case Study Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.17108 A very low complexity feature extractor called next iRDT is proposed and evaluated for the problem of keyword spotting (KWS). 29. A Systematic Survey on Event Camera Representation Learning Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2606.23078 Event cameras offer distinctive advantages, including microsecond-level latency and high dynamic range, rendering them promising for challenging perception tasks. 30. Low-complexity Soft-decision LLR Calculations for Next-generation IM-DD Systems with RIN Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.20102 The demand for higher speeds in intra-data center interconnects will eventually require high-order pulse amplitude modulation (PAM) combined with soft-decision (SD) forward error correction (FEC). 31. Cached LLM Probability Retrieval for Speech Recognition Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2608.16023 Large language models (LLMs) enhance automatic speech recognition (ASR) by providing linguistic priors; however, their direct rescoring is costly because it requires evaluating every N-best hypothesis. 32. FORCE-Interior: Measurement-Consistent Adaptation of a Poisson-Flow Generative Prior for Interior CT Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14320 Interior tomography reconstructs a region of interest (ROI) from truncated projections, an ill-posed problem with non-unique solutions and truncation-induced bias. 33. A Computationally Efficient Likelihood Approximation for Target Tracking in Time-Varying Multipath Channels Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2608.20206 Active sonar target tracking in shallow-water environments is challenging when weak target echoes are embedded in a time-varying background containing structured multipath components. Sources in this brief: eess.AS (Audio and Speech Processing); eess.IV (Image and Video Processing); eess.SP (Signal Processing). Selected 33 of 452 available items for this weekly brief.