Clay weekly context brief for the Systems category (ISO week 2026-W30). Clay tracks publications from the Systems feed list. Below are recent items from this category, each with its source and a short description of what the publication covers when one is available in the source feed. Recent publications: 1. MIMO Capacity Enhancement by Grating Walls: A Physics-Based Proof of Principle Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2604.01786 This paper investigates the passive enhancement of MIMO spectral efficiency through boundary engineering in a simplified two dimensional indoor proof of principle model. 2. WanSong v1.0 Technical Report Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.14749 Music generation foundation models have recently attracted significant industry attention. 3. A Hybrid Framework for Blood Vessel Morphology Classification: Discrete Geometry-based Tortuosity Feature Measurement, Information Gain-based Feature Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14195 Subjective visual grading of blood vessel tortuosity relies heavily on clinical experience, while traditional distance-based indices often fail to adequately characterize three-dimensional spatial deformation. 4. Denoising-Autoencoder-Assisted Physical Layer Secret Key Generation Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14505 In this paper, we propose denoising autoencoder (DAE)-assisted secret key generation (SKG), where channel noise reciprocity imperfections induced due to wireless channel measurements are suppressed, hence significantly enhancing the reliability and efficiency. 5. Towards Out-of-Distribution Detection in Vocoder Recognition via Latent Feature Reconstruction Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2406.02233 Advancements in synthesized speech have created a growing threat of impersonation, making it crucial to develop deepfake algorithm recognition. 6. FORCE-Interior: A Poisson Flow Generative Prior for Interior Tomography Reconstruction Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14320 Interior tomography reconstructs a region of interest (ROI) from truncated projection measurements. 7. SLIPT-Enabled Ground-to-UAV FSO Systems with Optical Reconfigurable Intelligent Surfaces Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14523 This paper proposes an optical reconfigurable intelligent surface (ORIS)-assisted ground-to-unmanned aerial vehicle (UAV) free-space optical (FSO) communication system empowered by simultaneous lightwave information and power transfer (SLIPT). 8. Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.13721 L2 speech assessment has traditionally focused on phonetic assessment, leaving the scoring of suprasegmental features such as rhythm and intonation underexplored. 9. ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14328 In proton therapy planning, respiratory-gated non-contrast CT (NCCT) is commonly used for lesion segmentation; however, accurate delineation remains challenging due to low lesion-to-background contrast. 10. Adaptive Score-Based VAMP: Self-Tuning Hyperparameters via Tilted EM Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14597 Approximate-message-passing methods offer fast Bayesian inference for high-dimensional inverse problems, but their performance and state-evolution predictions rely on correctly specified module parameters. 11. ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.12496 Zero-shot dialog TTS benefits from flow-matching, but minute-scale generation on dense mel-spectrograms causes severe memory bottlenecks, often forcing unnatural chunked synthesis. 12. Domain Adaptation of Mismatched Proximal Denoiser for Plug-and-Play Image Reconstruction Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14894 Plug-and-play proximal gradient descent (PnP-PGD) enables flexible image reconstruction by using denoisers as implicit priors. 13. Efficient Quantum Algorithm for Phase Optimization of 1-Bit RIS-Assisted MIMO Communication System Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14700 We propose a Quantum Approximate Optimization Algorithm with a deterministic linear ramp schedule (QAOA-LR) for phase optimization of a 1-bit RIS-assisted MIMO communication system. 14. CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.10142 Ultra-lightweight models are essential for the deployment of deep learning-based speech enhancement algorithms on edge devices. 15. Deep Scene-Driven Ordering of Hadamard Basis for Single-Pixel Spectral Imaging Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.15045 Spectral images are highly valuable for various applications, including environmental monitoring and precision agriculture. 16. Effect of Antenna Deployment on Achievable Rate in Cooperative Magnetic Induction Communication Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14701 Magnetic Induction (MI) communication can be applied in some through-the-earth scenarios such as mines and underground rivers. 17. Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.10146 Automatic Mean Opinion Score (MOS) prediction is essential for evaluating large-scale synthetic speech and audio enhancement systems, yet models frequently struggle with domain shift. 18. ESAR: Event-Based Synthetic Aperture Reconstruction Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.15073 Event cameras report asynchronous polarity events when changes in log--radiance exceed a fixed contrast threshold, producing signed temporal contrast measurements rather than conventional image frames. 19. Quantifying the complexity of trajectory ensembles with clustering-weighted multivariate multiscale sample entropy Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14738 Across the physical and life sciences, data increasingly appear as ensembles of trajectories, from chaotic flows and satellite constellations to clinical cohorts. 20. Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.10162 Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling. 21. WULPUS PRO: Multi-mode Ultra-Low-Power Wearable Ultrasound and Array Imaging with CMUT Support Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.12137 Wearable ultrasound enables continuous monitoring of physiological processes such as muscle dynamics, bladder volume, and cardiovascular activity. 22. Elliptic Range-Doppler Mapping for OFDM-ISAC under IQ Imbalance Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14775 Receiver in-phase/quadrature imbalance (IQI) couples each OFDM subcarrier with its mirror counterpart, creating ghost targets and degrading range-Doppler recovery in orthogonal frequency division multiplexing (OFDM) integrated sensing and communication (ISAC). 23. Perceived Annoyance in Multi-source Electric Vehicle AVAS Environments Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.10368 The increasing usage of electric vehicles in urban environments has resulted in a widespread presence of AVAS sounds. 24. 3D Lane Detection with Odometry for High-Speed Vehicle Racing Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14248 Lane boundary detection is a critical component in autonomous driving systems and has been rigorously studied in regular driving scenarios. 25. Conditional Generative Learning Enabled Wireless UAV Sensing and Tracking via Point Cloud Imaging Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14778 In this paper, we study an unmanned aerial vehicle (UAV) sensing and tracking problem, where a base station equipped with an antenna array continuously illuminates a flying UAV and exploits the reflected echoes for slot-wise point cloud imaging within its potential flight region. 26. GigaAM Multilingual: Foundation Model for Underrepresented Languages Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.10371 Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. 27. Emergent Region-Level Facial Correspondence in Frozen Vision Foundation Models Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14423 Frozen self-supervised vision models can align parts of generic objects, but it remains unclear whether this correspondence extends to human faces, where global layout is shared while identity-specific appearance varies sharply. 28. Jacobi Elliptic Chirps for Sub-Nyquist Multi-Target Ranging Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14783 Sub-Nyquist sampling is an attractive way to reduce the hardware cost of wideband pulse-compression radar, but it introduces coherent alias-induced replicas in the matched-filter range profile, producing spurious peaks known as ghost targets. 29. GigaChat Audio: Time-aware Large Audio Language Model Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.10387 Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. 30. Compression of 3D Gaussian Splatting Data Using GPU-friendly Graphics Texture Coding Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2607.14513 Techniques for modeling 3D scenes from image collections, such as 3D Gaussian Splatting (3DGS), are capable of generating high-quality novel views by leveraging graphics primitives with view-dependent appearance. 31. Learning-Driven Channel Representation for Wireless Localization: From Channel Observations to Location Inference Source: eess.SP (Signal Processing) Link: https://arxiv.org/abs/2607.14938 Wireless observations capture radio signal responses formed through interactions with propagation environments and spatial geometry. 32. FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation Source: eess.AS (Audio and Speech Processing) Link: https://arxiv.org/abs/2607.10421 While recent few-step sampling text-to-audio generation models like MeanAudio substantially accelerate generation by modeling average velocities, their strict one-step generation quality still lags significantly behind multi-step counterparts. 33. Converting T1-weighted MRI from 3T to 7T quality using deep learning Source: eess.IV (Image and Video Processing) Link: https://arxiv.org/abs/2507.13782 Ultra-high resolution 7 tesla (7T) magnetic resonance imaging (MRI) provides detailed anatomical views, offering better signal-to-noise ratio, resolution and tissue contrast than 3T MRI, though at the cost of accessibility. Sources in this brief: eess.AS (Audio and Speech Processing); eess.IV (Image and Video Processing); eess.SP (Signal Processing). Selected 33 of 445 available items for this weekly brief.