Separation-based video coding often fails in low-latency, noisy wireless scenarios due to channel sensitivity and lack of task awareness. Semantic communication, leveraging end-to-end task-oriented architectures such as Deep Joint Source–Channel Coding (DJSCC), offers a robust alternative. This paper introduces a lightweight video transmission framework that integrates a Semantic Attention (SemAtt) module with a Signal to Noise Ratio (SNR) modulation (SNRMod) for key-frame encoding and decoding. SemAtt uses semantic segmentation maps as compact priors to enhance feature discriminability with minimal computational overhead, while semantic priors also aid non-key frame reconstruction via lightweight generative models. Experiments on high-resolution urban scenes show improved robustness under low SNR, with a 5.2% reduction in LPIPS and a 6.3% increase in mIoU over conventional coding schemes, achieved with only a 9.9% increase in GFLOPs and 1.1% in trainable parameters. The framework is thus suitable for real-time, bandwidth-constrained wireless applications.
Semantic-Aware Attention-Driven JSCC for Efficient Video Transmission over Wireless Channels / Coppola, G., Narimani Kenari, M., Sciddurlo, G., Cordeschi, N., Grieco, L.A., Boggia, G.. - ELETTRONICO. - (2026). [10.1109/INFOCOM59046.2026.11571562]
Semantic-Aware Attention-Driven JSCC for Efficient Video Transmission over Wireless Channels
Giancarlo Sciddurlo;Nicola Cordeschi;Luigi Alfredo Grieco;Gennaro Boggia
2026
Abstract
Separation-based video coding often fails in low-latency, noisy wireless scenarios due to channel sensitivity and lack of task awareness. Semantic communication, leveraging end-to-end task-oriented architectures such as Deep Joint Source–Channel Coding (DJSCC), offers a robust alternative. This paper introduces a lightweight video transmission framework that integrates a Semantic Attention (SemAtt) module with a Signal to Noise Ratio (SNR) modulation (SNRMod) for key-frame encoding and decoding. SemAtt uses semantic segmentation maps as compact priors to enhance feature discriminability with minimal computational overhead, while semantic priors also aid non-key frame reconstruction via lightweight generative models. Experiments on high-resolution urban scenes show improved robustness under low SNR, with a 5.2% reduction in LPIPS and a 6.3% increase in mIoU over conventional coding schemes, achieved with only a 9.9% increase in GFLOPs and 1.1% in trainable parameters. The framework is thus suitable for real-time, bandwidth-constrained wireless applications.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


