Timbre-preserving Speech Transmission through Semantic Extracting and Recovery

Zhu, Shuai (contact); Qian, Cheng; Fu, Junjie; Dong, Yunquan

10.23919/JCN.2025.000134

Abstract : As a new communication paradigm, semantic communication is drawing increasing attentions from the research community due to its capability to enhance the efficiency of transmissions. In this paper, we investigate a speech transmission system with timbre restoration and semantic recovery, which is referred to as a Timbre and Semantic Recovery Semantic Communication system (TSR-SC). Specifically, to mitigate the inevitable degradation of audio signals during transmissions, a semantic extraction module is utilized to transcribe the audio into text sequences, which are then represented via semantic embeddings. To preserve acoustic characteristics such as intonation and speaking rate, during the transmission, the timbre of the original speech is extracted as prior knowledge to fine-tune the speech generation module at the receiver. This enables the generated speech to closely resemble the original one in tone and speed. Experimental results demonstrate that TSR-SC significantly outperforms traditional methods in terms of speech quality and timbre restoration. 

Index terms : Semantic Communication, Deep Learning, Speech Transmission, Semantic Recovery, Transformer, Tone Cloning