Warning: mkdir(): No space left on device in /www/wwwroot/Z3.com/func.php on line 127

Warning: file_put_contents(./cachefile_yuan/ymfswkjyxgs.com/cache/7e/7e154/af0c5.html): failed to open stream: No such file or directory in /www/wwwroot/Z3.com/func.php on line 115
Unisound's Technical Paper Officially Accepted by IEEE TASLP 2026, Pioneering Efficient Adaptation for Multilingual Speech-LLMs - Unisound


      1. 糖心VLOG官网入口,糖心LOGO入口,糖心视频网站,糖心视频在线免费观看

        Unisound's Technical Paper Officially Accepted by IEEE TASLP 2026, Pioneering Efficient Adaptation for Multilingual Speech-LLMs

        Unisound 24
        Unisound's Technical Paper Officially Accepted by IEEE TASLP 2026, Pioneering Efficient Adaptation for Multilingual Speech-LLMs

        Recently, the research work by the Unisound team, titled "Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition," has been officially accepted by IEEE Transactions on Audio, Speech and Language Processing (IEEE TASLP 2026), a top international journal in the field of audio, speech, and language processing.

        This work directly addresses the industry pain point of multilingual adaptation for Speech Large Language Models (Speech-LLMs), solving the core challenge of balancing cross-lingual interference and cross-domain knowledge transfer. It provides a novel and efficient fine-tuning paradigm for multilingual speech recognition in scenarioses with uneven data resources, while establishing a universal, reusable technical foundation for the large-scale deployment of low-resource language speech large models.

        IEEE TASLP, published by the IEEE Signal Processing Society, is the premier authoritative journal in the fields of audio, speech, and language processing. It focuses on three core directions—audio technology, speech processing, and natural language computing—and publishes high-quality, highly innovative cutting-edge academic research from around the world, with academic influence and industry recognition at the top level of the field.

        Core Information of the Selected Paper

        English Title: Zipper-LoRA: Dynamic Parameter Decoupling for Speech-LLM based Multilingual Speech Recognition

        Chinese Title: Zipper-LoRA:面向Speech-LLM多语种语音识别的动态参数解耦方法

        Authors: Yuxiang Mei, Delai Qiu, Shengping Liu, Jiaen Liang, Yanhua Long

        Research Areas: Multilingual Automatic Speech Recognition, Speech Large Language Models (Speech-LLMs), Low-Resource Language Adaptation, Parameter-Efficient Fine-Tuning (PEFT), LoRA Optimization

        Paper Abstract:

        For multilingual automatic speech recognition tasks with uneven resource distribution, existing speech large language models face the challenge of balancing cross-lingual interference and knowledge transfer during parameter-efficient fine-tuning: shared LoRA is easily dominated by high-resource languages, impairing low-resource language performance; while fully independent LoRA limits cross-lingual knowledge sharing.

        To address this issue, this paper proposes the Zipper-LoRA dynamic parameter decoupling framework, which divides LoRA adaptation capabilities into a shared subspace and a language-specific subspace, and uses a language-identity-aware router to dynamically fuse the two at the rank level, achieving fine-grained sharing of transferable knowledge and effective isolation of language conflicts. The paper further designs three variants—static, hard routing, and soft routing—and adopts a two-stage training strategy with Initial-B warm-start to improve convergence stability under imbalanced multilingual training.

        Experimental results show that Zipper-LoRA achieves excellent performance across all 12 target languages. Unlike methods such as FlyLoRA that achieve compression or efficiency optimization solely by adjusting the effective rank, Zipper-LoRA's proposed dynamic parameter decoupling strategy can effectively mitigate parameter interference between different languages while maintaining cross-lingual knowledge sharing. Furthermore, compared with Fully Shared Vanilla-LoRA and Fully Decoupled Independent-LoRA, Zipper-LoRA achieves more balanced performance across all languages, validating the effectiveness of dynamically fusing shared and language-specific knowledge.

        Figure 1. Zipper-LoRA Architecture Diagram

        Figure 2. Recognition Performance Comparison of Different LoRA Methods Across 12 Languages

        Paper Link: http://arxiv.org/abs/2603.17558

        Project Page: http://github.com/YuCeong-May/Zipper-LoRA

        From Speech Recognition to Speech-LLMs: Continuously Exploring the Boundaries of Speech Large Model Technology

        As one of the early domestic enterprises to enter the intelligent voice track, Unisound has been deeply engaged in speech recognition, speech synthesis, natural language processing, and large model research for many years, accumulating profound technical barriers and mature industrialization experience. The company's self-developed U2 speech large model already supports over 100 Chinese dialects and more than 15 international languages, and has been widely deployed in core vertical scenarioses such as healthcare, education, smart home, and intelligent in-vehicle systems, achieving large-scale commercialization of multi-scenario, multilingual speech AI.

        In the frontier field of Speech-LLMs, Unisound continues to deepen the efficient fusion architecture of speech encoders and large language models, continuously breaking through core technologies in end-to-end speech understanding and speech generation, and steadily strengthening multilingual and low-resource speech modeling capabilities.

        For a long time, the innovative achievements of Unisound's R&D team have been continuously published at top international conferences and journals such as ICASSP, Interspeech, ACL, and EMNLP, with broad international academic influence in fields including speech recognition, speaker verification, speech enhancement, and multilingual large model modeling.

        The acceptance of this work by IEEE TASLP 2026 is yet another authoritative recognition of Unisound's technical strength and innovation path. Looking ahead, we will adhere to the strategic direction of "strong foundation models, deep applications," continuously iterate model capabilities, promote the deep integration of speech large model basic research and industrial applications, and bring more advanced speech AI technology to more languages and scenarioses, serving a broader range of real-world applications.

        网站地图