About Me
I am a Ph.D. candidate in Computer Science at The Chinese University of Hong Kong, Shenzhen, advised by Prof. Haizhou Li. My research centers on speech-to-speech modeling, particularly accent conversion. I am also co-advised by Prof. Shuai Wang. For research discussions and collaborations, please contact me at qibingbai@link.cuhk.edu.cn.
Experience
- Tencent TEA-Lab Jun 2024 – Jul 2026Research Intern
Accent conversion, spoken dialogue model
- ByteDance AI Lab Feb 2022 – Mar 2023Research Intern
Speech-to-speech translation
Education
- The Chinese University of Hong Kong, Shenzhen 2023 – PresentPh.D. in Computer Science
- Southern University of Science and Technology 2020 – 2023M.Eng.
- Central South University 2016 – 2020B.Eng. in Electronic Engineering
Honors
- Best Paper Award 2024International Symposium on Chinese Spoken Language Processing (ISCSLP)
- SRIBD Ph.D. Fellowship 2023The Chinese University of Hong Kong, Shenzhen
- Excellent Teaching Assistant 2022Southern University of Science and Technology
Publications
- Controllable Accent Normalization via Discrete Diffusion
Interspeech 2026 (Long Paper)
- TokAN: Accent Normalization Using Self-Supervised Speech Tokens
Submitted to IEEE TASLP
- Bridging What the Model Thinks and How It Speaks: Expressive Speech Generation via Self-Aware Intent-Realization Alignment
Submitted to EMNLP 2026
- CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
ICASSP 2026
- Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
Interspeech 2025
- Diffusion-Based Method with TTS Guidance for Foreign Accent Conversion
ISCSLP 2024 Best Paper Award
- LLaST: Improved End-to-End Speech Translation System Leveraged by Large Language Models
Findings of ACL 2024
- Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition
ICASSP 2024
- A Study of Modeling Rising Intonation in Cantonese Neural Speech Synthesis
Interspeech 2022
- LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT
Interspeech 2022
- Leveraging Pseudo-Labeled Data to Improve Direct Speech-to-Speech Translation
Interspeech 2022
* Equal contribution.
