All Issue

2026 Vol.14, Issue 3 Preview Page

Research Article

30 September 2026. pp. 17-32
Abstract
감정 음성 데이터는 개인정보 보호 및 보안 제약으로 인한 수집의 어려움과 전문 인력이 요구되는 높은 구축 비용으로 인해 충분한 규모의 데이터 셋 확보가 어렵다. 또한 음성 품질 평가 메트릭의 성능 검증을 위한 표준화된 테스트 데이터 셋이 부재하여, 기존 메트릭의 신뢰성을 정량적으로 평가하기 어려운 실정이다. 본 연구에서는 이 두 가지 문제를 동시에 해결하기 위해 MetricGAN 기반의 품질 레벨 제어형 감정 음성 증강 프레임워크인 QC-MetricGAN을 제안한다. QC-MetricGAN은 목표 STOI값을 5단계 품질 레벨로 설정하고, 레벨별 차등 품질 저하 입력과 비대칭 손실 함수를 적용하여 Generator가 목표 품질 수준의 음성을 생성하도록 학습한다. 또한 생성자의 구조적 과향상 성향을 감지하여 입력 품질 저하 강도를 자동으로 강화하는 적응형 학습 전략을 통해 목표 레벨로의 안정적인 수렴을 유도한다. EMO-DB, RAVDESS, ESD(영어/중국어) 총 4개 데이터 셋에 대한 실험 결과, 전체 125개 실험 항목 중 113개(90.4%)가 수렴 기준을 만족하였으며, 데이터 규모가 증가할수록 수렴률과 생성 안정성이 향상되는 경향을 확인하였다. 제안한 프레임워크를 통해 생성된 다중 품질 데이터 셋은 음성 품질 평가 메트릭의 표준화 벤치마크로 활용될 수 있으며, SER 모델의 학습 데이터로 적용함으로써 실제 환경에 대한 모델의 강건성과 일반화 성능 향상에 기여할 수 있다.
Emotional speech data is difficult to collect in sufficient quantities due to privacy and security constraints, and its construction requires significant cost owing to the need for expert annotators. Furthermore, the absence of standardized test datasets for validating speech quality evaluation metrics makes it difficult to quantitatively assess the reliability of existing metrics. To address both of these challenges simultaneously, this paper proposes QC-MetricGAN, a MetricGAN-based emotional speech augmentation framework with controlled quality-level generation. QC-MetricGAN defines five discrete quality levels based on target STOI values, and trains the Generator to produce speech at each target quality level by applying level-wise differentiated quality degradation inputs and an asymmetric loss function. In addition, an adaptive training strategy is employed to detect the Generator's structural tendency toward over-enhancement and automatically strengthen the input quality degradation configuration, thereby guiding stable convergence toward the target quality level. Experiments conducted on four datasets — EMO-DB, RAVDESS, and ESD (English and Chinese, evaluated independently) demonstrate that 113 out of 125 experimental trials (90.4%) satisfy the convergence criterion and a consistent trend is observed in which larger dataset scales lead to improved convergence rates and generation stability. The multi-quality datasets generated by the proposed framework can serve as standardized benchmarks for evaluating speech quality assessment metrics, and can further contribute to improving the robustness and generalization performance of Speech Emotion Recognition (SER) models when applied as training data under diverse real-world acoustic conditions.
References
  1. J. L. Kröger, O. H. M. Lutz, and P. Raschke, “Privacy implications of voice and speech analysis Information disclosure by inference,” in Privacy and Identity Management. Data for Better Living: AI and Privacy, pp. 242-258, 2020.

    10.1007/978-3-030-42504-3_16
  2. R. Sharma, N. Munjal, and N. Grover, “Addressing data scarcity in speech emotion recognition: A comprehensive review,” ICT Express, Vol. 11, No. 1, pp. 1-14, 2024.

    10.1016/j.icte.2024.11.003
  3. Y. B. Singh and S. Goel, “A systematic literature review of speech emotion recognition approaches,” Neurocomputing, Vol. 492, pp. 245-263, 2022.

    10.1016/j.neucom.2022.04.028
  4. F. Burkhardt, A. Paeschke, M. Rolfes, W. Sendlmeier, and B. Weiss, “A database of German emotional speech,” in Proc. INTERSPEECH, pp. 1517-1520, 2005.

    10.21437/Interspeech.2005-446
  5. S.R. Livingstone and F.A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PLoS ONE, Vol. 13, No. 5, pp. 1-35, 2018.

    10.1371/journal.pone.0196391 29768426 PMC5955500
  6. K. Zhou, B. Sisman, R. Liu, and H. Li, “Emotional voice conversion: Theory, databases and ESD,” Speech Communication, Vol. 137, pp. 1-18, 2022.

    10.1016/j.specom.2021.11.006
  7. D.S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E.D. Cubuk, and Q.V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. INTERSPEECH, pp. 2613-2617, 2019.

    10.21437/Interspeech.2019-2680
  8. I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems (NeurIPS), pp. 2672-2680, 2014.

  9. J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems (NeurIPS), Vol. 33, pp. 6840-6851, 2020.

  10. C.H. Taal, R.C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time-frequency weighted noisy speech,” IEEE Transaction on Audio, Speech, and Language Processing, Vol. 19, Issue 7, pp. 2125-2136, 2011.

    10.1109/TASL.2011.2114881
  11. A.W. Rix, J. G. Beerends, M.P. Hollier, and A.P. Hekstra, “Perceptual evaluation of speech quality (PESQ): A new method for speech quality assessment of telephone networks and codecs,” 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing(ICASSP), Vol. 2, pp. 749-752, 2001.

    10.1109/ICASSP.2001.941023
  12. C.K.A. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” in Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing(ICASSP), pp. 6493-6497, 2021.

    10.1109/ICASSP39728.2021.9414878
  13. G. Mittag, B. Naderi, A. Chehadi, and S. Möller, “NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets,” in Proc. INTERSPEECH, pp. 2127-2131, 2021.

    10.21437/Interspeech.2021-299
  14. S.W. Fu, C.F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Proc. ICML, Vol. 97, 2031-2041, 2019.

  15. S.W. Fu, C. Yu, T.A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: An improved version of MetricGAN for speech enhancement,” in Proc. INTERSPEECH, pp. 201-205, 2021.

    10.21437/Interspeech.2021-599
  16. Z.Q. Wang and D. Wang, “Robust speech recognition from ratio masks,” 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5720-5724, 2016.

    10.1109/ICASSP.2016.7472773
  17. Z. Wang and X. Guo, “Research on Mandarin Chinese in speech emotion recognition,” in Processing 5th International Conference on Machine Learning and Natural Language Processing (MLNLP), pp. 99-103, 2022.

    10.1145/3578741.3578761
  18. F. Chen, and P.C. Loizou, “Predicting the intelligibility of vocoded and wideband Mandarin Chinese,” The Journal of the Acoustical Society of America, Vol. 129, Issue 5, pp. 3281-3290, 2011.

    10.1121/1.3570957 21568429 PMC3115276
  19. C. Xu, C. Brian, J. Moore, M. Diao, X. Li, and C. Zheng, “Predicting the intelligibility of Mandarin Chinese with manipulated and intact tonal information for normal-hearing listeners,” The Journal of the Acousticcal Society of America, Vol. 156, Issue 5, pp. 3088-3101, 2024.

    10.1121/10.0034233
Information
  • Publisher :The Society of Convergence Knowledge
  • Publisher(Ko) :융복합지식학회
  • Journal Title :The Society of Convergence Knowledge Transactions
  • Journal Title(Ko) :융복합지식학회논문지
  • Volume : 14
  • No :3
  • Pages :17-32