Research Article
-
10.1007/978-3-030-42504-3_16J. L. Kröger, O. H. M. Lutz, and P. Raschke, “Privacy implications of voice and speech analysis Information disclosure by inference,” in Privacy and Identity Management. Data for Better Living: AI and Privacy, pp. 242-258, 2020.
-
10.1016/j.icte.2024.11.003R. Sharma, N. Munjal, and N. Grover, “Addressing data scarcity in speech emotion recognition: A comprehensive review,” ICT Express, Vol. 11, No. 1, pp. 1-14, 2024.
-
10.1016/j.neucom.2022.04.028Y. B. Singh and S. Goel, “A systematic literature review of speech emotion recognition approaches,” Neurocomputing, Vol. 492, pp. 245-263, 2022.
-
10.21437/Interspeech.2005-446F. Burkhardt, A. Paeschke, M. Rolfes, W. Sendlmeier, and B. Weiss, “A database of German emotional speech,” in Proc. INTERSPEECH, pp. 1517-1520, 2005.
-
10.1371/journal.pone.0196391 29768426 PMC5955500S.R. Livingstone and F.A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PLoS ONE, Vol. 13, No. 5, pp. 1-35, 2018.
-
10.1016/j.specom.2021.11.006K. Zhou, B. Sisman, R. Liu, and H. Li, “Emotional voice conversion: Theory, databases and ESD,” Speech Communication, Vol. 137, pp. 1-18, 2022.
-
10.21437/Interspeech.2019-2680D.S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E.D. Cubuk, and Q.V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. INTERSPEECH, pp. 2613-2617, 2019.
-
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems (NeurIPS), pp. 2672-2680, 2014.
-
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems (NeurIPS), Vol. 33, pp. 6840-6851, 2020.
-
10.1109/TASL.2011.2114881C.H. Taal, R.C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time-frequency weighted noisy speech,” IEEE Transaction on Audio, Speech, and Language Processing, Vol. 19, Issue 7, pp. 2125-2136, 2011.
-
10.1109/ICASSP.2001.941023A.W. Rix, J. G. Beerends, M.P. Hollier, and A.P. Hekstra, “Perceptual evaluation of speech quality (PESQ): A new method for speech quality assessment of telephone networks and codecs,” 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing(ICASSP), Vol. 2, pp. 749-752, 2001.
-
10.1109/ICASSP39728.2021.9414878C.K.A. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” in Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing(ICASSP), pp. 6493-6497, 2021.
-
10.21437/Interspeech.2021-299G. Mittag, B. Naderi, A. Chehadi, and S. Möller, “NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets,” in Proc. INTERSPEECH, pp. 2127-2131, 2021.
-
S.W. Fu, C.F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Proc. ICML, Vol. 97, 2031-2041, 2019.
-
10.21437/Interspeech.2021-599S.W. Fu, C. Yu, T.A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: An improved version of MetricGAN for speech enhancement,” in Proc. INTERSPEECH, pp. 201-205, 2021.
-
10.1109/ICASSP.2016.7472773Z.Q. Wang and D. Wang, “Robust speech recognition from ratio masks,” 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5720-5724, 2016.
-
10.1145/3578741.3578761Z. Wang and X. Guo, “Research on Mandarin Chinese in speech emotion recognition,” in Processing 5th International Conference on Machine Learning and Natural Language Processing (MLNLP), pp. 99-103, 2022.
-
10.1121/1.3570957 21568429 PMC3115276F. Chen, and P.C. Loizou, “Predicting the intelligibility of vocoded and wideband Mandarin Chinese,” The Journal of the Acoustical Society of America, Vol. 129, Issue 5, pp. 3281-3290, 2011.
-
10.1121/10.0034233C. Xu, C. Brian, J. Moore, M. Diao, X. Li, and C. Zheng, “Predicting the intelligibility of Mandarin Chinese with manipulated and intact tonal information for normal-hearing listeners,” The Journal of the Acousticcal Society of America, Vol. 156, Issue 5, pp. 3088-3101, 2024.
- Publisher :The Society of Convergence Knowledge
- Publisher(Ko) :융복합지식학회
- Journal Title :The Society of Convergence Knowledge Transactions
- Journal Title(Ko) :융복합지식학회논문지
- Volume : 14
- No :3
- Pages :17-32
- DOI :https://doi.org/10.22716/sckt.2026.14.3.002


The Society of Convergence Knowledge Transactions






