DOI QR코드

DOI QR Code

Development of a Smartphone-Based Pronunciation Assessment and Feedback System Using Three-Dimensional Vowel Space Analysis and Lip Contour Extraction

스마트폰 기반 3차원 모음공간 분석 및 입술 윤곽 추출을 활용한 발음 평가·교정 시스템 개발

  • Hee-June Park (Department of Speech and Hearing Therapy, Catholic University of Pusan)
  • 박희준 (부산가톨릭대학교 언어청각치료학과)
  • Received : 2025.12.31
  • Accepted : 2026.01.20
  • Published : 2026.01.30

Abstract

The rapid globalization of Korean culture has precipitated a surge in Korean language learners, yet effective pronunciation training remains a significant pedagogical bottleneck. Traditional Computer-Assisted Pronunciation Training (CAPT) systems typically rely on two-dimensional acoustic analysis or automatic speech recognition confidence scores. These methods fundamentally fail to capture the temporal stability of vowel production-treating pronunciation as a static point rather than a dynamic distribution-and neglect the critical articulatory role of lip kinematics. This research presents the development of a novel smartphone-based biofeedback system that integrates 3D Vowel Space Analysis with Real-Time Lip Contour Extraction. By employing Kernel Density Estimation (KDE) on accumulated formant trajectories, the system visualizes the "density" of a learner's pronunciation in three dimensions. Simultaneously, a computer vision pipeline utilizes the MediaPipe Face Mesh to extract facial landmarks, providing immediate visual feedback on lip rounding. This multimodal approach effectively decouples acoustic errors from articulatory misconfigurations.

본 연구는 전 세계적으로 급증하는 한국어 학습 수요에 대응하여, 스마트폰만으로 정밀한 발음 교정이 가능한 시스템을 개발하는 것을 목적으로 한다. 기존의 2차원 평면 기반 음향 분석이 간과해 온 발화의 시공간적 안정성 (Stability)과 한국어 모음 변별의 핵심 기제인 입술 모양(lip rounding)을 통합적으로 분석한다. 연구 방법으로는 모바일 환경에서 선형 예측 부호화(LPC)를 통해 포먼트를 추출하고 커널 밀도 추정(KDE)을 적용하여 '3차원 포먼트 밀도' 지형도를 생성하며, 동시에 MediaPipe Face Mesh 기술로 입술의 개구도와 원순성을 정량화한다. 연구 결과, 개발된 시스템은 학습자가 생성하는 모음의 음향적 분산을 3차원 산맥 형태로 시각화하여 발음의 견고성을 인지하게 하였으며, 유사한 포먼트 값을 가지더라도 입술 모양이 잘못된 경우를 실시간으로 탐지하여 교정 효율을 높였다. 본 연구는 '점' 중심의 발음 평가를 '분포' 중심으로 전환하고, 음성학과 컴퓨터 비전 기술을 융합하여 모바일 컴퓨터 기반 컴퓨터 보조 발음훈련 (computer-assisted pronunciation training; CAPT) 시스템의 새로운 표준을 제시하였다.

Keywords

Acknowledgement

Following are results of a study on the "Busan Regional Innovation System & Education(RISE)" Project, supported by the Ministry of Education and Busan Metropolitan City

References

  1. Ministry of Education. (2024). Statistics of foreign students in higher education institutions in Korea. https://www.data.go.kr/data/15050054/fileData.do
  2. Lee, S., & Rhee, S. (2023). The relationship between vowel production and proficiency levels in L2 English produced by Korean EFL learners. Phonetics and Speech Sciences, 11(2), 1-10. DOI : 10.13064/KSSS.2019.11.2.001
  3. Benjamin, S., & Lee, H. Y. (2021). The perception and production of korean vowels by egyptian learners. Phonetics and Speech Sciences, 13(4), 23-34. DOI : 10.13064/KSSS.2021.13.4.023
  4. Neri, A., Mich, O., Gerosa, M., & Giuliani, D. (2008). The effectiveness of computer assisted pronunciation training for foreign language learning by children. Computer Assisted Language Learning, 21(5), 393-408 DOI : 10.1080/09588220802447651.
  5. Lu, X., & Dang, J. (2009). Vowel production manifold: Intrinsic factor analysis of vowel articulation. IEEE Transactions on Audio, Speech, and Language Processing, 18(5), 1053-1062. DOI : 10.1109/TASL.2009.2030939
  6. Sun, J., Yan, N., & Wang, L. (2013, October). Constructing a three-dimension physiological vowel space of the Mandarin language using electromagnetic articulography. In 2013 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (pp. 1-4). IEEE. DOI : 10.1109/APSIPA.2013.6694157
  7. Yang, B. (1996). A comparative study of American English and Korean vowels produced by male and female speakers. Journal of Phonetics, 24(2), 245-261. DOI : 10.1006/jpho.1996.0013
  8. Story, B. H., & Bunton, K. (2017). Vowel space density as an indicator of speech performance. The Journal of the Acoustical Society of America, 141(5), EL458-EL464.. DOI : 10.1121/1.4983342
  9. Google. (2025). MediaPipe Face Mesh Documentation. Retrieved December 23, 2025, https://developers.google.com/mediapipe/solutions/vision/face_mesh
  10. Suemitsu, A., Dang, J., Ito, T., & Tiede, M. (2015). A real-time articulatory visual feedback approach with target presentation for second language pronunciation learning. The Journal of the Acoustical Society of America, 138(4), EL382-EL387. DOI : 10.1121/1.4931827
  11. Yoon, K., & Kim, S. (2015). A comparative study on the male and female vowel formants of the Korean corpus of spontaneous speech. Phonetics and Speech Sciences, 7(2), 131-138. DOI: 10.13064/KSSS.2015.7.2.131