ORIGINAL ARTICLE
Figure from article: Evaluation of Azure...
 
KEYWORDS
TOPICS
ABSTRACT
Accurate reading aloud is crucial for Arabic-speaking learners because minor errors can alter meaning or hinder understanding. This study evaluated the Microsoft Azure Pronunciation Assessment (PA) tool for modern standard Arabic using the KSU Arabic Speech Database. We focused on male speakers recorded in silent rooms. We used the KSU database’s varied prompts words, phrases, and sentences to test Azure PA at the word, phoneme, and sentence levels. Because Azure PA is a closed source, we approximated its scoring by deriving equations for the Accuracy, Completeness, and Fluency metrics. We showed that a simple weighted combination closely matched the Pronunciation scores. We also applied a standard audio preprocessing pipeline (channel selection, denoising, pre-emphasis, silence trimming, and loudness normalization) to measure its impact. With preprocessing, for male speakers in the KSU database under the no-reference mode, Accuracy increased from 91.74% to 95.33%, Fluency from 60.84% to 97.84%, Completeness from 97.16% to 100.0%, and Pronunciation from 73.64% to 96.26%. The number of decoded word tokens increased fivefold. After preprocessing, all the consonant classes reached 90% accuracy, although the glottal segments remained weak. Sentence-level tests showed that Azure PA originally captured only brief fragments of long recordings. After preprocessing, it covered a greater portion of each sentence and recovered more of the target text. These results demonstrate both the strengths and limitations of Azure PA as a diagnostic tool for Arabic pronunciation based on KSU data.
Journals System - logo
Scroll to top