Robust Acoustic Pattern Recognition for Low Resource Automatic Speech Recognition System
Hassan Al Tamimi and Sara Hosseini
Abstract
ASR systems have been highly successful in high-resource languages yet less effective in low-resource languages due to limited labeled speech data, speaker variability, background noise, and variable acoustic conditions. The significant barrier to accessing and deploying speech-enabled technologies for underrepresented languages and dialects, thereby limiting human–computer interaction. To overcome these shortcomings, we suggest a sophisticated acoustic pattern recognition system that combines advanced acoustic feature extraction, data augmentation and deep neural network-based acoustic modelling to improve speech recognition in low-resource environments. The proposed approach exploits spectro-temporal representations and attention-driven learning mechanisms to learn discriminative acoustic patterns while making the model robust to noise and pronunciation variations. One important advantage of the framework is its effective generalization from limited training samples, enabled by optimized feature learning and transfer learning strategies. Experiments were conducted using a low-resource speech corpus comprising about 120 hours of transcribed speech from several speakers across a variety of acoustic conditions. The evaluation was conducted using Word Error Rate (WER), Character Error Rate (CER), and recognition accuracy. The proposed model's WER, CER, and overall recognition accuracy were 11.8%, 5.4%, and 92.6%, respectively, outperforming the conventional baseline ASR models, which had a WER of 18.7% and an accuracy of 84.3%. The experimental results demonstrated the effectiveness of the proposed framework in enhancing the robustness of recognition and acoustic pattern discrimination. The advances the development of reliable, scalable, and low-cost ASR systems for low-resource languages, thereby aiding greater language coverage and the practical implementation of speech technologies in real-world settings.