Authors: Assistant Professor K.Rajkumar, J.Akshaya

Abstract: When a person has trouble communicating, it may have a negative impact on their social life, academic performance, and mental health, all of which contribute to a worse quality of life overall. Stuttering, vocal tremor, and dysarthria are three of the most common speech disorders; these conditions have different acoustic features and need to be properly diagnosed in order to provide effective therapeutic treatment. The assessment method is subjective, time-consuming, and reliant on clinical skill in conventional diagnostic processes since it mostly relies on perceptual evaluation undertaken by speech-language pathologists. This has led to a rise in the need for sophisticated AI-powered diagnostic tools that can accurately and consistently categorise speech disorders. Opportunities for the development of automated diagnostic frameworks that analyse complicated audio patterns with higher accuracy have arisen due to recent breakthroughs in digital voice processing and Machine Learning (ML). This research introduces a system that uses machine learning to automatically categorise speech problems based on audio signal analysis. The suggested method augments model generalisability and increases dataset variety by combining real-world voice recordings with synthetically produced speech signals produced by MATLAB-based mathematical modelling, thus overcoming the barrier of limited clinical speech datasets. Important acoustic parameters, such as fundamental frequency (F0), formant frequencies (F1 and F2), and signal augmentation and normalisation are performed on speech recordings during preprocessing in order to extract them. These properties characterise normal and disordered speech patterns, respectively. After that, several supervised machine learning algorithms, such as Gradient Boosting, Support Vector Machine (SVM), and Random Forest, are trained on the retrieved characteristics to categorise speech samples into various disorders. Traditional regression-based performance metrics such as Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Coefficient of Determination (R²) are used to assess the built models, in conjunction with an 80% training and 20% testing approach. Under the specified experimental circumstances, Random Forest and Gradient Boosting outperform Support Vector Machine in classification performance, with an overall accuracy of 62.50%, according to the experimental study. Based on the results, ensemble learning approaches outperform traditional classifiers when it comes to complicated pathological speech patterns. Objective speech disorder diagnosis, early clinical intervention, and enhanced telemedicine-based speech assessment systems for improved patient care are all possible outcomes of the suggested framework's integration of machine learning, acoustic signal processing, and synthetic data generation.

DOI: https://doi.org/10.5281/zenodo.21626323