Valence-Arousal Modelling for Affective Analysis of Social Media Discourse on Sinhala Music Videos (Defence Practice)
Music Emotion Recognition and affective computing have become important research areas within Music Information Retrieval due to their applications in recommendation systems, human computer interaction, and computational social analysis. However, research in low resource languages such as Sinhala remains limited because of the scarcity of annotated datasets, linguistic resources, and robust computational models. This thesis investigates affective analysis of Sinhala social media discourse related to Sinhala music videos through a unified framework integrating transliteration, linguistic analysis, dataset construction, and continuous emotion modeling.
The study first addresses the challenge of Romanized Sinhala text commonly used in social media communication. Two transliteration approaches are explored, including a rule based framework and a Transformer based sequence to sequence model. Experimental results demonstrate that the deep learning based approach is more effective in handling informal and inconsistent Romanized Sinhala patterns.
A large scale Sinhala YouTube comment corpus was subsequently developed using comments collected from Sinhala music videos across multiple genres and eras. Through rigorous preprocessing, filtering, normalization, and transliteration, a high quality Sinhala language corpus was created for linguistic and affective analysis. Building upon this resource, the thesis introduces GeeSanBhava, a manually annotated emotion dataset labeled according to the Valence Arousal framework by three independent annotators with high inter annotator agreement.
The thesis further investigates computational emotion modeling for Sinhala social media comments. While previous studies primarily focused on discrete emotion classification, this research formulates emotion prediction as a continuous Valence Arousal regression problem to better capture nuanced emotional expressions. To address the challenges of low resource social media text, a hybrid multi task learning framework combining contextual transformer embeddings, static embeddings, sentence representations, and large language model based representations is proposed. Extensive experiments and ablation studies demonstrate the effectiveness of the proposed framework for continuous affect prediction in Sinhala.
Overall, this thesis contributes novel transliteration methodologies, large scale linguistic resources, annotated affective datasets, and hybrid representation learning approaches for Sinhala Natural Language Processing and Music Emotion Recognition. The findings establish a strong foundation for future research in low resource affective computing and multilingual emotion analysis.