Comparison of LDA and Topical Methods in Topic Modeling Against Emotion Labels on Twitter Data #Indonesiagelap
Keywords:
nlp, emotion, lda, bertopic, twitterAbstract
This study analyzes public discussions on Twitter regarding the #IndonesiaGelap hashtag to identify emotional patterns and major topics in socio-political discourse. A total of 10,866 tweets were collected and processed through standard preprocessing steps. Text representation was performed using TF-IDF and Sentence Transformer, which served as inputs for topic modeling with LDA and BERTopic. Emotion detection using EmoSense-ID revealed a dominance of negative emotions—particularly sadness, anger, and disgust—indicating widespread public concern and dissatisfaction. LDA produced 10 topics with an average coherence score of 0.342, while BERTopic identified 39 topics with generally higher coherence, demonstrating its ability to capture more nuanced subthemes. Cosine similarity evaluation further showed that embedding-based representations effectively captured semantic closeness among tweets. Overall, this study provides a concise yet comprehensive overview of public perceptions surrounding #IndonesiaGelap through combined emotion and topic modeling approaches.


