WEIGHTED LOSS STRATEGY FOR BERT-BASED TWITTER SENTIMENT ANALYSIS WITHOUT SYNTHETIC OVERSAMPLING
DOI:
https://doi.org/10.33480/jitk.v12i1.7680Keywords:
BERT, ChatGPT, Class Imbalance, Deep Learning, Natural Language ProcessingAbstract
The widespread adoption of ChatGPT has generated extensive public discourse across social media, necessitating robust sentiment analysis to understand collective opinions. Traditional approaches frequently employ the Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalance; however, its effectiveness on short-text data remains an open question. This study develops an optimized sentiment classification model and evaluates whether competitive performance can be achieved without synthetic data augmentation. The methodology encompasses comprehensive Natural Language Processing (NLP) preprocessing and stratified data partitioning to preserve distributional characteristics. A BERT-base architecture is fine-tuned using a class-weighted Cross-Entropy loss combined with weighted random sampling, deliberately avoiding SMOTE-based oversampling. The model is trained with the AdamW optimizer (learning rate: 3 × 10⁻⁵), batch size 32, and mixed-precision training for four epochs. On 198,639 preprocessed tweets, the proposed approach achieves 93.81% accuracy, with weighted precision, recall, and F1-score of 0.9365, 0.9381, and 0.9380 respectively, outperforming the baseline by 1.75 percentage points. Per-class analysis reveals strong performance for negative (F1-score: 0.96) and positive sentiment (F1-score: 0.94), with lower neutral classification (F1-score: 0.89), attributable to the inherent heterogeneity of neutral expressions. The training-validation gap remains below 5%, consistent with adequate regularization. These findings provide empirical evidence that, within the present experimental configuration, a properly optimized weighted loss strategy offers a viable and computationally efficient alternative to synthetic oversampling for BERT-based Twitter sentiment classification. Further controlled ablation studies and statistical validation are needed to establish generalizability.
Downloads
References
[1] L. Marron, “Exploring the potential of ChatGPT 3.5 in higher education: Benefits, limitations, and academic integrity,” in Handbook of Research on Redesigning Teaching, Learning, and Assessment in the Digital Era, IGI Global, 2023, pp. 326–349. doi: 10.4018/978-1-6684-8292-6.ch017.
[2] F. W. Putra, I. B. Rangka, S. Aminah, and M. H. R. Aditama, “ChatGPT in the higher education environment: Perspectives from the theory of high-order thinking skills,” Journal of Public Health (United Kingdom), vol. 45, no. 4, pp. e840-e841, Dec. 2023, doi: 10.1093/pubmed/fdad120.
[3] T. Adiguzel, M. H. Kaya, and F. K. Cansu, “Revolutionizing education with AI: Exploring the transformative potential of ChatGPT,” Contemporary Educational Technology, vol. 15, no. 3, 2023, doi: 10.30935/cedtech/13152.
[4] P. Shah, H. Patel, and P. Swaminarayan, “Multitask Sentiment Analysis and Topic Classification Using BERT,” ICST Transactions on Scalable Information Systems, vol. 11, Jul. 2024, doi: 10.4108/eetsis.5287.
[5] L. He, “Enhanced Twitter sentiment analysis with dual joint classifier integrating RoBERTa and BERT architectures,” Frontiers in Physics, vol. 12, 2024, doi: 10.3389/fphy.2024.1477714.
[6] A. Areshey and H. Mathkour, "Exploring transformer models for sentiment classification: A comparison of BERT, RoBERTa, ALBERT, DistilBERT, and XLNet," Expert Systems, vol. 41, no. 11, p. e13701, 2024, doi: 10.1111/exsy.13701.
[7] H. Murfi, S. T. Gowandi, G. Ardaneswari, and S. Nurrohmah, "BERT-based combination of convolutional and recurrent neural network for indonesian sentiment analysis," Applied Soft Computing, vol. 151, p. 111112, 2024, doi: 10.1016/j.asoc.2023.111112.
[8] S. Uyun, R. P. Rosalin, L. V. Sari, and H. H. Sucinta, "A Hybrid Classification Model Based on BERT for Multi-Class Sentiment Analysis on Twitter," Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, vol. 11, no. 2, pp. 1-12, Jun. 2025, doi: 10.26555/jiteki.v11i2.29267.
[9] F. M. Sinaga, R. Purba, S. J. Pipin, W. S. Lestari, and S. Winardi, “Optimization of Sentiment Analysis Classification of ChatGPT on Big Data Twitter in Indonesia using BERT,” JURNAL MEDIA INFORMATIKA BUDIDARMA, vol. 8, no. 3, p. 1665, Jul. 2024, doi: 10.30865/mib.v8i3.7861.
[10] A. Subakti, H. Murfi, and N. Hariadi, “The performance of BERT as data representation of text clustering,” Journal of Big Data, vol. 9, no. 1, Dec. 2022, doi: 10.1186/s40537-022-00564-9.
[11] A. Rajan and M. Manur, “Aspect based sentiment analysis using fine-tuned BERT model with deep context features,” IAES International Journal of Artificial Intelligence, vol. 13, no. 2, pp. 1250–1261, Jun. 2024, doi: 10.11591/ijai.v13.i2.pp1250-1261.
[12] J. Ma, “Using the BERT model and the attention mechanism to obtain an accurate sentiment analysis model,” Applied and Computational Engineering, vol. 43, 2024, doi: 10.54254/2755-2721/43/20241234.
[13] P. Akter et al., “Sentiment Analysis of Consumer Feedback and Its Impact on Business Strategies by Machine Learning,” The American Journal of Applied Sciences, vol. 7, no. 1, pp. 6–16, Jan. 2025, doi: 10.37547/tajas/Volume07Issue01-02.
[14] N. S. Suryawanshi, “Sentiment analysis with machine learning and deep learning: A survey of techniques and applications,” International Journal of Science and Research Archive, vol. 12, no. 2, pp. 005–015, Jul. 2024, doi: 10.30574/ijsra.2024.12.2.1205.
[15] S. Efendi and P. Sihombing, “Sentiment Analysis of Food Order Tweets to Find Out Demographic Customer Profile Using SVM,” MATRIK: Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer, vol. 21, no. 3, pp. 583–594, Jul. 2022, doi: 10.30812/matrik.v21i3.1898.
[16] F. M. Sinaga, S. J. Pipin, S. Winardi, K. M. Tarigan, and A. P. Brahmana, “Analyzing Sentiment with Self-Organizing Map and Long Short-Term Memory Algorithms,” MATRIK: Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer, vol. 23, no. 1, pp. 131–142, Nov. 2023, doi: 10.30812/matrik.v23i1.3332.
[17] S. J. Pipin, F. M. Sinaga, S. Winardi, and M. N. Hakim, “Sentiment Analysis Classification of ChatGPT on Twitter Big Data in Indonesia Using Fast R-CNN,” JURNAL MEDIA INFORMATIKA BUDIDARMA, vol. 7, no. 4, p. 2137, Oct. 2023, doi: 10.30865/mib.v7i4.6816.
[18] L. Geni, E. Yulianti, and D. I. Sensuse, "Sentiment Analysis of Tweets Before the 2024 Elections in Indonesia Using IndoBERT Language Models," Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, vol. 9, no. 3, pp. 746-757, Aug. 2023, doi: 10.26555/jiteki.v9i3.26490.
[19] J. Lu and H. Gweon, “Random k conditional nearest neighbor for high-dimensional data,” PeerJ Computer Science, vol. 11, 2025, doi: 10.7717/PEERJ-CS.2497.
[20] O. Ndama, I. Bensassi, and E. M. En-Naimi, “The impact of BERT-infused deep learning models on sentiment analysis accuracy in financial news,” Bulletin of Electrical Engineering and Informatics, vol. 14, no. 2, pp. 1231–1240, Apr. 2025, doi: 10.11591/eei.v14i2.8469.
[21] J. Sun, M. Wang, D. Ren, and D. Chen, “Research and Application of Text-Based Sentiment Analytics,” in Frontiers in Artificial Intelligence and Applications, IOS Press BV, 2024, pp. 619–629. doi: 10.3233/FAIA241391.
[22] Z. Su, “Applications of BERT in sentimental analysis,” Applied and Computational Engineering, vol. 92, no. 1, pp. 147–152, Oct. 2024, doi: 10.54254/2755-2721/92/20241711.
[23] A. Bello, S. C. Ng, and M. F. Leung, “A BERT Framework to Sentiment Analysis of Tweets,” Sensors, vol. 23, no. 1, Jan. 2023, doi: 10.3390/s23010506.
[24] P. Tisna Putra, A. Anggrawan, and H. Hairani, “Comparison of Machine Learning Methods for Classifying User Satisfaction Opinions of the PeduliLindungi Application,” MATRIK: Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer, vol. 22, no. 3, pp. 431–442, Jun. 2023, doi: 10.30812/matrik.v22i3.2860.
[25] F. Sinaga, S. Winardi, and Gunawan, “3SV-KNNC Optimization using SVR and LMKNN for Stock Price Prediction,” in 2022 4th International Conference on Cybernetics and Intelligent System (ICOSNIKOM), Jan. 2022, pp. 1–6. doi: 10.1109/ICOSNIKOM56551.2022.10034892.
[26] H. Mulyani, R. A. Setiawan, and H. Fathi, “Optimization of K Value in Clustering Using Silhouette Score (Case Study: Mall Customers Data),” Journal of Information Technology and Its Utilization, vol. 6, no. 2, pp. 45–50, Dec. 2023, doi: 10.56873/jitu.6.2.5243.
[27] Q. Hu, “A cross-language short text classification model based on BERT and multilayer collaborative convolutional neural network (MCNN),” MCB Molecular and Cellular Biomechanics, vol. 21, no. 3, 2024, doi: 10.62617/mcb739.
[28] Z. Wang, L. Wang, C. Huang, S. Sun, and X. Luo, "BERT-based Chinese text classification for emergency management with a novel loss function," Applied Intelligence, vol. 53, no. 9, pp. 10417-10428, 2023, doi: 10.1007/s10489-022-04138-3.
[29] S. Henning, W. Beluch, A. Fraser, and A. Friedrich, "A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language Processing," in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, 2023, pp. 523-540, doi: 10.18653/v1/2023.eacl-main.38.
[30] D. Elreedy, A. F. Atiya, and F. Kamalov, "A theoretical distribution analysis of synthetic minority oversampling technique (SMOTE) for imbalanced learning," Machine Learning, vol. 113, no. 7, pp. 4903-4923, 2024, doi: 10.1007/s10994-022-06296-4.
[31] K. Kelvin, F. M. Sinaga, W. S. Lestari, S. Winardi, K. H. Rambe, and R. Purba, “Enhancing Sentiment Analysis Accuracy with BERT and Silhouette Method Optimization,” JITK (Jurnal Ilmu Pengetahuan dan Teknologi Komputer), vol. 11, no. 1, Aug. 2025, doi: 10.33480/jitk.v11i1.6392.
[32] Z. Zhuang, M. Liu, A. Cutkosky, and F. Orabona, "Understanding AdamW through Proximal Methods and Scale-Freeness," Transactions on Machine Learning Research, 2022. [Online]. Available: https://openreview.net/forum?id=IKhEPWGdwK.
[33] A. R. Lubis, Y. Fatmi, and D. Witarsyah, "Comparing transformer-based and traditional models for sentiment analysis on social media datasets," in 2023 6th International Conference of Computer and Informatics Engineering (IC2IE), 2023, pp. 1-6, doi: 10.1109/IC2IE60547.2023.10331116.
[34] A. Joshy and S. Sundar, "Analyzing the Performance of Sentiment Analysis using BERT, DistilBERT, and RoBERTa," in 2022 IEEE International Power and Renewable Energy Conference (IPRECON), 2022, pp. 1-6, doi: 10.1109/IPRECON55716.2022.10059542.
[35] S. Kiritchenko and M. Rezapour, "A comparative study of sentiment classification with transformer-based models on social media text," in Proceedings of the 2024 International Conference on Computing and Informatics, 2024, pp. 1-8.
[36] N. R. Aljohani, A. Fayoumi, and S. -U. Hassan, "A novel focal-loss and class-weight-aware convolutional neural network for the classification of in-text citations," Journal of Information Science, vol. 49, no. 1, pp. 79-92, 2023, doi: 10.1177/0165551521991022.
[37] B. P. Nemade, V. Bharadi, S. S. Alegavi, and B. Marakarkandy, "A Comprehensive Review: SMOTE-Based Oversampling Methods for Imbalanced Classification Techniques, Evaluation, and Result Comparisons," International Journal of Intelligent Systems and Applications in Engineering, vol. 11, no. 9s, pp. 790-803, 2023.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Timbo Faritcan Siallagan, Riki Winanjaya, Juni Ismail

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






-a.jpg)
-b.jpg)











