IMPACT OF TEXT AUGMENTATION ON INDOBERT PERFORMANCE FOR HOSPITAL REVIEW SENTIMENT ANALYSIS
DOI:
https://doi.org/10.33480/jitk.v12i1.8327Keywords:
Easy Data Augmentation, Hospital Reviews, Imbalanced Dataset, IndoBERT, Sentiment AnalysisAbstract
Sentiment analysis of hospital patient reviews plays a critical role in evaluating healthcare service quality. However, limited labeled data and class imbalance often affect classification performance and reduce minority-class detection. This study empirically investigates the impact of text augmentation techniques on improving IndoBERT performance for sentiment classification of patient reviews at Dr. Mohammad Hoesin Palembang General Hospital. A total of 1,464 reviews were collected, preprocessed, and weakly labeled using a VADER-based approach, resulting in 1,168 positive and 296 negative instances. To address class imbalance, augmentation was applied exclusively to the training set using back-translation and Easy Data Augmentation (EDA), including synonym replacement, random insertion, random swap, and random deletion. IndoBERT was fine-tuned under consistent hyperparameter settings and evaluated using accuracy, precision, recall, F1-macro, and AUC. The baseline model achieved an F1-macro of 70.2%, indicating limited minority-class sensitivity. After augmentation, the random swap technique achieved the highest observed performance within the experimental setup, reaching 96.8% accuracy, 95.3% F1-macro, and 98.9% AUC. These results suggest improved performance within the experimental setting, particularly in minority-class detection. However, it should be noted that the labels were generated through a translation-based weak labeling approach, which may introduce noise and affect the accuracy of the labels. Therefore, the findings should be interpreted within the scope of this experimental setting.
Downloads
References
[1] T. Zitek, J. Bui, C. Day, S. Ecoff, and B. Patel, “A cross-sectional analysis of Yelp and Google reviews of hospitals in the United States,” JACEP Open, vol. 4, no. 2, p. e12913, 2023, doi: 10.1002/emp2.12913.
[2] A. R. Pandey, M. Seify, U. Okonta, and A. Hosseinian-Far, “Advanced Sentiment Analysis for Managing and Improving Patient Experience: Application for General Practitioner (GP) Classification in Northamptonshire,” Int. J. Environ. Res. Public Health, vol. 20, no. 12, Jun. 2023, doi: 10.3390/ijerph20126119.
[3] M. P. Tse, I. Dhalla, and D. Nayyar, “Google star ratings of Canadian hospitals: a nationwide cross-sectional analysis,” BMJ Open Qual., vol. 13, no. 3, Jul. 2024, doi: 10.1136/bmjoq-2023-002713.
[4] I. Villanueva-Miranda, Y. Xie, and G. Xiao, “Sentiment analysis in public health: a systematic review of the current state, challenges, and future directions,” Front. Public Heal., vol. 13, p. 1609749, 2025, doi: 10.3389/fpubh.2025.1609749.
[5] H. Imaduddin, F. Y. A’la, and Y. S. Nugroho, “Sentiment Analysis in Indonesian Healthcare Applications using IndoBERT Approach,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 8, pp. 113–117, 2023, doi: 10.14569/IJACSA.2023.0140813.
[6] O. S. Alkhnbashi, R. Mohammad, and M. Hammoudeh, “Aspect-Based Sentiment Analysis of Patient Feedback Using Large Language Models,” Big Data Cogn. Comput., vol. 8, no. 12, Dec. 2024, doi: 10.3390/bdcc8120167.
[7] Y. Mao, Q. Liu, and Y. Zhang, “Sentiment analysis methods, applications, and challenges: A systematic literature review,” J. King Saud Univ. - Comput. Inf. Sci., vol. 36, no. 4, p. 102048, 2024, doi: 10.1016/j.jksuci.2024.102048.
[8] N. C. Mei, S. Tiun, and G. Sastria, “Multi-label aspect-sentiment classification on Indonesian cosmetic product reviews with IndoBERT model,” International Journal of Advanced Computer Science and Applications, vol. 15, no. 11, 2024, doi: 10.14569/IJACSA.2024.0151168.
[9] J. R. Jim, M. A. R. Talukder, P. Malakar, M. M. Kabir, K. Nur, and M. F. Mridha, “Recent advancements and challenges of NLP-based sentiment analysis: A state-of-the-art review,” Nat. Lang. Process. J., vol. 6, p. 100059, Mar. 2024, doi: 10.1016/j.nlp.2024.100059.
[10] S. H. Park, C. P. Cheng, N. J. Buehler, T. Sanford, and W. Torrey, “A sentiment analysis on online psychiatrist reviews to identify clinical attributes of psychiatrists that shape the therapeutic alliance,” Front. Psychiatry, vol. 14, p. 1174154, 2023, doi: 10.3389/fpsyt.2023.1174154.
[11] L. Gandy, L. Ivanitskaya, L. Bacon, and R. Bizri-Baryak, “Public health discussions on social media: evaluating automated sentiment analysis methods,” JMIR Formative Research, vol. 9, Art. no. e57395, 2025, doi: 10.2196/57395.
[12] I. Aggarwal, S. Joseph, N. Jaganathan, A. Patel, V. Kumar, and M. Devarapalli, “Sentiment Analysis in Healthcare: A Comparison of VADER, BERT, and Flair NLP Models on Patient Reviews of Pain Management Physicians,” Cureus, vol. 17, no. 7, 2025, doi: 10.7759/cureus.88902.
[13] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, "IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP," in Proc. 28th Int. Conf. Comput. Linguist. (COLING), Barcelona, Spain, 2020, pp. 757–770, doi: 10.18653/v1/2020.coling-main.66.
[14] R. Kimera, D. N. Heo, D. N. Rim, and H. Choi, “Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda,” in NLPIR 2024 - 2024 8th International Conference on Natural Language Processing and Information Retrieval, 2025, pp. 142–148. doi: 10.1145/3711542.3711594.
[15] Natasya and A. S. Girsang, “Modified EDA and Backtranslation Augmentation in Deep Learning Models for Indonesian Aspect-Based Sentiment Analysis,” Emerg. Sci. J., vol. 7, no. 1, pp. 256–272, 2023, doi: 10.28991/ESJ-2023-07-01-018.
[16] J. Wei and K. Zou, "EDA: Easy data augmentation techniques for boosting performance on text classification tasks," in Proc. 2019 Conf. Empirical Methods Natural Lang. Process. and 9th Int. Joint Conf. Natural Lang. Process. (EMNLP-IJCNLP), Hong Kong, China, 2019, pp. 6382–6388, doi: 10.18653/v1/D19-1670.
[17] J. Chen, D. Tam, C. Raffel, M. Bansal, and D. Yang, “An Empirical Survey of Data Augmentation for Limited Data Learning in NLP,” Trans. Assoc. Comput. Linguist., vol. 11, pp. 191–211, 2023, doi: 10.1162/tacl_a_00542.
[18] S. Biswas, K. Young, and J. Griffith, "A Comparison of Automatic Labelling Approaches for Sentiment Analysis," in Proc. 14th Int. Joint Conf. on Knowledge Discovery, Knowledge Engineering and Knowledge Management (IC3K 2022) - KDIR, Valletta, Malta, 2022, pp. 195–202, doi: 10.5220/0011265900003269.
[19] J. Opitz, “A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation Practice,” Trans. Assoc. Comput. Linguist., vol. 12, pp. 820–836, 2024, doi: 10.1162/tacl_a_00675.
[20] M. Siino, I. Tinnirello, and M. La Cascia, “Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers,” Inf. Syst., vol. 121, Mar. 2024, doi: 10.1016/j.is.2023.102342.
[21] M. A. Palomino and F. Aider, “Evaluating the Effectiveness of Text Pre-Processing in Sentiment Analysis,” Appl. Sci., vol. 12, no. 17, Sep. 2022, doi: 10.3390/app12178765.
[22] S. Jahić and J. Vičič, “Impact of Negation and AnA-Words on Overall Sentiment Value of the Text Written in the Bosnian Language,” Appl. Sci., vol. 13, no. 13, Jul. 2023, doi: 10.3390/app13137760.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Yoga Anugrah Pratama.SY, Ali Ibrahim

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






-a.jpg)
-b.jpg)











