SARCASM DETECTION IN PALEMBANG LANGUAGE USING HYBRID CLASSICAL AND DEEP LEARNING APPROACHES

Authors

  • Johannes Petrus Universitas Multi Data Palembang
  • Antonius Wahyu Sudrajat Universitas Multi Data Palembang
  • Muhammad Rachmadi Universitas Multi Data Palembang

DOI:

https://doi.org/10.33480/jitk.v12i1.8222

Keywords:

Cross-Validation, Low-Resource NLP, Palembang, Sarcasm Detection, T-Friedman Test

Abstract

Research on sarcasm detection in sentences generally focuses on English and the national languages of certain countries, so that regional languages such as Palembang are still underrepresented in NLP research. This study aims to classify sarcasm in Palembang sentences by applying 12 models and evaluating their performance using a dataset containing 1,952 manually annotated samples. The experiment was also conducted using a 5-fold cross-validation scheme, and the statistical significance of the test results was analyzed using the Friedman t-test ( =41.96, FF=12.90). Contrary to common expectations, the Ensemble classifier achieved the best overall performance, attaining a mean F1-score of 0.877 and outperforming larger pre-trained and dialect-specific models. Even though it is still in the same language family, MelayuBERT's performance is very far behind (F1:0.691). Furthermore, the relatively lower performance of IndoBERT-1.5G (0.731) suggests that model scale alone does not guarantee effectiveness in low-resource language settings. These findings highlight the robustness of feature-engineered classical models compared to large-scale or dialect-specific pre-trained models for sarcasm detection in low-resource regional languages. Our research results with datasets sourced from varied domains have exceeded several national benchmarks. This study can be a baseline for further research.

Downloads

Download data is not yet available.

References

[1] E. Andina, “Implementasi dan Tantangan Revitalisasi Bahasa Daerah di Provinsi Lampung,” Aspir. J. Masal. Sos., vol. 14, no. 1, Jun. 2023, doi: 10.46807/aspirasi.v14i1.3859.

[2] Izzati, W. A. Rais, and H. Yustanto, “The Impact of Cultural Contact Between Javanese and Palembang Malay On the Language System In the City of Palembang,” Ichss, vol. 2, pp. 1–11, 2022.

[3] H. L. H. S. Warnars, J. Aurellia, and K. Saputra, “Translation Learning Tool for Local Language to Bahasa Indonesia using Knuth-Morris-Pratt Algorithm,” TEM J., vol. 10, no. 1, pp. 55–62, 2021, doi: 10.18421/TEM101-07.

[4] N. H. Jeremy, “The Impact of Text Preprocessing in Sarcasm Detection on Indonesian Social Media Contents,” Eng. Math. Comput. Sci. J., vol. 7, no. 2, pp. 183–189, 2025, doi: 10.21512/emacsjournal.v7i2.13503.

[5] Z. Wen et al., “Sememe knowledge and auxiliary information enhanced approach for sarcasm detection,” Inf. Process. Manag., vol. 59, no. 3, p. 102883, 2022, doi: 10.1016/j.ipm.2022.102883.

[6] A. Alqahtani, A. Alsheddi, and L. Alhenaki, “Text-based Sarcasm Detection on Social Networks: A Systematic Review,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 3, pp. 313–328, 2023, doi: 10.14569/IJACSA.2023.0140336.

[7] F. B. Kader, N. H. Nujat, T. B. Sogir, M. Kabir, H. Mahmud, and K. Hasan, “Computational Sarcasm Analysis on Social Media: A Systematic Review,” Sep. 2022, [Online]. Available: http://arxiv.org/abs/2209.06170

[8] Y. Y. Tan, C. O. Chow, J. Kanesan, J. H. Chuah, and Y. L. Lim, “Sentiment Analysis and Sarcasm Detection using Deep Multi-Task Learning,” Wirel. Pers. Commun., vol. 129, no. 3, pp. 2213–2237, 2023, doi: 10.1007/s11277-023-10235-4.

[9] A. C. Băroiu and Ștefan Trăușan-Matu, “Automatic Sarcasm Detection: Systematic Literature Review,” Inf., vol. 13, no. 8, pp. 1–17, 2022, doi: 10.3390/info13080399.

[10] S. H. Suhaimi, N. A. A. Bakar, and N. F. Nurulhuda, “Constructing and Analysing the MalaySarc Dataset: A Resource for Detecting and Understanding Sarcasm in Malay Language,” Proc. World Congr. Electr. Eng. Comput. Syst. Sci., pp. 1–10, 2023, doi: 10.11159/cist23.126.

[11] R. Misra and P. Arora, “Sarcasm detection using news headlines dataset,” AI Open, vol. 4, no. October 2022, pp. 13–18, 2023, doi: 10.1016/j.aiopen.2023.01.001.

[12] M. A. Rosid, M. A. K. Sambada, S. Busono, and F. Muharram, “Ensemble Machine Learning to Detect Sarcasm in English on Twitter Social Media,” J. Informatics Inf. Syst. Softw. Eng. Appl., vol. 6, no. 1, pp. 11–20, 2023, doi: 10.20895/inista.v6i1.1073.

[13] D. Šandor and M. Bagić Babac, “Sarcasm detection in online comments using machine learning,” Inf. Discov. Deliv., vol. 52, no. 2, pp. 213–226, 2024, doi: 10.1108/IDD-01-2023-0002.

[14] D. K. Sharma, B. Singh, S. Agarwal, N. Pachauri, A. A. Alhussan, and H. A. Abdallah, “Sarcasm Detection over Social Media Platforms Using Hybrid Ensemble Model with Fuzzy Logic,” Electronics, vol. 12, no. 4, p. 937, Feb. 2023, doi: 10.3390/electronics12040937.

[15] I. A. Ahmad, P. Gatla, and R. K. Mundotiya, “Sarcasm Identification and Classification in Hindi Newspaper Headlines,” ACM Trans. Asian Low-Resource Lang. Inf. Process., vol. 24, no. 4, pp. 1–21, Apr. 2025, doi: 10.1145/3714469.

[16] K. T. Ladoja and R. T. Afape, “Sarcasm Detection in Pidgin Tweets Using Machine Learning Techniques,” Asian J. Res. Comput. Sci., vol. 17, no. 5, pp. 212–221, 2024, doi: 10.9734/ajrcos/2024/v17i5450.

[17] B. R. Chakravarthi, “Sarcasm Detection in Tamil and Malayalam YouTube Comments,” Soc. Netw. Anal. Min., vol. 15, no. 1, pp. 1–18, 2025, doi: 10.1007/s13278-025-01486-z.

[18] D. Suhartono, W. Wongso, and A. Tri Handoyo, “IdSarcasm: Benchmarking and Evaluating Language Models for Indonesian Sarcasm Detection,” IEEE Access, vol. 12, no. May, pp. 87323–87332, 2024, doi: 10.1109/ACCESS.2024.3416955.

[19] J. Amalia, D. F. Matondang, G. E. M. Hutajulu, and A. Hasibuan, “Impact Of Sarcasm Detection on Sentiment Analysis Using Bi-LSTM and FastText,” J. Sist. Inf. Bisnis, vol. 14, no. 4, pp. 353–362, 2024, doi: 10.21456/vol14iss4pp353-362.

[20] R. Kusumastuti, E. Utami, and A. Yaqin, “Detection of Sarcasm Sentences in Indonesian Tweets using SentiStrength,” Proceeding - 6th Int. Conf. Inf. Technol. Inf. Syst. Electr. Eng. Appl. Data Sci. Artif. Intell. Technol. Environ. Sustain. ICITISEE 2022, pp. 93–98, 2022, doi: 10.1109/ICITISEE57756.2022.10057904.

[21] A. F. Aji et al., “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia,” Proc. Annu. Meet. Assoc. Comput. Linguist., vol. 1, pp. 7226–7249, 2022, doi: 10.18653/v1/2022.acl-long.500.

[22] J. Petrus, Ermatita, Sukemi, and Erwin, “An adaptable sentence segmentation based on Indonesian rules,” IAES Int. J. Artif. Intell., vol. 12, no. 3, pp. 1491–1499, 2023, doi: 10.11591/ijai.v12.i3.pp1491-1499.

[23] A. A. Aliero, B. S. Adebayo, H. O. Aliyu, A. G. Tafida, B. U. Kangiwa, and N. M. Dankolo, “Systematic Review on Text Normalization Techniques and its Approach to Non-Standard Words,” Int. J. Comput. Appl., vol. 185, no. 33, pp. 44–55, 2023, doi: 10.5120/ijca2023923106.

[24] G. Zeng, “Invariance Properties and Evaluation Metrics Derived from the Confusion Matrix in Multiclass Classification,” Mathematics, vol. 13, no. 16, 2025, doi: 10.3390/math13162609.

[25] A. R. Aditama and A. F. Wicaksono, “Classification of customer complaints on social media for e-commerce in Indonesia,” Int. J. Electr. Comput. Eng., vol. 15, no. 3, p. 2977, 2025, doi: 10.11591/ijece.v15i3.pp2977-2985.

[26] M. Martinović, K. Dokic, and D. Pudić, “Comparative Analysis of Machine Learning Models for Predicting Innovation Outcomes: An Applied AI Approach,” Appl. Sci., vol. 15, no. 7, pp. 1–44, 2025, doi: 10.3390/app15073636.

[27] J. Liu and Y. Xu, “T-Friedman Test: A New Statistical Test for Multiple Comparison with an Adjustable Conservativeness Measure,” Int. J. Comput. Intell. Syst., vol. 15, no. 1, pp. 1–19, 2022, doi: 10.1007/s44196-022-00083-8.

[28] Wella, R. I. Desanti, and Suryasari, “A Comparative Study of Machine Learning Approaches to Megathrust Earthquake Prediction in Subduction Zones,” J. Appl. Data Sci., vol. 6, no. 4, pp. 2421–2435, 2025, doi: 10.47738/jads.v6i4.904.

Downloads

Published

2026-08-11

How to Cite

[1]
“SARCASM DETECTION IN PALEMBANG LANGUAGE USING HYBRID CLASSICAL AND DEEP LEARNING APPROACHES”, jitk, vol. 12, no. 1, pp. 40–53, Aug. 2026, doi: 10.33480/jitk.v12i1.8222.