DETEKSI SPOILER PADA ULASAN BUKU BERBAHASA INDONESIA MENGGUNAKAN PENDEKATAN MACHINE LEARNING DAN DEEP LEARNING

Authors

  • Natasya Agustine Sadhi Tarumanagara University image/svg+xml
  • Hannah Larissa Halim Universitas Tarumanagara
  • Jessica Winola Universitas Tarumanagara
  • Viny C. Mawardi

DOI:

https://doi.org/10.33480/inti.v21i1.8587

Keywords:

BiLSTM, Class Imbalance, IndoBERT, Machine Learning, Spoiler Detection

Abstract

Spoiler detection in book reviews is a challenging text classification task because spoilers are not always identifiable from specific words but depend heavily on the narrative context in which information is revealed. This study compares six text classification models for detecting spoilers in Indonesian-language book reviews from Goodreads, namely Support Vector Machine (SVM), Random Forest, XGBoost, Bidirectional Long Short-Term Memory (BiLSTM), IndoBERT, and XLM-R (RoBERTa). Data were collected through Selenium-based web scraping and GraphQL API from 40 mystery and thriller book titles, resulting in 11,259 reviews with a class imbalance ratio of 1:10.4. All models were evaluated using AUC-ROC, PR-AUC, spoiler F1-score, and spoiler recall as primary metrics, with decision thresholds optimized through each model's validation set. XLM-R achieved the best overall performance with an AUC-ROC of 0.7017, a PR-AUC of 0.2263, and a spoiler F1-score of 0.2903, followed by IndoBERT, SVM, Random Forest, XGBoost, and BiLSTM. Overall, transformer-based models outperformed the traditional machine learning models and BiLSTM across most evaluation metrics. The results also indicate that each model exhibits different precision-recall characteristics, suggesting that model performance should not be evaluated using a single metric alone. These findings can serve as an initial reference for future research on Indonesian-language spoiler detection and support the development of automated spoiler detection systems for content moderation on digital book review platforms

Downloads

Download data is not yet available.

References

Al-Habib, H., Imah, E. M., Puspitasari, R. D. I., & Prahani, B. K. (2023). Text Processing Using Support Vector Machine for Scientific Research Paper Content Classification (pp. 273–282). https://doi.org/10.2991/978-94-6463-174-6_20

Bao, A., Ho, M., & Sangamnerkar, S. (2021). Spoiler Alert: Using Natural Language Processing to Detect Spoilers in Book Reviews. http://arxiv.org/abs/2102.03882

Brown, D. W. R. (2025). Spoilers, Narrative Pleasure, and Television. Projections (New York), 19(2), 25–43. https://doi.org/10.3167/proj.2025.190203

Chang, B., Lee, I., Kim, H., & Kang, J. (2021). “Killing Me” Is Not a Spoiler: Spoiler Detection Model using Graph Neural Networks with Dependency Relation-Aware Attention Mechanism. http://arxiv.org/abs/2101.05972

Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised Cross-lingual Representation Learning at Scale. https://github.com/facebookresearch/cc

Jamshidian, M. (2023). Evaluation of Text Transformers for Classifying Sentiment of Evaluation of Text Transformers for Classifying Sentiment of Reviews by Using TF-IDF, BERT (word embedding), SBERT Reviews by Using TF-IDF, BERT (word embedding), SBERT (sentence embedding) with Support Vector Machine Evaluation (sentence embedding) with Support Vector Machine Evaluation. https://arrow.tudublin.ie/scschcomdis

Karakaya, O., & Kilimci, Z. H. (2024). An efficient consolidation of word embedding and deep learning techniques for classifying anticancer peptides: FastText+BiLSTM. PeerJ Computer Science, 10. https://doi.org/10.7717/peerj-cs.1831

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. http://arxiv.org/abs/2011.00677

Lestari, V. B. , U. E. , & H. (2024). Combining Bi-LSTM And Word2vec Embedding For Sentiment Analysis Models of Application User Reviews. Indonesian Journal of Computer Science.

Lu, H., Ehwerhemuepha, L., & Rakovski, C. (2022). A comparative study on deep learning models for text classification of unstructured medical notes with various levels of class imbalance. BMC Medical Research Methodology, 22(1). https://doi.org/10.1186/s12874-022-01665-y

Lubis, A. R., Lase, Y. Y., Rahman, D. A., & Witarsyah, D. (2023). Improving Spell Checker Performance for Bahasa Indonesia Using Text Preprocessing Techniques with Deep Learning Models. Ingenierie Des Systemes d’Information, 28(5), 1335–1342. https://doi.org/10.18280/isi.280522

Suhaeni, C., Kamila, S. A., Fahira, F., Yusran, M., & Alfa Dito, G. (2025). Exploring a Large Language Model on the ChatGPT Platform for Indonesian Text Preprocessing Tasks. Indonesian Journal of Statistics and Its Applications, 9(1), 100–116. https://doi.org/10.29244/ijsa.v9i1p100-116

Wróblewska, A., Rzepiński, P., & Sysko-Romańczuk, S. (2021). Spoiler in a Textstack: How Much Can Transformers Help? http://arxiv.org/abs/2112.12913

Wyawhare, A. (2023). Comparative Analysis of Multilingual Text Classification & Identification through Deep Learning and Embedding Visualization.

Zevana, A., & Riana, D. (2024). TEXT CLASSIFICATION USING INDOBERT FINE-TUNING MODELING WITH CONVOLUTIONAL NEURAL NETWORK AND BI-LSTM. Jurnal Teknik Informatika (Jutif), 4(6), 1605–1610. https://doi.org/10.52436/1.jutif.2023.4.6.1650

Zhang, L. (2025). Features extraction based on Naive Bayes algorithm and TF-IDF for news classification. PLOS ONE, 20(7 July). https://doi.org/10.1371/journal.pone.0327347

Downloads

Published

2026-08-31

How to Cite

DETEKSI SPOILER PADA ULASAN BUKU BERBAHASA INDONESIA MENGGUNAKAN PENDEKATAN MACHINE LEARNING DAN DEEP LEARNING. (2026). INTI Nusa Mandiri, 21(1), 183-192. https://doi.org/10.33480/inti.v21i1.8587