Deteksi Ujaran Kebencian Bahasa Indonesia Menggunakan Fine-Tuning IndoBERT Berbasis Transformer
DOI:
https://doi.org/10.61805/fahma.v24i3.234Keywords:
Ujaran Kebencian, Bahasa Indonesia, IndoBERT, Transformer, Fine-tuningAbstract
Hate speech on Indonesian social media continues to increase, creating a need for automated systems capable of accurately classifying textual content. This study proposes hate speech detection using the pre-trained IndoBERT transformer model through fine-tuning. Unlike previous studies that combine transformers with additional architectures such as CNN or BiLSTM, this study directly evaluates IndoBERT without additional deep learning layers. The dataset consists of 13,169 Indonesian tweets annotated into Hate Speech (HS) and Non-Hate Speech (Non-HS) categories. The dataset was divided into 80% training, 10% validation, and 10% testing data. The model was trained for three epochs using the AdamW optimizer with a learning rate of 2e-5 and a batch size of 8. Experimental results show that IndoBERT achieved 90.43% accuracy, 90.11% precision, 90.36% recall, and a 90.23% F1-score. These results demonstrate that direct fine-tuning of IndoBERT can achieve strong classification performance with a simpler architecture, supporting its potential for automated digital content moderation.
Downloads
References
K. Saha, E. Chandrasekharan, and M. de Choudhury, “Prevalence and psychological effects of hateful speech in online college communities,” 2019. doi: 10.1145/3292522.3326032.
G. Kovács, P. Alonso, and R. Saini, “Challenges of hate speech detection in social media,” 2021, doi: 10.1007/s42979-021-00457-3.
H. Margono, M. Saud, and A. Ashfaq, “Dynamics of hate speech in social media: Insights from Indonesia,” 2024, doi: 10.1108/GKMC-11-2023-0464.
M. O. Ibrohim and I. Budi, “Multi-label hate speech and abusive language detection in Indonesian Twitter,” 2019. doi: 10.18653/v1/W19-3506.
S. D. A. Putri, M. O. Ibrohim, and I. Budi, “Abusive language and hate speech detection for Javanese and Sundanese languages in tweets,” 2021. doi: 10.18178/wcse.2021.02.011.
J. Patihullah and E. Winarko, “Hate speech detection for Indonesia tweets using word embedding and gated recurrent unit,” 2019, doi: 10.22146/ijccs.40125.
T. T. A. Putri and S. Sriadhi, “A comparison of classification algorithms for hate speech detection,” 2020. doi: 10.1088/1757-899X/830/3/032006.
P. S. B. Ginting, B. Irawan, and C. Setianingsih, “Hate speech detection on Twitter using multinomial logistic regression classification method,” 2019. doi: 10.1109/IoTaIS47347.2019.8980379.
T. L. Sutejo and D. P. Lestari, “Indonesia hate speech detection using deep learning,” 2018. doi: 10.1109/IALP.2018.8629154.
D. A. N. Erlani and E. B. Setiawan, “Hate comment detection on Twitter using Long Short Term Memory (LSTM) with Genetic Algorithm (GA),” 2024, doi: 10.59188/eduvest.v4i11.1758.
Q. Sifak and E. B. Setiawan, “Hate Speech Detection using CNN and BiGRU with Attention Mechanism on Twitter,” in 2023 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT), IEEE, Nov. 2023, pp. 170–175. doi: 10.1109/COMNETSAT59769.2023.10420628.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” 2019. doi: 10.18653/v1/N19-1423.
F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A pretrained language model for Indonesian Twitter with effective domain-specific vocabulary initialization,” 2021. doi: 10.18653/v1/2021.emnlp-main.833.
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Stroudsburg, PA, USA: International Committee on Computational Linguistics, 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66.
A. Marpaung, R. Rismala, and H. Nurrahmi, “Hate speech detection in Indonesian Twitter texts using Bidirectional Gated Recurrent Unit,” 2021. doi: 10.1109/KST51265.2021.9415760.
J. F. Kusuma and A. Chowanda, “Indonesian hate speech detection using IndoBERTweet and BiLSTM on Twitter,” 2023, doi: 10.30630/joiv.7.3.1035.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Minarwati, Meiskey Rambu Dini Naomi (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.








