Hate Speech Detection Leveraging Emoji Descriptions and BERT-Based Contextual Data Augmentation

Authors

  • Junita Amalia Institut Teknologi Del
  • Walker Valentinus Simanjuntak Institut Teknologi Del
  • Ruth Marelisa Hutagalung Institut Teknologi Del
  • Lamria Magdalena Tampubolon Institut Teknologi Del

DOI:

https://doi.org/10.21512/comtech.v17i2.14618

Keywords:

hate speech detection, emoji description, mBERT, data augmentation, Bi-LSTM

Abstract

This study evaluates the efficacy of emoji-to-text transformation and contextual data augmentation in improving hate speech detection on social media. Using a unified dataset from OLID, TweetEval, and the Sushil benchmark, we compared Multinomial Naive Bayes (MNB), Bidirectional LSTM (Bi-LSTM), and Multilingual BERT (mBERT). To address class imbalance and semantic loss, we implemented a preprocessing pipeline that converts emojis into textual descriptions and utilized a BERT-based contextual augmenter. Results indicate that textualizing emojis consistently improves performance by providing critical semantic anchors. The mBERT model achieved the highest performance with a Macro F1-score of 0.87 and a hate speech recall of 0.94, outperforming the Bi-LSTM (0.835) and MNB (0.83) baselines. Notably, the Bi-LSTM showed the highest sensitivity to the "Emoji Description" method, with a 5% performance increase compared to the "No Emoji" condition. These findings demonstrate that bridging the gap between visual symbols and textual intent significantly reduces model bias and improves the detection of nuanced toxicity. This research provides an optimized framework for context-aware moderation in imbalanced digital environments.

Dimensions

References

Alaoui, S. S., Farhaoui, Y., & Aksasse, B. (2022). Hate speech detection using text mining and machine learning. International Journal of Decision Support System Technology (IJDSST), 14(1), 1-20. https://doi.org/10.4018/IJDSST.286680

Amalia, F. S., & Suyanto, Y. (2024). Offensive language and hate speech detection using BERT model. IJCCS (Indonesian Journal of Computing and Cybernetics Systems), 18(4). https://doi.org/10.22146/ijccs.99841

Amalia, J., Tambunan, S. R., Purba, S. E. M., & Simanjuntak, W. V. (2025). Enhancing hate speech detection: leveraging emoji preprocessing with BI-LSTM model. Journal of Information Systems and Informatics, 7(2). https://doi.org/10.51519/journalisi.v7i2.1147

Aminu, E. F., Ayobami Ekundayo, Sarkibaka, S. D., Ojerinde, O. A., & Ugwuoke, U. C. (2024). Hate speech detector based on hybridized BERT-attention mechanism and context analyzer. Journal of Computer Sciences and Informatics, 1(1). 10.5455/JCSI.20240613103822

Barbieri, F., Camacho-Collados, J., Neves, L., & Espinosa-Anke, L. (2020). TweetEval: Unified benchmark and comparative evaluation for tweet classification. Computer Science, Computation and Language. https://doi.org/10.48550/arXiv.2010.12421

Chai, Y., & Xie, H. (2025). Text data augmentation for large language models: A comprehensive survey of methods, challenges, and opportunities. Artificial Intelligence Review, 59. https://doi.org/10.1007/s10462-025-11405-5

Christen, P., Hand, D. J., & Kirielle, N. (2023). A review of the f-measure: Its history, properties, criticism, and alternatives. ACM Computing Surveys, 56(3), 1 – 24. https://doi.org/10.1145/3606367

Dalavi, S. R., Nivelkar, T., Patil, S., & Sawant, A. (2023). Enhancing hate speech detection through emoji-based classification using bi-lstm and glove embeddings. 2023 6th International Conference on Advances in Science and Technology (ICAST). https://doi.org/10.1109/ICAST59062.2023.10455077

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1 (Long and Short Papers). https://doi.org/10.18653/v1/N19-1423

Dixon, S. J. (2025, November 10). Number of social media users worldwide 2030. Statista. Retrieved February 2, 2026, from https://www.statista.com/forecasts/278414/number-of-worldwide-social-network-users/

Gelber, K. (2021). Differentiating hate speech: A systemic discrimination approach. Critical Review of International Social and Political Philosophy, 24(4). https://doi.org/10.1080/13698230.2019.1576006

Handayani, T. P., & Hasyim, W. (2024). Comparative analysis of CNN-RNN models for hatespeech detection incorporating L2 regularization. International Journal of Engineering, Science, and Information Technology, 4(1).

Lazareva, S. (2023). Emojis as graphical hate speech markers. Journal of Language Works, 8(1). https://tidsskrift.dk/lwo/article/view/138019

Ma, E. (2019). nlpaug. Github. https://github.com/makcedward/nlpaug

Putra, C. D., & Wang, H.-C. (2023). Advanced BERT-CNN for hate speech detection. Procedia Computer Science, 234. https://doi.org/10.1016/j.procs.2024.02.170

Rawat, A., Kumar, S., & Samant, S. S. (2024). Hate speech detection in social media: Techniques, recent trends, and future challenges. Wires Computational Statistics, 16(2). https://doi.org/10.1002/wics.1648

Safawi, N. U. C. M., & Shafie, N. A. (2024). Performance of TF-IDF for text classification reviews on google play store: Shopee. Journal of Computing Research and Innovation, 9(2), 13-22. https://doi.org/10.24191/jcrinn.v9i2.410

Saikh, T., Barman, S., Kumar, H., Sahu, S., & Palit, S. (2024). Emojis trash or treasure: Utilizing emoji to aid hate speech detection. Proceedings of the 21st International Conference on Natural Language Processing (ICON). https://aclanthology.org/2024.icon-1.64/

Salsabila, N. S., Rahma, A., Agustin, A., & Nanda, D. (2025). The impact of emoji use on perception differences in online conversations among teenagers. LITERA Jurnal Bahasa Dan Sastra.

Telaumbanua, Y. A., Tealumbanua, N. T. N., Halawa, M. D., & Benedikta Gulo. (2024). The use of emojis in language communication on social media platforms. Journal of English Language and Education, 9(4). https://doi.org/10.31004/jele.v9i4.524

Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S., Farra, N., & Kumar, R. (2019). Predicting the type and target of offensive posts in social media. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 10.18653/v1/N19-1144

Downloads

Published

2026-09-28

How to Cite

Amalia, J., Simanjuntak, W. V., Hutagalung, R. M., & Tampubolon, L. M. (2026). Hate Speech Detection Leveraging Emoji Descriptions and BERT-Based Contextual Data Augmentation. ComTech: Computer, Mathematics and Engineering Applications, 17(2), 117–125. https://doi.org/10.21512/comtech.v17i2.14618
Abstract 54  .
PDF downloaded 50  .