Hate Speech Detection Leveraging Emoji Descriptions and BERT-Based Contextual Data Augmentation
DOI:
https://doi.org/10.21512/comtech.v17i2.14618Keywords:
hate speech detection, emoji description, mBERT, data augmentation, Bi-LSTMAbstract
This study evaluates the efficacy of emoji-to-text transformation and contextual data augmentation in improving hate speech detection on social media. Using a unified dataset from OLID, TweetEval, and the Sushil benchmark, we compared Multinomial Naive Bayes (MNB), Bidirectional LSTM (Bi-LSTM), and Multilingual BERT (mBERT). To address class imbalance and semantic loss, we implemented a preprocessing pipeline that converts emojis into textual descriptions and utilized a BERT-based contextual augmenter. Results indicate that textualizing emojis consistently improves performance by providing critical semantic anchors. The mBERT model achieved the highest performance with a Macro F1-score of 0.87 and a hate speech recall of 0.94, outperforming the Bi-LSTM (0.835) and MNB (0.83) baselines. Notably, the Bi-LSTM showed the highest sensitivity to the "Emoji Description" method, with a 5% performance increase compared to the "No Emoji" condition. These findings demonstrate that bridging the gap between visual symbols and textual intent significantly reduces model bias and improves the detection of nuanced toxicity. This research provides an optimized framework for context-aware moderation in imbalanced digital environments.
References
Alaoui, S. S., Farhaoui, Y., & Aksasse, B. (2022). Hate speech detection using text mining and machine learning. International Journal of Decision Support System Technology (IJDSST), 14(1), 1-20. https://doi.org/10.4018/IJDSST.286680
Amalia, F. S., & Suyanto, Y. (2024). Offensive language and hate speech detection using BERT model. IJCCS (Indonesian Journal of Computing and Cybernetics Systems), 18(4). https://doi.org/10.22146/ijccs.99841
Amalia, J., Tambunan, S. R., Purba, S. E. M., & Simanjuntak, W. V. (2025). Enhancing hate speech detection: leveraging emoji preprocessing with BI-LSTM model. Journal of Information Systems and Informatics, 7(2). https://doi.org/10.51519/journalisi.v7i2.1147
Aminu, E. F., Ayobami Ekundayo, Sarkibaka, S. D., Ojerinde, O. A., & Ugwuoke, U. C. (2024). Hate speech detector based on hybridized BERT-attention mechanism and context analyzer. Journal of Computer Sciences and Informatics, 1(1). 10.5455/JCSI.20240613103822
Barbieri, F., Camacho-Collados, J., Neves, L., & Espinosa-Anke, L. (2020). TweetEval: Unified benchmark and comparative evaluation for tweet classification. Computer Science, Computation and Language. https://doi.org/10.48550/arXiv.2010.12421
Chai, Y., & Xie, H. (2025). Text data augmentation for large language models: A comprehensive survey of methods, challenges, and opportunities. Artificial Intelligence Review, 59. https://doi.org/10.1007/s10462-025-11405-5
Christen, P., Hand, D. J., & Kirielle, N. (2023). A review of the f-measure: Its history, properties, criticism, and alternatives. ACM Computing Surveys, 56(3), 1 – 24. https://doi.org/10.1145/3606367
Dalavi, S. R., Nivelkar, T., Patil, S., & Sawant, A. (2023). Enhancing hate speech detection through emoji-based classification using bi-lstm and glove embeddings. 2023 6th International Conference on Advances in Science and Technology (ICAST). https://doi.org/10.1109/ICAST59062.2023.10455077
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1 (Long and Short Papers). https://doi.org/10.18653/v1/N19-1423
Dixon, S. J. (2025, November 10). Number of social media users worldwide 2030. Statista. Retrieved February 2, 2026, from https://www.statista.com/forecasts/278414/number-of-worldwide-social-network-users/
Gelber, K. (2021). Differentiating hate speech: A systemic discrimination approach. Critical Review of International Social and Political Philosophy, 24(4). https://doi.org/10.1080/13698230.2019.1576006
Handayani, T. P., & Hasyim, W. (2024). Comparative analysis of CNN-RNN models for hatespeech detection incorporating L2 regularization. International Journal of Engineering, Science, and Information Technology, 4(1).
Lazareva, S. (2023). Emojis as graphical hate speech markers. Journal of Language Works, 8(1). https://tidsskrift.dk/lwo/article/view/138019
Ma, E. (2019). nlpaug. Github. https://github.com/makcedward/nlpaug
Putra, C. D., & Wang, H.-C. (2023). Advanced BERT-CNN for hate speech detection. Procedia Computer Science, 234. https://doi.org/10.1016/j.procs.2024.02.170
Rawat, A., Kumar, S., & Samant, S. S. (2024). Hate speech detection in social media: Techniques, recent trends, and future challenges. Wires Computational Statistics, 16(2). https://doi.org/10.1002/wics.1648
Safawi, N. U. C. M., & Shafie, N. A. (2024). Performance of TF-IDF for text classification reviews on google play store: Shopee. Journal of Computing Research and Innovation, 9(2), 13-22. https://doi.org/10.24191/jcrinn.v9i2.410
Saikh, T., Barman, S., Kumar, H., Sahu, S., & Palit, S. (2024). Emojis trash or treasure: Utilizing emoji to aid hate speech detection. Proceedings of the 21st International Conference on Natural Language Processing (ICON). https://aclanthology.org/2024.icon-1.64/
Salsabila, N. S., Rahma, A., Agustin, A., & Nanda, D. (2025). The impact of emoji use on perception differences in online conversations among teenagers. LITERA Jurnal Bahasa Dan Sastra.
Telaumbanua, Y. A., Tealumbanua, N. T. N., Halawa, M. D., & Benedikta Gulo. (2024). The use of emojis in language communication on social media platforms. Journal of English Language and Education, 9(4). https://doi.org/10.31004/jele.v9i4.524
Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S., Farra, N., & Kumar, R. (2019). Predicting the type and target of offensive posts in social media. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 10.18653/v1/N19-1144
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Junita Amalia, Walker Valentinus Simanjuntak, Ruth Marelisa Hutagalung, Lamria Magdalena Tampubolon

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
a. Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License - Share Alike that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
b. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
c. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
 USER RIGHTS
 All articles published Open Access will be immediately and permanently free for everyone to read and download. We are continuously working with our author communities to select the best choice of license options, currently being defined for this journal as follows:

















