Combining Feature Embedding Based on the BERT Architecture and Random Forest Algorithm to Identify the Handling Section and Risk Priority Level in an ISP Company in Indonesia
DOI:
https://doi.org/10.21512/commit.v20i2.13022Keywords:
Sentiment Analysis, Bidirectional Encoder Representations from Transformers (BERT), Indonesian Bidirectional Representations from Transformers (BERT), Random Forest, Natural Language ProcessingAbstract
Revision Presently, Internet usage in society has grown considerably. It presents significant opportunities for companies to develop business related to internet service providers (ISPs). In Indonesia, both domestic and foreign ISPs are competing. To survive and remain popular with Indonesian customers, ISPs need to improve their products and services quickly. Hence, companies need to collect user feedback through customer loyalty applications and reviews to improve their service or technology. However, manually auditing thousands of textual reviews is labor-intensive and subject to human error, making it highly inefficient for agile, real-time corporate decision-making. The research aims to address this operational bottleneck by improving the automation process and providing decision-making information to ISP management. More specifically, the research contributes to implementing and comparing two transfer learning models: Bidirectional Encoder Representations from Transformer (BERT) and IndoBERT for tokenization and embedding. The embedding results are then fed into a random forest for classification. The dataset is collected via a scraping process from the Play Store application, focusing on user feedback and ratings. A total of 1,192 records of user feedback reviews from ISP loyalty apps on the Play Store in March 2024 were gathered. These are manually labeled by a general manager from the company into binary and multi-class categories. Splitting the dataset into 70% for training and 30% for testing achieves a classification accuracy of 70% for the problem domain with IndoBERT + Random Forest. However, multi-class classification for categorizing user feedback into a priority score model achieves an accuracy of 58%.
References
[1] W. Riani and S. Haryadi, “Effect of uneven distribution of broadband internet services in developing countries during economic recession,” Journal of Distribution Science, vol. 20, no. 10, pp. 1–9, 2022.
[2] B. H. Hayadi, H. Henderi, M. Budiarto, S. Sofiana,P. Padeli, D. Setiyadi, R. Swastika, and R. W. Arifin, “An extensive exploration into the multifaceted sentiments expressed by users of the myIM3 mobile application, unveiling complex emotional landscapes and insights,” Journal of Applied Data Sciences, vol. 5, no. 2, pp. 357–366, 2024.
[3] Y. Tirana and Sfenrianto, “Factors on mobile application user satisfaction in the largest Indonesian Internet Service Provider (ISP),” CommIT (Communication and Information Technology) Journal, vol. 17, no. 2, pp. 199–208, 2023.
[4] R. Yati, “Survey APJII: Pengguna Internet di Indonesia tembus 215 juta orang,” 2023. [Online]. Available: https://bit.ly/4gRLh8n
[5] S. Scherrer, S. Tabaeiaghdaei, and A. Perrig, “Quality competition among internet service providers,” Performance Evaluation, vol. 162, pp. 1–36, 2023.
[6] W. S. Ismail, M. M. Ghareeb, and H. Youssry, “Enhancing customer experience through sentiment analysis and natural language processing in e-commerce,” Journal of Wireless Mobile Networks, Ubiquitous Computing, and Dependable Applications (JoWUA), vol. 15, no. 3, pp. 60–72, 2024.
[7] R. I. Kurnia and A. S. Girsang, “Classification of user comment using Word2Vec and deep learning,” Advances in Science, Technology and Engineering Systems Journal, vol. 6, no. 2, pp. 566–572, 2021.
[8] I. J. Tae, A. Broillet-Schlesinger, and B. Y. Kim, “Effect of motivational factors on the use of integrated mobility applications: Behavioral intentions and customer loyalty,” Information, vol. 15, no. 9, pp. 1–19, 2024.
[9] T. H. Rohwiyati, A. I. Setiawan, L. Wahyudi, E. Dwi, and M. Amperawati, “E-trust and eservice quality on e-loyalty: Role of e-satisfaction and customer privacy,” Journal of Ecohumanism, vol. 3, no. 4, pp. 3130–3143, 2024.
[10] C. Y. Li, C. C. Huang, F. Lai, S. L. Lee, J. Wu, R. C. Chang, and H. W. Huang, “Mobile social service user identification framework based on action-characteristic data retention,” IEEE Access, vol. 8, pp. 127 748–127 767, 2020.
[11] M. Ariyanti, S. Widiyanesti, and W. H. Aprillia, “Service quality analysis of Telkomsel case study based on online customer reviews in Google Play Store,” in 2024 IEEE International Conference on Computing, Power and Communication Technologies (IC2PCT), vol. 5. Greater Noida, India: IEEE, Feb. 9–10, 2024, pp. 1669–1674.
[12] M. Dhotay, M. Dharrao, S. Deokate, A. Bongale, and D. Dharrao, “Multimodal sentiment analysis: Emerging innovations, core challenges, and future directions,” Discover Artificial Intelligence, vol. 6, pp. 1–28, 2026.
[13] E. Yulianti and N. K. Nissa, “ABSA of Indonesian customer reviews using IndoBERT: Single-sentence and sentence-pair classification approaches,” Bulletin of Electrical Engineering and Informatics, vol. 13, no. 5, pp. 3579–3589, 2024.
[14] N. Paltrinieri, L. Comfort, and G. Reniers, “Learning about risk: Machine learning for risk assessment,” Safety Science, vol. 118, pp. 475–486, 2019.
[15] W. Liao, Z. Liu, H. Dai, Z. Wu, Y. Zhang, X. Huang et al., “Mask-guided BERT for fewshot text classification,” Neurocomputing, vol. 610, 2024.
[16] D. Dharrao, M. R. Aadithyanarayanan, R. Mital, A. Vengali, M. Pangavhane, S. Rajput, and A. M. Bongale, “An efficient method for disaster tweets classification using gradient-based optimized convolutional neural networks with BERT embeddings,” MethodsX, vol. 13, pp. 1–10, 2024.
[17] Y. Xiong, G. Chen, and J. Cao, “Research on public service request text classification based on BERT-BiLSTM-CNN feature fusion,” Applied Sciences, vol. 14, no. 14, pp. 1–11, 2024.
[18] J. I. T. Krisna, A. Luthfiarta, L. D. Cahya, S. Winarno, and A. Nugraha, “Comparing optimizer strategies for enhancing emotion classification in IndoBERT models,” Advance Sustainable Science, Engineering and Technology (ASSET), vol. 6, no. 2, pp. 1–8, 2024.
[19] G. Z. Nabiilah, I. N. Alam, E. S. Purwanto, and M. F. Hidayat, “Indonesian multilabel classification using IndoBERT embedding and MBERT classification,” International Journal of Electrical & Computer Engineering (IJECE), vol. 14, no. 1, pp. 1071–1078, 2024.
[20] H. Imaduddin, F. Y. A’la, and Y. S. Nugroho, “Sentiment analysis in Indonesian healthcare applications using IndoBERT approach,” International Journal of Advanced Computer Science and Applications, vol. 14, no. 8, pp. 113–117, 2023.
[21] O. Galal, A. H. Abdel-Gawad, and M. Farouk, “Rethinking of BERT sentence embedding for text classification,” Neural Computing and Applications, vol. 36, no. 32, pp. 20 245–20 258, 2024.
[22] N. Darraz, I. Karabila, A. El-Ansari, N. Alami, and M. El Mallahi, “Integrated sentiment analysis with BERT for enhanced hybrid recommendation systems,” Expert Systems with Applications, vol. 261, 2025.
[23] V. A. Fitri, R. Andreswari, and M. A. Hasibuan, “Sentiment analysis of social media Twitter with case of anti-LGBT campaign in Indonesia using na¨ıve bayes, decision tree, and random forest algorithm,” Procedia Computer Science, vol. 161, pp. 765–772, 2019.
[24] Handrizal, T. H. F. Harumy, Herriyance, and M. A. Ilmi, “Sentiment analysis based on 7P marketing mix aspects of the Indriver application service using the BERT algorithm, based on user reviews on the Google Play Store,” Journal of Theoretical and Applied Information Technology, vol. 101, no. 19, pp. 6136–6144, 2023.
[25] K. Aziz, D. Ji, P. Chakrabarti, T. Chakrabarti, M. S. Iqbal, and R. Abbasi, “Unifying aspectbased sentiment analysis bert and multi-layered graph convolutional networks for comprehensive sentiment dissection,” Scientific Reports, vol. 14, pp. 1–22, 2024.
[26] D. Y. Yefferson, V. Lawijaya, and A. S. Girsang, “Hybrid model: IndoBERT and long short-term memory for detecting Indonesian hoax news,” IAES International Journal of Artificial Intelligence (IJ-AI), vol. 13, no. 2, pp. 1913–1924, 2024.
[27] T. Wang, “Improved random forest classification model combined with C5. 0 algorithm for vegetation feature analysis in non-agricultural environments,” Scientific Reports, vol. 14, pp. 1–13, 2024.
[28] S. Redjeki, S. Abadi, D. Kurniati, S. R. C. Nursari, A. Damayanti, and E. Iskandar, “Clustering on sentiment analysis: Effect of Twitter dataset,” Journal of advanced research, vol. 51, no. 1, pp. 39–51, 2024.
[29] Q. Xi and P. Jiang, “Design of news sentiment classification and recommendation system based on multi-model fusion and text similarity,” International Journal of Cognitive Computing in Engineering, vol. 6, pp. 44–54, 2025.
[30] I. Kanwal, F. Wahid, S. Ali, A.-U. Rehman, A. Alkhayyat, and A. Al-Radaei, “Sentiment analysis using hybrid model of stacked autoencoder-based feature extraction and long short term memory-based classification approach,” IEEE Access, vol. 11, pp. 124 181–124 197, 2023.
[31] K. Barik and S. Misra, “Analysis of customer reviews with an improved VADER lexicon classifier,” Journal of Big Data, vol. 11, pp. 1–29, 2024.
[32] L. A. Gaafar, A. Z. Ghalwash, A. A. Youssif, and H. A. Ghalwash, “Machine and deep learning models for multiclass sentiment classification,” Journal of Theoretical and Applied Information Technology, vol. 102, no. 22, pp. 8325–8339, 2024.
[33] A. Maiti, A. Abarda, and M. Hanini, “The impact of feature extraction techniques on the performance of text data classification models,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 35, no. 2, pp. 1041–1052, 2024.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Rhemzy Putra Maulana, Yulianto

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
a. Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License - Share Alike that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
b. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
c. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
Â
USER RIGHTS
All articles published Open Access will be immediately and permanently free for everyone to read and download. We are continuously working with our author communities to select the best choice of license options, currently being defined for this journal as follows: Creative Commons Attribution-Share Alike (CC BY-SA)

















