Comparative Evaluation of Modern Embeddings and Ensemble Learning for Sentiment Analysis of Mobile Banking Application Reviews

I Wayan Adi Wiratama, Putu Hendra Suputra, I Made Gede Sunarya

Abstract


User reviews of mobile banking applications provide an important source of information for evaluating the quality of digital services. However, conventional text representations such as Term Frequency–Inverse Document Frequency (TF-IDF) have limitations in capturing semantic context, while class imbalance can affect sentiment classification performance. This study aims to comparatively evaluate text representations and classification algorithms for sentiment analysis of mobile banking application reviews. The dataset consists of 9,885 user reviews of the Livin’ by Mandiri application collected from the Google Play Store. The study compares three text representation methods, namely TF-IDF, Multilingual E5 Large Instruct, and Nomic Embed v2, with four classification algorithms: Naive Bayes, Support Vector Machine (SVM), XGBoost, and CatBoost, resulting in 12 model combinations. Borderline-SMOTE was applied to the training data to address class imbalance, while hyperparameter optimization was performed using BayesSearchCV. Model performance was evaluated using accuracy, precision, recall, F1-score, and macro-F1 through stratified 5-fold cross-validation on the training data and evaluation on the test set. The results show that Multilingual E5 Large Instruct achieved the highest average macro-F1 score of 0.950. The combination of CatBoost and Multilingual E5 Large Instruct achieved the best individual performance, with an accuracy of 98% and a macro-F1 score of 0.97. These findings demonstrate that the choice of text representation and classification algorithm plays an important role in determining the performance of sentiment analysis for Indonesian-language mobile banking application reviews.

Keywords


ensemble learning; mobile banking; modern embeddings; oversampling; Sentiment analysis

Full Text:

PDF

References


B. Setiawan, “A Review of Sentiment Analysis Applications in Indonesia Between 2023-2024,” 2024. DOI: 10.26740/jieet.v8n2.p71-83.

L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei, “Multilingual E5 Text Embeddings: A Technical Report,” Feb. 2024, [Online]. Available: http://arxiv.org/abs/2402.05672

N. Muennighoff, N. Tazi, L. Magne, and N. Reimers, “MTEB: Massive Text Embedding Benchmark,” Mar. 2023, [Online]. Available: http://arxiv.org/abs/2210.07316

H. Han, W.-Y. Wang, and B.-H. Mao, “Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning,” 2005. DOI: 10.1007/11538059_91.

R. Wahyudi et al., “Analisis Sentimen pada Review Aplikasi Grab di Google Play Store Menggunakan Support Vector Machine,” Jurnal Informatika, Vol. 8, No. 2, 2021, [Online]. Available: http://ejournal.bsi.ac.id/ejurnal/index.php/ji

C. C. Yolanda, S. Syafriandi, Y. Kurniawati, and D. Fitria, “Sentiment Analysis of DANA Application Reviews on Google Play Store Using Naïve Bayes Classifier Algorithm based on Information Gain,” UNP Journal of Statistics and Data Science, Vol. 2, No. 1, pp. 48–55, Feb. 2024, DOI: 10.24036/ujsds/vol2-iss1/147.

W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, “A Survey on Imbalanced Learning: Latest Research, Applications and Future Directions,” Artif. Intell. Rev., Vol. 57, No. 6, Jun. 2024, DOI: 10.1007/s10462-024-10759-6.

M. C. Hinojosa Lee, J. Braet, and J. Springael, “Performance Metrics for Multilabel Emotion Classification: Comparing Micro, Macro, and Weighted F1-Scores,” Applied Sciences (Switzerland), Vol. 14, No. 21, Nov. 2024, DOI: 10.3390/app14219863.

A. Suharman and M. Kamayani Sulaeman, “Analisis Sentimen Pengguna Aplikasi Livin’ by Mandiri menggunakan Metode Support Vector Machine (SVM) dengan Ekstraksi Fitur TF-IDF dan Word2Vec,” Jurnal Pendidikan dan Teknologi Indonesia, Vol. 5, No. 8, pp. 2201–2212, Aug. 2025, DOI: 10.52436/1.jpti.941.

Q. H. Nguyen et al., “Influence of Data Splitting on Performance of Machine Learning Models in Prediction of Shear Strength of Soil,” Math. Probl. Eng., Vol. 2021, 2021, DOI: 10.1155/2021/4832864.

I. Nyoman Saputra Wahyu Wijaya, K. Agus Seputra, and N. Putu Novita Puspa Dewi, “Fine Tunning Model Indobert untuk Analisis Sentimen Berita Pariwisata Indonesia,” Jurnal Pendidikan Teknologi dan Kejuruan, Vol. 22, No. 2, 2025, DOI: 10.23887/jptk-undiksha.v22i2.104056.

Z. Nussbaum and B. Duderstadt, “Training Sparse Mixture Of Experts Text Embedding Models,” Mar. 2025, [Online]. Available: http://arxiv.org/abs/2502.07972

I. B. N. W. Manuaba, Gede Rasben Dantes, and Gede Indrawan, “Analisis Sentimen Data Provider Layanan Internet pada Twitter menggunakan Support Vector Machine (SVM) dengan Penambahan Algoritma Levenshtein Distance,” 2022.

A. Sasmita, G. A. Pradnyana, and D. G. H. Divayana, “Pengembangan Sistem Analisis Sentimen untuk Evaluasi Kinerja Dosen Universitas Pendidikan Ganesha dengan Metode Naïve Bayes,” JST (Jurnal Sains dan Teknologi), Vol. 11, No. 2, pp. 451–462, Sep. 2022, DOI: 10.23887/jstundiksha.v11i2.44384.

N. K. T. A. Saputri, I. G. A. Gunadi, and I. M. G. Sunarya, “Analisis Sentimen Pelayanan Daring di Fakultas Teknik dan Kejuruan Universitas Pendidikan Ganesha menggunakan Algoritma Naïve Bayes dan LSTM,” MALCOM: Indonesian Journal of Machine Learning and Computer Science, Vol. 4, No. 3, pp. 1120–1129, Jul. 2024, DOI: 10.57152/malcom.v4i3.1336.

P. W. Ariyani, I. Made, G. Sunarya, I. Gede, and A. Gunadi, “Analisis Sentimen Masyarakat terhadap Virus Corona berdasarkan Opini dari Twitter menggunakan Metode Naive Bayes dan K-Nearest Neighbor,” Jurnal Pendidikan Teknologi dan Kejuruan, Vol. 22, No. 2, 2025.

N. H. Cahyana, Y. Fauziah, W. Wisnalmawati, A. S. Aribowo, and S. Saifullah, “The Evaluation of Effects of Oversampling and Word Embedding on Sentiment Analysis,” JURNAL INFOTEL, Vol. 17, No. 1, pp. 54–67, Apr. 2025, DOI: 10.20895/infotel.v17i1.1077.

J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian Optimization of Machine Learning Algorithms,” 2012, pp. 2951–2959.

E. Lopez, J. Etxebarria-Elezgarai, J. M. Amigo, and A. Seifert, “The Importance of Choosing a Proper Validation Strategy in Predictive Models. A Tutorial with Real Examples,” Sep. 22, 2023, Elsevier B.V. DOI: 10.1016/j.aca.2023.341532.

T. Fearn, “Testing Differences in Predictive Ability: A Tutorial,” J. Chemom., Vol. 38, No. 8, Aug. 2024, DOI: 10.1002/cem.3549.

L. A. Yates, Z. Aandahl, S. A. Richards, and B. W. Brook, “Cross Validation for Model Selection: A Review with Examples from Ecology,” Ecol. Monogr., Vol. 93, No. 1, Feb. 2023, DOI: 10.1002/ecm.1557.




DOI: https://doi.org/10.32520/stmsi.v15i9.6912

Article Metrics

Abstract view : 0 times
PDF - 0 times

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.