Enhancing Hate Speech and Offensive Language Detection using CatBoost with RoBERTa-based Contextual Embeddings
Abstract
Keywords
Full Text:
PDFReferences
F. Alkomah and X. Ma, “A Literature Review of Textual Hate Speech Detection Methods and Datasets,” Jun. 01, 2022, MDPI. DOI: 10.3390/info13060273.
Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” Jul. 2019, [Online]. Available: http://arxiv.org/abs/1907.11692
T. Mandl et al., “Overview of the HASOC Subtrack at FIRE 2021: Hate Speech and Offensive Content Identification in English and Indo-Aryan Languages under Creative Commons License Attribution 4.0 International (CC BY 4.0),” 2021, [Online]. Available: http://ceur-ws.org
S. Agustian, R. Saputra, and A. Fadhilah, “‘Feature Selection’ with Pretrained-BERT for Hate Speech and Offensive Content Identification in English and Hindi Languages,” 2021. [Online]. Available: https://huggingface.co/surajp/RoBERTa-hindi-guj-san
T. Wolf et al., “Transformers: State-of-the-Art Natural Language Processing,” 2020. [Online]. Available: https://github.com/huggingface/
D. Quoc Nguyen, T. Vu, A. Tuan Nguyen, and V. Research, “BERTweet: A Pre-Trained Language Model for English Tweets,” 2020. [Online]. Available: https://pypi.org/project/emoji
E. Essa, K. Omar, and A. Alqahtani, “Fake News Detection based on a Hybrid BERT and LightGBM Models,” Complex and Intelligent Systems, Vol. 9, No. 6, pp. 6581–6592, Dec. 2023, DOI: 10.1007/s40747-023-01098-0.
Z. Guo, X. Wang, and L. Ge, “Classification Prediction Model of Indoor PM2.5 Concentration using CatBoost Algorithm,” Front. Built Environ., Vol. 9, 2023, DOI: 10.3389/fbuil.2023.1207193.
R. A. Muhammad et al., “Sistemasi: Jurnal Sistem Informasi Penerapan Metode Naïve Bayes dengan PSO pada Analisis Sentimen Pengguna Aplikasi X terhadap Bitcoin Sentiment Analysis of X Application Users on Bitcoin using the Naïve Bayes Method Optimized with Particle Swarm Optimization (PSO),” 2025. [Online]. Available: http://sistemasi.ftik.unisi.ac.id
T. Mandl, S. Modha, M. Anand Kumar, and B. R. Chakravarthi, “Overview of the HASOC Track at FIRE 2020: Hate Speech and Offensive Language Identification in Tamil, Malayalam, Hindi, English and German,” in ACM International Conference Proceeding Series, Association for Computing Machinery, Dec. 2020, pp. 29–32. DOI: 10.1145/3441501.3441517.
J. M. Ahn, J. Kim, and K. Kim, “Ensemble Machine Learning of Gradient Boosting (XGBoost, LightGBM, CatBoost) and Attention-based CNN-LSTM for Harmful Algal Blooms Forecasting,” Toxins (Basel)., Vol. 15, No. 10, Oct. 2023, DOI: 10.3390/toxins15100608.
A. Glazkova, M. Kadantsev, and M. Glazkov, “Fine-Tuning of Pre-Trained Transformers for Hate, Offensive, and Profane Content Detection in English and Marathi,” 2021. [Online]. Available: https://pypi.org/project/tweet-preprocessor
N. A. R. Putri and Ardiansyah, “Analisis Sentimen terhadap Kemajuan Kecerdasan Buatan di Indonesia menggunakan BERT dan RoBERTa,” Jurnal Sains dan Informatika, Vol. 9, No. 2, pp. 136–145, Nov. 2023, DOI: 10.34128/jsi.v9i2.649.
A. Kumar, P. K. Roy, and S. Saumya, “An Ensemble Approach for Hate and Offensive Language Identification in English and Indo-Aryan Languages,” 2021. [Online]. Available: https://support.google.com/youtube/answer/2801939.
G. B. Herwanto, A. M. Ningtyas, I. G. Mujiyatna, K. E. Nugraha, and I. N. Prayana Trisna, “Hate Speech Detection in Indonesian Twitter using Contextual Embedding Approach,” IJCCS (Indonesian Journal of Computing and Cybernetics Systems), Vol. 15, No. 2, p. 177, Apr. 2021, DOI: 10.22146/ijccs.64916.
S. Banerjee, M. Sarkar, N. Agrawal, P. Saha, and M. Das, “Exploring Transformer based Models to Identify Hate Speech and Offensive Content in English and Indo-Aryan Languages,” 2021. [Online]. Available: https://www.reuters.com/investigates/special-report/myanmar-facebook-hate
Z. M. Farooqi, S. Ghosh, and R. R. Shah, “Leveraging Transformers for Hate Speech Detection in Conversational Code-Mixed Tweets,” Dec. 2021, [Online]. Available: http://arxiv.org/abs/2112.09986
W. Yu and D. Kolossa, “Hybrid Representation Fusion for Twitter Hate Speech Identification,” 2021. [Online]. Available: https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html
N. Bölücü and P. Canbay, “Hate Speech and Offensive Content Identification with Graph Convolutional Networks,” 2021. [Online]. Available: http://ceur-ws.org
DOI: https://doi.org/10.32520/stmsi.v15i7.6637
Article Metrics
Abstract view : 13 timesPDF - 0 times
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.







