DETECTION OF OFFENSIVE CONTENT USING THE BERT MODEL AND MECHANISMS OF MONITORING

Authors

DOI:

https://doi.org/10.56132/2791-3368-2026-1-65-71-82

Keywords:

Berta model, obscene words, attention mechanism, social network, online content, corpus of the Kazakh language, Files Kappa

Abstract

This research paper considers methods for automatic detection of profanity and aggressive comments in Kazakh social networks. The relevance of the study is due to the increasing volume of profanity comments online and the lack of effective automated solutions for resource-limited languages. The Kazakh language is characterized by complex morphology and contextual specificity, which makes it difficult to use traditional natural language processing methods. The proposed study introduces a hybrid model based on the BERT transformer architecture, combining multihead attention and convolutional neural networks. This architecture enables the effective detection of offensive texts, including both explicit and subtle forms of comments exchanged among cadets and officers online. To train and test the model, a corpus of Kazakh-language comments annotated by experts was constructed, including opinions from cadets, commanders, and personnel. The quality of the annotation was assessed using Fleiss's Kappa coefficient, which gave 0.6649, with an overall agreement level of 91.67%, which indicates high reliability of the data. Experimental results showed that the proposed model outperforms the baseline methods (Naive Bayes, LSTM, FastText) in terms of accuracy, F1-score, and ROC-AUC. The resulting ROC-AUC value was 0.95. The research results confirm the effectiveness of transformer models with an attention mechanism for detecting profanity in the Kazakh language and can be used in the development of automated content moderation and information security monitoring systems.

Downloads

Download data is not yet available.

Author Biographies

  • Rustam Abdrakhmanov, International University оf Tourism And Hospitality

    сandidate of technical sciences, associate professor, Turkestan, Kazakhstan

  • Timur Kartbayev, Kazakh National Women's Teacher Training University

    PhD, professor, Digital Officer, Almaty, Kazakhstan 

  • Aigerim Toktarova, Auezov University

    PhD, senior lecturer, Shymkent, Kazakhstan

References

1. C. Van Hee, G. Jacobs, C. Emmery, B. Desmet, E. Lefever and B.Verhoeven, «Automatic detection of cyberbullying in social media text.» PloS one 13.10 (2018): e0203794.

2. G. Valle-Cano, «SocialHaterBERT: A dichotomous approach for automatically detecting hate speech on Twitter through textual analysis and user profiles» Expert Systems with Applications 216 (2023): 119446.

3. S. Mombekova, S. Akhmetova, G. Shaimerdenova, S. Alisheva, E. Mussirepova, A. Kydyrbekova and N. Torebay, «Eco-efficient industrial processes: Leveraging ai-powered management for reduced environmental footprint.» E3S Web of Conferences. Vol. 614. EDP Sciences, 2025.

4. K. Kumari, J.P. Singh, Y.K. Dwivedi and N.P. Rana, «Towards Cyberbullying-free social media in smart cities: a unified multi-modal approach» Soft computing 24 (2020): 11059-11070.

5. H. Elzayady, M.S. Mohamed, K.M. Badran, and G.I. Salama, «A hybrid approach based on personality traits for hate speech detection in Arabic social media» International Journal of Electrical and Computer Engineering 13.2 (2023): 1979.

6. Z. Makhanova, G. Beissenova, A. Madiyarova, M. Chazhabayeva, G. Mambetaliyev, M. Suimenova and A. Baiburin, «A Deep Residual Network Designed for Detecting Cracks in Buildings of Historical Significance.» International Journal of Advanced Computer Science & Applications 15.5 (2024).

7. J.L.Wang, L.A. Jackson, J. Gaskin and H.Z. Wang»The effects of Social Networking Site (SNS) use on college students’ friendship and well-being» Computers in Human Behavior 37 (2014): 229-236.

8. A. Toktarova, A. Abushakhma, E. Adylbekova, A. Manapova, B. Kaldarova and Y.Atayev, «Offensive language identification in low resource languages using bidirectional long-short-term memory network.» International Journal of Advanced Computer Science and Applications 14.6 (2023).

9. M. Rezvan, S. Shekarpour, L. Balasuriya and K. Thirunarayan, «A quality type-aware annotated corpus and lexicon for harassment research» Proceedings of the 10th acm conference on web science. 2018.

10. B.A. Talpur and O. Declan, «Cyberbullying severity detection: A machine learning approach» PloS one 15.10 (2020): e0240924.

11. A. Toktarova, D. Sultan, and Zh. Azhibekova, «Review of Machine Learning Models in Cyberbullying Detection Problem» 2024 IEEE 4th International Conference on Smart Information Systems and Technologies (SIST). IEEE, 2024.

12. C. Comito, F. Agostino, and C. Pizzuti, «Word embedding based clustering to detect topics in social media» IEEE/WIC/ACM International Conference on Web Intelligence. 2019.

Downloads

Published

2026-03-12

How to Cite

DETECTION OF OFFENSIVE CONTENT USING THE BERT MODEL AND MECHANISMS OF MONITORING. (2026). Bulletin of the Military Institute Named After S. Nurmagambetov, 1(65), 71-82. https://doi.org/10.56132/2791-3368-2026-1-65-71-82

Similar Articles

1-10 of 41

You may also start an advanced similarity search for this article.