Development of an Artificial Intelligence Model for Detecting Toxic (Harmful) Content in Texts

Received: 2026-05-30

Published: 2026-06-06

Abstract

This paper addresses the problem of automatic detection of toxic and harmful content in texts. The proliferation of toxic content on social networks and online platforms has a negative impact not only on users but also on society as a whole. The study proposes a hybrid artificial intelligence architecture that combines BERT-based transformers with bidirectional LSTM networks. The proposed model was trained on mixed Uzbek and Russian texts collected from Uzbek social media platforms. Experiments showed that the model achieved 94.7% accuracy, 93.2% recall, and 94.0% F1-score. In addition, the paper analyzes dataset preparation, model evaluation metrics, and real-time deployment capabilities.

List of references

  1. Nobata, C., Tetreault, J., Thomas, A., Mehdad, Y., & Chang, Y. (2016). Abusive Language Detection in Online User Content. Proceedings of WWW 2016, 145–153.

  2. Davidson, T., Warmsley, D., Macy, M., & Weber, I. (2017). Automated Hate Speech Detection and the Problem of Offensive Language. Proceedings of ICWSM 2017, 512–515.

  3. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of NAACL-HLT 2019, 4171–4186.

  4. Zhang, Y., Wallace, B. (2018). A Sensitivity Analysis of (and Practitioners' Guide to) Convolutional Neural Networks for Sentence Classification. Proceedings of IJCNLP, 253–263.

  5. Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S. (2019). Predicting the Type and Target of Offensive Posts in Social Media. Proceedings of NAACL-HLT 2019, 1415–1420.

  6. Kovačević, A., Dragović, I., Kosmajac, D. (2021). Offensive Language Detection in Serbian Texts Using Fine-Tuned BERT. Applied Sciences, 11(12), 5611.

  7. Mulki, H., Zaghouani, W. (2021). HabiRemoved: Hate Speech and Offensive Language Detection for Arabic. ACL Anthology Workshop Proceedings.

  8. Mamatov, A., Yusupov, K., Tursunov, O. (2020). Developing Word Embeddings for Uzbek Language Processing. Uzbek Journal of Information Technologies, 4(2), 31–44.

  9. Yusupov, B., Khoroshevsky, V. (2022). Social Media Text Analysis for Uzbek Language: Challenges and Opportunities. Proc. TürkLang 2022, 88–97.

  10. Shokh, A., Yulchiev, D. (2023). Code-Switching in Uzbek Digital Communication: Linguistic Patterns and Computational Challenges. Central Asian Computational Linguistics Conference, 14–22.

About the Authors

Javokhir Toshpulatov
Universitet

License

How to Cite

Development of an Artificial Intelligence Model for Detecting Toxic (Harmful) Content in Texts. (2026). MMIT Proceedings, 263-267. https://doi.org/10.61587/

Similar Articles

You may also start an advanced similarity search for this article.