DETECTION OF EMOTIONAL ABUSE IN DIGITAL TEXT USING NATURAL LANGUAGE PROCESSING
Main Article Content
Abstract
Emotional abuse in digital communication is difficult to detect because it relies on manipulation, gaslighting, invalidation, and coercive control rather than explicit profanity or hate speech. These patterns are psychologically harmful but linguistically subtle, and they pass through conventional content moderation systems without being flagged. This paper presents the development and evaluation of a Natural Language Processing (NLP) system for detecting emotionally abusive language in short-form digital text. Three classification models were compared: Multinomial Naive Bayes, Linear Support Vector Machine (SVM), and a fine-tuned RoBERTa transformer. The models were trained on a combined dataset of 85,487 labeled samples from the Davidson et al. [2] Hate Speech and Offensive Language dataset and the Jigsaw Toxic Comment Classification dataset. The SVM achieved the highest accuracy at 95.41% with a precision of 0.9558 and an F1-score of 0.8991. RoBERTa achieved 92.61% accuracy and an F1-score of 0.8852, while Naive Bayes recorded 92.30% and an F1-score of 0.8792. A qualitative error analysis identified systematic model failures on sarcasm, gaslighting, and coercive framing. The system was deployed as a real-time web application using Streamlit and FastAPI.
Article Details
Section
COPYRIGHT
Submission of a manuscript implies: that the work described has not been published before, that it is not under consideration for publication elsewhere; that if and when the manuscript is accepted for publication, the authors agree to automatic transfer of the copyright to the publisher.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work
- The journal allows the author(s) to retain publishing rights without restrictions.
- The journal allows the author(s) to hold the copyright without restrictions.
References
[1] Statista, "Number of social media users worldwide from 2017 to 2025," Statista Research Department, 2025.
[2] T. Davidson, D. Warmsley, M. Macy, and I. Weber, "Automated hate speech detection and the problem of offensive language," Proc. Int. AAAI Conf. Web and Social Media, vol. 11, no. 1, pp. 512-515, 2017.
[3] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," Proc. 2019 Conf. North American Chapter of the ACL, pp. 4171-4186, 2019.
[4] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, "RoBERTa: A robustly optimised BERT pretraining approach," arXiv preprint arXiv:1907.11692, 2019.
[5] E. Spertus, "Smokey: Automatic recognition of hostile messages in online discussions," Proc. Conf. Human Factors in Computing Systems, pp. 1058-1065, 1997.
[6] A. H. Razavi, D. Inkpen, S. Uritsky, and S. Matwin, "Offensive language detection using multi-level classification," Canadian Conf. Artificial Intelligence, pp. 16-27, 2010.
[7] W. Warner and J. Hirschberg, "Detecting hate speech on the world wide web," Proc. Second Workshop on Language in Social Media, pp. 19-26, 2012.
[8] B. Vidgen, T. Thrush, Z. Waseem, and D. Kiela, "Learning from the worst: Dynamically generated datasets to improve online hate detection," Proc. 59th Annual Meeting of the ACL, pp. 1667-1682, 2021.
[9] S. Rajamanickam, P. Mishra, H. Yannakoudakis, and E. Shutova, "Multi-task learning for emotion and abuse detection in online text," Proc. 60th Annual Meeting of the ACL, pp. 3456-3472, 2022.
[10] S. Giorgi, S. C. Guntuku, and L. H. Ungar, "Integrating psychological frameworks into transformer architectures for detecting gaslighting in digital communication," Proc. ACL, vol. 61, pp. 2345-2362, 2023.
[11] R. Pradhan, T. Chakraborty, and N. Goyal, "Modelling emotional trajectories for detecting manipulation and coercive control in digital conversations," ACM Trans. Intelligent Systems and Technology, vol. 15, no. 2, pp. 1-24, 2023.
[12] V. Kumar, R. Singh, and M. Gupta, "Conversation-level models for detecting relational abuse patterns in digital communication," Proc. ACM Conf. Fairness, Accountability, and Transparency, vol. 7, pp. 234-251, 2024.
[13] L. Zhang and Y. Chen, "Emotional abuse detection in digital text: A systematic review of approaches and research gaps," Computers in Human Behavior, vol. 150, pp. 107-128, 2024.
[14] C. Nobata, J. Tetreault, A. Thomas, Y. Mehdad, and Y. Chang, "Abusive language detection in online user content," Proc. 25th Int. Conf. World Wide Web, pp. 145-153, 2016.
[15] A. Lees, A. Borkowski, and M. Johnson, "Coercive control in digital spaces: Linguistic patterns of emotional abuse," J. Interpersonal Violence, vol. 37, no. 15-16, pp. 14567-14589, 2022.
[16] A. Adebayo and T. Ogunleye, "Digital emotional abuse among Nigerian university students: Prevalence and reporting patterns," J. African Digital Humanities, vol. 8, no. 2, pp. 45-62, 2023.
[17] C. Nobata, J. Tetreault, A. Thomas, Y. Mehdad, and Y. Chang, "Abusive language detection in online user content," Proc. 25th Int. Conf. World Wide Web, pp. 145-153, 2016.