1. A Big Data-empowered System for Real-time Detection of Regional Discriminatory Comments on Vietnamese Social Media
- Author
-
Huynh, An Nghiep, Do, Thanh Dat, and Do, Trong Hop
- Subjects
Computer Science - Computation and Language ,Computer Science - Computers and Society - Abstract
Regional discrimination is a persistent social issue in Vietnam. While existing research has explored hate speech in the Vietnamese language, the specific issue of regional discrimination remains under-addressed. Previous studies primarily focused on model development without considering practical system implementation. In this work, we propose a task called Detection of Regional Discriminatory Comments on Vietnamese Social Media, leveraging the power of machine learning and transfer learning models. We have built the ViRDC (Vietnamese Regional Discrimination Comments) dataset, which contains comments from social media platforms, providing a valuable resource for further research and development. Our approach integrates streaming capabilities to process real-time data from social media networks, ensuring the system's scalability and responsiveness. We developed the system on the Apache Spark framework to efficiently handle increasing data inputs during streaming. Our system offers a comprehensive solution for the real-time detection of regional discrimination in Vietnam., Comment: accepted by 2024 International Conference on Advanced Technologies for Communications (ATC) Program
- Published
- 2024