1. TTK: A toolkit for Tunisian linguistic analysis.
- Author
-
Mekki, Asma, Zribi, Inès, Ellouze, Mariem, and Belguith, Lamia Hadrich
- Subjects
- *
LINGUISTIC analysis , *NATURAL language processing , *SENTIMENT analysis , *DIALECTS - Abstract
Over the last two decades, many efforts have been made to provide resources to support the Arabic Natural Language Processing (NLP). Some of these resources target specific NLP tasks such as word tokenization, parsing, or sentiment analysis, while others attempt to tackle numerous tasks at once. In this paper, we present ¡¡TTK¿¿, a toolkit for Tunisian linguistic analysis. It consists of a collection of linguistic analysis tools for orthographic normalization, sentence boundaries detection, word tokenization, morphological analysis, parsing and named entity recognition. This paper focuses on the design and implementation of TTK tools. • NLP tools give us a better understanding of how the language may work in specific situations. • TTK contains text processing tools from orthographic normalization to comprehension. • All the proposed and used tools showed encouraging results. • All the proposed and used tools trait all the three forms of Tunisian dialect. [ABSTRACT FROM AUTHOR]
- Published
- 2024
- Full Text
- View/download PDF