Back to Search Start Over

Dependency Annotation of Ottoman Turkish with Multilingual BERT

Authors :
Özateş, Şaziye Betül
Tıraş, Tarık Emre
Genç, Efe Eren
Taşdemir, Esma Fatıma Bilgin
Source :
\c{S}aziye Bet\"ul \"Ozate\c{s}, Tar{\i}k T{\i}ra\c{s}, Efe Gen\c{c}, and Esma Bilgin Ta\c{s}demir. 2024. Dependency Annotation of Ottoman Turkish with Multilingual BERT. LAW-XVIII, pages 188-196, St. Julians, Malta
Publication Year :
2024

Abstract

This study introduces a pretrained large language model-based annotation methodology for the first de dency treebank in Ottoman Turkish. Our experimental results show that, iteratively, i) pseudo-annotating data using a multilingual BERT-based parsing model, ii) manually correcting the pseudo-annotations, and iii) fine-tuning the parsing model with the corrected annotations, we speed up and simplify the challenging dependency annotation process. The resulting treebank, that will be a part of the Universal Dependencies (UD) project, will facilitate automated analysis of Ottoman Turkish documents, unlocking the linguistic richness embedded in this historical heritage.<br />Comment: 9 pages, 5 figures. Accepted to LAW-XVIII

Details

Database :
arXiv
Journal :
\c{S}aziye Bet\"ul \"Ozate\c{s}, Tar{\i}k T{\i}ra\c{s}, Efe Gen\c{c}, and Esma Bilgin Ta\c{s}demir. 2024. Dependency Annotation of Ottoman Turkish with Multilingual BERT. LAW-XVIII, pages 188-196, St. Julians, Malta
Publication Type :
Report
Accession number :
edsarx.2402.14743
Document Type :
Working Paper