A Silver-Standard POS Corpus and Baseline Models for Tagin, a Low-Resource Language

Tungon Dugi, Koj Sambyo

Abstract


This study presents the design of a machinelearningbased partofspeech POS tagger and a silverstandard POS tagged corpus for the lowresource Tagin language Due to the lack of digitally available text data applying modern NLP techniques to the Tagin language remains extremely challenging To address this we constructed a POS tagged corpus comprising 5000 sentences collected from legacy Tagin dictionaries and grammar books with 23136 tokens and 4139 unique word forms We employed 21 UD v2–compatible POS tags with Taginspecific morphosyntactic adaptations For benchmarking we evaluated four classical sequencelabeling models HMM CRF Average Perceptron and Maximum Entropy Across all evaluation metrics the CRF model achieved the best performance over all the metrics with a word level accuracy of 8180 and OOV accuracy of 5676 The worstperforming model was the HMM with a word accuracy of 7718 and a notably low OOV accuracy of 3962 We also discuss recently proposed POS taggers for lowresource languages and other relevant NLP studies for comparative purposes To promote reproducibility and future research we have released the dataset and model details in a publicly accessible repository

Keywords


Part-of-speech tagging, low-resource language, Tagin language, conditional random fields, sequence labeling.

Full Text: PDF