AND-2500: A Novel Assamese News Corpus and Comparative Benchmarking of Language Models
Abstract
The rapid advancement of Large Language Models (LLMs) has transformed the landscape of Natural Language Processing (NLP), enabling remarkable progress in language understanding and generation. However, these advances remain unevenly distributed, particularly for low-resource languages such as Assamese, which suffer from a severe lack of high-quality open-sourced annotated corpora and standardized benchmarks. To address this gap, this paper introduces a novel AND-2500 dataset, which is manually curated by expert validation. The corpus comprises Assamese news articles annotated across 20 fine-grained categories following the International Press Telecommunications Council (IPTC) media topic taxonomy. This paper presents a comprehensive benchmarking framework that systematically evaluates both traditional encoder-only transformer models (MuRIL, IndicBERT, XLM-R etc.) and state-of-the-art generative LLMs (GPT2, Llama-3, Qwen2.5, Mistral and Aya-23) under supervised, zero-shot, and few-shot learning paradigms. Additionally, this paper assesses the generative capability of LLMs through a headline generation task, employing lexical and semantic evaluation metrics, including F1-score, BERTScore, and Semantic Answer Similarity (SAS). Our results demonstrate that supervised fine-tuned multilingual encoders like MuRIL remain the gold standard for classification (F1 score: 0.89) of Assamese text, whereas LLMs like Llama-3 can achieve near-supervised performance (F1: 0.88) with just twenty examples from AND-2500, offering a computationally efficient pathway for deploying LLMs in resource-constrained environments. This work establishes a vital baseline for Assamese NLP and promotes the open sourcing of linguistic resources for low-resource language.
Keywords
Dataset, LLMs, transformer language models, assamese language, low-resource NLP.