Automatic Post-Editing (APE) for low-resource languages (LRLs) often fails to improve Machine Translation (MT), and the score alone cannot say why: whether more training would help, or whether the training data is too inconsistent to learn from. We introduce a black-box, inference-time diagnostic that tells these two c...
Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya et al.· 0 citations
EnSiTa is presented, a trilingual multi-domain parallel dataset and benchmark for English, Sinhala and Tamil, and is the most extensive systematically documented multi-domain parallel data creation and benchmarking effort for low-resource MT.
Surangika Ranathunga, Nisansa de Silva, Aloka Fernando et al.· 1 citation· ⚡1
Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of...
Akesh Gunathilake, N. Karunarathna, Tharusha Bandaranayake et al.· Moratuwa Engineering Researc...· 0 citations
TripleBound is proposed, a hybrid framework for automated monolith-to-microservices decomposition that augments a heterogeneous graph neural network with weakly supervised triplet constraints derived from parser-inferred service groups based on package structure, naming conventions, and code location.
M. Weerasinghe, Himindu Kularathne, Methmini Madhushika et al.· 0 citations
Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and the limitations of static embedding alignments. This study investigates the diachronic evolution of the Sinhala language from the 13th to the 20th century using a multi-stage...
This research extends an existing multilingual LLM (Llama-3-8B) to get a better coverage for Sinhala and enhances the LLM tokenizer with Sinhala specific vocabulary and performs continual pre-training on a 10 million sentence Sinhala corpus, resulting in the SinLlama model.
The Colombo Tea Auction (CTA) plays a vital role in determining global tea prices, yet the relationship between local weather conditions and price behavior across different tea catalogues has not been thoroughly explored. In this study, we develop a novel, structured dataset by extracting information from 105 weekly br...
H. Mallawarachchi, Senilka Madurapperumage, Nadil Kulathunge et al.· 0 citations
A survey and comparative analysis of NLP-based Automatic Deception Detection focusing on the legal domain and the evolution from feature-based machine learning to Large Language Model (LLM) approaches are presented, showing strong domain sensitivity.
T. Samaradiwakara, Nisansa de Silva, George C. Lobb· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.