Automatic Post-Editing (APE) for low-resource languages (LRLs) often fails to improve Machine Translation (MT), and the score alone cannot say why: whether more training would help, or whether the training data is too inconsistent to learn from. We introduce a black-box, inference-time diagnostic that tells these two c...
Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya et al.· 0 citations
EnSiTa is presented, a trilingual multi-domain parallel dataset and benchmark for English, Sinhala and Tamil, and is the most extensive systematically documented multi-domain parallel data creation and benchmarking effort for low-resource MT.
Surangika Ranathunga, Nisansa de Silva, Aloka Fernando et al.· 1 citation· ⚡1
Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of...
Akesh Gunathilake, N. Karunarathna, Tharusha Bandaranayake et al.· Moratuwa Engineering Researc...· 0 citations
This research extends an existing multilingual LLM (Llama-3-8B) to get a better coverage for Sinhala and enhances the LLM tokenizer with Sinhala specific vocabulary and performs continual pre-training on a 10 million sentence Sinhala corpus, resulting in the SinLlama model.