Skip to content

Author

Surangika Ranathunga

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Aug 2026

Confident but Wrong: A Constrained Decoding Diagnostic for Low-Resource Automatic Post-Editing

Automatic Post-Editing (APE) for low-resource languages (LRLs) often fails to improve Machine Translation (MT), and the score alone cannot say why: whether more training would help, or whether the training data is too inconsistent to learn from. We introduce a black-box, inference-time diagnostic that tells these two c...

Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya et al. · 0 citations
#natural language process... Preprint Aug 2026

EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation

EnSiTa is presented, a trilingual multi-domain parallel dataset and benchmark for English, Sinhala and Tamil, and is the most extensive systematically documented multi-domain parallel data creation and benchmarking effort for low-resource MT.

Surangika Ranathunga, Nisansa de Silva, Aloka Fernando et al. · 1 citation · ⚡1
Conference Aug 2026

LMSpell: Spell Correction with Pre-Trained Language Models

Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of...

Akesh Gunathilake, N. Karunarathna, Tharusha Bandaranayake et al. · 0 citations
#natural language process... Conference Open access Aug 2025

SinLlama - A Large Language Model for Sinhala

This research extends an existing multilingual LLM (Llama-3-8B) to get a better coverage for Sinhala and enhances the LLM tokenizer with Sinhala specific vocabulary and performs continual pre-training on a 10 million sentence Sinhala corpus, resulting in the SinLlama model.

H.W.K. Aravinda, Rashad Sirajudeen, Samith Karunathilake et al. · 10 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.