Skip to content

Category

natural language processing

3,089 papers

#natural language process... Preprint Aug 2026

Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models

A systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs finds that at the task level, all paradigms improve over the non-finetuned baseline, but methods with comparable in-domain accuracy show substantially different knowledge transfer behavior.

Saksham Khatwani, He Cheng, Majid Afshar et al. · 0 citations
#natural language process... Preprint Jun 2026

ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

Experiments show ElementCheck consistently improves factuality verification across five backbone models while maintaining a favorable accuracy-cost trade-off, and further analyses demonstrate that complexity-aware verification reduces unnecessary re-verification and maintains stability across different backbones.

Xinming Wang, Hao-Ran Du, Yi Chen et al. · 0 citations
#natural language process... Preprint May 2026

TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

TreeGraft is a multi-drafter framework in which drafters of different costs jointly construct a shared draft tree that outperforms the better of the two fixed single-drafter endpoint strategies by 15.1% on average and reaches a maximum gain of 26.6%.

Jiaming Fan, D. Cao, Can-Chen Huang et al. · 0 citations
#natural language process... Preprint Aug 2026

One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

These results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.

Mu-Ge Zhang, Aaron Jencks, Krishna Badikela et al. · 0 citations
#computer vision Preprint Jun 2026

LV-ROVER-MLT: Low-Resource Maltese OCR by Synthetic Fine-Tuning and Multi-Stream Arbitration

Maltese has substantial text corpora and pretrained language models, but paragraph-scale OCR training data remains scarce; NOMOCRAT provides 57 verified annotated pages. LV-ROVER-MLT combines synthetic fine-tuning of Tesseract~5 with five complementary recognition streams and lexicon-gated word-level arbitration adapted to Maltese diacritics and hyphenation. In the DocEng~2026 Maltese OCR competition, the system placed first with held-out CER 0.0074; the next-ranked submission scored 0.0161 and NOMOCRAT scored 0.0163. The same approach produced a significant improvement over stock Tesseract on Luxembourgish, while the Hungarian result was inconclusive. A 36,803-pair Maltese OCR corpus constructed from EUR-Lex and Wikipedia provides an additional paragraph-level resource. Code, model weights, and corpus data are public.

Adam Darmanin · 1 citation

ProfileFoundry: A Synthetic Person-Object Substrate for Privacy, Memory, and Tool-Use Evaluation in LLM Agent

ProfileFoundry is a deterministic generator and fixed reference release of 100,000 adult synthetic Person Objects, a responsible synthetic source layer for constructing downstream foundation-model evaluations involving memory, privacy, document understanding, record linkage, and agent state while keeping the synthetic person behind each artifact inspectable.

Sriram Selvam, A. Ghosh · 0 citations
#natural language process... Preprint Jun 2026

Does Finetuning with Scientific Data Increase Hallucinations? A Multi-domain Factuality Evaluation of LLMs

SciFactCheck, a benchmark of 2,500 prompts across five scientific domains, is paired with a modular evaluation framework targeting three factuality hallucination types: unverifiability, overclaim, and attribution, and fundamentally challenge current methods of domain-specific fine-tuning for factuality and call for developing improved verification infrastructure for scientific content.

Raia Abu Ahmad, Nikolas Rauscher, Ekaterina Borisova et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.