This paper compares three constituency parsing representations derived from the Penn Korean Treebank, and shows that fine-grained morphological and XPOS representations provide valuable evidence for the evaluated parsers.
Jungyeul Park, Kyungtae Lim, Zihao Huang et al.· 0 citations
A systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs finds that at the task level, all paradigms improve over the non-finetuned baseline, but methods with comparable in-domain accuracy show substantially different knowledge transfer behavior.
Saksham Khatwani, He Cheng, Majid Afshar et al.· 0 citations
Experiments show ElementCheck consistently improves factuality verification across five backbone models while maintaining a favorable accuracy-cost trade-off, and further analyses demonstrate that complexity-aware verification reduces unnecessary re-verification and maintains stability across different backbones.
Xinming Wang, Hao-Ran Du, Yi Chen et al.· 0 citations
TreeGraft is a multi-drafter framework in which drafters of different costs jointly construct a shared draft tree that outperforms the better of the two fixed single-drafter endpoint strategies by 15.1% on average and reaches a maximum gain of 26.6%.
Jiaming Fan, D. Cao, Can-Chen Huang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
These results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.
Mu-Ge Zhang, Aaron Jencks, Krishna Badikela et al.· 0 citations
This time, the SHROOM-Visions task aims to tackle hallucinations through a model-agnostic detection task focused on large vision-language models, building on the recently introduced SHEEP dataset.
Raúl Vázquez, Aman Sinha, Chuyuan Li et al.· 0 citations
Results suggest that interpretable linguistic organization emerges within MoE routing patterns even without sequential language exposure, and suggest that staged bilingual exposure reduces single-language dominance.
A recurrent-style training strategy is proposed, which enables transformers to reuse their reasoning circuitry across input forms and substantially improves generalization on out-of-distribution two-hop queries.
Zili Zhang, Yilin Wang, Heng Wang et al.· 0 citations
Maltese has substantial text corpora and pretrained language models, but paragraph-scale OCR training data remains scarce; NOMOCRAT provides 57 verified annotated pages. LV-ROVER-MLT combines synthetic fine-tuning of Tesseract~5 with five complementary recognition streams and lexicon-gated word-level arbitration adapted to Maltese diacritics and hyphenation. In the DocEng~2026 Maltese OCR competition, the system placed first with held-out CER 0.0074; the next-ranked submission scored 0.0161 and NOMOCRAT scored 0.0163. The same approach produced a significant improvement over stock Tesseract on Luxembourgish, while the Hungarian result was inconclusive. A 36,803-pair Maltese OCR corpus constructed from EUR-Lex and Wikipedia provides an additional paragraph-level resource. Code, model weights, and corpus data are public.
ProfileFoundry is a deterministic generator and fixed reference release of 100,000 adult synthetic Person Objects, a responsible synthetic source layer for constructing downstream foundation-model evaluations involving memory, privacy, document understanding, record linkage, and agent state while keeping the synthetic person behind each artifact inspectable.
SciFactCheck, a benchmark of 2,500 prompts across five scientific domains, is paired with a modular evaluation framework targeting three factuality hallucination types: unverifiability, overclaim, and attribution, and fundamentally challenge current methods of domain-specific fine-tuning for factuality and call for developing improved verification infrastructure for scientific content.
Raia Abu Ahmad, Nikolas Rauscher, Ekaterina Borisova et al.· 0 citations
The Persuasion Index is proposed, a taxonomy of 15 dimensions grounded in persuasion theories from psychology and communication and one transparent implementation using 55 sub-features built from lexicons and rule-based detectors.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.