Skip to content

Category

natural language processing

2,926 papers

Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study

Multilingual sentence embeddings are increasingly used to estimate semantic similarity across languages, yet their sensitivity to fine-grained translation errors remains insufficiently understood. This study investigates whether general-purpose multilingual embedding models can distinguish correct English-Greek translations from minimally modified erroneous alternatives. A contrastive dataset was developed from FLORES+ sentence-aligned reference translations and reviewed by two translation experts. It contains 1,850 examples across ten core and five exploratory error categories, covering factual, lexical-semantic, grammatical, relational, referential, and discourse-level phenomena. Five multilingual sentence-embedding models (BGE-M3, Multilingual E5, Multilingual MPNet, LaBSE, and Jina Embeddings v3) were evaluated using cosine similarity between each English source sentence and its correct and erroneous Greek translations. A reference-free COMETKiwi model was also evaluated as an MT quality-estimation baseline. Performance was assessed through contrastive accuracy and score margins for category-specific sensitivity. BGE-M3 achieved the highest accuracy among embedding models at 89.30 percent, while COMETKiwi achieved 94.49 percent. Embedding models detected explicit factual and lexical changes more reliably than tense-and-aspect and pronoun-coreference errors. COMETKiwi improved performance on several difficult categories, including tense and aspect, pronoun and coreference, and semantic-role errors, but showed lower sensitivity to date-and-time errors and underperformed the embedding models on numbers. The results show complementary error-sensitivity profiles: multilingual sentence embeddings provide useful semantic adequacy signals but are better suited as components of broader translation-evaluation frameworks than as standalone metrics.

Eleftherios Kalogeros, Athanasios Ntalakas, M. Gergatsoulis et al. · 0 citations
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with questions requiring precise spatial reasoning and fine-grained visual understanding. These limitations often manifest as object, attribute, and spatial hallucinations, where models generate confident but visually unsupported responses due to insufficient region-level and fine-grained visual grounding. To address this challenge, we propose ReVA, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space. The image bridge maps final transformer block features into image tokens. The region bridge maps cropped features from enriched intermediate features across ViT blocks so early texture and later object cues are more evident, into K region tokens for every bounding box. ReVA uses a detector stack that supplies automatic zero-shot bounding boxes that are both question-agnostic and question-dependent, using RAM++ (Recognize Anything Model), spaCy, and Grounding DINO. The image tokens and region tokens are concatenated as an LLM prompt prefix to jointly encode scene-level context and fine-grained regional evidence when answering questions. Evaluated on VQAv2, MMBench, POPE, and SEED-Bench, ReVA achieves 82.85% mean F1 on POPE, compared with 81.14% for an image-token baseline without region tokens. These results demonstrate that explicit region-aware visual representations reduce object hallucination and improve the factual grounding of MLLMs.

A. Senthil · 0 citations
#artificial intelligence Preprint Aug 2026

GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon

GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs across three NLP tasks on an Apple M4 Pro with 48 GB unified memory, is presented.

R. Kannan, Rajendra P. Firke, Shreya Bengle et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?

This evaluation shows that while LLMs effectively uncover a substantial portion of implicitly-related touchpoints, significant room for improvement remains in their selection performance, and offers a new roadmap for transitioning conversion attribution from mechanical rule-matching to human-aligned semantic reasoning.

Jinqi Wu, Sishuo Chen, Zhangming Chan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Redesigning and Auditing Deep Research Writing for Faithful Reports

CLAIMPROBE is introduced, a claim-level audit that decomposes DR reports into claims and measures hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence and proposes CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to a query-derived outline, and drafts each section from a source-linked claim representation.

Hiroaki Hayashi, P. Venkit, Prafulla Kumar Choubey et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

Evaluation of six frontier models reveals that even the strongest model reaches only 63.1\% pass rate, with many tasks unsolved by any model, highlighting that multilingual coding competence is a distinct and underexplored capability axis.

Yunsu Kim, Kaden Uhlig, Ashwin Purohit et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework

The Prompt Phrases Prediction Network (PPN) is introduced, an encoder-decoder architecture designed to effectively extract keyword prompts embeddings and infuse the prompt embedding into the Prompt-guided KWS encoder by utilizing a Prompt-acoustic Multi-head Cross-attention (MHCA).

G. Xu, Cheng-Fei Li, Xian-Liang Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation

PAUSE (Pause-And-Update Strategy Editing) is an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long-form story adaptation, a structured artifact that can be inspected, edited, and then projected through downstream character, entity, and chapter-localization stages.

Taaha Kazi, Vasu Sharma, M. Saifullah et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework

This work proposes LPS-TC, a Lightweight Proactive Speech Turn Controller for plug-and-play integration, and introduces a two-tier evaluation scheme that assesses both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints.

Tianrui Pan, Qinglin Zhang, Chong Deng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models

This study establishes an end-to-end prototype from raw BIM data input, through defect identification, to repair suggestion generation, and establishes an end-to-end prototype to identify and repair various defects in BIM via domain-specific LLMs.

Jia-Rui Lin, Yunzhen Cai, Xiang Ni et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.