When Similar Means Different: Evaluating LLMs on Arabic-Hebrew Cognates
This work introduces SemCog Bench, a curated benchmark of 1,858 Arabic--Hebrew word pairs with sentence-level annotations for cognate identification and semantic disambiguation and finds that context and scale yield model-dependent gains, while original-script inputs generally perform best.