E-commerce search requires distinguishing products that are merely related to a query from those that directly satisfy the user's shopping intent. We augment query-product pairs with structured LLM-generated query and product attributes and human-validated relevance, explanations, and centrality judgments, and evaluate...
Girish A. Koushik, Swapnil Bhosale, Samarth Agrawal et al.· 0 citations
When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experime...
Girish A. Koushik, Diptesh Kanojia, Helen Treharne· 0 citations
The WMT 2020-2024 shared-task lineage with an extended English-Malayalam resource is consolidated into IndicQE-APE, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level tags and an error explanation, and a test set stratified over four difficulty axes.
Diptesh Kanojia, Archchana Sindhujan, S. Deoghare et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.