Skip to content
Review

Survey The Quest for the Right Mediator: Surveying Mechanistic Interpretability for NLP Through the Lens of Causal Mediation Analysis

· 0 citations · 205 references

TL;DR

The history and current state of interpretability taxonomized according to the types of causal units utilized, as well as methods used to search over mediators are described.

View source

Similar papers

#artificial intelligence Review Sep 2026

Contextual Causality with Large Language Models: A Survey

Understanding contextual causality is critical for large language models (LLMs), as it enables them to accurately identify causal relations in specific situations and support more reliable decision-making. Despite its significance, a systematic exploration of contextual causality with LLMs is still lacking. To fill thi...

Yi-Heng Zhao, Jun Yan, Chengming Hu · 0 citations

Beyond Good Intentions: When Does the Framing of Multilingual and Low-Resource NLP Research Become a Caricature?

Building language technologies and conducting NLP research for low-resource languages---particularly when led by native speakers or involving participatory research practices---are often framed as means of addressing inequality, serving local communities, and, at times, contributing to *decolonisation*. In this paper,...

Nedjma Djouhra Ousidhoum, Noopur Zambare, Mohamed Abdalla · 0 citations
Conference Open access Sep 2026

A Survey on Actionable Interpretability in Large Language Models

This survey reviews LLM interpretability through the lens of actionability, presenting a taxonomy of attributional and mechanistic approaches, along with emerging methods tailored to vision–language models (VLMs), and examining how actionable interpretability supports downstream objectives.

Jie Cai, Mafizur Rahman, James Enouen et al. · 0 citations
Open access 2026

The structure of reasoning: Inferring conceptual networks from text

For decades, public opinion scholars have argued for the need to go beyond measuring isolated political preferences to more richly examine how individuals reason about and justify the interconnections between their preferences. While early efforts used interviews and hand-coding to elicit the network structure of sub...

Sarah Shugars, Xin-Feng Gu · 0 citations
Review Sep 2026

Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers

Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique the...

YunHong Yang, Mike Thelwall, Guo-Xiu He · 0 citations
Preprint Aug 2026

Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpret...

Alona Strugatski, Licol Zeinfeld, Jason Cooper et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.