Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context
Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which laye...