Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Walk through how attention can resolve a pronoun to its antecedent with an example?
What you need to know
Step by step through the layers
- At "it", early layers — "it" has little meaning of its own. Its query looks for nearby nouns; "trophy" and "suitcase" score highest, so its context vector now carries "one of these two objects".
- At "fit in" — this position links "trophy" (the thing put inside) with "suitcase" (the container).
- At "big", later layers — "big" attends back to "it" and learns what it describes. Its query now encodes "the thing that was too big to fit", which matches the key of "trophy" better, because the trophy is the thing being put in.
- After that — the residual stream at "big" and later positions represents "the trophy was too big", which the rest of the sentence and the next-token prediction can use.
In a bidirectional encoder such as BERT, the position of "it" itself can see "big", so the same resolution can happen directly at "it" in later layers. In a decoder, it cannot, because "big" is in its future.
An illustrative picture of how one head's weights might look (made-up numbers):
| Position and layer | trophy | suitcase | other tokens |
|---|---|---|---|
| "it", early | 0.35 | 0.33 | 0.32 |
| "big", later | 0.71 | 0.12 | 0.17 |
The Winograd pair
"because it was too big" → trophy. "because it was too small" → suitcase. Only one word changes, and it is after the pronoun, so resolving "it" needs knowledge of the world (big things do not fit in small containers), not just grammar. This is the Winograd Schema Challenge (Levesque et al., 2012). Large models now do well on it; small ones often guess.
Two honest caveats
- No single "coreference head". Research finds some heads that track coreference well, but the resolution usually emerges across many heads and layers.
- Attention weight is not explanation. Jain and Wallace (2019), "Attention is not Explanation", showed attention weights can often be changed a lot without changing predictions. High weight from "it" to "trophy" suggests the link; ablation or probing tests it.
A real-life example
An Indian-language news app translates English stories into Hindi. Hindi marks grammatical gender on verbs and adjectives — for example "बड़ा था" for a masculine noun and "बड़ी थी" for a feminine one — so the translator must know which noun "it" refers to before it can write the ending.
A story says "The bridge replaced the old road because it was unsafe." The two nouns have different genders in Hindi. If the model links "it" to the wrong noun, the Hindi sentence is grammatical but says the wrong thing was unsafe. The team builds a test set of 300 such sentences from real stories, each with a clear referent, and measures the model on them before each upgrade. Their rule: a model change that lowers accuracy on that set does not ship, even if overall translation scores rise.
Follow-up questions to expect
- "How would you check that a model really uses 'trophy'?" — Patch or ablate: replace the value vector of "trophy" in the relevant heads and see whether the prediction at "it" changes.
- "Could a single layer solve it?" — Rarely. The clue ("too big") must first be linked to "it", and "it" to the nouns, before the right noun can be chosen; that takes several layers.
- "Why can't the "it" position decide on its own in GPT?" — The causal mask hides "big", which comes later; the decision is carried by later positions, which see the pronoun, both nouns and the clue.