Why Transformers Skip the Relay
Distance need not matter. In a Transformer, self-attention lets a word draw information directly from another permitted position, even when a sentence places many words between them, because the architecture does not force that information through every intervening word. Recurrence imposes that detour. A traditional recurrent network instead updates a hidden state sequentially, carrying earlier in...More
loading...