Conceptual

Transformer Architecture

Stacks of self-attention and position-wise feedforward blocks with residual connections and normalization, processing all positions in parallel with order injected via positional information. Recurrence is eliminated entirely, which is exactly what makes training parallelizable.