TL;DR

The Transformer (Vaswani et al., 2017, "Attention Is All You Need") replaces recurrence with self-attention, enabling parallelism and long-range modeling. It is the foundation of modern large language models.

Key Points

Further Notes

Sources

最后更新:2026-08-02