Curated academic reading list of 19 sources covering foundational architectures, efficient and long-context variants, theoretical analyses, and empirical benchmarks. Duplicate URLs and lower-relevance application papers were excluded.
Foundational Transformer papers and surveys that organize the major architectural families, including the original Transformer and vision Transformers.
Papers on sparse, efficient, long-context, and alternative sequence architectures, including canonical sparse-attention work and newer surveys or preprints.
Theoretical and mechanistic studies of Transformer expressivity, induction heads, linear attention, and length generalization.
Benchmark papers assessing attention efficiency, long-context modeling, and comparisons among Transformer and alternative sequence architectures.
Create your account to keep this answer and continue from it later.
Let's look at alternatives: