For nearly a decade, convolutional neural networks defined how machines see. Their architecture encoded a powerful assumption: that visual understanding emerges from local patterns, progressively composed into global meaning. This inductive bias worked remarkably well, but it also imposed a ceiling on what image models could learn from data alone.
The Vision Transformer, introduced by Dosovitskiy and colleagues in 2020, questioned this orthodoxy. It proposed something radical: treat an image as a sequence of patches and process it with the same transformer machinery that revolutionized natural language processing. No convolutions, no hand-crafted spatial hierarchies, just attention.
The results reshaped the field. Given sufficient data, ViT variants match or exceed CNN performance across benchmarks, from ImageNet classification to dense prediction tasks. But the deeper lesson lies in the architectural decisions that made this transfer possible. Understanding these choices reveals not just how ViTs work, but how architectural assumptions shape what a model can and cannot learn.
Patch Embedding Strategy
The core challenge in adapting transformers to vision is representational: transformers operate on sequences of tokens, while images are dense two-dimensional grids of pixels. A naive approach—treating every pixel as a token—would produce sequences of tens of thousands of elements, making self-attention's quadratic complexity computationally prohibitive.
ViT solves this through patch embedding. An input image, typically 224×224 pixels, is divided into non-overlapping patches of 16×16 pixels, yielding 196 patches. Each patch is flattened into a 768-dimensional vector and projected through a single linear layer into the transformer's embedding space. This projection is functionally equivalent to a convolution with stride equal to kernel size, but conceptually it establishes patches as the atomic units of visual reasoning.
The patch size represents a critical trade-off. Smaller patches (8×8) capture finer detail but produce longer sequences and higher compute costs. Larger patches (32×32) are efficient but sacrifice spatial granularity. ViT-B/16 became the reference configuration because 16×16 patches balance these concerns for standard resolution inputs, though variants like ViT-L/14 in CLIP demonstrate that smaller patches pay dividends when data and compute allow.
A learnable [CLS] token is prepended to the patch sequence, mirroring BERT's classification token design. After processing, this token's representation serves as the aggregate image embedding for downstream tasks. This unified tokenization strategy is what makes ViT genuinely architecture-agnostic across modalities.
TakeawayTokenization is not preprocessing—it is the first architectural decision that defines what your model can perceive. Choose the granularity of your tokens with the same care you choose your loss function.
Inductive Bias Trade-offs
CNNs encode two powerful inductive biases: locality (nearby pixels are correlated) and translation equivariance (a cat in the corner is still a cat). These assumptions are baked into the architecture through weight sharing and local receptive fields, meaning the model doesn't need to learn them from data.
Vision Transformers abandon these priors. Self-attention treats all patches as equally related a priori, allowing any patch to attend to any other from the first layer. This flexibility is both ViT's greatest strength and its most demanding constraint. Without built-in spatial priors, the model must learn visual structure entirely from data—which requires significantly more data to reach the same performance.
The empirical evidence is stark. On ImageNet-1k alone (roughly 1.3M images), CNNs like ResNet still outperform vanilla ViTs. But scale up to JFT-300M or LAION-scale datasets, and ViTs surpass CNNs decisively. The crossover point reveals a fundamental principle: inductive biases are a form of compressed prior knowledge that reduces sample complexity, but they also constrain the hypothesis space the model can explore.
Hybrid architectures like Swin Transformer and ConvNeXt navigate this trade-off pragmatically, reintroducing locality through windowed attention or convolution-inspired design. The lesson is that inductive bias exists on a spectrum, and the optimal position depends on your data regime, compute budget, and the diversity of patterns you need to represent.
TakeawayInductive biases are neither good nor bad—they are trades of flexibility for sample efficiency. The right architecture is the one whose priors match your data budget.
Positional Information for Images
Self-attention is permutation-invariant: shuffle the input tokens and the output changes only in ordering, not content. For text, this is corrected with positional encodings that inject sequence order. For images, the problem is more subtle—patches have not just an order, but a two-dimensional spatial relationship that carries semantic meaning.
The original ViT uses learned 1D positional embeddings, one per patch position, added to patch embeddings before the first transformer layer. Remarkably, ablation studies showed that 2D-aware embeddings offered minimal advantage over 1D on standard benchmarks. The model appears to recover 2D structure from the data itself, given enough examples.
However, this approach has limitations that later architectures address. Fixed learned embeddings tie the model to a specific input resolution, complicating transfer to higher-resolution images without interpolation. Relative positional encodings, as used in Swin Transformer, and rotary embeddings (RoPE), popularized in language models and adopted in newer ViT variants, offer resolution-independent and translation-aware alternatives.
The deeper insight is that positional information in vision is not just about location—it encodes the implicit geometry of the visual world. How you encode position determines whether your model treats images as sequences that happen to be 2D, or as genuinely spatial objects. Architectural decisions about position propagate through every attention computation, shaping what relationships the model can efficiently represent.
TakeawayPosition is not metadata—it is the coordinate system in which meaning unfolds. Every choice about encoding position is a choice about what geometry your model believes in.
Vision Transformers demonstrate that architectural assumptions are not neutral—they determine what a model can learn, how much data it needs, and which patterns it will discover. By replacing CNN priors with pure attention, ViT traded sample efficiency for representational flexibility, a trade that pays off at scale.
The broader principle extends beyond vision. As models generalize across modalities, the most successful architectures are those that make minimal, well-chosen assumptions about their input structure. Tokenization, inductive bias, and positional encoding form a triad of design decisions that recur in every domain.
Understanding these principles is not merely academic. Whether you're fine-tuning a pretrained ViT or designing a novel architecture, these are the levers you control.