Why GPT Ditched Half the Transformer Architecture

The original transformer had both an encoder and a decoder. GPT threw away the encoder entirely. Here's why decoder-only architecture became the backbone of modern generative AI — and what it means for the models powering today's synthetic media.

Share
Why GPT Ditched Half the Transformer Architecture

The transformer architecture introduced in the 2017 paper "Attention Is All You Need" was fundamentally a two-part machine: an encoder that read and understood an input sequence, and a decoder that generated an output sequence. It was designed for translation — read a sentence in English (encoder), produce a sentence in French (decoder). Yet the family of models that came to dominate generative AI, OpenAI's GPT series, quietly threw away half of that design. GPT is a decoder-only model. Understanding why reveals a lot about how today's generative systems — including the ones that power text, image captioning, and multimodal synthetic media pipelines — actually work.

The Original Two-Part Design

In the vanilla transformer, the encoder processes the entire input at once using bidirectional self-attention. Every token can attend to every other token, both to its left and right. This is ideal for understanding tasks where you have the full context available — like classifying a sentence or translating a complete phrase.

The decoder, by contrast, uses masked (causal) self-attention. Each token can only attend to tokens that came before it, never the future. This constraint exists because during generation, the model doesn't yet know what comes next — it must predict one token at a time. The decoder also included a cross-attention layer, which let it "look at" the encoder's representation of the input while generating output.

Why GPT Dropped the Encoder

The key insight behind GPT is that language generation can be framed entirely as next-token prediction. If your goal is to predict the next word given everything that came before, you don't need a separate encoder to "understand" a distinct input sequence. The input and the output live in the same stream of tokens. Prompt and completion are simply earlier and later parts of one continuous sequence.

This reframing collapses the architecture. Without a separate input sequence to encode, the cross-attention mechanism becomes unnecessary — there's nothing external to cross-attend to. What remains is a stack of decoder blocks using masked self-attention, predicting each token from the ones preceding it. The encoder, and the cross-attention plumbing that connected it to the decoder, simply vanish.

The Payoff: Scale and Simplicity

A decoder-only design has profound practical advantages. First, it is architecturally simpler — one type of block, repeated. This makes it far easier to scale to hundreds of billions of parameters, which is precisely the direction GPT-2, GPT-3, and successors took.

Second, it enables unsupervised pretraining on raw text at massive scale. Because the training objective is just "predict the next token," any body of text becomes training data without labels or paired input-output examples. This unlocked the internet-scale corpora that made large language models possible.

Third, decoder-only models proved to be surprisingly general. Tasks once thought to require an encoder — summarization, question answering, even translation — can all be expressed as next-token prediction given the right prompt. The in-context learning behavior that made GPT-3 famous emerged naturally from this unified formulation.

Encoders Didn't Disappear Entirely

It's worth noting that the encoder-only branch of the family tree is alive too. Models like BERT use only the encoder with bidirectional attention, excelling at understanding and classification tasks rather than generation. And encoder-decoder models like T5 remain strong for structured input-output tasks. The architecture you choose depends on the job: encoders for understanding, decoders for generating, both for translating between distinct sequences.

Why This Matters for Synthetic Media

The decoder-only paradigm isn't confined to text. The same autoregressive, next-token logic underpins many generative systems across modalities — from autoregressive image and audio models to the language backbones that steer multimodal generation. When a system generates a video caption, scripts a synthetic voice, or drives a text-to-image prompt, it often leans on decoder-only transformers doing exactly what GPT pioneered: predicting the next unit in a sequence.

Understanding this architectural choice helps demystify why modern generative AI behaves the way it does — why it's probabilistic, why context windows matter, and why scaling data and parameters has been so effective. The decision to "ditch half the transformer" wasn't a loss; it was a focusing of the architecture around the single task of generation, a bet that paid off and reshaped the entire field.


Stay informed on AI video and digital authenticity. Follow Skrew AI News.