GPT-style models use only the decoder with causal (masked) attention: each token sees only what came before, so they generate text one token at a time. Pretrained on next-token prediction over huge corpora, then instruction-tuned, they became the dominant LLM design — GPT, Claude, Llama, Gemini are all decoder-only. Generation, chat, code, reasoning all flow from this single objective.