One attention "head" can only focus one way. Multi-head attention splits the representation into several heads, each learning its own Q/K/V projections — one head might track syntax, another long-range references, another local context. Their outputs are concatenated and combined. Multiple parallel perspectives are why transformers capture rich, varied relationships in a single layer.