A model can only attend to a fixed number of tokens — its context window (from a few thousand to millions in modern models). Everything the model "knows" in a conversation must fit: system prompt, history, documents, and the reply. Exceed it and earlier content is truncated or must be summarized. Bigger windows cost more compute (attention scales with length), so there's a price for using them.