If many requests share a long prefix (a big system prompt, a document), recomputing it every time is wasteful. Prompt/context caching stores the processed prefix so repeated calls skip that work — major cost and latency savings for chatbots and RAG with stable instructions. Structure prompts with the stable, cacheable part first and the variable part last to maximize cache hits.