The mechanics: split documents into chunks, embed each into a vector, store them in a vector database. At query time, embed the question and retrieve the nearest chunks by vector similarity, then pass them to the LLM. Chunk size, overlap, and embedding quality make or break results — too-big chunks dilute relevance, too-small lose context. RAG is mostly a retrieval-quality problem.