An agent takes actions in an environment, receives rewards, and learns a policy that maximizes long-term reward through trial and error. It powers game-playing (AlphaGo), robotics, and (as RLHF) LLM alignment. It's powerful for sequential decisions but sample-hungry and sensitive to how you design the reward — a bad reward yields clever, unintended behavior.