Plain text prompts can't precisely control composition. ControlNet adds a conditioning input — an edge map, human pose skeleton, depth map, or scribble — that constrains the diffusion to match that structure while the prompt sets style/content. It's how you get a generated image in exactly the pose or layout you want. Part of a broader toolkit (img2img, inpainting, LoRAs) for directed, controllable generation.