The newest agent frontier: models that perceive a screen (via screenshots) and take UI actions — move the mouse, click, type, scroll — to complete tasks in real applications, including ones with no API. It generalizes automation beyond brittle scripts, but is slower, error-prone, and risky (it can take real actions), so it runs with confirmations, sandboxes, and human oversight for anything consequential.