Text-to-video models extend diffusion/transformer techniques to time, generating coherent short clips from a prompt while maintaining consistency across frames (the hard part — objects must persist and move plausibly). Sora, Veo, and others advanced rapidly. Costs are high and lengths short, but quality is climbing fast. Implications span filmmaking, ads, and synthetic media concerns (deepfakes), making provenance/watermarking important.