A multimodal model takes and/or produces multiple data types: describe an image, answer questions about a chart, transcribe audio, generate pictures from text. The trick is mapping each modality into a shared representation the model can reason over jointly. Modern frontier models are natively multimodal — vision + language in one model — which unlocks document understanding, accessibility, and richer assistants.