Multimodal AI
A model that can take in, or generate, more than one type of data — text, images, audio, video — instead of only text.
Early large language models only read and wrote text. A multimodal model can also look at a photo and describe it, listen to speech, watch video, or generate an image — because it's been trained to map different kinds of data into the same underlying representation the model reasons over.
Multimodality stops being a demo trick and becomes a hard requirement the moment AI has to act in the physical world: a robot or an autonomous vehicle needs to see, hear, and reason in the same model, not stitch together separate single-purpose systems — which is a big part of why 'physical AI' is treated as a harder, later milestone than chatbot-style text AI.