Multimodal AI
Multimodal AI refers to machine learning models capable of processing, understanding, and generating content across multiple data modalities—including text, images, video, audio, and code—rather than being limited to a single input or output type. These models bridge different types of information to achieve richer understanding than unimodal systems.