Models & Architecture
Multimodal AI
AI systems capable of processing and generating multiple types of data simultaneously, such as text, images, audio, and video, often within a single model.
Multimodal AI refers to models that can understand and generate across different data modalities. Unlike single-modality models (text-only or image-only), multimodal models can reason about information from multiple sources.
Prominent multimodal models:
- GPT-4V/GPT-4o: Text, image, and audio understanding
- Gemini: Text, image, audio, and video
- Claude 3: Text and image understanding
- Meta LLaMA 3.2: Text and image understanding
Multimodal capabilities enable applications like image captioning, visual question answering, document understanding, and video analysis. The integration of multiple modalities in a single model is considered a major step toward more general and capable AI systems.
Related Terms
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.