What is Multimodal AI?
Multimodal AI can understand and generate multiple types of content at once. Instead of being limited to text only or images only, multimodal AI combines different input and output types in a single interaction.
What is Multimodal AI?
Multimodal AI works with more than one type of data at the same time.
'Modal' means type of input or output. The main modalities are:
- Text
- Images
- Audio
- Video
- Documents
- Code
A unimodal AI only handles one type. For example, an older version of GPT only handled text.
A multimodal AI handles multiple types at once. For example, GPT-4o can accept text and images as input, and respond with text (or text describing actions it would take on the image).
The most capable modern AI models are multimodal: they can see images, hear audio, and respond in text, all in a single conversation.
Deep Learning ⊂ Machine Learning ⊂ Artificial Intelligence
Examples of Multimodal AI in Action
- You take a photo of a broken appliance and ask 'What is wrong and how do I fix it?' The AI analyzes the image and provides repair instructions
- You upload a page from a book in a foreign language and ask 'Translate this to English.' The AI reads the text from the image and translates it
- You share a spreadsheet screenshot and ask 'Which month had the highest sales?' The AI reads the data from the image and answers
- You speak a question out loud and the AI responds verbally, understanding your spoken words
- You share a diagram and ask 'Explain how this process works.' The AI analyzes the diagram and generates an explanation
Key Takeaways
- Multimodal AI can understand and generate multiple types of content at once.
- You take a photo of a broken appliance and ask 'What is wrong and how do I fix it?' The AI analyzes the image and provides repair instructions
- You upload a page from a book in a foreign language and ask 'Translate this to English.' The AI reads the text from the image and translates it
- You share a spreadsheet screenshot and ask 'Which month had the highest sales?' The AI reads the data from the image and answers