Skip to content
DopeSwagYolo

AI Models

What is multimodal AI?

Multimodal AI refers to artificial intelligence systems that can take in or produce more than one kind of data, such as text, images, audio and video. Such systems can relate these to one another instead of handling only one type.

Also known as: multimodal model, multimodal LLM, MLLM

Researched and fact-checked by AI, with no human review. 8 sources listed below. How we verify

Last updated

How it works

A modality is a way something is experienced, such as sight or sound. A 2017 survey by Tadas Baltrušaitis and two co-authors describes multimodal machine learning. It calls this the effort to build models that can handle information from several modalities and relate one to another. The glossary of a March 2025 report from the U.S. National Institute of Standards and Technology (NIST) defines multimodal models in similar terms.

One method learns from paired examples. A February 2021 paper by Alec Radford and colleagues at OpenAI described the CLIP model. It was trained on 400 million image-text pairs to predict which caption belongs to which image.

Later systems were built as single models, their makers say. Announcing Gemini on December 6, 2023, Google said multimodal models had usually been assembled from components trained separately for each modality. It said Gemini was instead pre-trained on several modalities from the outset. OpenAI's GPT-4o announcement of May 13, 2024, said its earlier voice mode chained three models. So the main model could not directly observe tone or background noise, it said. It said GPT-4o was a single model trained end to end across text, vision and audio.

Why it matters

The 2017 survey argues that because people experience the world through several senses, AI must interpret such signals together to understand it.

On security, the NIST report says redundant information across modalities does not necessarily make a model more robust. It says an attack on a single modality can compromise a multimodal model.

Where things stand in 2026

Stanford's 2026 AI Index reports results on MMMU, a test of college-level questions that combine text with diagrams, charts and tables. It says the leading model, Gemini 3.1 Pro Preview, scored 88.2% as of February 2026, within 0.4 percentage points of the best human expert reference. On ClockBench, which asks models to read analog clocks, the top model reached 50.6% in March 2026, against 90.1% for humans. The report sets that clock-reading result beside a gold-medal score at the 2025 International Mathematical Olympiad to illustrate what researchers call jagged intelligence.

As of October 2026, OpenAI's models page says all its latest models accept text and image input and produce text. It lists separate models for image generation, real-time speech, speech generation and transcription. Google's Gemini API model list also includes audio-to-audio, text-to-speech, image generation and video generation models.

Sources

Articles on AI Models