Skip to index

GLOSSARY

Multimodal

Models that work across multiple media types at once — reading images, generating audio, watching video — instead of text alone.

A multimodal model processes more than one kind of data: text plus images, audio, or video. You can show ChatGPT a photo and ask about it, dictate a prompt by voice, or feed a video to Gemini for analysis — all through the same interface.

Multimodality is quietly reshaping every category in this directory: assistants see your screenshots, video generators understand stills, and audio models clone voices from samples. It also unlocks accessibility use cases — describing images for blind users, transcribing speech — that text-only systems never could.

Related terms

Tools that use this

Related categories