Multimodal Systems

One model, many kinds of input

A multimodal model accepts more than text. Images, audio, documents and video are encoded into the same representation space as language, so a single model can read a chart, transcribe a meeting, describe a photograph or hold a spoken conversation without handing off to separate services.

How it works

What it unlocks

Practical notes

Further Reading