Chapter 09 · Sustainability & AISustainability & AI

Multimodal AI

Meaning statusEstablishedSource recordDirect document linkedWhy these are different

Definition

AI models that process and relate multiple types of data — text, images, audio, video, sensor streams — within a single system.

References

ICML 2021 / arXivRadford et al. — Learning Transferable Visual Models From Natural Language Supervision

This source provides part of the technical or institutional basis for the definition.

ESAΦ-lab — AI for Earth observation innovation lab

This source supports the explanation of how the term is applied, measured or governed in practice.

Overview

What it means

Multimodal models learn shared representations across data types; OpenAI's CLIP (2021), which jointly embeds images and text, was a landmark. The approach allows a model to answer questions about an image, caption a chart, or cross-reference a report with satellite evidence.

How it is used

Sustainability applications exploit complementary evidence: combining field photographs with acoustic records and satellite scenes for biodiversity assessment, or cross-checking textual claims against imagery in supply-chain due diligence.

Why it matters

Environmental reality is multimodal — it is seen, heard, measured and described. Models that integrate these channels enable richer monitoring, but also make it harder to trace which evidence drove a conclusion.

Have evidence, context, or a correction to share? Every suggestion is considered by an editor before publication.

Meaning status
Established
Verification date
Not recorded
Last updated
21 Aug 2026
What the classifications mean

Meaning status: Established

EstablishedCurrentMultiple definitionsContestedEmergingIndexed