Chapter 09 · Sustainability & AISustainability & AI

Synthetic data

Meaning statusEstablishedSource recordDirect document linkedWhy these are different

Definition

Artificially generated data that mimics the statistical properties of real data, used to train or test AI models where real observations are scarce, expensive or sensitive.

References

UK ICOPrivacy-enhancing technologies guidance (incl. synthetic data)

This source provides part of the technical or institutional basis for the definition.

NISTAI RMF Generative AI Profile (NIST AI 600-1, 2024)

This source supports the explanation of how the term is applied, measured or governed in practice.

Overview

What it means

Synthetic data can come from simulators, statistical models or generative AI. It fills gaps — rare species sightings, extreme weather events, equipment failures — but its fidelity to reality must be validated, since models trained on flawed synthetic data inherit those flaws.

How it is used

Applications include augmenting rare-event training sets for ecological and climate models, stress-testing financial and grid models, and sharing privacy-preserving versions of personal energy-use data — a use case covered in UK ICO guidance on privacy-enhancing technologies.

Why it matters

Many sustainability problems are data-poor exactly where stakes are high. Synthetic data can unblock model development, but treating generated data as equivalent to observation is an analytical and governance risk that must be managed explicitly.

Have evidence, context, or a correction to share? Every suggestion is considered by an editor before publication.

Meaning status
Established
Verification date
Not recorded
Last updated
21 Aug 2026
What the classifications mean

Meaning status: Established

EstablishedCurrentMultiple definitionsContestedEmergingIndexed