Chapter 08 · Finance, data & evidenceSustainability Language

Representativeness

Meaning statusEstablishedSource recordDirect document linkedWhy these are different

Definition

The extent to which evidence reflects the population, places, periods and conditions about which a conclusion is being made.

References

Overview

“A sample can be large, precise and still describe the wrong people. ”

Representativeness is the bridge between what was observed and what is claimed. A survey may contain thousands of records, a satellite model may cover millions of hectares and a programme database may list every registered participant. None of that establishes that the evidence reflects the wider population to which the conclusion is applied. The concept begins with a target population.

Who or what is the evidence intended to represent? All coffee farmers in a country, members of participating cooperatives, farms within mapped districts, workers present during an audit, or households reachable by mobile telephone? Each is a different population. A conclusion becomes unreliable when the population described in the claim is broader than the population that had a realistic chance of entering the data.

The 1936 Literary Digest poll remains a useful warning.

It received more than two million responses and confidently predicted that Alf Landon would defeat Franklin D. Roosevelt in the United States presidential election. Roosevelt won decisively. The problem was not sample size. The magazine's lists overrepresented wealthier citizens, and those who chose to respond differed from those who did not. A very large sample amplified confidence without repairing selection.

Sustainability systems reproduce the same problem in quieter ways. Farmers registered with a cooperative are easier to locate than isolated producers. Households with smartphones are easier to survey than households without connectivity. Farms beside roads are easier to visit than farms several hours away. Workers on permanent contracts are easier to interview than seasonal or migrant labour.

The people most exposed to risk may be precisely those least likely to appear in the system. Representativeness has several dimensions. Geographic coverage matters where environmental and livelihood conditions vary across landscapes. Seasonal coverage matters where labour, income or pesticide use changes during harvest. Demographic coverage matters where gender, age, migration status or land tenure shapes experience.

Enterprise coverage matters where small suppliers operate differently from large ones. A sample may be representative in one dimension and distorted in another. Probability sampling provides a defensible basis for generalisation because units have known or calculable chances of selection. It does not guarantee a perfect result.

Coverage gaps, non-response, inaccurate frames and measurement error can still distort estimates.

Non-probability samples can also be useful for rapid learning, qualitative insight or finding rare cases, but the claim must remain proportionate to the design. Convenience data do not become population evidence because they are stored in a dashboard. Weighting can reduce known imbalances when reliable information about the population exists.

It cannot recreate groups that were never observed or correct characteristics that were not measured. A dataset of accessible farms cannot be made fully representative of inaccessible farms by multiplying records from the accessible group. Statistical adjustment is a tool, not a substitute for field design. Time also changes representativeness.

A farmer registry assembled three years ago may no longer reflect current production. A baseline collected before a drought may not describe the population after migration, crop failure or conflict. Repeated measurement must account for who leaves, who enters and whose data disappear.

Attrition is not only a technical issue; it can remove the households experiencing the greatest difficulty. The practical discipline is to make the inference visible. State the target population, sampling frame, selection process, response rate, exclusions and known differences between participants and non-participants. Disaggregate results where averages conceal important groups.

Test whether conclusions change under alternative weights or assumptions. Where representativeness cannot be established, narrow the wording: “among surveyed participants” is more credible than “farmers in the region. ” Representativeness is not achieved once. It is maintained through updated frames, deliberate inclusion, field quality control and honest limits on generalisation.

The objective is not to make every dataset universal.

It is to ensure that the reach of the claim never exceeds the reach of the evidence.

Practical application

Define the target population before designing collection. Build or evaluate the sampling frame, identify groups with weak coverage and choose a selection method suited to the intended inference. Track contact, eligibility, refusal and non-response separately, and compare the achieved sample with known population characteristics. Report the boundaries of the evidence alongside the result.

Use weighting and sensitivity analysis where justified, but document assumptions. Commission supplementary qualitative or purposive work for groups likely to remain invisible, without presenting those findings as statistically representative.

Why it matters

Sustainability decisions distribute resources, scrutiny and opportunity. Unrepresentative evidence can direct support towards those already visible while overlooking remote producers, informal workers or vulnerable communities. It can also create false confidence in averages that do not describe the people most affected.

Common misconception

Representativeness is often confused with sample size. A larger sample reduces random sampling error under an appropriate design; it does not automatically correct selection, coverage or non-response bias. A small, well-designed probability sample can support stronger inference than millions of self-selected records.

Connections

Sampling determines how observations enter a dataset. Representativeness asks whether those observations support the intended generalisation. Bias explains systematic distortion within the design or measurement process. Data Quality addresses the wider fitness of information for use, while Inclusion asks whose experience is missing and why.

A question worth asking

Who had little or no chance of appearing in this evidence, and how would the conclusion change if their experience were visible?

Selected references

United Nations Statistics Division. 2005. Household Sample Surveys in Developing and Transition Countries: Design, Implementation and Analysis. Kish, L. 1965. Survey Sampling. Groves, R. M. et al. 2009. Survey Methodology, Second Edition. Bethlehem, J. 2010. Selection Bias in Web Surveys. International Statistical Review 78(2): 161-188. ISO 20252:2019.

Market, Opinion and Social Research, Including Insights and Data Analytics - Vocabulary and Service Requirements.

How it is used

The term appears in legislation, policies, governance systems, contracts, oversight and compliance decisions, where policymakers, regulators, legal teams, boards and organisations use it to classify, assess or communicate the extent to which evidence reflects the population, places, periods and conditions about which a conclusion is being made.

Its correct use depends on the applicable jurisdiction, legal or policy text, effective date, scope and responsible actor.

Have evidence, context, or a correction to share? Every suggestion is considered by an editor before publication.

Meaning status
Established
Last verification recorded
22 Aug 2026
Last updated
22 Aug 2026
What the classifications mean

Meaning status: Established

EstablishedCurrentMultiple definitionsContestedEmergingIndexed