
Dutch Data Prize 2024 win: Making population and family data FAIR for future research
The Generations and Gender Programme (GGP) is an international research infrastructure that provides researchers and policymakers with data on population and family dynamics. Its commitment to making these data Findable, Accessible, Interoperable and Reusable (FAIR) was recognised with the Dutch Data Prize 2024. Olga Grunwald, postdoctoral researcher at the Netherlands Interdisciplinary Demographic Institute (NIDI-KNAW), explains how GGP works to make its data FAIR, why this matters for research and what winning the Dutch Data Prize means to the programme.
Founded in 2000 and headquartered at NIDI-KNAW since 2009, GGP provides freely accessible, cross-nationally comparable data on demographic change. Its datasets include the Generations and Gender Survey (GGS) and the Fertility and Family Surveys (FFS), covering more than 300,000 individuals in over 30 countries.
The GGS, GGP’s main survey, follows the life courses of people aged 18 to 79 in different countries. The data enable researchers to study questions around topics such as fertility, the financial and social circumstances of young adults, family formation, and relationships between generations.
Making data accessible and reusable
GGP applies the FAIR principles to both its microdata and metadata. Researchers can access the microdata through the GGP Data Portal after registration. The data are available in widely used formats including SPSS, Stata and CSV. Because the datasets contain sensitive information, GGP follows the principle of ‘as open as possible, as closed as necessary’, balancing accessibility with appropriate privacy safeguards.
The metadata are documented using the DDI-Lifecycle 3 standard and are openly available through GGP’s Colectica portal in JSON and DDI-XML formats. GGP also uses controlled vocabularies from CESSDA to improve consistency and interoperability across datasets. Further work includes developing provenance metadata, assigning DOIs to microdata and working towards CoreTrustSeal certification.
Much of this work is carried out by colleagues at the GGP Central Hub at NIDI, in collaboration with colleagues at the French Institute for Demographic Studies (INED).
Keeping data understandable over time
For GGP, FAIR data are particularly important because datasets can remain valuable for decades. Its experience with the Fertility and Family Surveys from the 1990s illustrates this. These surveys remain an important research resource, but because they predate the FAIR principles, some decisions made during their creation are difficult to reconstruct.
This experience has demonstrated the importance of detailed and structured documentation. Applying the FAIR principles helps ensure that future researchers can understand the context, methodology and structure of data, even many years after they were collected. It also supports transparency, trust and effective reuse.
Winning the Dutch Data Prize is therefore an important recognition of the work GGP has undertaken to improve its data practices. ‘Over the past two years, we’ve invested significant time and energy into understanding what FAIR principles truly mean and how we can apply them to GGP and our data,’ says Grunwald. ‘Receiving this recognition reassures us that we’re on the right track.’
Building on the Dutch Data Prize
GGP is considering several ways to use the prize money to further improve its data. One possibility is a hackathon bringing together researchers and experts to improve interoperability between datasets in the GGP Data Portal. Another is to make older datasets, including the FFS, more FAIR by improving their metadata, documentation and accessibility.
Grunwald also encourages other researchers to nominate their datasets for the Dutch Data Prize. Preparing a nomination provides an opportunity to reflect on existing data practices and consider how they align with the FAIR principles.
‘Data management efforts often go unnoticed. The Dutch Data Prize is an opportunity to highlight the value of FAIR data and the contribution that well-managed and reusable data make to the wider research community,’ says Grunwald.