Harmonisation Toolbox
The following web pages present harmonisation scripts for measurement instruments for the topics: Make it Green!, Make it Digital!, Make it Healthy!, Make it Strong!, and Make it Equal!.
The scripts are based on a selection of variables taken from the Variable Database, which gives an overview of all variables on these topics with extensive coverage across international survey programmes, countries and time points.
The scripts are prepared in the open-source statistical software R (R Core Team, 2024) and embedded into detailed guidelines that explain every data preparation step on the way to the harmonised data set. During the data preparation process, we use structuring elements based on Kołczyńska (2022). Since users stay in complete control of the process when using the scripts, they can modify them according to their needs or apply them to other potential variable candidates for harmonisation taken from the Variable Database. At the end of the process, users receive a customised harmonised dataset.
The key business of data harmonisation is to make them comparable. In our case, that means 1. the judgement of whether the question texts measure the same thing and 2. harmonising the response scales. There is a huge body of literature on data harmonisation approaches originating from different scientific fields (Kolen & Brennan, 2014; Roth & Singh, 2024; Tomescu-Dubrow et al., 2024). However, a harmonised data set, which is a combination of different data sets across survey programmes, time, and countries, does not meet the requirements for the more eloquent methods. Nevertheless, by applying different harmonization methods to survey data, Heizmann (2025) demonstrates very impressively that substantive conclusions tend to stay the same regardless of the harmonization approach chosen. For these reasons, the harmonisation procedure we suggest in our scripts is based on the linear stretch approach (de Jonge et al., 2014), which is a rather rough method that assigns the source variable’s lowest response option to the lowest value of the target scale and the highest to the highest value and all intermediate options are given equally distanced numbers in between. In our scrips, we are choosing target scales that suit the data best. Consequently, harmonised target response scales may differ within and across topics.
For a walkthrough of the Harmonisation Toolbox, see the short video below.