The NextGen Harmonised Data Gateway
  • Home
  • Variable Database
    • Variable Database
    • Make it Green!
    • Make it Digital!
    • Make it Equal!
    • Make it Healthy!
    • Make it Strong!
  • Harmonisation
    • Harmonisation Toolbox
    • Make it Green!
    • Make it Digital!
    • Make it Equal!
    • Make it Healthy!
    • Make it Strong!
  • Data
  • Training
  • Contact
    • User Support
    • Team
  1. Harmonisation Toolbox
  • Harmonisation Toolbox
  • Make it Green!
  • Make it Digital!
  • Make it Equal!
  • Make it Healthy!
  • Make it Strong!

Harmonisation Toolbox

The following web pages present harmonisation scripts for measurement instruments for the topics: Make it Green!, Make it Digital!, Make it Healthy!, Make it Strong!, and Make it Equal!.

The scripts are based on a selection of variables taken from the Variable Database, which gives an overview of all variables on these topics with extensive coverage across international survey programmes, countries and time points.

The scripts are prepared in the open-source statistical software R (R Core Team, 2024) and embedded into detailed guidelines that explain every data preparation step on the way to the harmonised data set. During the data preparation process, we use structuring elements based on Kołczyńska (2022). Since users stay in complete control of the process when using the scripts, they can modify them according to their needs or apply them to other potential variable candidates for harmonisation taken from the Variable Database. At the end of the process, users receive a customised harmonised dataset.

The key business of data harmonisation is to make them comparable. In our case, that means 1. the judgement of whether the question texts measure the same thing and 2. harmonising the response scales. There is a huge body of literature on data harmonisation approaches originating from different scientific fields (Kolen & Brennan, 2014; Roth & Singh, 2024; Tomescu-Dubrow et al., 2024). However, a harmonised data set, which is a combination of different data sets across survey programmes, time, and countries, does not meet the requirements for the more eloquent methods. Nevertheless, by applying different harmonization methods to survey data, Heizmann (2025) demonstrates very impressively that substantive conclusions tend to stay the same regardless of the harmonization approach chosen. For these reasons, the harmonisation procedure we suggest in our scripts is based on the linear stretch approach (de Jonge et al., 2014), which is a rather rough method that assigns the source variable’s lowest response option to the lowest value of the target scale and the highest to the highest value and all intermediate options are given equally distanced numbers in between. In our scrips, we are choosing target scales that suit the data best. Consequently, harmonised target response scales may differ within and across topics.

For a walkthrough of the Harmonisation Toolbox, see the short video below.

Back to top

References

de Jonge, T., Veenhoven, R., & Arends, L. (2014). Homogenizing responses to different survey questions on the same topic: Proposal of a scale homogenization method using a reference distribution. Social Indicators Research, 117, 375–300. https://doi.org/10.1007/s11205-013-0335-6
Heizmann, B. (2025). Good enough? A comparison of different harmonization procedures and their substantive consequences using the example of life satisfaction. Quality & Quantity. https://doi.org/10.1007/s11135-025-02060-7
Kołczyńska, M. (2022). Combining multiple survey sources: A reproducible workflow and toolbox for survey data harmonization. Methodological Innovations, 15(1), 62–72. https://doi.org/10.1177/20597991221077923
Kolen, M. J., & Brennan, R. L. (2014). Test equating, scaling, and linking. Springer. https://doi.org/10.1007/978-1-4939-0317-7
R Core Team. (2024). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/
Roth, M., & Singh, R. K. (2024). Questionlink: Harmonizing single item survey questions on the same construct. https://matroth.github.io/questionlink/
Tomescu-Dubrow, I., Wolf, C., Slomczynski, K. M., & Jenkins, J. C. (2024). Survey data harmonization in the social sciences. John Wiley & Sons, Ltd. https://onlinelibrary.wiley.com/doi/abs/10.1002/9781119712206
Make it Green!

© EU Funded Project 101131118 — Infra4NextGen

 

Subject to Terms of Use | Privacy notice | Cookie policy | Built with R and Quarto