Skip to main content
Data in Brief logoLink to Data in Brief
. 2025 Dec 3;64:112330. doi: 10.1016/j.dib.2025.112330

A time-series dataset on sustainable development Goals compliance in Europe

Javier Fernández-Macho 1
PMCID: PMC12765241  PMID: 41492552

Abstract

The dataset compiles time series of Sustainable Development Goals (SDG) indicators for NUTS2 regions across EU Member States, candidate countries, and EFTA members from 1980 to 2024. To support multivariate analyses—e.g. with methods based on Data Envelopment Analysis (DEA), which requires complete data—extensive filtering and imputation were performed. Missing values were addressed through temporal and geographical methods, followed by either a conservative or a neutral final imputation to ensure dataset completeness. The processed dataset may also support broader applications, such as econometric modeling, policy evaluation, or index construction, which require a complete set of data with no missing values.

Keywords: Economic development, Environmental policy, Sustainability measurement, Sustainable Economy


Specifications Table

Subject Social Sciences
Specific subject area Time series of statistical indicators for SDGs in Europe at three levels of geographical distribution: regions (NUTS2), supra-regions (NUTS1), and countries (NUTS0).
Type of data Tables (zipped .csv format), Raw, Filtered, Processed.
Data collection To get from the actual Eurostat sources to our raw dataset, some rearrangement of multidimensional code names and cleaning of alphanumerical flags was needed. Furthermore, when several units of measurement are available, we selected relative units, typically existing percentages or rates. Exceptionally, we had to transform some variables from their Eurostat unit—million PPS, with no other direct relative value present—to percentage of the EU27 average obtained from the same source.
Data source location Eurostat; European Union Member States, EFTA (Iceland, Norway, Switzerland), EU candidates (Albania, Montenegro, North Macedonia, Serbia, Turkey).
Data accessibility Repository name: Mendeley Data
Data identification number: 10.17632/fsjv4vk9gv.3
Direct URL to data: https://data.mendeley.com/datasets/fsjv4vk9gv/3
Related research article

1. Value of the Data

  • The dataset offers a response to a growing need for more granular, long-term analysis of SDG compliance at the regional level.

  • Their geographical distribution provides an appropriate scale for capturing regional dynamics, enabling a deeper understanding of their development trends and imbalances.

  • The processed dataset may support broader applications, such as econometric modeling, policy evaluation, or index construction, which require a complete dataset with no missing data.

  • The datasets are intended to serve as a foundation for future research into regional SDG performance and sustainable development policy evaluation.

  • Possible scenarios include constructing composite regional indices of sustainable development, alignment policy, or training SDG compliance forecasting models. Examples of research usage are DEA-based efficiency analysis, regional sustainable development comparisons, policy evaluation (EU / national / regional), visualization and mapping of SDG progress, econometric models and descriptive statistics, and machine learning applications (e.g., clustering or prediction of regional trends).

  • The final dataset can serve policy analysts and regional planners who could apply it to monitor and compare SDG progress across territories in order to support evidence-based policy design. In this respect, international organizations may leverage it for comparative studies of sustainable development in Europe and neighboring countries. Besides, academic researchers in economics, geography, and environmental sciences may use it for DEA-based efficiency analysis, econometric modeling, or machine learning approaches to regional sustainability. Finally, civil society groups, NGOs, and journalists can draw on the dataset for descriptive statistics, mapping, and visualization of inequalities and trends in regional sustainable development.

2. Background

While existing studies on SDGs have largely concentrated on global impacts and aggregate patterns, there is a growing need for more granular, long-term analysis at the regional level. National averages often mask substantial spatial disparities in how regions experience and respond to socio-economic and environmental challenges. In the European context, NUTS2 regions provide an appropriate scale for capturing these dynamics, enabling a deeper understanding of uneven development trajectories and regional imbalances.

However, accessing comprehensive time-series data on SDG indicators at such level of geographical disaggregation remains a challenge due to the incompleteness of data in many regions. To address this, we present a structured and accessible dataset that incorporates all the processing details necessary to estimate the missing values. This is important in cases that use a methodology incompatible with missing values, e.g. the DEA1-based impact assessment and forecasting of SDG compliance in [1]. This Data-in-Brief article not only facilitates access to the dataset but also offers a transparent account of its construction. Beyond supporting the analysis in those situations where a complete set of data is needed, these resources are intended to serve as a foundation for future research into regional SDG performance and sustainable development policy evaluation.

3. Data Description

The raw dataset collects time series of statistical indicators available for the SDGs in Europe at three NUTS2 levels of geographical distribution: regions (NUTS2), supra-regions (NUTS1), and countries (NUTS0). The latter two levels are included to supplement missing information at the primary regional level, enabling the dataset to be used in contexts where a complete set of data, free of missing values, is required.

In summary, the raw dataset is made up of a multivariate set of 19 SDG time series, obtained from Eurostat as the primary data source [2], for each of the 639 NUTS geographical units—434 at level 2, 162 at level 1, and 43 at the country level—across EU Member States, candidate countries, and EFTA3 countries during the period 1980–2024. Table 1 gives a detailed description of the indicators involved.

Table 1.

SDG indicators.

Indicator Description Eurostat dataset Unit
SDG 1: No poverty
sdg_01_10 Persons at risk of poverty or social exclusion ilc_peps11n %
sdg_01_20 Persons at risk of monetary poverty after social transfers ilc_li41 %
sdg_01_31 Severe material and social deprivation rate ilc_mdsd18 %
sdg_01_40 Persons living in households with very low work intensity ilc_lvhl21n %
SDG 3: Good health and well-being
sdg_03_60 Self-reported unmet need for medical examination and care hlth_silc_08b_r %
SDG 4: Quality education
sdg_04_10 Early leavers from education and training edat_lfse_16 %
sdg_04_20 Tertiary educational attainment edat_lfse_04 %
sdg_04_31 Participation in early childhood education educ_uoe_enra22 %
sdg_04_60 Adult participation in learning trng_lfse_04 %
SDG 5: Gender equality
sdg_05_30 Gender employment gap tepsr_lm220 %
SDG 8: Decent work and economic growth
sdg_08_20 Young people NEET (neither in employment nor in education and training) edat_lfse_22 %
sdg_08_30 Employment rate lfst_r_lfe2emprt %
sdg_08_40 Long-term unemployment rate lfst_r_lfu2ltu % active population
SDG 9: Industry, innovation and infrastructure
sdg_09_10 Gross domestic expenditure on R&D rd_e_gerdreg Million PPS (2005 prices) as % EU27–2020 average
sdg_09_30 R&D personnel rd_p_persreg % active population in full-time equivalent
SDG 10: Reduced inequalities
sdg_10_10 Purchasing power adjusted GDP per capita nama_10r_2gdp PPS per inhabitant as % EU27–2020 average
SDG 11: Sustainable cities and communities
sdg_11_40 Road traffic deaths tran_r_acci per million inhabitants
SDG 15: Life on land
sdg_15_50 Area at risk of severe soil erosion by water aei_pr_soiler %
SDG 16: Peace, justice and strong institutions
sdg_16_10 Deaths due to homicide hlth_cd_asdr2 standardized rate

With this information, a raw dataset, SDGTS_DB_NUTS_raw, was prepared covering the period 1980–2024. Also, two SDG datasets, SDGTS_DB_NUTS_nona and SDGTS_DB_NUTS_nona_alt, with the complete multivariate time-series sets for the period 2019–2024 were constructed following the steps described in the “Experimental design, materials and methods” section using two different final imputation methods. The first uses a conservative imputation method, tailored for DEA analysis, while the second uses a more neutral imputation method, more suitable for other analytical purposes, as described below in more detail.

4. Experimental Design, Materials and Methods

To get from the actual Eurostat sources to our raw dataset, some rearrangement of multi- dimensional code names and cleaning of alphanumerical flags was needed [3]. Furthermore, when several units of measurement are available, we selected relative units, typically existing percentages or rates. Exceptionally, we had to transform some indicator from its Eurostat unit, with no other direct relative value present, to percentage of the EU27 average obtained from the same source.

With a view to a subsequent analysis that necessitates a complete dataset, some data processing was required to supplement missing information.

While the primary motivation for supplementing missing data was to enable studies such as those using a DEA-based analysis, which cannot typically accommodate missing values, complete time series also enhance the overall utility of the dataset. For example, they may allow for robust statistical analysis and econometric modeling (e.g. panel regression or time-series forecasting), improve the performance of machine learning algorithms (e.g. for classification or clustering), and support consistent comparisons of regional trends with no time disruptions. Moreover, complete data are essential for important tasks in policy monitoring and evaluation, such as for conducting simulation or scenario analysis, and for constructing composite indicators such as the SDG index in [1]. In this sense, the ensuing dataset can be broadly applicable across a wide range of analytical and policy-oriented uses. As a result of this data processing, 15 complete time series of SDG indicators were obtained for 280 NUTS2 regions in Europe for the period 2019–2024. What follows is a detailed description of the steps taken.

step 1.- Selection criteria and coverage. SDG indicators were selected based on the availability of data at the NUTS2 level for the period from 2019 onward. Two indicators were excluded due to insufficient regional coverage: sdg_03_60 (self-reported unmet need for medical examination and care) was excluded as it lacked data for most regions prior to 2021, and sdg_15_50 (area at risk of severe soil erosion by water) was excluded due to the absence of data beyond 2018. After this first step, the dataset initially included NUTS2 regions from 37 countries, including all EU Member States plus members of the EFTA (Switzerland, Iceland, Liechtenstein, and Norway), official candidates for EU membership (Albania, Montenegro, North Macedonia, Serbia, and Turkey) and United Kingdom.

step 2.- Selection of DEA input and output indicators. The selection of input and output indicators for implementing a DEA method was based on data availability and coverage criteria. DEA inputs were required to have data available for the year 2019, and all indicators satisfied this criterion. DEA outputs were selected among those indicators that had data available for at least 70 % of NUTS2 regions in at least one year since 2019. Based on this threshold, two indicators were excluded from the output list due to insufficient regional coverage: sdg_04_31 (Participation in early childhood education), and sdg_09_30 (R&D personnel). It is important to note, however, that this distinction is pertinent only to the specific objectives of a DEA-based analysis and can be ignored for most other potential uses of the dataset.

step 3.- Missing values: temporal and geographical imputation. Following the initial data processing step, a total of approximately 34 % missing values were identified across NUTS2 code-year-indicator combinations, necessitating further filtering and imputation. As a first measure, Extra-Regio and overseas territories were excluded from the dataset. This reduced the number of missing values to approximately 33 %.

Temporal imputation. To address missing values over time, a simple carry-forward method was applied: any missing value was replaced by the corresponding value from the previous year. This procedure reduced the number of missing entries to approximately 8 %.

Geographical imputation (NUTS1 level). Subsequently, missing values were imputed using data from the corresponding NUTS1 region, where available. Despite these imputation efforts, 49 NUTS2 regions—representing nearly 15 % of all regions—still exhibited >40 % missing data. These regions were deemed insufficiently represented and were therefore excluded from further analysis. The excluded regions included Åland (FI20), Svalbard (NO0B), and all of Albania, Liechtenstein, Montenegro, Iceland, and the UK, the latter because Eurostat stopped updating those territorial units since Brexit. After this, the number of missing values decreased to approximately 4 %.

Geographical imputation (NUTS0 level). A final round of imputation was conducted using national-level (NUTS0) data to fill remaining gaps. This step further reduced the number of missing values to approximately 3 % of the original dataset. step 4.- Final imputation. As previously noted, several intended applications of the dataset—particularly those involving multivariate analysis—require a complete set of time series without missing values. For example, this is especially critical for the DEA methodology, which cannot accommodate missing data.

Conservative imputation. An initial option considered was the outright removal of all code–year combinations containing missing values. However, this approach was deemed excessively severe, as it would result in a substantial loss of information [4,5]. For similar reasons, we also rejected extreme imputation strategies—such as substituting missing output values with zeros or imputing missing inputs with arbitrarily large values—as these could unduly distort the analysis [6,4].

Instead, we adopted a more conservative imputation strategy: remaining missing values in input indicators were replaced with the maximum observed value for the respective indicator, while missing values in output indicators were replaced with the minimum observed value [4,5]. This approach was designed to avoid artificially inflating efficiency scores while preserving the completeness of the dataset.

Neutral imputation. It should be noted, however, that the final imputation strategy described above is specifically aligned with the objectives of a DEA-based method, and may not be necessarily suitable for all potential uses of the dataset [see 5]. For alternative applications—such as descriptive statistics, econometric models, visualization, or even machine learning—a more neutral strategy replacing missing values with the indicator-wise mean may be more appropriate, which resulted in an alternative version of the dataset.

Table 2 gives a brief account of aggregated figures from the imputation steps described in this section.

Table 2.

Imputation summary table.

2019–2024 Total observations Total regions Missing obs. % missing
Steps 1, 2 Original data 27,456 356 9220 33.6 %
Step 3 excluding Extra-Regio's and Overseas 26,952 329 8976 33.3 %
temporal imputation 26,952 329 2184 8.1 %
geo imputation from NUTS1 24,210 280 1153 4.1 %
geo imputation from NUTS0 25,200 280 799 2.8 %
Step 4 Final imputation(*) 25,200 280 0 0 %

(*) For comparison purposes, only output indicators were considered in the conservative imputation (SDGTS_DB_NUTS_nona).

Table 3 compares the two final versions, with the conservative (SDGTS_DB_NUTS_nona) and neutral imputations (SDGTS_DB_NUTS_nona_alt) mentioned above. As mentioned, the first is tailored for DEA analysis, while the second is better suited for other analytical tasks such as econometric analysis.

Table 3.

Summary statistics for the two final versions.

SDGTS_DB_NUTS_nona
SDGTS_DB_NUTS_nona_alt
indicator mean median stdev mean median stdev
edat_ lfse_04 32.1 31.6 10.6 32.1 31.6 10.6
edat_ lfse_16 10.7 9.0 6.7 10.8 9.1 6.6
edat_lfse_22 13.5 11.0 7.5 13.5 11.0 7.5
hlth_cd_ asdr2 0.7 0.6 0.5 0.7 0.6 0.5
ilc_li41 16.4 14.9 6.6 16.7 15.3 6.2
ilc_ lvhl21n 8.1 8.0 4.1 8.3 8.3 3.9
ilc_ mdsd18 7.0 5.2 6.1 7.2 5.5 6.0
ilc_ peps11n 21.6 19.8 8.4 22.0 20.2 8.0
lfst_r_ lfe2emprt 72.6 75.2 9.5 72.6 75.2 9.5
lfst_r_ lfu2ltu 2.7 1.8 2.8 2.7 1.8 2.8
nama_10r_ 2gdp 92.0 86.0 38.2 92.2 86.0 38.0
rd_e_ gerdreg 1.1 0.2 2.9 1.2 0.2 2.9
tepsr_lm220 13.1 9.6 10.0 13.1 9.6 10.0
tran_r_ acci 51.7 47.0 25.6 52.1 48.0 25.3
trng_ lfse_04 11.3 9.2 7.8 11.3 9.2 7.8

After the final imputation step, the SDG dataset was rendered fully complete, with no remaining missing values. Therefore, it can be used in a broad selection of applications where missing values must be avoided. For example, this procedure enabled us to implement the DEA-based method as described in [1], resulting in the computation and forecasting of SDG index scores as presented in Table A.2 in the repository (see DATA AVAILABILITY below), which may serve as illustration for future research into regional SDG performance and sustainable development policy evaluation.

Fig. 1 includes a flowchart with the steps followed from the original Eurostat data to the final full array.

Fig. 1.

Fig. 1:

Flowchart of the imputation steps.

Limitations

The dataset covers European countries, including EU Member States, candidate countries, and EFTA members; however, some NUTS2 territories were excluded due to insufficient data availability. Missing values were addressed through temporal and geographical imputation procedures, followed by a final imputation stage to ensure complete time series. While this approach preserves analytical usability, it may introduce uncertainty in trend estimation where imputation was extensive, but the procedure is justified and documented.

Ethics Statement

The author has read and follow the ethical requirements for publication in Data in Brief and confirms that the current work does not involve human subjects, animal experiments, or any data collected from social media platforms.

Data Availability

The 1980–2024 raw dataset with missing values present, and the two 2019–2024 SDG datasets with the complete multivariate time-series sets, i.e. with no missing values after the imputation process described above (SDGTS_DB_NUTS_raw, SDGTS_DB_NUTS_nona and SDGTS_DB_NUTS_nona_alt zipped csv files) can be accessed from the Mendeley Data repository here [7]. Additionally, Table A.2 of SDG index scores can be obtained from the same source.

Credit Author Statement

J.F.M.: Conceptualization, Methodology, Software, Writing – original draft, Writing – review & editing, Data curation, Visualization, Project administration, Funding acquisition.

Acknowledgments

This work benefited from work previously supported by the EU Interreg Atlantic Area Programme 2014–2020 and European Regional Development Fund (ERDF) under Grant EAPA 224/2016 MOSES. In addition, financial support from UPV/EHU Econometrics Research Group (Basque Government grants IT1359–19, IT1508–22) and Spanish Ministry of Science and Innovation (grant PID2020–112951GB-I0) is also acknowledged.

Acknowledgments

Declaration of Competing Interest

The author declares that he has no known competing financial interests or personal relationship that could have appeared to influence the work reported in this paper.

Footnotes

1

Data Envelopment Analysis.

2

Eurostat’s Nomenclature of Territorial Units for Statistics.

3

European Free Trade Association.

Data Availability

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The 1980–2024 raw dataset with missing values present, and the two 2019–2024 SDG datasets with the complete multivariate time-series sets, i.e. with no missing values after the imputation process described above (SDGTS_DB_NUTS_raw, SDGTS_DB_NUTS_nona and SDGTS_DB_NUTS_nona_alt zipped csv files) can be accessed from the Mendeley Data repository here [7]. Additionally, Table A.2 of SDG index scores can be obtained from the same source.


Articles from Data in Brief are provided here courtesy of Elsevier

RESOURCES