Abstract
The dataset compiles time series of Sustainable Development Goals (SDG) indicators for NUTS2 regions across EU Member States, candidate countries, and EFTA members from 1980 to 2024. To support multivariate analyses—e.g. with methods based on Data Envelopment Analysis (DEA), which requires complete data—extensive filtering and imputation were performed. Missing values were addressed through temporal and geographical methods, followed by either a conservative or a neutral final imputation to ensure dataset completeness. The processed dataset may also support broader applications, such as econometric modeling, policy evaluation, or index construction, which require a complete set of data with no missing values.
Keywords: Economic development, Environmental policy, Sustainability measurement, Sustainable Economy
Specifications Table
| Subject | Social Sciences |
| Specific subject area | Time series of statistical indicators for SDGs in Europe at three levels of geographical distribution: regions (NUTS2), supra-regions (NUTS1), and countries (NUTS0). |
| Type of data | Tables (zipped .csv format), Raw, Filtered, Processed. |
| Data collection | To get from the actual Eurostat sources to our raw dataset, some rearrangement of multidimensional code names and cleaning of alphanumerical flags was needed. Furthermore, when several units of measurement are available, we selected relative units, typically existing percentages or rates. Exceptionally, we had to transform some variables from their Eurostat unit—million PPS, with no other direct relative value present—to percentage of the EU27 average obtained from the same source. |
| Data source location | Eurostat; European Union Member States, EFTA (Iceland, Norway, Switzerland), EU candidates (Albania, Montenegro, North Macedonia, Serbia, Turkey). |
| Data accessibility | Repository name: Mendeley Data Data identification number: 10.17632/fsjv4vk9gv.3 Direct URL to data: https://data.mendeley.com/datasets/fsjv4vk9gv/3 |
| Related research article |
1. Value of the Data
-
•
The dataset offers a response to a growing need for more granular, long-term analysis of SDG compliance at the regional level.
-
•
Their geographical distribution provides an appropriate scale for capturing regional dynamics, enabling a deeper understanding of their development trends and imbalances.
-
•
The processed dataset may support broader applications, such as econometric modeling, policy evaluation, or index construction, which require a complete dataset with no missing data.
-
•
The datasets are intended to serve as a foundation for future research into regional SDG performance and sustainable development policy evaluation.
-
•
Possible scenarios include constructing composite regional indices of sustainable development, alignment policy, or training SDG compliance forecasting models. Examples of research usage are DEA-based efficiency analysis, regional sustainable development comparisons, policy evaluation (EU / national / regional), visualization and mapping of SDG progress, econometric models and descriptive statistics, and machine learning applications (e.g., clustering or prediction of regional trends).
-
•
The final dataset can serve policy analysts and regional planners who could apply it to monitor and compare SDG progress across territories in order to support evidence-based policy design. In this respect, international organizations may leverage it for comparative studies of sustainable development in Europe and neighboring countries. Besides, academic researchers in economics, geography, and environmental sciences may use it for DEA-based efficiency analysis, econometric modeling, or machine learning approaches to regional sustainability. Finally, civil society groups, NGOs, and journalists can draw on the dataset for descriptive statistics, mapping, and visualization of inequalities and trends in regional sustainable development.
2. Background
While existing studies on SDGs have largely concentrated on global impacts and aggregate patterns, there is a growing need for more granular, long-term analysis at the regional level. National averages often mask substantial spatial disparities in how regions experience and respond to socio-economic and environmental challenges. In the European context, NUTS2 regions provide an appropriate scale for capturing these dynamics, enabling a deeper understanding of uneven development trajectories and regional imbalances.
However, accessing comprehensive time-series data on SDG indicators at such level of geographical disaggregation remains a challenge due to the incompleteness of data in many regions. To address this, we present a structured and accessible dataset that incorporates all the processing details necessary to estimate the missing values. This is important in cases that use a methodology incompatible with missing values, e.g. the DEA1-based impact assessment and forecasting of SDG compliance in [1]. This Data-in-Brief article not only facilitates access to the dataset but also offers a transparent account of its construction. Beyond supporting the analysis in those situations where a complete set of data is needed, these resources are intended to serve as a foundation for future research into regional SDG performance and sustainable development policy evaluation.
3. Data Description
The raw dataset collects time series of statistical indicators available for the SDGs in Europe at three NUTS2 levels of geographical distribution: regions (NUTS2), supra-regions (NUTS1), and countries (NUTS0). The latter two levels are included to supplement missing information at the primary regional level, enabling the dataset to be used in contexts where a complete set of data, free of missing values, is required.
In summary, the raw dataset is made up of a multivariate set of 19 SDG time series, obtained from Eurostat as the primary data source [2], for each of the 639 NUTS geographical units—434 at level 2, 162 at level 1, and 43 at the country level—across EU Member States, candidate countries, and EFTA3 countries during the period 1980–2024. Table 1 gives a detailed description of the indicators involved.
Table 1.
SDG indicators.
| Indicator | Description | Eurostat dataset | Unit |
|---|---|---|---|
| SDG 1: No poverty | |||
| sdg_01_10 | Persons at risk of poverty or social exclusion | ilc_peps11n | % |
| sdg_01_20 | Persons at risk of monetary poverty after social transfers | ilc_li41 | % |
| sdg_01_31 | Severe material and social deprivation rate | ilc_mdsd18 | % |
| sdg_01_40 | Persons living in households with very low work intensity | ilc_lvhl21n | % |
| SDG 3: Good health and well-being | |||
| sdg_03_60 | Self-reported unmet need for medical examination and care | hlth_silc_08b_r | % |
| SDG 4: Quality education | |||
| sdg_04_10 | Early leavers from education and training | edat_lfse_16 | % |
| sdg_04_20 | Tertiary educational attainment | edat_lfse_04 | % |
| sdg_04_31 | Participation in early childhood education | educ_uoe_enra22 | % |
| sdg_04_60 | Adult participation in learning | trng_lfse_04 | % |
| SDG 5: Gender equality | |||
| sdg_05_30 | Gender employment gap | tepsr_lm220 | % |
| SDG 8: Decent work and economic growth | |||
| sdg_08_20 | Young people NEET (neither in employment nor in education and training) | edat_lfse_22 | % |
| sdg_08_30 | Employment rate | lfst_r_lfe2emprt | % |
| sdg_08_40 | Long-term unemployment rate | lfst_r_lfu2ltu | % active population |
| SDG 9: Industry, innovation and infrastructure | |||
| sdg_09_10 | Gross domestic expenditure on R&D | rd_e_gerdreg | Million PPS (2005 prices) as % EU27–2020 average |
| sdg_09_30 | R&D personnel | rd_p_persreg | % active population in full-time equivalent |
| SDG 10: Reduced inequalities | |||
| sdg_10_10 | Purchasing power adjusted GDP per capita | nama_10r_2gdp | PPS per inhabitant as % EU27–2020 average |
| SDG 11: Sustainable cities and communities | |||
| sdg_11_40 | Road traffic deaths | tran_r_acci | per million inhabitants |
| SDG 15: Life on land | |||
| sdg_15_50 | Area at risk of severe soil erosion by water | aei_pr_soiler | % |
| SDG 16: Peace, justice and strong institutions | |||
| sdg_16_10 | Deaths due to homicide | hlth_cd_asdr2 | standardized rate |
With this information, a raw dataset, SDGTS_DB_NUTS_raw, was prepared covering the period 1980–2024. Also, two SDG datasets, SDGTS_DB_NUTS_nona and SDGTS_DB_NUTS_nona_alt, with the complete multivariate time-series sets for the period 2019–2024 were constructed following the steps described in the “Experimental design, materials and methods” section using two different final imputation methods. The first uses a conservative imputation method, tailored for DEA analysis, while the second uses a more neutral imputation method, more suitable for other analytical purposes, as described below in more detail.
4. Experimental Design, Materials and Methods
To get from the actual Eurostat sources to our raw dataset, some rearrangement of multi- dimensional code names and cleaning of alphanumerical flags was needed [3]. Furthermore, when several units of measurement are available, we selected relative units, typically existing percentages or rates. Exceptionally, we had to transform some indicator from its Eurostat unit, with no other direct relative value present, to percentage of the EU27 average obtained from the same source.
With a view to a subsequent analysis that necessitates a complete dataset, some data processing was required to supplement missing information.
While the primary motivation for supplementing missing data was to enable studies such as those using a DEA-based analysis, which cannot typically accommodate missing values, complete time series also enhance the overall utility of the dataset. For example, they may allow for robust statistical analysis and econometric modeling (e.g. panel regression or time-series forecasting), improve the performance of machine learning algorithms (e.g. for classification or clustering), and support consistent comparisons of regional trends with no time disruptions. Moreover, complete data are essential for important tasks in policy monitoring and evaluation, such as for conducting simulation or scenario analysis, and for constructing composite indicators such as the SDG index in [1]. In this sense, the ensuing dataset can be broadly applicable across a wide range of analytical and policy-oriented uses. As a result of this data processing, 15 complete time series of SDG indicators were obtained for 280 NUTS2 regions in Europe for the period 2019–2024. What follows is a detailed description of the steps taken.
step 1.- Selection criteria and coverage. SDG indicators were selected based on the availability of data at the NUTS2 level for the period from 2019 onward. Two indicators were excluded due to insufficient regional coverage: sdg_03_60 (self-reported unmet need for medical examination and care) was excluded as it lacked data for most regions prior to 2021, and sdg_15_50 (area at risk of severe soil erosion by water) was excluded due to the absence of data beyond 2018. After this first step, the dataset initially included NUTS2 regions from 37 countries, including all EU Member States plus members of the EFTA (Switzerland, Iceland, Liechtenstein, and Norway), official candidates for EU membership (Albania, Montenegro, North Macedonia, Serbia, and Turkey) and United Kingdom.
step 2.- Selection of DEA input and output indicators. The selection of input and output indicators for implementing a DEA method was based on data availability and coverage criteria. DEA inputs were required to have data available for the year 2019, and all indicators satisfied this criterion. DEA outputs were selected among those indicators that had data available for at least 70 % of NUTS2 regions in at least one year since 2019. Based on this threshold, two indicators were excluded from the output list due to insufficient regional coverage: sdg_04_31 (Participation in early childhood education), and sdg_09_30 (R&D personnel). It is important to note, however, that this distinction is pertinent only to the specific objectives of a DEA-based analysis and can be ignored for most other potential uses of the dataset.
step 3.- Missing values: temporal and geographical imputation. Following the initial data processing step, a total of approximately 34 % missing values were identified across NUTS2 code-year-indicator combinations, necessitating further filtering and imputation. As a first measure, Extra-Regio and overseas territories were excluded from the dataset. This reduced the number of missing values to approximately 33 %.
Temporal imputation. To address missing values over time, a simple carry-forward method was applied: any missing value was replaced by the corresponding value from the previous year. This procedure reduced the number of missing entries to approximately 8 %.
Geographical imputation (NUTS1 level). Subsequently, missing values were imputed using data from the corresponding NUTS1 region, where available. Despite these imputation efforts, 49 NUTS2 regions—representing nearly 15 % of all regions—still exhibited >40 % missing data. These regions were deemed insufficiently represented and were therefore excluded from further analysis. The excluded regions included Åland (FI20), Svalbard (NO0B), and all of Albania, Liechtenstein, Montenegro, Iceland, and the UK, the latter because Eurostat stopped updating those territorial units since Brexit. After this, the number of missing values decreased to approximately 4 %.
Geographical imputation (NUTS0 level). A final round of imputation was conducted using national-level (NUTS0) data to fill remaining gaps. This step further reduced the number of missing values to approximately 3 % of the original dataset. step 4.- Final imputation. As previously noted, several intended applications of the dataset—particularly those involving multivariate analysis—require a complete set of time series without missing values. For example, this is especially critical for the DEA methodology, which cannot accommodate missing data.
Conservative imputation. An initial option considered was the outright removal of all code–year combinations containing missing values. However, this approach was deemed excessively severe, as it would result in a substantial loss of information [4,5]. For similar reasons, we also rejected extreme imputation strategies—such as substituting missing output values with zeros or imputing missing inputs with arbitrarily large values—as these could unduly distort the analysis [6,4].
Instead, we adopted a more conservative imputation strategy: remaining missing values in input indicators were replaced with the maximum observed value for the respective indicator, while missing values in output indicators were replaced with the minimum observed value [4,5]. This approach was designed to avoid artificially inflating efficiency scores while preserving the completeness of the dataset.
Neutral imputation. It should be noted, however, that the final imputation strategy described above is specifically aligned with the objectives of a DEA-based method, and may not be necessarily suitable for all potential uses of the dataset [see 5]. For alternative applications—such as descriptive statistics, econometric models, visualization, or even machine learning—a more neutral strategy replacing missing values with the indicator-wise mean may be more appropriate, which resulted in an alternative version of the dataset.
Table 2 gives a brief account of aggregated figures from the imputation steps described in this section.
Table 2.
Imputation summary table.
| 2019–2024 | Total observations | Total regions | Missing obs. | % missing | |
|---|---|---|---|---|---|
| Steps 1, 2 | Original data | 27,456 | 356 | 9220 | 33.6 % |
| Step 3 | excluding Extra-Regio's and Overseas | 26,952 | 329 | 8976 | 33.3 % |
| temporal imputation | 26,952 | 329 | 2184 | 8.1 % | |
| geo imputation from NUTS1 | 24,210 | 280 | 1153 | 4.1 % | |
| geo imputation from NUTS0 | 25,200 | 280 | 799 | 2.8 % | |
| Step 4 | Final imputation(*) | 25,200 | 280 | 0 | 0 % |
(*) For comparison purposes, only output indicators were considered in the conservative imputation (SDGTS_DB_NUTS_nona).
Table 3 compares the two final versions, with the conservative (SDGTS_DB_NUTS_nona) and neutral imputations (SDGTS_DB_NUTS_nona_alt) mentioned above. As mentioned, the first is tailored for DEA analysis, while the second is better suited for other analytical tasks such as econometric analysis.
Table 3.
Summary statistics for the two final versions.
| SDGTS_DB_NUTS_nona |
SDGTS_DB_NUTS_nona_alt |
|||||
|---|---|---|---|---|---|---|
| indicator | mean | median | stdev | mean | median | stdev |
| edat_ lfse_04 | 32.1 | 31.6 | 10.6 | 32.1 | 31.6 | 10.6 |
| edat_ lfse_16 | 10.7 | 9.0 | 6.7 | 10.8 | 9.1 | 6.6 |
| edat_lfse_22 | 13.5 | 11.0 | 7.5 | 13.5 | 11.0 | 7.5 |
| hlth_cd_ asdr2 | 0.7 | 0.6 | 0.5 | 0.7 | 0.6 | 0.5 |
| ilc_li41 | 16.4 | 14.9 | 6.6 | 16.7 | 15.3 | 6.2 |
| ilc_ lvhl21n | 8.1 | 8.0 | 4.1 | 8.3 | 8.3 | 3.9 |
| ilc_ mdsd18 | 7.0 | 5.2 | 6.1 | 7.2 | 5.5 | 6.0 |
| ilc_ peps11n | 21.6 | 19.8 | 8.4 | 22.0 | 20.2 | 8.0 |
| lfst_r_ lfe2emprt | 72.6 | 75.2 | 9.5 | 72.6 | 75.2 | 9.5 |
| lfst_r_ lfu2ltu | 2.7 | 1.8 | 2.8 | 2.7 | 1.8 | 2.8 |
| nama_10r_ 2gdp | 92.0 | 86.0 | 38.2 | 92.2 | 86.0 | 38.0 |
| rd_e_ gerdreg | 1.1 | 0.2 | 2.9 | 1.2 | 0.2 | 2.9 |
| tepsr_lm220 | 13.1 | 9.6 | 10.0 | 13.1 | 9.6 | 10.0 |
| tran_r_ acci | 51.7 | 47.0 | 25.6 | 52.1 | 48.0 | 25.3 |
| trng_ lfse_04 | 11.3 | 9.2 | 7.8 | 11.3 | 9.2 | 7.8 |
After the final imputation step, the SDG dataset was rendered fully complete, with no remaining missing values. Therefore, it can be used in a broad selection of applications where missing values must be avoided. For example, this procedure enabled us to implement the DEA-based method as described in [1], resulting in the computation and forecasting of SDG index scores as presented in Table A.2 in the repository (see DATA AVAILABILITY below), which may serve as illustration for future research into regional SDG performance and sustainable development policy evaluation.
Fig. 1 includes a flowchart with the steps followed from the original Eurostat data to the final full array.
Fig. 1.
Flowchart of the imputation steps.
Limitations
The dataset covers European countries, including EU Member States, candidate countries, and EFTA members; however, some NUTS2 territories were excluded due to insufficient data availability. Missing values were addressed through temporal and geographical imputation procedures, followed by a final imputation stage to ensure complete time series. While this approach preserves analytical usability, it may introduce uncertainty in trend estimation where imputation was extensive, but the procedure is justified and documented.
Ethics Statement
The author has read and follow the ethical requirements for publication in Data in Brief and confirms that the current work does not involve human subjects, animal experiments, or any data collected from social media platforms.
Data Availability
The 1980–2024 raw dataset with missing values present, and the two 2019–2024 SDG datasets with the complete multivariate time-series sets, i.e. with no missing values after the imputation process described above (SDGTS_DB_NUTS_raw, SDGTS_DB_NUTS_nona and SDGTS_DB_NUTS_nona_alt zipped csv files) can be accessed from the Mendeley Data repository here [7]. Additionally, Table A.2 of SDG index scores can be obtained from the same source.
Credit Author Statement
J.F.M.: Conceptualization, Methodology, Software, Writing – original draft, Writing – review & editing, Data curation, Visualization, Project administration, Funding acquisition.
Acknowledgments
This work benefited from work previously supported by the EU Interreg Atlantic Area Programme 2014–2020 and European Regional Development Fund (ERDF) under Grant EAPA 224/2016 MOSES. In addition, financial support from UPV/EHU Econometrics Research Group (Basque Government grants IT1359–19, IT1508–22) and Spanish Ministry of Science and Innovation (grant PID2020–112951GB-I0) is also acknowledged.
Acknowledgments
Declaration of Competing Interest
The author declares that he has no known competing financial interests or personal relationship that could have appeared to influence the work reported in this paper.
Footnotes
Data Envelopment Analysis.
Eurostat’s Nomenclature of Territorial Units for Statistics.
European Free Trade Association.
Data Availability
References
- 1.Fernández-Macho J. Research Square. 2025. DEA-based impact assessment and forecasting: the case of SDG compliance in Europe after Covid-19. [Google Scholar]; 10.21203/rs.3.rs-6614879/v1.
- 2.Eurostat, sustainable development indicators by region, WebPage, information note. (2025). URL https://ec.europa.eu/eurostat/databrowser-backend/api/public/explanatory-notes/get/Info-note_sdg_reg_20250206.pdf.
- 3.Eurostat, statistical symbols, abbreviations and units of measurement, WebPage. (2025). URL https://ec.europa.eu/eurostat/statistics-explained/index.php?title=Tutorial:Symbols_and_abbreviations#Statistical_symbols.2C_abbreviations_and_units_of_measurement.
- 4.Kuosmanen T. Modeling blank data entries in data envelopment analysis, Econ-WPA working paper at WUSTL 210001. 2002 https://econwpa.ub.uni-muenchen.de/econ-wp/em/papers/0210/0210001.pdf URL. [Google Scholar]
- 5.Zha Y., Song A., Xu C., Yang H. Dealing with missing data based on data envelopment analysis and halo effect. Appl. Math. Model. 2013;37(9):6135–6145. doi: 10.1016/j.apm.2012.11.015. [DOI] [Google Scholar]
- 6.Thompson R.G., Dharmapala P.S., Thrall R.M. Importance for DEA of zeros in data, multipliers, and solutions. J. Productiv. Anal. 1993;4(4):379–390. doi: 10.1007/bf01073546. [DOI] [Google Scholar]
- 7.Fernández-Macho J. Mendeley Data. 2025. 1980-2024 Europe SDG datasets and 2019-2030 DEA-based scores. [Google Scholar]; 10.17632/fsjv4vk9gv.4.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The 1980–2024 raw dataset with missing values present, and the two 2019–2024 SDG datasets with the complete multivariate time-series sets, i.e. with no missing values after the imputation process described above (SDGTS_DB_NUTS_raw, SDGTS_DB_NUTS_nona and SDGTS_DB_NUTS_nona_alt zipped csv files) can be accessed from the Mendeley Data repository here [7]. Additionally, Table A.2 of SDG index scores can be obtained from the same source.

