Skip to main content
PLOS One logoLink to PLOS One
. 2023 Mar 15;18(3):e0282323. doi: 10.1371/journal.pone.0282323

Connecting higher education to workplace activities and earnings

Hung Chau 1, Sarah H Bana 2,3, Baptiste Bouvier 4, Morgan R Frank 1,3,5,*
Editor: Simona Lorena Comi6
PMCID: PMC10016720  PMID: 36920887

Abstract

Higher education is a source of skill acquisition for many middle- and high-skilled jobs. But what specific skills do universities impart on students to prepare them for desirable careers? In this study, we analyze a large novel corpora of over one million syllabi from over eight hundred bachelors’ granting US educational institutions to connect material taught in higher education to the detailed work activities in the US economy as reported by the US Department of Labor. First, we show how differences in taught skills both within and between college majors correspond to earnings differences of recent graduates. Further, we use the co-occurrence of taught skills across all of academia to predict the skills that will be taught in a major moving forward. Our unified information system connecting workplace skills to the skills taught during higher education can improve the workforce development of high-skilled workers, inform educational programs of future trends, and enable employers to quantify the skills of potential workers.

1 Introduction

Education plays a critical role in economic growth and social progress. College degrees are generally associated with higher potential lifetime earnings, larger professional networks, and more adaptable careers [1, 2]. Higher education is a major part of US workforce development but information on the skills and expertise taught during higher education remain absent—even as recent research highlights the critical role of skills in shaping labor trends [35]. However, most empirical work relies on coarse labor distinctions, such as college major and institutional information (e.g., school brands), to explain these occupational trends [69]. While useful, these coarse educational and labor categories may hide further insights into the skills of “high-skilled” workers that contribute to positive career outcomes [10].

Many workers acquire skills through higher education that shape their careers. Studies have shown that social-cognitive skills and sensory-physical skills are correlated to high- and low-wage occupations, respectively, and that skill polarization divides workers with and without higher education [11]. Discrepancies between skills demanded, taught, and researched have been identified by applying textual matching techniques to job advertisements, course syllabi, and research publications in Computer Science [12]. These analyses of skills reveal gaps between the workforce and educational/training systems. Understanding the sources of these gaps, across all fields of study, may improve curriculum design, inform educational policy, and improve student outcomes when they enter the workforce.

In this work, we analyze the recently-available Open Syllabus Project (OSP) dataset, which contains over 1.4 million course syllabi from more than 3,000 US colleges and universities from 2008 to 2017. While relatively new, this data source has proven useful for modeling higher education. For example, one study quantified the skill (mis-)alignment between academic research, industry, and educational offerings in data science and data engineering [12]. They used Burning Glass (BG) skill taxonomy and applied matching techniques to extract skills appearing in job titles and descriptions, course syllabi, and publication titles and abstracts. Another study proposed a new measure for the “education-innovation gap” using the textual similarity between course syllabi and academic journals to model the dissemination of frontier knowledge into college classrooms while relating these dynamics to students’ graduation rates and incomes [13].

Our work is the first attempt to connect workplace activities to higher education through course syllabi; here, we use the granular workplace activities designed and produced by the U.S. Department of Labor (i.e., O*NET Detailed Work Activity (DWA) taxonomy described in Section Materials) to explain the underlying knowledge structures across college majors (i.e., fields of study (FOS)) and among US universities. We use word embeddings to represent textual documents [14, 15], and explore different distance metrics to measure the similarity of two embedded skill vectors. Consequently, we are able to apply agglomerative hierarchical clustering techniques to the DWA-based vector representations of FOS and universities to discover their clusters. Hierarchical clustering [16] produces a nested sequence of cluster, and the hierarchy of clusters enables us to explore clusters at any level of detail without the need of identifying a specific number of topics as would be the case with K-means clustering techniques. Motivated by the principle of relatedness [17], we model the relationships between pairs of skills across academia to forecast how skills change over time. Based on our out-of-sample earnings prediction evaluation with 5-fold cross validation, we also discover that differences in acquired skills help to explain the variance of graduates’ earnings. Our results offer an approach that connects college education to future careers. These insights may enable educational policy and academic programs to adapt to the skill dynamics in the labor market. For example, information systems that bridge between higher education and workforce skill data may inform updates to course design that prepare students with the necessary skills for their desired careers.

In summary, this paper attempts to answer these following research questions:

  • Q1. Can the granular workplace activities used by the Department of Labor to describe the US workforce also distinguish between different college majors and institutions?

  • Q2. How do the DWAs taught in a curriculum or field of study evolve over time? Can the relationships between pairs of skills across all of academia help to predict the skill evolution?

  • Q3. Do the differences in taught skills during higher education predict graduates’ earnings? Similarly, do differences in taught skills within college majors correspond to earnings differences of recent graduates?

In the next section, we describe multiple datasets that enable us to answer aforementioned research questions. We then describe our methodology in detail, present our analysis and discuss its implications and potential weaknesses to conclude the paper.

2 Materials

Open Syllabus Project Dataset (https://opensyllabus.org (OSP)) is one of the largest corpora of syllabi in the world. As of October of 2019, it contains over eight million syllabi, collected from 5,381 colleges and universities, including over three million syllabi taught at 3,186 US institutions. OSP’s fields-of-study classifier draws heavily from the Classification of Instructional Programs (CIP) taxonomy used by the National Center for Education Statistics to determine the academic field of study (e.g., Economics, Business, Computer Science) best associated with each syllabus. It includes 62 fields of study. Each syllabus has a unique identifier and the text assignment data including a description of its content, a list of references and recommended readings, and course requirements (such as assignments and exams). Syllabi can be directly mapped to graduation and enrollment statistics from the US Department of Education’s Integrated Postsecondary Education Data System (IPEDS). Syllabi are annotated with metadata including the institution, department, and academic year associated with the course. We extract and concatenate course titles, course descriptions and learning objectives from syllabi’s textual data to create “course descriptions.” More details can be found in SI Section 1 in S1 File. We limit the data from 2008 and 2017 (the ten most recent years in OSP), resulting in roughly 1.4 million syllabi representing college courses from 1,481 institutions. More about courses statistics per year and/or per field of study (FOS) can be found in S12 and S13 Figs in S1 File.

O*NET Detailed Work Activity (DWA) Taxonomy (https://www.onetonline.org/help/online/dwa). O*NET is designed and produced by the U.S. Department of Labor/Employment and Training Administration. The O*NET database allows snapshots of the relationships between occupations and skills. It has 2070 DWAs (e.g., “develop methods of social or economic research.”, “design integrated computer systems.”, “design public or employee health programs.”) representing specific work activities performed across a small to moderate number of occupations within a job family. For example, the occupations with related activities to DWA “design public or employee health programs.” include “Preventive Medicine Physicians”, “Occupational Health and Safety Specialists”, “Occupational Health and Safety Technicians”, “Dietitians and Nutritionists”, and “Dentists, General”.

Integrated Postsecondary Education Data System (https://nces.ed.gov/ipeds/) (IPEDS) is the core postsecondary education data collection program of the U.S. Department of Education’s National Center For Education Statistics (NCES). It annually collects information from all providers of postsecondary education, including public institutions, private nonprofit institutions, and private for-profit institutions, in fundamental areas such as enrollment, program completion and graduation rates. Providing data is required for any institution that applies for or participates in any Federal financial assistance program. IPEDS also includes a wide range of information about institution and institution groups, such as Degree-granting status, Institutional category, and Carnegie classifications. The Carnegie Classification, or more formally, the Carnegie Classification of Institutions of Higher Education (https://carnegieclassifications.iu.edu/), is a framework for categorizing all accredited, degree-granting institutions in the United States. It is designed to group colleges and universities based on their research activities.

College Scorecard (https://data.ed.gov/) is a U.S. Department of Education data initiative providing transparency and consumer information related to individual institutions of higher education and individual fields of study (e.g., majors) within those institutions. College Scorecard provides information about post-college earnings including median earnings of graduates working and not enrolled after completing highest credential in their first and second years for the two graduation cohorts of years 2016 and 2017. We only use the first year earnings of graduates. We process the data for Baccalaureate colleges and universities, and create the mapping between College Scorecard CIP code and OSP CIP code (the mapping can be found in this GitHub folder (https://github.com/HungChau/OSP-connect-higher-education/tree/main/cip_code_mapping)). As a result, we obtain 9007 earnings records for 832 institutions in 54 fields-of-study.

3 Methods and results

3.1 Modeling course syllabi with workplace skills

Are the workplace activities tracked by the US Department of Labor robust and effective to describe the knowledge in higher education? The O*NET database is produced by the US Bureau of Labor Statistics and details the labor market trends of workplace skills and activities by occupation. Specifically, detailed work activities (DWAs) are elements in the O*NET database that provide information about occupations’ labor requirements. This data has been used to analyze several labor market dynamics including job polarization [11, 18] and the economic resilience of cities [3, 19]. Although O*NET relates occupations to skills in the workforce, similar data is not reported for educational programs even though many high-skilled workers obtain skills in college before entering the workforce.

We bridge this gap by detecting O*NET’s detailed work activities from syllabus course descriptions. Each syllabus in the OSP data contains a description of the course content, a list of references and recommended readings, and course requirements, such as assignments and exams. Given a syllabus, we extract the course’s title, description, and learning objectives from the text and concatenate them to form the course descriptions (details are in SI Section 1A in S1 File). We apply word embeddings [20] and document similarity techniques from natural language processing to represent each DWA and syllabus as continuous vectors distributed in the same pre-trained language embedding space. Language embedding models enable us to describe the semantic similarity between two textual documents or sentences; here, we compare syllabus course descriptions to DWAs. We choose pre-trained fastText word embeddings from [21], which is constructed from all Wikipedia pages in 2017, the UMBC webbase corpus, and the statmt.org news data. We choose these word embeddings because the semantic diversity of Wikipedia and news articles should capture the semantic diversity of topics taught across FOS. This model has been used in several applications [2224], and achieves better performance than simple bag-of-words and TF-IDF [15]. We compute the relationship (0 < = rs(dwa)< = 1) between a syllabus s and a DWA by comparing their word embedding vector representations with soft cosine measure [25] (details are in SI Section 1B). As a result, syllabi are represented based on their relationships with the DWAs (called the DWA-based syllabus representation). We provide an example of the most and least prevalent DWAs detected for a political science syllabus at Harvard University in 2013 (see Fig 1A).

Fig 1. The work activities inferred syllabi reveal key differences among universities and fields of study.

Fig 1

(A) An example political science syllabus from Harvard University and the activities that are most and least strongly associated with its course description. DWA-syllabus similarity scores range from 0 (not detected) to 1 (strongly detected). (B) The DWAs that most significantly distinguish Accounting syllabi from Medicine syllabi. (C) The DWAs that most strongly separate MIT syllabi from Harvard syllabi. (D) The DWAs that most strongly separate Special Focus 4-Year Medical Schools syllabi from Engineering Schools syllabi. More examples can be found in S1-S4 Figs in S1 File.

In addition to course descriptions, syllabi are annotated with metadata about where and when the course was taught. Metadata includes the institution, department/major/FOS, and academic year. OSP’s field classifier is trained and tested on the IPEDS 2010 CIP taxonomy to determine the academic field (i.e., FOS) best associated with each syllabus. This enables us to calculate the relationship between each pair of DWAs based on the co-occurrence of dwa1 and dwa2 in any set of course syllabi S; for example, the set of all syllabi within a given FOS, simf(dwa1, dwa2) for fFOS, or across all of academia, sim(dwa1, dwa2). We experiment with various semantic distance metrics to compute DWA relationships through syllabi including Jaccard similarity, Cosine similarity, Euclidean distance, and Manhattan distance (see SI Section 2 in S1 File). We find Jaccard similarity to be the most predictive and we present those results in the main text. It is worth noting that relationships between two DWAs can be directly computed by measuring the cosine similarity of their embedding vectors. However, this approach measuring a static relationship between DWAs fails to distinguish the dynamics of how one DWA relates to another locally (i.e., within a FOS or a university) and globally (i.e., across all of academia) overtime, which will be discussed in Section Predicting the change in taught skills. For example, social skills and computer programming skills may be semantically different but co-taught as complementary skills across syllabi (e.g., computational social science, social network analysis, or econometrics).

The syllabus-DWA relationships (rs(dwa)) also enable us to model a FOS f and a university u in terms of their relationship to each of the DWAs according to, respectively,

rf(dwa)=1|Sf|sSfrs(dwa)andru(dwa)=fFOSsSf,uαf,u·rs(dwa)fFOSαf,u·|Sf,u|. (1)

These relevance scores are a measure of how strongly the skill (i.e., dwa) is represented in a field or university. While rf(dwa) (the relevance score of the dwa to FOS f) is the average over the similarity scores of that DWA across sSf, ru(dwa) (the relevance score of the dwa to university u) is the mean similarity score of that DWA across syllabi weighted by the estimated graduation rates (αf,u) of the syllabus’s field of study at that university. In the absence of course enrollment data, we use graduation rates for each FOS at each university to approximate the number of students who learn from each syllabus. Sf represents all of the syllabi within a given FOS f, and Sf,u represents all of the syllabi within a given FOS at a university u.

These tools enable us to compare pairs of syllabi, FOS, or universities based on their most common DWAs. We publish the DWA similarities by different metrics, DWA scores for each FOS and for each university by year from 2008 to 2017 in a Github repository (https://github.com/HungChau/OSP-connect-higher-education). Specifically, we compare entities of the same type (e.g., one FOS to another) by subtracting its DWA vector representation from the other’s and rank the resulting vector in descending order. We visualize the top 15 DWAs of each entity that contribute most to the difference of the pair in Fig 1B–1D. For example, the DWAs “refer patients to other healthcare practitioners or health resources” and “administer basic health care or medical treatments” most strongly distinguish Medicine from Accounting, while “analyze budgetary or accounting data” and “analyze business or financial data” identify Accounting from Medicine (see Fig 1B). Similarly, we compare pairs of universities based on their taught DWAs. As an example, “design integrated computer systems” and “design alternative energy systems” most strongly distinguish Massachusetts Institute of Technology (MIT) from Harvard University, while “forecast economic, political, or social trends” and “develop financial or business plans” more strongly identify Harvard from MIT (see Fig 1C). These results match our intuition as MIT is the world-leading engineering university and Harvard is in the top ten universities in each social science area according to U.S. News rankings. Building on this, we can group universities based on their Carnegie classification to identify the major differences in taught DWAs. We compare Medical Schools to Engineering Schools in Fig 1D. More examples can be found in S1-S4 Figs in S1 File.

3.2 Identifying Field-of-Study and university clusters

Do DWAs capture the focal knowledge offered by an academic field or a university? To further compare education among FOS, we use agglomerative hierarchical clustering on DWA-based vector representations of each FOS. Hierarchical clustering [16] produces a nested sequence of clusters like a tree (also called a dendrogram). Agglomerative clustering builds the dendrogram from the bottom level, and merges the most similar (or nearest) pair of clusters at each level to go one level up. Hierarchical clustering can take any form of distance or similarity function, and the hierarchy of clusters enables us to explore clusters at any level of detail without the need of picking a number of topics k as would be the case with K-means clustering. Pairs of FOS are similar if they are associated with similar types of work activities. For instance, Accounting is clustered together with Business and Marketing; Medicine is clustered together with Nursing, Nutrition, Health Technician, Dentistry and Veterinary Medicine; the STEM cluster includes Mathematics, Physics, Astronomy, Biology, Earth Sciences, Atmospheric Sciences and Chemistry; and the Social Science cluster includes Social Work, Political Science, History, Sociology, Women Studies, Anthropology and Religion (see Fig 2).

Fig 2. The similarity of FOS based on the prevalence of DWAs in syllabi from within those fields.

Fig 2

The dendrogram and heatmap show similar FOS clustered together based on their DWA-vector representations.

Similarly, we compare all US universities in our data set using agglomerative hierarchical clustering performed on the weighted DWA-based vector representation of each institution in Fig 3. We see that similar universities are clustered together. For example, The University of Texas Medical Branch, The University of Texas Health Science Center, and Oregon Health and Science University are clustered together. Although our dataset contains a large number of universities, we select a subset of Ivy Plus universities and universities from various IPEDS Carnegie Classifications to visualize in Fig 3. We filter out universities that have less than 100 syllabi or were missing syllabi in any year from 2008 to 2017. Carnegie classifications are mostly recovered by the clusters (see colors in Fig 3). Additionally, engineering schools like California Institute of Technology, Massachusetts Institute of Technology, and Carnegie Mellon University, are clustered together. Similarly, liberal arts schools including Cornell University, Harvard University, and University of Pennsylvania are clustered together.

Fig 3. The similarity of universities based on the graduation-weighted prevalence of DWAs offered in their course syllabi.

Fig 3

The dendrogram and heatmap reveals the hierarchical clustering of the Ivy Plus group and Special Focus Four-Year groups from the Carnegie Classification 2018 based on DWA vector representations.

3.3 Predicting the change in taught skills

How do the DWAs taught in a field of study evolve over time? In particular, which new skills or topics will emerge in a field’s syllabi? Forecasting these educational trends enables proactive course design by educators and could inform educational incentives from policy makers. Here, we use the principle of relatedness [17] to hypothesize that DWAs that occur together across all of higher education are more likely to be co-taught within a given FOS in the future. If correct, then modeling the relationships between pairs of DWAs across all of academia should forecast the introduction of new topics within a FOS even if that topic has not been part of that FOS historically. As an illustrative example, although largely absent from Economics syllabi today, machine learning may become more common in Economics because Economics already teaches linear regression which is commonly taught as an example of machine learning in Computer Science courses. As a more specific example from our data, DWAs that relate to machine learning, such as “analyze website or related online data to track trends or usage” may become more prevalent in Economics syllabi moving forward (e.g., in studies of online job postings [12, 26]).

We test our hypothesis using OSP data to predict which DWAs become important in a FOS (f). We use the relevance scores (rf(dwa)) calculated from the syllabi of each FOS in two different years (i.e., 2008 and 2017). We recast this problem as predicting the score difference (Δr) of a DWA between the two years:

Δrdwa,f=rf2017(dwa)-rf2008(dwa) (2)

We also perform classification analysis for predicting DWAs becoming important in future, which can be found in SI Section 3B in S1 File. We run several ordinary least squares (OLS) regressions to predict Δrdwa,f using the relevance scores of the dwa to FOS f (rf(dwa)) and various models of inter-DWA relationships (described in Section Modeling course syllabi with workplace skills). As a baseline, we first consider Model 1 using only the current relevance scores of DWA within each FOS with FOS fixed effects (denoted λf) according to

Δrdwa,f=β0+β1rf2008(dwa)+λf. (3)

Next, we additionally include a variable representing the co-occurrence of DWAs across syllabi within a FOS (denoted Rf) to create Model 2

Δrdwa,f=β0+β1rf2008(dwa)+β2(dwaDWAsimf(dwa,dwa)rf2008(dwa)|DWA|)Rf+λf (4)

and yet another similar Model 3 using DWA pair co-occurrences across syllabi from every FOS (denoted R)

Δrdwa,f=β0+β1rf2008(dwa)+β2(dwaDWAsim(dwa,dwa)rf2008(dwa)|DWA|)R+λf. (5)

Model 4 includes an interaction term between DWA’s relevance score within a FOS (i.e., Rf) and DWA pair co-occurrences within that FOS according to

Δrdwa,f=β0+β1rf2008(dwa)+β2Rf+β3(rf2008(dwa)*Rf)+λf (6)

and, in Model 5, using DWA pair co-occurrence across all FOS

Δrdwa,f=β0+β1rf2008(dwa)+β2R+β3(rf2008(dwa)·R)+λf (7)

As robustness checks, we run Models 2, 3, 4 & 5 with the two different methods and four distance metrics aforementioned in Section Modeling course syllabi with workplace skills for computing the DWA relationships. Although we could compare DWA pairs based solely on their semantic similarity using their word embedding vectors, this approach would miss DWA pairs that capture complementary topics. For example, Models 2 and 3 would be identical to Models 4 and 5, respectively. The results (see SI Section 3A in S1 File) show that modeling DWA relationships based on their co-occurrence in syllabi with Jaccard similarity yields the best performances across all the models involving inter-DWA relationships. We discuss these results in the main text.

We compare model performance using root mean squared error (RMSE) with 5-fold cross validation in Fig 4 (R-squared metric is reported in S11A Fig in S1 File). First, including variables representing DWA relationships decreases RMSE (i.e., Model 2 (R2 = 0.231) & Model 3 (R2 = 0.239) are statistically significantly better than Model 1 (R2 = 0.191)). Second, measuring DWA co-occurrences across all of academia (i.e., using R) instead of only within a single FOS (i.e., using Rf) improves model predictions. Specifically, Model 3 (R2 = 0.239) outperforms Model 2 (R2 = 0.231) and Model 5 (R2 = 0.244) outperforms Model 4 (R2 = 0.231).

Fig 4. Workplace activities detected from syllabi predicting teaching dynamics within a field of study.

Fig 4

We perform 5-fold cross validation and repeat 40 times (i.e., 200 trials in total) for each model and measure RMSE by the resulting model applied to the test set. Asterisks indicate the statistically significant difference between two models’ performances with Bonferroni correction. Predicting the importance of DWAs changing in nine years (2008 vs. 2017). As a baseline, model 1 only considers the current DWA score and FOS fixed effects. The other models consider the relationships between DWAs, how they interact with each other to predict how they may change in future.

These results suggest that FOS educational trends within a FOS correspond to global educational trends across all of academia. In particular, this evidence supports our hypothesis that DWAs tend to be co-taught more within a given FOS if they are bundled together across all of higher education. (e.g., Computer Science may increasingly teach “analyze green technology design requirements” since it is commonly taught with “identify information technology project resource requirements” in other FOS including Engineering). Although Model 4 does not outperform Model 2, including the interactions between current DWA relevance scores and the average of the proximity of global DWA relationships does yield a significant improvement (i.e., Model 5 outperforms Model 3). In conclusion, the best performing model is Model 5 which leverages the information about the current score of the DWA, their relationships with other DWAs across academia, and the interaction of these two variables. Model 5 improves 3.3 percent (27.5 percent) in terms of RMSE (R-squared) over Model 1, which only uses the 2008 DWA relevance scores. Therefore, we train Model 5 using the entire data, and use it to predict the relevance scores of DWAs in a FOS nine years later. Table 1 shows some examples of DWAs that became important within a FOS —in terms of ranking DWAs—in nine years. The full list of DWAs that are predicted to increase their ranks by at least five units and ranked in the top 50 in 9 years can be found in the aforementioned Github repository.

Table 1. Examples of DWAs that are predicted to increase their ranks in 9 years in particular fields.

We only select DWAs that are ranked in top 50 in future. The full list of predicted DWAs can be found in the same Github folder.

Field-of-Study Detailed Work Activity Rank (2017) Rank (2026)
Computer Science analyze green technology design requirements. 40 33
apply information technology to solve business or other applied problems. 46 40
Economics evaluate plans or specifications to determine technological or environmental implications. 37 27
develop marketing plans or strategies for environmental initiatives. 58 50
Journalism gather information about work conditions or locations. 37 24
prepare scientific or technical reports or presentations. 48 42
Medicine develop healthcare quality and safety procedures. 28 23
operate laboratory equipment to analyze medical samples. 65 50
Physics develop procedures for data entry or processing. 43 33
develop performance metrics or standards related to information technology. 41 34

3.4 Predicting graduate earnings

Do detected DWAs predict the variation in graduates’ earnings? Most—if not all—educational programs aim to provide students with the skills and abilities to successfully enter the workforce (e.g., to gain employment and maximize earnings). Most empirical work relies on coarse labor distinctions such as college major and institutional information (e.g., school brands) to correlate to graduate earnings [7, 9, 27, 28], but none have provided insights into the skills students learn that could contribute to their future earnings. Our analysis of DWAs in university course syllabi provides the first data set connecting taught skills to students’ earnings after graduation. We collect earnings of graduates from the College Scorecard earnings data from the U.S. Department of Education. Though large, the OSP course syllabus data is not distributed evenly across fields-of-study and institutions. Some fields and institutions have much less course syllabi. Thus, to sufficiently estimate work activities taught in a FOS at a university, we limit earnings records for FOS (in an institute) that have at least 10 course syllabi; and perform Kolmogorov-Smirnov statistical test to make sure the remaining earnings records representative for the entire population of the field at the institute (more details on the selection process and criteria are in SI Section 4 in S1 File). We build several OLS regression models to predict average graduate earnings across FOS (f) at a university (u) based on the relevance scores of the DWAs across fields (DWA) and within field (FOS*DWA), FOS fixed effects (FOS), school brands (i.e., school ranks (Historical U.S. News and World report rankings are compiled by Andy Reiter and available at https://andyreiter.com/datasets/) if available) fixed effects (RANK), and geography fix effects (GEO). Due to the limited availability of earnings data, we use groups of 10 ranks (i.e., 1–10, 10–20) for national universities and 15 ranks (i.e., 1–15, 15–30) for liberal arts colleges. For geographical features, we group universities together based on their divisions (U.S. Geographic Levels are available at https://www.census.gov/programs-surveys/economic-census/guidance-geographies/levels.html) (e.g., New England Division, West North Central Division). These groups are represented using indicator variables in the regression analyses.

To avoid model over-fitting, we perform 5-fold cross validation and LASSO feature selection on the models that include DWA features. LASSO [29] is one of the most popular methods for feature selection; it minimizes the residual sum of squares subject to the sum of the absolute value of coefficients being less than a constant. This constraint tends to “regularize” large models by producing some 0 coefficients when variables are co-linear. In other words, the penalty factor determines how many features are retained; using cross-validation to choose the penalty factor helps assure that the model will generalize well to future data samples. As a result, we find that DWAs improve predictions of graduate incomes (see Fig 5 for RMSE metric and S11B Fig in S1 File for R-squared metric according to 5-fold cross validation). Including DWAs improves predictions of earnings compared to FOS fixed effects (i.e., smaller RMSE). Also, R2 = 0.684 of the DWA model is significantly better than that of FOS model R2 = 0.677). Controlling for university rankings and geography further improves the FOS model (i.e., FOS+RANK+GEO (R2 = 0.757) model is significantly better than FOS (R2 = 0.677) model). But combining DWA variables with RANK and GEO variables and FOS fixed effects yields even further improvement (FOS+RANK+GEO+DWA model (R2 = 0.761) is statistically significantly better than that of FOS+RANK+GEO model). This evidence suggests that some of the information about graduate earnings represented in university rankings is also encoded the DWA variables (e.g., a LASSO regression model containing DWA variables accounts for 48% of the variation in college rankings; year and FOS fixed effects account for 7.9%). Finally, the best model (FOS+RANK+FOS*DWA) is found when we allow DWA variables to interact with FOS fixed effects which suggests that different DWAs correspond to earnings variation in different FOS (R2 = 0.779). The geographic variables also help to improve the best model’s performance but not significant (R2 = 0.782).

Fig 5. Workplace activities detected from syllabi predicting median first-year earnings of college graduates across fields of study.

Fig 5

We perform 5-fold cross validation and repeat 40 times (i.e., 200 trials in total) for each model and measure RMSE by the resulting model applied to the test set. Asterisks indicate the statistically significant difference between two models’ performances with Bonferroni correction. As a baseline, we consider the FOS, school ranking, and geographic fixed effects to predict earnings.

3.5 Within Field-of-Study skill variation and the earnings of recent college graduates

Do differences in taught skills within college majors correspond to earnings differences of recent graduates? To study how DWAs relate to earnings of graduates of a specific field of study, we perform separate regression analyses for each FOS with at least 100 institution-year observations. We employ LASSO feature selection for DWAs and report model performance using 40 independent trials of 5-fold cross-validation to mitigate over-fitting. The remaining DWAs are used to predict earnings. As can be seen from Fig 6, the DWA+GEO models perform significantly better than the baseline GEO models in terms of RMSE. Due to the limited earnings data within FOS to perform cross validation, the school ranking is omitted; the baseline models only include geographic variables (GEO). We obtain similar performance when alternatively using the model variance explained (R2) (see S11C Fig in S1 File). This result again shows that the DWAs complement the FOS information by increasing the share of the earnings explained by the model and improving the model’s predictions. However, DWA+GEO model performance varies across FOS. For example, the DWA+GEO model improves 27.2% RMSE over the GEO model for Business compared to a more modest improvement of 4.2% for Psychology. Although O*NET DWAs improve predictions in general, this varied performance across FOS could be because DWAs represent key skills and activities better in some FOS than in others. Nevertheless, our methodology shows that using granular workplace skills helps to identify important features contributing to earnings of graduates beyond course educational and labor categories.

Fig 6. Workplace activities detected from syllabi predicting median first-year earnings of college graduates within a field of study.

Fig 6

We perform 5-fold cross validation and repeat 40 times (i.e., 200 trials in total) for each model and measure RMSE by the resulting model applied to the test set. The baseline GEO model only includes geographic variables. The performances of the DWA+GEO models are statistically significantly better than the GEO models with the p-values < 0.05 for all of the reported FOS (the school ranking is omitted due to the limited earnings data).

Identifying DWAs that correspond to increased earnings after graduation could inform students’ course selection based on the demand for skills in the labor market. To demonstrate this, we analyze the regression of FOS Business as an example. After performing 5-fold cross validation on the model determined by LASSO feature selection, there are 57 DWAs remaining. Based on our statistical regression analysis, the 57 DWA features are able to explain 69.2% of the variance of the earnings in Business. Among those, 10 DWAs have significant coefficients with the p-values below 0.05. DWAs “complete documentation required by programs or regulations,” “evalutate program effectiveness,” and “advise others on career or personal development” are positively associated with earnings while “conduct health or safety training programs” is negatively associated with earnings (regression coefficients estimated with pvalue < 0.01 in each case). The list of DWAs have significant coefficients for all the 10 FOS can be found in S2 Table in S1 File. The full list of all the selected DWAs including the coefficients and statistics can be found in this GitHub folder (https://github.com/HungChau/OSP-connect-higher-education/tree/main/selected_DWAs).

4 Discussion

Knowledge, skills, and abilities shape workers’ careers, and so, quantifying their sources may impact workforce development and our understanding of the labor market. Largely, higher education is a source of skill acquisition for many middle and high-skilled jobs in America. However, there is a disconnect between work and learning in the US; higher education can fail to meet the skill demands of the labor market thus creating “skill gaps” across the country. A labor market information system where work skills are shared across entities, connecting education to work, could help students know what skills they need, educators know what skills to instruct for, employers know what skills workers have, and policy makers more effectively impact workforce development. This study demonstrates a methodology to bridge material taught in U.S. colleges and universities with the detailed work activities (DWAs) used by the Department of Labor to describe the US workforce. This creates new opportunities to track changes in the evolution of higher education and workforce development; for example, the emergence of DWAs within the syllabi of a field of study (FOS), or major, corresponds to the co-occurrence of DWA pairs across all of academia (see Fig 4). As an illustrative example, discussions of green technology design requirements may become more prominent in Computer Science programs because they go hand-in-hand with information technology project resource requirements, commonly taught in courses across academia. Educators, educational policy, and course recommendation systems could use these insights to design educational programs and to advise students towards the classes offering the experience that will be most valuable for their career goals. Following our example, proactive curriculum design might include green technology topics to prepare students for jobs in Computer Science.

However, it is likely not the case that every FOS will teach every skill or ability, in part, because labor market incentives for specific DWAs vary by industry, region, and employer. Thus, insights into the course topics that correspond to increased, or decreased, earnings after graduation (see Fig 6 for example) may increase the relevance of an educational program or policy and increase students’ success when they enter the workforce. For example, academic programs might grow to include new high-demand skills while decreasing emphasis on outdated topics. Such insights could inform goal-based learning [30] in course recommendation systems while improving explanations of recommendations. Increasingly-personalized course recommendations can identify relevant topics based on students’ predefined goals (e.g., maximizing job earnings). For example, recommending Business courses that include “complete documentation required by programs or regulations” work activities might proactively prepare today’s students to meet the growing demand for Business Analytics in the labor market.

4.1 This study has a few limitations

This study demonstrates how novel syllabus data and natural language processing (NLP) techniques can connect labor market data to higher education by predicting the change in taught skills within a FOS and linking DWAs to graduate earnings. Future work might build on our study by analyzing the causal implications of skill-level adjustments to course content. In particular, our study’s approach is unable to address selection bias when students choose a university in which to enroll. But future work may study natural experiments that overcome this barrier. Potential examples include the hiring, firing, or retirement of new faculty, the creation of a new school or department, the emergence of a large employer (e.g., resulting from new tax credit), or large donations focused on specific learning outcomes. For example, future work might augment our analysis of graduate’s recent earnings with other career outcome measures. Our analysis of the College Scorecard earnings data is limited to only two graduation cohorts and similar Post-Secondary Employment Outcomes data is limited to only a few institutions. Furthermore, we only consider earnings one year after graduation, which may not capture the full career trajectory [31]. However, future analysis involving workers’ resumes will enable direct connections between workers’ educational foundations during college and their career dynamics (e.g., worker adaptability, tenure, and mobility) in addition to earnings. Similarly, job postings analysis might compare employer demands to the DWAs detected in our study thus identifying the most or least adaptive educational programs (e.g., [12]). Future research along this dimension will offer new insights into the sources and sinks of the high-skilled workers that shape job polarization [11] and urbanization today [4, 19].

We have demonstrated, using mean cohort level graduate earnings, that there is already detectable variation in earnings based on skills taught in courses offered. Our approach has focused on outcomes for groups of graduates (e.g., by major or university). Future work with alternative data might investigate variations in labor market outcomes for individuals. For example, students studying the same major could take different courses offered, thus learning different skills. Whether the course selection by individual students leads to different occupations and different earnings, and how much learned skills could explain individual career variation are interesting questions left to be discovered. One challenge in undertaking such research is the availability and accessibility of this type of datasets at scale due to privacy concerns. Further, our analyses focused on students with bachelor’s degrees, but future work might study the skills of graduate education or the undergraduate education that lead to graduate school admission.

Our study relied on simple off-the-shelf techniques in combination with novel data sources, but future work might expand our methods with more sophisticated approaches. For example, this study used pre-trained static word embeddings and standard document similarity techniques to detect work activities from syllabi, but more complex NLP techniques could yield further insights. Static word embeddings are a powerful tool for capturing syntactic and semantic regularities in language, but each word is represented by a single vector regardless of context. That is, all senses of a polysemous word have to share the same representation. Contextualized word representations, such as Transformer-based embeddings, overcome those issues and have yielded significant improvements on many NLP tasks. Additionally, our study relies on the O*NET taxonomy used by US Department of Labor to describe labor market trends. These granular DWAs reveal core differences between courses, fields and universities. For example, DWA relevance scores improved predictions of graduate earnings within many fields of study, but not all. This suggests that “skill” differences may impact the effectiveness of college education (in terms of earnings) but O*NET DWAs may not be the most precise taxonomy to describe the granular level of knowledge expressed in courses. This is in part because O*NET data is not designed to describe higher education, but to describe workers. There is no standard knowledge base describing more granular concepts and skills in higher education and the labor market. This highlights an urgent need for future educational research that builds a knowledge base that could standardize and advance insights into how educational foundations shape workforce development and the skills of workers. With the advances of text mining methods, one could extract skills described in course syllabi and job postings, and align those skills to connect educational contents with the demands of the labor market. There are some existing job skill taxonomies to describe job postings’ requirements such as BG’s or LinkedIn’s proprietary skill taxonomies. Börner et al. (2018) analyze course syllabi and BG’s job postings focusing on areas of Data Science and Data Engineering. They use BG’s skill taxonomy instead of the one used by the U.S. Bureau of Labor Statistics to analyze skill discrepancies between research, education and jobs. Modeling job postings with NLP techniques has also been shown to be useful in understanding wage premia [32]. Although our study focuses on the work side of job seeking, we acknowledge that the demand from the employer side is also important to understand the holistic picture from skill offerings in higher education to skill demands in the labor market; which could benefit many applications such as identifying potential curricular gaps or recommending courses to meet jobs’ requirements.

Increasingly, researchers and policy makers use workers’ skills and abilities to describe labor market outcomes in addition to workers’ educational attainment based on their occupation [5]. But, similar data and methods are only just being developed and applied to workforce development and, in particular, to higher education. This study offers an approach and a methodology to connect higher education to workplace skills thus enabling new strategies for course recommendation, curriculum design, and education policy that prepare students to meet their career goals.

Supporting information

S1 File

(PDF)

Acknowledgments

We thank Erik Brynjolfsson, Seth Benzell, Daniel Rock, Nabeel Gillani, and Peter Brusilovsky for their feedback throughout this project.

Data Availability

All relevant data are within the manuscript and its Supporting information files.

Funding Statement

This research is supported in part by the University of Pittsburgh Pitt Momentum Fund and the Center for Research Computing. This work has been supported (in part) by # 2109-33808 from the Russell Sage Foundation. Any opinions expressed are those of the principal investigator(s) alone and should not be construed as representing the opinions of the Foundation. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Raj Chetty, John Friedman, Emmanuel Saez, Nicholas Turner, Danny Yagan. Mobility Report Cards: The Role of Colleges in Intergenerational Mobility. 2017.
  • 2. Witteveen Dirk, Attewell Paul. The earnings payoff from attending a selective college. Social Science Research. 2017;66:154–169. doi: 10.1016/j.ssresearch.2017.01.005 [DOI] [PubMed] [Google Scholar]
  • 3. Moro Esteban, Frank Morgan R, Pentland Alex, Rutherford Alex, Cebrian Manuel, Rahwan Iyad. Universal resilience patterns in labor markets. Nature Communications. 2021;12(1):1–8. doi: 10.1038/s41467-021-22086-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.David Autor. Work of the Past, Work of the Future. National Bureau of Economic Research. 2019.
  • 5. Frank Morgan R., Autor David, Bessen James E., Brynjolfsson Erik, Cebrian Manuel, Deming David J., et al. Toward understanding the impact of artificial intelligence on labor. Proceedings of the National Academy of Sciences. 2019;14:6531–6539. doi: 10.1073/pnas.1900949116 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Arcidiacono Peter. Affirmative Action in Higher Education: How Do Admission and Financial Aid Rules Affect Future Earnings?. Econometrica. 2005;73(5):1477–1524. doi: 10.1111/j.1468-0262.2005.00627.x [DOI] [Google Scholar]
  • 7. Cellini Stephanie Riegg, Turner Nicholas. Gainfully Employed? Assessing the Employment and Earnings of For-Profit College Students Using Administrative Data. Journal of Human Resources. 2019;54:342–370. doi: 10.3368/jhr.54.2.1016.8302R1 [DOI] [Google Scholar]
  • 8. Chetty Raj, Friedman John N, Saez Emmanuel, Turner Nicholas, Yagan Danny. Income Segregation and Intergenerational Mobility Across Colleges in the United States. The Quarterly Journal of Economics. 2020;135(3):1567–1633. doi: 10.1093/qje/qjaa005 [DOI] [Google Scholar]
  • 9. Bleemer Zachary, Mehta Aashish. Will Studying Economics Make You Rich? A Regression Discontinuity Analysis of the Returns to College Major. American Economic Journal: Applied Economics. 2022;14(2):1–22. [Google Scholar]
  • 10.Xiaoxiao Li, Sebastian Linde, Hajime Shimao. Major Complexity Index and College Skill Production. Available at SSRN 3791651. 2021.
  • 11. Alabdulkareem Ahmad, Frank Morgan R., Sun Lijun, AlShebli Bedoor, Hidalgo César, Rahwan Iyad. Skill discrepancies between research, education, and jobs reveal the critical need to supply soft skills for the data economy. Science Advances. 2018;4(7). [Google Scholar]
  • 12. Börner Katy, Scrivner Olga, Gallant Mike, Ma Shutian, Liu Xiaozhong, Chewning Keith, et al. Skill discrepancies between research, education, and jobs reveal the critical need to supply soft skills for the data economy. Proceedings of the National Academy of Sciences. 2018;115(50):12630–12637. doi: 10.1073/pnas.1804247115 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Barbara Biasi, Song Ma. The Education-Innovation Gap. National Bureau of Economic Research. 2022.
  • 14. Giabelli Anna and Malandri Lorenzo and Mercorio Fabio and Mezzanzanica Mario and Seveso Andrea. Skills2Job: A recommender system that encodes job offer embeddings on graph databases. Applied Soft Computing. 2021;101. doi: 10.1016/j.asoc.2020.107049 [DOI] [Google Scholar]
  • 15.Zachary A. Pardos, Hung Chau, Haocheng Zhao. Data-Assistive Course-to-Course Articulation Using Machine Translation. Proceedings of the Sixth (2019) ACM Conference on Learning @ Scale. 2019.
  • 16. Johnson S C. Hierarchical clustering schemes. Psychometrika. 1967;32:241–254. doi: 10.1007/BF02289588 [DOI] [PubMed] [Google Scholar]
  • 17.César A. Hidalgo, Pierre-Alexandre Balland, Ron Boschma, Mercedes Delgado, Maryann Feldman, Koen Frenken, et al. The Principle of Relatedness. Unifying Themes in Complex Systems IX. 2018:451–457.
  • 18. Acemoglu Daron, Autor David. Chapter 12—Skills, Tasks and Technologies: Implications for Employment and Earnings. Handbook of Labor Economics. 2011;4:1043–1171. doi: 10.1016/S0169-7218(11)02410-5 [DOI] [Google Scholar]
  • 19. Frank Morgan R, Sun Lijun, Cebrian Manuel, Youn Hyejin, Rahwan Iyad. Small cities face greater impact from automation. Journal of the Royal Society Interface. 2018;15(139):20170946. doi: 10.1098/rsif.2017.0946 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean. Efficient Estimation of Word Representations in Vector Space. CoRR. 2013.
  • 21.Mikolov, Tomas and Grave, Edouard and Bojanowski, Piotr and Puhrsch, Christian and Joulin, Armand. Advances in Pre-Training Distributed Word Representations. Proceedings of the International Conference on Language Resources and Evaluation (LREC 2018). 2018:1–4.
  • 22.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, Kristina Toutanova. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. NAACL. 2019:1–13.
  • 23. Trotzek Marcel, Koitka Sven, Friedrich Christoph M. Utilizing Neural Networks and Linguistic Metadata for Early Detection of Depression Indications in Text Sequences. IEEE Transactions on Knowledge and Data Engineering. 2020;32(3):588–601. doi: 10.1109/TKDE.2018.2885515 [DOI] [Google Scholar]
  • 24. Kastrati Zenun, Imran Ali Shariq, Kurti Arianit. Integrating word embeddings and document topics with deep learning in a video classification framework. Pattern Recognition Letters. 2019;128:85–92. doi: 10.1016/j.patrec.2019.08.019 [DOI] [Google Scholar]
  • 25. Sidorov Grigori, Gelbukh Alexander, Gomez-Adorno Helena, Pinto David. Soft Similarity and Soft Cosine Measure: Similarity of Features in Vector Space Model. Computacion y Sistemas. 2014;18(3):491–504. [Google Scholar]
  • 26.Emile ZCammeraat, Mariagrazia Squicciarini. Burning Glass Technologies’ data use in policy-relevant analysis: An occupation-level assessment. OECD. 2021.
  • 27. Eide Eric R., Hilmer Michael J., Showalter Mark H. Is it where you go or what you study? The relative influence of college selectivity and college major on earnings. Contemporary Economic Policy. 2016;34(1):37–46. doi: 10.1111/coep.12115 [DOI] [Google Scholar]
  • 28. Kim ChangHwan, Tamborini Christopher R., Sakamoto Arthur. Field of Study in College and Lifetime Earnings in the United States. Sociology of Education. 2015;88(4):320–339. doi: 10.1177/0038040715602132 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Tibshirani Robert. Regression Shrinkage and Selection Via the Lasso. Journal of the Royal Statistical Society: Series B (Methodological). 1996;58(1):267–288. [Google Scholar]
  • 30.Weijie Jiang, Zachary A. Pardos, Qiang Wei. Goal-Based Course Recommendation. Proceedings of the 9th International Conference on Learning Analytics & Knowledge. 2019:36–45.
  • 31. Deming David J, Noray Kadeem. Earnings Dynamics, Changing Job Skills, and STEM Careers*. The Quarterly Journal of Economics. 2020;135(4):1965–2005. doi: 10.1093/qje/qjaa021 [DOI] [Google Scholar]
  • 32.Sarah H. Bana. work2vec: Using language models to understand wage premia. Stanford Digital Economy Lab. 2022.

Decision Letter 0

Simona Lorena Comi

4 Nov 2022

PONE-D-22-21083Connecting Higher Education to Workplace Activities and EarningsPLOS ONE

Dear Dr. Frank,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

The reviewer and I see value in this paper, given the novelty that it brings to the literature on skills acquisition in higher education and how they are related to the skills demanded in the labor market. I agree with the reviewer that this topic is very interesting and important.

After carefully reading the paper, I believe that the current version of the paper suffers from several limitations and needs some additional work. Here are my main comments on this version of your article.

The abstract should be re-written following the list of results found in the paper. First, relationships between syllabi and DWA; second, prediction of the evolutions of skills; third, the relationship between DWA associated with syllabi and earnings.  

  1. Introduction. I agree with reviewer 1 that the introduction of the paper is not very effective. You start very far from your empirical exercises- even mentioning America’s position in the global economy. Then, you pose questions to which you provide no answer: “ what educational foundations best enable students to achieve their career goals?” You look at entry wage, not careers. Another example:  “How do academic majors change their curriculum over time?” while you describe some changes, do not explain how. And so on. Please, be more focused and effective in introducing your empirical study; after a succinct (3-4 lines) about the importance of tertiary education, you should skip directly to line 44. In writing the introduction, you should also try and be more precise. You do not “discover the temporal dynamics of skills and differences in skills that correspond to workers’ earnings (lines 75 -76)”. Instead, you study how skills changed over time, use your model to predict future changes, and then study how much of the variance in earning is explained by differences in acquired skills. Furthermore, is “differentiate” the appropriate word in Q1 (line 83)? Are the expressions “patterns of skills across academia” and  “pattern of topics” correct in Q2? To predict the future, you exploit the variation over time of DWA, not across academia. Wouldn’t it be better a simple “How do the DWAs taught in a curriculum evolve over time?” or something like this? Lastly, do you really predict (=out of sample prediction) graduates’ earnings in your paper?

  2. Numbering the sections would help readers follow the article’s narration.

  3. Some of the evidence reported in the Supporting Information should be moved into the main text, while some figures could be added in an Appendix. For instance, part of the discussion on how you built and selected the sample used to analyze how the DWA described earning variation should be reported in the main text (line 315). The details reported now are not sufficient to understand your empirical strategy.

  4. Kolmogorov test to select subsets of observations. You lost me here. As far as I know, the K-S test is a test of dominance between two distributions. How do you use it to select single observations?  You wrote in the Supporting Information: “It is possible that a subset of earnings records of a FOS does not effectively represent the distribution of the entire population. We perform the Kolmogorov-Smirnov (KS) statistical test for the subset population against the entire population. (lines 102-103” Specifically, how do you group observations into subsets? Which rule are you followings? I suggest using some tests to detect outliers in the earnings distribution rather than arbitrary group observations and compare their “performance” against the “population.”

  5. How exactly do you model average earnings? Your econometric models should be reported in the Supporting Information if you prefer. As it is, I do not understand how you add DWA to the earning equations: in lines 316-321, you say that you add DWAs across fields weighted (?) by propensity scores- but arent DWAs dummies at this point of the paper? – and then interacted with FOS. You lost me here. If I understood correctly, you are using dummies, specifically, a dummy for each DWA. This means that you are adding to a model with 2872 observations, 2070 DWA (SI, line 48) dummies, and 2070*47 (= n° FOS) and running a LASSO procedure to choose the best specification. Am I correct? If so, please try to be more precise in explaining this procedure, and provide some more details, possibly non-technical. For instance, is there any regularity in excluded DWA? Can some of these steps be interpreted economically?

  6. One of the ceteris paribus that you would like to add to your analysis of how much earning variation is explained by DWA detected from syllabi are local labor market conditions. Thus you should probably add to your models in Figure 4B also the state or geographical area of the university fixed effects.

  7. Which variables are included in the Baseline in figure 4C?

  8. In identifying DWAs that correspond to increased earnings after graduation (lines 362-374), you are implying that you have unbiased coefficients, which may not be the case in your setting. You must be aware that many potential threats (omitted variables, collinearity, etc.) may bias the coefficients, and warn the reader to take your results cautiously since they are conditional correlations.

  9. You have no limit to the number of figures, so why did you group them? For instance, Figure 4A should be named Figure 4 and inserted around line 293; Figure 4B should be called Figure 5  and reported around line 340, and so on. This will allow you to put titles to each panel of your figure, which actually are missing, and avoid using a general but ineffective title like that of Figure 4.

Minor points:

  1. Line 198: what do you mean by “ most prevalent”?

  2. Line 354: you probably mean: “ by increasing the “share” of the earnings explained,” not the “variation.”

  3. I would use the term predict in predicting something out of the sample. When you study the association between sets of variables, you are explaining, describing, etc. Specifically, in the section predicting graduate earnings, you are looking at how much DWAs explain variation in earnings.

  4. Line 242: “Predicting educational trend” what do you mean by the word “educational”?

  5. Please, report all the coefficients, statistics, and information of the regressions behind SI Table S2; they may be informative for the reader.

Congratulations on the work so far; I look forward to reading the revision.

Please submit your revised manuscript by Dec 19 2022 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Simona Lorena Comi

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf   

2. Thank you for stating the following financial disclosure:

“This research is supported in part by the University of Pittsburgh Pitt Momentum Fund (MRF) and the Center for Research Computing (MRF). This work has been supported (in part) by Grant # 2109-33808 from the Russell Sage Foundation (MRF & SHB). Any opinions expressed are those of the principal investigator(s) alone and should not be construed as representing the opinions of the Foundation.”

Please state what role the funders took in the study.  If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

3.  We note that you have stated that you will provide repository information for your data at acceptance. Should your manuscript be accepted for publication, we will hold it until you provide the relevant accession numbers or DOIs necessary to access your data. If you wish to make changes to your Data Availability statement, please describe these changes in your cover letter and we will update your Data Availability statement to reflect the information you provide.

4. Please update your submission to use the PLOS LaTeX template. The template and more information on our requirements for LaTeX submissions can be found at http://journals.plos.org/plosone/s/latex.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: PONE-D-22-21083

Connecting Higher Education to Workplace Activities and Earnings

Reviewer’s report

The evaluated paper addresses a relevant and current topic. The discussion section is well constructed and summarizes quite well the objectives and achievements of the article. I leave some comments hoping that they will be useful for the improvement of the manuscript.

1. The introduction should be expanded to better contextualize the paper. Please, cite works that have used the algorithm suggested in the paper (or text mining algorithms in general) in the context of the labor market and personnel selection.

2. The general and specific objectives, the research questions, and the novelty of the study should be better exposed in the first pages of the manuscript.

3. I wonder if the knowledge and skills currently described by the U.S. Department of Labor are connected to the actual requirements of the jobs (i.e., employers’ demands from job seekers). To identify potential curricular gaps, the ideal would also be to connect the course contents (knowledge and skills gained in college) with the demands of employers (for example, reviewing job offers posted on online job portals through text mining methods). In the discussion section, the authors should give their arguments in this regard.

4. K-means clustering is a classical way for text categorization, but the clustering algorithm should be better explained in the manuscript. In particular, how were the degrees and universities in figures 2 and 3 grouped?

5. On Lines 263 to 265, the authors state that: “We run several ordinary least squares (OLS) regressions to predict Δrdwa,f using various models of DWA propensity scores and inter-DWA relationships.” However, I do not get to see the results in the documents delivered. Typically, in the statistical literature, "propensity scores" refer to the probabilities predicted by a logistic regression model. Do they have the same interpretation in the context of the article? This needs clarification.

Minor issues

6. The wording of lines 322-340 should be improved and better explain the LASSO methodology.

7. Define “inconsistent student achievement.”

8. Explain better the intention of Figure 1.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2023 Mar 15;18(3):e0282323. doi: 10.1371/journal.pone.0282323.r002

Author response to Decision Letter 0


22 Dec 2022

RESPONSE TO REVIEWERS

We thank the reviewers for their critical and extensive assessment of our work "Connecting Higher Education to Workplace Activities and Earnings” submitted for consideration for publication PLOS ONE - PONE-D-22-21083. In the following we address their concerns point by point. We also specify revisions to the main paper and the Supplementary Materials. We were able to address all of the reviewers’ concerns and we feel like the manuscript is significantly improved.

REVIEWER COMMENTS

Editor (Remarks to the Author):

Introduction. I agree with reviewer 1 that the introduction of the paper is not very effective. You start very far from your empirical exercises- even mentioning America’s position in the global economy. Then, you pose questions to which you provide no answer: “ what educational foundations best enable students to achieve their career goals?” You look at entry wage, not careers. Another example: “How do academic majors change their curriculum over time?” while you describe some changes, do not explain how. And so on. Please, be more focused and effective in introducing your empirical study; after a succinct (3-4 lines) about the importance of tertiary education, you should skip directly to line 44. In writing the introduction, you should also try and be more precise. You do not “discover the temporal dynamics of skills and differences in skills that correspond to workers’ earnings (lines 75 -76)”. Instead, you study how skills changed over time, use your model to predict future changes, and then study how much of the variance in earning is explained by differences in acquired skills. Furthermore, is “differentiate” the appropriate word in Q1 (line 83)? Are the expressions “patterns of skills across academia” and “pattern of topics” correct in Q2? To predict the future, you exploit the variation over time of DWA, not across academia. Wouldn’t it be better a simple “How do the DWAs taught in a curriculum evolve over time?” or something like this? Lastly, do you really predict (=out of sample prediction) graduates’ earnings in your paper?

Authors: We thank the editor for the feedback and suggestions. We have made significant changes in the Introduction to be more focused on our study. Specifically, we have removed some background about U.S. education, jobs and earnings, and added introduction for methods and techniques (e.g., natural language processing, text mining and network principles) used in the study, including references. We have also revised research question 2 to the following:

Q2. How do the DWAs taught in a curriculum or field of study evolve over time?

Can the relationships between pairs of skills across all of academia help to predict the skill evolution?

For the out-of-sample prediction question, we did perform 5-fold cross validation for both sections. We describe this analysis in the caption of SI Figure S8. But, we agree that this information was not readily apparent in the main text and we have updated the fifth paragraph of Section 3.3 to clearly state that the models are evaluated with 5-fold cross validation. We also mention it in the Introduction.

Numbering the sections would help readers follow the article’s narration.

Authors: We thank the editor for the suggestion. We have added numbers to the section titles.

Some of the evidence reported in the Supporting Information should be moved into the main text, while some figures could be added in an Appendix. For instance, part of the discussion on how you built and selected the sample used to analyze how the DWA described earning variation should be reported in the main text (line 315). The details reported now are not sufficient to understand your empirical strategy.

Authors: We thank the editor for the comment. We provide more details in our response below, but our sample is simply the College Scorecard earnings records for FOS-university-year combinations with at least 10 syllabi in the OSP data. . We have updated the first paragraph of Section 3.4, SI Section 4 and added the following sentences to the main text to better explain our strategy:

“Though large, the OSP course syllabus data is not distributed evenly across fields-of-study and institutions. Some fields and institutions have much less course syllabi. Thus, to sufficiently estimate work activities taught in a FOS at a university, we limit earnings records for FOS (in an institute) that have at least 10 course syllabi; and perform Kolmogorov-Smirnov statistical test to make sure the remaining earnings records representative for the entire population of the field at the institute (more details on the selection process and criteria are in SI Section 4)”.

Kolmogorov test to select subsets of observations. You lost me here. As far as I know, the K-S test is a test of dominance between two distributions. How do you use it to select single observations? You wrote in the Supporting Information: “It is possible that a subset of earnings records of a FOS does not effectively represent the distribution of the entire population. We perform the Kolmogorov-Smirnov (KS) statistical test for the subset population against the entire population. (lines 102-103” Specifically, how do you group observations into subsets? Which rule are you followings? I suggest using some tests to detect outliers in the earnings distribution rather than arbitrary group observations and compare their “performance” against the “population.”

Authors: We want to clarify our intention of using KS test. We did not use the Kolmogorov test to select subsets of observations. To sufficiently estimate work activities taught in a FOS at a university, we apply a filtering process which limits earnings records for FOS (in an institution) that have at least 10 course syllabi; therefore, some earning records are removed from the FOS when taken in aggregate. We want to test whether the remaining records are representative for the entire population (i.e., all the earning records) for the FOS at the institute.

We thank the editor for the questions. We have changed “a subset of earnings records” to “the remaining earnings records” (after the filtering process) to reduce the misunderstanding. We have updated the manuscript and SI as stated in the previous response.

How exactly do you model average earnings? Your econometric models should be reported in the Supporting Information if you prefer. As it is, I do not understand how you add DWA to the earning equations: in lines 316-321, you say that you add DWAs across fields weighted (?) by propensity scores- but arent DWAs dummies at this point of the paper? – and then interacted with FOS. You lost me here. If I understood correctly, you are using dummies, specifically, a dummy for each DWA. This means that you are adding to a model with 2872 observations, 2070 DWA (SI, line 48) dummies, and 2070*47 (= n° FOS) and running a LASSO procedure to choose the best specification. Am I correct? If so, please try to be more precise in explaining this procedure, and provide some more details, possibly non-technical. For instance, is there any regularity in excluded DWA? Can some of these steps be interpreted economically?

Authors: We agree that we should clarify our modeling procedure. Our earnings data comes from the US Department of Education College Scorecard database which details the median earnings of graduating classes by FOS, institution, and year. College Scorecard earnings data is described in Section 2 - Materials and Methods “College Scorecard provides information about post-college earnings including median earnings of graduates working and not enrolled after completing highest credential in their first and second years for the two graduation cohorts of years 2016 and 2017.”

Another clarification is that the variables representing DWAs are not dummies. Instead, the variables are the propensity scores of DWAs, which vary by university, and even by FOS at a specific university. The propensity score of each DWA for a FOS (r_f(dwa)) and for a university (r_u(dwa)) are explained in Equation (1). These propensity scores are averages of document similarity scores based on language embeddings (i.e., DWA text compared to syllabus text) and are real-valued between 0 (dissimilar) and 1 (similar).

For the models involving DWAs, we perform LASSO feature selection for 2070 DWA features. LASSO is one of the most popular methods for feature selection; it minimizes the residual sum of squares subject to the sum of the absolute value of coefficients being less than a constant. This constraint tends to “regularize” large models by producing some 0 coefficients when variables are co-linear. In other words, the LASSO feature selection helps us narrow the set of DWAs down to the ones that improve the prediction of earnings. In the end, 331 DWA variables remain in the model.

Similarly, we perform LASSO feature selection for 53,820 (2070*26) FOS*DWA features, which reduces to 594 features. 26 is the number of the remaining FOS after the filtering process explained in SI lines 98-100 “Furthermore, we select FOS that have at least 30 earnings records across institutions for prediction tasks, resulting to the remaining 2601 earnings records in 26 FOS at 343 institutions (see Table S1 for details of numbers of observations of FOS in our analysis before and after filtering).”

This has now been clarified in the second paragraph of Section 3.4:

“LASSO [30] is one of the most popular methods for feature selection; it minimizes the residual sum of squares subject to the sum of the absolute value of coefficients being less than a constant. This constraint tends to “regularize” large models by producing some 0 coefficients when variables are co-linear. In other words, the penalty factor determines how many features are retained; using cross-validation to choose the penalty factor helps assure that the model will generalize well to future data samples.”

One of the ceteris paribus that you would like to add to your analysis of how much earning variation is explained by DWA detected from syllabi are local labor market conditions. Thus you should probably add to your models in Figure 4B also the state or geographical area of the university fixed effects.

Authors: We thank the editor for raising this critical question. We have come up with a solution to measure the variation across geographical areas. We group universities together based on their divisions (e.g., New England Division, West North Central Division), using U.S. Geographic Levels. These groups are represented using indicator variables in the regression analyses. We have added a discussion about this new variable at the end of the first paragraph of Section 3.4, and the new results have been reported in the second paragraph of Section 2.4 and SI Figure S11B.

Which variables are included in the Baseline in figure 4C?

Authors: The baseline models only use average earnings of the FOS (there are no covariance). Explained in Lines 350-352 “the DWA models perform significantly better than the baseline models in terms of RMSE (i.e., comparing the mean earnings of graduates of the target FOS; due to the limited data within FOS to perform cross validation, the school ranking is omitted)”; and the caption of Figure 4.C “...The baseline model is the mean earnings of graduates of that FOS…”.

In this revision, we have added geographical variables (GEO) to earnings prediction thanks to the editor’s suggestion. The baseline GEO model now includes geographical fixed effects. We have updated the first paragraph of section 3.5 “...Due to the limited earnings data within FOS to perform cross validation, the school ranking is omitted; the baseline models only include geographic variables (GEO)...”; and the caption of Figure 4.C “...The baseline GEO model only includes geographic variables…”

We have also updated section 3.5 and Figure 4C with the updated performances which involve GEO variables.

In identifying DWAs that correspond to increased earnings after graduation (lines 362-374), you are implying that you have unbiased coefficients, which may not be the case in your setting. You must be aware that many potential threats (omitted variables, collinearity, etc.) may bias the coefficients, and warn the reader to take your results cautiously since they are conditional correlations.

Authors: We thank the editor for raising the concerns. We have updated the third paragraph of the Discussion section to highlight the limitations of our work including selection bias, omitted variables, etc. We state:

Future work might build on our study by analyzing the causal implications of skill-level adjustments to course content. In particular, our study's approach is unable to address selection bias when students choose a university in which to enroll. But future work may study natural experiments that overcome this barrier. Potential examples include the hiring, firing, or retirement of new faculty, the creation of a new school or department, the emergence of a large employer (e.g., resulting from new tax credit), or large donations focused on specific learning outcomes.

You have no limit to the number of figures, so why did you group them? For instance, Figure 4A should be named Figure 4 and inserted around line; Figure 4B should be called Figure 5 and reported around line, and so on. This will allow you to put titles to each panel of your figure, which actually are missing, and avoid using a general but ineffective title like that of Figure 4.

Authors: We thank the editor for the suggestion. We have split Figure 4 into Figure 4, 5 & 6 and added the caption for each of the figures in the main text.

Line 198: what do you mean by “ most prevalent”?

Authors: We thank the editor for the question. It means “most common”. We have changed it to “most common”.

Line 354: you probably mean: “ by increasing the “share” of the earnings explained,” not the “variation.”.

Authors: We thank the editor for the suggestion. We have changed “variation” to “share”.

I would use the term predict in predicting something out of the sample. When you study the association between sets of variables, you are explaining, describing, etc. Specifically, in the section predicting graduate earnings, you are looking at how much DWAs explain variation in earnings.

Authors: We did perform out-of-sample prediction with 5-fold cross validation in our study. We compute RMSE and R-square for out-of-sample data points. We explain it in the first paragraph of Section 3.5 “We employ LASSO feature selection for DWAs and report model performance using 40 independent trials of 5-fold cross-validation to mitigate over-fitting.”.

We thank the editor for the question. We have revised the Introduction to emphasize the out-of-sample prediction evaluation with cross validation in our study.

Line 242: “Predicting educational trend” what do you mean by the word “educational”?

Authors: We thank the editor for the question. We have changed “Predicting educational trend” to “Predicting the change in taught skills”

Please, report all the coefficients, statistics, and information of the regressions behind SI Table S2; they may be informative for the reader.

Authors: We thank the editor for the suggestion. We have reported the coefficients and statistics for all the selected DWAs for the regressions in this GitHub folder. We have also added the sentence “The full list of all the selected DWAs including the coefficients and statistics can be found in this GitHub folder[13].” at the end of Section 3.5.

Reviewer #1 (Remarks to the Author):

The introduction should be expanded to better contextualize the paper. Please, cite works that have used the algorithm suggested in the paper (or text mining algorithms in general) in the context of the labor market and personnel selection. The general and specific objectives, the research questions, and the novelty of the study should be better exposed in the first pages of the manuscript.

Authors: We thank the reviewer for the suggestion. We have made significant changes in the Introduction to highlight the novelty and be more focused on our study. We have added introduction for methods and techniques (e.g., natural language processing, text mining and network principles) used in the study (including references) in the fourth paragraph of Introduction section.

I wonder if the knowledge and skills currently described by the U.S. Department of Labor are connected to the actual requirements of the jobs (i.e., employers’ demands from job seekers). To identify potential curricular gaps, the ideal would also be to connect the course contents (knowledge and skills gained in college) with the demands of employers (for example, reviewing job offers posted on online job portals through text mining methods). In the discussion section, the authors should give their arguments in this regard.

Authors: We thank the reviewer for the suggestion. We have added a discussion at the end of the fifth paragraph in Discussion section to address this:

“With the advances of text mining methods, one could extract skills described in course syllabi and job postings, and align those skills to connect educational contents with the demands of the labor market. There are some existing job skill taxonomies to describe job postings' requirements such as BG's or LinkedIn's proprietary skill taxonomies. Börner et al. (2018) analyze course syllabi and BG's job postings focusing on areas of Data Science and Data Engineering. They use BG's skill taxonomy instead of the one used by the U.S. Bureau of Labor Statistics to analyze skill discrepancies between research, education and jobs. Modeling job postings with NLP techniques has also been shown to be useful in understanding wage premia [33]. Although our study focuses on the work side of job seeking, we acknowledge that the demand from the employer side is also important to understand the holistic picture from skill offerings in higher education to skill demands in the labor market; which could benefit many applications such as identifying potential curricular gaps or recommending courses to meet jobs' requirements.”

K-means clustering is a classical way for text categorization, but the clustering algorithm should be better explained in the manuscript. In particular, how were the degrees and universities in figures 2 and 3 grouped?

Authors: We thank the reviewer for this clarifying question. The FOS are part of the syllabus metadata using the IPEDS FOS taxonomy used by the US Department of Education; that is, they are determined by experts and not through clustering techniques. K-means clustering requires us to predetermine a number of clusters that will be found (i.e., what is K?) and so we instead opt for a more flexible clustering approach. Hierarchical clustering produces a nested sequence of clusters like a tree (or dendrogram); which enables us to explore different clusters under different similarity thresholds. The groups of universities presented by different colors in Fig. 3 are defined by Carnegie classification; and these groups are mostly recovered by the clusters produced by the agglomerative hierarchical clustering method for most similarity thresholds. The novel OSP syllabus data enables us to present FOS and universities as embedding vectors using NLP techniques; these numerical vectors are used to construct the dendrograms with the clustering method.

We have added a brief explanation about the agglomerative hierarchical clustering in the first paragraph of section 3.2. “Hierarchical clustering [29] produces a nested sequence of clusters like a tree (also called a dendrogram). Agglomerative clustering builds the dendrogram from the bottom level, and merges the most similar (or nearest) pair of clusters at each level to go one level up. Hierarchical clustering can take any form of distance or similarity function, and the hierarchy of clusters enables us to explore clusters at any level of detail without the need of picking a number of topics as would be the case with K-means clustering.”

On Lines 263 to 265, the authors state that: “We run several ordinary least squares (OLS) regressions to predict Δrdwa,f using various models of DWA propensity scores and inter-DWA relationships.” However, I do not get to see the results in the documents delivered. Typically, in the statistical literature, "propensity scores" refer to the probabilities predicted by a logistic regression model. Do they have the same interpretation in the context of the article? This needs clarification.

Authors: We thank the reviewer for this clarifying question. In this report, we only apply one method to compute the DWA propensity scores for FOS and universities (explained in Section 3.1). We do have various models of inter-DWA relationships including Jaccard similarity, Cosine similarity, Euclidean distance, and Manhattan distance, and direct similarity from DWAs’ embedding vectors. The results are reported in SI Section 3A, and are mentioned in the main text in lines 269-272 “The results (see SI Section 3A) show that modeling DWA relationships based on their co-occurrence in syllabi with Jaccard similarity yields the best performances across all the models involving inter-DWA relationships. We discuss these results in the main text.”

To clarify, "propensity scores" refers to how similar the DWAs are to a syllabus using cosine distance between their language embedded vectors (explained in Section 3.1, Equation (1)), not the probabilities predicted by a logistic regression model. That is, you can consider propensity score to be the “inferred semantic distance” between DWAs and syllabi course descriptions.

To avoid the confusion, we have updated the fourth paragraph of Section 3.1. “While r_f (dwa) (the dwa propensity score of FOS f ) is the average over the similarity scores of that DWA across s ∈ S_f , r_u(dwa) (the dwa propensity score of university u) is the mean similarity score of that DWA across syllabi weighted by the estimated graduation rates (αf,u) of the syllabus’s field of study at that university.”; and the third paragraph of Section 3.3. “We run several ordinary least squares (OLS) regressions to predict ∆rdwa,f using the DWA propensity scores of FOS f (r_f (dwa)) and various models of inter-DWA relationships (described in Section 3.1).”

The wording of lines 322-340 should be improved and better explain the LASSO methodology.

Authors: We thank the reviewer for the suggestion. We have added these new sentences to the second paragraph of Section 3.4 to explain the LASSO methodology:

“LASSO [30] is one of the most popular methods for feature selection; it minimizes the residual sum of squares subject to the sum of the absolute value of coefficients being less than a constant. This constraint tends to “regularize” large models by producing some 0 coefficients when variables are co-linear. In other words, the penalty factor determines how many features are retained; using cross-validation to choose the penalty factor helps assure that the model will generalize well to future data samples.”

Define “inconsistent student achievement.”

Authors: We have revised the Introduction section based on the previous comments. The revision does not include “inconsistent student achievement”.

Explain better the intention of Figure 1.

Authors: We thank the reviewer for the comment. The intention of Figure 1 to show evidence supporting the content of Section 3.1. Thanks to the novelty of OSP syllabi dataset, our work is the first attempt to connect workplace activities to higher education through course syllabi using The O*NET database that is produced by the US Bureau of Labor Statistics and details the labor market trends of workplace skills and activities by occupation. The visualizations in Figure 1 show that, using natural language processing technique (i.e., word embeddings), we are able to infer work activities in syllabi, revealing key differences among universities and fields of study. We have added the following paragraph as the beginning of the second paragraph in Section 3.1 to provide additional details about the process of modeling detailed work activities and course syllabi with word embeddings.

“We bridge this gap by detecting O*NET’s detailed work activities from syllabus course descriptions. Each syllabus in the OSP data contains a description of the course content, a list of references and recommended readings, and course requirements, such as assignments and exams. Given a syllabus, we extract the course’s title, description, and learning objectives from the text and concatenate them to form the course descriptions (details are in SI Section 1A). We apply word embeddings [24] and document similarity techniques from natural language processing to represent each DWA and syllabus as continuous vectors distributed in the same pre-trained language embedding space. Language embedding models enable us to describe the semantic similarity between two textual documents or sentences; here, we compare syllabus course descriptions to DWAs. We choose pre-trained fastText word embeddings from [25], which is constructed from all Wikipedia pages in 2017, the UMBC webbase corpus, and the statmt.org news data.”

Attachment

Submitted filename: Plos One - RESPONSE TO REVIEWERS.pdf

Decision Letter 1

Simona Lorena Comi

30 Jan 2023

PONE-D-22-21083R1

Connecting Higher Education to Workplace Activities and Earnings

PLOS ONE

Dear Dr. Frank,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Overall, the reviewer and I are quite happy with the progress made so far. However, reviewer 1 still wants to see a few additional changes – minor points- to your article before recommending an unconditional acceptance. I agree with that assessment – you have responded very well to most of the issues raised, and as a result, the manuscript has taken a big step toward a final publication. However, I would still ask you to address the remaining comments of reviewer 1 in the last revision of your paper.

Please submit your revised manuscript by Mar 16 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Simona Lorena Comi

Academic Editor

PLOS ONE

Journal Requirements:

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

********** 

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

********** 

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

********** 

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

********** 

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

********** 

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Connecting Higher Education to Workplace Activities and Earnings

PONE-D-22-21083R1

The reviewer thanks the authors for the significant revision they have performed on the manuscript. In general, they adequately respond to all my concerns.

Some minor remarks:

- The sections and subsections are not numbered as the authors claim in their responses.

- The Materials and Methods Section (Lines 68 to 119) only shows materials (databases used). Actually, the methods appear starting from line 120 in the Results Section. Check this out.

- I keep seeing the term "propensity score" as confusing. It should be better clarified in the manuscript. Is it really the term that is used in computer science literature?

********** 

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2023 Mar 15;18(3):e0282323. doi: 10.1371/journal.pone.0282323.r004

Author response to Decision Letter 1


3 Feb 2023

RESPONSE TO REVIEWERS

We thank the reviewers for their critical and extensive assessment of our work "Connecting Higher Education to Workplace Activities and Earnings” submitted for consideration for publication PLOS ONE - PONE-D-22-21083. In the following we address their concerns point by point. We also specify revisions to the main paper. We were able to address all of the reviewers’ concerns and we feel like the manuscript is significantly improved.

REVIEWER COMMENTS

Reviewer #1 (Remarks to the Author):

The sections and subsections are not numbered as the authors claim in their responses.

Authors: We have numbered the sections in the manuscript as suggested.

[Reviewer] The Materials and Methods Section (Lines 68 to 119) only shows materials (databases used). Actually, the methods appear starting from line 120 in the Results Section. Check this out.

Authors: We thank the reviewer for pointing it out. We have updated the section title as follows:

- “Materials and Methods” to “Materials”

- “Results” to “Methods and Results”

[Reviewer[ I keep seeing the term "propensity score" as confusing. It should be better clarified in the manuscript. Is it really the term that is used in computer science literature?

Authors: To reduce the confusion, we have changed “propensity score” to “relevance score” throughout the manuscript, and also added the definition of “relevance score” after the Equation (1) as follows “These relevance scores are a measure of how strongly the skill (i.e., dwa) is represented in a field or university”.

Attachment

Submitted filename: Response to Reviewers.pdf

Decision Letter 2

Simona Lorena Comi

14 Feb 2023

Connecting Higher Education to Workplace Activities and Earnings

PONE-D-22-21083R2

Dear Dr. Frank,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Simona Lorena Comi

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Thanks for the reviews. The manuscript has been substantially improved. I hope you can continue in this line of research.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

**********

Acceptance letter

Simona Lorena Comi

20 Feb 2023

PONE-D-22-21083R2

Connecting Higher Education to Workplace Activities and Earnings

Dear Dr. Frank:

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department.

If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org.

If we can help with anything else, please email us at plosone@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Professor Simona Lorena Comi

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 File

    (PDF)

    Attachment

    Submitted filename: Plos One - RESPONSE TO REVIEWERS.pdf

    Attachment

    Submitted filename: Response to Reviewers.pdf

    Data Availability Statement

    All relevant data are within the manuscript and its Supporting information files.


    Articles from PLOS ONE are provided here courtesy of PLOS

    RESOURCES