Skip to main content
Nicotine & Tobacco Research logoLink to Nicotine & Tobacco Research
. 2025 Dec 19;28(6):982–989. doi: 10.1093/ntr/ntaf257

Identifying Key Predictors of Smoking Cessation Success: Text-Based Feature Selection Using a Large Language Model

Thuy T T Le 1,✉, Jiongxuan Yang 2, Zimo Zhao 3, Kaidi Zhang 4, Wenjun Li 5, Yan Hu 6,7
PMCID: PMC12747148  NIHMSID: NIHMS2130205  PMID: 41416738

Abstract

Introduction

The most effective way to reduce mortality and morbidity among current smokers is to quit smoking. Although about half of smokers attempted to quit, only one-tenth succeeded in 2022. Understanding key predictors of smoking cessation success would inform smoking cessation interventions and increase quitting rates.

Methods

We analyzed data from waves 5 and 6 of the Population Assessment of Tobacco and Health (PATH) study (December 2018 to November 2021). Using OpenAI’s GPT-4.1, we identified the top 45 variables from wave 5 that are highly predictive of 12-month smoking abstinence in wave 6, based on descriptions of survey variables. We then validated the predictive power of the GPT-4.1-selected variables by comparing the performance of eXtreme Gradient Boosting (XGBoost) trained on different sets of variables. Finally, we derived insights into the top 10 variables, ranked according to their SHapley Additive exPlanations values.

Results

The performance of XGBoost trained with all possible wave 5 variables and the 45 selected variables was almost identical (AUC:0.749 vs AUC:0.752). The top 10 variables included past 30-day smoking frequency, minutes from waking up to smoking first cigarette, important people’s views on tobacco use, prevalence of tobacco use among close associates, daily electronic nicotine product use, emotional dependence, and health harm concerns.

Conclusions

The high predictive performance of XGBoost, when trained on the selected variables, underscores the efficiency and efficacy of GPT-4.1-based feature selection. The top 10 variables include various risk factors that have been previously reported in the literature for their influence on smoking behavior.

Implications

Our findings do not establish causal relationships between the selected predictors and 12-month smoking abstinence. However, identifying these key predictors provides valuable insights into the factors highly associated with smoking cessation success. This study demonstrates the ability of OpenAI’s GPT-4.1 to perform feature selection using only the textual descriptions of variables. The efficient and successful application of GPT-4.1 for variable selection highlights the potential of integrating artificial intelligence tools into tobacco research to guide resource-efficient and targeted intervention strategies.

Introduction

Smoking cessation remains a public health priority due to the well-documented adverse health effects associated with tobacco use, including cardiovascular diseases, chronic obstructive pulmonary disease, and various forms of cancer.1,2 Despite substantial efforts, achieving sustained smoking cessation continues to be a significant public health challenge. In 2022, 67.7% of the 28.8 million US adults who smoked expressed a desire to quit, and 53.3% attempted to do so. However, only 8.8% successfully quit smoking.3 This limited success highlights the complexity of the cessation process, driven by a multitude of biological, psychological, and social factors. Advanced research methodologies are needed to better identify the key predictors of cessation behaviors and improve the efficacy of future interventions.

Smoking cessation success is influenced by a complex interplay of various factors, including genetic,4,5 psychological,6,7 social and environmental elements,8-12 which interact in dynamic and multifaceted ways. A better understanding of the key predictors of cessation success is essential for enhancing quit rates. In recent years, machine learning has been increasingly applied to study tobacco research, allowing researchers to leverage large survey datasets and uncover hidden patterns not easily detected through traditional statistical approaches.13 In a machine learning workflow, feature selection plays a critical role in identifying the most informative variables for predicting the outcome of interest. Effective feature selection can improve model interpretability, reduce overfitting, and enhance predictive accuracy. Feature selection techniques can be categorized into filters, wrappers, and embedded methods, depending on how the feature selection process interacts with the modeling/learning algorithm.14 These methods have demonstrated their efficiency in automatically identifying variables significantly associated with smoking behaviors within extensive datasets.15,16 However, they require rigorous training on empirical data to be effective.

Large Language Models (LLMs) are deep learning models with a massive number of parameters that are pre-trained on vast amounts of data, allowing them to perform a variety of tasks such as language translation, text summarization, sentiment analysis, question answering, text generation, information retrieval, etc.17 They have a wide range of potential applications across domains, including healthcare, medicine, engineering, biology, finance, and marketing, among others.17,18 By leveraging their advanced natural language processing capabilities, LLMs have the potential to transform data interpretation and inform decision-making processes in public health and clinical settings.

Unlike classical machine learning models, which typically rely on structured datasets for feature selection, recent studies have indicated the potential of large language models (LLMs) in performing feature selection tasks.19-21 In this emerging approach, LLMs are prompted with textual descriptions of survey variables and can identify those most relevant to a given outcome (e.g., tobacco use behavior) without direct access to the individual-level dataset. By directly querying LLMs to evaluate the importance of each predictor based solely on its description, this method offers the advantages of simplicity and scalability across different tasks. This query-based strategy offers unique opportunities for rapid variable prioritization, but it also carries important limitations, including sensitivity to prompt design, inherited biases from pretraining data, risks of hallucination, and challenges in reproducibility - issues common to any LLM-based approaches.19,21,22 Recognizing these limitations and applying appropriate measures to address them is critical for robust application. However, when appropriately implemented, this LLM capacity enables researchers to prioritize the collection of data most relevant to addressing a specific research question, thus improving efficiency and reducing respondent burden.

The Population Assessment of Tobacco and Health (PATH) study is a nationally representative longitudinal cohort study that collects comprehensive data on participant socio-demographic information, tobacco use patterns and habits, tobacco risk perceptions, and health status, among others.23 It is an invaluable resource for studying smoking cessation behavior. In this study, we investigate the use of an LLM – OpenAI’s GPT-4.1 – in identifying the most important predictors of smoking cessation success in the PATH data. We first employed GPT-4.1 to perform feature selection based solely on the text descriptions of the variables, without accessing the actual data. Then we used XGBoost to evaluate how well these GPT-4.1-selected features are at predicting smoking cessation success in 2 years using the PATH data. XGBoost is an ensemble learning method that builds multiple decision trees and integrates them to enhance prediction accuracy, while also handling large datasets efficiently.24 By leveraging the robustness and efficiency of XGBoost, we seek to 1) assess the efficacy of GPT-4.1 for text-based feature selection, 2) gain deeper insights into the factors that facilitate or impede smoking cessation success. The pioneering use of LLMs to study smoking cessation behavior will pave the way for future exploration of this innovative approach to address other tobacco-related issues. This study’s findings could inform targeted interventions and policies designed to support individuals in their efforts to quit smoking, ultimately contributing to better health outcomes at the population level.

Materials and Methods

Data

For this study, we utilized data from the adult PATH survey23 waves 5 (December 2018 to November 2019) and 6 (March 2021 to November 2021). The adult PATH study employs a four-stage stratified sampling design to collect data from individuals aged 18 and older, both tobacco users and non-users, who live in U.S. civilian, non-institutionalized settings.25 The combination of these two waves provides us with a robust sample size for assessing smoking cessation success, defined as 12-month smoking abstinence. The cohort included current smokers who had smoked 100 or more cigarettes in their lifetime and were smoking some day or every day at wave 5. Among 34 309 adults who completed wave 5, 25 743 (75%) also participated in wave 6. Of these, 6120 (24%) were classified as current smokers in wave 5 and had a non-missing 12-month smoking status in wave 6. Smokers were considered to have successfully quit smoking if they had not smoked a cigarette in the 12 months leading up to the wave 6 survey. Among 6120 smokers from wave 5, 496 (8%) successfully quit by wave 6.

We started with a merged wave 5 and wave 6 dataset that consists of 2315 wave 5 variables, encompassing participants’ various social demographic characteristics, heath status, and tobacco-related information among others. We first filtered the data to include only current smokers from wave 5, removing not directly relevant variables such as survey weights, random variables, and those with more than 2.5% missing data or low variation. Missing values were imputed using mean for numerical variables and mode for factor variables. Subsequently, we eliminated highly correlated variables with a cutoff probability of 0.75 to further reduce the number of independent variables and enhance the performance of downstream models.

Our final dataset contains 6120 wave 5 current smokers, of whom 496 individuals were abstinent from smoking for 12 months by wave 6, and 303 wave 5 variables, which were used for further analysis. Of these variables, we employed GPT-4.1, an LLM from OpenAI,26 to select the top 45 most important variables from wave 5 in predicting smoking cessation success in wave 6. This selection was based solely on the textual descriptions of the variables, rather than their values. Subsequently, we trained XGBoost on the data with GPT-4.1-selected variables to evaluate the efficacy of GPT-4.1 in the context of feature selection and gain insights into quitting behavior. The dataset was divided into training data (80%), which was used for fine-tuning XGBoost’s hyperparameters and training XGBoost, and test data (20%) which was exclusively for evaluating the performance of XGBoost on unseen data. We used GPT-4.1 to perform feature selection based solely on textual descriptions of survey variables, i.e., without access to the actual dataset.

Feature Selection

To enhance the performance of XGBoost and decrease computational costs, feature selection is typically performed prior to training machine learning models. In this study, we leveraged GPT-4.1 to conduct feature selection using prompting techniques to identify a subset of the most informative wave 5 variables that are highly predictive of smoking status in wave 6. Specifically, we crafted prompts that presented the model with descriptions of PATH survey variables one at a time, using the model to assign a numerical score ranging from 1 to 100, with 1 being the least influential and 100 being the most influential, to reflect each variable’s importance in predicting smoking cessation status. Notably, while the model processed the descriptive text of each variable, it did not have direct access to any actual data within the dataset. Additionally, we did not provide the model with any survey details or context. Our objective was to examine how effectively LLMs can identify the most important variables for predicting smoking cessation status based solely on the model’s pre-existing knowledge. Due to the random nature of LLMs, we requested GPT-4.1 to rank each variable 50 times, and each variable’s final score is the average of 50 scores to obtain robust rankings of these variables. The efficacy of GPT-4.1 in the context of text-based feature selection would be evaluated using XGBoost in the following section. For this task, we set the GPT-4.1 parameter temperature to 0.5 to balance coherence and creativity.

Below is the prompt that we used in this study:

Prompt = “As a researcher specializing in tobacco cessation behavior analysis, after thoroughly reviewing the literature, your task is to score the importance of the variable, variable_desciption, in predicting whether an individual will abstain from smoking during the 12-month period preceding their interview, which is scheduled to take place two years from now. Provide a numeric score between 1 and 100 to reflect the variable’s importance, with 1 being the least important and 100 being the most important. Ensure your response is formatted precisely as: ‘Score: XX’, where XX represents your numeric score. Following the score, include a concise reasoning of your rating starting with ‘Reasoning:’, and limited to one or two sentences.”

where variable_desciption is the description of each variable from Table A2.

Model Development

After GPT-4.1 provided a list of scores for 303 wave 5 variables, we trained XGBoost on the training data using (i) all 303 variables and (ii) the 45 highest-scored variables. To evaluate the performance of XGBoost on the test data in each case, we conducted a series of experiments with 1000 different splits of the training dataset into training and validation subsets. For each experiment, several hyperparameters of XGBoost27 – such as learning rate (eta), maximum tree depth (max_depth), minimum child weight (min_child_weight), subsampling (subsample), and column sampling (colsample_bytree) – were optimized using Bayesian optimization28 within predefined parameter ranges to enhance model performance, see Table A1 in the Appendix. Bayesian optimization was performed using only the training subset. These parameters helped balance learning speed, control model complexity, and ensure generalization by using subsets of features and data at each iteration. The scale_pos_weight parameter, which addresses class imbalance in the dataset, was set to 3.4, calculated as the square root of the ratio of the number of successful quitters to the number of unsuccessful quitters at the time of the wave 6 survey. The model was trained for a specified number of boosting rounds (n_rounds = 1000), with early stopping (early_stopping_rounds = 400) implemented to cease training if performance did not improve after a certain number of iterations, thus preventing overfitting and conserving computational resources.

Throughout this process, Area Under the Receiver Operating Characteristic Curve29 (AUC) was used as the evaluation metric to monitor model performance. Averaged AUC and its standard deviation were computed after 1000 runs. Additionally, SHapley Additive exPlanations (SHAP) values30 were calculated after each iteration to determine the contribution of each variable to the model’s predictions. SHAP, a technique based on Shapley values from cooperative game theory, is used to explain the output of machine learning models.31 For a given observation, SHAP values quantify each feature’s contribution to the deviation of the individual prediction of that observation from the average model prediction.31 Features with positive SHAP values positively influence the prediction, while those with negative values exert a negative impact. By aggregating SHAP values from the 1000 runs, we aimed to gain insights into the directional impact of each instance on the model’s predictions and to identify the top 10 most important variables for predicting participants’ 12-month smoking abstinence status.

Results

Table 1 presents the basic demographic characteristics of our study sample. The final dataset for analysis includes 6120 individuals who were current established smokers in wave 5. Of these, 2329 (38.1%) were aged 18 to 34 years, 2166 (35.4%) were aged 35 to 54 years, and 1625 (26.5%) were aged 55 years or older. The sample consists of 3265 females (53.3%) and 2854 males (46.7%). In terms of race/ethnicity, 855 (14.0%) identified as Hispanic, 3731 (61.0%) as non-Hispanic White, 1002 (16.4%) as non-Hispanic Black, 426 (7.0%) as non-Hispanic other race, and 106 had unspecified ethnicities.

Table 1.

Basic Demographic Characteristics of the Study Sample

Characteristics Wave 5 current established smokers
Unweighted number Unweighted percentage
Overall 6120 100
Age
18 to 34 years old 2329 38.1
35 to 54 years old 2166 35.4
55 or more years old 1625 26.5
Sex
Female 3265 53.3
Male 2854 46.7
Missing 1 0.0
Race/Ethnicity
Hispanic 855 14.0
Non-Hispanic and White 3731 61.0
Non-Hispanic and Black 1002 16.4
Non-Hispanic and Other 426 7.0
Missing 106 1.6
Highest education
> = College 676 11.0
Some college 2059 33.6
≤ High school/GED 3353 54.8
Missing 32 0.6
Past 12-month total household income
> = $100 000 438 7.2
$50 000 - < $100 000 1098 17.9
< $50 000 4310 70.4
Missing 274 4.5

Leveraging only pre-trained knowledge, GPT-4.1 provided robust scores for all 303 variables in predicting 12-month smoking abstinence. Table A2 in the Appendix presents averaged scores and standard deviations of the top 45 variables across 50 different runs. These selected variables encompass various factors, including smoking frequency and habits, social and environmental influences, physical and psychological nicotine dependence, health concerns, cessation attempt and intention, and patterns of other tobacco product usage.

The performance of XGBoost is nearly identical whether considering all 303 wave 5 variables or only the top 45 features selected by GPT-4.1. Specifically, the average AUCs of XGBoost, calculated over 1000 runs, are 0.75 (95% confidence interval: 0.72 - 0.78) when trained on the 303 variables, and 0.75 (95% confidence interval: 0.73 - 0.78) when using the 45 selected features. When XGBoost is trained on the top 45 variables with the highest SHAP values, its performance is slightly lower than in the two aforementioned scenarios (average AUC: 0.74; 95% confidence interval: 0.70 – 0.77).

Figure 1 presents the top 10 most important predictors of smoking cessation success among 45 selected variables, based on their mean SHAP values - their contribution to the model’s prediction. The SHAP value of each variable for each observation indicates its directional impact on predicting the likelihood of achieving 12-month smoking abstinence. When training XGBoost, 12-month smoking abstinence status is encoded as 1 if an individual reported having not smoked in the past 12 months at the time of wave 6 survey, and 0 otherwise. Therefore, variables’ values with negative SHAP values are associated with a decreased likelihood of having not smoked in the past 12 months, while those with positive SHAP values indicate an increased likelihood. For example, Figure 1 shows that individuals with a high “past 30-day cigarette smoking frequency” in wave 5, which indicates greater nicotine dependence, were less likely to achieve 12-month smoking abstinence in wave 6. Individuals who currently used electronic nicotine products every day in wave 5 were highly associated with successfully quitting smoking in wave 6.

Figure 1.

Figure 1

The distribution of SHAP values for the top 10 most important predictors of 12-month smoking abstinence. Each plot represents a different predictor variable, including: (1) “Usually crave tobacco right after waking up,” (2) “Consider yourself a smoker,” (3) “Tobacco use for mood improvement,” (4) “Current everyday electronic nicotine product user,” (5) “Most people I spend time with are tobacco users,” (6) “Minutes from waking up to smoking first cigarette,” (7) “Important people’s tobacco views,” (8) “Everyday tobacco user,” (9) “Health concerns from past tobacco use,” and (10) “Past 30 day cigarette smoking frequency.” Each plot visualizes the SHAP value density along the x-axis and the categories or response scale for each variable along the y-axis, demonstrating the impact of each predictor on model output for smoking abstinence.

The performance of XGBoost with different feature set sizes is presented in Table A3 in the Appendix.

Discussion

Leveraging its extensive pre-trained knowledge, GPT-4.1 was able to assign importance scores to wave 5 PATH survey variables for predicting 12-month smoking abstinence status in wave 6, using only the variables’ text descriptions and without direct access to the data. With the 45 highest-scored variables, XGBoost predicted smoking cessation status in the outcome wave with high accuracy. In particular, the performance of XGBoost trained with the 45 GPT-4.1-selected variables was similar to that of the model trained with all 303 variables, and slightly better than the model trained on the top 45 variables selected using average SHAP values (AUCs: 0.75, 0.75, and 0.74, respectively). This suggests the high quality of the selected variables, demonstrating the promising capability of GPT-4.1 and LLMs in general in performing text-based feature selection. Based on its provided importance scores, GPT-4.1 can help identify key factors that should be the focus of data collection to address specific research questions. This enables researchers to prioritize this information when designing survey questionnaires.

The list of the top 45 wave 5 variables that are significantly associated with 12-month smoking abstinence in wave 6 encompasses various factors. These include information about participants’ smoking frequency and habits,12,32,33 social and environmental influences,8-12,33 nicotine dependence,34,35 health concerns,33 cessation attempt and intention,10,12,36 and other tobacco product usage10,12 (see Table A1). It is worth noting that most of the variables used to compute the tobacco dependence score in Strong et al.34,35 are included in our top 45 variables. The association of these diverse factors with successful smoking cessation aligns well with previous research findings. These alignments reinforce the relevance of the selected variables in predicting smoking cessation success, underscoring the efficacy of GPT-4.1’s text-based feature selection.

The SHAP value of each variable for each observation offers valuable information into how individual variables impact the likelihood of achieving 12-month smoking abstinence. The following insights regarding the influence of the top 10 variables are extracted from Figure 1.

Figure 1 indicates that individuals who smoked fewer days in the past 30 days, waited longer to smoke their first cigarette after waking, or usually did not want to smoke/use tobacco products right after waking up were more likely to quit successfully compared to those who smoked more frequently and sooner.12,32 In addition, individuals who identified themselves as smokers or current everyday tobacco users (i.e., those who currently used any form of tobacco product, including electronic nicotine products, combustible tobacco products such as traditional cigarettes, and smokeless tobacco products every day) were associated with a lower probability of having abstained from smoking in the past year at the follow-up.12,36 It is worth noting that our baseline population of interest consists of individuals who had smoked at least 100 cigarettes in their lifetime and currently smoked cigarettes either some days or every day. This implies that current everyday tobacco users include a subset of our baseline population who smoked every day. These factors, directly or indirectly, reflect individuals’ nicotine dependence levels, indicating that those with lower dependence were more likely to successfully quit smoking within two years.12,32,34

Furthermore, Figure 1 shows that current every day electronic nicotine product users were linked to an increased probability of being 12-month smoking abstinent. The association between daily electronic nicotine product user and smoking cessation success has been controversial in the literature with previous studies reported different findings, see12,37,38 and references therein. More research into this issue is needed to reach a consensus on the role of electronic nicotine products in smoking cessation. Individuals whose important people had negative views on tobacco use and who did not spend time mostly with smokers were more likely to quit smoking compared to their counterparts.10,39,40 This suggests that social environment plays a significant role in smoking cessation. In addition, smokers are moderately or very concerned about their future health due to their tobacco use were linked to an increased likelihood of achieving 12-month smoking abstinence. However, this association has not been well-reported in the literature (for example Lesueur’s study33 reported similar findings to ours, while Li’s41 did not). Finally, our findings also indicate that current smokers who completely agreed with the statement “smoking/using tobacco product(s) really helps me feel better if feeling down” in wave 5 were correlated with successfully quitting smoking by wave 6. In summary, further studies are required to verify the associations that are not yet well-established in the literature.

In addition to GPT-4.1, we tested older versions of OpenAI’s GPTs (e.g., GPT- 3.5-TURBO and GPT- 4-0613). While these models generally performed reasonably well, GPT-4.1 provided more robust importance scores compared to the others. In this work, GPT-4.1 was prompted executively with textual descriptions of survey variables. It is worth noting that varying prompt formulations may yield different results.21 Future studies that systematically test and compare different available LLMs and explore diverse prompt designs would be valuable for identifying optimal strategies for text-based feature selection. With classical machine learning methods such as Random Forest and XGBoost, we have a clear understanding of the criteria (e.g., SHAP values,15,16 mean decrease in accuracy, and mean decrease in impurity15,42) that are used to provide importance scores for variables. However, it is important to note that for LLMs, we do not know the criteria on which their provided scores are based.19 In this study, feature selection was conducted via the OpenAI API, rather than through personal ChatGPT interfaces, to ensure that outputs were not influenced by account-level personalization or usage history. However, as with any LLM-based method, the GPT-4.1-based feature selection has limitations. These include sensitivity to prompt design, reproducibility, risks of hallucinations, and potential inherited biases.19,21,22 For example, exposure of GPT-4.1 to PATH survey instruments or related literature during pretraining may have influenced variable rankings. While our study demonstrates promising results with a specific survey dataset, these findings may not generalize to other survey instruments due to variations in variable descriptions and context. LLM-based feature selection may also perform differently on non-English survey variables, as most LLMs are mainly pretrained on English-language data.43 More research is indispensable to better characterize the behavior of LLMs across diverse datasets and contexts. Leveraging the power of LLMs with appropriate measures to mitigate their limitations represents best practice. Because the findings of this study are based on the PATH data, it also bears the limitations of the PATH study.15

In conclusion, this study highlights the promising capability of LLMs, particularly GPT-4.1, in identifying the most important variables for predicting 12-month smoking abstinence using solely variable descriptions, without direct access to the dataset itself. Leveraging this ability among many others, LLMs can be used to aid researchers in designing survey questionnaires by allowing them to focus on the most relevant factors, thereby enhancing the efficiency and effectiveness of data collection.

Supplementary Material

Appendix_ntaf257
appendix_ntaf257.docx (26.6KB, docx)

Contributor Information

Thuy T T Le, Department of Health Management and Policy, University of Michigan School of Public Health, Ann Arbor, MI, 48109, United States.

Jiongxuan Yang, Department of Biostatistics, University of Michigan School of Public Health, Ann Arbor, MI, 48109, United States.

Zimo Zhao, School of Data Science, The Chinese University of Hong Kong Shenzhen, Shenzhen, Guangdong, 518172, China.

Kaidi Zhang, School of Data Science, The Chinese University of Hong Kong Shenzhen, Shenzhen, Guangdong, 518172, China.

Wenjun Li, Department of Public Health and Center for Health Statistics, University of Massachusetts Lowell, Lowell, MA, 01854, United States.

Yan Hu, School of Data Science, The Chinese University of Hong Kong Shenzhen, Shenzhen, Guangdong, 518172, China; National Health Data Institute, Shenzhen, Guangdong, 518172, China.

Author Contributions

Thuy TT Le (Conceptualization [lead], Data curation [lead], Formal analysis [lead], Investigation [lead], Methodology [lead], Software [lead], Validation [lead], Visualization [lead], Writing - original draft [lead], Writing - review & editing [lead]), Jiongxuan Yang (Methodology [equal], Writing - original draft [supporting]), Zimo Zhao (Conceptualization [equal], Methodology [equal], Writing - review & editing [supporting]), Kaidi Zhang (Conceptualization [equal], Methodology [equal], Writing - review & editing [supporting]), Wenjun Li (Validation [equal], Writing - review & editing [lead]), Yan Hu (Conceptualization [equal], Methodology [equal], Writing - review & editing [supporting]).

Funding

T.T.T.L was supported by the National Cancer Institute of the National Institutes of Health and the Food and Drug Administration Center for Tobacco Products (Award Number 2U54CA229974).

Declaration of Interests

None.

Data availability

All data produced are available online at https://www.icpsr.umich.edu/web/NAHDAP/studies/36498.

Disclaimer

The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH or the Food and Drug Administration.

References

  • 1. U.S. Department of Health and Human Services . The Health Consequences of Smoking: 50 Years of Progress. A Report of the Surgeon General. Atlanta, GA: U.S. Department of Health and Human Services, Centers for Disease Control and Prevention, National Center for Chronic Disease Prevention and Health Promotion, Office on Smoking and Health; 2014. [Google Scholar]
  • 2. U.S. Department of Health and Human Services . How Tobacco Smoke Causes Disease: The Biology and Behavioral Basis for Smoking-Attributable Disease: A Report of the Surgeon General. Atlanta, GA: U.S. Department of Health and Human Services, Centers for Disease Control and Prevention, National Center for Chronic Disease Prevention and Health Promotion, Office on Smoking and Health; 2010. [Google Scholar]
  • 3. VanFrank  B. Adult smoking cessation—United States, 2022. MMWR Morb Mortal Wkly Rep. 2024;73(29):633–641. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Coley  K, Wang  Q, Packer  R, et al.  Genome-wide association study of varenicline-aided smoking cessation. Nicotine Tob Res. 2025;27(10):ntaf009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Salloum  NC, Buchalter  EL, Chanani  S, et al.  From genes to treatments: a systematic review of the pharmacogenetics in smoking cessation. Pharmacogenomics.  2018;19(10):861–871. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Joly  B, Perriot  J, d’Athis  P, Chazard  E, Brousse  G, Quantin  C. Success rates in smoking cessation: psychological preparation plays a critical role and interacts with other factors such as psychoactive substances. PLoS One. 2017;12(10):e0184800. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Lawrence  D, Mitrou  F, Zubrick  SR. Non-specific psychological distress, smoking status and smoking cessation: United States National Health Interview Survey 2005. BMC Public Health. 2011;11:1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Jackson  SE, Squires  H, Shahab  L, et al.  Associations of close social connections with smoking and vaping: a population study in England. Nicotine Tob Res. 2025;27(3):447–456. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Aschbrenner  KA, Naslund  JA, Gill  L, et al.  Qualitative analysis of social network influences on quitting smoking among individuals with serious mental illness. J Ment Health. 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Lee  C-w, Kahende  J. Factors associated with successful smoking cessation in the United States, 2000. Am J Public Health. 2007;97(8):1503–1509. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Akter  S, Rahman  MM, Rouyard  T, Aktar  S, Nsashiyi  RS, Nakamura  R. A systematic review and network meta-analysis of population-level interventions to tackle smoking behaviour. Nat Hum Behav. 2024;8(12):2367–2391. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Berry  KM, Reynolds  LM, Collins  JM, et al.  E-cigarette initiation and associated changes in smoking cessation and reduction: the population assessment of tobacco and health study, 2013–2015. Tob Control. 2019;28(1):42–49. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Fu  R, Kundu  A, Mitsakakis  N, et al.  Machine learning applications in tobacco research: a scoping review. Tob Control. 2021;32(1):99–109. [DOI] [PubMed] [Google Scholar]
  • 14. Guyon  I, Elisseeff  A. An introduction to variable and feature selection. J Mach Learn Res. 2003;3(Mar):1157–1182. [Google Scholar]
  • 15. Le  TTT, Issabakhsh  M, Li  Y, et al.  Are the relevant risk factors being adequately captured in empirical studies of smoking initiation? a machine learning analysis based on the population assessment of tobacco and health study. Nicotine Tob Res. 2023;25(8):1481–1488. 10.1093/ntr/ntad066. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Le  TTT. Key risk factors associated with electronic nicotine delivery systems use among adolescents. JAMA Netw Open. 2023;6(10):e2337101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Raiaan  MAK, Mukta  MSH, Fatema  K, et al.  A review on large language models: architectures, applications, taxonomies, open issues and challenges. IEEE Access. 2024;12:26839–26874. [Google Scholar]
  • 18. Busch  F, Hoffmann  L, Rueger  C, et al.  Current applications and challenges in large language models for patient care: a systematic review. Commun Med. 2025;5(1):26. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Jeong  DP, Lipton  ZC, Ravikumar  P. LLM-select: Feature selection with large language models. arXiv preprint arXiv:240702694. 2024.
  • 20. Li  D, Tan  Z, Liu  H. Exploring large language models for feature selection: a data-centric perspective. SIGKDD Explor Newsl. 2025;26(2):44–53. [Google Scholar]
  • 21. Sclar  M, Choi  Y, Tsvetkov  Y, Suhr  A. Quantifying language models' sensitivity to spurious features in prompt design or: how i learned to start worrying about prompt formatting. arXiv preprint arXiv:231011324. 2023.
  • 22. Ouyang  L, Wu  J, Jiang  X, et al.  Training language models to follow instructions with human feedback. Adv Neural Inf Process Syst. 2022;35:27730–27744. [Google Scholar]
  • 23. FDA and NIH Study . Population Assessment of Tobacco and Health. https://www.fda.gov/tobacco-products/research/fda-and-nih-study-population-assessment-tobacco-and-health. Accessed June 2024.
  • 24. Chen  T, Guestrin  C. Xgboost: A scalable tree boosting system. 2016:785–794.
  • 25. Hyland  A, Ambrose  BK, Conway  KP, et al.  Design and methods of the population assessment of tobacco and health (PATH) study. Tob Control. 2017;26(4):371–378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. OpenAI . Introducing GPT-4.1 in the API. https://openai.com/index/gpt-4-1/. Accessed May 2025.
  • 27. XGBoost Parameters . https://xgboost.readthedocs.io/en/stable/parameter.html. Accessed June 2025.
  • 28. Snoek  J, Larochelle  H, Adams  RP. Practical bayesian optimization of machine learning algorithms. Adv Neural Inf Process Syst. 2012;25. [Google Scholar]
  • 29. Nahm  FS. Receiver operating characteristic curve: overview and practical use for clinicians. Korean J Anesthesiol. 2022;75(1):25–36. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Lundberg  SM, Lee  S-I. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30. [Google Scholar]
  • 31. Ponce-Bobadilla  AV, Schmitt  V, Maier  CS, Mensing  S, Stodtmann  S. Practical guide to SHAP analysis: explaining supervised machine learning model predictions in drug development. Clin Transl Sci. 2024;17(11):e70056. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Baker  TB, Piper  ME, McCarthy  DE, et al.  Time to first cigarette in the morning as an index of ability to quit smoking: implications for nicotine dependence. Nicotine Tob Res. 2007;9(Suppl_4):S555–S570. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Lesueur  FE-K, Bolze  C, Melchior  M. Factors associated with successful vs. unsuccessful smoking cessation: data from a nationally representative study. Addict Behav. 2018;80:110–115. [DOI] [PubMed] [Google Scholar]
  • 34. Strong  DR, Leas  E, Noble  M, et al.  Predictive validity of the adult tobacco dependence index: findings from waves 1 and 2 of the population assessment of tobacco and health (PATH) study. Drug Alcohol Depend. 2020;214:108134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Strong  DR, Pearson  J, Ehlke  S, et al.  Indicators of dependence for different types of tobacco product users: descriptive findings from wave 1 (2013-2014) of the population assessment of tobacco and health (PATH) study. Drug Alcohol Depend. 2017;178:257–266. 10.1016/j.drugalcdep.2017.05.010. [DOI] [PubMed] [Google Scholar]
  • 36. U.S. Department of Health and Human Services . Smoking Cessation. A Report of the Surgeon General. Atlanta, GA: U.S. Department of Health and Human Services, Centers for Disease Control and Prevention, National Center for Chronic DiseasePrevention and Health Promotion, Office on Smoking and Health; 2020. [Google Scholar]
  • 37. Quach  NE, Pierce  JP, Chen  J, et al.  Daily or nondaily vaping and smoking cessation among smokers. JAMA Netw Open. 2025;8(3):e250089. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Chen  R, Pierce  JP, Leas  EC, et al.  Effectiveness of e-cigarettes as aids for smoking cessation: evidence from the PATH study cohort, 2017–2019. Tob Control. 2023;32(e2):e145–e152. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Hitchman  SC, Fong  GT, Zanna  MP, Thrasher  JF, Laux  FL. The relation between number of smoking friends, and quit intentions, attempts, and success: findings from the international tobacco control (ITC) four country survey. Psychol Addict Behav. 2014;28(4):1144. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Nagawa  CS, Pbert  L, Wang  B, et al.  Association between family or peer views towards tobacco use and past 30-day smoking cessation among adults with mental health problems. Prev Med Rep. 2022;28:101886. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Li  L, Borland  R, Cummings  KM, et al.  Are health conditions and concerns about health effects of smoking predictive of quitting? Findings from the ITC 4CV survey (2016–2018). Tob Prev Cessat. 2020;6:60. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Han  H, Guo  X, Yu  H. Variable selection using mean decrease accuracy and mean decrease gini based on random forest. IEEE. 2016;219–224. [Google Scholar]
  • 43. Zhang  X, Li  S, Hauer  B, Shi  N, Kondrak  G. Don't trust ChatGPT when your question is not in English: a study of multilingual abilities and types of LLMs. arXiv preprint arXiv:230516339. 2023.

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Appendix_ntaf257
appendix_ntaf257.docx (26.6KB, docx)

Data Availability Statement

All data produced are available online at https://www.icpsr.umich.edu/web/NAHDAP/studies/36498.


Articles from Nicotine & Tobacco Research are provided here courtesy of Oxford University Press

RESOURCES