Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Aug 8.
Published before final editing as: J Psychopathol Clin Sci. 2026 Aug 6:10.1037/abn0001160. doi: 10.1037/abn0001160

Large Language Models Can Supplement the Assessment of Clinical High Risk for Psychosis

Luz Maria Alliende 1,*, Rob Voigt 2, Maeve Hoffman 1, Gregory P Strauss 3, Lauren M Ellman 4, Elaine F Walker 5, Philip Corlett 6, Jason Schiffman 7, Scott W Woods 6, Albert Powers 6, Steven M Silverstein 8, James A Waltz 9, Richard Zinbarg 1, Shuo Chen 9, Trevor F Williams 10, Joshua Kenney 6, James M Gold 9, Vijay A Mittal 1,11,12,13
PMCID: PMC13449246  NIHMSID: NIHMS2193888  PMID: 42560867

Abstract

Capturing psychosis risk before illness onset is an ongoing challenge for psychosis-spectrum studies. Natural language processing (NLP) tools can harness information embedded in the notes generated during clinical interviews to obtain more objective markers of psychosis risk; for this project, we used readily available assessor notes. The following project acts as proof of concept on the usefulness of AI-based tools to capture latent psychosis risk information. We used assessor notes for 2077 Structured Interview for Psychosis-Risk Syndromes interviews as input for different AI-based NLP models to produce metrics for psychosis risk. Namely: a trained encoder-only model (ModernBERT), and zero-shot and few-shot instantiation of two decoder-only models (LLaMA-4 and GPT-4.1 mini). AI-based risk metrics were compared to gold-standard ratings of clinical high risk for psychosis (CHR). The AI-based risk metrics’ relationship to traditional psychosis risk scores (SHARP and NAPLS) was also assessed. Finally, we explored the added benefit of adding AI-based risk metrics to models predicting future participant conversion. All models performed above chance in classifying interview notes for the presence or absence of CHR syndromes. The trained encoder model performed the best out of all models determining the presence of CHR syndromes (accuracy = 82.67%, κ = 0.63). Positive CHR classification by the encoder model resulted in a 0.69 and 0.85 standard deviation increase in SHARP and NAPLS risk scores (p < .001). A one standard deviation increase in decoder-generated risk scores increased traditional risk scores between 0.24 to 0.47 standard deviations (all ps < .001). In exploratory analyses, decoder generated risk scores incrementally improved models predicting conversion including SHARP but not NAPLS risk scores. While our results need to be considered in the context of one consortium, AI-based NLPs show potential as an aid for the diagnosis of CHR syndromes, evaluating psychosis risk, and even predicting future conversion, even with sub-optimal but readily available inputs (i.e., assessor notes). Future projects could utilize AI-based tools’ potential in augmenting psychosis risk screenings and risk predictors.

General Summary:

Early identification of individuals at risk for psychosis is an ongoing challenge in mental healthcare and research. This project suggests that generating AI-based risk markers from the notes assessors write during assessment can improve the identification of psychosis risk.

Introduction

People at clinical high risk (CHR) for psychosis, those who meet criteria for attenuated psychosis syndromes, are at greater risk of developing a future threshold psychotic disorder (Cannon et al., 2008; Yung & McGorry, 1996), supporting the utility of early identification of this population to enhance prevention. However, timely determination of whether youth are experiencing these syndromes requires specialized assessments that are not widely (Kotlicka-Antczak et al., 2020) or equitably (DeLuca et al., 2022) available. Further, even when determining that youth are experiencing a CHR syndrome, the majority of these youth will never develop a future threshold psychosis spectrum disorder, with conversion rates averaging around 25% according to a recent meta-analysis (Salazar de Pablo et al., 2021). Regardless of future conversion status, people that meet criteria for a CHR syndrome will tend to have significant and persistent symptom distress and functional impairment even if they no longer experience subthreshold positive symptoms (Addington et al., 2011). More accurately capturing psychosis risk is an ongoing task for psychosis spectrum researchers and clinicians.

Established risk metrics, such as those developed by the Shanghai-At-Risk-for-Psychosis (SHARP) (Zhang et al., 2019) and North American Prodrome Longitudinal Study (NAPLS) (Cannon et al., 2016) working groups, are invaluable tools and offer comparable accuracy to risk calculators for cardiovascular disease. However, these scores rely on lengthy symptom, functional, and cognitive evaluations by trained assessors.

AI-based natural language processing (NLP) tools could bridge some of the gaps in accurate and scalable CHR diagnosis protocols. First, NLP tools can provide novel insights and increase the accuracy of psychosis-risk assessment by tapping into information embedded in unstructured, freely generated assessor notes. Second, the automated nature of these tools offers the potential for cost-effective, scalable assessments that are not bound by the same bottlenecks as traditional tools.

The present project serves as proof-of-concept that AI-based tools can extract novel psychosis risk information from readily available language-based data: assessor notes. These notes - unstructured free text generated by interviewers while assessing psychosis risk- are idiosyncratic to each assessor as they are not intended to be used as data but rather as an aid for assessment. Despite the heterogeneity of these notes and how much information is lost if we were to compare them to interview transcripts, assessor notes capture clinically relevant information and are easily accessible as data. Systematically analyzing these notes by hand, would be prohibitively costly. However, using AI-based NLP tools could allow researchers and clinicians to efficiently and effectively use unstructured language as a source of risk information.

There is a current surge of transformer-based artificial intelligence NLP tools, with proliferating novel applications and accelerated improvements in quality and capabilities. The “transformer” is an influential neural network architecture based on self-attention mechanisms that allow corresponding models to capture and integrate contextual relationships across positions in an input sequence; this allows each unit of meaning (i.e., token representation) to be informed by relevant parts of the input simultaneously (Vaswani et al., 2017). Transformer-based NLP models are not homogenous in the tasks they are best suited for and, broadly, tend to focus on either encoding or decoding. Encoder models (ex. BERT-based models) are designed and optimized to “understand” input text, making them especially well-suited for tasks like classification (Devlin et al., 2019). Conversely, decoder models (ex. OpenAI’s ChatGPT) focus on producing relevant language outputs, hence why they are sometimes called generative AIs (Radford et al., 2018). Decoder models use what is called “causal” or “masked” attention, which restricts the attention-based information incorporated into each token to only the tokens that preceded it; such models are well-suited for producing relevant language outputs, which has given rise to the common parlance of “generative AIs” (Radford et al., 2018).

Additionally, models can be exposed to different amounts of data examples as part of their tailoring for a specific task; this runs the full gamut, as there can be anywhere from no exposure to extensive training that modifies model parameters. On one hand, “zero” and “few-shot” learning are forms of “in-context learning” in which models adapt behavior without parameter updates (Brown et al., 2020). Zero-shot learning refers to cases where there are no examples of the expected outcome, just task instructions. In few-shot learning, a few input-output examples are provided as part of the prompt, allowing the models to use these examples in addition to their pre-trained knowledge to better respond to a new input without needing to modify model parameters. Contrastingly, fine-tuning relies on extensive exposure to examples to update a model’s pre-trained parameters to better adjust to a specific application (Ziegler et al., 2020). For encoder models, this means providing multiple inputs and their respective labels to modify the final layers of the model to increase accuracy for the dataset. Decoder models aim to predict the output’s next token based on the input data; exposing the models to the expected outcomes for fine-tuning allows this completion process to better suit the task at hand. Figure 1 provides a diagram for the encoder and decoder models used in this study, and the exposure to examples from the dataset; all models are thoroughly described in the methods section.

Figure 1. Study’s Model Architecture and Exposure to Training Data.

Figure 1.

Beyond broad architectural differences and exposure to data, individual models differ widely technically: how they are constructed, the data used to train them, and their performance across tasks. These differences also expand to logistical, security, and ethical considerations, such as differences in cost, their proprietary nature, how user data is processed and stored, or redlining of sensitive topics. The current project tested how models from different sources, with different architectures and different exposures to training data can aid in the prediction of psychosis risk.

The present project aims to explore the ways different NLP tools can be useful in estimating psychosis risk using readily available assessor notes in a sample with a psychiatrically diverse control group. For the first aim of this study, we assessed each NLP model’s performance classifying interview notes by CHR status. We expected all models to determine CHR status significantly above chance. We also expect the reliability between our NLP-based CHR classifications and clinician’s CHR rating to be below the inter-test reliability between well-established clinician-based tests to determine CHR. For the purposes of this study, we will set a benchmark of κ = 0.78 as this is the agreement between the two most commonly used tests to determine CHR (i.e. the Structured Interview for Psychosis-risk Syndromes (SIPS) and Comprehensive Assessment of At-Risk Mental States (CAARMS)) (Fusar-Poli et al., 2016). Additionally, we expected that trained encoder models would outperform zero and few-shot generative models given their parametric tailoring to assessor notes and ratings and that this model’s architecture is especially well-suited for classification. As our second aim, we assessed the relationship between NLP models’ risk outcomes and those of well-established dimensional psychosis-risk algorithms that quantify the probability of future conversion to full psychosis. We expected that positive cases and higher risk scores from NLP models would have significantly higher standard psychosis-risk scores. Lastly, as an exploratory aim, we assessed whether adding NLP-based risk metrics aid the prediction of future conversion status. We expected that models, including traditional risk scores, would significantly improve when incorporating LLM-based risk metrics. By accomplishing these aims, we evaluated whether NLP is useful for enhancing commonly implemented clinical tools for psychosis risk prediction. This study acted as a proof-of-concept for the potential scalable use of these tools with readily available assessor notes.

Methods

Participants

Participants were recruited as part of the CAPR consortium- a large multisite study aimed at determining the computational underpinnings of psychosis risk. Northwestern University’s Institutional Review Board approved all protocols and procedures (Protocol: STU00211351), and they were acknowledged by the Institutional Review Boards of all participating sites. This consortium recruited CHR participants, healthy controls (HC) and clinical controls (CC) who had never had a psychotic disorder between the ages of 12 and 34. Participants were assessed at baseline and at 12, 24-month follow-ups. Participants from all clinical groups were included in this study (total N = 1057, CHR = 426, CC = 384, HC = 247) if there were any notes available for SIPS-P interviews (N = 2077) and had at least a SHARP or NAPLS risk score measure available. Participant demographics can be found in Table 2.

Table 2.

Sample Descriptive Statistics

Total Participants (N = 1057) CHR (N = 426) CC (N = 384) HC (N = 247) CHR vs. Both control groups
Demographics
Participant’s Sex Assigned at Birth (N Female, % Female) 656 (62.06%) 266 (62.44%) 244 (63.54%) 146 (59.11%) X2 (1) = 0.02
p = .89
Age at baseline (Mean years, sd) 23.05 (4.10) 23.00 (4.16) 23.10 (4.03) 23.10 (4.13) t (1055) = 0.65
p = .52
Clinical
SHARP Risk Score at baseline (Mean, sd, N) 3.41 (3.09)
N = 779
4.60 (3.06)
N = 403
2.17 (2.61)
N = 355
1.33 (1.40)
N = 21
t (777) = 12.18
p < .001*
NAPLS 1 year Risk Score at baseline
(Mean, sd, N)
6.90 (7.45)
N = 747
11.90 (10.20)
N = 272
4.51 (2.60)
N = 273
3.41 (1.53)
N = 202
t (745) = 16.05
p < .001*
Interviews Notes N = 2077
N = 801
N = 778
N = 498
Visits Baseline = 965
12m = 691
24m = 421
Baseline = 371
12m = 263
24m = 167
Baseline = 355
12m = 264
24m =159
Baseline = 239
12m = 164
24m = 95
X2 (2) = 0.30
p = .86
Interview Word Count
(Mean, sd)
1542.68 (723.17) 1946.28
(875.29)
1362.87 (462.27) 1131.53 (285.35) t (1032) = 20.37
p < .001*

To determine participants’ clinical group at baseline, all potential participants were screened for psychosis-risk syndromes (for a detailed study protocol see: Mittal et al., 2021). If participants met criteria for attenuated positive symptom syndrome, brief intermittent psychotic syndrome, and/or genetic risk and deterioration syndrome according to criteria from the Structured Interview for Psychosis-risk Syndromes (SIPS) (Miller et al., 2003) they were included in the CHR group. Participants that did not meet criteria for a CHR syndrome were included as either HCs or CCs. HCs had no lifetime history of DSM disorders as assessed by the SCID-5 (First et al., 2016) and have experienced only questionable (SIPS-P score = 1) or less severe positive symptoms in their lifetime. CCs had either: 1. a history of or currently met criteria for DSM disorders, 2. Increased vulnerability for psychosis because they had a first-degree family member with a psychosis spectrum disorder or met criteria for schizotypal personality disorder, or 3. Experienced sub-syndromic positive symptoms. Sub-syndromic positive symptoms were operationalized as experiencing either frequent mild positive psychosis spectrum symptoms (SIPS-P score = 2) or more severe SIPS-P symptoms happening infrequently or that are better explained by another DSM disorder. We opted to include CCs’ participants and their interviews in this study to better replicate clinical or research populations and further stress-test model capabilities. In research and clinical settings, an ongoing challenge is truly capturing psychosis risk rather than people experiencing similar symptoms due to a different disorder (ex. perceiving hostility from others due to social anxiety rather than paranoia) or sub-syndromic psychotic-like experiences due to normative individual variability or minimal underlying trait risk factors (ex. familial risk or environmental stress exposure) (Lundin et al., 2024 ; Staines et al., 2022; van Os & Reininghaus, 2016).

Participants that met criteria for a threshold psychotic disorder at baseline according to SIPS-criteria were excluded from the study.

Conversion to psychosis in future interviews was determined using the SIPS criteria for presence of psychotic symptoms.

Measures

SIPS Positive Symptoms Assessor Notes

All participants in the CAPR study were assessed using the positive subscale of the SIPS (Miller et al., 2003) to assess them for psychosis risk positive symptoms. This semi-structured interview consists of 48 questions that aim to measure: unusual thought content or delusional ideas, suspiciousness or persecutory ideas, grandiose ideas, perceptual abnormalities or hallucinations, and disorganized communication. Through this symptom assessment the interview is used to determine if there is current psychosis, if the person meets criteria for a psychosis-risk syndrome and the status of said symptoms, as well as the severity of psychosis-risk symptoms.

During the SIPS interview, assessors were provided with free-text fields on Redcap after each question that they could use to take notes throughout the assessment. Some sample questions from the SIPS-P subscale and notes in response to this question are included in Table 1, Supplement 1, 2 and 3. Notably, these notes are only meant to help assessors make more accurate ratings and are not intended to be used as a systematic form of data collection. For the purposes of this study, we included all questions in the positive subscale of the SIPS followed by the respective assessor notes as input for the relevant model. On average this input text was 1542.68 words long (sd = 723.17, see Table 2 for clinical group breakdown). Expectedly, interview notes for CHR participants were significantly longer than for the control comparison groups (t (1032) = 20.37, p < .001). All notes were de-identified by hand before analyzing, removing any personal health information, or information that could determine a participant’s identity or personal health information (See Table 1 for an example on how deidentification was handled). See Supplement 1 for more example notes for one of the scale’s questions. For this study, 79 assessors across 7 sites conducted 2077 interviews for 1057 participants. Of these participants, 362 had one available interview, 370 had two, and 325 had three interviews with available notes. In this sample, 20 CHR participants with any follow-up data (N = 292) later converted to a psychotic disorder (6.85%) for whom 42 interviews prior to conversion were available.

Table 1.

Example Input Text for Mock Participant: SIPS Positive Questions and Interviewer Notes

Mock exemplary notes for SIPS Positive questions
All participants were asked every question in the interview and all notes were included. The following are mock notes that capture typical interviewer notes to preserve participant privacy.

(…) Do you ever feel that some person or force may be controlling or interfering with your thinking?
No.
Do you ever feel as if your thoughts are being said out loud so that other people can hear them?
As if friends could hear thought. Not really.
Do you ever think that people might be able to read your mind?
Best friend knows what they are thinking. [name] knows her well and can predict her answers.
Not essence of question.
Do you ever think that you can read other people’s minds?
No.(…)
(…)Have you ever found yourself feeling mistrustful or suspicious of other people?
Sometimes, about guy’s intentions. Has happened in past month. Vague feeling not worried about specific harm. “all men cheat” takes a long time to trust guys, trusts after knowing them better. Worse in past year, at least 3 times a week, sometimes more ex partner cheated, thinks this is their nature. specific to her.
Do you ever feel that you have to pay close attention to whaťs going on around you in order to feel safe?
Lives in [neighborhood], thinks it is not very safe. on subway at night. Friends have similar concern but less intense. More than peers. X2 month after going out. Nothing specific just more (robbed)
Do you ever feel like you are being singled out or watched?
Men leering, specific to her. Does not fear harm. Trying to get with her and then cheat. Assessed above.
(…)
(…) Do you ever think you hear sounds and then realized that there is probably nothing there?
Started 2 years ago. Would hear sirens outside home but nobody is there. Could be sound traveling, could be mind, stress. Only at home. Roommate did never heard them, could not check with them since late at night. Worse since moving. Twice a week, sometimes 3. Ok with it, wakes up to check. No.
Do you ever hear your own thoughts as if they are being spoken outside your head?
Sometimes before sleep. Only falling asleep. Once every 2 month, none in past month, not always. Inner monologue (…)

Note: These notes are representative of the style and content of interview notes. These examples are fictitious to preserve participant privacy. See Supplement 1 for true, deidentified, example answers to the same question by different participants.

Risk Scores

We included two risk score measures to test whether NLP-generated risk scores or classifications could provide additional value to established predictors of psychosis-risk.

SHARP Risk Score

The ShangHai-At-Risk-for-Psychosis (SHARP) program developed a simple individualized risk calculator based on metrics obtained from assessment using the SIPS (Miller et al., 2003) to determine risk of conversion to full psychosis (Zhang et al., 2019). This risk score consists of four dimensions: 1. Functional decline (higher risk scores for participants with a drop in Global Assessment of Functioning score in the past year of 25 points or more), 2. Severity of positive symptoms (higher risk scores for participants with higher unusual thought content/delusional ideas, suspiciousness/paranoia, and disorganized communication scores), 3. Severity of Negative symptoms (higher risk scores for participants with higher social anhedonia, expression of emotion, ideational richness scores), 4. Severity of general symptoms (lower risk scores for participants with higher dysphoric mood scores). SHARP risk scores were available for 1509 participants.

NAPLS Risk Calculator

One of the outcomes from the second wave of the North American Prodrome Longitudinal Study (NAPLS-2) was the development of an individual conversion to psychosis risk calculator based on profiles of risk indicators (Cannon et al., 2016). This risk calculator also includes measures of severity of positive symptoms (higher risk associated with higher scores in the unusual thought content/delusional ideas and suspiciousness/paranoia SIPS domains) (Miller et al., 2003), and decline in functioning (greater decline in social functioning in the past year as measured by the Global Functioning: Social scale (Cornblatt et al., 2007) produce higher risk scores). In contrast to the SHARP SIPS-based risk calculator, the NAPLS risk calculator incorporates cognitive measures, namely, processing speed as measured by the Brief Assessment of Cognition in Schizophrenia symbol coding test (Keefe et al., 2004) and verbal learning as measured by the Hopkins Verbal Learning Test–Revised (Brandt, 1991); where worse cognitive performance equates to higher risk scores. This risk score calculator also includes metrics of stressful life events, measured using the Research Interview Life Events Scale (Dohrenwend et al., 1978) and the Childhood Trauma and Abuse Scale (Janssen et al., 2004), respectively, with a higher number of lifetime stressors leading to higher risk scores. Lastly, having a first-degree family member with a psychotic illness and younger age is also associated with higher risk scores for the NAPLS risk calculator. NAPLs risk scores were available for 747 participants.

Natural Language Processing Models

The following section outlines the general architecture for each of our chosen models, model parameters and each model’s exposure study data. This information is also presented as a schematic in Figure 1. Supplements 2 and 3 illustrate the type of prompt used for these models.

Trained Encoder-Only: ModernBERT

Encoder-only language models specialize in “understanding” text rather than generating it. Language inputs are tokenized and converted into detailed representations in high-dimension vector space, called embeddings, that capture the input’s meaning, structure, and context. These models are well-suited for classification tasks based on language properties. Bidirectional Encoder Representations from Transformers (BERT) models process passages of text bidirectionally (left-to-right and right-to-left) in order to capture words’ contexts more fully. ModernBERT is a state-of-the-art bidirectional encoder-only model with faster performance and better performance metrics. The base version of ModernBERT is pre-trained on 2 trillion tokens in English and uses up to 8192 tokens as context.

In order to further train ModernBERT on determining CHR status and conversion, we used 80% of the interviews for training and held-out 20% of the interviews for testing, selected at random. This process was iterated using 5-fold validation, namely the process is repeated five times so that all datapoints were part of four training sets and one testing set (see Figure 1 for a visual representation of 5-fold validation). For both training and testing, the model used to determine interviewees’ CHR status, prompts included the SIPS’s manual description of CHR syndromes and all interview questions and notes. A total of 2077 interviews were classified to determine whether they met criteria for CHR status. Out of these interviews 1715 had available SHARP or NAPLS risk scores (CHR = 773, CC = 713, HC = 229).

The model used to determine conversion status was trained and tested exclusively on interviews for CHR participants. This model was prompted by the SIPS’s description of the Presence of Psychotic Symptoms criteria and all interview questions and notes. The trained models classified a total of 801 CHR interviews using the same 5-fold validation process, 773 of which had available SHARP or NAPLS risk-score data.

For all analyses, models were evaluated and saved every 25 steps of training and processed in batches of 8 for both training and evaluation. The models' learning rate was 0.00002, and weight decay was 0.01 as per ModernBERT standard usage. Training took place over 4 epochs, namely, all inputs in the training subsample were processed 4 times. Within each fold, the code evaluates accuracy every 25 steps on the held-out data, then reloads the checkpoint that achieved the highest accuracy on that held-out set using the most accurate model for the final classification task. All analyses were performed on Northwestern’s High-Performance Computing cluster.

LLaMA 4: Few-Shot and Zero-Shot Decoder

Generalist large language models are deep learning models using transformer architecture and neural networks that are trained on an extensive and varied corpora of text with the purpose of understanding and generating text for a wider variety of tasks. Decoder-only models take a prompt and then produce tokens in response using information from the prompt and self-attention mechanisms on previously generated tokens. For the purposes of our analyses we used Large Language Model Meta AI, Scout version 4 (LLaMA 4 Scout) accessed through the open-source platform Ollama (llama4:latest; instruct version; Q4_K_M quantization,~108.6B total parameters; 10,485,760-token context window). All analyses were performed and predictions generated locally on a powerful consumer-grade laboratory computer.

We implemented both few-shot and zero shot learning for this model. For the purposes of this project, both types of models were fed prompts providing instructions and descriptions of the relevant clinical outcome (CHR status or conversion) and instructed to provide a risk score from 0 to 100% on how likely each interview notes were to belong to the respective clinical category (CHR syndrome or future conversion). Few-shot prompting for this model included providing labeled examples for notes from one clinical case (CHR or future converter) and one case that did not meet these criteria (CC or HC, or non-converting CHR). Zero-shot prompting included no labeled examples.

The prompt provided to determine risk for CHR status included the SIPS’s description of CHR syndromes, all questions and notes for a participant’s SIPS-P assessment, and an instruction to provide a percent rating of how likely the participant is to be in a CHR syndrome category. The few-shot version of this model produced a risk score for 2075 interviews, 1713 of which had SHARP or NAPLS risk scores (CHR = 772, CC = 712, HC= 229) and the zero-shot version produced risk scores for 2072 interviews, 1710 of which had SHARP or NAPLS risk scores (CHR = 769, CC = 712, HC= 229). The few-shot prompted version produced risk scores for all available interviews, but the 2 notes used as examples within the prompt were excluded. The zero-shot version prompt failed to produce risk-scores for 5 interview notes due to either model degenerate behavior or model redlining (ex. stating that the model could not be used for medical diagnosis).

The prompt for the model used to produce a risk score for conversion to psychosis included the SIPS’s description of Presence of Psychotic Symptoms criteria, all SIPS-P interview questions and notes, and an instruction to provide a percent rating of how likely the participant is to develop psychosis. This model was only used on CHR participants’ interviews. The few-shot prompt version of this model generated conversion risk scores for 796 interviews (5 missing, 2 used in the prompt, 3 due to model degenerate behavior or redlining), 768 of which had SHARP or NAPLS risk scores; the zero-shot prompt version generated risk scores for all 801 interviews, 773 of which had SHARP or NAPLS risk scores.

GPT-4.1-mini: Few-Shot and Zero-Shot Decoder

We also conducted analyses to obtain risk scores using a few-shot and zero-shot approach using Open AI’s GPT4.1 mini. For these analyses, we used OpenAI’s o4-mini model through Azure’s AI Foundry platform under Northwestern University’s Business Associate Agreement with Microsoft. This agreement provides HIPAA-compliant data storage and processing so that, for example, de-identified interview data or proprietary SIPS manual text is not used to further train OpenAI’s models.

Mirroring our analyses using LLaMA, the model used to generate risk scores for CHR risk was prompted with syndrome descriptions. For few-shot prompting we used the same CHR positive and negative cases used in LLaMA analyses as labeled examples. The zero-shot version of this prompt also included no labeled examples. Following the syndrome description and, if included, labeled examples, we also included instructions to provide a percentage rating of how likely the participant is to be in a CHR syndrome category. The few-shot version prompt of this model generated a CHR risk score for 2075 interviews, all but the excluded prompts, 1714 of which had SHARP or NAPLS risk scores (CHR = 773, CC = 712, HC= 229). The zero-shot version of this model produced a risk score for 2070 interviews, 7 interviews were not classified due to model redlining, 1710 of which had SHARP or NAPLS risk scores (CHR = 770, CC = 712, HC= 228).

The model used to generate conversion risk scores for CHR participant’s interviews included the SIPS’s description of Presence of Psychotic Symptoms criteria, interview questions and their respective notes and instruction to provide a percent rating of how likely the participant is to develop psychosis. For few-shot prompting we used the same converting and non-converting cases used in LLaMA analyses as labeled examples. The zero-shot version of this prompt also included no labeled examples. The few-shot version of this model generated a conversion risk score for 797 interviews (2 excluded since they were part of the prompt, 2 not classified due to model redlining), 769 of which had SHARP or NAPLS risk scores. The zero-shot version produced risk scores for all 801 CHR cases, 773 of which had SHARP or NAPLS risk scores.

Exploratory Analyses: Fine-Tuned Decoder - GPT-4.1-mini

Lastly, fine-tuning a decoder model refers to the customization of a pre-trained model’s parameters through training with tailored data so it performs better in a particular context, task or type of data. For the purposes of these analyses, we fine-tuned OpenAI’s GPT 4.1 mini using the platform's default hyperparamenters for batch size, learning rate, and number of epochs. Similar to the training process used for the encoder model, we used 80% of the pertinent interviews for fine-tuning and 20% for validation. This split was selected randomly. We did not perform 5-fold validation for this model given that this process was exponentially more costly than our prior GPT-based analyses. Given this reduced dataset analyses using these fine-tuned models should be considered exploratory.

The prompt in the training and validation model used to determine CHR status included the SIPS description on CHR syndromes, all interview questions and notes, as well as instructions to produce a 1 if the interview corresponded to a CHR assessment and 0 if it did not. Training labels were then provided in the format of 1 for CHR cases and 0 for CC and HC interviews. We opted for binary classification rather than a risk score for this model, as this better matched the clinical outcome data available for training. The final fine-tuned model classified 414 held-out interviews by CHR status (CHR = 159, CC = 155, HC = 100).

The fine-tuned model used to determine conversion in the CHR group was formatted in the same way but instead included the SIPS description for Presence of Psychotic Symptoms and instructions to determine whether the interviewee would (1) or would not (0) develop psychosis in the future, and labels corresponding to psychosis transition. The final fine-tuned model classified 167 held-out CHR participants’ interviews.

Statistical Analysis

All analyses were performed using R version 4.2.3 (R Core Team, 2023).

Aim 1: Model Performance Predicting CHR Status

The four models we selected were tested for their raw accuracy predicting clinical status, namely, percent of interviews classified correctly by CHR status. We also calculated Cohen’s Kappa comparing NLP-models’ performance determining CHR status to the gold standard of SIPS-based clinician assessment. For the purposes of this study, a kappa value of 0.78 is considered the gold standard benchmark as this was the established agreement between two established clinical high risk measures (SIPS and CAARMS) (Fusar-Poli et al., 2016).

The four models were also assessed for sensitivity/recall, positive predictive value/precision, and their F1 statistics. For models that produce a binary outcome, namely our encoder model and fine-tuned encoder model, raw model predictions were compared to binary clinical outcomes (1 for cases, 0 for non-CHR interviewees). We opted for a face-value approach to CHR classification for our models that produced a linear likelihood of risk rather than selecting the best performing cut-off score. Interviews where risk was predicted as 50% or higher were labeled as CHR, and those with risk scores lower than 50% were considered non-CHR interviewees.

Aim 2: NLP Metrics Relationship to Traditional Risk Scores

In order to assess the relationship between NLP-generated risk metrics and SHARP and NAPLS risk-scores we constructed mixed linear regression models. These models included sample normalized SHARP and NAPLS risk-scores as the outcome measure and included fixed effects for the NLP risk metric and assessment site, and a random intercept effect for the assessment’s interviewer. NLP risk metrics were binary for the encoder-only and fine-tuned decoder-only model (CHR = 1, Non-CHR = 0). The continuous risk-scores for few-shot and zero-shot models were converted from 0 to 100, to a 0 to 1 proportion and sample normalized for ease of comparison between models. We also tested analyses using binarized outcomes for the 0-shot and few-shot approaches to further ease model comparison. Similar to our approach for Aim 1, risk scores under 0.5 were scored as not CHR and those at or over that value scored as CHR cases. Finally, given how differences in word count between CHR and CC or HC interviews could be driving differences between groups, we included sensitivity analyses testing whether the value of LLM-produced risk scored remained after including word count as a covariate.

Exploratory Aim 3: Improving Conversion Risk Models

Lastly, we constructed binomial general linear models using conversion status as an outcome variable. These final analyses aim to determine whether the addition of NLP-based metrics can significantly improve models predicting conversions that include well-established risk metrics (i.e. the sample normalized SHARP and NAPLS risk-scores). These analyses were conducted exclusively with CHR participants’ interviews. Given that conversion to psychosis is a relatively rare event in CHR samples, and that in the present sample has less than 10 events per variable for our outcome, this aim is exploratory in nature.

Models included a binary outcome for conversion (Converted = 1, No conversion during follow ups = 0), an intercept value, interviewee’s normalized SHARP or NAPLS risk score, and a fixed effect for assessment site and a random effect for interviewer. These models were then compared to models including a fixed effect for an NLP-based metric. For our encoder and fine-tuned decoder models the NLP-based metric was binary (NLP predicted conversion =1, NLP predicted not converting = 0). For our few-shot and zero shot decoder models we used a sample normalized continuous NLP-based predictor. For ease of comparison across models, percent risk scores were converted to proportions (0 to 1); in the same spirit as for Aim 2 we also present results were risk-scores from 0-shot and few-shot approaches were binarized (risk scores < 0.5 = 0, non-converter; risk scores ≥ 0.5 = 1, converter). We also included word count sensitivity analyses for models in this aim, including word count as a covariate. To address the relative rarity of conversion outcomes, all Aim 3 models were additionally estimated using Bayesian mixed-effects logistic regression via the blme package in R (Chung et al., 2013). In this model we placed a weakly informative normal prior (SD = 2) on fixed effects. This prior regularizes coefficient estimates toward zero, reducing the risk of sparse-data inflation while remaining permissive of clinically plausible effect sizes. The parameter of interest in these models is the estimate associated with the NLP metric. Next, to assess the added predictive utility of the NLP-generated conversion probability scores and classification metrics, we performed a likelihood ratio test using ANOVA with a Chi-squared test comparing models with and without an NLP-based metric.

Transparency and Openness

Each model’s output data as well as R code used for statistical analysis are included as supplementary materials. Interview notes cannot be shared due to the sensitive nature of their content. Similarly, prompts including materials from the SIPS manual cannot be shared due to the proprietary nature of this instrument. This study was not preregistered. Preliminary results from these project were presented at the 2025 Society for Research in Psychopathology annual meeting.

Results

Aim 1: Model Performance Predicting CHR Status

In line with our hypothesis, all models correctly classified more than 50% of interviews. However, also following our hypothesis, no models had reliability of 0.78 or above. The best performing model in terms of reliability was the encoder model, which showed moderate agreement with clinician ratings of CHR (κ = 0.63). Generative models using a zero-shot (LLaMA κ = 0.34, GPT κ = 0.44) or few-shot (LLaMA κ = 0.52, GPT κ = 0.41) approach showed weak to minimal agreement with gold-standard CHR classification based on the SIPS (McHugh, 2012).

In addition to showing the best reliability and in line with our hypothesis, classification by CHR status accuracy was the highest for the encoder model (82.67%). The encoder model also outperformed other models in sensitivity/recall (0.76), which was paired with good positive predictive value/precision (0.78). Lastly, this encoder model had the best F1 statistic (0.77) out of all tested models, showcasing its comparative strengths, in line with our hypothesis.

In contrast, decoder models using a few-shot or zero-shot approach had more substantial drawbacks. In agreement with our hypothesis, both LLaMA (few-shot:77.73%, zero-shot: 72.54%) and GPT decoder models (few-shot: 75.70%, zero-shot: 74.68%) had comparatively lower accuracy. The main issue across these models was missing true CHR cases, which is reflected in their relatively low sensitivity/recall (LLaMA few-shot: 0.68, LLaMA zero-shot: 0.33, GPT few-shot: 0.42, GPT zero-shot: 0.48) and higher positive predictive value/precision (LLaMA few-shot: 0.72, LLaMA zero-shot: 0.88, GPT few-shot: 0.85, GPT zero-shot: 0.81). In line with our hypotheses, these models had worse performance than the encoder model in terms of their F1 statistic (LLaMA few-shot: 0.70, LLaMA zero-shot: 0.48, GPT few-shot: 0.56, GPT zero-shot: 0.60).

Figure 2 presents a visual summary comparing model performance for all NLP models classifying interviews by clinical high-risk status.

Figure 2. Model Performance for CHR Classification.

Figure 2.

Note: This figure showcases model performance detecting CHR cases. Panel A presents summary metrics comparing models’ performance. The remaining panels show the number of true negative and true positive cases (in green) and false positive and false negative (in red) classifications.

Exploratory analysis on a sub-set of data using a fine-tuned GPT decoder model showed weak accuracy (79.95%) and reliability (κ = 057), lower than those of the trained encoder model but higher than those for decoder models using zero or few-shot approaches. This is in line with our hypothesis that training would improve model accuracy but that encoder models would outperform decoder models. Just like with untrained decoders the main issue with this fine-tuned classification was relatively low sensitivity/recall to true CHR cases (0.69). Likewise, positive predictive value was comparatively better (0.76) as was F1 (0.70).

Aim 2: NLP Metrics Relationship to Traditional Risk Scores

In line with our hypothesis, all NLP models produced either binary outcomes or continuous risk scores that positively and significantly relate to the SHARP and NAPLS traditional risk scores. A positive CHR classification using our trained encoder model resulted in a significant increase of 0.69 standard deviations in SHARP risks scores (See Figure 3A and Table 3) and 0.85 standard deviations increase in NAPLS risk scores (See Figure 3B and Table 3) risk scores in models that account for a fixed effect of site and a random effect for interviewer. Results direction and significance remained after including normalized interview word count as a covariate (see Supplement 4).

Figure 3. BERT Classifications’ Relationship to SHARP and NAPLS Risk Scores.

Figure 3.

Note: This figure illustrates the relationship between traditional psychosis risk scores (i.e.. sample normalized SHARP and NAPLS scores) and CHR classification using an encoder BERT model. Interviews for CHR participants classified using the SIPS clinical interview are represented in blue and non-CHR participants in red.

Table 3.

Models Predicting Risk Scores

SHARP Risk Score as Outcome NAPLS Risk Score as Outcome
LLM Model LLM β [95% CI], p LLM β [95% CI], p
BERT 0.69 [0.60 – 0.79], p < .001* 0.85 [0.72 – 0.98], p < .001*
LLaMA Few-Shot 0.39 [0.34 – 0.44], p < .001* 0.45 [0.38 – 0.51], p < .001*
LLaMA Zero-Shot 0.30 [0.25 – 0.35], p < .001* 0.46 [0.40 – 0.52], p < .001*
GPT Few-Shot 0.24 [0.20 – 0.29], p < .001* 0.47 [0.41 – 0.52], p < .001*
GPT Zero-Shot 0.28 [0.23 – 0.32], p < .001* 0.47 [0.41 – 0.53], p < .001*

Note: All reported beta values are standardized. Significant results in the relevant parameters are boldened and marked with an asterisk.

All continuous risk scores generated using few-shot or zero-shot approaches significantly related to traditional risk scores. A standard deviation increase in the model’s LLM-based metric led to 0.24 to 0.47 standard deviation increases in risks scores after accounting for site and interviewer effects. Few-shot approaches with both LLaMA and GPT were significantly and positively related to both SHARP (See Figure 4A, 4C and Table 3) and NAPLS risk scores (See Figure 4B and 4D). This was also the case for results obtained using a zero-shot approach with either LLaMA or GPT for both SHARP (See Figure 4E, 4G and Table 3) and NAPLS risk scores (See Figure 4F, 4H and Table 3). All results remained significant after including normalized interview word count as a covariate (see Supplement 4).

Figure 4. NLP-based Risk Scores Relationship to SHARP and NAPLS Risk Score.

Figure 4.

Note: This figure illustrates the relationship between traditional psychosis risk scores (i.e. sample normalized SHARP and NAPLS) and sample normalized NLP-generated risk metrics. Interviews for CHR participants classified using the SIPS clinical interview are represented in blue and non-CHR participants in red.

Converting continuous risk scores produced by our few-shot and zero-shot models to binary outcomes did not change the direction or significance of effects. A positive CHR classification accounted for an increase between 0.59 and 1.16 standard deviations in risk scores among participants. This was the case for SHARP and NAPLS risk scores (see Supplement 5).

Finally, our exploratory analyzes using a fine-tuned decoder-only model produced binary outcomes that were significantly related to both SHARP (𝛽 = 0.62, 95% CI: 0.39 – 0.83, p < .001) and NAPLS (𝛽 = 1.02, 95% CI: 0.69 – 1.32, p < .001) risk scores when controlling for site and interviewer effects. Positive classifications using this model led to a 0.62 and 1.02 standard deviation increase in SHARP and NAPLS risks cores respectively. Results direction and significance remained for fine-tuned classifications when including normalized interview word count as a covariate (see Supplement 4).

Exploratory Aim 3: Improving Conversion Risk Models

Given the imbalance in conversion outcomes in the CHR sample, the following results are exploratory in nature. Likely due to this imbalance, our trained encoder-only and fine-tuned GPT models did not predict any positive conversion cases (i.e., all CHR interview notes were classified as non-conversions). Hence, we excluded these results from further exploratory analysis.

All remaining decoder models produced risk scores that both acted as significant predictors of conversion in models that included SHARP risk scores, as well as site and interviewer effects, and significantly improved model performance when compared to models that only included SHARP risk scores and covariates (see Table 4). After converting continuous risk scores to binary conversion predictions, these classifications were also significant predictors of conversion in models including SHARP risk scores (See Supplement 5). The significance and direction of results remained after including normalized interview word count as a covariate (see Supplement 4) and after applying a Bayesian mixed-effects approach using weakly informative priors to correct for sparse-data inflation (see Supplement 6).

Table 4.

Incremental Validity in Models Predicting Conversion

SHARP Risk Score Only SHARP Risk Score + LLM Score Model Comparison
LLM Model SHARP β [95% CI], p SHARP β [95% CI], p LLM β [95% CI], p Δχ2(1), p
LLaMA Few-Shot 0.21 [0.00, 0.42], p = .045 0.20 [−0.02, 0.42], p = .073 0.67 [0.27, 1.07], p = .001* 12.11, p < .001*
LLaMA Zero-Shot 0.20 [−0.01, 0.41], p = .064 0.19 [−0.03, 0.41], p = .095 0.58 [0.27, 0.88], p < .001* 13.54, p < .001*
GPT Few-Shot 0.21 [−0.00, 0.42], p = .052 0.23 [−0.01, 0.46], p = .057 0.89 [0.60, 1.18], p < .001* 40.20, p < .001*
GPT Zero-Shot 0.20 [−0.01, 0.41], p = .064 0.13 [−0.11, 0.36], p = .286 0.69 [0.42, 0.96], p < .001* 21.06, p < .001*
NAPLS Risk Score Only NAPLS Risk Score + LLM Score
NAPLS β [95% CI], p NAPLS β [95% CI], p LLM β [95% CI], p
LLaMA Few-Shot 0.29 [−0.12, 0.70], p = .169 0.21 [−0.22, 0.64], p = .335 0.50 [−0.25, 1.25], p = .188 1.89, p = .169
LLaMA Zero-Shot 0.32 [−0.06, 0.70], p = .098 0.25 [−0.15, 0.64], p = .217 0.34 [−0.16, 0.84], p = .185 1.68, p = .195
GPT Few-Shot 0.29 [−0.11, 0.69], p = .161 0.16 [−0.27, 0.58], p = .474 0.60 [0.12, 1.08], p = .014* 5.58, p = .018 *
GPT Zero-Shot 0.32 [−0.06, 0.70], p = .098 0.30 [−0.10, 0.69], p = .138 0.14 [−0.37, 0.66], p = .586 0.28, p = .598

Note: All reported beta values are standardized. Significant results in the relevant parameters are boldened and marked with an asterisk.

Results using NAPLS risk scores varied by model. According to our hypotheses, the few shot version of the GPT model produced LLM-generated risk scores that significantly predicted conversion in models including NAPLS risk scores increasing conversion by 0.60 for every standard deviation increase in risk-score. This model significantly improved conversion prediction over a model solely including NAPLS risk scores, as well as interviewer and site effects (see Table 4). The remaining few-shot and zero-shot decoder models’ risk scores did not have a significant effect in models including NAPLS risk scores (see Table 4). For conversion prediction models including NAPLS risk scores, transforming continuous risk scores into binary predictors did not change the direction or significance of results (see Supplement 5). This pattern of results was replicated when including normalized interview notes' word count as a covariate (see Supplement 4) and applying a weak-priors correction for sparse-data inflation (see Supplement 6).

Discussion

The present study showcases the potential of AI-based NLP tools to support traditional evaluations of psychosis risk. For our first aim, we found that, compared to gold standard clinical assessments, AI-based risk metrics were performing above chance but below most acceptable test metrics (Trevethan, 2017). Our findings also highlight the variability among models, showcasing how training and encoder architectures could improve model performance. While trained models performed well on held-out interview data, it is unknown whether these results would generalize to the same extent to notes collected using different notes formatting or with a different set of assessors. Additionally, our encoder model, generally better-suited for classification, outperformed our models relying on decoder architectures. Future efforts could focus on encoder models for CHR classification until they become useful and scalable clinical tools, tailoring them for accuracy and generalizability with larger and more diverse datasets. While decoder models performed worse in terms of accuracy and reliability than the encoder-only model, it is notable that our zero and few-shot decoder models performed above chance given their null or limited exposure to CHR interview data. While under-performing in terms of reliability the generalizability of these ready-to-use models is likely to translate to similar projects given their minimal exposition to CHR interview data.

When assessing the relationship between established, traditional risk scores and these AI-based metrics of risk, we found a consistent positive and significant relationship across all models for both SHARP and NAPLS risk scores. The opacity of AI models makes it difficult, or impossible, to elucidate what is being captured by AI-based NLP tools (Dobson, 2023). Previous work using natural language processing tools on CHR participants’ interview transcripts suggests that the specific semantic content of interviews, such as talking about voices and sounds, as well a formal property akin to poverty of speech, semantic density, could provide clinically relevant data (Rezaii et al., 2019). However, models could be capturing differences in assessor notes of less clinical significance, such as note length. Despite model opacity, the convergence between AI-based metrics highlights how AI-based risk measurement is capturing a similar construct as well-established demographic, symptomatic, functional, and neurocognitive risk signals.

Predicting conversion remains a pervasive challenge for researchers and clinicians working with people at heightened risk for psychosis, in part due to the relative rarity of psychosis onset during follow-up as an outcome (Salazar de Pablo et al., 2021). While this unbalance also limits the applicability of AI-based tools and can act as a barrier in training models to identify positive cases, our exploratory findings suggest that NLP tools could potentially augment traditional risk metrics. Given the limited power of our sample, these finding need to be taken solely as a proof-of-concept and need to be further tested for replicability and generalizability. In line with this concern, augmentation was only present for SHARP risk scores and not NAPLS risk-scores. While this could signal that our AI models provided a more useful complement to some clinical protocols that others, this could also be interpreted as AI-models capitalizing on noise or spurious findings. Future work can use training sets augmented for conversion either through oversampling or through statistical amplification, to enhance these measures' sensitivity to true positive cases and assess the replicability of these findings.

Our findings need to be interpreted within the bounds of some technical limitations. First, is the nature of assessor notes as a source of information. While readily available, assessor notes are a sub-par source of clinical information that is not systematic, consistent across assessors, or intended to capture what is most clinically relevant. Further, formal differences in notes for participants who are at higher risk for psychosis, such as specific assessor wording, could be captured by models - rather than clinically relevant language properties. Notably, our analyses were robust to differences in the length of interview notes. More complete textual information, such as interview transcripts, could enhance the performance of AI-based tools. Regardless, it is encouraging that models were able to capture clinically relevant information even when their input data was not optimal. In addition, since all notes were thoroughly de-identified before processing, some information could have been meaningfully changed through this process to either enhance or, more likely, hinder AI’s capacity to capture clinically relevant information after de-identification. De-identification is a necessary and unavoidable step to ensure privacy and confidentiality unless all analyses are run locally, which is often computationally costly and not feasible. Additionally, de-identification is itself a time-consuming process. However, this process requires very limited training and requires substantially less human power than any form of hand-coding for textual data. Next, models showed different rates of degenerate behavior and redlining due to sensitive content. Any model that would be expanded to future clinical use would need to perform consistently and permit clinical usage. Further, these differences also mean that models’ differences in performance could also be driven by underlying characteristics in the interviews that could be classified- rather than redlined or triggering degenerate model behavior. In a similar vein, exposure to specific interviews, such as those used as examples in few-shot learning, could also be driving model results.

While this study has considerable limitations and no models performed at a rate where we could consider using them for diagnosis as-is, the present work highlights AI-based metrics as a fruitful avenue to augment psychosis risk assessments. Future work, using more complete data as input as well as trying prompt and model tailoring can provide us with models that can improve current CHR detection protocols and conversion prediction. This work can also elucidate the type of language information that is most informative to accurately detect psychosis-risk. Perhaps more important, well-thought AI-based tools can provide the opportunity for scalable screening protocols that can be implemented with minimal personnel and staff training and reach underserved populations.

Supplementary Material

Supplemental Material
Supplemental Material 2

Acknowledgements

This research was made possible by NIMH grants R01MH1120088, R01MH120089, R01MH120090, R01MH120091, R01MH120092, and F32MH133302. We want to thank M. Zwiebach and E. Kizilbash for their careful work in data cleaning and deidentification, Northwestern’s “Bring Your Own Data” Social Sciences group for their peer support and technical troubleshooting, and A. Alliende for sharing his passion on working with artificial intelligence.

Footnotes

Conflicts of Interest

The authors have no conflicts of interest to declare.

References

  1. Addington J, Cornblatt BA, Cadenhead KS, Cannon TD, McGlashan TH, Perkins DO, Seidman LJ, Tsuang MT, Walker EF, Woods SW, & Heinssen R (2011). At Clinical High Risk for Psychosis: Outcome for Nonconverters. The American Journal of Psychiatry, 168(8), 800–805. 10.1176/appi.ajp.2011.10081191 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Brandt J (1991). The hopkins verbal learning test: Development of a new memory test with six equivalent forms. Clinical Neuropsychologist, 5(2), 125–142. 10.1080/13854049108403297 [DOI] [Google Scholar]
  3. Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, & Askell A (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901. [Google Scholar]
  4. Cannon TD, Cadenhead K, Cornblatt B, Woods SW, Addington J, Walker E, Seidman LJ, Perkins D, Tsuang M, McGlashan T, & Heinssen R (2008). Prediction of Psychosis in Youth at High Clinical Risk. Archives of General Psychiatry, 65(1), 28–37. 10.1001/archgenpsychiatry.2007.3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Cannon TD, Yu C, Addington J, Bearden CE, Cadenhead KS, Cornblatt BA, Heinssen R, Jeffries CD, Mathalon DH, McGlashan TH, Perkins DO, Seidman LJ, Tsuang MT, Walker EF, Woods SW, & Kattan MW (2016). An Individualized Risk Calculator for Research in Prodromal Psychosis. The American Journal of Psychiatry, 173(10), 980–988. 10.1176/appi.ajp.2016.15070890 [DOI] [PMC free article] [PubMed] [Google Scholar]
  6. Chung Y, Rabe-Hesketh S, Dorie V, Gelman A, Liu J (2013). “A nondegenerate penalized likelihood estimator for variance parameters in multilevel models.” Psychometrika, 78(4), 685–709. doi: 10.1007/s11336-013-9328-2. [DOI] [PubMed] [Google Scholar]
  7. Cornblatt BA, Auther AM, Niendam T, Smith CW, Zinberg J, Bearden CE, & Cannon TD (2007). Preliminary Findings for Two New Measures of Social and Role Functioning in the Prodromal Phase of Schizophrenia. Schizophrenia Bulletin, 33(3), 688–702. 10.1093/schbul/sbm029 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. DeLuca JS, Novacek DM, Adery LH, Herrera SN, Landa Y, Corcoran CM, & Walker EF (2022). Equity in Mental Health Services for Youth at Clinical High Risk for Psychosis: Considering Marginalized Identities and Stressors. Evidence-Based Practice in Child and Adolescent Mental Health, 7(2), 176–197. 10.1080/23794925.2022.2042874 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Devlin J, Chang M-W, Lee K, & Toutanova K (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein J, Doran C, & Solorio T (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 4171–4186). Association for Computational Linguistics. 10.18653/v1/N19-1423 [DOI] [Google Scholar]
  10. Dobson JE (2023). On reading and interpreting black box deep neural networks. International Journal of Digital Humanities, 5(2), 431–449. 10.1007/s42803-023-00075-w [DOI] [Google Scholar]
  11. Dohrenwend BS, Krasnoff L, Askenasy AR, & Dohrenwend BP (1978). Exemplification of a method for scaling life events: The Peri Life Events Scale. Journal of Health and Social Behavior, 19(2), 205–229. [PubMed] [Google Scholar]
  12. First MB, Williams JBW, Karg RS, & Spitzer R (2016). Structured Clinical Interview for DSM-5® Disorders—Clinician Version (SCID-5-CV). American Psychiatric Association Publishing. [Google Scholar]
  13. Fusar-Poli P, Cappucciati M, Rutigliano G, Lee TY, Beverly Q, Bonoldi I, Lelli J, Kaar SJ, Gago E, Rocchetti M, Patel R, Bhavsar V, Tognin S, Badger S, Calem M, Lim K, Kwon JS, Perez J, & McGuire P (2016). Towards a Standard Psychometric Diagnostic Interview for Subjects at Ultra High Risk of Psychosis: CAARMS versus SIPS. Psychiatry Journal, 2016, 7146341. 10.1155/2016/7146341 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Janssen I, Krabbendam L, Bak M, Hanssen M, Vollebergh W, de Graaf R, & van Os J (2004). Childhood abuse as a risk factor for psychotic experiences. Acta Psychiatrica Scandinavica, 109(1), 38–45. 10.1046/j.0001-690x.2003.00217.x [DOI] [PubMed] [Google Scholar]
  15. Keefe RSE, Goldberg TE, Harvey PD, Gold JM, Poe MP, & Coughenour L (2004). The Brief Assessment of Cognition in Schizophrenia: Reliability, sensitivity, and comparison with a standard neurocognitive battery. Schizophrenia Research, 68(2–3), 283–297. 10.1016/j.schres.2003.09.011 [DOI] [PubMed] [Google Scholar]
  16. Kotlicka-Antczak M, Podgórski M, Oliver D, Maric NP, Valmaggia L, & Fusar-Poli P (2020). Worldwide implementation of clinical services for the prevention of psychosis: The IEPA early intervention in mental health survey. Early Intervention in Psychiatry, 14(6), 741–750. 10.1111/eip.12950 [DOI] [PubMed] [Google Scholar]
  17. Lundin NB, Blouin AM, Cowan HR, Moe AM, Wastler HM, & Breitborde NJK (2024). Identification of Psychosis Risk and Diagnosis of First-Episode Psychosis: Advice for Clinicians. Psychology Research and Behavior Management, 17, 1365–1383. 10.2147/PRBM.S423865 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. McHugh ML (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282. [PMC free article] [PubMed] [Google Scholar]
  19. Miller TJ, McGlashan TH, Rosen JL, Cadenhead K, Cannon T, Ventura J, McFarlane W, Perkins DO, Pearlson GD, & Woods SW (2003). Prodromal assessment with the structured interview for prodromal syndromes and the scale of prodromal symptoms: Predictive validity, interrater reliability, and training to reliability. Schizophrenia Bulletin, 29(4), 703–715. 10.1093/oxfordjournals.schbul.a007040 [DOI] [PubMed] [Google Scholar]
  20. Mittal VA, Ellman LM, Strauss GP, Walker EF, Corlett PR, Schiffman J, Woods SW, Powers AR, Silverstein SM, Waltz JA, Zinbarg R, Chen S, Williams T, Kenney J, & Gold JM (2021). Computerized Assessment of Psychosis Risk. Journal of Psychiatry and Brain Science, 6(3), e210011. 10.20900/jpbs.20210011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. R Core Team (2023). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. url: https://www.R-project.org/. [Google Scholar]
  22. Radford A, Narasimhan K, Salimans T, & Sutskever I (2018). Improving Language Understanding by Generative Pre-Training. Open AI: Language Understanding Paper. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf [Google Scholar]
  23. Rezaii N, Walker E, & Wolff P (2019). A machine learning approach to predicting psychosis using semantic density and latent content analysis. Npj Schizophrenia, 5(1), 9. 10.1038/s41537-019-0077-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  24. Salazar de Pablo G, Radua J, Pereira J, Bonoldi I, Arienti V, Besana F, Soardo L, Cabras A, Fortea L, Catalan A, Vaquerizo-Serrano J, Coronelli F, Kaur S, Da Silva J, Shin JI, Solmi M, Brondino N, Politi P, McGuire P, & Fusar-Poli P (2021). Probability of Transition to Psychosis in Individuals at Clinical High Risk: An Updated Meta-analysis. JAMA Psychiatry, 78(9), 970–978. 10.1001/jamapsychiatry.2021.0830 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Staines L, Healy C, Coughlan H, Clarke M, Kelleher I, Cotter D, & Cannon M (2022). Psychotic experiences in the general population, a review; definition, risk factors, outcomes and interventions. Psychological Medicine, 52(15), 3297–3308. 10.1017/S0033291722002550 [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Trevethan R (2017). Sensitivity, Specificity, and Predictive Values: Foundations, Pliabilities, and Pitfalls in Research and Practice. Frontiers in Public Health, 5. 10.3389/fpubh.2017.00307 [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. van Os J, & Reininghaus U (2016). Psychosis as a transdiagnostic and extended phenotype in the general population. World Psychiatry, 15(2), 118–124. 10.1002/wps.20310 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, & Polosukhin I (2017). Attention Is All You Need (No. arXiv:1706.03762). arXiv. 10.48550/arXiv.1706.03762 [DOI] [Google Scholar]
  29. Yung AR, & McGorry PD (1996). The Prodromal Phase of First-episode Psychosis: Past and Current Conceptualizations. Schizophrenia Bulletin, 22(2), 353–370. 10.1093/schbul/22.2.353 [DOI] [PubMed] [Google Scholar]
  30. Zhang T, Xu L, Tang Y, Li H, Tang X, Cui H, Wei Y, Wang Y, Hu Q, Liu X, Li C, Lu Z, McCarley RW, Seidman LJ, Wang J, & Group, on behalf of the S. (ShangHai ARfor P. S(2019). Prediction of psychosis in prodrome: Development and validation of a simple, personalized risk calculator. Psychological Medicine, 49(12), 1990–1998. 10.1017/S0033291718002738 [DOI] [PubMed] [Google Scholar]
  31. Ziegler DM, Stiennon N, Wu J, Brown TB, Radford A, Amodei D, Christiano P, & Irving G (2020). Fine-Tuning Language Models from Human Preferences (No. arXiv:1909.08593). arXiv. 10.48550/arXiv.1909.08593 [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Material
Supplemental Material 2

RESOURCES