Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2021 Jun 1.
Published in final edited form as: Schizophr Res. 2020 Apr 1;220:141–146. doi: 10.1016/j.schres.2020.03.043

Ambulatory digital phenotyping of blunted affect and alogia using objective facial and vocal analysis: Proof of concept

Alex S Cohen a, Tovah Cowan a, Thanh P Le a, Elana K Schwartz a, Brian Kirkpatrick b, Ian M Raugh c, Hannah C Chapman c, Gregory P Strauss c
PMCID: PMC7306442  NIHMSID: NIHMS1581419  PMID: 32247747

Abstract

Negative symptoms reflect one of the most debilitating aspects of one of the most debilitating diseases known to humankind. As yet, our treatments for negative symptoms are palliative at best and our understanding of their causes is relatively superficial. To address this, we are developing objective ambulatory tools for digitally phenotyping their severity which can be used outside the confines of the traditional clinical and research settings. The present study evaluated the feasibility, reliability and validity of ambulatory vocal acoustic and facial emotion expression analysis. Videos were provided by 25 patients with schizophrenia or schizoaffective disorder and 27 nonpsychiatric controls using inexpensive, non-invasive ambulatory recording methods. Controls provided 411 video recordings, and patients provided 377 video recordings; an average of 15.22 and 14.50 per participant per group respectively. The vast majority (over 80%) of these videos were usable for analysis. An empirically-supported, limited-feature vocal (7 features) and facial (3 features) set was examined. Within participants, these features varied considerably over time, but showed moderate to good test-retest reliability in many cases once contextual factors (e.g., activity involved in at the time of testing) were accounted for. Vocal and facial features showed statistically significant convergence with a “gold standard” negative symptom measure. Ambulatory vocal/facial features were more strongly associated with engagement in social or work activities in patients than negative symptom ratings. These data support the use of ambulatory vocal/facial analytic technologies for digital phenotyping of these negative symptoms.

Keywords: ambulatory assessment, negative symptoms, digital phenotyping, vocal acoustic expression, facial emotion expression

INTRODUCTION

Negative symptoms are a major source of human suffering, disability and economic cost (Carpenter and Buchanan, 2017; Daniel, 2013; Kirkpatrick et al., 2006). As yet, our understanding of their causes, maintaining factors, and underlying pathophysiology is poor and we lack evidence-based treatments for ameliorating, curing or preventing them. A major obstacle for understanding negative symptoms involves our dependence on measuring them using clinical rating scales. While integral to understanding negative symptoms, rating scales are limited in that they reflect global impressions about a patient based on behavior and collateral information obtained during brief and constrained clinical interactions. Hence, clinical ratings likely offer insufficient “resolution” for meaningfully capturing patients’ social behavior, emotional experience and expression, motivation and related phenomena as they ebb and flow over brief periods of time and during daily life (Cohen, 2018). To observe potential treatment effects, measurement approaches with better resolution are needed. Two technological and methodological innovations hold the potential for addressing these limitations: 1) automation and objectification of certain negative symptoms using audio and video analysis (Cohen et al., 2016a, Cohen et al., 2017, 2008; Raugh et al., 2019). and 2) ambulatory data collection methods to expand the assessment domain beyond the clinical setting (Ben-Zeev, 2017; Moran et al., 2017; Strauss et al., 2018a). The present project evaluated ambulatory audio and video analysis for digital phenotyping of negative symptoms. Digital phenotyping involves the quantification of in situ phenotypes using personal digital devices (Onnela, Rauch, 2016; Torous, Onnela, Keshavan, 2017).

Automatic feature extraction from audio and video recordings of patient behavior has long been conducted to understand psychiatric symptoms and psychological functions in research (see Andreasen et al., 1981). The potential for this methodology has increased dramatically within the last decade, owing in large part to increased availability of mobile devices with audio/video recording capability and increased computational methods for feature extraction and definition. That being said, the application of these technologies for understanding negative symptoms, particularly using ambulatory data collection, has been limited. Audio analysis potentially captures a variety of vocal production, pragmatic, prosodic and physical functions/signs associated with speech that conceptually map onto alogia, blunted vocal affect and even asociality and amotivation. Video analysis captures a variety of facial expressive functions/signs, and hence, can be used to capture blunted facial affect, paucity of gestures and potentially anhedonia, asociality and amotivation.

To date, a number of studies have employed vocal analysis to understanding negative symptoms, though most of them have been applied within highly constrained laboratory and clinical settings (Cohen and Elvevåg, 2014). For example, large studies of both nonpsychiatric and psychiatric groups have revealed a limited feature set of vocal features that shows convergence with clinically-rated negative symptoms in a number of studies (Cohen et al., 2016b; Cohen et al., 2014). Video analysis has been used much less frequently in schizophrenia and negative symptom research (Cohen et al., 2017; Hamm et al., 2011; Wang et al., 2008), though there is a rich literature using various other facial behavior coding and electromyography approaches (Kring & Moran, 2008). To our knowledge, there are no studies examining patient behavior using ambulatory video recording technologies collected over multiple testing sessions. While there are studies using these technologies in nonpatient populations, there are limited studies employing facial and vocal analysis to understand psychological states. For these reasons, little is known about compliance or psychometrics of these tools for measuring psychiatric states.

The present study is part of a larger attempt to develop highly sophisticated, objective, scalable and psychometrically-supported tools to measure negative symptoms that are efficient with respect to patient and clinic/research staff temporal and financial resources. In this study, we examine feasibility (i.e., quantity of usable data procured), temporal reliability (i.e., over various self-reported activities) and convergent (i.e., relationship with ratings from a “gold-standard” clinical interview measure) and incremental (i.e., independent contribution to social/occupational functioning beyond the clinical interview-based ratings) validity. Data reflect a limited feature set extracted from automated analysis of video media procured from stable adult outpatients with schizophrenia-spectrum disorders and nonpsychiatric controls.

METHODS

Participants

Data was collected from two participant groups: (a) 25 individuals with DSM-5 (APA, 2013) diagnoses of schizophrenia (SZ; n = 8) or schizoaffective disorder (n = 17); and (b) 27 control participants free of demonstrable psychiatric history or clinical diagnosis (CN). Groups did not significantly differ on age, ethnicity, sex, or parental education; however, SZ had lower personal education than CN (see Table SS2.).

Individuals with SZ were recruited from local community outpatient mental health centers and advertisements. Clinical diagnosis was determined via the Structured Clinical Interview for DSM5 (SCID; First, et al., 2002). CN participants were recruited from the local community using posted flyers and electronic advertisements. CN had no current psychiatric diagnoses as established by the SCID-I and SCID-II (First, et al., 2002), no family history of psychosis, and were not taking psychotropic medications. All participants were free from lifetime neurological disease and substance use disorders within the last 6 months. All participants received monetary compensation for their participation and provided written informed consent for a protocol approved by the University of Georgia and Louisiana State University Institutional Review Boards.

Clinical Measures

Negative symptoms were measured using the Brief Negative Symptom Scale (BNSS) We examined five-factors identified in recent largescale, multinational factor analysis of multiple negative symptom measures (Ahmed et al., 2018; Strauss et al., 2018b): blunted affect, alogia, anhedonia, asociality, and avolition. The BNSS was selected for this study because it has received considerable psychometric support for measuring negative symptoms. The Positive and Negative Symptom Scale (PANSS) (Kay et al., 1987) was used to measure general symptoms. The Level of Function Scale (Hawk et al., 1975) was used to measure community based social and vocational functional outcomes. Clinical ratings were made by the PI or research staff trained to reliability standards (alpha >.80) on gold standard training tapes developed by GPS. Patients also completed neuropsychological testing (Wechsler Test of Adult Reading (Wechsler, 2001) and the MATRICS Consensus Cognitive Battery (Nuechterlein et al., 2008).

EMA assessment

Participants were provided with Blu Vivo 5R phones with an Android operating system and underwent EMA training. EMA training consisted of instructions on how to use the phone, how to initiate and complete surveys, how to record a video of their recent events using the app, and basic troubleshooting. Participants received a follow-up call the first day of Phase 2 to ensure surveys were delivering properly.

Surveys were preprogrammed and delivered on smartphones provided to participants using the mEMA application from Ilumivu (https://ilumivu.com/). Momentary surveys were quasi-randomly scheduled within 90 minute epochs between 9:00 and 21:00. Momentary surveys were programmed to not occur within 18 minutes of each other and no more than 3 hours apart. Each momentary survey was available for a 25-minute window; surveys became available 10 minutes before their scheduled delivery and were available for 15 minutes after delivery. Surveys recorded current contextual variables such as social interaction, activity, and location, coded as binary (i.e., yes, no) and not mutually exclusive “social”, “work”, “recreation” or “no” activity variables. Videos were recorded at the end of the surveys but were optional. For the videos, participants were instructed to record themselves while giving a step by step description of their past hour and to talk for 30 seconds. Participants were compensated $1 for each survey completed with no additional incentive for completing videos. Patients and controls completed videos for an average of 23% and 27% (standard deviations = 20% and 19% respectively) of the surveys; rates that did not significantly differ by group (t = 0.29, p = 0.56). These videos formed the basis of facial and vocal analyses used in the current study.

Facial and Vocal analysis (Figure 1)

Figure 1. Facial and vocal features examined in this study.

Figure 1.

This illustration contains a list of the features examined in this study with a brief description of each. To the right of this list is a photo of a man with the face and the mouth highlighted, to emphasize where the features are derived from.

FaceReader version 7.0 (Noldus, 2018), a commercially-available program, was used to measure facial expressions. FaceReader analyzes individual video frames for landmark features, and integrates individual features using predefined algorithms. In this study, we report data from three super-ordinate features: neutral (i.e., “neutral” valence expressions), positive (i.e., “happiness” expressions) and negative (i.e., as a sum of “anger”, “sadness”, “disgust” and “fear” expressions used in prior studies of schizophrenia-spectrum pathology (Cohen et al., 2013, 2017). Scores reflect a measure of confidence that an individual emotion “type” is being shown, ranging from 0 (not at all) to 100 (perfect match) based on predefined algorithms. Acoustic analysis was conducted using the Computerized assessment of Affect from Natural Speech (CANS; Cohen et al., 2010; Cohen et al., 2016a). Digital audio files are organized into “frames” for analysis (i.e., 100 per second). During each frame, physical properties of speech are quantified, including fundamental frequency (i.e., frequency or “pitch”) and intensity (i.e., volume). In this study, we report data for five commonly used acoustic measures derived from our prior Principal Component Analysis of 1350 adults not known to have a psychiatric diagnosis (Cohen et al., 2015) and those with SMI diagnoses using laboratory (n = 309; (Cohen et al., 2016b) and EMA technologies over a week-long epoch (n = 25; Cohen et al., 2019). To ensure adequate data for analysis, audio recordings with less than five utterances and video recordings with less than 10% of frames were analyzable by Facereader were excluded (as in Cohen et al., 2013; 2017).

Analyses

Analyses were conducted in six steps. First, we reported and compared the number and quality of videos produced by patients versus controls to evaluate feasibility of ambulatory video data collection for patient research. Second, we evaluated potential demographic and clinical confounds affecting consequent analyses. Third, we evaluated the temporal stability of our vocal/facial features using Intra-Class Correlation Coefficients (ICC). Vocal and facial expression is highly variable across contexts within individuals (as noted in at least some best practice recommendations; Gabriel, et al., 2019, and in studies examining smartphone-recorded acoustic features, e.g., Cohen et al., 2019), so we examined ICC values overall and as a function of reported activity (i.e., “yes/no” engagement in mutually-exclusive “social”, “work”, “recreational”, and “no” activities at the time of the data collection). We expected ICC values of the latter, but not the former, to be in the moderate and good ranges (i.e., 0.50 to 0.75 and 0.75 to 0.90 (Koo and Li, 2016). Fourth, we compared vocal/facial features between patients and controls using logistic regression, with all ten features entered in a single step to predict group status. Fifth, we evaluated convergence between BNSS factor scores and facial and vocal features using linear regression. All features were entered in a single step to predict one of five BNSS factor scores (entered as dependent variables). Finally, we evaluated the incremental validity of vocal/facial features beyond that of BNSS ratings for predicting work and social activities. Work and social activities were measured using a) SLOF subscale scores, and b) EMA-based self-report activities based on whether the patient was engaged in “social” or “work” activities at the time of the assessment. We used multidimensional criteria for this analysis since the SLOF shares method variance with clinical ratings (i.e., they are both based on clinical judgements from similar data collected during the same clinical interview). Analyses were conducted in R (R Core Team, 2017) using base, psych (Revelle, 2015), lme4 (Bates et al., 2015), and ICC (Wolak et al., 2012) packages. We were unable to nest data within individuals because we were modelling “2nd order” variables (e.g., diagnostic group, BNSS score). All Variance Inflation Factor scores were below 2.50 (suggesting multi-collinearity was not an issue). All vocal/facial features scores were standardized and Winsorized (i.e., values exceeding 3.50 standard deviations replaced with a value of 3.50), with the consequent skew scores being below 2.

RESULTS

Data collection tolerance and completion

Controls and patients did not significantly differ in the 411 and 377 video recordings they provided respectively; an average of 15.22 and 14.50 per participant per group. There was considerable variability in the number of video recordings completed by patients (Table SS1). The quality of the videos was similar for patients and controls, with about 68% and 69% of total video frames for each recording for patients and controls respectively being analyzable. In approximately 19% of videos (N = 147 of 788 videos), the quality was insufficient for any facial analysis. For an additional 1 audio recording, there was insufficient speech for audio analysis. In total, there were 339 and 271 valid videos available for analysis for controls and patients respectively.

Data structure and considerations (Tables SS2 & SS3)

The patient and control groups were not statistically different with respect to age or reading level (Table SS2; p’s > 0.19). Patients had significantly less education and poorer neurocognitive functioning than controls (p’s < 0.05). Demographic and neurocognitive variables were not generally significantly related with facial or acoustic features. Increasing age was associated with higher intensity perturbation (p < 0.01). Increasing education was associated with higher vocal fundamental frequency and lower shimmer (p’s < 0.01). Males versus females had lower pitch and intonation values (p’s < 0.01). Vocal and facial features were generally not highly inter-correlated (i.e., r values < 0.10; Table SS3).

Temporal stability (Table 1)

Table 1.

Intra-Class Correlation Coefficients, computed as a function of group (i.e., patient, control) and while engaged in various activities.

Feature Controls Patients While Working While in Recreation While Socializing Doing “Nothing”
Facial Expression: Neutral 0.22 0.18 0.28 0.20 0.35 0.33
Facial Expression: Positive 0.31 0.46 0.43 0.55 0.46 −0.01
Facial Expression: Negative 0.28 0.21 0.35 0.26 0.30 0.64
Vocal Production: M Pause Time 0.00 0.36 0.12 0.08 0.10 0.76
Vocal Production: N Utterances 0.35 0.41 0.49 0.68 0.27 0.91
Vocal Variability: Pitch 0.48 0.77 0.81 0.71 0.34 0.34
Vocal Variability: Intonation 0.44 0.46 0.42 0.29 0.39 0.50
Vocal Variability: Jitter 0.25 0.68 0.78 0.44 0.14 0.50
Vocal Variability: Emphasis 0.12 0.33 0.20 0.40 0.30 0.54
Vocal Variability: Shimmer 0.14 0.52 0.37 0.64 0.26 0.38
K samples 339 271 131 56 86 56

Notes: Values in the moderate range or higher (i.e., exceeding 0.50 (Koo and Li, 2016) are boldfaced).

ICC values for the patients (range = 0.18 to 0.46) and controls (range = 0.00 to 0.35) were both low for the vocal production and the facial expression features. ICC values were similarly low for the vocal variability features for controls (range = 0.12 to 0.48), but were more consistent for patients for some variables (range = 0.33 to 0.77). When ICC values were computed as a function of reported activity at the time of the assessment, greater stability was observed. For example, number of utterances, pause times and negative facial expressions were highly stable when participants reported doing “nothing”.

Group differences (Tables 2, SS2)

Table 2.

Logistic regressions predicting patient (1; N = 25; K = 271) versus control (0; N = 27; K = 339) group status as a function of ambulatory vocal/facial features

Chi-Square Log Likelihood Pseudo-R2, a

Model Summary 179.94* 89.97* 0.34/0.26

Video/Audio features vif Coefficient z value

Facial Expression: Neutral 2.26 −0.67 (0.15) 4.54*
Facial Expression: Positive 1.84 −0.67 (0.14) 4.91*
Facial Expression: Negative 2.00 −0.42 (0.14) 3.09*
Vocal Production: Mean Pause Time 1.05 −0.01 (0.08) 0.18
Vocal Production: N Utterances 1.14 0.40 (0.11) 3.68*
Vocal Variability: Pitch 1.48 −1.24 (0.13) 9.60*
Vocal Variability: Intonation 1.07 −0.12 (0.10) 1.24
Vocal Variability: Jitter 1.80 −0.75 (0.14) 5.23*
Vocal Variability: Emphasis 1.62 0.71 (0.13) 5.59*
Vocal Variability: Shimmer 1.80 −0.05 (0.13) 0.37

Notes:

a

Nagelkerke and Snell values respectively.

*

= p < 0.05

Collectively, the facial and acoustic feature sets significantly discriminated patient and control groups using logistic regression (log likelihood = 89.97, X2 = 179.94, p’s < 0.001), and were associated with pseudo-R2 values of 34% and 26% of the variance using Nagelkerke and Snell estimates respectively (Table 2). Inspection of coefficient weights suggested that patients had lower neutral, positive and negative facial expressions, more utterances, lower pitch, less jitter and more emphasis. Descriptive statistics are included in Table SS2.

Convergence with clinical ratings (Table 3, Figure 1)

Table 3.

Linear regressions predicting clinical ratings as a function of ambulatory vocal/facial features, for patients only (N = 25, K = 271).

Alogia Blunted Affect Anhedonia Avolition Asociality
ΔF 3.03* 6.80* 4.63* 5.43* 4.78*

ΔR (ΔR2) Adjusted 0.26 (0.07) 0.42 (0.18) 0.35 (0.12) 0.37 (0.14) 0.35 (0.12)

Features
  Facial Expression: Neutral −0.08 (0.08) 0.03 (0.08) 0.17 (0.08)* −0.10 (0.08) 0.19 (0.08)*
  Facial Expression: Positive −0.20 (0.07)* −0.08 (0.07) 0.20 (0.07)* −0.14 (0.07)* 0.29 (0.07)*
  Facial Expression: Negative 0.08 (0.08) 0.26 (0.07)* 0.23 (0.07)* 0.12 (0.07) 0.21 (0.07)*
  Vocal Production: M Pause Time 0.09 (0.07) 0.24 (0.07)* 0.29 (0.07)* 0.17 (0.07)* 0.15 (0.07)*
  Vocal Production: N Utterances −0.17 (0.07)* −0.06 (0.07) 0.12 (0.07) 0.09 (0.07) −0.09 (0.07)
  Vocal Variability: Pitch −0.08 (0.08) 0.24 (0.07)* 0.17 (0.08)* 0.20 (0.08)* 0.05 (0.08)
  Vocal Variability: Intonation −0.15 (0.06)* −0.10 (0.06) −0.05 (0.06) −0.20 (0.06)* −0.14 (0.06)*
  Vocal Variability: Jitter −0.06 (0.09) 0.05 (0.09) 0.02 (0.09) 0.14 (0.09) −0.12 (0.09)
  Vocal Variability: Emphasis −0.14 (0.07) −0.02 (0.07) −0.08 (0.07) −0.23 (0.07)* 0.07 (0.07)
  Vocal Variability: Shimmer 0.12 (0.09) −0.06 (0.08) 0.18 (0.09)* 0.2 (0.08)* 0.18 (0.09)*

Notes:

*

= p < 0.05

The facial and vocal feature sets showed significant convergence with all negative symptom ratings (F values = 3.03 to 6.80, p’s < 0.05). The variance explained was highest for blunted affect ratings (18%), with a more modest amount explained by anhedonia, avolition and asociality (12% to 14%) and the least explained by alogia ratings (7%). Evaluation of the coefficient weights suggested there was convergence between vocal/facial features and clinical ratings, though this varied as a function of both feature and rating. For example, facial expression features showed divergent patterns in alogia and avolition versus blunted affect, anhedonia and asociality ratings. That is, alogia and avolition were associated with decreased positive expressions whereas anhedonia and asociality were associated with increased negative and positive expressions. Moreover, features generally, but not always, significantly corresponded to their conceptually-related clinical ratings. For example, increased pause times were related to each of the clinically-rated negative symptoms except for alogia, whereas number of utterances was associated only with alogia. Similarly, vocal variability features converged with a variety of negative symptom ratings, though generally, not with clinically-rated blunted affect. Correlations supporting these analyses are included in Figure SS4.

Incremental validity (Table 4)

Table 4.

Regressions comparing negative symptom ratings versus vocal/facial features in predicting work and social functioning from Clinical Interview and EMA-based assessments.

Dependent Measure ΔR (ΔR2): NEGATIVE SYMPTOM RATINGS ΔR (ΔR2): VOCAL/FACIAL FEATURES
Interview: Work 0.48 (0.23) * 0.25 (0.06) *
Interview: Social 0.57 (0.33) * 0.40 (0.16) *
EMA: Social 0.14 (0.02) * 0.30 (0.09) *
EMA: Work 0.14 (0.02) * 0.35 (0.12) *

Notes: Negative symptoms ratings and vocal/facial features entered independently, and in varying orders to compute ΔF and ΔR2 Statistics; Analyses are for patients only (N = 25; K = 222) group; ΔR/ΔR2 values adjusted for the number of predictors, k = 123, k = 47

*

= p < .05

Using linear regression models, clinical ratings (BNSS total scores) and ambulatory facial and vocal features collectively explained 48% and 30% of variance in SLOF work and social functioning scores respectively. While the majority of this reflected negative symptom ratings (not surprising due to method variance), ambulatory audio/video features explained significant additional variance beyond this. Using logistic regressions, predicting activity at the time of assessment, a significant, but comparatively more modest amount of variance was explained by BNSS and audio/video features (5% to 15% overall). For both social and work activities, the vast majority of this variance reflected the audio and video features (over 80%). The BNSS explained very little variance with social or work engagement (~ 2%). In sum, ambulatory facial/vocal features made important incremental contribution to both interview and EMA-based measures of activity (Table 4).

DISCUSSION

The present study evaluated automated ambulatory facial and vocal analysis for measuring negative symptoms. To our knowledge, this study is the first to analyze video data collected from schizophrenia patients using mobile phones and is one of the largest of its kind examining patient behavior based on video analysis using any methodology. Feasibility for this methodology was established in that a) the total number did not differ between patients and controls and b) the vast majority of the videos submitted were usable for analysis. Psychometric characteristics of the facial and vocal analyses were favorable. Although not formally examined here, the “inter-rater” reliability of video/audio analysis using the same video, software and software parameters is typically near perfect (Cohen, 2018). Temporal stability varied, but was moderate to good in many cases once context (e.g., participant activity engaged in while speaking) was accounted for. Facial and vocal features showed statistically significant convergence with a “gold standard” negative symptom measure, and also showed unique contributions to variance in functioning/activities above and beyond that associated with this “gold standard” measure.

Consistent with other studies using EMA to understand schizophrenia symptoms (e.g. Ben-Zeev et al., 2009; Moran et al., 2017; Strauss et al., 2018a), we found that behaviors underlying negative symptoms were highly dynamic across testing sessions within individuals. This is in contrast to findings that symptoms measured using traditional clinical ratings are highly stable over time (Kay et al., 1987; Kirkpatrick et al., 2011). This potential discrepancy in temporal stability between clinical ratings and EMA/objective technologies is not surprising given their fundamental differences in “resolution” – defined as the ability to detect changes. Clinical ratings offer fairly low resolution in terms of being able to detect symptom changes over time and context given their ordinal scaling and their ambiguously and variably-defined assessment window. Clinical ratings of blunted facial affect, for example, can conceivably vary as a function of assessment window (e.g., interview lengths of 10 minutes versus several hours) and interviewer demographic/interpersonal characteristics (e.g., perceived “warmth” of interviewer). This is not stated to demean the potential utility of clinical rating scales, as their precision has been adequate for detecting changes over weeks and months from many psychosocial and pharmacological interventions. Nonetheless, ambulatory vocal/facial technologies offer the ability to precisely isolate changes over user-defined temporal or contextual epochs (Cohen, 2018). From this perspective, it is not surprising that clinical ratings were relatively unrelated to social/work activities when assessed using EMA technologies – as these reflect a much higher resolution and different context than that covered using clinical ratings. While further work remains in optimizing the reliability and validity of vocal and facial analysis given their inherently dynamic nature, the present study reflects an important step in this process.

Supporting the convergent and criterion validity of our ambulatory technologies, vocal/facial features showed convergence with traditional interview-based negative symptom ratings. In a few important cases, these relationships were counter-intuitive. For example, average pause length – a measure conceptually related to alogia was significantly related to each of the clinically-rated negative symptoms except for alogia, whereas number of utterances was associated only with alogia. Similarly, clinically-rated blunted affect was not associated with flatter vocal expression (e.g., intonation, emphasis), but was associated with increasing negatively and positively valenced facial expressions. The reasons for these discrepancies are unclear (and potentially echo findings in Brown et al, 2007), but it does seem reasonable to conclude that clinically-rated symptoms do not necessarily reflect abnormalities in all relevant behavioral features at all times. Clinical symptoms can manifest across a myriad of potentially independent behaviors, and these behaviors can be quantified using a myriad of features. Consider that the winner of the 2013 INTERSPEECH competition predicting psychological states using machine learning of acoustic speech features employed 6,373 vocal features (Schuller et al., 2014). While many of these features are redundant with each other, many are physiologically and functionally distinct. Hence, it stands to reason that a particular patient on a particular occasion may express alogia through long pauses, few utterances, both or neither. This highlights another limitation of clinical rating scales, as they are unable to effectively distinguish between these various features; and highlights the computational challenge in integrating these voluminous streams of vocal/facial features together for measuring negative symptoms.

The present project was limited in some key regards. First, while the total number of videos examined was quite large for a study of this kind, the number of participants recruited was modest. Thus, phenotypic heterogeneity of schizophrenia many not have been fully captured. Second, the number of videos produced by each participant varied considerably. This may be because participants were not compensated for providing videos. Our compliance rates likely resemble those seen in the outpatient treatment setting, at least, insofar as monetary compensation is unlikely in these settings for most patients. While compliance did not vary as a function of obvious clinical or demographic characteristics, it is possible that certain clinical characteristics were under or over-represented in our videos. Understanding factors affecting compliance is a critical issue for future research. Third, the patients were medicated and psychiatrically stable. While typical for this kind of research, it is possible that the results would differ for unmedicated patients. Fourth, we examined summary vocal and facial features in this study, and there are countless other features that could be examined in future research. These limitations notwithstanding, objective vocal and facial analysis using ambulatory technologies appears a promising means of measuring negative symptoms.

Supplementary Material

1

Acknowledgements

The authors would like to thank the study participants and lab members for their aid in collecting and processing data.

Role of funding source

This research was funded by NIH grant R21 MH112925.

Footnotes

Conflicts of interest

The authors report no conflict of interest.

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  1. Ahmed AO, Kirkpatrick B, Galderisi S, Mucci A, Rossi A, Bertolino A, Rocca P, Maj M, Kaiser S, Bischof M, Hartmann-Riemer MN, Kirschner M, Schneider K, Garcia-Portilla MP, Mane A, Bernardo M, Fernandez-Egea E, Jiefeng C, Jing Y, Shuping T, Gold JM, Allen DN, Strauss GP (2018). Cross-cultural Validation of the 5-Factor Structure of Negative Symptoms in Schizophrenia. Schizophr. Bull. 45(2), 305–314. 10.1093/schbul/sby050 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Andreasen NC, Alpert M, Martz MJ (1981). Acoustic Analysis: An Objective Measure of Affective Flattening. Arch. Gen. Psychiatry. 38(3), 281–285. 10.1001/archpsyc.1981.01780280049005 [DOI] [PubMed] [Google Scholar]
  3. American Psychiatric Association; (2013). Diagnostic and statistical manual of mental disorders (5th ed.). Arlington, VA: 10.1176/appi.books.9780890425596.744053 [DOI] [Google Scholar]
  4. Bates D, Mächler M, Bolker BM, & Walker SC (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software. 10.18637/jss.v067.i01 [DOI] [Google Scholar]
  5. Ben-Zeev D (2017). Technology in Mental Health: Creating New Knowledge and Inventing the Future of Services. Psychiatr. Serv. 68, 107–108. 10.1176/appi.ps.201600520 [DOI] [PubMed] [Google Scholar]
  6. Ben-Zeev D, Young MA, Madsen JW (2009). Retrospective recall of affect in clinically depressed individuals and controls. Cogn. Emot. 23, 1021–1040. 10.1080/02699930802607937 [DOI] [Google Scholar]
  7. Brown LH, Silvia PJ, Myin-Germeys I, & Kwapil TR (2007). When the need to belong goes wrong: The expression of social anhedonia and social anxiety in daily life. Psychological Science, 18(9), 778–782. [DOI] [PubMed] [Google Scholar]
  8. Calamia MR (2018). Practical Considerations for Evaluating Reliability in Ambulatory Assessment Studies. Psychol. Assess. In Press. [DOI] [PubMed] [Google Scholar]
  9. Carpenter WT, Buchanan RW (2017). Negative Symptom Therapeutics. Schizophr. Bull. 43, 681–682. 10.1093/schbul/sbx054 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Cohen AS, Alpert M, Nienow TM, Dinzeo TJ, Docherty NM (2008). Computerized measurement of negative symptoms in schizophrenia. J. Psychiatr. Res. 42, 827–836. 10.1016/j.jpsychires.2007.08.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Cohen AS, Lee Hong S, Guevara A (2010). Understanding emotional expression using prosodic analysis of natural speech: Refining the methodology. J. Behav. Ther. Exp. Psychiatry 41, 150–157. 10.1016/j.jbtep.2009.11.008 [DOI] [PubMed] [Google Scholar]
  12. Cohen AS, Morrison SC, Callaway DA (2013). Computerized facial analysis for understanding constricted/blunted affect: Initial feasibility, reliability, and validity data. Schizophr. Res. 148, 111–116. 10.1016/j.schres.2013.05.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Cohen AS, Elvevåg B (2014). Automated computerized analysis of speech in psychiatric disorders. Curr. Opin. Psychiatry 27, 203–209. 10.1097/YCO.0000000000000056 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Cohen AS, Mitchell KR, Elvevåg B (2014). What do we really know about blunted vocal affect and alogia? A meta-analysis of objective assessments. Schizophr. Res. 159, 533–538. 10.1016/j.schres.2014.09.013 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Cohen AS, Dinzeo TJ, Donovan NJ, Brown CE, Morrison SC (2015). Vocal acoustic analysis as a biometric indicator of information processing: Implications for neurological and psychiatric disorders. Psychiatry Res. 226, 235–241. 10.1016/j.psychres.2014.12.054 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Cohen AS, Mitchell KR, Docherty NM, Horan WP (2016a). Vocal expression in schizophrenia: Less than meets the ear. J. Abnorm. Psychol. 125, 299–309. 10.1037/abn0000136 [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Cohen AS, Renshaw TL, Mitchell KR, Kim Y (2016b). A psychometric investigation of “macroscopic” speech measures for clinical and psychological science. Behav. Res. Methods 48, 475–486. 10.3758/s13428-015-0584-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Cohen AS, Mitchell KR, Strauss GP, Blanchard JJ, Buchanan RW, Kelly DL, Gold J, McMahon RP, Adams HA, Carpenter WT (2017). The effects of oxytocin and galantamine on objectively-defined vocal and facial expression: Data from the CIDAR study. Schizophr. Res. 188, 141–143. 10.1016/j.schres.2017.01.028 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Cohen AS ( 2018). Advancing ambulatory biobehavioral technologies beyond “proof of concept”: Introduction to the special section. Psychol. Assess. In Press. [DOI] [PubMed] [Google Scholar]
  20. Cohen AS, Fedechko TL, Schwartz EK, Le TP, Foltz PW, Bernstein J, Cheng J, Holmlund TB, Elvevåg B (2019). Ambulatory vocal acoustics, temporal dynamics, and serious mental illness. J. Abnorm. Psychol. 128(2), 97 10.1037/abn0000397 [DOI] [PubMed] [Google Scholar]
  21. Daniel DG (2013). Issues in Selection of Instruments to Measure Negative Symptoms. Schizophr. Res. 150(2–3), 343–345. 10.1016/j.schres.2013.07.005 [DOI] [PubMed] [Google Scholar]
  22. First MB, Spitzer RL, Gibbon M, and Williams JB. (2002). Structured Clinical Interview for DSM-IV-TR Axis I Disorders-Patient Edition (SCID-I/P, 1/2007 revision), in: Biometrics Research. [Google Scholar]
  23. Gabriel AS, Podsakoff NP, Beal DJ, Scott BA, Sonnentag S, Trougakos JP, & Butts MM (2019). Experience sampling methods: A discussion of critical trends and considerations for scholarly advancement. Organizational Research Methods, 22(4), 969–1006. [Google Scholar]
  24. Hamm J, Kohler CG, Gur RC, Verma R (2011). Automated Facial Action Coding System for dynamic analysis of facial expressions in neuropsychiatric disorders. J. Neurosci. Methods. 200(2), 237–256. 10.1016/j.jneumeth.2011.06.023 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Hawk AB, Carpenter WT, Strauss JS (1975). Diagnostic Criteria and Five-Year Outcome in Schizophrenia: A Report From the International Pilot Study of Schizophrenia. Arch. Gen. Psychiatry. 32(3), 343–347. 10.1001/archpsyc.1975.01760210077005 [DOI] [PubMed] [Google Scholar]
  26. Kay SR, Fiszbein A, Opler LA (1987). The Positive and Negative Syndrome Scale (PANSS) for schizophrenia. Schizophr. Bull. 13, 261–276. [DOI] [PubMed] [Google Scholar]
  27. Kirkpatrick B, Fenton WS, Carpenter WT, Marder SR (2006). The NIMH-MATRICS consensus statement on negative symptoms, Schizophr. Bull. 32(2), 214–219.. 10.1093/schbul/sbj053 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Kirkpatrick B, Mucci A, Galderisi S (2017). Primary, Enduring Negative Symptoms: An Update on Research. Schizophr. Bull. 43, 730–736. 10.1093/schbul/sbx064 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Kirkpatrick B, Strauss GP, Nguyen L, Fischer BA, Daniel DG, Cienfuegos A, Marder SR (2011). The brief negative symptom scale: Psychometric properties. Schizophr. Bull. 37, 300–305. 10.1093/schbul/sbq059 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Koo TK, Li MY (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. J. Chiropr. Med. 15(2), 155–163. 10.1016/j.jcm.2016.02.012 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Llabre MM, Ironson GH, Spitzer SB, Gellman MD, Weidler DJ (1988). How Many Blood Pressure Measurements are Enough?: An Application of Generalizability Theory to the Study of Blood Pressure Reliability. Psychophysiology 25, 97–106. 10.1111/j.1469-8986.1988.tb00967.x [DOI] [PubMed] [Google Scholar]
  32. Moran EK, Culbreth AJ, Barch DM (2017). Ecological momentary assessment of negative symptoms in schizophrenia: Relationships to effort-based decision making and reinforcement learning. J. Abnorm. Psychol. 126, 96–105. 10.1037/abn0000240 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Noldus (2018) Facial expression recognition software : FaceReader [WWW Document]. Noldus Int. Technol. 10.1103/Physics.4.30 [DOI] [Google Scholar]
  34. Nuechterlein KH, Green MF, Kern RS, Baade LE, Barch DM, Cohen JD, Essock S, Fenton WS, Frese FJ, Gold JM, Goldberg T, Heaton RK, Keefe RSE, Kraemer H, Mesholam-Gately R, Seidman LJ, Stover E, Weinberger DR, Young AS, Zalcman S, Marder SR (2008). The MATRICS consensus cognitive battery, part 1: Test selection, reliability, and validity. Am. J. Psychiatry 165, 203–213. 10.1176/appi.ajp.2007.07010042 [DOI] [PubMed] [Google Scholar]
  35. Onnela JP, & Rauch SL (2016). Harnessing smartphone-based digital phenotyping to enhance behavioral and mental health. Neuropsychopharmacology, 41(7), 1691. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. R Core Team (2017). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria: URL https://www.R-project.org/. [Google Scholar]
  37. Revelle W (2015). Package “psych” - Procedures for Psychological, Psychometric and Personality Research. R Package. [Google Scholar]
  38. Raugh IM, Chapman HC, Bartolomeo LA, Gonzalez C, & Strauss GP (2019). A comprehensive review of psychophysiological applications for ecological momentary assessment in psychiatric populations. Psychological assessment, 31(3), 304. [DOI] [PubMed] [Google Scholar]
  39. Schuller B, Steidl S, Batliner A, Epps J, Eyben F, Ringeval F, Marchi E, Zhang Y, (2014). The INTERSPEECH 2014 computational paralinguistics challenge: Cognitive & physical load, in: Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH. [Google Scholar]
  40. Strauss GP, Cohen AS (2017). A Transdiagnostic Review of Negative Symptom Phenomenology and Etiology. Schizophr. Bull. 43(4), 712–719. 10.1093/schbul/sbx066 [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Strauss GP, Frost Visser K, Esfahlani FZ, Sayama H (2018a). An ecological momentary assessment evaluation of emotion regulation abnormalities in schizophrenia. Psychol. Med. 48(14), 2337–2345 10.1017/S0033291717003865 [DOI] [PubMed] [Google Scholar]
  42. Strauss GP, Nuñez A, Ahmed AO, Barchard KA, Granholm E, Kirkpatrick B, Gold JM, Allen DN (2018b). The Latent Structure of Negative Symptoms in Schizophrenia. JAMA Psychiatry. 75(12), 1271–1279. 10.1001/jamapsychiatry.2018.2475 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Torous J, Onnela JP, & Keshavan M (2017). New dimensions and new tools to realize the potential of RDoC: digital phenotyping via smartphones and connected devices. Translational psychiatry, 7(3), e1053–e1053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Wang P, Barrett F, Martin E, Milonova M, Gur RE, Gur RC, Kohler C, Verma R (2008). Automated video-based facial expression analysis of neuropsychiatric disorders. J. Neurosci. Methods. 168(1), 224–238. 10.1016/j.jneumeth.2007.09.030 [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Wechsler D, 2001. Wechsler Test of Adult Reading: WTAR. Springer; 10.1007/978-1-4419-1698-3_257 [DOI] [Google Scholar]
  46. Wolak ME, Fairbairn DJ, & Paulsen YR (2012). Guidelines for estimating repeatability. Methods in Ecology and Evolution 10.1111/j.2041-210X.2011.00125.x [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

1

RESOURCES