Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2011 Aug 6.
Published in final edited form as: Vision Res. 2010 May 25;50(17):1693–1711. doi: 10.1016/j.visres.2010.05.024

A Computational Shape-based Model of Anger and Sadness Justifies a Configural Representation of Faces

Donald Neth 1, Aleix M Martinez 1
PMCID: PMC2912412  NIHMSID: NIHMS216950  PMID: 20510267

Abstract

Research suggests that configural cues (second-order relations) play a major role in the representation and classification of face images; making faces a “special” class of objects, since object recognition seems to use different encoding mechanisms. It is less clear, however, how this representation emerges and whether this representation is also used in the recognition of facial expressions of emotion. In this paper, we show how configural cues emerge naturally from a classical analysis of shape in the recognition of anger and sadness. In particular our results suggest that at least two of the dimensions of the computational (cognitive) space of facial expressions of emotion correspond to pure configural changes. The first of these dimensions measures the distance between the eyebrows and the mouth, while the second is concerned with the height-width ratio of the face. Under this proposed model, becoming a face “expert” would mean to move from the generic shape representation to that based on configural cues. These results suggest that the recognition of facial expressions of emotion shares this expertise property with the other processes of face processing.

1. INTRODUCTION

Emotions play a major role in human behavior and are major factors for the study of evolution, cognition, and consciousness (Izard, 2009). A fundamental problem in the study of emotions, and of cognitive science, is to uncover the dimensions of the computational space of facial expressions of emotion. Facial expressions of emotion are generated by moving each of the facial muscles of the faces to certain positions (Duchenne, 1862; Darwin, 1872; Ekman & Friesen, 1978). This causes the facial features and skin of the face to move and deform in ways that observers then interpret as an expressed emotion. Each muscle, or group of muscles, employed to create these constructs is referred to as an Action Unit (AU). Ekman and Friesen (1978) identified those AUs responsible to generate the most common (sometimes called universal (Ekman, 1992), due to their production similarity across cultures) emotions. The generative function of a facial expression of an emotion is thus a combination of AUs; although different sets of AUs may express the same emotion, each set is generally associated to a unique emotion.

How does the brain interpret (i.e. represent, analyze and classify) the image of a facial expression then? Since one uses a combination of AUs to produce expressions of emotion, it seems natural to assume that observers may try to resolve the inverse problem when analyzing an image; i.e., determining which AUs are used to produce the expression. If this hypothesis were true, the inverse process described above would help us understand what/how others feel (Preston & de Waal, 2002), which is related to the theory of mind (Carruthers & Smith, 1996) and mirror neurons (Gallese & Goldman, 1998; Rizzolatti et al. 2001). In this model, the dimensions of the computational (or cognitive) space defining facial expressions of emotion correspond to each of the AUs (Pantic & Bartlett, 2008). The amount of movement in each AU (i.e., the value about each dimension of the computational space) can be used to interpret the intensity of the expressed emotion. However, some neurophysiological studies suggest a dissociation of the processes of empathy and theory of mind (Vollm et al., 2006). Hence, one may not necessarily need to resolve the inverse problem for recognizing emotions, even if these are then feed into or shared with other cognitive processes.

Furthermore, uncovering the AUs being used by another person from an image is not an easy task. Imaging the image shown in Figure 1a. This expression includes all the AUs necessary to produce the (genuine) perception of anger. According to the model described in the preceding paragraph, to interpret this emotion correctly, we need to determine which AUs are active and by how much. One option is to learn which texture patterns in the face image are most descriptive for each AU, yet are not present in other AUs. For example, AU 4 creates a wrinkle of the skin around the middle section of the forehead, which, if visible, can be readily detected. Research in computer vision has studied this approach, for the detection of several of the AUs (Pantic & Bartlett, 2008). These computational studies have also demonstrated that to detect the textural patterns indicating the movement of an AU, high resolution or high quality images are necessary. Yet, this does not seem to be a problem for human observers. As illustrated in Figure 1b–c, we can recognize emotions in low quality, in low resolution, when the texture pattern caused by the AU is not visible (due to, for example, illumination), or even in sketches (where there is no texture present). To do recognition in these impoverished stimuli, the human visual system needs to make use of simple image cues, which ought to be: i) easily extractable from images, and ii) robust to image manipulation. Furthermore, the behavior of and predictions drawn from the computational model that uses these image features need to correlate with those observed with human subjects. Wilbraham et al. (2008) have shown that even where textural features can achieve good classification results, these do not necessarily correlate with the results obtained with human subjects. Figure 1b–c is a clear example of this – these images would not be correctly classified by a texture-based model, but are readily interpreted by humans.

Figure 1.

Figure 1

Variability of images showing angry expressions. a) An angry face; from Ekman & Friesen (1978). Note that each AU creates a clearly identifiable texture pattern. b) A low resolution image of the same angry faces. In this image, the AUs are not clearly identifiable based on their associated texture patterns. c) A sketch illustrating an angry expression. Here, there is no texture. Recognition of emotions must be done using other image cues. d,e) Subject responses from the study of Neth & Martinez (2009). Images were shown in pairs. Images could correspond to the original face, or a modified face with a reduced or enlarged eyebrow-mouth distance (example images are in Figure 2). The abscissa in the plots above reflects the difference in positive or negative change level between the first and second image shown to the subject. For example, if the first image represents a configural change of 75% (increase with respect to the original) and the second image is at a level of 25%, the subject's response would fall into the −50% grouping (negative change). The ordinate reflects the percentage of less, same, and more responses made at each grouping, which were the three response options given to the subjects. Positive changes led to an increase of the perception of sadness/anger. Negative changes to a decrease of that perception.

A large number of research studies suggest that faces are process differently than other objects (Diamond & Carey, 1986; Bruce & Young, 1986; Tanaka & Farah, 1993; Moscovitch et al., 1997; Leder & Bruce, 2000), with neurophysiological studies seemingly confirming this view (Tsao et al., 2006). Most studies have identified configural cues as the predominant features underlying face recognition. Bouvrie and Sinha (2007) have recently shown that to construct a computational model that emulates recognition in congenitally blind children, configural cues also play a major role. These studies generally addressed the identity problem in face recognition though. In a recent paper (Neth & Martinez, 2009), we have shown that the vertical distance between the eyebrows and the mouth biases the perception of anger and sadness. When this distance is made shorter than “normal,” faces are perceived as angry, Figure 2a,c,e. When this same distance is made longer than average, the perception of sadness increases, Figure 2b,d,f. By “normal” (and average) we refer to the normative distance within a population, i.e., the average brow-mouth distance in the population present in our environment (Valentine, 1991). This result suggests that configural cues (i.e., second-order relations) compose (at least) part of the computational face space representing facial expressions of emotion. Figure 1d–e summarizes the results of Neth & Mertinez (2009). In this study, images were shown in pairs. The relationships among internal facial features in these images were manipulated without affecting the overall shape of the face or head, resulting in a pure configural change defined by the eyebrow to mouth distance, Figure 2. The x-axis in the plots in Figure 1d–e defines the difference in brow-mouth distance between the first and second images shown to the subject. At 0%, the eyebrow-mouth distances in the pair are identical. In Figure 1d, the positive percentages correspond to an increase of the distance in the second displayed image. In Figure 1e it corresponds to a decrease. The reverse applies to the negative percentages. We see that as the configural change increases, the perception of sadness and anger grows. As the configural change is deemphasized, the perception of the emotion diminishes. In summary, an increase of the eyebrow to mouth distance leads to the perception of sadness, while a decrease of this distance leads to the perception of anger – even in otherwise neutral expressions. This suggests that facial expressions of emotions are (at least in part) also directed by configural changes. Note that these configural cues are directly related to AUs. In particular, moving the inner corners of the eyebrows is controlled by AUs 1 and 4, which are used to express anger and sadness, respectively.

Figure 2.

Figure 2

Samples of each type of configural displacement. a) Eyes-Brows Down; b) Eyes-Brows Up; c) Nose-Mouth Up; d) Nose Mouth Down; e) Nose-Half-Mouth-Full Up; f) Nose-Half-Mouth-Full Down.

Configural cues seem ideal for the analysis of facial expressions, because they can be readily obtained from almost any resolution or when the texture pattern caused by a given AU is not visible. Note how the above defined feature (eyebrow-to-mouth distance) suffices to successfully classify all the images in Figure 1a–c, even the sketch image where the distance between the inner corner of the brows and the mouth is the only clearly marked feature. It is important to note that, in real faces, these configural cues generally occur when there is a facial muscle change and, hence, these can be employed to determine the AUs associated to the perceived emotion.

However, configural cues alone may not be able to account for the recognition of all the AUs responsible in producing anger and sadness. There should exist other cues (e.g., curvature of the brows or mouth) which also influence our perception of an emotion. In the present paper, we show how configural cues, such as the one illustrated above, emerge naturally from a shape-based representation of the face. We then demonstrate how additional configural cues for the recognition of anger and sadness can be obtained from this proposed model. In particular, we show that the height-width ratio of the face is a second dimension of this computational space. We demonstrate how this second dimension can attenuate the relevance of the one defining the brow to mouth distance and even change its interpretation. In Figure 3a, we illustrate this two-dimensional computational face space. The x-axis describes a decrease (left of the origin) or increase (right) of the distance between eyebrows and mouth. The y-axis represents the height-width ratio of the face; thinner faces are above the origin, wider faces below it. The origin represents the normative (average) face. As we can see from this model, a decrease of the brow-mouth distance results in an increase of the perception of anger, Figure 3b. Then, an increase in the width of the face reduces this perception without the need to decrease the intra brow-mouth distance, Figure 3c.

Figure 3.

Figure 3

Computational space of facial expressions of anger and sadness. Anger is represented on the left hand-side of the origin. Sadness is on the right-hand side. The origin represents the norm (average) neutral face of a population. a) The two dimensions defined in this paper, which correspond to the distance between the baseline of the inner corner of the eyebrows and the mouth and the height-width ratio of the face contour (given by the jaw line and crest). b) The feature vector representing an angry face moves further to the left when the brow-mouth distance is decreased. c) The same feature vector is move downward and to the right when the when the height-width ration decreases (i.e., the height decreases of the width increases). In this example, the vector moves in a way to maintain the 2-norm distance to the origin. The 2-norm is represented by the circle. The norm that will be followed depends on the stimuli images we use for training, as it will be shown later.

The computational face space just introduced makes a series of important predictions. First, if one of the dimensions of our cognitive space really describes a configurational change, then this effect should also be present in sketches of the face where only this feature is manipulated. In Experiment 1, we will show that this is indeed the case. Second, we note that the perception of anger/sadness should disappear with a simple inversion, since inversion of the stimuli is known to eliminate configural cues (Yin, 1969; Rhodes, 1988; Farah et al., 1998; Leder & Bruce, 2000). We demonstrate this effect in Experiment 2. Third, if the second dimension of the computational space describes the height-width ratio of the face, then the first dimension should not have such a large influence when western subjects look at Asian face, because these are generally wider than those commonly seen in western countries. This hypothesis provides new insights into the other race effect, where subjects have a hard time identifying faces of uncommonly seen races (Valentine 1991). We validate this hypothesis in Experiment 3, where we show that the attenuation of the perception of anger and sadness is proportional to the increase in width and/or decrease in height.

The results summarized above illustrate how configural cues can be learned (obtained) from a representation of the shape of the face. As such, these results suggest that shape analysis is a major mechanism underlying the coding and recognition of facial expressions of emotion. It also shows how simple yet robust features can be extracted from this representation, which justify our ability to do recognition of emotions from impoverish stimuli, such as sketches or low resolution images.

The results reported in the present paper suggest that the recognition of facial expressions of emotion has more in common to other processes of face recognition than previously thought. They also provide evidence on what it means to become a “face expert.” Under the proposed model, an “expert” would be that who moves from the shape-based representation to the configural one. These results suggest that one could also become “expert” to other objects where shape leads to configural cues, e.g. Greebels (Gauthier & Tarr, 1997). Our results also have important implications in developmental models (Freire and Lee, 2002; Le Grand et al., 2004; Pascalis et al., 2002; Bouvrie and Sinha, 2007) and the study of some mental illnesses associated to/with possible perceptual abnormalities (Chambon et al. 2006; Shin et al., 2008; Deruelle et al., 1999). Furthermore, we will illustrate how our results can be used to justify some predictions made by the continuous and categorical views of emotion representation.

2. COMPUTATIONAL MODEL

As shown in Figures 1e–d and 2, faces with a longer eye to mouth distance are perceived as sadder and those with a reduced distance as angrier. Nonetheless, the faces shown in the figure bear neutral emotions. This effect is most probably due to an overgeneralization that arises from the analysis of sad and angry faces (Neth & Martinez, 2009; Zebrowitz, 1997). We have observed (Neth & Martinez, 2009) that the distance between the inner corner of the two eyebrows and the mouth increases when one expresses sadness and decreases while expressing anger. This is made evident in Figure 1c. Therefore, one hypothesis is that when designing a shape-based algorithm for the classification of anger and sadness, the eyebrow to mouth distance will be among the ones providing most discrimination. Here, by “discrimination” we mean that this feature has a clear, unique position for each emotion and that this can be readily detected under varying illumination and pose or impoverished conditions.

If the above mentioned hypothesis were true, then a shape analysis of the faces shown in Figure 2 would identify the brow to mouth and eye to mouth distances as the most important for representation and discrimination. These results are clear from the computational analysis of shape to be described next. This shape-based representation will then be taken to define (at least part of) the computational model of sadness and anger.

To define a shape-based face representation and facilitate comparison between individual faces, a stimuli set consisting of 60 images was generated. To construct this set, twelve face images were selected from the AR face database (Martinez & Benavente, 1998), six males and six females with no large moles or any clearly identifiable feature that could cause distraction over the other facial components. The images are of 768 × 576 pixels. Each of the twelve images was then carefully modified using a morphing software (FantaMorph©, by Abrosoft) to increase or decrease the vertical eye-mouth distance without effecting the shape or texture of each of the facial features and their surroundings. The increase (decrease) in distance corresponded to an addition (subtraction) of 20 pixels between eyes and mouth. To obtain this increment (decrement) of the eye-to-mouth distance, placement of the brows, eyes, nose, and mouth were varied along the vertical axis. Four types of configural change were applied, corresponding to a downward/upward displacement of the eye and brows (Figure 2a,c) or a downward/upward displacement of the mouth and nose (Figure 2b,d). Those configurations which correspond to decrease on the distance between the brows and the mouth are perceived as angrier and will conform the angry set. The images resulting in a larger eyebrows-mouth distance are grouped in the sad set.

2.1 METHODS

Our shape representation is given by eight separate facial features comprising the contours for the nose, mouth, jaw, crown, left eye, right eye, left brow, and right brow, Figure 4a. These shape contours can be automatically obtained using the computational method of Ding and Martinez (2010). From these contours, the set of facial markers shown in Figure 4b were obtained. Fiducial points on the face images were used to ensure consistency between images. For example, the nose contour begins and ends at the nasal canthi, while the jaw contour begins and ends at the temporal canthi. The facial markers were then used to calculate cubic splines, each describing one of the facial features. The length of each spline curve was determined by summing the Euclidean distances between each point of the spline. The curve length was then used to resample the contour at equal intervals resulting in equally spaced points along each feature contour. Both the nose and mouth contours are represented by 21 points while the remaining contours are represented by 11 points, Figure 4b.

Figure 4.

Figure 4

Definition of facial features. a) Feature contours; b) Feature points.

Our shape representation is then based on the classical statistical modeling of shape (Dryden & Mardia, 1998): Each of the 108 fiducial points defined above is given by its (xj,i,yj,i) image coordinates, i={1,…,108}, and j is a unique label assigned to each of the images in our stimulus set, j={1,…,60}. Shapes are thus represented by 216 parameters. This means that each face image is now represented by a 216-dimensional feature vector, each feature representing the x and y coordinates of the face fiducials, (xj,1,…,xj,108, yj,1,…,yj,108)T. To make this representation invariant to translation, we calculate the mean over the x values, μj,x, and the mean over all y, μj,y. Subtracting the mean form the feature vectors, provides translation invariance, i.e., xj=(xj,1−μj,x,…,xj,108−μj,x, yj,1−μj,y,…,yj,108−μj,y)T. In fact, this maps the previously defined feature vectors to the null space of the vectors of all ones, which describes all possible translations. This implies that the new representation X (a matrix whose columns xj are the feature vectors just defined) is rank deficient, since all the vectors lie in a 215-dimensional space.

The resulting normalized set of feature points representing the twelve original face images (xj, j=1,…,12), were used to compute a set of feature points describing the mean face. The centroid and principal axes were determined for each individual set of feature points (for all images, including that of the mean face). Each individual set of feature points was then rotated to align the principal axis with that of the average set of feature points. The area encompassed by the jaw and crown contours was calculated for the individuals as well as the mean point sets. Each individual set of points was scaled such that the area matched that of the mean. This process made the representation invariant to in-plane rotations and scale.

Before we apply this representation to images expressing real anger and sadness, we need to demonstrate that the shape-based model defined above is consistent with the configural cues described in the earlier and with Figure 3. To show this, Principal Component Analysis (PCA) was performed on the normalized set X (i.e., the normalized feature points of all 60 faces). Figure 5a shows the computational (shape-based) space obtained by the two Principal Components (PCs) associated to the largest eigenvalues. Recall that the PCs associated to the largest eigenvalues are those which carry most of the variance of the data, i.e., those PCs that best describe the data in the sense of keeping most of the energy (Martinez and Kak, 2001).

Figure 5.

Figure 5

a) Principal components for fully shifted images. Eyes Up (black); Eyes Down (green); Mouth Down (magenta); Mouth Up (blue); Neutral (red). b) The images of the Ekman database have been described using the shape model defined in the text and then projected onto the space defined in (a). We see that all the images are correctly classified using a single dimension – that measuring the relative distance between brows and mouth. Note that the origin (0,0) of the plot corresponds to the normative measures. We also see that sad faces are generally thinner than average. This means that 70% of the sad and angry faces can be successfully classified using the width-height ratio.

The center of this PCA space corresponds to the mean face. All other faces are given as deviations from this mean (norm) face (Valentine, 1991; Neth & Martinez 2009). Our twelve original neutral faces described earlier are shown in red in Figure 5a. The modified images are shown using other colors: black for eyes-bows up, green of eyes-brows down, blue for mouth-nose up and magenta for mouth-nose down. Each of the numbers corresponds to one of the 12 possible identities.

The key result is to employ this same computational model in the classification of anger and sadness. For this, we use the 15 images of anger and the 15 images of sadness in the Ekman & Friesen (1978) set. The same 216 facial markers are used to defined each face Aj and Sj, where j=1,…,15, and “A” is used for Angry and “S” for sad. These feature vectors are now projected onto the 2-dimensional space defined above. The result is in Figure 5b.

2.2 RESULTS

The first principal component (PC1) of the shape model shown in Figure 5a carries 55.3% of the data variance. This means that this dimension of the computational (shape-based) space is more important than the rest combined. We see that (as expected) this dimension separates the images into two groups, corresponding to the angry and sad sets. The second principal component (PC2) has less impact on the discrimination of angry and sad expressions, accounting for 14.7% of the variance. It separates each group into two subgroups; i.e., eyes-bows down and mouth-nose up for the angry group, and eyes-bows up and mouth-nose down in the sad set.

According to the shape model just defined, it is clear that PC1 is the dimension that most affects the perception of sadness and anger in our image set. Hence, our next concern is to determine what facial features are most represented by PC1. To do this, the relative contribution made by each of the weights in both PC1 and PC2 was examined. First, each element in every PC was normalized to the value of its maximum weight. The relative size of each weight comprising PC1 is depicted in Figure 6a. The weights for all x- and y-image-components are shown in their respective spatial locations, i.e., the contributions of the xj,i and yj,i in X are shown in their corresponding physical location of the face. The contribution of the x-component is plotted with an asterisk, while that of the y-component is plotted as a circle. The size of each symbol reflects its relative weight, normalized to the maximum weighting.

Figure 6.

Figure 6

Relative importance of PC weights: a) PC1, and b) PC2.

It is clear from the figure that the y-components of the mouth and the inner corners of the brows are major contributors, with the eye corners being also relevant. This shows that PC1 encodes (primarily) the distance between the eyebrows and the mouth. However, it is not limited to this. Further analysis shows that the weights of the bottom of the nose are also important. This suggests that modifying the vertical distance between the mouth and the nose would also have an effect on the perception of anger/sadness. To illustrate this effect, we generated another set of images, where the mouth was (as above) moved 20 pixels up or down, but the nose was only moved 10 pixels, Figure 2e–f. A projection of the shape representations of these new images shows that they fall between the neutral and sad group, when the distance is increased, and between the neutral and angry set, when the distance is decreased. This results demonstrate that a reduction in the displacement of the nose have a direct effect on the perception of anger and sadness.

Also, we notice that the vertical distance between the jaw line and the crest (i.e., heights of the face) is also encoded by PC1. In fact, as this distance is reduced, faces moved toward the mean face and, as such, their perception of anger or sadness is diminished. For example, the subject identified by the label “10,” has a longer faces, making his original and modified “sad” expression look sadder, but his modified “angry” expression seem less angry, Figure 5a.

The relevance of the size of the face is further illustrated by the second PC defining our computational space. The relative size of each weight comprising PC2 is depicted in Figure 6b. We see that the y-components of the jaw and both the x- and y-components of the crown are the highest contributors to PC2. This reinforces the importance that the height, and to some extent the width, of the face influences the perception of anger and sadness in these face images. Note that the height and width are correlated and thus the relevant term is their ratio.

To further verify these results, we carry a discriminant analysis of the shape feature vectors described above. Discriminant Analysis (DA) is a dimensionality reduction (more specifically, a feature extraction) method that identifies the linear combination of features that are most important to discriminate two or more classes (Martinez & Zhu, 2005), rather than those that carry most of the variance as in PCA. Hamsici & Martinez (2008) have derived an algorithm to extract (i.e., obtain) the 1-dimensional space where the Bayes classification error is the smallest from all possible 1-dimensional spaces. Recall that the Bayes error is the smallest error a classifier can attain in a given space (Fukunaga, 1990). Hence, the Bayes optimal DA (BDA) of Hamsici & Martinez (2008) identifies the linear combination of features that can best discriminate between classes.

We first apply the BDA algorithm to faces expressing anger and sadness from the Ekman-Friesen set. To visualize the results, we followed the same procedure described above for PCA. The results are in Figure 7a. In these results, it is even clearer that the y-component of the brow's features carry most of the discriminant information. It is also evident from these results that the y-component of the brows is mostly compared to that of the mouth. In Figure 7b, we show the features that are most useful to discriminate between a neutral expression (i.e., a face that does not convey any emotion). This DA analysis shows that the most relevant feature is the height/width ratio. In Figure 7c, we show the features that best discriminate between anger and sadness, neutral, and other commonly observed emotions (happiness, surprise, fear and disgust). The most important features being the eyebrows, mouth and height/width ratio. Hence, again, these three facial features seem instrumental for discriminating anger from sadness and necessary for distinguishing them from other typical emotions.

Figure 7.

Figure 7

Discriminability of each of the facial features of the shape model for classifying between: a) anger versus sadness, b) anger versus sadness versus neutral, and c) anger and sadness versus neutral, happiness, surprise, fear and disgust. Overall, the most discriminant features are the eyebrows, mouths and height/width ratio.

The final and possibly most important result is shown in Figure 5b. Here, we see that the first dimension of this computational space correctly classified all the 30 images in the Ekman-Friesen set. That is the reason why this dimension carries most of the variance, because it undoubtedly serves as a great prediction of sadness versus anger. The second dimension (width-height ratio) is also a good prediction, since it correctly classified 70% of the images. And, although it is not as useful as the first, combining it with the first dimension increases or decreases the perception of anger and sadness of the percept.

2.3 DISCUSSION

The results obtained above suggest a computational space where two of its dimensions correspond to configural cues. The first of them describes the vertical distance between the eyebrows and the mouth. The second one defines the height-width ratio. This space works as follows: When one expresses anger, the distance between the inner corners of the eyebrows and the mouth decreases. The same distance increases when expressing sadness. The first dimension of the cognitive space uses this information to discriminate between sad, angry and neutral expressions. However, when an individual face carries an uncanny large brow (or eye) to mouth distance, this face is (incorrectly) interpreted as sad even if the person bears a neutral expression. When this distances is shorter than average, the face appears angry. This is illustrated in Figure 3. The second dimension in this figure represents variations of the height-width ratio of the face. That is, when the vertical distance between the top and bottom of the face decreases, the effect on the perception of sadness and anger is reduced/increased. Note that reducing the heights of the face has a similar effect to increasing its width. For instance, an increase of the width also reduces the perception of sadness in faces with a large or small eye-mouth distance.

In summary, faces with specially large or short brow-to-mouth distance and particular thin or wide faces are perceived as being sadder or angrier than more normative faces. While this may seem a disadvantage, very few faces bear such a large deviation from that of the average face – the male character in Wood's famous painting American Gothic being an example of an uncannily long face and a large brow to mouth distance. Within each ethnicity and race, faces tend to be quite normative (Zebrowitz, 1997). This carries the already mentioned advantage of being able to classify actual expressions of anger and sadness very reliably. We have seen how the angry and sad faces of the Ekman and Friesen (1978) set can be successfully classified with this computational space. This is because there is a direct relationship between the configural changes described above and the shape changes caused by the movement of the AUs involved in sadness and anger. For example, anger is mostly marked by lowering the inner corners of the brows with AU 4. Similarly, sadness can be express with AU 1, which will raise the brows and AU 15, which can lower the bottom of the mouth.

If the computational face space indeed includes the two configural dimensions described in this section, then there is a set of hypotheses that can be made from it and should be observed with human subject. First, since the first (and most important) dimension represents a simple brow-to-mouth distance, its effects should also be noticeable from a simple line drawing of the face. In experiment 1, we will show that even when faces are reduced to very simple sketches, subjects still associate a deviation of the brow-mouth distance to anger and sadness. Yet, since the information used by the subjects to do this classification is assumed to be configural, then a simple inversion of the stimuli should eliminate the perception of any emotion. In experiment 2, we prove that this is indeed the case.

We then study the role of the second dimension of our computational space in another experiment. In experiment 3 we show that when the height-width ratio of faces decreases, the perception of sadness and anger diminish. We do this by including faces of other races. In particular, we show Asian faces to Caucasian subjects. Asian faces tend to be shorter and wider, reducing the height-width ratio. According to our computational model, manipulation of the brow-mouth distance should have a milder effect in the perception of sadness and anger in these images. The results of this experiment are consistent with this prediction.

3. EXPERIMENT 1

The computational (cognitive) space defined in the preceding section suggests that configural cues are used to code and recognize anger and sadness from images of faces. For the model to hold true, one should perceive anger and sadness in simple drawings (i.e., schematics) of neutral faces where the distance between the eyes and mouth have been made shorter or larger than average.

3.1 METHODS

Images

The schematic faces used in this experiment were derived from the average of the twelve face identities used in the previous section. The eyes were represented by small squares of 3×3 pixels, and were placed at the centroid for each eye contour. The eyebrows and mouth were represented by horizontal lines passing through the centroids of their respective contours. The left and right endpoints were taken from the left and right maxima of each contour. The nose was represented by a single vertical line passing through the centroid of the nose contour. The upper y value was set as the y value for the eyes. The lower y value was set to correspond to the lower boundary at the tip of the nose. Finally, the jaw and crown were represented by the values obtained for the average of all faces in the previous section. This process eliminated all textural information and local shape, leaving exclusively configural cues in the stimuli image. In this first experiment, the jaw and crown contours will remain fixed across all configural changes. In the previous study, facial components were displaced by 20 pixels. Small differences like these are sufficient to make real faces look strikingly different (Mondloch et al., 2006). When the texture and (shape) details of the facial components are eliminated, these changes need to be exaggerated – as it is typical in cartoons. The current study utilized a maximum displacement of 40 pixels. To achieve this, the schematic face for the neutral average is shown in Figure 8a. For each type of configural change, individual images were created at the following percentages of displacement: 0% (i.e., original image), 25% (10 pixel displacement), 50% (20 pixels), 75% (30 pixels), and 100% (40 pixels), Figure 8b–g. All images are of 768 × 576 pixels. Image sets from these four configurations were divided into two groups. The angry group included face images in which the mouth was displaced upward and face images in which the eyes and brows were displaced downward. The sad group included face images in which the mouth was displaced downward and face images in which the eyes and brows were displaced upward.

Figure 8.

Figure 8

Schematic Faces. a) Neutral schematic face. b) Eyes-Brows Down; c) Eyes-Brows Up; d) Nose-Mouth Up; e) Nose Mouth Down; f) Nose-Half-Mouth-Full Up; g) Nose-Half-Mouth-Full Down.

Procedure

Face images from the sad schematic group and the angry schematic group were used in two independent sessions. Each session was run on different days. In each session, face images were presented sequentially in 108 pairs and no image pair was used more than once. The type and magnitude of configural change were randomly selected for the initial image; the magnitude of configural change applied to the second image was randomly selected. Prior to the display of the first image of each pair, an initial visual cue (crosshair) was shown for 600 ms in a randomly selected location to alert the subject as to where the first image would be displayed. The first image was then displayed for 600 ms. This was followed by a mask with a duration of 500 ms. A second visual cue was then displayed in another randomly selected location for 600 ms followed by the display of the second image for another 600 ms. To preclude any confounding effects due to implied motion, the screen position at which each image was displayed was randomly varied. Subjects were asked to indicate whether the second image seemed more, the same, or less angry (or sad) relative to the first image displayed. A typical stimulus time-line is shown in Figure 9. The duration for each of the two sessions was around 10 minutes.

Figure 9.

Figure 9

Stimulus time-line for schematic faces. An initial visual cue is shown in a randomly selected location to alert the subject as to where the first schematic face image will be displayed. After 600 ms, the first schematic face image is displayed for a period of 600 ms. A mask (blank screen) is then presented for 500 ms. A second visual cue is displayed in another randomly selected location for 600 ms. The second schematic face image is then displayed for 600 ms. Finally, a blank screen with a large question mark (“?”) at the center is displayed until the subject responds with a key-press.

Subjects

Twenty-one subjects with normal or corrected-to-normal vision were drawn from the population of faculty, staff, and students at The Ohio State University. Subjects were seated at a personal computer with a nineteen inch LCD monitor. The typical viewing distance was 50 centimeters, providing a percept of approximately 15 degrees vertically and 8 degrees horizontally.

3.2 RESULTS

Responses to the configural changes comprising the sad schematic and angry schematic groups are shown in Figure 10a,b. In each case, the abscissa reflects the difference in configural change between the first and second image in each presentation. For example, if the first image is at a level of 25% and the second image is at a level of 50%, the subject's response would fall into the 25% grouping. Likewise, if the first image is at 100% and the second image is at 50%, the response would fall into the −50% grouping. The ordinate reflects the percentage of less, same, and more responses made at each grouping.

Figure 10.

Figure 10

Cumulative responses for schematic faces: a) Sad schematic group; b) angry schematic group. Regression for schematic faces: c) Sad schematic group; d) Angry schematic group. The results suggest a mixture of Gaussian responses which can be justified by a mean-center RBF neural representation.

We see that as the distance between the brows and the mouth increases/decreases, the perception of sadness/anger increases with it. This increase is clearly sigmoidal – as one would expect from a categorical perception of emotions. To study this effect quantitatively, a regression analysis was performed on the more responses to increasing configural change and the less responses to decreasing configural change. In addition, same responses for both increasing and decreasing configural change were also estimated. To accurately represent the trends for each response, a third-order (cubic) polynomial was employed. The results are shown in Figure 10c,d. The data shows that the effect of increasing or decreasing the configural change tails off at the extremes. This is due to a ceiling effect on the maximum number of faces that are likely to be classified in each emotion. This result suggests a computational space where faces are perceived as sadder/angrier as we move away from the normative (mean) face, but with a categorical (threshold) effect limiting the influence of emotion perception by each of the feature (dimensions) in our representation. This means that when the distance between brows and mouth increases by x, the computational model will estimate the increase of the perception of the given emotion as f(x), where f(.) is a sigmoidal function. Figure 10c,d show estimated functions of the subject responses.

To demonstrate that there is a statistical significance between the results observed with the schematic faces above, a χ2 goodness-of-fit test was performed on each of the levels of configural change. For each of the sad schematic and angry schematic groups, the responses for the 0% displacement were used as the expected values. Thus, the null hypothesis (H0) states that there is no difference between the response profile for any given level of configural change when compared to the response profile for no configural change. All comparisons for both the sad schematic group and angry schematic group yielded a significant χ2 value for p≤0.01 with two degrees of freedom (χ2cv= 9.21; df = 2). Hence, the response profiles for each level of comparison are significantly different from the respective response profile of no configural change. For both the angry schematic group and sad schematic group, the residuals indicate that the less responses are the major contributors to the significant differences for the negative deformations, which is consistent with the results obtained from our computational model. Similarly, the more responses are the major contributors to the significant differences for the positive deformations – also as expected.

As seen in Figure 8, the length of the nose changes between categories. To make sure this did not serve as a diagnostic feature in classification of anger and sadness, we repeated the above experiment with a static nose. The results were equivalent to those shown in Figure 10. This result support the observation that the brow and mouth distance is a diagnostic feature.

3.3 DISCUSSION

Line drawings of faces are commonly used by artists to depict identity, expression, and other aspects of the face. White (1999) investigated the representation of facial expressions by using line drawings modified to represent happy, sad, angry and neutral expressions. In all conditions, the eyes, nose, and external contour of the face remained fixed. Both the sad and angry condition used a downward curve for the mouth. The brows were angled downward at the nose for the angry condition and upward at the nose for the sad condition. The face stimuli were used in an inversion study to assess the degree of configural processing involved in the cognitive representation of expressions. If facial expressions are represented explicitly (i.e., by featural information) then recognition should not be greatly affected by inversion. However, the results of the study suggested otherwise. It was concluded that the configural information that defines the brows and mouth as an expression, along with the configural information that defines the line drawing as a face, is encoded as an undecomposed whole. This result is consistent with the computational (shape-base) model defined in the present paper.

In contrast to White's study, the present experiment explored the possibility that the shape-based representation yields pure configural representations and may thus bias the perception in netrual expressions. We utilized line drawings of faces in which the local features of each component remained unchanged in all images. Only the relationships among the identical features were varied, resulting in a pure configural (second-order) change. The results of this study support the notion that configural differences bias the perception of emotion in neutral faces and are thus likely dimensions of the computational space of facial expressions of emotion. The general form of this pattern is similar to that identified in with full-color face images (Neth & Martinez, 2009). The responses tend to roll off at the extremes, suggesting a Gaussian response for each set of changes. This can be justified by a mean-center radial basis function (RBF) neural representation. RBF have been previously used to model different visual processes, including face perception (e.g., Howell & Buxton, 1998).

The computational space defined thus far would represent the neutral face at its center, as it is common in the contiuous view (Russell, 1980). This model plays a central role in Russell's (2003) concept of core affect which is the neurophysiological state accessible as the simplest, non-reflective feelings. Core affect is always present but may subside into the background of consciousness. The prototypical emotions described in categorical approaches map into the core affect model but hold no special status. According to Russell (2003), prototypical emotions are rare; what is typically categorized as prototypical may in fact reflect different patterns than other states classified as the same prototype. In this sense, emotional life comprises the continuous fluctuations in core affect, the on-going perception of affective qualities, frequent attribution of core affect to an external stimulus, and instrumental behaviors in response to that external stimulus. The many degrees and variations of these components will rarely fit the pattern associated with a prototype and will more likely reflect a combination of various prototypes to varying extents.

While Russell makes a compelling argument for the continuous nature of emotion, humans appear predisposed to experience certain continuously varying stimuli as belonging to distinct categories. In most psychological processes, the density in the computational space plays a major role in the perception of a category (Krumhansl, 1978). For example, within a group (e.g., ethnicity, race), most faces are alike and, hence, most faces are represented around the origin of the computational space. It is well-known that distinguishing two entities becomes more difficult as the density (i.e., the amount of stimuli present in this area) increases. Therefore, face images that are closer to the mean are perceived as similar to one another, since the density there is high. Faces that are away from the mean are seen as more distinctive, because the density is low. This effect is shown in Figure 11. In Figure 11a–b, we show the subject responses to variations of 25%. In Figure 11c–d we see the responses at 50% change and in Figure10e–f those at 75%. The variation of change is shown in the x-axis of the plots in this figure, which is indicated as nm%, with n specifying the percentage of configural change applied to the first image and m the change applied to the second. For example, the labels 0–25%, 25–50%, 50–75% and 75–100% all represent variations of 25% between the first and second image and are thus in Figure 11a,b. In these plots, we see that the percentage of times that subjects considered the second image was sadder/angrier than the first increases as we move away from the mean face. That is, the image at 100% change is seen as sadder/angrier than that at 75% more times than when the image at 25% is compared to that of 0% -- even though the increase/decrease of the configural change is the same. Close analysis of all the plots in Figure 11 revels that this is true for all cases; i.e., the perception of anger and sadness increases as we move away from the mean (i.e., the center of the computational space).

Figure 11.

Figure 11

Responses for a–b) 50% configural change in schematic faces, c–d) 100% configural change, e–f) 150% configural change.

This last effect is similar to that observed with caricaturing. It has been shown that while exact line drawings of a face are difficult to identify, caricatures of the face are generally easier (Sinha, 2002; Benson and Perrett, 1994). Caricaturing has also been shown to improve automatic recognition of faces by computers (Craw et al., 1999) and neural responses in the macaque anterior inferotemporal cortex increase sigmoidally as the stimulus face is made more distinct from the norm (Leopold et al., 2006).

Figure 11 also shows what happens when the first image is at a greater configural change than the second, i.e., n>m. In this case, we see that there is no clear increase as we move away from the mean face. This is in fact expected, because when n>m, the second image is closer to the mean than the first and thus more difficult to be interpreted – we move from a less to a more dense region of the face space. Therefore, this result suggests that the important variables are the individual density of each of the images employed. In the computational model, this means that the response f(x) needs to be combined with the inverse of the density δ(.) at the point x, e.g., f(x) + αδ(x), with α>0 defining the relevance of the density term.

4. EXPERIMENT 2

The purpose of this study is to assess the impact of inversion on the perception of emotion found in the previous experiments with upright schematics. If the proposed computational space were correct, it would be expected that the inversion will disrupt the effect seen with upright face images. Hence, the hypothesis of this study is that configural changes in the relative position of the nose, mouth, eyes, and eyebrows in inverted faces images will not affect the perception of emotional expression in an otherwise neutral face.

4.1 METHODS

Images

The current study will utilize the manipulated face images shown in Figure 2, with the difference that all face images have now been inverted. As in the initial experiment, the relationships among internal features of the face are manipulated without affecting the overall shape of the face or head, resulting in a configural change. Placement of the eyes, brows, nose, and mouth were varied along the vertical axis. For each type of configural change, individual images were created at the following percentages of displacement: 0% (i.e., original image), 25% (5 pixel displacement), 50% (10 pixels), 75% (15 pixels), and 100% (20 pixels). Two types of configural changes scaled the displacement of the lower aspect of the nose to half of the displacement of the mouth. Image sets from the six configurations were divided into two groups. The angry inverted group included face images in which the mouth was displaced upward and face images in which the eyes and brows were displaced downward. The sad inverted group included face images in which the mouth was displaced downward and face images in which the eyes and brows were displaced upward. All face images were inverted by applying a 180 degree rotation. The reader can see the stimuli images by rotating the page with Figure 2 by 180°.

Procedure

The method of presentation was identical to that of the first experiment. Face images from the sad inverted group and the angry inverted group were used in two independent sessions which were run on different days. Faces were presented sequentially in 300 pairs with no image pair being used more than once. Subjects were given the opportunity to take a brief rest at the 1/3 and 2/3 completion points.

In each session, pairs of face images were displayed sequentially. The first image was randomly selected; the identity and type of configural change applied to the second image remained the same as in the first image while the magnitude of configural change applied to the second image was randomly selected. An initial visual cue (crosshair) was shown for 600 ms in a randomly selected location to alert the subject as to where the first image would be displayed. The first image was then displayed for 600 ms. This was followed by a mask with a duration of 500 ms. A second visual cue was then displayed in another randomly selected location for 600 ms followed by the display of the second image for another 600 ms. To preclude any confounding effects due to implied motion, the screen position at which each image was displayed was randomly varied. Subjects were asked to indicate whether the second image seemed more, the same, or less angry (or sad) relative to the first image displayed.

Subjects

Seventeen subjects with normal or corrected-to-normal vision were drawn from the population of faculty, staff, and students at The Ohio State University. None of the subjects had participated in the previous study. Subjects were seated at a personal computer with a nineteen inch LCD monitor. The typical viewing distance was 50 centimeters, providing a percept of approximately 15 degrees vertically and 8 degrees horizontally.

4.2 RESULTS AND DISCUSSION

As in the previous study, an ANOVA (p≤0.01) was performed on each of the less-same-more response profiles at each level of configural change for both the sad inverted and angry inverted groups, Figure 12. The nose & mouth down, nose half down & mouth full down, and eyes & brows up responses were combined to ensure continuity. Likewise, the nose & mouth up, nose half up & mouth full up, and eyes & brows down responses were combined. Results indicate that the number of same responses are significantly different from their respective less and more responses in all cases. None of the less responses is significantly different from the more responses either.

Figure 12.

Figure 12

Responses to inverted face images. a) Sad inverted group; b) angry inverted group.

A χ2 goodness-of-fit test was performed on each of the levels of configural change. For each of the sad inverted and angry inverted groups, the responses for the 0% displacement were used as the expected values. Thus, the null hypothesis (H0) states that there is no difference between the response profile for any given level of configural change when compared to the response profile for no configural change. None of the comparisons for the sad inverted group yielded a significant χ2 value for p≤0.01 (χ2cv= 9.21; df = 2). Therefore, in both the sad inverted and angry inverted sets there was no perception of sadness or anger. This shows that all cues, which previously made the “neutral” faces of Figure 2 appear as sad and angry, disappear when the stimuli are inverted. This provides strong support for the notion that the biases observed in the previous experiments are in fact due to configural rather than featural processes. This is congruent with the conclusions drawn from the studies of facial recognition utilizing inverted face images (Yin, 1969; Farah et al., 1998; Valentine, 1991; Rhodes et al., 1993).

While studying identity recognition on faces, Leder & Bruce (2000) showed that most of the information lost during inversion is indeed the spacing (second-order relations) between features rather than other holistic properties. The results reported in this section suggest that the same is true for the recognition of facial expressions of emotion. Evidence indicates that configural processing applied to the recognition of identity is learned and is not fully developed until ~8 years of age (Le Grand et al., 2004; Pascalis et al., 2002). This suggests that configural information is also learned. Our computational model demonstrates how this configural information can be learned (extracted) from a shape-based representation of faces. In this model, faces are first represented in a shape space. Becoming a “face expert” implies extracting configural cues that are easy to compute and robust to image changes. These extracted features are however sensitive to inversion and to faces of other races where the configural cues are different. This section has provided evidence for the former of these two effects, we turn to the latter effect in the section to follow.

5. EXPERIMENT 3

In this study, the interaction between the other-race effect and biases induced by configural changes is examined. There has been minimal direct research on the perception of expressions in other-race faces, especially as related to configural features. A reduction in holistic processing, along with increased reliance on featural analysis, has been observed in studies with faces of uncommonly seen races (Tanaka, Kiefer, & Bukack, 2004; Michel et al., 2006). This is usually known as the other-race effect. However, the cited studies dealt with the recognition of identity and did not address the recognition of facial expressions of emotion.

In this paper, we have argued that (at least part of) the computational face space of expressions is similar to that used to encode identity, since this also uses configural cues to represent expressions. Hence, in the present experiment, it is expected that configural changes applied to Asian faces will result in the perception of emotion in manipulated neutral faces. Specifically, it is hypothesized that configural changes in the relative position of the nose, mouth, eyes, and eyebrows will affect the perception of emotional expression in an otherwise neutral Asian face, but to a lesser degree than that observed for in-group (Caucasian) faces. This reduction will be shown to be given by the different height-width ratio of Asian faces as compared to Caucasian faces. This result will be obtained by means of a statistical analysis of the computational face space defined above.

5.1 METHODS

Images

As in the experiments above, six types of configural change were applied to each of 12 individual face images taken from the Japanese and Caucasian Neutral Faces of the Ekman & Matsumoto database (1988). For each type of configural change, individual images were created at the following percentages of displacement: 0% (i.e., original image), 25% (5 pixel displacement), 50% (10 pixels), 75% (15 pixels), and 100% (20 pixels). Two types of configural changes scaled the displacement of the lower aspect of the nose to half of the displacement of the mouth. One of the unchanged originals is shown in Figure 13a. Examples of each type of configural change are shown at 100% displacement in Figure 13b–g. All images were cropped to 768 × 576 pixels. As above, image sets from the six configurations were divided into two groups. The angry Asian group included face images in which the mouth was displaced upward and face images in which the eyes and brows were displaced downward. The sad Asian group included face images in which the mouth was displaced downward and face images in which the eyes and brows were displaced upward.

Figure 13.

Figure 13

Samples of each type of configural displacement. a) Original face image; b) eyes-brows down; c) eyes-brows up; d) nose-mouth up; e) nose mouth down; f) nose half up, mouth full up; g) nose half down, mouth full down.

Procedure

The method of presentation for this third experiment was identical to that of the previous two presented above. Face images from the sad and the angry Asian groups were used in two independent sessions. Each session was run on different days. Faces were presented sequentially in 300 pairs with the identity, type and magnitude of displacement randomly selected for the initial image. No image pair was used more than once. Subjects were given the opportunity to take a brief rest at the 1/3 and 2/3 completion points.

Subjects

Twenty non-Asian subjects with normal or corrected-to-normal vision were drawn from the population of faculty, staff, and students at The Ohio State University. The subjects had not participated in the previous studies. Subjects were seated at a personal computer with a nineteen inch LCD monitor, with a typical viewing distance providing a percept of approximately 15 degrees vertically and 8 degrees horizontally.

5.2 RESULTS AND DISCUSSION

Responses for the sad Asian and angry Asian groups are shown in Figure 14a–b. As we have done before, the abscissa in each plot reflects the difference in configural change between the first and second image in each presentation. The ordinate represents the percentages of the number of responses for each level of change.

Figure 14.

Figure 14

Asian responses. a) Sad Asian group. b) Angry Asian group. c,d) Linear fits to the responses shown in (a) and (b).

The overall pattern of responses for both Asian groups reflects a greater discriminability for larger changes. However, this pattern is not as distinct as that noted with the in-group (own-race) faces in the first experiment (see Figure 1d–e). Nonetheless, a change from maximum to neutral still leads to a perception that the level of sadness or anger is less in the second image relative to the first. Similarly, a change from neutral to maximum results in a perception of increased sadness or anger, as seen by the increase in more responses.

A χ2 goodness-of-fit test was performed on each of the levels of configural change. For each of the sad and angry Asian groups, the responses for the 0% displacement were used as the expected values. Thus, the null hypothesis (H0) states that there is no difference between the response profile for any given level of configural change when compared to the response profile for no configural change. All comparisons for the sad Asian group yielded a significant χ2 value for p≤0.01 (χ2cv= 9.21; df = 2). All comparisons except for −25% in the angry group also yielded a significant χ2 value.

A regression analysis was performed on the more responses to increasing configural change and the less responses to decreasing configural change. In addition, same responses for both increasing and decreasing configural change were analyzed, Figure 14c–d. The percentage of more responses increases linearly with the amount of positive differences of change. For the sad Asian group (Figure 14c), as the distance between the baseline of the eyes and the mouth increases, the percentage of more responses increases linearly (r2=0.994, where r2 measures the correlation of a linear function g(.) and the subjects responses). This linear modeling corresponds to the mid section of a sigmoidal response, as we have seen in previous sections. For the angry Asian group (Figure 14d), when the distance between eyes decreases, the more response increases linearly (r2=0.996). Similarly, the percentage of less responses increases linearly with the amount of negative change. In this case, for the sad Asian group, the r2 value is 0.983, while for the angry Asian group r2=0.969. The percentage of same responses (i.e., identical perception of sadness or anger in the first and second image) also decreases linearly. For the sad Asian group, the percentage of same responses for negative change has an r2 value of 0.956, while the percentage of same responses for positive change has an r2 value of 0.989. For the angry Asian group, these are r2=0.974 and 0.961, respectively.

It is important to emphasize that the relationship between configural change and perception of expression observed with full-color faces is not as marked when employing Asian faces. This can be readily seen by examining the respective slopes associated with each linear fit. Slopes from the study of Neth & Martinez (2009) are shown alongside those of the current study in Table 1. In both the sad and angry Asian groups, the slopes are about 2/3 the value of their corresponding slopes in the sad and angry groups in the study of Neth & Martinez (2009).

Table 1.

a) Slopes for the plots of the sad group (from Neth & Martinez (2009), Figure 1d) and sad Asian group from Experiment 3 (Figure 14a). b) Slopes for angry group (Figure 1e) and angry Asian group (Figure 14b).

a
Full-Color Faces Asian Faces
Less −63.19 −42.17
Same - Decreasing 59.85 42.85
Same - Increasing −58.56 −43.32
More 62.56 46.08
b
Full-Color Faces Asian Faces
Less −61.56 −36.58
Same - Decreasing 60.42 38.49
Same - Increasing −57.89 −47.46
More 58.76 44.74

Our concern is to understand why there is this loss in the perception of sadness and anger in the Asian stimuli. Recall that the computational face space defined above included two dimensions. The first dimension was associated to the eyebrow-moth distance, while the second corresponded to the height-width ratio. Since the values of the first dimensions have been kept fixed, our hypothesis is that the decrease on the values in the second dimension (i.e., the height-width ratio) is responsible for the loss is the perception of the emotion. A comparison of the average height-width ration of the faces in our first experiment and those in the Asian group reveals a significant (ANOVA, p≤0.00013) difference. The mean ratio for the Caucasian faces is .68 while that of the Asian group is .72. This shows that the Asian faces have a 18% decrease. If we keep 92% of the values in Figure 1d–e, we obtain values in the range of those showed in Figure 14a–b, suggesting that this second dimension of the computational space serves can attenuate the relevance of the first dimension.

The proposed shape model was previously employed to justify the emergence of configural cues in the representation and recognition of facial expressions of emotion. We now want to see whether this model is consistent with the results reported in this section. Feature points were selected in each of the twelve neutral Asian face images as well as in all of the configural change conditions; as had been previously done with the other image set, Figure 4. PCA was then performed using the 108 normalized points for each face. Figure 15 depicts the distribution of all twelve neutrals (shown in red) and each type of configural change at 100% displacement (other colors) relative to the first two principal components, PC1, PC2. Both the nose half down – mouth full down and nose half up – mouth full up conditions are omitted for clarity. The center of the PCA distribution represents the mean (norm) of all twelve Asian face images used in this study. Here, the PC1 (which accounts for 48.22% of the data variance) separates the shifted images into groups that correspond to the angry Asian and sad Asian groups; as had been the case with the other image set. PC2, with 17.6% of the variance, differentiates the two types of conditions in each group.

Figure 15.

Figure 15

Asian face space given by the first two principal components of fully shifted Asian images: eyes up (black), eyes down (green), mouth down (magenta), mouth up (blue), neutral (red).

The above result illustrates how the Asian computational space (Figure 15) is as appropriate as that learned with Caucasian faces (Figure 5). However, in our experiment, we only employed non-Asian subjects. Hence, our subjects have not learned the Asian cognitive space shown in Figure 15, but rather that previously shown in Figure 5. If we plot the Asian face images into the Caucasian face space, we observe that feature vectors moves toward the center about PC1, Figure 3. To see this, let us go back to Figure 6a. We see that the shape representation of PC1 does in fact encode information of the height and width, as does PC2. Thus, when the value of height-width ratio decreases, the feature vectors move down along PC2 and toward the origin along PC1. Since it is PC1 which carries most of the variance, the perception of sadness/anger is reduced substantially. Again, our results suggest that the second dimension, modulates the placement of the stimulus along the first dimension.

To see that the features used to classify the Asian faces are still configural, we can redo the experiment reported in this section but with the stimuli of Figure 13 inverted. We asked ten subjects that had not participated in the previous experiments to rate the perception of anger and sadness. The results showed that this inversion eliminates the perception of emotion from the Asian faces.

7. GENERAL DISCUSSION

An understanding of the underlying computational face space is fundamental not only for research in face and object recognition but to understand behavior and other cognitive processes. Similarly, an understanding of how emotions are represented and recognized is essential for the understanding of behavior and cognition and plays a major role in studies of evolution and consciousness (Izard, 2009; Darwin, 1872).

Past research has demonstrated that face recognition of identity (i.e., where the task is to name the individual we see in a picture) involves configural cues (Diamond & Carey, 1986; Rhodes et al, 1989; Tanaka & Sengco, 1997). In a series of experiments, Leder & Bruce (2000) show how faces are mostly encoded by configural (second-order) cues, rather than holisticly (in the sense of being processed as a “Gestalt” where the parts are not generally decomposable from the whole). These second-order relations measure the relative spacing between features (e.g., eyes and nose), as opposed to first-order relations which would only specify the ordering of the features (e.g., eyes over the nose) (Diamond & Carey, 1986). Rhodes et al. (1989) showed that the use of these configural cues is related to expertise, since these are mostly used to code and recognize faces of one's own-race rather than those belonging to other races. Pascalis et al. (2002) demonstrated that young children may in fact be capable of making use of configural cues even in faces of other species. But, as our expertise in human faces and, in particular, those of our own-race increases, the face space becomes more tuned to small variations in our group (Mondloch et al., 2006) and less useful for others. Schwaninger et al. (2004) have proposed a computational model of configural perception consistent with how humans recognize identity.

The papers summarized above all relate to the recognition of identity in faces. The results reported in the present paper suggest that the same applies to the recognition of facial expressions of emotion. That is, configural cues constitute (at least part of) the dimensions of the computational space used to represent and recognize emotions. Our observation is in agreement with the work of Calder et al. (2000) who showed that recognition of facial expressions of emotion is faster and more accurately on “whole” faces than on faces whose top and bottom components have been cropped from two different expressions and presented in isolation or attached to form a new “whole.” They also show that this effect is eliminated when the stimuli images were inverted. These results also support the idea of a configural coding of facial expressions.

Our results have several important implications. First they suggest that the underlying mechanisms involve in the recognition of expression and identity are very similar, if not the same, as previous models have advocated (Martinez, 2003). In fact, other face analysis tasks, such as the perception of attractiveness (Abbas & Duchaine, 2008) and gender (Baudouin and Humphreys, 2006), also seem to employ configural cues – suggesting that the same or a similar computational analysis is applied to faces regardless of the classification task.

Kagian et al. (2008) have recently shown that a similar shape-based descriptor to ours can also justify the perception of attractiveness of female faces. Their results reinforce the claim that shape-based representations are important in face analysis. In another recent paper, Hammal et al. (2009) show that similar descriptions of faces can classify expressions of emotion similarly to human subjects (even when the percept is partially occluded). In the present paper, we have shown how such a shape-based representation justifies the emergence of these configural cues in the recognition of anger and sadness. It suggests that shape is a common underlying component of face analysis (similar to the conclusions of Riesenhuber et al. (2004)), but also suggests that face “expertise” emerges when one employs this representation to extract simple yet robust features for the analysis of highly similar objects -- yielding a configural space. Although these configural features are easy and robust under many image changes, one of the side effects of this “expertise” learning is that these features may appear in uncommon places. For example, a person may have an uncanny large distance between brows and mouth, making this person's face look sad even in neutral position (Neth & Martinez, 2009). This is the effect we have exploited in the present work, Figures 2 and 7. This corresponds to an over-generalization effect (Zebrowitz, 1997). Hess et al. (2009) and Zebrowitz and Fellous (2010) present a related result where angry faces are shown to over-generalize as male faces rather than women and White rather than Black or Korean. These results can also be explained by configural changes. What remains unclear is which brain pathways are responsible for such computations. The ventral pathway connecting the primary visual cortex to more specialized areas known to respond to faces, biological movement and emotions (e.g., the superior temporal sulcus, the fusiform face area or the amygdalae) is one option that has received considerable attention. This possibility makes sense if we consider shape a precursor of configural cues. However, recent research has reignited the importance of the pathway connecting the retina to the superior colliculus and then to the amygdala. It is still unclear which pathway contributes to what, but the model proposed in this paper could help test new hypotheses.

Also recently, Balas and Sinha (2008) and Sinha and Poggio (1996) have demonstrated the importance of the outline of the face. In the present study we have shown that the height-width ratio can serve as a regularizing term of the configural distance between internal facial features. These results are not only important to understand the underlying mechanisms of face and emotion recognition. A recent result by Lebrecht at al. (2009) shows that gaining expertise in faces of other races can attenuate implicit social biases attributed to them. Another study by Pollak et al. (2009) shows that abuse children are more acute at recognizing emotions, suggesting a higher degree of expertise to some image features. And, misreading faces may have important consequences in court and elections (Zebrowitz, 1997).

Our results further suggest that although the recognition of expressions of emotion is categorical, the underlying face space of is continuous. This point requires careful clarification. By continuous we do not mean to imply that expressions are not perceived categorically. In fact, our results strongly support the categorical view. This is, for example, made clear by the sigmoidal responses of the subjects. Note that while the responses to more, same and less would be expected to be linear if emotions were encoded in a continuous manner, these responses are expected to be sigmoidal under the categorical model. These categorical perceptions can be readily learned by simple Radial Basis Functions (RBFs), since their cdf (cumulative density function) is sigmoidal. In this model, each RBF would be mean-centered with the covariance defining the degree of variability allowed in each of the expressions. Thus, as we move away from the norm face and toward the configuration defining one of the emotions, the perception of that emotion increases, Figure 3b. As we move away from that position along another dimension that perception diminishes, Figure 3c. Thus, the underlying space may be continuous, allowing for the coding and interpretation of a variety of emotions and their combinations. This, for example, facilitates the perception of composites and multiple emotions.

Another important point to study is the effectiveness of the coding and transmission of the emotional signal and of its decoding by the cognitive system. Ideally, the coding of (categorically) different signals should be orthogonal (Smith et al., 2005). The computational face space defined in the present paper seems to agree with this view. The first identified dimension of the face space is the same for anger and sadness, but each category is defined on opposite sides of the normative face shape which minimizes any overlap between them. This maximizes the channel capacity (following Shannon's information theory), since the noise term can only have a minimum influence on the perception of emotion around the mean face. In fact, we have already shown that as we move toward the mean the percept becomes less distinguishable. This is consistent with this information theory argumentation. A transmitted signal x is received as y after this is passed through a noisy channel with conditional distribution. Assuming a Gaussian distribution, this model predicts that each category becomes clearer (i.e., less affected by noise) as we move away from the mean. This is exactly what we have observed above, and is consistent with the RBF model defined earlier in this section. Therefore, encoding two categories on opposite sides of a “normative” value on a single feature, allows for good communication between the messenger and the receiver in a noisy channel. The second dimension of the proposed computational space is orthogonal to the first one. It works by modulating the first dimension.

These results are fundamental not only for the understanding of the underlying computational face space, but for its emulation by computers. Face recognition is of primary importance in many areas of computational intelligence – ranging from human-computer interaction to content-based retrieval. If we are to build computers that interact with users in a more natural manner, it would be preferable if the representation and processing of faces was similar to that used by humans. Similarly, if computers are to retrieve or manage large amounts of data automatically, it would be preferable that be done in a manner similar to ours. The results reported in the present paper demonstrate that facial expressions of emotion can, at least in part, be recognized using simple configural cues. These image cues can be readily obtained from most images, even at low resolutions and from sketches. This will facilitate the development of computational approaches and brings us closer to understanding the cognitive mechanisms underlying face representation and processing.

Acknowledgements

This research was supported in part by a grant from the National Science Foundation. DN was supported in part by a fellowship from Ohio State's Center for Cognitive Sciences.

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

References

  1. Abbas ZA, Duchaine B. The role of holistic processing in judgments of facial attractiveness. Perception. 2008;37:1187–1196. doi: 10.1068/p5984. [DOI] [PubMed] [Google Scholar]
  2. Ackerman JM, Shapiro JR, Neuberg SL, Kenrick DT, D.V., Griskevicius V, Maner JK, Schaller M. They all look the same to me (unless they're angry) Psychological Science. 2006;17:836–840. doi: 10.1111/j.1467-9280.2006.01790.x. [DOI] [PubMed] [Google Scholar]
  3. Attneave F. Dimensions of similarity. The American Journal of Psychology. 1950;63:516–556. [PubMed] [Google Scholar]
  4. Balas BJ, Sinha P. Portraits and perception: configural information in creating and recognizing face images. Spatial Vision. 2008;21:119–135. doi: 10.1163/156856807782753949. [DOI] [PubMed] [Google Scholar]
  5. Baudouin JY, Humphreys GW. Configural information in gender categorization. Perception. 2006;35(4):531–540. doi: 10.1068/p3403. [DOI] [PubMed] [Google Scholar]
  6. Beale JM, Keil FC. Categorical effects in the perception of faces. Cognition. 1995;57:217–239. doi: 10.1016/0010-0277(95)00669-x. [DOI] [PubMed] [Google Scholar]
  7. Benson PJ, Perrett DI. Visual processing of facial distinctiveness. Perception. 1994;23:75–93. doi: 10.1068/p230075. [DOI] [PubMed] [Google Scholar]
  8. Bouvrie JV, Sinha P. Object concept learning: Observations in congenitally blind children and a computational model. Neurocomputing. 2007;70:2218–2233. [Google Scholar]
  9. Bruce V, Yong A. Understanding face recognition. British Journal of Psychology. 1986;77:305–327. doi: 10.1111/j.2044-8295.1986.tb02199.x. [DOI] [PubMed] [Google Scholar]
  10. Calder AJ, Young AW, Keane J, Dean M. Configural Information in Facial Expression Perception. Journal of Experimental Psychology: Human Perception and Performance. 2000;26:527–551. doi: 10.1037//0096-1523.26.2.527. [DOI] [PubMed] [Google Scholar]
  11. Carruthers P, Smith PK, editors. Theories of theories of mind. Cambridge University Press; 1996. [Google Scholar]
  12. Chambon V, Baudouin JY, Franck N. The role of configural information in facial emotion recognition in schizophrenia. Neuropsychologia. 2006;44:2437–2444. doi: 10.1016/j.neuropsychologia.2006.04.008. [DOI] [PubMed] [Google Scholar]
  13. Craw I, Costen N, Kato T, Akamatsu S. How Should We Represent Faces for Automatic Recognition? IEEE Trans. Pattern Analysis and Machine Intelligence. 1999;21:725–736. [Google Scholar]
  14. Darwin C. The Expression of the emotions in man and animal. J. Murray; London: 1872. [Google Scholar]
  15. Deruelle C, Mancini J, Livet MO, Cassé-Perrot C, de Schonen S. Configural and Local Processing of Faces in Children with Williams Syndrome. Brain and Cognition. 1999;41:276–298. doi: 10.1006/brcg.1999.1127. [DOI] [PubMed] [Google Scholar]
  16. Diamond R, Carey S. Why faces are and are not special: An effect of expertise. Journal of Experimental Psychology: General. 1986;115:107–117. doi: 10.1037//0096-3445.115.2.107. [DOI] [PubMed] [Google Scholar]
  17. Ding L, Martinez AM. Features versus Context: An approach for precise and detailed detection and delineation of faces and facial features. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2010 doi: 10.1109/TPAMI.2010.28. 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Duchenne G. The Mechanism of Human Facial Expression. Cambridge University Press; 1862. Reprinted in 1990. [Google Scholar]
  19. Dryden IL, Mardia KV. Statistical Shape Analysis. John Wiley & Sons; West Sussex, England: 1998. [Google Scholar]
  20. Ekman PJ, Friesen WV. The facial Action coding system: A technique for the measurement of facial movement. Consulting Psychology Press; San Diego: 1978. [Google Scholar]
  21. Ekman P. An argument for basic emotions. Cognition and Emotion. 1992;6:169–200. [Google Scholar]
  22. Farah MJ, Wilson KD, Drain M, Tanaka JN. What is “special” about face perception? Psychological Review. 1998;105:482–498. doi: 10.1037/0033-295x.105.3.482. [DOI] [PubMed] [Google Scholar]
  23. Freire A, Lee K. Face recognition in 4-to 7-year-olds: Processing of configural, featural, and paraphernalia information. Journal of Experimental Child Psychology. 2001;80:347–371. doi: 10.1006/jecp.2001.2639. [DOI] [PubMed] [Google Scholar]
  24. Gallese V, Goldman A. Mirror neurons and the simulation theory of mind-reading. Trends in Cognitive Sciences. 1998;2:493–501. doi: 10.1016/s1364-6613(98)01262-5. [DOI] [PubMed] [Google Scholar]
  25. Gauthier I, Tarr MJ. Becoming a “Greeble” expert: exploring mechanisms for face recognition. Vision Research. 1997;37:1673–1682. doi: 10.1016/s0042-6989(96)00286-6. [DOI] [PubMed] [Google Scholar]
  26. Hammal Z, Arguin M, Gosselin F. Comparing a novel model based on the transferable belief model with humans during the recognition of partially occluded facial expressions. Journal of Vision. 2009;9(2):22, 1–19. doi: 10.1167/9.2.22. [DOI] [PubMed] [Google Scholar]
  27. Hamsici OC, Martinez AM. Bayes Optimality in Linear Discriminant Analysis. IEEE Trans. on Pattern Recognition and Machine Intelligence. 2008;30:647–657. doi: 10.1109/TPAMI.2007.70717. [DOI] [PubMed] [Google Scholar]
  28. Hassin R, Trope Y. Facing faces: Studies on the cognitive aspects of physiognomy. Journal of Personality and Social Psychology. 2000;78:837–852. doi: 10.1037//0022-3514.78.5.837. [DOI] [PubMed] [Google Scholar]
  29. Hess U, Adams RB, Grammer K, Kleck RE. Face gender and emotion expression: Are angry women more like men? Journal of Vision 24. 2009;9(12) doi: 10.1167/9.12.19. Article 19. [DOI] [PubMed] [Google Scholar]
  30. Howell AJ, Buxton H. Learning identity with radial basis function networks. Neurocomputing. 1998;20:15–34. [Google Scholar]
  31. Izard CE. Emotion theory and research: Highlights, unanswered questions, and emerging issues. Annual Review of Psychology. 2009;60:1–25. doi: 10.1146/annurev.psych.60.110707.163539. [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Kagian A, Dror G, Leyvand T, Meilijson I, Cohen-Or D, Ruppin E. A machine learning predictor of facial attractiveness revealing human-like psychophysical biases. Vision Research. 2008;48:235–243. doi: 10.1016/j.visres.2007.11.007. [DOI] [PubMed] [Google Scholar]
  33. Krumhansl CL. Concerning the applicability of geometric models to similarity data: the interrelationship between similarity and spatial density. Psychological Review. 1978;85:445–463. [Google Scholar]
  34. Lebrecht S, Pierce LJ, Tarr MJ, Tanaka JW. Perceptual other-race training reduces implicit racial bias. PLoS One. 2009;4:e4215. doi: 10.1371/journal.pone.0004215. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Leder H, Bruce V. When inverted faces are recognized: The role of configural information in face recognition. Quarterly Journal of Experimental Psychology A – Human Experimental Psychology. 2000;53:513–536. doi: 10.1080/713755889. [DOI] [PubMed] [Google Scholar]
  36. Le Grand R, Mondloch CJ, Maurer D, Brent HP. Impairment in holistic face processing following early visual deprivation. Psychological Science. 2004;15:762–768. doi: 10.1111/j.0956-7976.2004.00753.x. [DOI] [PubMed] [Google Scholar]
  37. Leopold DA, Bondar IV, Giese MA. Norm-based face encoding by single neurons in the monkey inferotemporal cortex. Nature. 2006;442:572–575. doi: 10.1038/nature04951. [DOI] [PubMed] [Google Scholar]
  38. Lyons MJ, Campbell R, Plante A, Coleman M, Kamachi M, Akamatsu S. The Noh mask effect: vertical viewpoint dependence of facial expression perception. Proceedings of the Royal Society of London B. 2000;267:2239–2245. doi: 10.1098/rspb.2000.1274. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Martinez AM, Benavente R. The AR-face database. Vol. 24. Computer Vision Center; 1998. Technical Report. [Google Scholar]
  40. Martinez AM. Matching expression variant faces. Vision Research. 2003;43:1047–1060. doi: 10.1016/s0042-6989(03)00079-8. [DOI] [PubMed] [Google Scholar]
  41. Martinez AM, Kak AC. Pca versus lda. IEEE Trans. on Pattern Recognition and Machine Intelligence. 2001;23:228–233. [Google Scholar]
  42. Martinez AM, Zhu M. Where are linear feature extraction methods applicable? IEEE Trans. on Pattern Recognition and Machine Intelligence. 2005;27:1934–1944. doi: 10.1109/TPAMI.2005.250. [DOI] [PubMed] [Google Scholar]
  43. Matsumoto D, Ekman P. Japanese and Caucasian facial expressions of emotion (JACFEE) Intercultural and Emotion Research Laboratory, Department of Psychology, San Francisco State University; San Francisco, CA: 1988. [Google Scholar]
  44. Michel C, Rossion B, Han J, Chung CS, Caldara R. Holistic processing is finely tuned for faces of one's own race. Psychological Science. 2006;17:608–615. doi: 10.1111/j.1467-9280.2006.01752.x. [DOI] [PubMed] [Google Scholar]
  45. Mondloch CJ, Maurer D, Ahola S. Becoming a face expert. Psychological Science. 2006;17:930–934. doi: 10.1111/j.1467-9280.2006.01806.x. [DOI] [PubMed] [Google Scholar]
  46. Montepare JM, Dobish H. The contribution of emotion perception and their overgeneralization to trait impressions. J. Nonverbal Behavior. 2003;27:237–254. [Google Scholar]
  47. Moscovitch M, Winocur G, Behrmann M. What is special about face recognition? Nineteen experiments on a person with visual object agnosia and dyslexia but normal face recognition. Journal of Cognitive Neuroscience. 1997;9:555–604. doi: 10.1162/jocn.1997.9.5.555. [DOI] [PubMed] [Google Scholar]
  48. Neth D, Martinez AM. Emotion perception in emotionless face images suggests a norm-based representation. Journal of Vision. 2009;9(1):5, 1–11. doi: 10.1167/9.1.5. [DOI] [PubMed] [Google Scholar]
  49. Pantic M, Bartlett MS. Machine Analysis of Facial Expressions. In: Delac K, Grgic M, editors. Face Recognition. I-Tech Education and Publishing; Vienna, Austria: 2009. pp. 377–416. [Google Scholar]
  50. Paunonen SV, Ewan K, Earthy J, Lefave S, Goldberg H. Facial features as personality cues. Journal of Personality. 1999;67:555–583. [Google Scholar]
  51. Pascalis O, de Haan M, Nelson CA. Is face processing species-specific during the first year of life? Science. 2002;296:1321–1323. doi: 10.1126/science.1070223. [DOI] [PubMed] [Google Scholar]
  52. Pollak SD, Messner M, Kistler DJ, Cohn JF. Development of perceptual expertise in emotion recognition. Cognition. 2009;110:242–247. doi: 10.1016/j.cognition.2008.10.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Preston SD, de Waal FBM. Empathy: Its ultimate and proximate bases. Behavioral and Brain Sciences. 2003;25:1–20. doi: 10.1017/s0140525x02000018. [DOI] [PubMed] [Google Scholar]
  54. Rhodes G. Looking at faces: First-order and second-order features as determinants of facial appearance. Perception. 1988;17:43–63. doi: 10.1068/p170043. [DOI] [PubMed] [Google Scholar]
  55. Rhodes G, Brake S, Atkinson AP. What's lost in inverted faces? Cognition. 1993;47:25–57. doi: 10.1016/0010-0277(93)90061-y. [DOI] [PubMed] [Google Scholar]
  56. Rhodes G, Tan S, Brake S, Taylor K. Expertise and configural coding in face recognition. British journal of psychology. 1989;80:313–31. doi: 10.1111/j.2044-8295.1989.tb02323.x. [DOI] [PubMed] [Google Scholar]
  57. Rizzolatti G, Fogassi L, Gallese V. Neurophysiological mechanisms underlying the understanding and imitation of action. Nature Reviews Neuroscience. 2001;2:661–670. doi: 10.1038/35090060. [DOI] [PubMed] [Google Scholar]
  58. Riesenhuber M, Jarudi I, Gilad S, Sinha P. Face processing in humans is compatible with a simple shape-based model of vision. Proc. R. Soc. Lond. B. 2004;271(suppl.):S448–450. doi: 10.1098/rsbl.2004.0216. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Rotshtein P, Henson RNA, Treves A, Driver J, Dolan RJ. Morphing Marilyn into Maggie dissociates physical and identity face representations in the brain. Nature Neuroscience. 2006;8:107–113. doi: 10.1038/nn1370. [DOI] [PubMed] [Google Scholar]
  60. Russell JA. A circumplex model of affect. Journal of Personality and Social Psychology. 1980;39:1161–1178. [Google Scholar]
  61. Russell JA. Core affect and the psychological construction of emotion. Psychological Review. 2003;110:145–172. doi: 10.1037/0033-295x.110.1.145. [DOI] [PubMed] [Google Scholar]
  62. Schwaninger A, Wallraven C, Bülthoff HH. Computatonal modeling of face recognition based on psychophysical experiments. Swiss Journal of Psychology. 2004;63:207–215. [Google Scholar]
  63. Shin YW, Na MH, Ha TH, Kang DH, Yoo SY, Kwon JS. Dysfunction in configural face processing in patients with schizophrenia. Schizophrenia Bulletin. 2008;34:538–543. doi: 10.1093/schbul/sbm118. [DOI] [PMC free article] [PubMed] [Google Scholar]
  64. Smith ML, Cottrell GW, Gosselin F, Schyns PG. Transmitting and decoding facial expressions. Psychological Science. 2005;16(3):184–189. doi: 10.1111/j.0956-7976.2005.00801.x. [DOI] [PubMed] [Google Scholar]
  65. Sinha P. Recognizing complex patterns. Nature Neuroscience. 2002;5(suppl.):1093–1097. doi: 10.1038/nn949. [DOI] [PubMed] [Google Scholar]
  66. Sinha P, Poggio T. I think I know that face…. Nature. 1996;384:404. doi: 10.1038/384404a0. [DOI] [PubMed] [Google Scholar]
  67. Tanaka JW, Farah MJ. Parts and Wholes in face recognition. Quarterly Journal of Experimental Psychology A – Human Experimental Psychology. 1993;46:225–245. doi: 10.1080/14640749308401045. [DOI] [PubMed] [Google Scholar]
  68. Tanaka JW, Kiefer M, Bukach C. A holistic account of the own-race effect in face recognition: evidence from a cross-cultural study. Cognition. 2004;93:B1–B9. doi: 10.1016/j.cognition.2003.09.011. [DOI] [PubMed] [Google Scholar]
  69. Tanaka JW, Sengco J. Features and their configuration in face recognition. Memory & Cognition. 1997;25:583–592. doi: 10.3758/bf03211301. [DOI] [PubMed] [Google Scholar]
  70. Tsao DY, Freiwald WA, Tootell RBH, Livingstone MS. A cortical region consisting entirely of face-selective cells. Science. 2006;311:670–674. doi: 10.1126/science.1119983. [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Tversky A. Features of similarity. Psychological Review. 1977;84:327–352. [Google Scholar]
  72. Valentine T. A unified account of the effects of distinctiveness, inversion, and race in face recognition. The Quarterly Journal of Experimental Psychology. 1991;43A:161–204. doi: 10.1080/14640749108400966. [DOI] [PubMed] [Google Scholar]
  73. Vollm BA, Taylor ANW, Richardson P, Corcoran R, Stirling J, McKie S, Deakin JFW, Elliott R. Neuronal correlates of theory of mind and empathy: A functional magnetic resonance imaging study in a nonverbal task. Neuroimage. 2006;29:90–98. doi: 10.1016/j.neuroimage.2005.07.022. [DOI] [PubMed] [Google Scholar]
  74. White M. Representation of facial expressions of emotion. American Journal of Psychology. 1999;112:371–381. [PubMed] [Google Scholar]
  75. Wilbraham DA, Christensen JC, Martinez AM, Todd JT. Can low level image differences account for the ability of human observers to discriminate facial identity? Journal of Vision. 2008;8(15):5, 1–12. doi: 10.1167/8.15.5. [DOI] [PubMed] [Google Scholar]
  76. Yin RK. Looking at upside-down faces. Journal of Experimental Psychology. 1969;81:141–145. [Google Scholar]
  77. Young AW, Rowland D, Calder AJ, Etcoff NL, Seth A, Perrett DI. Facial expression megamix: Tests of dimensional and category accounts of emotion recognition. Cognition. 1997;63:271–313. doi: 10.1016/s0010-0277(97)00003-6. [DOI] [PubMed] [Google Scholar]
  78. Zebrowitz LA. Reading faces: window to the soul? Westview Press; Boulder, CO: 1997. [Google Scholar]
  79. Zebrowitz LA, Fellous J. Facial resemblance to Emotion: Group Differences, Impression Effects, and Race Stereotypes. Journal of Personality and Social Psychology. 2010 doi: 10.1037/a0017990. [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES