Skip to main content
The Journal of the Acoustical Society of America logoLink to The Journal of the Acoustical Society of America
. 2019 May 21;145(5):EL423–EL429. doi: 10.1121/1.5103191

Differentiating post-cancer from healthy tongue muscle coordination patterns during speech using deep learning

Jonghye Woo 1,a),, Fangxu Xing 1, Jerry L Prince 2, Maureen Stone 3, Jordan R Green 4, Tessa Goldsmith 5, Timothy G Reese 6, Van J Wedeen 6, Georges El Fakhri 1
PMCID: PMC6530633  PMID: 31153323

Abstract

The ability to differentiate post-cancer from healthy tongue muscle coordination patterns is necessary for the advancement of speech motor control theories and for the development of therapeutic and rehabilitative strategies. A deep learning approach is presented to classify two groups using muscle coordination patterns from magnetic resonance imaging (MRI). The proposed method uses tagged-MRI to track the tongue's internal tissue points and atlas-driven non-negative matrix factorization to reduce the dimensionality of the deformation fields. A convolutional neural network is applied to the classification task yielding an accuracy of 96.90%, offering the potential to the development of therapeutic or rehabilitative strategies in speech-related disorders.

1. Introduction

The human tongue is a highly complex muscular structure composed of several paired intrinsic and extrinsic muscles that are critical for speaking, swallowing, and breathing (Kier and Smith, 1985). In order to produce intelligible speech, a variety of internal tongue muscle groupings, i.e., functional units, are temporarily formed, rapidly and nimbly, in a highly coordinated fashion (Green and Wang, 2003; Ramanarayanan et al., 2013; Stone et al., 2004; Woo et al., 2019). For individuals who have difficulty with speech production following tongue cancer, the location and size of functional units obtained from the same speech movements are likely to be different from those of healthy controls because of surgical modifications or compensatory strategies used to maximize speech intelligibility (Stone et al., 2014). In addition, the extent to which tongue function for speech is impaired due to resected muscle groups as in tongue cancer (Goldsmith and Jacobson, 2015) varies depending on the location and size of tumors resected, which leads to a large variability in the local muscle coordination patterns during speech. Due to the large variability and complexity of tongue motion during speech for both healthy controls and patients, it remains challenging to detect and characterize the similarities and differences of functional units between controls and patients. Differentiating post-cancer from healthy tongue muscle coordination patterns is, however, an important part of the development and verification of speech motor control theories and in the development of new therapeutic, surgical, or rehabilitative strategies.

In recent years, machine/deep learning has produced state-of-the-art results in a large number of tasks such as recognition, prediction, and classification from complex sets of data (LeCun et al., 2015). In the context of tongue motion analysis for speech using various imaging and motion capture techniques, several researchers have studied the classification of speech movements using machine learning. For example, Tang et al. (2011) investigated spatiotemporal gestural descriptors from ultrasound for classifying speech movements using a support vector machine (SVM). Wang et al. (2016) examined a set of flesh points on the tongue and lips that are optimal to classify speech movements via a SVM from electromagnetic articulographs. To the best of our knowledge, no attempts have been made for the classification task using tagged-magnetic resonance imaging (MRI) previously. For an extensive review on scientific and clinical applications of speech movement analysis using machine learning, readers can refer to the paper by Green (2015) and the references therein.

In this work, we propose a deep learning classification framework to differentiate post-cancer from healthy motion coordination patterns after tongue cancer treatment using tagged-MRI (Parthasarathy et al., 2007) and voxel level tracking (Xing et al., 2017) for the first time. To achieve this goal and mitigate the aforementioned challenges, we use a four-dimensional (4D) (three-dimensional space with time) atlas of tongue motion from tagged-MRI (Woo et al., 2017) and atlas-driven non-negative matrix factorization (NMF) (Woo et al., 2019), to characterize the standardized muscle coordination patterns, and a deep learning framework to build an accurate and interpretable classification framework. The use of a 4D atlas of tongue motion in combination with NMF is crucial because it provides a benchmark to compare differences of internal muscle coordination patterns in an objective manner. We show that this classification occurs within the convolutional neural network (CNN) with the implicit yet interpretable structures of muscle coordination patterns derived from atlas-driven NMF.

2. Materials and methods

2.1. Subjects and MRI acquisition

A total of 26 subjects including 18 healthy controls and 8 tongue cancer patients after treatment participated in this study. Each subject was trained prior to the scan to speak a simple utterance (“ə-suk”) in synchrony with a repetitive (metronome-like) sound. During the scan, the subjects repeated the utterance following the repetitive sound while MRI scans were performed. Both T2-weighted two-dimensional (2D) dynamic cine-MRI and tagged-MRI images were acquired using a Siemens 3.0 T Tim Trio system with a 16-channel head and neck coil (Siemens Medical Solutions, Erlangen, Germany). These dynamic acquisitions were acquired 3 times each—in axial, coronal, and sagittal orientations—at 26 frames per second for a duration of 1 s on identical slice positions. Table 1 shows the characteristics of the patients including speech intelligibility, tumor size, and closure type.

Table 1.

Characteristics of patients.

Subjects Speech Intelligibility Tumor Size Closure Type
Patients 1–5 100% T1 Primary
Patient 6 99% T2 Radial forearm free flap
Patient 7 95% T2 Primary
Patient 8 98% T2 Primary

2.2. Our method

Tissue point tracking from tagged-MRI. In order to track each voxel within the tongue from tagged-MRI, we use the previously developed method called phase vector incompressible registration algorithm (PVIRA) (Xing et al., 2017). First, cubic B-spline interpolation is used to generate dense 2D slices from the sparsely acquired tagged-MRI slices. All three orientations (axial, coronal, and sagittal acquisitions) are interpolated onto the same digital grid. A harmonic phase filter (Osman et al., 1999) is then used to yield three harmonic phase (HARP) volumes from the interpolated results. Finally, the HARP phase volumes are input into an iLogDemons algorithm (Mansi et al., 2011) (modified to accept three phase volumes) to find accurate correspondences across time frames. This tracking method yields accurate and incompressible motion fields that capture the dynamic motion of the tongue during speech.

4D atlas building of tongue motion. We build a Lagrangian 4D atlas of tongue motion (ə-suk) from 18 healthy controls as follows. First, we use a super-resolution volume reconstruction technique (Woo et al., 2012) to combine three orthogonal stacks of cine images, yielding 26 volumes with isotropic resolution for each subject. Second, we establish a common Lagrangian coordinate system across subjects by using the first time frame (representing a neutral position) of the 18 cine super-resolved volumes. This coordinate system is found using a symmetric diffeomorphic groupwise registration with a cross-correlation similarity measure (Avants et al., 2008). Third, we transform the motion fields computed by PVIRA into this common coordinate system, thus putting them into correspondence. Finally, we average all the motion fields registered at each voxel to yield the final 4D Lagrangian atlas.

Weighting map estimation from atlas-driven NMF. Once we obtain the 4D atlas in the common spatial coordinate system, we extract motion features from 26 time frames including the magnitude and angle of point tracks (Woo et al., 2019). Those features are aggregated to form an input matrix for graph-regularized sparse non-negative matrix factorization (GS-NMF). The GS-NMF produces two output matrices: (1) spatial building blocks reflecting muscle groups, which changes over time, and (2) the associated weighting map. The weighting map is then used to identify the coherent regions, thus revealing spatiotemporally varying functional units (Woo et al., 2019). The 4D motion atlas of tongue motion using the same speech task in a healthy population represents standardized motion patterns. Therefore, the building blocks and their weighting maps identified from the 4D atlas motion can be seen as standardized building blocks with weighting maps that can generate the standardized motion patterns. More specifically, let UA be a non-negative feature matrix constructed by the motion features from the 4D atlas as shown in Fig. 1, where K ≤ min(m, n), and let VA=[vik]+m×K be the building blocks and WA=[wkj]+K×n be their weighting map, respectively. The goal of GS-NMF is to find two matrices while maintaining intrinsic and geometric relationship between features given by

UAVAWA. (1)

Fig. 1.

Fig. 1.

(Color online) A schematic illustration of obtaining weighting maps used for the classification task. Once we obtain the building blocks from the 4D motion atlas, we fix the building blocks to obtain objective weighting maps across controls and patients. The obtained objective weighting maps are used for the subsequent classification task.

Because the building blocks obtained from the 4D atlas and GS-NMF can be interpreted as the standardized aggregated information that defines the muscle group types given the speech task, the weighting map provides the tissue coordinate location and weights of each muscle group type to generate the phonemic movements that produce the speech. We then compute weighting maps for each subject by using the same building blocks from the 4D atlas so that the obtained weighting maps provide objective weights, i.e., a proxy structure containing implicit muscle coordination patterns, associated with the standardized muscle groups given by

USAVAWSA, (2)

where USA represents an input feature matrix derived in the atlas space for each subject and WSA represents the obtained weighting map, while VA remains fixed. This facilitates the objective comparison of muscle coordination patterns across controls and patients. We obtain the weighting maps using different combinations of phonemes including (1) /ə/ - /s/, (2) /ə/ - /u/, (3) /ə/ - /k/, (4) /s/ - /u/, (5) /s/ - /k/, and (6) /u/ - /k/, each subject having six unique weighting maps.

Deep learning-based classification. The numbers of the weighting maps estimated for controls and patients are 108 and 48, respectively. With this limited training data, it is likely that many of the complicated relationships could be attributed to sampling noise and therefore they will exist in the training data only but not in test data, which thus leads to overfitting. To address this, we perform data augmentation on weighting maps using combinations of scaling and rotation, thereby yielding 936 weighting maps in total (648 weighting maps for controls and 288 weighting maps for patients).

The original and augmented weighting maps are then input into a CNN classification system. The architecture of our CNN classification system is summarized in Fig. 2. The architecture contains six learned layers—five convolutional and one fully-connected. We follow each convolutional layer with a ReLU activation function, a batch normalization layer, and a max-pooling layer, except for the fully-connected layer where a softmax activation function is used. Dropout (Hinton et al., 2012) is not used, following the suggestion in Ioffe and Szegedy (2015), and for optimization we use Adam optimization with the default learning rate. We use binary cross-entropy as the loss function. Our CNN is trained on an GTX 1080 Ti GPU (NVIDIA, Santa Clara, CA) using 936 weighting maps from controls and patients. Our implementation is based on Keras with Tensorflow as a backend.

Fig. 2.

Fig. 2.

(Color online) A schematic illustration of our deep learning architecture.

3. Validation

Our method was compared to three other classification methods including a SVM with a radial basis function kernel, a random forest (RF), and a multilayer perceptron (MLP). All three were optimized independently using the same training data that we used to train our CNN in order to achieve the best classification performance. To measure the classification accuracy, we used five-fold cross-validation. The results using our approach and other competing approaches were statistically compared using the Student t-test with a level of significance set at p < 0.05.

4. Results

Example motion fields from our 4D atlas of ə-suk are shown in the top row of Fig. 3 and three clusters of example functional units for the transitions /ə/ - /s/, /s/ - /u/, and /u/ - /k/ are shown in the bottom row of Fig. 3. The bottom row shows the location of three functional units that reflect and characterize the averaged motions in the top row. The motion from /ə/ to /s/ includes minimal motion in the tongue base, forward motion in the posterior tongue, and forward/upward motion in the tip and central body. The functional units capture these three motion regions and create clear borders between them. The /s/ to /u/ motion shows functional divisions between the anterior and posterior tongue. The movement regions in the posterior are not well delineated, suggesting integrated regions of motion in the back that differ from the one in the front. The functional units for the /u/ to /k/ motion capture three different functional units, showing the upper tongue divided into an anterior and posterior region, each moving toward the center as the tongue body is elevated.

Fig. 3.

Fig. 3.

(Color online) Illustration of the motion fields of a 4D motion atlas, ə-suk, in a Lagrangian coordinate system (top row), and identified functional units (bottom row). The motion fields and the functional unit results are depicted relative to the neutral tongue position. The cone size is the magnitude of the motion fields, and red, green, and blue colors are left–right, front–back, and up–down directions of the motion fields in the top row, respectively. Of note, the colors in the bottom row represent different units, not directions.

Figure 4 illustrates the motion fields, functional units, and associated weighting maps for the transitions /ə/ - /s/, /s/ - /u/, and /u/ - /k/ for a control and a patient (patient 4). The second and third row display the location of three functional units and associated weighting maps that reflect and characterize the averaged motions in the first row, respectively. For the weighting maps, each axis represents a different parameter. The left axis represents spatial locations within the tongue, the right axis represents reduced motion features (in this work, the number was set to 20), and the height axis represents the magnitude of weights. For the control, similar to Fig. 3, the motion from /ə/ to /s/ includes motion in the tongue base, forward motion in the posterior tongue, and forward/upward motion in the tip. The /s/ to /u/ motion shows functional divisions between the anterior and posterior tongue. The functional units for the /u/ to /k/ motion show the upper tongue divided into an anterior and posterior region. For the patient, the motion from /ə/ to /s/ includes motion in the tongue base, forward motion in the posterior tongue, and forward/upward motion in the tip and lower body. The /s/ to /u/ motion shows functional divisions among the tip, posterior region, and base of the tongue. The functional units for the /u/ to /k/ motion capture three different functional units, showing the tip, base, and body of the tongue. As shown in Fig. 4, it can be clearly observed that the tongue's tip and base of the patient have a tendency to act as separate functional units over time. However, for the control, at certain time periods (such as from /s/ to /u/), the tongue's tip and base were identified as one unified unit. This could be contributed by the fact that the control's internal tongue muscles are more active than the patient so that they could cooperate more freely, even connecting the tip and base. However, the patient appeared to lose that capability for a broader muscle connection due to the surgical removal of tumors.

Fig. 4.

Fig. 4.

(Color online) Illustration of the motion fields (first row), three clusters of functional units (second row), and associated weighting maps (third row) for the transitions /ə/ - /s/, /s/ - /u/, and /u/ - /k/ in a Lagrangian coordinate system for a control and a patient.

We validated the effectiveness of our approach using five-fold cross-validation. In short, we first divided the whole sample into five disjoint subsets of approximately equal size. We then trained each model 5 times in which for each run, four of the subsets were used for training, while leaving out one of the subsets from the training, and using only the left out subset to compute the predictive ability of the model. Accuracy was measured by how effectively we could classify motion patterns into healthy versus patient groups. Our approach achieved an accuracy of 96.90 ± 2.12%, while SVM, RF, and MLP achieved an accuracy of 83.33 ± 1.92%, 84.40 ± 2.66%, and 84.72 ± 2.90%, respectively. The Student's t-test results demonstrate that there is a statistically significant difference (p < 0.05) in the classification accuracy between our method and each competing method.

5. Discussion and summary

In this work, we presented a method to discriminate tongue motion patterns during speech in tongue cancer patients and healthy controls from tagged-MRI via a deep learning classification approach. To overcome the greater complexity and variability inherent in internal tongue muscle coordination patterns, a 4D atlas from tagged-MRI and atlas-driven NMF was used to yield a low-dimensional yet interpretable and referential structure to compare the motion patterns of the tongue between patients with tongue cancer and healthy controls. Our results show that our method achieves over 95% classification accuracy and is statistically superior to competing approaches. Upon analysis of the different speech mechanisms revealed by our approach, we may better understand the underlying mechanisms of muscle coordination patterns and potentially contributing to the development of therapeutic, rehabilitative, and surgical strategies.

In order to learn physiologically interpretable feature representation as an input to our classification model instead of a variety of individual motion features, an NMF-based approach (Woo et al., 2019) was used for dimension reduction while retaining interpretability through weighting maps which implicitly contain muscle coordination patterns. This is partly attributed to the fact that NMF offers parts-based representation of motion features resulting from muscle activity that are inherently non-negative. Of note, it is, however, difficult to visually assess the number of clusters, the locations and sizes of the function units directly from the weighting maps since the weighting maps are low-dimensional structures. Also due to the high variability in the identified functional units, it is challenging to directly compare and contrast the weighting maps between the control and the patient as shown in Fig. 4. It is, therefore, necessary to use computerized algorithms including the spectral clustering algorithm and deep learning to identify the size and location of the functional units and classify between controls and patients, respectively. This is in contrast with other dimension reduction methods such as Principal Component Analysis (PCA) (Xing et al., 2016) that is used as a versatile tool for dimension reduction and visualization and Independent Component Analysis (ICA) that is used for separation of the data into a set of statistically independent components. The decomposition results derived from PCA and ICA are, however, limited to orthogonal and independent basis vectors, thereby lacking physiological interpretability of the underlying data.

In the present work, we employed a Lagrangian configuration as in Woo et al. (2017) based on material coordinates. This Lagrangian coordinate system in conjunction with the 4D atlas greatly simplified all the computations of motion fields, weighting maps, and subsequent comparison using classifiers. The improved classification accuracy can, in part, be attributed to the minimization of across-subject difference in tongue motor function, which presumably allowed for the identification of common patterns of tongue muscle coordination within each participant group. The Lagrangian coordinate system used to simplify tongue motion data analysis is analogous with the coordinate system previously established in other organs including the “motionless movies” concept (Wedeen et al., 1995) to analyze myocardial strain-rates of the heart or referential structures of the brain (Wedeen et al., 2012) to map multimodal and multiscale data to characterize brain structure and function at different scales from a population of interest.

Although the patient population studied in this work was limited to patients who underwent a partial glossectomy, we anticipate that the novel methods used to identify and characterize tongue impairment will be relevant to a large number of individuals with anatomic and neurologically-based speech impairments such as Amyotrophic Lateral Sclerosis (ALS). For example, the extent to which tongue function is impaired due to muscle weakness as in ALS varies depending on the severity of the disease, which also leads to large variability in the local muscle coordination patterns. Likewise, our proposed method could be applied to the ALS population for characterizing functional units and subsequently distinguishing between persons with ALS and healthy controls.

There are a few ways to expand on the present work. First, deep learning methods are typically trained on a large number of datasets and deep learning architectures have many layers. However, the present work is limited by the number of subjects used for training and testing. In future work, we will increase the number of subjects to build our model and test our algorithm. In addition, we will stratify the patient group based on the location and size of tumors resected for tongue cancer patients, or different disease states of ALS. This more granular participant stratification will allow us to investigate the multi-class categorization problem. Second, in this work, we used one speech word to build the model and to test our algorithm. In future work, we will also diversify speech words to establish a larger database and further build training models for various machine learning tasks such as classification, recognition, and prediction. Third, we only used motion quantities including magnitude and angle derived from displacements. In future work, we will use principal strains to investigate various cause-and-effect relationships among muscle groups (Buchaillard et al., 2008).

Acknowledgments

This work was partially supported by the National Institute of Health (NIH) Grant Nos. R00DC012575, R21DC016047, R01DC014717, R01CA133015, and P41EB022544.

Contributor Information

Jonghye Woo, Email: .

Fangxu Xing, Email: .

Jerry L. Prince, Email: .

Maureen Stone, Email: .

Jordan R. Green, Email: .

Tessa Goldsmith, Email: .

Timothy G. Reese, Email: .

Van J. Wedeen, Email: .

Georges El Fakhri, Email: .

References and links

  • 1. Avants, B. B. , Epstein, C. L. , Grossman, M. , and Gee, J. C. (2008). “ Symmetric diffeomorphic image registration with cross-correlation: Evaluating automated labeling of elderly and neurodegenerative brain,” Med. Image Anal. 12(1), 26–41. 10.1016/j.media.2007.06.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Breiman, L. (2001). “ Random forests,” Machine Learn. 45(1), 5–32. 10.1023/A:1010933404324 [DOI] [Google Scholar]
  • 3. Buchaillard, S. , Perrier, P. , and Payan, Y. (2008). “ To what extent does tagged-mri technique allow to infer tongue muscles' activation pattern? A modelling study,” in 9th Annual Conference of the International Speech Communication Association (Interspeech 2008), pp. 2839–2842. [Google Scholar]
  • 4. Goldsmith, T. , and Jacobson, M. C. (2015). “ Revisiting swallowing function following contemporary surgical interventions for oral/oropharyngeal cancer: Key underlying issues,” Perspectives Swallowing Swallowing Disorders (Dysphagia) 24(3), 89–98. 10.1044/sasd24.3.89 [DOI] [Google Scholar]
  • 5. Green, J. R. (2015). “ Mouth matters: Scientific and clinical applications of speech movement analysis,” Perspectives Speech Sci. Orofacial Disorders 25(1), 6–16. 10.1044/ssod25.1.6 [DOI] [Google Scholar]
  • 6. Green, J. R. , and Wang, Y.-T. (2003). “ Tongue-surface movement patterns during speech and swallowing,” J. Acoust. Soc. Am. 113(5), 2820–2833. 10.1121/1.1562646 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Hinton, G. E. , Srivastava, N. , Krizhevsky, A. , Sutskever, I. , and Salakhutdinov, R. R. (2012). “ Improving neural networks by preventing co-adaptation of feature detectors,” arXiv:1207.0580.
  • 8. Ioffe, S. , and Szegedy, C. (2015). “ Batch normalization: Accelerating deep network training by reducing internal covariate shift,” arXiv:1502.03167.
  • 9. Kier, W. M. , and Smith, K. K. (1985). “ Tongues, tentacles and trunks: The biomechanics of movement in muscular-hydrostats,” Zoological J. Linnean Soc. 83(4), 307–324. 10.1111/j.1096-3642.1985.tb01178.x [DOI] [Google Scholar]
  • 10. LeCun, Y. , Bengio, Y. , and Hinton, G. (2015). “ Deep learning,” Nature 521(7553), 436–444. 10.1038/nature14539 [DOI] [PubMed] [Google Scholar]
  • 11. Mansi, T. , Pennec, X. , Sermesant, M. , Delingette, H. , and Ayache, N. (2011). “ ilogdemons: A demons-based registration algorithm for tracking incompressible elastic biological tissues,” Int. J. Comp. Vis. 92(1), 92–111. 10.1007/s11263-010-0405-z [DOI] [Google Scholar]
  • 12. Osman, N. F. , Kerwin, W. S. , McVeigh, E. R. , and Prince, J. L. (1999). “ Cardiac motion tracking using cine harmonic phase (HARP) magnetic resonance imaging,” Magnetic Resonance Med. 42(6), 1048–1060. 10.1002/(SICI)1522-2594(199912)42:6<1048::AID-MRM9>3.0.CO;2-M [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Parthasarathy, V. , Prince, J. L. , Stone, M. , Murano, E. Z. , and NessAiver, M. (2007). “ Measuring tongue motion from tagged cine-MRI using harmonic phase (HARP) processing,” J. Acoust. Soc. Am. 121(1), 491–504. 10.1121/1.2363926 [DOI] [PubMed] [Google Scholar]
  • 14. Ramanarayanan, V. , Goldstein, L. , and Narayanan, S. S. (2013). “ Spatio-temporal articulatory movement primitives during speech production: Extraction, interpretation, and validation,” J. Acoust. Soc. Am. 134(2), 1378–1394. 10.1121/1.4812765 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Stone, M. , Epstein, M. A. , and Iskarous, K. (2004). “ Functional segments in tongue movement,” Clinical Ling. Phonetics 18(6–8), 507–521. 10.1080/02699200410003583 [DOI] [PubMed] [Google Scholar]
  • 16. Stone, M. , Langguth, J. M. , Woo, J. , Chen, H. , and Prince, J. L. (2014). “ Tongue motion patterns in post-glossectomy and typical speakers: A principal components analysis,” J. Speech, Lang., Hear. Res. 57(3), 707–717. 10.1044/1092-4388(2013/13-0085) [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Tang, L. , Hamarneh, G. , and Bressmann, T. (2011). “ A machine learning approach to tongue motion analysis in 2d ultrasound image sequences,” in International Workshop on Machine Learning in Medical Imaging ( Springer, New York: ), pp. 151–158. [Google Scholar]
  • 18. Wang, J. , Samal, A. , Rong, P. , and Green, J. R. (2016). “ An optimal set of flesh points on tongue and lips for speech-movement classification,” J. Speech, Lang., Hear. Res. 59(1), 15–26. 10.1044/2015_JSLHR-S-14-0112 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Wedeen, V. J. , Rosene, D. L. , Wang, R. , Dai, G. , Mortazavi, F. , Hagmann, P. , Kaas, J. H. , and Tseng, W.-Y. I. (2012). “ The geometric structure of the brain fiber pathways,” Science 335(6076), 1628–1634. 10.1126/science.1215280 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Wedeen, V. J. , Weisskoff, R. M. , Reese, T. G. , Beache, G. M. , Poncelet, B. P. , Rosen, B. R. , and Dinsmore, R. E. (1995). “ Motionless movies of myocardial strain-rates using stimulated echoes,” Magnetic Reson. Med. 33(3), 401–408. 10.1002/mrm.1910330313 [DOI] [PubMed] [Google Scholar]
  • 21. Woo, J. , Murano, E. Z. , Stone, M. , and Prince, J. L. (2012). “ Reconstruction of high-resolution tongue volumes from MRI,” IEEE Trans. Biomed. Eng. 59(12), 3511–3524. 10.1109/TBME.2012.2218246 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Woo, J. , Prince, J. , Stone, M. , Xing, F. , Gomez, A. , Green, J. , Hartnick C., Brady, T. , Reese, T. , Wedeen, V. , and El Fakhri, G. (2019). “ A sparse non-negative matrix factorization framework for identifying functional units of tongue behavior from MRI,” IEEE Trans. Med. Imag. 38(3), 730–740. 10.1109/TMI.2018.2870939 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Woo, J. , Xing, F. , Stone, M. , Green, J. , Reese, T. G. , Brady, T. J. , Wedeen, V. J. , Prince, J. L. , and El Fakhri, G. (2017). “ Speech map: A statistical multimodal atlas of 4d tongue motion during speech from tagged and cine MR images,” Comp. Methods Biomech. Biomed. Eng. 1–13. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Xing, F. , Woo, J. , Gomez, A. D. , Pham, D. L. , Bayly, P. V. , Stone, M. , and Prince, J. L. (2017). “ Phase vector incompressible registration algorithm for motion estimation from tagged magnetic resonance images,” IEEE Trans. Med. Imag. 36(10), 2116–2128. 10.1109/TMI.2017.2723021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Xing, F. , Woo, J. , Lee, J. , Murano, E. Z. , Stone, M. , and Prince, J. L. (2016). “ Analysis of 3-d tongue motion from tagged and cine magnetic resonance images,” J. Speech, Lang., Hear. Res. 59(3), 468–479. 10.1044/2016_JSLHR-S-14-0155 [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from The Journal of the Acoustical Society of America are provided here courtesy of Acoustical Society of America

RESOURCES