Abstract
Modern radiotherapy stands to benefit from the ability to efficiently adapt plans during treatment in response to setup and geometric variations such as those caused by internal organ deformation or tumor shrinkage. A promising strategy is to develop a framework, which given an initial state defined by patient-attributes, can predict future states based on pre-learned patterns from a well-defined patient population. Here, we investigate the feasibility of predicting patient anatomical changes, defined as a joint state of volume and daily setup changes, across a fractionated treatment schedule using two approaches. The first is based on a new joint framework employing quantum mechanics in combination with deep recurrent neural networks, denoted QRNN. The second approach is developed based on a classical framework, which models patient changes as a Markov process, denoted MRNN. We evaluated the performance characteristics of these two approaches on a dataset of 125 head and neck cancer patients, which was supplemented by synthetic data generated using a generative adversarial network. Model performance was evaluated using area under the receiver operating characteristic curve (AUC) scores. The MRNN framework had slightly better performance than the QRNN framework, with MRNN(QRNN) validation AUC scores of 0.742 ± 0.021 (0.675 ± 0.036), 0.709 ± 0.026 (0.656 ± 0.02l), 0.724 ± 0.036 (0.652 ± 044), and 0.698 ± 0.016 (0.605 ± 0.035) for system state vector sizes of 4, 6, 8, and 10, respectively. Of these, only the results from the two higher order states had statistically significant differences (p < 0.05). A similar trend was also observed when the models were applied to an external testing dataset of 20 patients, yielding MRNN(QRNN) AUC scores of 0.707 (0.623), 0.687 (0.608), 0.723 (0.669), and 0.697 (0.609) for states vectors sizes of 4, 6, 8, and 10, respectively. These results suggest that both stochastic models have potential value in predicting patient changes during the course of adaptive radiotherapy.
Keywords: Adaptive Radiotherapy, Head and Neck Cancer, Deep Learning, Quantum Computing, Markov Process
1. Introduction
1.1. Adaptive radiotherapy
Radiation therapy (RT) has been widely established as the standard of care for numerous cancer types, with as many as 50% of cancer patients who receive RT as part of their treatment each year (1). Innovations in treatment delivery, such as the use of multi-leaf collimators during intensity modulated radiation therapy (IMRT) and volumetric arc therapy (VMAT) to dynamically move and shape the beam during treatment, have made it possible to achieve dose distributions that conform tightly around the tumor with steep dose gradients, allowing for the delivery of higher, more curative doses to treat the disease while minimizing the dose delivered to the surrounding healthy tissues (2, 3).
Despite these advancements, a challenge remains in that over the course of fractionated treatment, factors such as weight loss, tumor shrinkage, or daily anatomical variations can cause geometrical changes in the patient’s anatomy such that the original treatment plan is no longer optimal – either due to inadequate target coverage or due to increased exposure to organs at risk (OARs). Clinics often track anatomical changes through daily cone-beam computed tomography (CBCT) images taken prior to each treatment; if a patient’s plan is deemed to stray too far from the original treatment objectives, alterations can be made, which at their most extreme involve a full replan. The process of using imaging data to make adaptations to a patient’s plan during treatment is called image-guided adaptive radiotherapy (IG-ART).
While IG-ART has been shown to improve the accuracy and quality of patient treatment plans (4, 5), several challenges continue to inhibit wider implementation:
Clinical time and resource constraints necessitate that IG-ART methods be efficient (6).
In addition, because CBCT images suffer from higher noise and a limited field of view (7), daily anatomical changes can be difficult to quantify with accuracy – necessitating IG-ART plans to be robust to measurement uncertainties.
Finally, while it has been shown that IG-ART can improve treatment accuracy, there exists no standardized set of criteria for how much a patient’s plan needs to deviate from its intended objectives in order to trigger the use of ART methods (8).
A consequence of these challenges and their time-consuming processes is that some patients do not undergo necessary plan adaptations who might otherwise benefit from them to optimize their treatment outcomes. One approach which addresses the first of these obstacles is to develop a framework which, based on relevant patient characteristics, predicts the likelihood of patient anatomical changes occurring in future fractions. Such a framework would allow for a proactive approach to ART in contrast to passive image deformation approaches, effectively buying time for busy clinics by allowing them to anticipate and prepare for plan adaptations before they occur.
1.2. Predictive Models for ART
The development of predictive models to aid in decision making for ART is an active research topic. Predictive models typically utilize imaging data (image-guided models) and/or clinical, dosimetric, biological, and radiomic data (knowledge-based models) (9). Given the prevalence of daily imaging for patient setup, models based on image guidance are arguably the easiest to implement into a clinical workflow.
There is currently no set consensus regarding what types of predictions are most valuable or necessary for an ART paradigm. Recent studies have reported on models for predicting TCP/NTCP (10, 11), dosimetric deviations from the original treatment plan (12), as well as incidence of radiation-induced toxicity (13) and tumor failure (14). This variety of approaches is due in part because the objectives for implementing ART may differ across cancer-type and/or treatment clinic, and can vary from improving dosimetric accuracy, reducing tumor failure, or preventing specific radiation induced toxicities such as radiation pneumonitis or xerostomia. Adding to this challenge is that there is still an incomplete understanding between the relationship of anatomical/dosimetric deviations during treatment and clinically meaningful outcomes (15).
Recent studies have reported on the use of geometric anatomical changes (measured via CT imaging) as a condition for triggering plan adaptation and linked anatomical changes to significant, undesirable dosimetric outcomes (16, 17). Some studies have also begun to look at predicting geometric changes during radiotherapy. Limitations of these geometric predictive models include small sample sizes (N=10–34), predicting only one fraction into the future, and the exclusion of patients who did not experience tumor shrinkage (18–20).
1.3. Quantum Predictive Models and Quantum Processes
Recent work has found that quantum mechanics can be applied successfully as a predictive framework for systems outside of fundamental physics – namely human cognition and decision making (21). In the next two sections we provide a brief review of the quantum predictive framework as well as the Markov predictive framework. For a more complete formulation of quantum mechanics postulates, interested readers are referred to textbooks on the subject (22).
A quantum system can be fully described by its unitary state vector, which can be written as a weighted sum of orthonormal vectors:
| (1) |
Where the set of vectors |vi∈N〉 form an orthonormal basis in a finite N-dimensional Hilbert space. The time evolution of a stationary quantum state is described by the time-independent Schrödinger equation:
| (2) |
Whose solution takes the form:
| (3) |
| (4) |
Where |Ψ(t)〉 is the state at time, t, and U(t + n) is a unitary transition matrix defined by a Hermitian operator, H, called the Hamiltonian. (Note that in Eqn. 4, Plank’s constant (denoted by ħ) has been absorbed into H.)
1.4. Classical Markov Processes
A Markov process refers to a classical stochastic process which has no memory of its past and is influenced only by its current state (23). Given a finite number, N, of potential states, the state probability distribution can be written as an N × l column vector, M(t). Each entry of M(t) represents the probability of finding the system in each state at time t, such that:
| (5) |
Similar to a quantum process, a Markov process evolves over time according to a transition matrix, P(t), such that:
| (6) |
The time evolution of this state transition matrix is described by the Kolmogorov forward and backward equations:
| (7) |
| (8) |
An expression for P(t) can therefore be written as:
| (9) |
Where Q is referred to as the generator matrix, which satisfies the conditions:
| (10) |
| (11) |
The results from Eqns. (6) and (9) can be used to describe the time evolution of the Markov system:
| (12) |
1.5. Deep Learning and Recursive Neural Networks
Deep learning refers to a broad class of machine learning methods which are capable of learning latent/higher order characteristics of a dataset in order to achieve a user-defined goal (typically the minimization of an objective function). While there are no structural requirements for deep learning algorithms, the most widely successful implementations of deep learning have been realized through artificial neural networks (ANNs), which have enjoyed widespread success across a diverse range of applications in medicine and hold the potential to revolutionize the field of radiation oncology (24).
Recursive neural networks (RNNs) are neural networks which can handle sequence-data (25). RNNs differ from traditional ANNs in that all of the nodes in an RNN share the same weights and biases. In addition, because they take an entire sequence as input, the width of an RNN (number of nodes per layer) is usually set as the number of events (or timesteps) per sequence. Each node in an RNN takes two categories of information as input: information about the current state of the sequence (either from the input data or from the output of the node in the previous layer) and information about the previous states, represented by a hidden state, ht−1. The flow of information in an RNN can thus be thought of as traveling both across the sequence and upwards through the layers.
1.6. Study objectives
In this study, we present a novel method of predicting patient changes over the course of treatment which models the patient as a stationary quantum state whose timepoints are discretized by fraction number. The motivation for modeling the dynamics of patient anatomy as a quantum system is informed by the nature of quantum systems. Specifically, it is impossible for human observers to predict the behavior of a quantum system with absolute certainty. Similarly, in a radiation oncology paradigm, the aforementioned limitations in CBCT image quality in combination with daily setup variations, physiological changes such as tumor response, and anatomical motion, means that there are inherent uncertainties in knowing a patient’s “state” during the course of treatment. The stochastic nature of these two systems lends to the notion that modeling patient changes after the evolution of a quantum system may lead to more accurate predictions than those which could be achieved using a classical predictive framework.
To test this hypothesis, we designed a predictive framework which models patient changes as a stationary quantum system defined by the time-independent Schrodinger equation evolving according to a patient-specific Hamiltonian. The parameters of the Hamiltonian are derived from an RNN, which takes data from the first several fractions of a patient as input in order to predict later fractions. In addition, we also developed a Markov-based predictive framework which models patient changes according to the Kolmogorov backward equation and a patient-specific generator matrix. We trained and evaluated both frameworks on a head and neck cancer (HNC) radiotherapy dataset, which was selected because HNCs represent a patient population which commonly experiences geometric changes leading to plan adaptations (7).
2. Methods
2.1. Patient Population and Data Acquisition
Training data for this study were acquired retrospectively from dataset of 128 head and neck cancer (HNC) patients who received volumetric arc therapy (VMAT) at one institution between the years of 2013 and 2016. Of the original 128, three patients were excluded due to abnormalities which occurred during their treatment sequence. The remaining 125 patients consisted of 102(23) males(females) with an average age of 58.3 (range = [10, 84]). Patients received a target prescription dose between 60–70 Gy delivered over 30–35 fractions, with 17 patients undergoing replanning. Oropharyngeal cancers which are positive for human papilloma virus (HPV) have been shown to have higher radiosensitivity (26, 27). Of the patients whose tumor site was located in the oropharynx (n=68), 52 were recorded as positive for HPV, 5 were negative, and 1 was unknown. Patient immobilization of the upper body was achieved using 5-point masks and daily setup errors were corrected using CBCT image guidance.
An additional 20 HNC patients treated with 30–35 fractions of VMAT radiotherapy between 2018–2020 were selected to serve as an external testing dataset. The patients consisted of 14(6) males(females) with an average age of 60 (range = [22, 86]). Deformable image registration was performed on daily CBCT images for two fractions per treatment week and volume/table shifts were calculated from these fractions in the same manner as the training dataset.
2.2. Data preprocessing
The data used to define the patient state at each treatment fraction consisted of fractional primary clinical target volumes (CTVs) (with respect to the volume at the first fraction) as well as daily table shifts (lateral, longitudinal, vertical) used for patient alignment (Figure 1). The primary CTV volume changes were selected for the system state because they can act as an indicator for plan adaptation, while daily table shifts were selected because they represent the measurements performed at every fraction, as part of standard clinical workflow, thus they can provide a ground truth to assess the predictive performance of the proposed models. Further support for the use of table shifts to complement the tumor volume data was informed by previous literature on the main sources of errors and margin changes in radiotherapy (28, 29). Patient rotation angles (defined as the yaw angle of the treatment couch) were initially recorded but were ultimately omitted from the model because most patients did not undergo rotational shifts during treatment. Pitch and roll angles were also omitted because all of the patients in the training dataset were treated on 4D couches (which perform translations along all three axes and rotations around the yaw axis).
Figure 1.

Schematic of data collected from each patient across treatment fractions. Primary CTV volumes were acquired from daily CBCTs and daily table shifts (described by the coordinate space shown in the treatment system) were collected for each available fraction and used to define an orthonormal state vector.
Deformable image registration (DIR) was performed between daily CBCT images and the planning CT as described previously (12, 13). In the training dataset, 60% of patients had DIR performed for every fraction, while another 12% had DIR performed for all but 1 fraction. The rest of the patients in this study (including the external testing set) had DIR performed on at least two fractions per treatment week. CTV volumes were obtained from the registered daily CBCT images and normalized with respect to the volume of the first fraction in their respective treatment course. The motivation for reporting fractional (or relative) volumes as opposed to raw volumes was in part due to the wide range of volume sizes across patients [2.5cc – 456cc] and the need to relate meaningful tumor volume changes between patients (i.e., depending on tumor size, a 2cc reduction in tumor volume could represent an 80% reduction in tumor volume in one patient and only a 0.5% reduction in another). Daily table shifts and rotations were calculated as the difference between table coordinates recorded for the daily CBCT image and the DICOM treatment record (RTRECORD).
Fractions within a treatment sequence were also omitted from the model if they did not meet the inclusion criteria described below. The recorded CTV volumes and table shifts occasionally contained anomalies which did not represent meaningful changes across the treatment fractions. For example, some CTV volumes were measured which appeared as large spikes in the overall fraction sequence. Upon inspection of the original CBCT structure sets in Aria (Varian), it was found that these spikes in volume could be attributed to an error in the deformable image registration. To resolve this issue, fractions were discarded whose volume differed from both the fraction immediately prior and immediately preceding by > 10 cc. Abnormal table shifts were defined as a shift which exceeded 15 mm in any direction. These instances were attributed to large setup deviations from the original planning CT. Fractions with shifts exceeding this cut-off were excluded when they occurred sparsely throughout a treatment regimen. In two patients, CBCT shifts were recorded which exceeded 40 mm for every fraction. Because these represented a consistent setup variation from the planning CT, the shifts in these fractions were not discarded, and were instead normalized by subtracting off the shift in the first fraction to more faithfully represent position modifications that were not due to setup error. This ultimately resulted in the exclusion of 99 fractions across 35 patients for volume spikes and 8 fractions across 7 patients for table shift abnormalities (out of 3657 total fractions). Data preprocessing was performed in the same manner for the external testing dataset, with a total of 7 fractions discarded (from a total of 678 fractions) for failing to meet the inclusion criteria defined above.
After cleaning, the dataset underwent data augmentation, in which missing fractions were filled in by their nearest neighbor. Even fractions were discarded to provide a balance between the sparsely and densely sampled fractions; the final dataset used for training and validation had a sequence length of 18 and spanned fractions 1–35.
2.3. Mapping to discrete orthonormal state vectors
In order to model patient anatomical changes as a quantum system, the patient characteristics (volume + table shift values) at each point in time (fraction #) were mapped to an orthonormal state vector, Ψi, where Ψi∈[1,N] form a basis in an N-dimensional Hilbert space, and N is the number of potential states the patient can be in at a given time. It is desirable to select a value for N such as to balance the need to ensure that the Hilbert space contains enough dimensions to meaningfully capture the complexity of the patient population but is not so large as to result in large numbers of sparsely used discrete states. The dataset was encoded using several different values of N to assess the predictive models’ performance under this trade off.
The primary CTV volumes and daily table shifts were converted to orthonormal state vectors by first mapping them to discrete state values using vector quantization (30). A generalization of Lloyd-Max scalar quantization, vector quantization is a data reduction technique which finds an optimal codebook of vectors such that for a given dataset, the total difference between each multivariate sample and its nearest codebook neighbor is minimal. Once the codebook has been calculated, each sample vector is mapped to its nearest codebook neighbor and the dimensionality of the data is effectively reduced to the number of vectors (or states) contained within the codebook. Vector quantization was performed via the Vector Quantization Design Tool (vqdtool) available through MATLAB’s DSP system toolbox (Mathworks, Inc.), after which each state was encoded as a one-hot vector. Figure 2 displays an example of the distribution of lateral shifts recorded across all fractions from the original dataset (Fig. 2a) and after encoding to 4 (Fig. 2b) and 10 states (Fig. 2c), as well as a subset of the lateral shifts recorded from the original dataset (orange) and after encoding (blue) (Fig. 2d & 2e). The 10-state encoding, while computationally more expensive, provides a closer approximation to the original data. The external testing dataset was quantized using the same codebook as the training data. Note that the 10-state encoding only includes 6 discrete states for the lateral shifts (Fig. 2c). This is because the vector quantizer looks for the optimal representation of the distribution of vectors within a multidimensional dataset, and a single parameter may thus appear multiple times within a codebook of unique vectors.
Figure 2.

Distribution of lateral shifts recorded across all fractions from the original dataset (a) and after encoding to 4 (b) and 10 states (c), as well as a subset of the lateral shifts recorded from the raw, unquantized dataset (orange) and after encoding (blue) (d & e). Note that in Figure 2d, 2 of the 4 states, while distinct, are difficult to distinguish due to their close grouping.
2.4. Generation of additional synthetic data via Generative Adversarial Network (GAN)
One challenge for this study was the limited number of patients available to train the proposed models. To address this challenge, a Time-series generative adversarial network (TimeGAN) was employed to generate additional synthetic data (31). Generative adversarial networks describe a class of deep learning algorithms which learn the latent distributions of a dataset by pitting two competing networks against each other. The generator network seeks to generate realistic “fake” data samples to fool a discriminator network, which seeks to distinguish real samples from fakes.
Training is successful when the generator has reached the point where it can produce fake samples which a well-trained discriminator identifies correctly only 50% of the time. TimeGAN learns the distributions of features across time by utilizing RNNs for both the generator and the discriminator, as well as introducing a supervised loss term in addition to the traditional unsupervised adversarial loss. Finally, TimeGAN also utilizes an additional embedding network which learns a reversible mapping between latent space and features. Our implementation of TimeGAN followed the same architecture and hyperparameters as the original publication (31). Synthetic data was generated in the form of discrete state values produced from vector quantization (i.e., after discrete state mapping but prior to one-hot encoding). In order to prevent model leakage, TimeGAN was trained on five separate training folds used to train outer loop of the predictive models during nested cross validation. This resulted in an additional 200 samples per fold, for a total of 325 patients for model evaluation.
In order to assess the quality of the synthetic data, histograms counting all instances of each discrete state across all fractions and patients were calculated, and t-SNE as well as PCA analyses were performed (flattened in the time dimension) to visually compare the original and synthetic data. In addition, two quantitative performance metrics, a discriminative score and a predictive score, were obtained by training two separate neural networks. The discriminative score was calculated by training a discriminator (in the form of an RNN) to classify real vs. synthetic samples, and then reporting the classification error from a holdout testing dataset, defined as the absolute value of 0.5 minus the classification accuracy. Thus, the discriminative score is a number bounded between 0 and 0.5, where the closer to 0, the harder it is for an independent classifier to distinguish between real and synthetic data. The predictive score was calculated by training an RNN network using the synthetic dataset to predict the state in the next time step, and then reporting the mean absolute error of the model when tested on the original data. Thus, for the predictive score, a lower value indicates greater ability of an independent predictor to learn time-dependant features of the original dataset from the synthetic dataset. Figure 3 displays the results of the data visualizations (Figs. 3a, 3b, 3d, & 3e) as well as the testing scores for the discriminator and predictive networks (Fig. 3c), which are compared against results reported in the original TimeGAN study in which the model was trained on another sequence-based, medical dataset consisting of a private lung-cancer pathways dataset (31).
Figure 3.

Assessment of synthetic data generated by TimeGAN for the 10-state vector encoding trained using original data from fold 5: (a) shows the distribution of the original 10-state encoded data, (b) shows the distribution of the synthetic, (c) displays discriminative and predictive scores for the HN synthetic data (H&N) and the scores from the lung cancer pathways data (Lung*) reported in the original TimeGAN paper, (d) and (e) display t-SNE and PCA plots respectively for the original 10-state data and the synthetic 10-state data.
2.5. Quantum-based prediction algorithm
After mapping patient attributes to state vectors, the remaining key piece of information necessary to predict future patient states is to acquire an expression for the Hamiltonian, which determines the time evolution of the quantum system according to Eqns. (3) and (4). The Hamiltonian was derived from the output of a gated recurrent unit (GRU) neural network (32), which we refer to as the QRNN. In order to preserve the unity length of predicted states as well as ensure that the squared probability amplitudes sum to 1, the predicted Hamiltonian must be Hermitian (i.e., equal to its conjugate transpose). However, because PyTorch is currently limited in its use of complex numbers within ANNs, we assumed a real-valued Hamiltonian, and therefore needed only to build a symmetric matrix from the GRU output. State vectors from the first several fractions are used as input to the GRU, and the network output parameters necessary to construct the Hamiltonian in the form of vector of length (N + 1) × N/2, where N is the total possible number of states. This parameter vector was then used to construct a symmetric (and by default Hermitian) matrix according to Section 1.2.7 in Golub and Van Loan, 2013 (33). In order to calculate the unitary transition matrix defined by Eqn. 4, the expression for the transition matrix was reformulated using Euler’s formula and spectral decomposition was used to perform the necessary operations on the Hamiltonian matrix:
| (13) |
| (14) |
Where v and λ are the respective eigenvectors and eigenvalues of H. The transition matrix was split into its real and complex components such that:
| (15) |
Where t is expressed as an integer and determines which future fraction to be predicted. (Note that we can parameterize t, as t = 2x such that |Ψ(t0 + t)〉 = |Ψ(t0 + 2x)〉, to yield odd numbered fractions). Future fraction predictions, were acquired by calculating the probability amplitude of each state after applying the unitary transition matrix to the known system state of fraction t0:
| (16) |
| (17) |
| (18) |
Where is the ith element in the vector given by:
| (19) |
| (20) |
Where “. 2” refers to element-wise multiplication. The predicted future state was then defined as the basis vector with the largest probability. The model loss function was then defined as the negative log likelihood of the algorithm predicting the correct future state:
| (21) |
A schematic of the flow of information through the neural network is shown in Figure 4.
Figure 4.

Schematic of QRNN Algorithm. A set number of initial states are fed into a recurrent neural network, which produces a vector of parameters necessary to build a Hamiltonian matrix. The Hamiltonian is then used to construct a transition matrix, which in turn is used in conjunction with the oldest initial state to predict the probability of each remaining fraction existing in a given state. The fraction state predictions are then used in conjunction with the ground truth to calculate the model loss.
2.6. Markov-based prediction algorithm (MRNN)
Similar to the quantum model, the Markov model also utilized a GRU neural network (which we refer to as the MRNNR) and took the first 13 odd fractions as its input. The output of the MRNN was a vector of length l = N × (N − 1), where again N is the total number of possible states. A generator matrix, Q, was then built from this parameter vector such that Eqns. (10) and (11) were satisfied. The probability transition matrix defined by Eqn. (9) was then calculated using reference (34). The probability distribution of future states was thus calculated as:
| (22) |
And the predicted future state was assigned as the state with the maximum probability. The loss function for the MRNN was then defined as the negative log likelihood between the Markov-predicted state and the known future state:
| (23) |
2.7. Model Architecture and Training
The models were built and trained with Python using the PyTorch library on a Windows 10 desktop computer (AMD Ryzen 5 3600 6-Core Processor, 3.59 GHz, NVIDIA GeForce GTX 1660 SUPER). The neural networks for both models utilized Adam optimization with a learning rate of 0.001 and a batch size of 30. The QRNN and MRNN models were trained and assessed using a nested cross validation (NCV) scheme to mitigate overfitting effects and assess generalizability (35). NCV involves performing an inner loop of cross validation (CV) on each training fold for parameter tuning in order to avoid introducing additional bias into the model assessment. The inner loop consisted of 5-fold CV and was trained for 50 epochs in order to train the number of layers in the GRU as well as the dropout values. Optimal parameters were based on the minimum loss model loss achieved over ever combination of parameter values. The outer loop consisted of 5-fold CV and was trained for 100 epochs, which was determined experimentally. Model loss, area under the Receiver Operating Curve (AUC), and confusion matrices were recorded to assess model performance.
3. Results
For each system state representation N ∈ {4, 6, 8, 10}, models were trained using the first 13 odd fractions as input to predict the remaining odd fractions (15–35). To assess model performance, AUC scores were calculated for each model using the one-vs-one algorithm (36), which calculates the average AUC of all possible pairwise combinations of classes, with each class weighted by its support size. AUC scores were computed both across all fractions (15–35) and on a per fraction basis to assess both general model performance as well as model performance at specific timepoints.
3.1. Overall Results
Table 1 displays the validation AUC scores calculated using all future fractions (15–35) averaged over the 5-folds of cross validation (outer fold of NCV), as well as z-scores and corresponding p-values calculated from the DeLong test for each system state representation (37). From the results in Table 1, the MRNN model was observed to have a generally higher performance on the validation set than the QRNN model; however, these differences were only statistically significant (p < 0.05) for the 8-state and 10-state representations. In addition, both models seemed to experience reduced performance with increasing state-vector size (with the exception of the 8-state MRNN results).
Table 1.
Validation AUCs averaged over 5-folds of cross validation recorded for MRNN and QRNN models after 100 epochs. (mean ± 95% CI), as well as z-scores and corresponding p-values from the DeLong test.
| Validation AUC after 100 Epochs | ||||
|---|---|---|---|---|
| State-vector size | MRNN | QRNN | Z-score | p-value |
| 4 | 0.742 ± 0.021 | 0.675 ± 0.036 | 1.35 | 0.177 |
| 6 | 0.709 ± 0.026 | 0.656 ± 0.021 | 1.83 | 0.068 |
| 8 | 0.724 + 0.036 | 0.652 ± 0.044 | 2.53 | 0.011 |
| 17 | 0.698 ± 0.016 | 0.605 ± 0.035 | 2.67 | 0.007 |
Figure 5 displays the AUC scores on the internal validation dataset of the outer loop of NCV validation after 100 epochs of training. The results in Figure 5 suggest that for both the MRNN and QRNN models the AUC score appears to decrease for increasing system state sizes. In addition, there also appears to be a slight decrease in AUC score for fractions at later timepoints.
Figure 5.

MRNN and QRNN validation AUC scores vs fraction number for each system state representation. Shaded region indicates the 95% confidence interval.
One challenge of calculating AUC scores on a per-fraction basis is that this leads to fewer samples used to determine each score. In order to calculate the AUC, it is necessary for the sample set to contain instances (true labels) of every possible class. Because of this requirement and the fact that it is not possible to stratify patients across folds such that each fraction has an equal distribution of state values, it is possible for some folds to have undefined AUC values at specific fractions. For such instances, the mean AUC values and confidence intervals were reported based on the remaining folds. This resulted in models sometimes artificially appearing to have smaller measurement uncertainties – for example, the MRNN and QRNN models trained using system state vectors of size 10 had 4 out of 5 folds which contained undefined AUC scores at fraction 25. Table 2 summarizes the number of folds left out for each fraction during the per-fraction AUC calculations. The higher order system state representations (8 and 10) were found to suffer from a higher number of undefined folds.
Table 2.
Number of folds omitted from each fraction during AUC score calculation due to lack of samples for each possible class.
| Number of folds omitted from AUC calculation | ||||
|---|---|---|---|---|
| Fraction # | 4-state | 6-state | 8-state | 10-state |
| 15 | 0 | 1 | 2 | 2 |
| 17 | 0 | 1 | 1 | 1 |
| 19 | 0 | 1 | 2 | 3 |
| 21 | 0 | 1 | 3 | 3 |
| 23 | 0 | 1 | 2 | 3 |
| 25 | 0 | 0 | 1 | 4 |
| 27 | 0 | 0 | 2 | 1 |
| 29 | 0 | 0 | 1 | 2 |
| 31 | 0 | 0 | 1 | 1 |
| 33 | 0 | 1 | 0 | 3 |
| 35 | 0 | 0 | 3 | 2 |
3.2. 4-state predictions
Figure 6 displays the results for the 4-state quantum and Markov models, including the mean training and testing AUCs recorded over 100 epochs during the outer loop of 5-fold cross validation (Fig. 6a & 6b), the corresponding model loss (Fig. 6c & 6d), as well as confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) across each of the 4-states (Fig. 6e & 6f). The training, validation AUCs and model loss displayed in Fig. 6a–6d suggest that the Markov model exhibited higher performance than the Quantum model throughout training, while both models continued to learn over the entire course of 100 epochs. Both models exhibited high specificity and NPVs for all 4 system states. System states with less support (i.e. fewer instances of that given class) exhibited lower sensitivity and PPVs.
Figure 6.

QRNN and MRNN model results for 4-state system predicting fractions 15–35: AUC scores (mean and 95% CI) (a) & (b), model loss (c) & (d), and confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) averaged across 5-folds after 100 epochs of training (e) & (f).
3.3. 6-state predictions
Figure 7 displays the results for the 6-state quantum and Markov models, including the mean training and testing AUCs recorded over 100 epochs during the outer loop of 5-fold cross validation (Fig. 7a & 7b), the corresponding model loss (Fig. 7c & 7d), as well as confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) across each of the 6-states (Fig. 7e & 7f). The training and validation AUCs and model loss displayed in Fig. 7a–7d suggest that the MRNN model exhibited slightly better performance over the course of training. In addition, both models achieved maximum performance after 15–20 epochs. Both models exhibited high specificity and NPVs across all system states, and for both models the highest sensitivity and PPV scores occurred in the state with the largest support (number of instances of that class).
Figure 7.

QRNN and MRNN model results for 6-state system predicting fractions 15–35: AUC scores (mean and 95% CI) (a) & (b), model loss (c) & (d), and confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) averaged across 5-folds after 100 epochs of training (e) & (f).
3.4. 8-state predictions
Figure 8 displays the results for the 8-state quantum and Markov models, including the mean training and testing AUCs recorded over 100 epochs during the outer loop of 5-fold cross validation (Fig. 8a & 8b), the corresponding model loss (Fig. 8c & 8d), as well as confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) across each of the 8-states (Fig. 8e & 8f). The training and validation AUCs and model loss displayed in Fig. 8a–8d suggest that the MRNN model exhibited higher performance over the course of training. In addition, both models achieved maximum performance after 15–20 epochs. Both models exhibited high specificity and NPVs across all system states, while the sensitivity and PPV is reduced from that of the 4 and 6-state models. Note that the mean PPV is not defined for state 7 because one at least one of the five folds did not predict any patient to be in this state (i.e. no true positives or false positives).
Figure 8.

QRNN and MRNN model results for 8-state system predicting fractions 15–35: AUC scores (mean and 95% CI) (a) & (b), model loss (c) & (d), and confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) averaged across 5-folds after 100 epochs of training (e) & (f).
3.5. 10-state predictions
Figure 9 displays the results for the 10-state quantum and Markov models, including the mean training and testing AUCs recorded over 100 epochs during the outer loop of 5-fold cross validation (Fig. 9a & 9b), the corresponding model loss (Fig. 9c & 9d), as well as confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) across each of the 8-states (Fig. 9e & 9f). The training and validation AUCs and model loss displayed in Fig. 9a–10d suggest that the MRNN model exhibited higher performance over the course of training. In addition, both models achieved maximum performance after 15–20 epochs. Both models exhibited high specificity and NPVs across all system states. As the support for each state is spread out over more total states (8–10 states vs 4–6 states) the model sensitivity and PPV for each individual state is reduced. Note that for the MRN model, the mean PPV is not defined for states 3 and 6 because one at least one of the five folds did not predict any patients for these states (i.e., no true positives or false positives).
Figure 9,

QRNN and MRNN model results for 10-state system predicting fractions 15–35: AUC scores (mean and 95% CI) (a) & (b), model loss (c) & (d), and confusion matrix metrics: sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) fractions averaged across 5-folds after 100 epochs of training (e) & (f).
3.6. External Testing Results
Final models for each system-state representation were trained using all of the real training data (N = 125) as well as 250 synthetic samples (50 samples randomly selected from each of the 5 outer cross validation folds) so that the ratio of real to synthetic training data was the same as used during the cross validation process (1:2). For comparison, an additional set of final models were also trained using only real data. The hyperparameters for each model were selected as those chosen most frequently during nested cross validation. All models had an RNN depth of 2 layers and dropout coefficients, which ranged from 0.1–0.5. Model testing was performed using the external dataset of 20 patients described in Section 2.5. Table 3 displays the final testing AUC scores for each model, which are consistent with the training cross-validation results, where the MRNN approach was found to achieve higher AUC scores than the QRNN, though model performance was relatively less sensitive to the order of the system state representation.
Table 3.
Testing AUC scores calculated as the highest achieved over the course of 100 epochs of training.
| Testing AUC | ||||
|---|---|---|---|---|
| State-vector size | MRNN | QRNN | MRNN (noGAN) | QRNN (noGAN) |
| 4 | 0.729 | 0.650 | 0.707 | 0.623 |
| 6 | 0.677 | 0.651 | 0.687 | 0.608 |
| 8 | 0.716 | 0.683 | 0.723 | 0.669 |
| 10 | 0.671 | 0.607 | 0.697 | 0.609 |
3.7. Auxiliary study predicting 1 week ahead
The primary scope of this study was to assess the full potential of each model to predict changes in the patient state over the entire course of treatment. Our motivation for this goal was both to compare the models’ performance on a challenging problem as well as to investigate whether patient information from the first 13 fractions had any predictive power for patient states by the end of treatment. For many clinics however, the ability to predict patient changes only a week in advance may be sufficient for plan adaption. For this reason, we performed an additional study in which the MRNN and QRNN models were trained to predict fractions 19–25 based on input from the first 13 fractions. This corresponded to providing the model with data up through the middle of the third week of treatment and asking it to predict the patient states and the end of week 4 and through week 5.
The validation and external testing results of this auxiliary study are displayed in Tables 4 and 5 respectively. For the validation results, performance for both MRNN and QRNN models were improved for all but one case. As with the previous studies, the MRNN model exhibited stronger performance than the QRNN model for both validation and testing. However, the difference between the two models was significant only for the highest order system state.
Table 4.
Validation AUCs averaged over 5-folds of cross validation recorded for MRNN and QRNN models after 100 epochs. (mean ± 95% CI), as well as z-scores and corresponding p-values from the DeLong test.
| Validation AUC after 100 Epochs | ||||
|---|---|---|---|---|
| State-vector size | MRNN | QRNN | Z-score | p-value |
| 4 | 0.745 ± 0.031 | 0.706 ± 0.031 | 1.02 | 0.309 |
| 6 | 0.729 ± 0.022 | 0.673 ± 0.026 | 1.69 | 0.091 |
| 8 | 0.733 ± 0.021 | 0.671 ±0.043 | 1.21 | 0.230 |
| 10 | 0.701 ± 0.009 | 0.615 ± 0.024 | 3.13 | 0.002 |
Table 5.
Testing AUC scores calculated as the highest achieved over the course of 100 epochs of training.
| Testing AUC | ||||
|---|---|---|---|---|
| State-vector size | MRNN | QRNN | MRNN (noGAN) | QRNN (noGAN) |
| 4 | 0.711 | 0.666 | 0.701 | 0.620 |
| 6 | 0.709 | 0.653 | 0.706 | 0.621 |
| 8 | 0.742 | 0.726 | 0.768 | 0.701 |
| 10 | 0.694 | 0.637 | 0.694 | 0.591 |
For the external testing results, model performance was improved when the synthetic data was included in the training data for all models except the 6-state and 10-state MRNN models. In addition, the performance for both MRNN and QRNN models was consistently improved over that of models predicting every future fraction only when the synthetic data was included.
4. Discussion
In this work we presented two approaches for radiotherapy geometric adaptation based on quantum computing and Markov methods within a deep learning framework. Overall, the Markov predictive model was found to exhibit slightly higher performance (though not always statistically significant) than the quantum predictive model—both across all fractions (Table 1 and Table 2) and on a per-fraction basis (Fig. 5). One possible explanation for this outcome is the difference in the number of parameters necessary to construct the quantum-based Hamiltonian and the Markov generator matrix. The Markov generator matrix is required to have positive off-diagonal values and rows which sum to 0—requiring N × N − 1) parameters for N states. The Hamilton matrix is required to be Hermitian (in this case symmetric due to the additional assumption that it contained only real values)—requiring N/2 × (N + 1) parameters for N states. This resulted in the Markov model having two additional learned parameters for the 4-state representation, and 35 additional parameters for the 10-state. In future studies, (once complex numbers are supported by the autograd function in Pytorch) we plan to reformulate the QRNN model such that the Hamiltonian can more realistically contain any value in the complex plane. This would represent an increase in the degrees of freedom for the QRNN model and could help to better improve its performance compared to the MRNN model.
For external testing, final MRNN and QRNN models were trained in one of two ways – either with both real and synthetic data, or with only real data. For the MRNN testing results, which predicted states all the way to end of treatment, it was found that the models trained using only real data outperformed those with synthetic data for three of the four system state vectors, while the opposite trend was observed in the QRNN model results. However, if we estimate that the uncertainty associated with the testing results is at least as large as the 95% confidence intervals reported for the cross validation then these trends do not appear to be significant. For the testing results which only predicted 1–2 weeks into the future, it was found that the addition of the synthetic data during training improved the QRNN model performance above the range of estimated model uncertainty, suggesting that there are conditions in which the use of synthetic data to aid training provides value.
A key challenge for the MRNN and QRNN predictive models is to determine an adequate size of the system state vector. Higher-dimensional vectors are preferable from a clinical standpoint because they can more faithfully represent the original data, as demonstrated by Fig. 2d and 2e. In addition, higher dimensional system state vectors can more easily allow for the incorporation of additional types of patient data—such as biomarkers, radiomics, or dosimetric information—allowing for a more complete and individualized representation of patients. However, the results of this study suggest that there is a tradeoff between model performance and state vector complexity. This observation is not surprising; because the problem is formulated as a multi-class classification problem, an increase in system-state vector size increases the number of classes which the model must learn to distinguish without increasing the number of samples—effectively forcing each class to be more sparsely represented. This challenge makes it particularly difficult to draw confident conclusions about both models’ predictive abilities across different fractions (Fig. 5). Based on the 4-state and 6-state results it appears there may be a slight reduction in model performance as the fraction number to be predicted grows further from the last initial state. This observation is further supported by the improved results of the auxiliary study, in which the models were trained to predict only fractions 1–2 weeks into the future. Finally, it is worth noting that of all the system state vector lengths investigated in this study, only the 4-state models continued to learn over the entire course of 100 epochs (Figure 6). The other three higher order models all reached their maximum performance within the first 20–50 epochs and began overfitting from that point afterwards (Figures 7–9).
The current workflow for implementing plan adaptations on HN patients involves identification of anatomical changes through visual inspection of daily CBCT images. When dramatic changes are identified by the clinical team, a new CT simulation can be performed to better quantify those changes and to assist in adaptive replanning. This current ART workflow is hindered by the time constraints inherent in a clinical environment. Models which can predict patient changes during treatment could help to alleviate this challenge by allowing clinical teams to anticipate and prepare for adaptions before they are necessary. Several previous studies have reported on the use of predictive models for HN cancer. Rosen et al. reported a model which predicted radiation-induced xerostomia using dose volume histogram metrics along with CBCT image information from the last two weeks of treatment (13). In another study, McCulloch et al. reported a model which predicted the need for replan based on estimated dose deviations calculated using daily CBCT images from the 5th, 10th, and 15th fractions (12). To our knowledge, the models presented in this study are the first which predict anatomical changes in head and neck cancer beyond the next treatment fraction using early treatment data (the first 13 fractions). Recent studies have reported different approaches for predicting tumor shrinkage in lung cancer. Zhang et al. reported on the validation of a principal-component-based predictive atlas using CBCTs of 34 locally advanced lung cancer patients from two institutions who experienced at least 20% tumor shrinkage over the course of radiotherapy (19), while Wang et al. reported the use of a deep learning algorithm to predict the spatial distribution of lung tumors using weekly T2-weighted MRI scans on 10 lung cancer patients (20). This approach allows for predictions to be made far enough in advance for clinical intervention and could also be used to improve future prediction studies. For example, geometric predictions could be coupled with daily dose calculations to provide a more accurate estimation of dosimetric deviations. The models presented in this work further distinguish themselves from previous studies in that they can be formulated to encode any aspect of the patient state (geometrical, physiological, dosimetric, or higher order latent features).
The most accurate data driven models for outcome prediction are deep learning models because they can recognize complex, nonlinear relationships in data. However, machine learning is sensitive to noise in the data, and in our case because there are inherent uncertainties in knowing the state of the patient during any given fraction, we can only offer noisy data to the algorithm for training. The motivation for utilizing Markov/quantum models is that both of rely on probabilistic transition matrices to predict changes over time and are formulated deal with systems that have a stochastic component/inherent uncertainties. The results of this study suggest that future work will be necessary to define and achieve an appropriate balance between predictive performance and providing a clinically meaningful and actionable description of the patient state. In their current formulation, the models for all 4 state-vector sizes achieve very high specificity but suffer from low sensitivity. One could envision a paradigm in which a subset of the patient states is designated as highly actionable (due to high volume change, large daily shifts, or perhaps based on expected clinical outcomes), with the model used as a warning system to alert the clinical team when a patient is likely to transition to the actionable state later during their treatment course. Given the resource constrains associated with plan adaption, such a scenario might add value to the clinical workflow because although the model cannot be trusted to predict every actionable state, the situations in which it does flag a patient as necessitating replan in the future can be relied on to justify replan.
The limitations associated with this study are summarized as follows: the patient variations predicted in this study were limited to geometric anatomical changes over the course of treatment. While there is limited consensus on which exact patient changes are relevant and/or necessary to track for the purpose of ART, dosimetric information (in particular dose deviations from the original treatment plan) represents another category of data commonly used for informing ART decisions. In addition, because many HNC patients experience xerostomia induced by radiation toxicity, it would be valuable to also include data related to the parotid glands in future studies.
Due to a number of factors, including the CTV structure boundaries ending outside the field of view of the daily CBCTs as well as occasional failure of the DIR algorithm, it was not possible to include every treatment fraction for each patient. While generally at least 2 fractions per week were included, the RNN models might be more successful at learning relevant time-dependent patterns in the patient data if all fractions could be included. Finally, this study did not consider the impact of “measurements” on the patient system. In a quantum framework, any measurement on the system changes the state of the system itself (often referred to as “collapsing the system’s wave function”). While the values predicted in this study were treated as artificial measurements (primarily because the quantitative value of the day-to-day volume of the primary CTV and treatment couch position, while potentially valuable, are not currently explicitly used in the clinical workflow for ART decision-making), it would be interesting in future studies to consider how these models can predict the impact of clinical decisions (such as the choice to modify a plan) over the course of treatment. Such a study may reveal additional benefits that are more unique to the quantum-based QRNN model.
5. Conclusion
In this study we report on the design and assessment of two novel deep machine learning frameworks for predicting anatomical changes in radiotherapy patients. A comparison between the performance of the quantum-inspired and Markov-inspired models indicated that the Markov models tend to exhibit higher performance. In addition, it was found that there was a tradeoff between model performance and the size of the system state vector. A major challenge of this study was that as the system state dimension increased, so did the sparsity of class representation, leading to a decrease in model performance. Further studies utilizing both greater numbers of patients with diverse treatment outcomes, information regarding clinical decisions over the course of treatment, as well as additional datatypes beyond basic imaging features such as tumor volume are necessary to fully evaluate the predictive power of these models.
6. Acknowledgement
This research was supported financially in part by the University of Michigan Regents’ Fellowship and by National Institutes of Health (NIH) grant R01-CA233487.
7. References
- 1.Baskar R, Lee KA, Yeo R, Yeoh K-W. Cancer and radiation therapy: current advances and future directions. Int J Med Sci. 2012;9(3):193–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Bortfeld T. IMRT: a review and preview. Physics in Medicine and Biology. 2006;51(13):R363–R79. [DOI] [PubMed] [Google Scholar]
- 3.Teoh M, Clark CH, Wood K, Whitaker S, Nisbet A. Volumetric modulated arc therapy: a review of current literature and clinical use in practice. Br J Radiol. 2011;84(1007):967–96. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Yan D, Vicini F, Wong J, Martinez A. Adaptive radiation therapy. Physics in Medicine and Biology. 1997;42(1):123–32. [DOI] [PubMed] [Google Scholar]
- 5.Schwartz DL, Garden AS, Thomas J, Chen Y, Zhang Y, Lewin J, et al. Adaptive Radiotherapy for Head-and-Neck Cancer: Initial Clinical Outcomes From a Prospective Trial. International Journal of Radiation Oncology*Biology*Physics. 2012;83(3):986–93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Wu QJ, Li T, Wu Q, Yin F-F. Adaptive Radiation Therapy: Technical Components and Clinical Applications. The Cancer Journal. 2011;17(3). [DOI] [PubMed] [Google Scholar]
- 7.Lim-Reinders S, Keller BM, Al-Ward S, Sahgal A, Kim A. Online adaptive radiation therapy. International Journal of Radiation Oncology* Biology* Physics. 2017;99(4):994–1003. [DOI] [PubMed] [Google Scholar]
- 8.Sonke J-J, Aznar M, Rasch C. Adaptive Radiotherapy for Anatomical Changes. Seminars in Radiation Oncology. 2019;29(3):245–57. [DOI] [PubMed] [Google Scholar]
- 9.Tseng H-H, Luo Y, Ten Haken RK, El Naqa I. The Role of Machine Learning in Knowledge-Based Response-Adapted Radiotherapy. Front Oncol. 2018;8:266–. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Zhen X, Chen J, Zhong Z, Hrycushko B, Zhou L, Jiang S, et al. Deep convolutional neural network with transfer learning for rectum toxicity prediction in cervical cancer radiotherapy: a feasibility study. Physics in Medicine & Biology. 2017;62(21):8246. [DOI] [PubMed] [Google Scholar]
- 11.Luo Y, El Naqa I, McShan DL, Ray D, Lohse I, Matuszak MM, et al. Unraveling biophysical interactions of radiation pneumonitis in non-small-cell lung cancer via Bayesian network analysis. Radiother Oncol. 2017;123(1):85–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.McCulloch MM, Lee C, Rosen BS, Kamp JD, Lockhart CM, Lee JY, et al. Predictive models to determine clinically relevant deviations in delivered dose for head and neck cancer. Practical radiation oncology. 2019;9(4):e422–e31. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Rosen BS, Hawkins PG, Polan DF, Balter JM, Brock KK, Kamp JD, et al. Early Changes in Serial CBCT-Measured Parotid Gland Biomarkers Predict Chronic Xerostomia After Head and Neck Radiation Therapy. Int J Radiat Oncol Biol Phys. 2018;102(4):1319–29. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Vallières M, Kay-Rivest E, Perrin LJ, Liem X, Furstoss C, Aerts HJWL, et al. Radiomics strategies for risk assessment of tumour failure in head-and-neck cancer. Scientific Reports. 2017;7(1):10117. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Heukelom J, Fuller CD. Head and Neck Cancer Adaptive Radiation Therapy (ART): Conceptual Considerations for the Informed Clinician. Seminars in Radiation Oncology. 2019;29(3):258–73. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Møller DS, Holt MI, Alber M, Tvilum M, Khalil AA, Knap MM, et al. Adaptive radiotherapy for advanced lung cancer ensures target coverage and decreases lung dose. Radiotherapy and Oncology. 2016;121(1):32–8. [DOI] [PubMed] [Google Scholar]
- 17.Bhide SA, Davies M, Burke K, McNair HA, Hansen V, Barbachano Y, et al. Weekly volume and dosimetric changes during chemoradiotherapy with intensity-modulated radiation therapy for head and neck cancer: a prospective observational study. International Journal of Radiation Oncology* Biology* Physics. 2010;76(5):1360–8. [DOI] [PubMed] [Google Scholar]
- 18.Kamrani A, Azimi M, Nasr EA, editors. Predictive modeling of tumors using RP2015: IEEE. [Google Scholar]
- 19.Zhang P, Yorke E, Mageras G, Rimner A, Sonke J-J, Deasy JO. Validating a Predictive Atlas of Tumor Shrinkage for Adaptive Radiotherapy of Locally Advanced Lung Cancer. Int J Radiat Oncol Biol Phys. 2018;102(4):978–86. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wang C, Rimner A, Hu Y-C, Tyagi N, Jiang J, Yorke E, et al. Toward predicting the evolution of lung tumors during radiotherapy observed on a longitudinal MR imaging study via a deep learning algorithm. Medical physics. 2019;46(10):4699–707. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Busemeyer JR, Wang Z, Pothos E. Quantum models of cognition and decision. Oxford handbook of computational and mathematical psychology. 2015:369–89. [Google Scholar]
- 22.Nielsen MA, Chuang IL. Quantum computation and quantum information.
- 23.Norris JR. Markov Chains: Cambridge University Press; 1998. [Google Scholar]
- 24.Cui S, Tseng H-H, Pakela J, Ten Haken RK, El Naqa I. Introduction to machine and deep learning for medical physicists. Medical Physics. 2020;47(5):e127–e47. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Lipton ZC, Berkowitz J, Elkan C. A critical review of recurrent neural networks for sequence learning. arXiv preprint arXiv:150600019. 2015. [Google Scholar]
- 26.Fakhry C, Westra WH, Li S, Cmelak A, Ridge JA, Pinto H, et al. Improved survival of patients with human papillomavirus-positive head and neck squamous cell carcinoma in a prospective clinical trial. J Natl Cancer Inst. 2008;100(4):261–9. [DOI] [PubMed] [Google Scholar]
- 27.Lindel K, Beer KT, Laissue J, Greiner RH, Aebersold DM. Human papillomavirus positive squamous cell carcinoma of the oropharynx: a radiosensitive subgroup of head and neck carcinoma. Cancer. 2001;92(4):805–13. [DOI] [PubMed] [Google Scholar]
- 28.van Herk M. Errors and margins in radiotherapy. Semin Radiat Oncol. 2004;14(1):52–64. [DOI] [PubMed] [Google Scholar]
- 29.Santanam L, Esthappan J, Mutic S, Klein EE, Goddu SM, Chaudhari S, et al. Estimation of setup uncertainty using planar and MVCT imaging for gynecologic malignancies. Int J Radiat Oncol Biol Phys. 2008;71(5):1511–7. [DOI] [PubMed] [Google Scholar]
- 30.Gersho A, Gray RM. Vector quantization and signal compression: Springer Science & Business Media; 2012. [Google Scholar]
- 31.Yoon J, Jarrett D, van der Schaar M, editors. Time-series generative adversarial networks2019.
- 32.Cho K, Van Merriënboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, et al. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:14061078. 2014. [Google Scholar]
- 33.Golub GH, Van Loan CF. Matrix Computations: Johns Hopkins University Press; 2013. [Google Scholar]
- 34.Skafte N Cuda_exmp 2018. [updated Nov, 2018. Nov, 2018:[Available from: https://github.com/SkafteNicki/cuda_expm.
- 35.Cawley GC, Talbot NLC. On over-fitting in model selection and subsequent selection bias in performance evaluation. The Journal of Machine Learning Research. 2010;11:2079–107. [Google Scholar]
- 36.Hand DJ, Till RJ. A Simple Generalisation of the Area Under the ROC Curve for Multiple Class Classification Problems. Machine Learning. 2001;45(2):171–86. [Google Scholar]
- 37.DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. Biometrics. 1988;44(3):837–45. [PubMed] [Google Scholar]
